Introduction to Vega-Lite

Vega-Lite is a declarative language for interactive data visualization. Vega-Lite offers a powerful and concise visualization grammar for quickly building a wide range of statistical graphics.

By declarative, we mean that you can provide a high-level specification of what you want the visualization to include, in terms of data, graphical marks, and encoding channels, rather than having to specify how to implement the visualization in terms of for-loops, low-level drawing commands, etc. The key idea is that you declare links between data fields and visual encoding channels, such as the x-axis, y-axis, color, etc. The rest of the plot details are handled automatically. Building on this declarative plotting idea, a surprising range of simple to sophisticated visualizations can be created using a concise grammar.

To learn more about the motivation and basic concepts behind Vega-Lite, watch the Vega-Lite presentation video from OpenVisConf 2017.

This notebook walks through the basic process of creating visualizations with Vega-Lite. At its core, Vega-Lite uses JavaScript Object Notation (JSON) as a stand-alone file format for describing charts. However, to create Vega-Lite visualizations in a more convenient and programmatic way, we will be using the Vega-Lite API, a set of JavaScript methods that produce Vega-Lite JSON specifications as output.

Are you new to JavaScript and/or Observable? This notebook assumes basic familiarity with both. For a crash course, see A Minimal Introduction to JavaScript and Observable!


Imports

To start, we will import the vega-lite-api library. We will use Vega-Lite version 5:

import { vl } from "@vega/vega-lite-api-v5";

... and a printTable utility for formatting data tables:

import { printTable } from "@uwdata/data-utilities";

Data

Data in Vega-Lite is assumed to be formatted as a data table (or data frame) consisting of a set of named data columns. We will also regularly refer to data columns as data fields. Once loaded, the default table representation is an array of JavaScript objects.

When using Vega-Lite, datasets may be pre-loaded as an array of objects, or provided via a URL to load a network-accessible dataset. As we will see, the named columns of the data frame are an essential piece of plotting with Vega-Lite.

In these notebooks, we will often use datasets from the vega-datasets repository:

data = Object {annual-precip.json: ƒ(), ...};
Name Miles_per_Gallon Cylinders Displacement Horsepower Weight_in_lbs Acceleration Year Origin
"chevrolet chevelle malibu" 18 8 307 130 3504 12 "1970-01-01" "USA"
"buick skylark 320" 15 8 350 165 3693 11.5 "1970-01-01" "USA"
"plymouth satellite" 18 8 318 150 3436 11 "1970-01-01" "USA"
"amc rebel sst" 16 8 304 150 3433 12 "1970-01-01" "USA"
"ford torino" 17 8 302 140 3449 10.5 "1970-01-01" "USA"

All datasets in the vega-datasets collection can also be accessed via URLs:

"https://cdn.jsdelivr.net/npm/vega-datasets@1.31.1/data/cars.json"

Open the URL above in a separate browser tab if you want to examine the raw data file!


Weather Data

Statistical visualization in Vega-Lite begins with "tidy" data frames. Here, we'll start by creating a simple data frame (df) containing the average precipitation (precip) for a given city and month — written in JSON format:

df = [
  {"city":"Seattle","month":"Apr","precip":2.68},
  {"city":"Seattle","month":"Aug","precip":0.87},
  {"city":"Seattle","month":"Dec","precip":5.31},
  {"city":"New York","month":"Apr","precip":3.94},
  {"city":"New York","month":"Aug","precip":4.13},
  {"city":"New York","month":"Dec","precip":3.58},
  {"city":"Chicago","month":"Apr","precip":3.62},
  {"city":"Chicago","month":"Aug","precip":3.98},
  {"city":"Chicago","month":"Dec","precip":2.56}
];
city month precip
"Seattle" "Apr" 2.68
"Seattle" "Aug" 0.87
"Seattle" "Dec" 5.31
"New York" "Apr" 3.94
"New York" "Aug" 4.13
"New York" "Dec" 3.58
"Chicago" "Apr" 3.62
"Chicago" "Aug" 3.98
"Chicago" "Dec" 2.56

Marks and Encodings

Given the weather data above, we can specify how we would like the data to be visualized with Vega-Lite. We first indicate what kind of graphical mark (geometric shape) we want to use to represent the data. We can create a new Vega-Lite mark instance using the vl.mark* methods.

We can create a point mark using vl.markPoint(), and then pass the data to the data() method. Finally, we invoke render() to draw the plot:

vl.markPoint()
  .data(df)
  .render();

To visually separate the points, we can map various encoding channels (or just channels for short) to fields in the dataset. For example, we can encode the data field city using the y channel, which represents the y-axis position of the points. To specify this, use the encode() method, passing it definitions for specific channels.


Data Transformation: Aggregation

To allow for more flexibility in how data are visualized, Vega-Lite has a built-in syntax for aggregation of data. For example, we can compute the average of all values by specifying an aggregation function along with the field name:


Changing the Mark Type

Let's say we want to represent our aggregated values using rectangular bars rather than circular points. We can do this by replacing vl.markPoint with vl.markBar:


Customizing a Visualization

By default Vega-Lite makes some choices about properties of the visualization, but these can be changed using methods to customize the look of the visualization. For example, we can modify scale properties using the scale property, set titles using the title property, and we can specify the color of the mark by passing an object to the mark* method with a color property containing a valid CSS color string:


Multiple Views

As we've seen above, a basic Vega-Lite visualization represents a plot with a single mark type. What about more complicated diagrams, involving multiple charts or layers? Using a set of view composition operators, Vega-Lite can take multiple chart definitions and combine them to create more complex views.


Interactivity

In addition to basic plotting and view composition, one of Vega-Lite's more exciting features is its support for interaction.


Aside: Examining JSON Output

The Vega-Lite API converts plot specifications to a JSON-compatible object format that conforms to the Vega-Lite schema. Using the toObject() method, we can inspect the specification that is sent to Vega-Lite:

{
  "mark": {
    "type": "circle"
  },
  "encoding": {
    "x": {
      "field": "precip",
      "type": "quantitative",
      "aggregate": "average"
    },
    "y": {
      "field": "city",
      "type": "nominal"
    }
  }
}

Coda: Exporting and Publishing a Visualization

Once you have visualized your data, perhaps you would like to publish it somewhere else on the web. To save an exported image, you can simply right-click a visualization and select "Save Image As..." from the context menu, assuming the default canvas rendering option is used.

To include a Vega-Lite visualization on your own web page, you can use the vega-embed JavaScript package. Take an exported Vega-Lite JSON string, and incorporate it in a web page that imports vega-embed. Here is a basic HTML template, where the JSON specification for your plot produced by chart.toObject() should be stored in the spec JavaScript variable:

<!DOCTYPE html>
<html>
<head>
  <script src="https://cdn.jsdelivr.net/npm/vega@5"></script>
  <script src="https://cdn.jsdelivr.net/npm/vega-lite@5"></script>
  <script src="https://cdn.jsdelivr.net/npm/vega-embed@6"></script>
</head>
<body>
  <div id="vis"></div>
  <script type="text/javascript">
    const spec = {};  /* JSON output for your chart's specification */
    const opt = {renderer: "canvas", actions: false};  /* Options for the embedding */
    vegaEmbed("#vis", spec, opt);
  </script>
</body>
</html>

Next Steps

🎉 Hooray, you've completed the introduction to Vega-Lite! In the next notebook, we will dive deeper into creating visualizations using Vega-Lite's model of data types, graphical marks, and visual encoding channels.