Dear friends, Alberto here.
This is the second in a series of articles emerging from the many notes I’ve taken in the past few years for my fifth book. Read the first post here.
Philosophizing
Back in 2020 I wrote the foreword for Data Visualization in Society, an open access book edited by Martin Engebretsen and Helen Kennedy. I titled it ‘The dawn of a philosophy of visualization’; it’s a call to expand the horizons of our field by borrowing and translating conceptual and practical tools from other disciplines, mainly philosophical ones. Here’s the key paragraph:
Writing about visualization doesn’t mean just thinking about how to design visualizations, but also about what visualization is, why it is the way it is—and what it could be. Data visualization is a technology—or set of technologies—and, like artifacts such as the clock, the compass, the abacus, or the map, it transforms the way we see and relate to reality. As Langdon Winner suggested in The Whale and the Reactor (1986), a foundational book in the phenomenological philosophy of technology, to create technologies doesn’t consist just of crafting stuff; rather, when technologies come about ‘new worlds are being made’. What ‘new worlds’ does visualization generate? That’s a question for a potential philosophy of visualization.
I’ve spent the last six years musing about my own words, and trying to act accordingly.
This semester I asked my friend and University of Miami colleague, Otávio Bueno, for permission to sit in one of his PhD seminars. Otávio specializes in the philosophy of mathematics, modeling, and scientific representation, which is precisely the topic he chose for the seminar that I inquired about. It’s been a life-changing experience.
Otávio’s scholarly stature is on par with other analytic philosophers of scientific representation I was already vaguely familiar with before joining his seminar, such as Bas van Fraassen, Nancy Cartwright, Mauricio Suárez, Tarja Knuuttila, or Steven French. Their work is fertile ground for anyone interested in questions such as “how does visualization represent?” or, more fundamentally, “what do we even mean when we say that a visualization represents?”
Representation
To get a glimpse of what the literature on scientific representation has to offer, take a look at this terrific overview by Roman Frigg and James Nguyen, authors of the also excellent book Modeling Nature. Let me give you just a quick taste of how I think we can bring some of their ideas to information design and data visualization (for a fully comprehensive treatment, you’ll have to wait until I finish writing the damn book!)
My favorite definition of representation—any type of representation, including artistic representation, and not just visual representation—comes from a 1993 paper by Bas van Fraassen and Jill Sigman:
Representation of an object involves producing another object which is intentionally related to the first by a certain coding convention which determines what counts as similar in the right way.
Here are some key components of this definition:
Properly speaking, the representation is not the object that does the representing—a chart, for example. Representation is a process; the object we create is the carrier or vehicle of the representation; the object represented is usually called the target of the representation.
Representation encompasses the act of representing, and also the act of connecting the vehicle back to the target, f.i. by making inferences about the latter based on the former (more about this below).
There’s no representation without intention. The author/viewer of the representing object bestows on it its role (“This chart represents this data set I have about X”).
The relation between target and vehicle depends on a coding convention. In data visualization, this can be a grammar of graphics.
Representation also depends on some sort of similarity between the target and the vehicle.
Similarity is a tricky word. In this context, it isn’t restricted to external resemblance, like in a realistic painting or sculpture; after all, our charts don’t resemble the data they encode! External resemblance is just a very specific type of “similarity”. Similarity can also be conventional or, in the case of charts, structural.
Take the scatter plot below. To design it, we have to build it in a certain way, positioning its marks on the X and Y axis in relation to the magnitudes they encode. There’s a one-to-one structural relation between the values of the two quantitative variables—GDP per capita and life satisfaction scores—and their X and Y positions. (In this case, by the way, we’re dealing with a specific and highly restrictive type of structural similarity, an isomorphism; however, as Michael Ende would say, “that is another story and shall be told another time.”)

Scatter plot of self-reported life satisfaction versus GDP per capita country by country, colored by continent. The graph shows a positive correlation between the two variables. Source: OurWorldInData
Visualization as epistemic representation
Not all visualizations are scientific representations, as not all visualizations represent scientific content, and visualization has applications in many fields besides the sciences.
However, at least since I wrote The Truthful Art (2016), where I first played with the idea that visualizations are, in essence, artifacts that make data and data models more concrete, I’ve been convinced that information design and data visualization can be conceptualized and studied as part of the much broader world of epistemic representations (all scientific representations are also epistemic representations, but the reverse isn’t true.) Here’s a Venn diagram about all this:

The large universe of representations contains many subsets. One of them is epistemic representations. This smaller set of epistemic representations contains even smaller subsets, among them scientific representations and visualizations. Some visualizations are also scientific representations, but many others aren’t.
Gabriele Contessa, whose writings have influenced me greatly, explains epistemic representation this way:
A vehicle is an epistemic representation of a certain target for a certain user if and only if the user takes the vehicle to denote the target and she adopts an interpretation of the vehicle (in terms of the target).
There’s a bit to unpack here. To Contessa, epistemic representation is inseparable from surrogative reasoning; what he means is that the condition for a representation to be an epistemic representation of a target, you have to be able to make inferences about this target based on inferences you first make about the vehicle.
(Representations in general, a majority of them, don’t meet this surrogative reasoning condition. For instance, a national flag represents a country but you can’t infer anything about the country just by looking at its flag.)
This may sound a bit obscure. Let’s explain epistemic representations with a diagram loosely inspired by Applying Mathematics, a 2018 book co-authored by Otávio himself:

Based on Bueno & French (2018); adapted and modified after a long conversation with the former. The earliest diagram of this kind appears in the landmark paper ‘Models and representation’ (1997) by R.I.G. Hughes.
The cycle depicted in the diagram involves three steps, immersion, derivation, and interpretation. Here’s an explanation using the Our World in Data scatter plot above:
Immersion consists of embedding certain features of the target into the vehicle. In the scatter plot (our vehicle), the quantities in the data set (the target) are “translated” (encoded) as positions (X and Y) on a coordinate plane, and the categorical variable “continent” is translated as color hue.
The second step is derivation. If you know how to read a scatter plot—this is the “for a certain user” in Contessa’s account of epistemic representation— you can draw inferences about the scatter plot itself. For example, you may notice that the two quantitative variables are positively correlated, or that African countries tend to be at the lower end of both life satisfaction and GDP per capital.
The third step is interpretation, which means that (again, if you know how to read a scatter plot), you can “translate” what you’ve learned from the vehicle (the visualization) back to the target (the data). This closes the representation / surrogative reasoning cycle.
To make the point clearer, imagine that instead of a scatter plot, you’re reading a road map. You first make inferences about the vehicle (“the scale on this map is 1:100,000, and the distance between this point and this other point is 1 centimeter”), and then you interpret those inferences back into the target, the place represented (“the two dots on the map are cities; the real distance between them is 100,000 centimeters, or 1 kilometer.”) We use the map to first reason about the map itself, and after about the geography we wish to navigate.
At this point, I shall introduce a complication. I just wrote that the scatter plot represents our data, but what we’re usually interested in is not data itself, which is just a mediator between us and the phenomenon that we truly want to learn about. How then, does a visualization relate to the phenomenon? Here’s the thing, measurement is also a form of epistemic representation! And so is data modeling. We can therefore envision a forward chain of immersion and a backward chain of interpretation. Like this:

This is a highly simplistic way to illustrate the fact that obtaining data by measuring a phenomenon is a kind of epistemic representation. We can build models from this data, and the models are also representations. Finally, we can use data visualization to depict either the “raw” data (if the data isn’t too messy or complex) or the model.
I hope your brain hasn’t exploded yet. I’m just scratching the surface of a much more intricate, nuanced, and interesting story that I’m hoping to tell at length soon...
—
Enough theorizing for today! I leave you with ‘Name in Blood’ by Black Label Society, the band of former Ozzy Osbourne guitarist Zakk Wylde. Besides playing the guitar, Wylde also sings, and he sounds a bit like the Prince of Darkness himself.


