What I mean by computationally entangled interpretive research
Surrealness nothwithstanding, I have to engage with my new chosen academic field of IS/Information Systems. And so I've proposed a TREO talk at ICIS 2026 in Lisbon called "Computationally-Entangled Interpretive Methods: Using Computation In, and In Service Of, Interpretation." ICIS is the flagship IS conference, and a TREO talk is short, five-ish minutes, an opinion piece more than a paper. That format forces compression, and compression tends to leave a position under-explained. This post is my attempt to say more plainly what I mean, mostly for myself.
The starting problem is size. Organizational and social life increasingly leaves traces in corpora, digital trace data, social media archives, news archives, that are too large to read closely. A researcher used to sitting with transcripts or field notes now faces millions of documents.
The compounding problem is the tendency in my field, information systems, and in management research more broadly, to let computation absorb that problem by doing the interpreting for us. A topic model produces clusters, an embedding produces a similarity space, and the output gets treated as an analytic finding in its own right. Machine summarization and classification stand in for interpretation. The human interpreter is kept "in the loop" at the edge of the process, or ejected entirely in the name of scale and efficiency. Speed is king, publish or perish, etc.
I want to center the interpreter instead. Computation can be performed in service of interpretation, not merely as a replacement for it. Specifically, I think there are two uses of computation that hold up well under this standard: The first is purposive sampling: using computation to help the interpreter find the texts worth reading closely, based on some understanding of the social process and structure underlying the texts, out of a corpus no one could read in full. The second is intentional representation: using computation to construct representations, a visual, a cluster, an embedding space, a network, that the interpreter can participate in creating and only then interpret, rather than accept as a finished result. Both center human reading and reasoning. I suspect there is room for more uses like these than the field currently exploits, and part of what I want to work out in conversation with interested colleagues is what else belongs in that category.
Underneath both is a point about method itself. Choices of data, data structuring, and algorithm are theoretical choices. Where you draw a corpus boundary, how you define an actor category, what unit you model at, how permissive you let a clustering algorithm be, all of these encode assumptions about what the phenomenon is, who counts as a participant in it, how discourse typically unfolds, how narrative are created. This is close to Karen Barad's idea of the agential cut: an empirical method doesn't just describe a phenomenon that was already there, it enacts what counts as the phenomenon in the first place. So I pay close attention to the cut, and to the role method plays in producing the phenomena we then claim to study.
I've written about this before. Two co-authors and I are analyzing years of Twitter and Factiva discourse around an emerging market category. When we tuned a topic model, we found that Twitter's community discourse forms a few large, persistent attractors, so a small number of maximally stable clusters served it well. Factiva's journalist-authored, beat-structured text has no equivalent dominant narrative, and the same default badly underrepresented it. We instead pushed toward the finest-grained partition available, because we already knew, from understanding how wire journalism gets produced, that it should recover many discrete regulatory and institutional beats. That choice came from what we knew qualitatively about the communities that produced the two corpora, what we understood from theory about how categories form, and what we knew about the difference between excess-of-mass and leaf clustering methods. And the payoff was interpretive, discourse from one category of actors trending differently from other categories, impossible to notice unless in the light of our entangled knowings.
Where does this leave LLMs? They are remarkable tools, and I used one to help form the text of this very post. But I don't think they are the right tool for computationally entangled interpretive research, at least not as a stand-in for the interpreter. I resist the framing of a lone scholar with an LLM as collaborator, because that framing still outsources a good bit of judgment to a general-purpose model. I also think LLMs are the gas guzzlers of the algorithmic world, and I'm not about to roll coal for my beloved qual research. Got to draw a line somewhere.
My position is that this work is best done in interdisciplinary collaboration: people who separately hold deep knowledge of theory, phenomenon, and method, in dialogue with each other. It's people willing to keep talking to each other as the corpus grows past what any one of us can read. Will this be faster, more efficient, yield more publications than if a lone author mobilized a frontier model and a bajillion usage credits? I don't know. But I do know that doing it this way yields, to me at least, a deeper understanding of phenomenon, theory, and method. It yields new questions and complications and pipelines of research beyond the question at hand. I am a better scholar and teacher for it.
ADDENDUM August 25th.
Having squinted at this post for a couple days, I think I have a couple more things to say, or rather clarify what I am NOT saying with the above. After all, how could I possibly know what I want to write until after I have written it?
1. I'm not trying to change computational social science writ large, although the methology could use some reflexivity around the human-in-the-loop conceit, how "in the loop" is the human anyway? No this is about what people might recognize as more or less "traditional" qualitative research, both the more positivist-ish researcher-independent inductive case flavor, and the more interpretive and narrative-driven (abductive?) flavor (If you care, I get pedantic about these categories of qual here). It is this type of research that is most directly threatened by both ballooning corpora and the increasing prevalence of consumer-grade generative AI. Which brings me to...
2. Rolling coal analogy notwithstanding, I'm not anti-AI. I use generative AI (Claude these days), both at work and at home. I dabble in implementing LLMs. I think genAI has a lot of interesting, innovative organizational use cases. I am however concerned about the corporate, profit-driven nature of most commercial genAI and firmly believe in a future where we have a mix of corporate-run AI systems as well as locally run open-source and adjacent systems. The closest analogy is the respective roles of Windows, Mac, and Linux operating systems in the modern digital era. Here, though, I want to emphasize that genAI systems are not the right tool for the job specifically when it comes to qual research, both in the narrow sense of performing the research at hand, and in the broader sense of an unfolding career as a scholar who conducts qualitative research. It doesnt feel right, it leaches the joy out of qual, and people who I respect as excellent qualitative scholars (who have enthusiastically experimented with AI and qual) dont think it does a good job... and they would know.
Comments
Post a Comment