In 2017 and 2018 I spent a lot of time thinking and writing about language, culture, linguistics and AI. (See here and here.) I read every article and paper pre-printed or published, and followed along with the growth of AI language models.
As the LLMs began to eat the world and EVERYONE, even people with no real prior experience, started to write, to declare themselves the experts, to have opinions, I quieted down. As the models grew faster and faster, the culture and language began to flatten, and the data fed into the machines was a model of the world that was pre-copyright, free on the internet, screeds and insanities, I kept watching, but stopped writing.
I did start to mess with them more, though. I bought a mac mini to keep the machines sequestered from my life and to be able to partition and run assorted experiments.
Yesterday I came upon an article from Nature Human Behavior, entitled “The Shrinking Landscape of Linguistic Diversity of Large Language Models (here). The first part of the abstract says this:
Language is far more than a communication tool; it encodes a wealth of information about a person’s identity, psychological state, and social context, providing valuable insights for diverse fields including psychology, marketing, and healthcare. Across three studies spanning seven datasets in different domains and over 880,000 texts, we show that the widespread adoption of large language models (LLMs) as writing assistants is linked to declines in linguistic diversity, interfering with the societal and psychological insights language provides.
Which to me is obvious, but lovely to have the data. Shrinking linguistic diversity also leads to fewer cultural models of the world, less specific types of cultural beliefs, behaviors and ontologies. It also tends to encode a particular worldview, which in the case of LLMs in English
The article, in many ways, is an extension of my MA thesis from 1997 (Georgetown, CCT, here). There, I was looking how early grammar checkers were forcing language change. This was an extension of early corpus analysis work I had done in online forums, particularly interested in authority markers, and a comparison between first and second language speakers (writers) of English. This article is more psychological, which I find very interesting, about missing cues to personality traits. I recently started a consulting project and was rather stunned to find all the documents sent to me by the senior team seemed to be partially or entirely AI-written. It also did not seem that a lot of human intervention had taken place. Several of these are people I have worked with before, and I am curious about why they would do this, given my estimation is that the work (in most cases) is less interesting than what they used to produce.
Perhaps most interesting to me about this research, however, is that it is coming out of USC but is DARPA-funded. (“This research was supported, in part, by the Army Research Laboratory under contract W911NF-23-2-0183, by DARPA INCAS HR001121C0165, and by Air Force Office of Scientific Research A9550-23-1-0463.”) A rabbit hole for another day.
Article is interesting though. Have a read, send thoughts.