One (good?) thing about moving houses is that we move stuff from its usual place, open boxes and drawers that had not been opened in a long time and find stuff that we had forgotten about. One such find, in last month’s move, were notes for work that Jan Kietzmann and I did about 15 years ago, about how to assess the value of social media to support marketing decision making. We presented the idea at the 2012 Academy of Marketing Conference, and even received a Best Paper in Track Award. Though (and this is one of my biggest regrets), we never got round to writing a full journal paper. You can find the slides for that conference paper, entitled “A Conceptual Investigation of the Value of Social Media Data as a Source of Customer Insight”, here.
At the time, our focus was social media data. Organisations were becoming increasingly interested in using what people posted online to understand customers, identify trends and inform marketing decisions.
We were developing a framework for thinking about the value of social media data, which forced users to look beyond the fact of whether the data were accurate or not not. Our framework asked managers to look beyond the intrinsic characteristics of the data and consider whether those characteristics made the data suitable for the particular decision or application.
Fast forward to 2026
We have considerably more data, and more ways of collecting, combining and analysing it. And, yet, as noted in the paper “The Myth of Good Data: Data Integrity in a Shifting Ecosystem”, co-authored by Kathleen H. Pine and Melissa Mazmanian, and published in MIS Quarterly, many organisations still can’t get the benefits of data-driven initiatives. So, Pine and Mazmanian set to out to find out why.
In the paper, Pine and Mazmanian argue that it is wrong to talk about data quality as an intrinsic property of data – i.e., that it is wrong to say that data are good or bad. Instead, they argue that data do not exist independently of the organisational ecosystem in which they are produced, categorised, stored, interpreted and used. And, so, rather than conceptualise data quality as a permanent, inherent characteristic of data, we should consider it as situated and relational attribute. They call this data integrity, and define it as the “the alignment between data resources and the situated needs, expectations, and organizational and technological infrastructures of a particular use scenario” (p. 616).
The authors illustrate this through a study in the healthcare context. They illustrate how “data that were considered to have sufficient completeness, validity, consistency, and accuracy by the people relying on it to take action or assess a situation were suddenly (and often unexpectedly) seen as no longer having these properties when the expectations, projections, or desired use of the data shifted”. That is, Pine and Mazmanian show that data can remain accurate, complete and consistent, yet no longer be fit for purpose.
The paper goes on to document how organisations must engage in “data rehabilitation” to solve the problem. This is a labour-intensive effort that depends heavily on retraining staff, changing daily collection habits, and altering workflows, and which falls disproportionately on frontline administrative staff.

From social media data to AI
Looking back at that 2012 presentation, I am struck by how much the context has changed (the corpus of data and how easy it is to access it as now bigger than ever), but the fundamental problem hasn’t: if we want to use data, we need to look beyond accuracy and assess how well it fits the problem / decision in question.
And, as Pine and Mazmanian show, this fit needs to be revisited every time the organisation, the technology or the purpose of the data changes. The faster data moves through organisations, the more easily assumptions about its meaning can become invisible. And AI introduces another complication: increasingly, the data ecosystem itself includes AI-generated or AI-transformed data. This makes questions about provenance, interpretation and fitness for purpose even harder to ignore.
Implications for managers
Pine and Mazmanian highlight the following practical implications:
- Acknowledge that the intensive work of data rehabilitation often falls on staff at the bottom of the hierarchy. The weight of organidational success is increasingly held by frontline workers who, typically, are not compensated for their increased efforts with higher salaries or organizational recognition
- Recognise that data rehabilitation has costs, and that organisations with limited financial resources may struggle to invest in doing so, possibly undermining high-stakes data-driven initiatives.
- Depending on the perceived criticality of data rehabilitation efforts, personnel at the top of the hierarchy may need to shift their work practices, to align with the needs of a shifting data ecosystem. It is possible that some will resist this, which undermines results.
To these, I would add one more: Do not force an existing dataset onto a different problem. As the organisation, technology and decision context change, the fit between data and its intended use needs to be reassessed.
Looking back at the question Jan and I were asking in 2012, I am even more convinced that we were asking the right question. If only we had managed to prioritise writing that paper!

