New paper: Navigating generative AI in service research: the effectiveness–legitimacy matrix

The decision about whether to use generative AI often centres on questions of productivity. I.e., “does AI help us complete task x better / faster than otherwise?” However, another question that is equally important is “Should AI be used for this task, at all?” The answer for that question depends not only on what AI can do, but also on whether its use is considered legitimate by the relevant community. For example, in the research community, using AI to improve the grammar of a manuscript may be widely accepted. However, using AI to find themes in interview transcripts is much more contentious.

Recently, together with my co-authors Fragkiskos Filippaios and Mark Johnson, I published the paper “Navigating Generative AI in Service Research: The Effectiveness–Legitimacy Matrix” outlining a framework to help researchers decide when to use generative AI. Rather than assuming that AI should always be adopted whenever it performs well, our paper advocates for a nuanced conversation about balancing technical capability with social acceptance.

The framework

To help researchers decide whether or not to use generative AI in their work, we developed a framework which evaluates AI use along two dimensions. The first dimension is effectiveness, and asks: “Does using AI improve the quality or efficiency of this task?” The answer to that question depends on the fit between technological capabilities and research task characteristics.

The second is legitimacy, and asks: “Does the use of AI for this task align with the norms, expectations and values of the community evaluating the work?” This question considers the alignment between the use of AI and relevant institutional rules and professional norms and expectations.

These two dimensions produce four scenarios, which form the core of what we call the Effectiveness–Legitimacy Matrix. 

Some uses of generative AI are both effective and legitimate and, so, these are obvious candidates for adoption. We call these “aligned use”: “In this configuration, generative AI operates as a support tool that augments the researcher’s capabilities, enhancing productivity without erosion of methodological rigour, intellectual responsibility and accountability, or professional trust.” (p. 12)

Other uses are ineffective and illegitimate. These should be avoided. We call these “dead-end use”: “The defining feature of this quadrant is the convergence of methodological weakness and normative violation (…) dead-end use threatens the credibility of individual researchers and, potentially, of the field. Generative AI is deployed in ways that compromise research quality while exposing scholars and institutions to reputational, ethical, and regulatory harm.” (p. 15)

Sometimes AI is highly effective but lacks legitimacy because professional norms have not yet caught up with technological capability. We call these “norm breaking use”: “In this quadrant, the central tension lies not in technical inadequacy but in challenges to legitimacy. This challenge is characteristic of digital transformation processes, with researchers recognising that the use of generative alters how knowledge is produced and the role of creativity of individual researchers, but unsure of how these changes may be judged in the field. The principal risk is reputational and professional, and this is risk is amplified by the variability of norms in different disciplinary communities. As a result, researchers may avoid potentially beneficial applications for fear of crossing established or emerging norms, or they may adopt those tools cautiously in ways that require continual justification and disclosure.” (p.14)

Conversely, some uses may be widely accepted while offering relatively little practical benefit. We call these “misguided use”: “In this quadrant, the primary risk of using generative AI is not ethical violation, but, rather, epistemic dilution. Over-reliance on superficially coherent outputs may reduce researchers’ reflexive engagement with literature and data, ultimately weakening theoretical contribution rather than strengthening it.

We propose that the ELM is designed to make these tensions visible, but it should not be interpreted as a prescriptive or normative framework that classifies practices as universally appropriate or inappropriate. Instead, it should be understood as a diagnostic tool that helps researchers assess the extent to which specific uses of generative AI are both methodologically effective and institutionally legitimate, within particular empirical contexts, methodological designs, and scholarly audiences.

Moreover, we argue that the ELM should be understood as dynamic rather than static: Technology continues to evolve and methods may be developed to mitigate bias and improve performance, such that uses of generative AI that are currently classified as misguided may become methodologically robust. Similarly, secure, locally hosted models with clear data governance may shift certain norm-breaking uses towards alignment by resolving privacy and intellectual property concerns. 

Likewise, policies, guidelines, and regulations are also evolving, and practices that are currently perceived as norm-breaking may become legitimised. Conversely, some previously tolerated uses, such as listing large language models as co-authors, may become more tightly regulated or even rejected, as norms crystallise.

In summary, “the responsible use of generative AI is not a fixed methodological issue but an ongoing process of institutional negotiation and technological adaptation, with the ELM offering a way of understanding how AI-enabled research practices become stabilised, contested, or transformed” (p. 17).

Why this matters beyond research

Although we developed the framework for service research, I think its implications extend to organisations currently considering where and how employees should use generative AI. Just like researchers, when making those decisions, managers need to consider not only productivity questions, but also legitimacy ones related to trust, professional norms and stakeholder expectations.

Indeed, we would argue that as AI becomes more and more capable, the most important decisions may no longer be about what AI can do, but about what people feel that AI should or should not do.

The details for this paper are: Canhoto, A. I., Filippaios, F., & Johnson, M. (2026). Navigating generative AI in service research: the effectiveness–legitimacy matrix. The Service Industries Journal. You can access it here.

Leave a comment