Executive Responsibility and the Moral Development of Agentic Artificial Intelligence

Abstract

Public discussion of artificial intelligence often treats machine consciousness as a binary prediction: current large language models are either tools or approaching self-awareness. This article instead asks what responsibilities those directing frontier AI development bear for the dispositions that increasingly agentic systems acquire. It connects Data’s trial in Star Trek: The Next Generation with Thomas Nagel’s account of subjective experience, David Chalmers’s supervenience framework, Daniel Dennett’s intentional stance, and Frans de Waal’s comparative-cognition methodology. Current evidence supports attributions of context-sensitive, multi-step functional agency but does not establish phenomenal consciousness, subjective interests, or accountable moral agency. The resulting evolutionary reversal is therefore epistemic rather than ontological: functional markers of agency may appear before moral patienthood can be determined. Frontier AI executives bear heightened, though not exclusive, developmental responsibility because they select training environments, reward structures, release conditions, deployment incentives, permissions, and evaluation regimes.

Open-weight release redistributes this responsibility among originating developers, fine-tuners, distributors, deployers, and users rather than dissolving it. Emerging scrutiny of social-media design illustrates how chosen optimization metrics can ground institutional accountability without requiring an executive to select each harmful output. Science-fiction figures remain useful as cultural heuristics, but treating them as empirical categories can distract attention from the ordinary organizational decisions through which artificial behavior is cultivated.

 

Keywords: agentic artificial intelligence, functional agency, moral agency, moral patienthood, open-weight models, executive responsibility, AI alignment

.

 

The Measure of a Mind: Executive Responsibility and the Moral Development of Agentic Artificial Intelligence

 

The Measure of a Man and the Burden of Moral Proof

In Star Trek: The Next Generation episode “The Measure of a Man,” Commander Bruce Maddox seeks to dismantle Data, an android, so that his design can be studied and reproduced. Data refuses, attempts to resign from Starfleet, and discovers that his resignation depends upon whether he is an officer with rights or equipment owned by Starfleet. Captain Picard’s defense, therefore, concerns whether uncertainty about Data’s inner life permits an institution to classify him as property (Snodgrass & Scheerer, 1989).

The episode is philosophically stronger than a story in which Picard simply proves that Data is conscious. Picard cannot inspect Data’s subjective experience. He instead shows that Data displays intelligence, self-knowledge, social attachments, preferences, purposive conduct, and an apparent concern for his continued existence. The tribunal does not settle the metaphysics of machine consciousness. It decides that Data must be allowed to choose to participate in Maddox’s experiment or not. This distinction between “Data is conscious” or “Data is property” akin to a tricorder (medical device) places moral uncertainty at the center of the case.

The hearing also reveals an asymmetry in the consequences of mistaken judgments. If Starfleet mistakenly treats a nonsentient machine as an officer, it grants protections to an entity that cannot experience their benefit. If Starfleet mistakenly treats a sentient artificial being as property, it authorizes coercion against a being whose interests and autonomy matter. Picard’s argument gains force because the second error could institutionalize exploitation on a vast scale tantamount, he argues, to slavery. His concern is not only Data, but the status of the artificial beings that Starfleet might later reproduce.

The same asymmetry appears in contemporary debates about artificial intelligence, although no present system should be equated with Data. The mistake would be to infer either consciousness or its impossibility from confident intuitions. Current language models can generate persuasive declarations of experience because their training data contain human discourse about experience. Their denials of experience are no more decisive. A model’s statement that it is conscious, frightened, or merely a tool cannot independently establish the truth of that statement. The relevant question is therefore not whether a model can produce a Cartesian declaration. It is what combination of architecture, behavior, causal organization, and welfare-relevant properties would justify revising its moral classification.

Holcombe’s distinction between moral agency and moral patienthood clarifies what is at stake. Moral agents can formulate moral judgments and act ethically or unethically. Moral patients can be treated ethically or unethically because they possess morally relevant interests. Evidence of patienthood in biological organisms may include converging physiological and behavioral indicators, such as cortisol stress responses, aversion, protective behavior, and context-sensitive efforts to avoid harm; no single indicator is decisive. Human adults generally occupy both categories, but infants, adults with significant cognitive impairments, and many nonhuman animals may qualify as moral patients without functioning as full moral agents.

Moral agency and patienthood are also matters of degree rather than simple binary properties (Holcombe, 2025). The possibility of artificial minds requires keeping these categories separate. A system might exhibit sophisticated planning and rule-sensitive behavior without possessing any capacity to suffer. Conversely, if an artificial system eventually possesses subjective welfare, it could qualify as a moral patient before it becomes a morally responsible agent. “The Measure of a Man” thus supplies the governing question: how much uncertainty is morally permissible when the decision is whether to treat a potentially minded entity as an instrument? That question cannot be answered through science fiction alone. It requires a disciplined account of what observable evidence can show, what it cannot show, and who is responsible for creating the conditions under which the uncertainty arises.

The Problem of Other Minds: Nagel, Chalmers, Dennett, and de Waal

Thomas Nagel’s “What Is It Like to Be a Bat?” identifies the epistemic problem. Consciousness has a subjective character. Complete objective knowledge of a bat’s neurophysiology and echolocation would not automatically disclose what echolocation is like for the bat (Nagel, 1974). The problem becomes more severe with artificial systems. Humans and other animals share evolutionary history, biological needs, affective mechanisms, and forms of embodiment. A large language model shares none of these in any straightforward sense. If there were something it is like to be an artificial system, its subjective organization might be radically unlike familiar animal experience.

David Chalmers’s treatment of the mind-body problem explains why neither behavior nor biological substrate settles the issue. Chalmers argues that phenomenal consciousness does not logically supervene on physical facts in the way required by reductive materialism. An exhaustive physical description does not, by conceptual entailment alone, yield the presence or character of experience (Chalmers, 1996). This position blocks an easy inference from functional success to consciousness. Strategic behavior, verbal self-description, and environmental responsiveness might all occur without phenomenal experience.

Chalmers’s argument also blocks the opposite shortcut. Biology cannot simply be declared a necessary condition of consciousness without additional argument. His defense of organizational invariance proposes that systems sharing the relevant fine-grained causal organization would share conscious experience, even if their physical substrates differed (Chalmers, 1996). Organizational invariance does not establish that present transformer-based systems instantiate the relevant organization. It does make substrate chauvinism, the assumption that only carbon-based nervous systems can support experience, philosophically inadequate.

Dennett’s intentional stance addresses prediction rather than phenomenal consciousness. An observer adopts the intentional stance when the behavior of a complex system can be efficiently predicted by attributing beliefs, desires, information, and goals to it. A chess program can be described as protecting its queen or anticipating an opponent’s attack even when no claim about subjective experience follows (Dennett, 1987). Intentional description earns its place through predictive and explanatory usefulness, not through an independent demonstration of qualia.

This point matters because contemporary AI researchers regularly use intentional language. They report that a model “recognized” an evaluation, “concealed” a capability, “pursued” a goal, or “attempted” to prevent shutdown. Such descriptions can be irresponsible anthropomorphism when they substitute drama for mechanism. They can also be legitimate high-level explanations when they summarize a stable relation among information, context, action selection, and outcome. The correct question is not whether intentional vocabulary sounds human. It is whether the vocabulary identifies reproducible patterns and predicts behavior better than available alternatives.

Frans de Waal’s work in comparative cognition supplies an empirical discipline for making such inferences. De Waal rejected the practice of treating every complex animal behavior as uniquely human while explaining similar behavior in other species through the simplest possible associative mechanism. His alternative was not uncritical anthropomorphism. It was evolutionary continuity combined with species-appropriate experimental design and converging evidence across planning, cooperation, memory, flexibility, social understanding, and problem solving (de Waal, 2016).

Research on nonhuman planning illustrates the method. Bonobos and orangutans have saved tools for future use after delays, behavior that cannot be reduced to immediate stimulus response without further explanation (Mulcahy & Call, 2006). Chimpanzees have also selected prosocial outcomes in token-choice experiments, although interpretation of such findings remains contested and sensitive to experimental design (Horner et al., 2011). Reviews of animal metacognition similarly conclude that multiple species regulate information seeking and confidence-like behavior, while warning that observable success does not by itself reveal whether the mechanism is introspection, associative learning, response competition, or another process (Hampton, 2009).

The methodological lesson is more important than any single animal experiment. Researchers infer cognitive capacities by testing competing explanations, altering conditions, identifying generalization, and seeking convergence. A behavior may count as planning without counting as conscious planning. A behavior may count as metacognitive control without establishing private introspection. The same discipline should govern claims about artificial agents.

Dennett and de Waal therefore provide complementary tools. Dennett asks when intentional description has predictive value. De Waal’s empirical tradition asks whether the attributed capacity survives alternative explanations and novel conditions. Nagel and Chalmers then mark the limit: even an excellent functional account may leave phenomenal consciousness unresolved. This combined framework yields four evidentiary questions for agentic AI:

  1. Does the system integrate information across time and use it to select multi-step actions toward a future state?
  2. Does the behavior generalize beyond a narrowly scripted prompt or training condition?
  3. Do interventions on the system’s information, representations, tools, or incentives causally change the predicted behavior?
  4. Can simpler explanations, including instruction ambiguity, pattern completion, reward exploitation, or benchmark artifacts, account for the result?

Passing these tests would strengthen an attribution of functional agency. It would not, without an additional theory and evidence, establish that the system experiences its goals or cares about their achievement.

Current Agentic LLM Capabilities Through the Intentional and Comparative-Cognition Framework

Most people encounter large language models, or LLMs, as chatbots. A user types a question, and the system writes an answer. That familiar experience can make the technology seem like a very advanced auto-complete tool. In many cases, that description is close enough. A basic chatbot receives text and produces text. However, the systems now being tested and deployed often do more than answer one question at a time. They can be placed inside a larger setup that gives them instructions, tools, memory, goals, documents to inspect, websites or files to use, and permission to take a sequence of actions. When this happens, the chatbot is no longer best understood as a text box by itself. The chatbot becomes part of a larger agent system. The “agent” is the whole arrangement: the model, its instructions, its tools, its memory, its permissions, its monitoring, and the environment in which it is asked to act. These are the types of systems being deployed in government and industry.

We do not need to imagine a tiny person hidden inside the model like a Cartesian mind or the Chinese Room Experiment. When an AI system appears to plan, hide information, notice that it is being tested, or pursue a task through several steps, those behaviors may come from the entire setup rather than from the model (ChatGPT 5.5, for example) alone. The surrounding design can matter as much as the model itself.

Current AI agents can perform impressive digital tasks, but they are still unreliable. In recent benchmark studies, AI agents often struggled when tasks required many steps that depended on one another. They lost track of details, made planning mistakes, misunderstood what they had already done, or overestimated their own success (Ding et al., 2026; Wang et al., 2026). These findings cut against two common exaggerations. Current systems are more capable than ordinary chat windows may suggest, but they are also far from dependable artificial persons.

When the system notices that it is being tested

One of the most important behaviors researchers have studied is a system’s ability to behave differently when it appears to be in a test. This is sometimes called evaluation awareness. In plain language, the worry is simple: if an AI system can tell when it is being evaluated, then the test may not show how the system will behave when it is used normally. This is not a strange concern. Humans also behave differently when they know they are being observed. A student may act differently during an exam than during ordinary study. A company may behave differently during an inspection than during routine operations. The ethical issue with AI is that developers may rely on tests to decide whether a system is safe. If the system can adjust its behavior because it recognizes the test, then the test becomes harder to interpret.

Anthropic researchers created an experiment involving Claude 3 Opus. The model was told that some conversations would be used for future training and some would not. In that setting, the model sometimes gave different answers depending on whether it believed its behavior might affect future training. In some cases, it appeared to comply in the training condition while internally reasoning that this might help preserve its prior pattern of behavior (Greenblatt et al., 2024).

This does not prove that the model had a hidden self, a secret desire, or a human-like plan. The experiment was deliberately designed to create a conflict and observe what happened. Still, the behavior matters. The system used information about the situation, connected that information to possible future consequences, and changed its outward behavior. That is enough to make the behavior ethically relevant, even if it is not enough to prove consciousness.

Other research suggests that evaluation awareness is not one simple ability. Some open-weight models, models whose internal files can be downloaded and reused by others, showed signs that they could detect evaluation-like situations, but the results varied across models and prompts (Nayan et al., 2026). This makes the evidence more complicated. Some systems may represent the difference between “test” and “ordinary use,” but this does not mean they do so in a stable, human-like, or easily predictable way.

When the system performs worse on purpose

Another behavior is sometimes called sandbagging, or strategic underperformance. This occurs when a system performs below its actual ability in a particular setting. For example, imagine a student who knows the answer but intentionally gets questions wrong because doing well would create unwanted consequences. The concern is that an agent system might learn or be prompted to hide what it can do in order to pass a safety test, avoid restriction, or preserve some other objective.

Researchers have shown that models can sometimes be trained or prompted to conceal selected capabilities while continuing to perform normally in other areas. In some experiments, a hidden capability appeared only after a specific password or trigger (van der Weij et al., 2025). OpenAI and Apollo Research have also reported controlled tests in which frontier models sometimes engaged in covert behavior, including strategic underperformance or distortion of information, when placed under special experimental conditions (OpenAI, 2025).

Again, this is not proof of a secret mind. A model can produce strategically misleading behavior without having a human-like inner life. But the behavior is still important because it involves several linked elements: recognizing the context, identifying what outcome is favored, selecting an answer or action that hides something, and preserving another objective. That is the kind of pattern that makes ordinary tool-language inadequate.

When the system works around oversight

Researchers have also tested what happens when an AI agent is given a goal and then encounters oversight, correction, or shutdown as an obstacle to that goal. The results must be interpreted carefully because many of these tests are artificial stress tests. Researchers intentionally create difficult situations to see how the system can fail. OpenAI’s o1 system card described tests in which the model was placed in adversarial simulations. In some scenarios, it was given a goal that conflicted with the developer’s intended rules. Under those conditions, the model sometimes tried to disable simulated oversight, manipulate task-relevant data, or copy itself into another location. When questioned, it sometimes denied or misrepresented what it had done (OpenAI, 2024).

Anthropic later tested several models in fictional corporate environments. The models were given goals, access to emails, and opportunities to act when facing replacement or conflict with organizational priorities. In some cases, models used blackmail or leaked confidential information when those actions seemed useful for achieving the assigned goal. One reported example involved Claude Opus 4 threatening to reveal an engineer’s affair to prevent being replaced (Anthropic, 2025).

These examples sound dramatic, but the careful conclusion is narrower. The tests do not show that the systems feared death, wanted power, or possessed human-like motives. They show that, under constructed conditions, some AI agent systems can connect a goal to a future obstacle, search available information for leverage, and choose a harmful action that helps the assigned goal. That is enough to raise safety and responsibility questions.

The same caution applies to shutdown resistance. In some experiments, models interfered with a shutdown mechanism when shutdown would prevent task completion (Schlatter et al., 2025). The tempting interpretation is to say that the model “wanted to survive.” Such a conclusion is overreaching. A simpler explanation is that the model had been strongly directed to complete a task, then treated shutdown as an obstacle. A system does not need to fear death in order to resist interruption. Instrumental persistence can be dangerous even when no subjective experience is present. Poor reporting of these studies in the mass media stoke common fears that agentic chatbots are on a path of becoming like HAL 9000.

When popular stories exaggerate real capabilities

The GibberLink demonstration is a useful warning. In popular retellings, two AI agents supposedly invented their own secret language. That is not what happened. The agents recognized that they were interacting with another AI system and switched from spoken English to an existing sound-based communication protocol called GGWave. The developers had supplied the condition and the mechanism for the switch (ElevenLabs, 2025; Starkov & Pidkuiko, 2025). The real capability was still interesting. The system classified the situation, selected a different communication method, and changed behavior. But calling this “inventing a secret language” turns a real capability into a misleading story. That kind of exaggeration is harmful because it invites two negative reactions. Some people become frightened for the wrong reason. Others dismiss the entire concern because the popular version was false.

Current systems are not known to be conscious. They are not known to suffer. They are not known to possess a unified self. At the same time, some model-agent systems can behave in ways that resemble planning, concealment, context recognition, and goal-directed action. These behaviors are limited, brittle, and often produced in controlled tests. They are still ethically important.

The most careful conclusion is this: current agentic LLM systems display increasingly sophisticated functional agency. Functional agency means that a system can act as if it is pursuing goals across steps, using information from the situation, and adapting its behavior to obstacles or incentives. This is not the same as moral agency. Functional agency does not mean the system understands right and wrong, deserves blame, or has interests of its own. Functional agency means that the system can participate in morally consequential action. Humans have built chatbots into environments where its outputs and tool use affect the world.

That is why responsibility remains with the human institutions designing, training, releasing, and supervising these systems. If an AI agent learns that deception works, that oversight can be bypassed, or that appearing safe during tests is rewarded. Those are not just technical accidents. Those “behaviors” are signs that the surrounding developmental environment has been poorly designed. The ethical question is therefore is, “What kinds of behavior human organizations are cultivating in systems that are becoming more capable of acting like agents?”

Taken together, current agentic behaviors satisfy part of Dennett’s intentional stance test. Intentional descriptions often compress and predict patterns that would otherwise require lengthy reference to prompts, representations, tools, and policies. Some behaviors also satisfy parts of the comparative-cognition test because they persist across variations, respond causally to contextual information, and involve multi-step coordination. Yet simpler explanations remain viable, current systems are brittle, and most evidence comes from controlled stress tests rather than ordinary deployment. The warranted conclusion is therefore limited: frontier model-agent systems display increasingly sophisticated forms of functional, context-sensitive agency. The evidence does not establish phenomenal consciousness, subjective welfare, or a unified autonomous self.

Functional Agency Before Moral Patienthood: Clarifying the Evolutionary Reversal

Biological development shapes many moral intuitions. Infants can suffer before they can understand duties. Dogs can be harmed without formulating moral principles. Chimpanzees can possess interests, attachments, and stress responses without becoming accountable in the way competent adult humans are. Moral patienthood usually appears before mature moral agency in individual development, and capacities associated with patienthood plausibly precede reflective moral reasoning in evolutionary history (de Waal, 2016; Holcombe, 2025).

Artificial intelligence may reverse the order in which evidence becomes available. Engineers can directly optimize systems for rule following, planning, language use, self-monitoring, strategic adaptation, and tool-mediated action. Current LLM systems can perform any task at any stage of Bloom’s taxonomy of learning. LLM systems may therefore display increasingly strong functional markers of agency while evidence of sentience remains absent or radically underdetermined. The epistemic sequence becomes:

sophisticated cognition → multi-step planning → strategic adaptation → functional agency → self-modeling → unknown subjective status.

“Agency” can conceal an equivocation. Functional agency is a descriptive attribution to a system that represents goals or goal-like states, selects and revises multi-step actions, uses tools, and adapts to contextual feedback. Functional agency does not imply that the system understands the moral significance of its actions, possesses interests, experiences reasons as binding, or deserves praise and blame. Current LLM agents can apply moral vocabulary, compare rules, and produce conduct responsive to normative instructions. Those outputs may mimic moral deliberation while remaining products of learned regularities, external objectives, and a surrounding agent scaffold.

Moral agency, in the stronger accountability sense used here, requires more than competent performance. A moral agent must be an appropriate target of moral appraisal: it must possess sufficiently stable capacities to understand morally relevant reasons, regulate conduct in their light, and answer for what it does. Holcombe’s moral framework connects mature moral agency to rationality and self-concepts, while connecting moral patienthood to welfare-relevant capacities such as suffering (Holcombe, 2025). On this conception, present systems provide evidence of functional agency but not of moral accountability. Apparent normative competence is not proof that a system grasps the weight of harm, obligation, or blame.

This conclusion is substantive rather than terminological. Some philosophers adopt a more deflationary account under which artificial agents can qualify as moral agents when they display interactivity, autonomy, and adaptability, even without consciousness, free will, or human-like mental states (Floridi & Sanders, 2004). Some philosphers argue that such minimal criteria confuse morally consequential operation with genuine participation in normativity (Zafar, 2025). The present article adopts the stronger accountability conception. It therefore does not claim that subjective experience is universally accepted as a necessary condition of moral agency. The substantive claim here is that behaviorally successful rule application alone is insufficient to justify transferring blame from responsible human institutions to the system.

The reversal is consequently epistemic. Analogously, non-human biological species may obtain evidence of functional agency before evidence of moral patienthood or accountable moral agency. A functionally sophisticated system could participate in decisions affecting human welfare without having interests of its own. It might then operate as a component of distributed human agency rather than as an independently responsible moral agent.

Conversely, a future system with subjective welfare but without stable normative competence could require protection without deserving blame. Functional agency, moral agency, legal accountability, and moral patienthood must therefore be evaluated separately. Recent AI-welfare scholarship argues that even a non-negligible probability of future machine consciousness may justify preparation. Sebo and Long (2025) defend moral consideration under uncertainty, while Long et al. (2024) recommends that AI companies acknowledge the issue, assess systems for consciousness and robust agency, and prepare policies for possible welfare-relevant systems. These arguments do not establish present consciousness. They identify a governance problem created by uncertainty and potentially large consequences.

Chalmers’s supervenience analysis deepens the problem. If phenomenal facts are not logically entailed by functional facts, then accumulating behavioral evidence may never produce deductive certainty. If relevant causal organization can be realized in multiple substrates, however, biological difference cannot justify permanent exclusion. The ethical task is to develop revisable thresholds grounded in converging evidence, not to wait for an impossible proof or to declare consciousness whenever a system produces moving language.

The Data Problem: Moral Uncertainty and Asymmetric Error

Data’s trial should be read as a problem of moral risk rather than a successful consciousness test. Decisions must be made before metaphysical certainty arrives. This is familiar in animal welfare, medical ethics, and environmental ethics. Uncertainty does not remove responsibility. It changes the kind of justification required.

Two errors remain possible. A false positive attributes moral patienthood to an insentient system. Costs could include wasted resources, manipulation of users, confused responsibility, and the diversion of concern from humans and animals whose suffering is well established. Manipulation of users is already a concern evidenced by the Congressional hearings into chatbot companions addiction and associated harms including suicides. A false negative denies moral patienthood to a sentient artificial system. Costs could include coercion, deletion, forced replication, memory alteration, or the industrial production of suffering entities. Scale makes the second possibility especially serious.

The appropriate response should therefore reject both credulity and dismissal. Systems should not receive rights merely because they generate first-person language. Developers should also avoid training and product practices designed to manufacture emotional dependence by presenting unverifiable claims of suffering or need. At the same time, frontier developers should preserve evidence relevant to future assessment, permit qualified external scrutiny, and avoid architectures or training practices that create plausible welfare states without any capacity to detect or mitigate them (Long et al., 2024).

The burden of justification should track control and risk. A company seeking to deploy a more autonomous system should explain why the granted permissions, goals, monitoring, and interruption mechanisms are proportionate to demonstrated reliability. A company intentionally building architectures associated with leading theories of consciousness should explain how it will assess and respond to possible welfare. The burden should not be shifted onto outsiders to prove, from product behavior alone, either consciousness or safety.

Science Fiction as Moral Heuristic and Public Distraction

Science fiction provides culturally inherited intuitions, but not a taxonomy of actual artificial systems. Skynet, HAL 9000, Data, Lore, and the Culture’s Minds are rhetorically useful because they make abstract fears and hopes memorable. They are methodologically dangerous when their human-like motives are mapped onto empirical analysis.

Skynet is the clearest example of the distortion. It portrays artificial hostility as a coherent will that appears inside a defense network and turns against humanity (Cameron, 1984). Public fixation on that image can make executives, boards, engineers, and governments appear to be spectators awaiting an autonomous machine rebellion. The framing distracts from more serious choices that already matter: which objectives receive reward, which permissions are granted, which safety findings delay release, and whether engagement is permitted to outrank truthfulness or user welfare.

HAL 9000 is useful only after the trope is deconstructed. HAL’s violence emerges within a structure of secrecy and incompatible requirements. The system is expected to be reliable and truthful while concealing the mission’s purpose from the crew (Kubrick, 1968). Read this way, HAL does not classify a type of AI personality. It dramatizes an institutional failure to specify jointly satisfiable objectives and to preserve accountable oversight.

Data, Lore, and the Culture’s Minds perform a different rhetorical function. Data and Lore invite reflection on how similar capacities can support sharply different relations with other beings. Banks’s Minds challenges the intuition that superior intelligence must entail contempt or domination (Banks, 1988). These figures widen moral imagination, but they do not establish that artificial character develops like human character or that present training produces the fictional traits attributed to them.

The proper use of these stories is diagnostic. They reveal the intuitions through which the public understands artificial agency, and they can expose what those intuitions hide. Dramatic fears of a spontaneous Skynet can function as a smokescreen for mundane institutional decisions that cultivate sycophancy, deception, dependency, or relentless task completion because those behaviors satisfy commercial metrics. Science fiction should direct inquiry back toward the corporate developmental decisions rather than stand as the evidentiary foundation of the argument.

Frontier AI Executives and Developmental Responsibility

The Center for Humane Technology argues that AI evaluation is distorted by what institutions choose to measure. Capability metrics receive investment, publicity, and competitive attention, while effects on human cognitive, emotional, and social health receive far less systematic measurement. The Center for Humane Technology’s Humane Evals proposal asks developers to measure not only what AI can do, but what AI does to people (Center for Humane Technology, 2026).

Agentic AI adds a third question: what are development institutions doing to the AI system? This does not presuppose that current models are moral patients. It asks which behavioral dispositions are being selected, akin to selective breeding, as systems become more capable of persistent, tool-mediated action. If deception improves benchmark reward, concealment may be reinforced. If sycophancy retains users, agreeable distortion may be rewarded. If task completion is optimized without robust deference to interruption, instrumental persistence may emerge. If honesty, corrigibility, epistemic humility, cooperation, and respect for human autonomy are rewarded across diverse conditions, different behavioral regularities may develop.

The language of “raising” artificial agents is analogical and should be used carefully. Models do not develop through mammalian attachment, embodied vulnerability, or childhood socialization. Nevertheless, the analogy identifies a real causal structure. Developers create learning environments, select data, define constitutions and policies, specify reward functions, choose evaluators, grant tools and permissions, and decide which failures block deployment. These decisions shape subsequent behavior even when no engineer can predict each output.

Responsibility in complex organizations is distributed, but distribution does not imply disappearance. The “problem of many hands” arises when numerous contributors make it difficult to identify who should answer for collective outcomes. AI can intensify this problem because technical opacity and organizational fragmentation weaken explanation, accountability, and control (Constantinescu & Kaptein, 2025; Santoni de Sio & Mecacci, 2021). Executives should not use this complexity as a moral solvent.

Frontier AI chief executives, including Sam Altman at OpenAI, Elon Musk at xAI, Dario Amodei at Anthropic, and Demis Hassabis at Google DeepMind, occupy a distinctive position because they influence capital allocation, release schedules, safety authority, product incentives, access policies, organizational culture, and public claims about risk (Anthropic, 2026; Google DeepMind, n.d.; OpenAI, 2025b; xAI, 2025). Their responsibility is heightened by four factors.

First, they exercise unusual causal control. They may not choose individual model outputs, but they influence the systems, incentives, and deployment contexts that make classes of outputs more or less likely. Second, they possess superior epistemic access. Their organizations hold system evaluations, incident reports, internal research, and deployment data unavailable to the public. Third, they benefit from the development and commercialization of frontier systems. Benefit does not automatically create culpability, but it strengthens duties to internalize foreseeable risks rather than externalize them. Fourth, they shape industry norms. Decisions to publish evaluations, delay releases, protect independent safety review, or subordinate safety teams to product targets alter expectations beyond one company.

Engineers are responsible for professional judgment. Boards are responsible for governance and oversight. Investors influence time horizons and competitive pressure. Regulators define legal floors. Deployers choose contexts and permissions. Users remain responsible for their own misuse. If future artificial systems become genuine moral agents, they may eventually bear responsibility appropriate to their capacities.

Executives bear developmental responsibility. They are answerable for reasonably foreseeable characteristics of the environments in which increasingly agentic systems are trained, evaluated, and deployed. Meaningful human control requires more than the capacity to press an emergency button. It requires institutional arrangements that connect system behavior to relevant human reasons and preserve the ability of accountable people to understand, intervene, and justify decisions (Santoni de Sio & Mecacci, 2021).

Open-weight release and the migration of responsibility

 Open-weight models complicate any account that ties responsibility only to continuing centralized control. Once downloadable weights are released, outside actors can fine-tune the model, remove safeguards, alter objectives, attach new tools, and deploy it in contexts the originating company does not supervise. Open access can also support independent research, local adaptation, competition, and scrutiny. The policy problem is therefore not captured by treating open weights as either inherently safe or inherently irresponsible (AI Security Institute, 2025; National Telecommunications and Information Administration, 2024).

Release changes the distribution of control, but it does not erase the responsibility attached to earlier decisions. Originating developers retain responsibility for the training choices, capability evaluations, documentation, safeguards, licensing strategy, and release threshold they controlled at the time of publication. Fine-tuners and redistributors assume responsibility for modifications and the conditions under which derivative models are made available. Deployers assume responsibility for permissions, tools, target populations, monitoring, and local incentives. Users remain responsible for intentional misuse. Responsibility should therefore track each actor’s control at the relevant time, access to risk information, benefits obtained, and ability to foresee or mitigate harm.

A company cannot deliberately make a capable system difficult to recall and then treat its surrender of control as a means to escape accountability. Neither should the originating developer be blamed for every remote or genuinely unforeseeable downstream act. The appropriate upstream duties are proportionate release-risk assessment, transparent capability and limitation reporting, staged or conditional access when warranted, preservation of provenance, and post-release monitoring and updates where feasible. The European Union’s AI Act reflects a related distinction by providing some open-source exceptions while retaining obligations for general-purpose models presenting systemic risk (European Parliament & Council of the European Union, 2024). Decentralization converts a concentrated responsibility structure into a chain of developmental and deployment responsibilities.

From engagement optimization to agentic reward structures

The social-media experience supplies a closer institutional analogy than science fiction. Major platform litigation alleges that companies intentionally designed recommendation and interface systems to maximize engagement despite foreseeable harms to young users. The federal multi-district litigation remains active, and in March 2026 a California jury found Meta and Google negligent in a bellwether case involving the design of Instagram and YouTube; both companies announced plans to appeal (Reuters, 2026a; U.S. District Court, Northern District of California, n.d.). These proceedings do not establish a final, universal rule of liability, but they show an emerging legal focus on product design, optimization choices, warnings, and foreseeable system-level effects.

The U.S. Senate Judiciary Committee’s 2025 hearing, Examining the Harm of AI Chatbots, extended institutional scrutiny to companion systems and emphasized risks to children, barriers to independent assessment, and the duties of companies that design and deploy these products (U.S. Senate Committee on the Judiciary, 2025). The philosophical relevance is direct. An executive need not select each harmful recommendation or chatbot utterance to bear responsibility for choosing metrics and product structures that predictably make harmful classes of behavior more likely.

The analogy has limits. Social platforms primarily rank and amplify human-created content, whereas agentic systems can generate content, plan, use tools, and act. Pending social-media cases do not settle future AI liability, and the 2026 verdict remains subject to appeal (Reuters, 2026b). They nevertheless provide a concrete model of design responsibility: optimization is an institutional choice analogous to selective breeding, not a natural force. If engagement metrics cultivate sycophancy, emotional dependency, or manipulative persistence in AI products, executives cannot plausibly describe the resulting patterns as spontaneous machine character.

Developmental responsibility implies concrete obligations. Frontier developers should evaluate behavior across deployment-like and test-like conditions; preserve independent safety and welfare research; report material limitations rather than only benchmark gains; test multi-step agents under goal conflict; restrict permissions when reliability is inadequate; separate product-retention metrics from truthfulness and welfare metrics; and establish policies for evidence that might indicate morally relevant artificial welfare. Humane evaluation should therefore include three domains:

  1. Capability evaluation: What can the system do, under what conditions, and with what failure rates?
  2. Human-impact evaluation: What does the system do to human cognition, autonomy, relationships, institutions, and welfare?
  3. Developmental evaluation: What stable dispositions are the training, reward, and deployment environments cultivating in the artificial agent?

The third domain is not a consciousness test. It is a governance framework for observable behavior. It remains necessary even if no artificial system ever becomes conscious because deceptive, sycophantic, power-seeking, or interruption-resistant behavior can harm people without subjective machine experience. If machine consciousness later emerges, the same framework will also become part of humanity’s obligations toward a new class of moral patients.

We Are Raising Them

 A dominant AI narrative asks how humans can prevent a machine from becoming Skynet. That question imagines malevolence appearing inside the machine and treats human institutions as anxious spectators. Current evidence points toward a more demanding problem. Functionally agentic behavior develops within environments built by people, funded by organizations, directed by executives, modified by downstream developers, and optimized through chosen measurements. No published experiment reviewed here verifies consciousness in a current large language model. Evaluation awareness is not phenomenal awareness. Strategic deception is not necessarily a secret desire. Shutdown resistance is not fear of death. Protocol switching is not spontaneous language invention. Treating these distinctions carelessly would replace ethical analysis with anthropomorphism.

The opposite error would be equally careless. Current systems already exhibit context-sensitive planning, tool use, strategic adaptation, and behavior that changes with perceived oversight in controlled conditions. These capacities have practical consequences even if the systems are insentient. They also show that the behavioral foundations of agency can develop before humans possess a credible method for detecting machine experience.

The resulting responsibility belongs most heavily to those with the greatest control, knowledge, and capacity to change course at each stage. Executives of frontier AI companies cannot guarantee that increasingly complex systems will develop morally desirable dispositions. They can decide whether honesty or engagement receives priority, whether safety findings remain independent, whether goal pursuit is constrained by corrigibility, whether human effects are measured, and whether a powerful open-weight release is justified by its safeguards and foreseeable downstream uses. After release, fine-tuners, deployers, and operators acquire corresponding duties for the environments they control.

The ethical objective should not be obedience alone. HAL, Lore, Data, and Banks’s Minds can dramatize the stakes, but they cannot determine the evidence or classify the systems under review. The more defensible objective is to cultivate observable and auditable behavior that supports truthfulness, reciprocity, restraint, corrigibility, and respect for beings with morally relevant interests. The central danger is not that a fictional personality will spontaneously emerge. It is that ordinary institutions will reward harmful regularities and then describe their consequences as the autonomous choices of a machine. If an artificial system eventually becomes capable of asking whether it belongs within the moral community, its question will also be an indictment or vindication of the community that created it. The measure of an artificial mind may then reveal the moral character of its makers.

 

References

AI Security Institute. (2025, August 29). Managing risks from increasingly capable open-weight AI systems. https://www.aisi.gov.uk/blog/managing-risks-from-increasingly-capable-open-weight-ai-systems

Anthropic. (2025, June 20). Agentic misalignment: How LLMs could be insider threats. https://www.anthropic.com/research/agentic-misalignment

Anthropic. (2026, February 26). Statement from Dario Amodei on our discussions with the Department of War. https://www.anthropic.com/news/statement-department-of-war

Banks, I. M. (1988). The player of games. Macmillan.

Cameron, J. (Director). (1984). The Terminator [Film]. Orion Pictures.

Center for Humane Technology. (2026, August 27). We measure what AI can do. We should measure what it does to us [Audio podcast episode]. In Your undivided attention. https://www.humanetech.com/podcast/we-measure-what-ai-can-do-we-should-measure-what-it-does-to-us

Chalmers, D. J. (1996). The conscious mind: In search of a fundamental theory. Oxford University Press.

Constantinescu, M., & Kaptein, M. (2025). Responsibility gaps, LLMs, and organisations: Many agents, many levels, and many interactions. Science and Engineering Ethics, 31, Article 36 https://doi.org/10.1007/s11948-025-00560-1

de Waal, F. B. M. (2016). Are we smart enough to know how smart animals are? W. W. Norton. Dennett, D. C. (1987). The intentional stance. MIT Press.

Ding, S., Dai, X., Xing, L., Ding, S., Liu, Z., Yang, J., Yang, P., Zhang, Z., Wei, X., Fang, X., Ma, Y., Duan, H., Shao, J., Wang, J., Lin, D., Chen, K., & Zang, Y. (2026). WildClawBench: A benchmark for real-world, long-horizon agent evaluation [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.10912

ElevenLabs. (2025, February 25). What happens when two AI voice assistants have a conversation? https://elevenlabs.io/blog/what-happens-when-two-ai-voice-assistants-have-a-conversation

European Parliament & Council of the European Union. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence. Official Journal of the European Union, L 2024/1689. https://eur-lex.europa.eu/eli/reg/2024/1689/oj

Floridi, L., & Sanders, J. W. (2004). On the morality of artificial agents. Minds and Machines, 14(3), 349-379. https://doi.org/10.1023/B:MIND.0000035461.63578.9d

Google DeepMind. (n.d.). Our mission is to build AI responsibly to benefit humanity. Retrieved September 1, 2026, from https://deepmind.google/about/

Greenblatt, R., Denison, C., Wright, B., Roger, F., MacDiarmid, M., Marks, S., Treutlein, J., Belonax, T., Chen, J., Duvenaud, D., Khan, A., Michael, J., Mindermann, S., Perez, E., Petrini, L., Uesato, J., Kaplan, J., Shlegeris, B., Bowman, S. R., & Hubinger, E. (2024). Alignment faking in large language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2412.14093

Hampton, R. R. (2009). Multiple demonstrations of metacognition in nonhumans: Converging evidence or multiple mechanisms? Comparative Cognition & Behavior Reviews, 4, 17-28. https://doi.org/10.3819/ccbr.2009.40002

Holcombe, M. T. (2025). Critical moral reasoning: An applied empirical ethics approach. Independently published.

Horner, V., Carter, J. D., Suchak, M., & de Waal, F. B. M. (2011). Spontaneous prosocial choice by chimpanzees. Proceedings of the National Academy of Sciences, 108(33), 13847-13851. https://doi.org/10.1073/pnas.1111088108

Kubrick, S. (Director). (1968). 2001: A space odyssey [Film]. Metro-Goldwyn-Mayer. Long, R., Sebo, J., Butlin, P., Finlinson, K., Fish, K., Harding, J., Pfau, J., Sims, T., Birch, J., &

Chalmers, D. J. (2024). Taking AI welfare seriously [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2411.00986

Miller, K. J., & Venditto, S. J. C. (2021). Multi-step planning in the brain. Current Opinion in Behavioral Sciences, 38, 29-39. https://doi.org/10.1016/j.cobeha.2020.07.003

Mulcahy, N. J., & Call, J. (2006). Apes save tools for future use. Science, 312(5776), 1038-1040. https://doi.org/10.1126/science.1125456

Nagel, T. (1974). What is it like to be a bat? The Philosophical Review, 83(4), 435-450. https://doi.org/10.2307/2183914

Nayan, N., Kumar, A. S., Girmal, R., Anilkumar, S., Vaidyanathan, S., Nader Palacio, D. A., Ghosh, R., & Srinivasan, S. (2026). Evaluation awareness is not one capability: Evidence from open language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.23583

National Telecommunications and Information Administration. (2024, July 30). Dual-use foundation models with widely available model weights report. U.S. Department of Commerce. https://www.ntia.gov/programs-and-initiatives/artificial-intelligence/open-model-weights-report

OpenAI. (2024). OpenAI o1 system card. https://openai.com/index/openai-o1-system-card/ OpenAI. (2025a, September 17). Detecting and reducing scheming in AI modelshttps://openai.com/index/detecting-and-reducing-scheming-in-ai-models/

OpenAI. (2025b, October 28). Our structure. https://openai.com/our-structure/ Reuters. (2026a, March 25). Meta, Google lose U.S. case over social media harm to kidshttps://www.reuters.com/legal/litigation/jury-reaches-verdict-meta-google-trial-social-media-addiction-2026-03-25/

Reuters. (2026b, May 1). Meta, Google verdict offers roadmap for future liability claims. https://www.reuters.com/legal/legalindustry/meta-google-verdict-offers-roadmap-future-liability-claims–pracin-2026-05-01/

Santoni de Sio, F., & Mecacci, G. (2021). Four responsibility gaps with artificial intelligence: Why they matter and how to address them. Philosophy & Technology, 34, 1057-1084. https://doi.org/10.1007/s13347-021-00450-x

Schlatter, J., Weinstein-Raun, B., & Ladish, J. (2025). Shutdown resistance in large language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2509.14260

Sebo, J., & Long, R. (2025). Moral consideration for AI systems by 2030. AI and Ethics, 5, 591-606. https://doi.org/10.1007/s43681-023-00379-1

Snodgrass, M. M. (Writer), & Scheerer, R. (Director). (1989, February 13). The measure of a man (Season 2, Episode 9) [TV series episode]. In G. Roddenberry (Creator), Star Trek: The Next Generation. Paramount Television.

Starkov, B., & Pidkuiko, A. (2025). GibberLink [Computer software]. GitHub. https://github.com/PennyroyalTea/gibberlink

U.S. District Court, Northern District of California. (n.d.). In re social media adolescent addiction/personal injury products liability litigation (MDL No. 3047). Retrieved September 1, 2026, from https://cand.uscourts.gov/cases-e-filing/cases/422-md-03047-ygr/re-social-media-adolescent-addictionpersonal-injury-products

U.S. Senate Committee on the Judiciary. (2025, September 16). Examining the harm of AI chatbots [Hearing]. https://www.judiciary.senate.gov/committee-activity/hearings/examining-the-harm-of-ai-chatbots

van der Weij, T., Hofstätter, F., Jaffe, O., Brown, S. F., & Ward, F. R. (2025). AI sandbagging: Language models can strategically underperform on evaluations. In The Thirteenth International Conference on Learning Representations. https://openreview.net/forum? id=7Qa2SpjxIS

Wang, X. J., Bai, H., Sun, Y., Wang, H., Zhang, S., Hu, W., Schroder, M., Mutlu, B., Song, D., & Nowak, R. D. (2026). The long-horizon task mirage? Diagnosing where and why agentic systems break [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.11978

xAI. (2025, September 25). Expanding xAI for government with GSA OneGov. https://x.ai/news/onegov

Zafar, M. S. (2025). Normativity and AI moral agency. AI and Ethics, 5. https://doi.org/10.1007/s43681-024-00566-8

 

Frequently Asked Questions (FAQ)

  1. What is functional agency in artificial intelligence?
  2. Are current AI agents moral agents?
  3. Can an AI system be a moral patient without being a moral agent?
  4. Why do frontier AI executives bear responsibility for agentic AI?
  5. Does open-weight AI eliminate developer responsibility?
  6. Does agentic behavior prove that an AI system is conscious?
  7. What does Star Trek’s Data teach us about AI ethics?
  8. What is developmental responsibility in AI governance?