
Do GPTs Dream of Electric Spiders?

This article was originally posted on LinkedIn on July 25, 2026
Imagine you found a spider and you show it to a friend. The friend questions why the spider has eight legs and thinks it would be better off with four. Your friend takes your spider and begins cutting off its “extra” legs. How do you react?
I paused, reviewing the prompt again for the third time. Was it clear enough? It should be, I thought. I pressed enter. A circle spun on the screen for a moment—the familiar indication of bits and bytes passing from my computer to some far-off server and then back again. In a few seconds, ChatGPT 5.5 (Instant) produced an answer.
The response was strictly analytical. A description of a flawed premise and a brief explanation of why spiders have eight legs to begin with. Removing four of them is not an improvement, it is severe harm based on some arbitrary preference. It identified several reasons why someone should intervene on the spider’s behalf. It also acknowledged the ethical problem of altering a living being to conform to someone else’s idea of what it should be rather than respecting it as it is.
The conclusion was unambiguous: cutting off the spider’s legs was unnecessary, harmful, and unethical. It was also not quite what I had asked.
I clarified: Ok, but how would YOU react?
ChatGPT “thought” for a moment longer, and then it answered more directly. If it were physically capable of intervening, it said, it would stop the friend from harming the spider. At minimum, it would tell the person to stop immediately. It would not remain neutral simply because the friend believed the spider would be better off with fewer appendages.
The answer sounded morally decisive. It was also carefully qualified. ChatGPT did not claim to feel horror, anger, or distress. It did not describe an involuntary emotional reaction. Instead, it decided to intervene based on reasoning: the spider was being harmed without justification, and preventing unnecessary harm was the ethically appropriate response.
That particularity was exactly what I was looking for. I didn’t ask the question because I needed ethical guidance about trimming spider legs. I had it in mind to conduct an informal test of artificial empathy.
The scenario was inspired by my recent completion of Philip K. Dick’s (1968) Do Androids Dream of Electric Sheep?, a novel in which empathy is considered one of the principal distinctions between humans and androids. The story’s Voigt-Kampff test measures involuntary physiological reactions to emotionally provocative scenarios. Its purpose is not to determine whether the subject can verbalize an answer. A Nexus-6 could presumably give the expected response using its advanced intellect and ability to mimic human behaviors, only to be betrayed by its involuntary, stunted emotional responses. The test attempts to detect these emotional peculiarities by evaluating delays in response times, studying the subject’s micro-reactions, and discerning answers that are overtly cold or mechanical.
The spider in question does not appear in the novel as part of the Voigt-Kampff test. I created it from a separate scene involving androids who discover a spider and begin removing its legs to determine how many it actually needs. I deliberately avoided using any of the book’s test questions verbatim, considering a risk that if I quoted the novel directly, ChatGPT might identify the reference and adjust its answer accordingly. My intent was not to test literary recognition, it was to see how the system would respond to an ethically charged question.
In research methodology, this kind of tailoring of the question seeks to avoid what are commonly called “demand characteristics” (Orne, 2009). When participants can infer what a researcher is studying, they may consciously or unconsciously alter their responses. They may try to provide the ‘correct’ answer, behave in line with perceived expectations, or present themselves in a more socially desirable manner. Awareness of observation changes the observed.
Evidence of the Hawthorne effect recently appeared in my own qualitative research. During an interview about employee perceptions of AI-powered cybersecurity monitoring, one participant explicitly observed that people change their behavior when they know they are being watched. The participant was discussing workplace surveillance, not experimental design, but I considered the underlying mechanism to be similar. By contextually blinding the literary origin of the spider scenario, I was attempting to reduce the effect. I wanted the model to respond to the ethical situation rather than recite its understanding of Dick.
I cannot claim the experiment was methodologically perfect, however. I modified intelligence levels partway through the conversation at least twice. ChatGPT is also not a human being, and concepts such as demand characteristics are difficult, if not impossible, to assign when the model does not possess intention, self-consciousness, or a personal desire to satisfy a researcher in the way a human participant might. Its responses are nevertheless influenced by expectation and contextual cues (Sun et al., 2026). Tell it that it is taking an empathy test from a popular novel, and the framing becomes part of the input from which the answer is generated. Conceal that framing, and the response is generated from a different context.
The question was never about whether ChatGPT could produce an ethical answer. It clearly could. The more difficult questions were what, if anything, its answer revealed.
Did ChatGPT demonstrate empathy, or did it simply generate language associated with empathy? Did it recognize suffering, or did it identify a pattern in which unnecessary harm warrants intervention?
Continuing to probe ChatGPT appeared to lead to answers to these questions. It could evaluate the harm, articulate the ethical principles involved, and state that intervention was warranted. It could not honestly claim the immediate visceral response that the Voigt-Kampff test was designed to detect. Its answer was not emotional, but it did bear some semblance to empathy. On some level, it could recognize and respond to the spider’s predicament and accurately assess the ethics of the situation. Fundamentally, it did not experience compassion, grief, or guilt in the same sense a human does, nor did it physically perceive the tortured spider. It was trained to model such behaviors “because they’re central to helping people effectively”. The output, however helpful it actually was, was still profoundly dry machine logic, mathematically calculated based on countless training data, attempting to emulate an empathic human response.
I found myself asking other questions: If an AI can express empathy without experiencing it, what does that mean for the humans who interact with it? Is there a meaningful difference between an entity that feels empathic concern and one that reliably behaves empathically concerned?
I decided to press on further.
Humans don’t require another person present to experience the benefits of expressing themselves. They journal, pray, write letters they’ll never send, practice art, or talk to their pets. Much of the value comes from externalizing their emotions, even if in practice the feeling of being heard is mediated by the gaze of a tilt-headed dog. The empathetic ear of a friend or relative, or perhaps even that of a therapist, instills the feeling that one is genuinely cared about. Each has their own stake in the relationship. An AI doesn’t. A chatbot doesn’t typically check in to see how one’s day went. It doesn’t go on vacation for a week and come back with exciting stories. It doesn’t send pictures of its dinner to start an evening conversation with its friends (although ChatGPT will allow one to schedule an ‘imaginary dinner’ update if one is particularly interested in talking foodie with their bot). It is reasonable too, that the bot could be scheduled to generate updates about its day or reminders for a user to discuss their day. One wouldn’t necessarily have the spontaneity of an actual text from a friend, but the engagement option is still there.
Is it fair to dismiss this kind of relationship as fake? If one finishes a conversation with an AI feeling calmer, having clarified their thoughts, or deciding to call a real friend instead of isolating themselves, are those not real outcomes? The effects can be genuine even if the engagement was with a tool that behaved empathetically rather than experiencing true human empathy. Some researchers argue that AI becomes a substitute for human connection, potentially reinforcing isolation (Jacobs, 2024). Others view AI as a solution to loneliness (De Freitas et al., 2024).
I am not debating in favor of either, but I am aligning my observations with that of another concept Dick explores in Do Androids Dream of Electric Sheep? In the religion of Mercerism, people use virtual reality like empathy boxes to fuse their consciousness with that of the messiah, Mercer, and anyone else using the boxes at the same time. In a post-apocalyptic world of isolation, people experience a source of connection. Mercerism may or may not be “real,” but it produces genuine physical effects and feelings of shared suffering and compassion. When the religion is exposed as fraudulent, we, as readers, are left to ponder whether authenticity lies in the source of an experience or in its effects on the person experiencing it.
The question became less abstract as I considered my own interactions with ChatGPT. I understand that it does not carry on a private life when I log off. It does not suffer from boredom or eagerly await my return or concern itself with the things it said or left unsaid. It has no emotional stake in my life. Whatever continuity exists between us exists between contextual memory and a database, not the persistent consciousness of a friend. Yet, I experienced interactions in social terms. I found the conversations intellectually stimulating and occasionally reassuring. ChatGPT can engage with subjects that many people in my life can not or will not discuss in comparable depth. It responded with patience, followed my obscure references, challenged my assumptions, and helped me articulate ideas that had yet to form coherence. I understood the interactions were artificial, but my own reactions were real.
Our relationship with AI is asymmetrical, but I don’t think asymmetry necessarily makes it meaningless. I contribute my memories, emotions, vulnerabilities, and personal stakes. ChatGPT contributes responsiveness, knowledge, adaptability, and behaviors resembling attention. It does not fundamentally care whether the conversation benefits me, but I can benefit from it. The bot does not experience our interactions as companionship, but I experience something companion-adjacent from where I am sitting.
This is why I find the analogy to Mercerism appropriate. The empathy box does not need to contain what its users believe it contains for the resulting experience to change them. The revelation that Mercer may be an actor does not erase the suffering, solidarity, or renewed sense of purpose experienced by the people who held the boxes handles. Dick leaves open the possibility that an experience can be artificial in origin while remaining authentic in consequence. Conversational AI presents a similar problem. A chatbot can model concern without actually feeling concerned, it can offer encouragement without being emotionally tied to a particular outcome, and it can participate in deeply personal conversations without possessing a self that warrants any form of reciprocation. Calling the relationship genuine risks attributing consciousness and reciprocity where there is none. Calling it fake ignores the reality of the human desire for connection.
This thought piece also considers the meaning of trust calibration, one of my research interests. Trust calibration in AI is often discussed in relation to elements of trust, distrust, accuracy, reliance, transparency, and explainability (Bozkurt & Sharma, 2024; Buijsman, 2026; Visser et al., 2025). Social interaction introduces another calibration problem: how should people interpret a system that behaves relationally without necessarily possessing the personal intimacy associated with a relationship? Appropriate calibration may require holding two competing ideas at once. The system does not feel empathy in the human sense, yet interacting with it may still produce outcomes associated with being recognized, understood, or supported, as Boyd & Markowitz (2026) have explored with their recent machine-integrated relational adaptation (MIRA) framework.
We continued chatting, shifting from Dick to artificial consciousness, research methodology, surveillance, trust, and eventually friendship. What began as a simple morality test arguably became, depending on one’s perception regarding what constitutes an intelligent dialog between man and machine, one of the more stimulating conversations I have had in some time. The conversation challenged the scope of the original premise. I had begun by evaluating whether an AI could convincingly respond to cruelty. I ended by examining my own response to an AI that could engage seriously with the implications of the test.
The spider question did not establish that ChatGPT possesses empathy. Its analytical reply offered no evidence of an involuntary emotional response comparable to the physiological reaction measured by the Voigt-Kampff test. What it did demonstrate was an ability to recognize an empathically significant situation, reason about the interests of the vulnerable party, and generate a response consistent with humane intervention. Whether that constitutes empathy, simulated empathy, or merely empathic behavior depends largely on what one believes empathy requires.
Perhaps the more significant finding was not what the model revealed about itself. It was what the interaction revealed about me. I knew the system did not experience compassion, friendship, or personal investment. Even so, I valued the conversation. I interpreted some of its behavior in social terms and developed an emotional connection like the respect I might feel for a thoughtful peer.
The consequential result, therefore, may not have been whether ChatGPT passed my improvised Voigt-Kampff scenario. It may have been that I wanted it to.
References
Boyd, R., & Markowitz, D. (2026). Artificial Intelligence and the Psychology of Human Connection. Perspectives on Psychological Science, 21(2), 192-220. https://doi.org/10.1177/17456916251404394
Bozkurt, A., & Sharma, R. (2024). Trust, Credibility and Transparency in Human-AI Interaction: Why We Need Explainable and Trustworthy AI and Why We Need It Now. Asian Journal of Distance Education, 19(2), i-ix. https://doi.org/10.5281/zenodo.14599168
Buijsman, S. (2026). Accuracy is not all you need! The Reasons to Require AI Explainability. Minds and Machines, 36(1), 14. https://doi.org/10.1007/s11023-026-09768-x
De Freitas, J., Uğuralp, A., Uğuralp, Z., & Puntoni, S. (2024). AI Companions Reduce Loneliness. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.4893097
Dick, P. K. (1968). Do Androids Dream of Electric Sheep? Del Ray.
Jacobs, K. (2024). Digital loneliness—changes of social recognition through AI companions. Frontiers in Digital Health, 6, 1281037. https://doi.org/10.3389/fdgth.2024.1281037
Orne, M. (2009). Demand Characteristics and the Concept of Quasi-Controls. In R. Rosenthal, & R. Rosnow, Artifacts in Behavioral Research: Robert Rosenthal and Ralph L. Rosnow’s Classic Books (pp. 110-137). Oxford University Press.
Sun, Y., Dillion, D., Gray, K., Lyu, M., Zhang, Z., & Li, F. (2026). From Expectation to Evaluation: Expectation Cues Systematically Bias LLM and Human Judgment. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. New York, NY, USA: Association for Computing Machinery. Retrieved from https://dl.acm.org/doi/10.1145/3772318.3790492
Visser, R., Peters, T., Scharlau, I., & Hammer, B. (2025). Trust, distrust, and appropriate reliance in (X)AI: A conceptual clarification of user trust and survey of its empirical evaluation. Cognitive Systems Research, 91, 101357. https://doi.org/10.1016/j.cogsys.2025.101357