The Day You Can No Longer Trust a Phone Call

The Day You Can No Longer Trust a Phone Call

AI can now clone a familiar voice, fake a face and create evidence that never existed. As deepfake technology spreads, the simple act of hearing someone may no longer be enough to know they are real.

The Call

Your phone rings. It is your son. His voice is shaking. He says he has been in an accident and needs money sent somewhere in the next twenty minutes. You know that voice. You have known it since it first learned to speak. There is no doubt in your mind that it is him, because doubt requires a reason, and a familiar voice has never given you one.

Except it is not him.

This already happened to the Trapp family in the Bay Area. Fox News reported that they received a frantic call from what sounded exactly like their son, saying he had caused a car accident, injured a pregnant woman, and was about to be arrested. A second voice on the line, posing as a police officer, told the mother to withdraw 15,000 dollars in cash and hand it to a courier already on the way, and she withdrew it and handed it over. Her son had never been in any accident. Every voice on that call had been synthetic.

Until recently, reproducing a specific person’s voice well enough to sustain a real-time emergency call was beyond the reach of ordinary criminals. Faking it would have required an actor who could imitate someone’s reflexes and specific way of being surprised, not just their sound. With enough usable audio pulled from a video or a voicemail greeting, software can now generate new sentences in that voice, saying things the person never said, to someone who has every reason to believe them. Criminals running this scam do not need to become better con artists. They need a voice sample and a plausible emergency, and both have gotten much easier to get.

When Seeing Stops Being Believing

The next stage is harder to dismiss, because it involves seeing, not just hearing.

In January 2024, a finance employee at the engineering firm Arup, working out of its Hong Kong office, received an email that appeared to come from the company’s UK based chief financial officer, asking him to carry out a confidential transaction. He hesitated, then joined a video call where people who looked and sounded like his CFO and several colleagues directed him to make fifteen transfers totaling around 25 million dollars to accounts in Hong Kong. Hong Kong police later said the attackers had built the deepfakes from video and audio of the real executives pulled from past conferences and webinars. Nothing was confirmed as fraudulent until he called head office afterward and was told no such transaction had ever been authorized.

Arup’s own investigation traced the entire loss to one thing: an employee’s confidence that the people on his screen were real. The company’s systems and data were never touched. The exploit was the same one that worked on the Trapps, just scaled up. The loss was simply far larger.

The Perfect Alibi

There is a second effect that may end up mattering more than the fakes themselves. Once people know convincing fakes exist, real evidence becomes easier to dismiss.

A genuine recording surfaces showing someone saying or doing something they would rather deny. A few years ago, that recording would have been damaging regardless of what the person claimed. Today they have a new option that costs nothing and requires no proof: they can simply say it was generated by AI. Sometimes they will be right. Other times they will not be, and there may be no fast or universally trusted way for the public to tell the difference.

Researchers have started calling this the liar’s dividend, a world where the mere existence of convincing fakes gives cover to people caught doing something real, because doubt is now cheap to manufacture and expensive to resolve. The damage done by deepfakes reaches past the fake recordings themselves and into every genuine recording that can now be waved away with a single sentence.

The clearest real-world case is Gabon’s, from 2018. President Ali Bongo, out of public view for months after a stroke, released a New Year’s video meant to prove he was still capable of governing. His stiff movements and irregular blinking led political opponents to declare the video a deepfake. A week later, the military cited that claim as part of the justification for an attempted coup. Forensic analysts who later examined the footage found no clear evidence it had been altered, and most concluded it was probably genuine, a man recovering from a stroke, not a fabrication. The coup failed, but the sequence became the standard example researchers point to when they explain the liar’s dividend, since no one ever had to prove the video was fake for the claim alone to do the damage.

The Institutions Built on Evidence

Courts are dealing with a version of this from the opposite direction, evidence faked rather than real evidence denied. In the case of Mendones v. Cushman & Wakefield, NBC News reported that a judge on California’s Alameda County Superior Court grew suspicious of a video the plaintiffs had submitted to support their own claim, in which a witness’s face looked strangely fuzzy and her voice sounded flat and disjointed. The judge determined the footage had been generated with AI, and dismissed the case for submitting fabricated evidence. The fabrication came from inside the case itself, a plaintiff trying to manufacture proof for their own argument, and it was caught only because the fake still had visible flaws, flaws that may not remain as the technology improves. A University of Colorado Boulder-led report found that more than 80 percent of court cases already hinge to some degree on video evidence, which is exactly what makes this shift consequential.

Elections face a more public version of the same threat. Two days before the 2024 New Hampshire presidential primary, thousands of voters received a robocall in a voice that sounded exactly like President Biden’s, urging them to skip the primary. The consultant behind it was indicted and fined six million dollars by the Federal Communications Commission, which went on to ban AI generated voices in robocalls entirely. The call proved how little effort it now takes to put a convincing lie in a familiar public voice into thousands of phones at once.

Newsrooms carry a quieter version of the same exposure. Verifying a photograph or a clip used to mean confirming the source. It increasingly means examining the file itself for signs of synthesis. In 2025, Washington Post reporter Drew Harwell posted a demonstration showing how quickly a tool called Sora 2 could generate convincing fake police bodycam footage, exactly the kind of material a newsroom might once have treated as beyond dispute.

Your Identity Becomes an Attack Surface

Underneath the fraud and the legal questions sits a more basic shift. People have spent the last two decades giving their voices and faces away freely, through videos, podcasts, interviews, and voice notes sent to friends. A recording used to be just a recording. It can now become raw material for synthetic imitation, and the more of it exists publicly, the easier a convincing clone becomes to build.

That raises a question the law is still slow to answer. If public footage of someone can be used to produce new material of them saying things they never said, what has actually been taken from them, and who is responsible for it? This is not limited to executives or public figures. Anyone with a handful of public videos, a wedding livestream, a school presentation, has enough material circulating to become a plausible target.

Children are a distinct case. Parents routinely post years of a child’s voice and face online long before that child has any say in it, and none of that footage disappears. It sits there as a permanent, public record, exactly the kind of raw material a cloning tool draws on. That permanence is why some of the defenses described next are already being built pre-emptively, in households that have never been targeted.

The Cost of Proving Reality

None of this stays convincing forever on its own, and it does not need to. A fabricated clip can reach millions of people within minutes of being posted, while confirming or debunking it typically takes hours or days, and a fake released shortly before a market announcement or an election only needs to be believed long enough to trigger the reaction its creator wanted. By the time a correction arrives, the reaction it was meant to prevent has usually already happened. Companies with the resources to build dedicated fraud teams can absorb that lag. Most individuals cannot, and are left with their own judgment in a moment engineered to overwhelm it.

That gap is now being closed inside companies with specific procedures rather than general caution: callbacks before acting on financial instructions received by phone or video, multi-person approval for unusual transfers instead of trusting a single voice on a call, and out-of-band verification, confirming a request through a channel separate from the one it arrived on. The scale being defended against is documented, not hypothetical. The Deloitte Center for Financial Services has tracked generative AI enabled fraud losses in the US rising from 12.3 billion dollars in 2023 toward a projected 40 billion dollars by 2027. The FBI’s Internet Crime Complaint Center recorded more than 3,100 fraud complaints from adults over 60 in 2025 that specifically referenced AI, with losses exceeding 352 million dollars, more than 5 million of it tied to voice-cloned distress calls of exactly the kind that fooled the Trapps.

Inside families, the response looks a lot less technical. Some are agreeing on a private word or question, something only a real family member would know, for the moment someone calls claiming to be in danger and needing money fast. One of the more sophisticated forms of mass impersonation ever built is being met, in some households, with the same kind of solution people used before phones existed at all. Every version of this response costs something: slower transactions, extra security steps, a general tax on interactions that used to require none of it. All of it is spent protecting one thing, confidence that the person on the other end is real, and the money is simply what follows from that.

The Authenticity Economy

For much of the internet’s history, finding information was the hard part. Now information can be generated faster than anyone can consume it, and the volume will only grow from here. Proof has become the scarce thing: a verified recording, a confirmed live conversation, a photograph that can be traced back to an actual moment and an actual camera, worth more now than the sheer volume of content surrounding it.

That proof is starting to take shape as its own infrastructure: cryptographic signatures attached to media the moment it is captured, cameras that can certify their own output, provenance records that travel with a file wherever it goes. Banks are adding callback steps. Courts are learning to spot the seams in fabricated evidence. Families are agreeing on passwords. None of these are the same mechanism, but they are all answers to the same problem: the old shortcut to trust, hearing a voice, seeing a face, no longer works on its own.

The Trapps have a password now, for the one call they hope never comes again. That is what it costs now to answer a phone, or watch a video, the way people used to.

Yogendra Singh
Yogendra Singh

Yogendra Singh is the founder and editor of Structural Signals, an independent publication covering long-term trends in technology, economics, energy, geopolitics and society.

Articles: 67

Leave a Reply

Your email address will not be published. Required fields are marked *