Editor’s Note: This piece is part of my ongoing tech advocacy, systems-change analysis, and reporting for Unembedded, informed by more than 15 years working at the intersection of technology, media, and human rights. I examine complex technical and policy developments through an intersectional, public-interest lens—asking how systems of power translate into real-world consequences, who they protect, who they fail, and what those failures demand of companies, policymakers, and institutions.
“Refugees are welcome here, even if they rape our women, because white people do that too.”
Imagine opening your phone to discover a video of you delivering that exact, venomous statement. Within hours, you are the centerpiece of a viral internet storm, framed by a hyper-realistic, AI-generated clone of your own face giving a vicious speech you never uttered. This was the horrifying reality for Dorothy McHugh, a 77-year-old Scottish Labour councillor in Dundee. The malicious video, posted in November 2025, weaponized her likeness during a highly sensitive local protest over asylum-seeker housing, altering public trust and severely damaging her personal integrity.
Before the shock of that political hit could even settle, a second deepfake launched a parallel attack on a young Muslim woman just across the English Channel. A dedicated health advocate working to de-stigmatize menstrual hygiene for minority girls awoke to find her face stolen. A malicious actor had taken a real television news interview she gave, stolen her likeness, and used generative AI to clone her face and manufacture this viral circus.
Through a series of deepfakes, the young Muslim woman was depicted engaging in bizarre, physically absurd gym exercises, while shouting contradictory health advise to supposedly demonstrate how to "stay in shape”. The impact was swift and devastating. The deeply humiliating videos quietly racked up tens of millions of views across Meta’s social media ecosystem, unleashing a weaponized flood of abusive, xenophobic, and misogynistic harassment that targeted her weight and her hijab.
Both women did what any of us would do in this situation: they immediately reported the content to Meta. Yet, Meta missed both opportunities to recognize and prevent harm, and failed to remove either video.
When the young Muslim advocate tried to report the video for bullying and harassment, Meta’s automated system completely closed her initial reports and subsequent appeals without ever escalating them to a human being for review. The platform maintained the content stayed up because its internal definition of “unwanted manipulated imagery” was too narrow. Because the bad actors fabricated her conduct and actions rather than explicitly altering her physical body, the algorithms let it slide. It ruled in the perpetrator’s favor.
Dorothy McHugh faced the exact same digital brick wall. The automated systems repeatedly refused to prioritize her reports, or reach a human reviewer. As a result, her deepfake remained online for ten agonizing months, undermining her ability to fulfill her public duties as an elected official ahead of her upcoming council election. The constant, looming presence of the video—coupled with vile, violent user comments discovered later during the investigation, including one explicitly stating she "wants [to be] raped"—deeply re-traumatized the 77-year-old councillor. Exhausted by the platform’s automated apathy, Dorothy had to step back entirely while a colleague fought Meta on her behalf.
Imagine reporting that someone has stolen your face, invented your words, and circulated the result to provoke hostility and real-world harm, only to discover that a multi-billion-dollar platform does not consider your safety urgent enough for a real human being to examine.
Worse yet, you quickly realize that the avenues to fight back are not merely unhelpful, they are entirely non-functional.
When you navigate Meta’s safety interface, you enter an administrative void. There is no emergency hotline, no direct support email, and no path to reach a living person during a crisis. Instead, you are forced to deal with an adversarial maze of dead ends. You click a generic “Report” button, only to have a context-blind algorithm send you cold, automated rejection script, before closing your support ticket within seconds.
If you attempt to appeal, you’re met with automated chat tools with broken loops of unhelpful text, support dashboards that freeze, and ticketing systems that quietly close your case without explanation trapping you and any other victim in a labyrinth of automated rejections while the defamatory media continues to accumulate thousands of views and generate active engagement unchecked. It is a hollow facade of corporate compliance designed to absorb your complaints, burn you out, and leave you completely abandoned while the toxic machine continues to profit off your deepfake.
These are not isolated glitches in an otherwise functional matrix; they are part of an industrialized crisis of protracted systemic platform neglect. For over a decade, platforms’ automated systems have systematically prioritized rigid policy blocks over human nuance—a pattern that became glaringly visible following the Arab Spring and the Syrian conflict. As platforms became vital repositories for human rights monitoring, automated tools operating under blanket policy triggers routinely censored and erased critical evidence, civilian testimony, and war-crimes documentation. Yet, while the algorithms scrubbed history, they simultaneously failed to catch actual weaponized hate speech, targeted disinformation, and coordinated mass-reporting campaigns.

It is the same systemic disfunction that led to the suppression of vital documentation during Black Lives Matter protests in the U.S., where automated algorithms repeatedly flagged street journalism and anti-racism activism as graphic violence, erasing the record of real-world events, allowing harmful narratives about the movement to flourish. It mirrors the digital censorship faced by the LGBTQ+ community, whose essential mental health resources, community organizing, and educational spaces are routinely mislabeled, shadowbanned, and blocked as “adult content” by context-blind filters.
This crisis has escalated with synthetic media, and its collidision with virtually non-existent defense mechanisms. Global monitoring shows the total volume of deepfakes circulating online has exploded to over 8 million unique files—a staggering 16-fold increase in just two years. Yet, when these fakes hit the internet, Meta and its peers collapse under the weight of their own automated apathy.
According to a Reporters Without Borders (RSF) analysis of targeted deepfakes, a vast majority of malicious creators operate with absolute impunity. This lawlessness is directly fueled by a total lack of cooperation from social media platforms, which routinely leave the offending content active long after it has been flagged. The fallout from these unaddressed deepfakes is devastating, inflicting severe harms that range from financial fraud and character defamation to direct, terrifying threats to victims’ physical safety.
Crucially, this is not a gender-neutral crisis; this is an engine of gender-based violence. Women accounted for a disproportionate 74 percent of the cases studied by RSF, proving that bad actors have successfully turned generative AI into a specialized tool for silencing and terrorizing women online.
It is the exact same dynamic that played out across the two deepfake cases with the Chancellor and Health Advocate in the UK and Europe.
Eventually, on September 17, Meta’s Oversight Board stepped in, forcing Meta to take down the videos of the women. But the significance of their investigation and rulings extends beyond these two bad moderation decisions and failures. It confirmed that the platform’s reporting channels are structurally broken, relying on automated triage scripts that shift the burden (including the burden of proof) onto the victim. It is a system built not to resolve harm, but to automate corporate indifference.
Furthermore, it reveals that Meta and other platforms remain trapped in outdated, checklist-driven frameworks that treat synthetic media merely as an issue of transparency: whether a piece of content fake. But they are missing the bigger picture.
By reducing complex human threats to simple binary equations—asking only “is this AI-generated?” rather than “how is this being weaponized?”—they remain fundamentally blind to localized intent. And they act like putting a label on an AI-generated video solves the problem. But it’s not just about labeling, this is about power.
The real problem isn’t just asking “is this AI-generated?” The real questions are: Who is being targeted? What is the deeper intent? How does the content spread, and how does it tap into real-world inequalities to yield real world consequences?
These videos were never merely lies. They hijacked and appropriated women’s identities and weaponized them as delivery systems to pump out political hate, humiliation, and threats and intimidation. In both cases, the synthetic media was not the harm’s endpoint. Deepfakes like these don’t just invent new problems—they supercharge existing ones, making the attacks more real, easily distributable, and harder to stop. And it leaves the target with nowhere to escape.
Yet, the tech industry’s current defense strategy completely ignores this human toll, opting instead for superficial fixes.

Meta’s fix: Adding warning labels and AI disclosure tags for transparency.
The Reality: A label merely provides technical metadata—it tells the viewer if and how a piece of content was created or altered. It does absolutely nothing to disrupt the behavioral violence of the attack itself.
Much of the tech industry’s emerging response to synthetic media relies on provenance, detection, and disclosure. Think of provenance as a digital watermark baked into a file when it is made, tracking its birth certificate and history to prove how it was created. Detection acts like a security scanner, using automated software to hunt for hidden AI pattern glitches in uploaded content. Finally, disclosure is the warning label slapped on the post, shifting the burden onto everyday viewers to figure out what is real and what is fake.
Sure, it is useful to let people know when a photograph, recording or video has been altered. But provenance and disclosure alone, as an approach, is based on a dangerously narrow theory of harm.
Their thinking goes: If you tell the viewer, “Hey, this is fake,” then you’ve solved the problem; This framework assumes the main threat is simply a confused viewer mistaking a fabrication for reality. It implies that once a warning label pops up to fix the misunderstanding, the system has supposedly done its job, and the platform’s responsibility ends. The viewer can now make a better judgment. But in reality, it leaves the victim completely unprotected from the ongoing attack.
It might address a fleeting moment of deception, but it completely fails to stop networked abuse. A labeled deepfake can still reach millions of people. It can provoke racist and misogynistic harassment, contaminate search results, jump between platforms and circulate in private groups where the original warning label and tag disappears entirely. Copies can survive long after the target has issued a denial, and short clips can be stripped of context and repackaged as viral memes, reaction videos or alleged “evidence.”
Virality does not wait for verification.
Furthermore, a passive label cannot restore a victim’s control over their own face or voice, undo the emotional humiliation, expose the perpetrators coordinating the campaign, or stop algorithmic recommendation systems from extending its reach. Crucially, it fails to neutralize the implicit threat delivered to everyone watching: speak publicly, and this will happen to you too.
This creates a devastating chilling effect, forcing women and marginalized communities to think twice before appearing online simply because their identities are being weaponized against them. Frequent targets of online abuse already carry a disproportionate cost for participating in public debate, and synthetic media dramatically raises that cost by allowing harassers to manufacture incriminating material faster than any target or fact-checker can possibly rebut it.
Therefore, the baseline technical question cannot simply be: Is this AI-generated? Platforms must force their systems to address structural product questions:
Was this identity used without consent?
Is the manipulation exploiting and inciting hatred against a vulnerable community?
Is the content being amplified by coordinated networks?
Has the content created a credible, irreversible safety risk for the target—and others?
Did the platform’s own recommendation algorithm, and lack of intervention, accelerate its spread?
Can the victim immediately reach a qualified human reviewer before the damage spreads, in a timely and effective way?
These are not media-literacy questions about whether users know how to spot a deepfake. They are fundamental questions about product design, corporate responsibility, and platform power.
Meta’s Fix: Meta claims to safeguard political expression and legitimate criticism of powerful officials through their public figure moderation exemption.
The Reality: Meta weaponizes the public figure rule to sidestep corporate responsibility, transforming a legacy free-speech protection into an active, dangerous vector for identity theft and AI abuse.
Meta initially concluded that the deepfake video of a Scottish local councillor didn't violate its bullying and harassment rules simply because she is a politician, and therefore categorized her as a public figure under the exemption. This policy distinction originally made sense when platforms wanted to separate legitimate criticism of powerful officials from personal bullying. But in the age of generative AI, it functions as a glaring loophole, effectively creating a permission structure for wholesale impersonation and AI abuse.
Politicians, journalists, and human rights defenders are still human beings. They do not forfeit the right to control their own bodies, faces, and voices the moment they enter the public eye. Completely fabricating someone's words out of thin air isn't criticism—it is an assault.
This loophole becomes downright lethal in conflict zones and crisis environments, where the consequences of synthetic media collapse instantly from reputational harm into physical danger. Imagine a synthetic video depicting a journalist praising an armed group, an human-rights defender confessing to foreign espionage, or a humanitarian worker making a fabricated sectarian statement. This content is routinely weaponized to undermine independent reporting, mobilize digital lynch mobs, justify arbitrary state detention, or poison the credibility of authentic documentation. In many environments, it can also result in direct physical harm and even death.
A weaponized deepfake does not need to convince a global audience or survive a month-long fact-check or investigation. It only needs to deceive a single armed checkpoint or reach a hostile actor before the correction catches up. By the time a tech company in Silicon Valley figures out what happened, debates whether to slap a passive label on it, the real-world physical damage is already done.
For targets operating outside the Western hemisphere, the danger is multiplied exponentially. If you work in a language or culture Meta doesn’t prioritize, reports are routinely missed, misclassified, or ignored. For those in Arabic-language environments, this risk is deeply compounded by platforms’ systemic failures in dialect recognition, context-aware moderation, and access to rapid human escalation channels. A global policy written in English offers zero protection when an Arabic-language report is processed by automated filters that completely lack the political and cultural literacy required to recognize a threat to someone’s life. The corporate policy may exist on paper, but the actual protection does not.

Meta’s Fix: Meta evaluates content through isolated checklists—treating hate speech, bullying, and privacy issues as independent technical violations—to process millions of posts objectively and at a massive scale.
The Reality: Platforms treat interconnected harms as separate violations of their policies, which abusers easily exploit by combining multiple tactics into a single assault that slips through the platform’s structural blind spots.
Social media networks love corporate organization. If a piece of content is uploaded, their systems evaluate it through isolated lenses: hate speech lives in one department, bullying in another, privacy violations down the hall, and misinformation across the street.
But real-world abusers do not respect corporate organizational charts or policy thresholds.
A malicious deepfake is a multi-pronged weapon of convergence. It doesn’t just do one thing; it simultaneously combines identity theft, gendered harassment, political disinformation, and coordinated bot amplification. It permeates different ideologies, weaponizes historical social fractures, and crosses digital borders to hit a target from every angle at once.
Yet, because Meta assesses each rule in absolute isolation, its automated systems frequently conclude that no single rule has been breached severely enough to justify a takedown. Nothing ever seems bad enough on its own to trigger action. Abusers understand this structural blind spot perfectly, exploiting it to orchestrate attacks that fly entirely under the safety radar.
The Oversight Board, at least, recognized this systemic failure. By evaluating the content within its real-world context, they saw that the deepfake targeting the politician was not merely inaccurate political commentary—it was a coordinated strike that hijacked her likeness to incite hatred against refugees. Similarly, the video targeting the Muslim activist was not just an unwanted digital manipulation; it was a compounded attack drawing power from the toxic intersection of sexism and Islamophobia.
Perhaps the most disturbing fact in these cases is not that Meta’s automated systems missed the videos. Automated systems will always make mistakes. It is that the reporting and review infrastructure did not reliably correct those mistakes. This is where platform accountability becomes measurable.
To counter this, platforms must adopt a cumulative-harm standard for synthetic media. When digital manipulation intersects with identity-based abuse, structural intimidation, or coordinated harassment, safety responses must immediately escalate rather than vanish between corporate policy silos. The burden of proof cannot remain on the victims whose identities have been stolen.
A meaningful remedy system would give people whose likenesses have been manipulated a dedicated reporting route, rapid human review, preservation of relevant data and a way to identify duplicates and coordinated reposts. It would address reach as well as removal: limiting recommendations, preventing monetization and reducing recirculation while a credible impersonation report is investigated.
It would also preserve evidence of the campaign. Removal is necessary, but deletion without documentation can erase information needed by researchers, courts, journalists or the target herself to establish patterns of abuse.
For human-rights defenders and organizations documenting conflict, the preservation question is critical. Platforms must be capable of restricting harmful circulation while retaining authenticated records, associated metadata and enforcement histories through rights-respecting processes. Safety and accountability should not be forced into a false choice between leaving abusive material online and making the evidence disappear.
In summary
The Oversight Board’s recommendations point in the right direction: stronger labeling, warning screens, and distribution limits for high-risk AI content, alongside broader protection against unwanted manipulated imagery. But Meta must go further.
It should immediately implement a six-point upgrade to its architecture:
1. Expedite impersonation channels. Create an expedited reporting route and impersonation channel accessible to public figures, journalists, activists and ordinary users alike whose likeness has been stolen.
2. Institute cumulative evaluation. Evaluate synthetic content cumulatively across the intersections of harassment, hate, privacy, authenticity and coordinated-harm policies.
3. Automate networked takedowns. Proactively detect and act on duplicate and near-duplicate versions across Facebook, Instagram and Threads simultaneously.
4. Audit global performance disparities. Audit safety response quality by language, dialect, gender and region—and publicly publish the disparities.
5. Establish secure data archives. Preserve removed synthetic media and relevant metadata when it may document coordinated abuse or other rights violations. The deepfake can be used as evidence itself, so treat it as such when collecting and archiving the content and its data.
6. Provide target transparency reports. Provide victims with a clear record of what action was taken, what viral or targeted reach the content obtained and whether copies remain active on the network.
Most importantly, Meta must stop treating virality as something that happens around harmful content rather than something its own systems actively produce and incentivize.
Deepfakes are governance tests. The next generation of synthetic-media policy will fail if it focuses only on whether a piece of content carries an “AI-generated” label.
Deepfakes are becoming highly efficient instruments for redistributing power: taking control of a person’s face, voice, and credibility and handing that control over to anonymous strangers, political actors, or hostile networks. The people most exposed will always be those already forced to fight for equal visibility, safety and credibility online—women, Muslims, migrants, racialized communities, journalists, and human rights defenders.
The ultimate measure of a platform’s response is therefore not whether its software can identify an artificial file. It is whether it can recognize, intercept, and dismantle the real system of harm built around it.
Meta did not initially do that in either of these cases. The Oversight Board corrected the individual decisions. Now the question is whether Meta will correct the broken architecture that produced them. A digital label may tell us that a video is fake. It cannot, by itself, stop the power power behind the lie.




