10NEWS
World

Navigating the AI Dilemma: South African News Publishers Face Content Scraping Challenges

By Editor • September 1, 2026 • 2 min read

South African news publishers are grappling with a significant dilemma as a recent study reveals that less than one-third of local news websites actively block AI crawlers from accessing their content. This issue is highlighted in a new report titled "The Protocol Gap: South Africa," which was launched by the Journalism Relay Project and partners on Tuesday.

The study analyzed the robots.txt files of 263 news websites to determine how they manage AI crawlers—automated programs used by tech giants like OpenAI and Google to harvest online content. While 74.1% of the sites had a robots.txt file, only 30.4% explicitly denied access to at least one AI crawler, a figure stagnant since December 2025.

The report underscores a troubling reality: larger, well-funded publishers have the resources to block AI access, while smaller, independent, and community-focused outlets are often left vulnerable. As publishers weigh the trade-offs between visibility and control, the pressure to allow AI access for online exposure is palpable.

Robots.txt files serve as a basic mechanism for websites to signal their preferences regarding bot access, yet they lack legal enforceability. According to MLTT director Michael Markovitz, this tool functions more as a public declaration than a strict barrier, exposing all South African publishers to AI scraping.

Despite its limitations, the robots.txt mechanism can at least record a publisher's refusal of AI access, which could be useful in legal contexts. The report noted instances where such directives have been cited in legal complaints in various countries.

One of the most challenging aspects of this decision-making process is the relationship with Google. Publishers can instruct Google's crawlers not to use their content for specific AI functions without impacting their visibility in standard search results. However, the rise of Google’s AI-generated summaries complicates matters, potentially forcing publishers to restrict access to their pages entirely if they wish to prevent their work from being used in these AI Overviews.

This situation is particularly dire for publishers reliant on search traffic, as many may find themselves in a precarious position, forced to keep their doors open to AI crawlers despite their misgivings. MLTT program manager Ompha Tshamano highlighted that this dilemma reflects a structural issue in the market rather than merely a technical one.

The report also noted that leading South African news sites experienced a significant drop in daily page views—around 20%—from May 2025 to May 2026. While it doesn’t directly link AI to these losses, the data suggests that AI platforms currently deliver minimal traffic back to publishers. Between January and May 2025, popular AI chatbots generated less than 1% of visits to South African news sites.

As the landscape evolves, the need for industry solidarity grows. IFPIM’s Irene Jay Liu argued that media organizations must collectively advocate for specific remedies to combat these challenges effectively. The conversation surrounding these issues is critical as publishers strive to protect their work in an era increasingly dominated by AI.

Source: www.dailymaverick.co.za

#AI #crawlers #journalism #media #South Africa

Similar posts