StoriesIO goes beyond standard headlines to deliver deep dives into IoT, Android, smartphones, browsers, and the technologies shaping how we connect.
Coverage That Goes Deeper
Four core areas where StoriesIO spends extra time cutting through the noise.
News
Daily headlines, launches, and industry shifts across the technology sector.
Devices
Smartphone and gadget coverage including Honor, Vivo, Nokia, and modular concepts.
Reviews
Long-form reviews of phones, smartwatches, and apps with honest verdicts.
IoT
How connected devices, browsers, and wearables are reshaping everyday tech.
Web scraping for AI training and the law
Web scraping has become a central way for AI companies to collect text, images, code, product data and public conversations. A crawler can gather billions of pages quickly, then turn that material into datasets used to train large language models, image generators and recommendation systems. The technical process may look routine, yet the legal position depends on what was copied, where it came from, how it was used and whether people’s information was involved. Learn more about How To Run Linux Apps.
For Australian publishers, developers and startups, the issue sits at the intersection of copyright, privacy, contract law, computer access rules and consumer protection. A page being visible without a login does not automatically make it free to copy for commercial machine learning. Equally, an automated request is not automatically unlawful. The facts and the jurisdiction do much of the work. Learn more about Android 16 S First Developer.
Why public access does not settle the question
A web page can be publicly readable while still carrying several legal interests. Copyright may protect the wording of an article, the selection and arrangement of material, photographs, illustrations, software and some original datasets. A crawler that stores a page is making a reproduction, even if the copy exists briefly in a cache or is later transformed into numerical model weights.
Australia’s Copyright Act 1968 does not contain a broad US-style fair use defence. It has narrower fair dealing exceptions, including research or study, criticism or review, reporting news, parody or satire and access for people with a disability. Commercial AI training will not automatically qualify simply because the resulting model is used for research, or because the source material was publicly available.
That distinction matters to a technology site covering smartphones, browsers or wearables. Copying a short factual specification may raise fewer copyright concerns than ingesting an entire review, image gallery and carefully structured buying guide. Facts themselves generally receive limited protection, while the original expression and editorial selection around those facts can be protected.
Australian law also lacks the European Union’s separate sui generis database right. This does not create a free pass for bulk extraction. Copyright, confidentiality, contract terms, passing off, privacy and computer misuse rules may still apply, particularly where a company has invested heavily in curating and maintaining the dataset.
Copyright, licences and training copies
The main copyright question is often framed too narrowly: “Did the AI reproduce the article in its answer?” Training can involve several stages before an output appears. A service may download a work, store it, clean it, break it into tokens, create intermediate copies and use those representations to adjust model parameters. Each step may need a legal basis, although the analysis depends on the relevant statute and the technical details.
Licensing is the clearest route when a business wants predictable rights. A licence can define which pages may be collected, whether images and code are included, permitted model types, retention periods, attribution, security controls, audit rights and whether the model may generate competing content. Some publishers may prefer a collective licensing arrangement, while others may reserve content for subscriptions or direct partnerships.
A useful distinction is between training data and retrieval. A retrieval-augmented system may fetch an article at answer time and quote or summarise it under a separate permission model. That does not eliminate copyright issues, but it can make source tracking, takedowns and attribution easier than embedding every article into a general-purpose model.
Technology publishers should catalogue what they own before negotiating. For example, a site’s devices coverage may combine staff writing, manufacturer press images, screenshots, agency photographs and specifications supplied by third parties. A licence covering editorial text may not cover every image or quoted dataset on the same page.
Terms of use and the meaning of unauthorised access
Website terms can prohibit automated collection, commercial reuse or model training. If a crawler continues after a clear prohibition, the operator may argue breach of contract, especially where the user accepted terms through an account, API agreement or paid service. The strength of that argument varies when a visitor simply reads an open page without actively agreeing to anything.
Robots.txt is important operationally, but it is not generally a statute. It is a machine-readable signal about crawling preferences. Ignoring it may support an argument that the crawler acted contrary to the publisher’s stated wishes, yet robots.txt alone does not create a universal private right to payment or automatically turn copying into a criminal offence.
The line between scraping and hacking matters. Circumventing a login, bypassing a paywall, defeating rate limits, exploiting an unpatched endpoint or using stolen credentials creates much greater risk than downloading openly served pages at a reasonable rate. Australian computer offence provisions can apply to unauthorised access or modification, while overseas systems may trigger laws in the country where the server or company is located.
Courts overseas have produced mixed signals. The United States decision in hiQ Labs v LinkedIn treated some public-profile scraping claims differently from access behind authentication, but it did not resolve copyright or privacy questions globally. A ruling about access under US computer law should not be treated as a general licence for an Australian startup to harvest any public website.
Privacy turns public data into personal data
Personal information is a separate issue from copyright. A name, email address, phone number, location history, profile handle, photograph or combination of clues may identify an individual. Public availability does not remove privacy obligations. A dataset assembled from many harmless-looking pages can become more sensitive when linked, enriched or used to infer age, health, politics, employment or behaviour.
The Australian Privacy Act 1988 and the Australian Privacy Principles are central for organisations covered by them, although coverage depends on factors such as turnover, sector and statutory exemptions. The Office of the Australian Information Commissioner has repeatedly emphasised that privacy risks must be considered when personal information is used with generative AI. Collection should be reasonably necessary, transparent and handled with appropriate security and retention controls.
For an Australian service, collecting local posts about a person in Melbourne, Brisbane or Perth may still involve personal information even when those posts are publicly viewable. A dataset that captures children’s comments, Indigenous community information, precise locations or inferred health conditions presents heightened risks. Overseas transfers and cloud providers add questions about disclosure, governance and breach response.
De-identification is not a magic eraser. Removing names may leave unique writing styles, usernames, photographs or location patterns that allow re-identification. Companies should document the source, purpose, legal basis, fields collected, deletion process and access controls. They should also provide a channel for correction or removal where appropriate, without promising that every model can instantly forget a learned pattern.
Regulation, lawsuits and cross-border exposure
Legal risk often arrives through several channels at once. A publisher may send a cease-and-desist letter, a photographer may claim image infringement, a person may complain to the privacy regulator and a platform may suspend the crawler’s account. The company collecting the material may be based in Sydney while its servers, customers and affected creators are spread across California, London and Singapore.
Australia’s Competition and Consumer Act 2010 can become relevant if a business makes misleading claims about training data, licences, privacy protections or the originality of generated material. A claim that an AI system was trained only on permissioned content needs evidence. Marketing language such as “fully compliant” can create its own exposure if the underlying dataset includes unlicensed or improperly collected material.
The European Union adds another layer. The EU Copyright Directive includes text-and-data-mining rules, with different treatment for research organisations and commercial users and an ability for rights holders to reserve rights in machine-readable ways. The EU AI Act also creates transparency and governance duties for certain general-purpose AI models. A company serving European users may need controls that exceed the minimum position in Australia.
The safest legal assessment therefore asks precise questions rather than relying on slogans. Was the source open or restricted? Was there an agreement? Was the copy substantial? Were personal details collected? Did the crawler bypass a technical barrier? Is the model commercial? Can the operator identify and remove affected material? These questions are more useful than simply labelling the activity “public web scraping”.
A responsible approach for publishers and AI builders
Publishers can make their position clearer through layered controls. Terms of use should address automated collection and machine learning in plain language. Robots.txt can communicate crawler preferences, while an AI-specific policy or licensing page can identify acceptable uses. APIs with authentication, rate limits and usage logs provide more control than leaving a high-value feed exposed through ordinary HTML.
Technical measures should match the stated policy. Blocking every bot can interfere with search engines, accessibility tools and legitimate research. Conversely, an organisation that says it prohibits AI collection but leaves unrestricted bulk endpoints available may still have legal rights, though its evidence about notice and enforcement could be weaker. Logging requests, preserving notices and responding consistently can help establish a defensible record.
AI companies should perform source due diligence before ingestion. They can prioritise licensed corpora, public-domain material, opt-in contributions and datasets with clear provenance. A crawl should respect authentication barriers, rate limits and deletion requests. Filters can reduce the collection of emails, health information and children’s data, while memorisation testing can identify whether the model is reproducing long passages or confidential records.
The same discipline applies to smaller Australian startups operating from an incubator in Sydney or a shared office in Adelaide. They should identify who owns each dataset, where data is stored, which jurisdictions apply and who handles complaints. A short legal review before a large crawl can be cheaper than rebuilding a model, notifying thousands of people or defending an urgent court application.
| Issue | Questions to ask | Practical control |
|---|---|---|
| Copyright | What expression, images, code or curated selection is being copied? | Obtain a licence, use public-domain material or narrow the dataset |
| Contract | Did the operator accept terms, an API agreement or a paid account condition? | Review terms, preserve consent records and obey access limits |
| Computer access | Was a login, paywall, rate limit or technical barrier bypassed? | Never circumvent controls; use authorised APIs |
| Privacy | Does the dataset identify people or reveal sensitive traits? | Minimise fields, screen data, document purpose and support deletion |
| Cross-border rules | Will Australian data be stored or used overseas? | Map transfers and apply the stricter relevant safeguards |
| Model behaviour | Can the system reproduce long passages or personal records? | Run memorisation tests, add filters and maintain incident procedures |
Clear provenance is becoming a competitive advantage. A model trained on documented, permissioned material may be easier to sell to government agencies, schools and large Australian businesses than one built from an opaque crawl. Publishers can also treat licensing as a new digital rights channel rather than relying only on blocking.
For readers and creators, the practical point is simple: the legal issues around web scraping for AI training are rarely decided by a single setting or a single court case. Public visibility, copyright ownership, privacy expectations, contractual notice and technical conduct interact. Careful collection, honest disclosure and permission where needed offer a more durable path than assuming that anything reachable by a browser is free for an AI dataset.
From The StoriesIO Archive
A look at devices, platforms, and experiments covered across recent reporting.