Is Scraping Public Social Media Data Legal for AI Training?
Introduction
Businesses developing artificial intelligence systems often collect publicly visible social media posts, images, profiles, comments, and other user-generated content for data analysis and machine learning. The fact that information can be viewed online, however, does not automatically mean that it may be copied, aggregated, sold, or used to train commercial AI models without legal restrictions.
Under Philippine law, the legality of scraping depends on the nature of the information, the purpose and manner of collection, the applicable lawful basis, the safeguards implemented, and whether the processing respects the rights and reasonable expectations of data subjects. The principal legal regime is the Data Privacy Act of 2012, supplemented by its implementing rules and National Privacy Commission issuances.
Does Public Availability Mean Consent?
No. Public availability is not the same as consent to unrestricted processing. The National Privacy Commission has expressly stated that publicly available personal data remain protected by the Data Privacy Act, its implementing rules, and NPC issuances ([NPC Advisory No. 2026-01 (2026)](#I1.0); [NPC Advisory No. 2024-04 (2024)](#I6.5)).
Accordingly, a social media user’s name, photograph, username, location, contact information, opinions, employment details, or other identifying information may remain personal information even when displayed on a public account. The Data Privacy Act applies to the processing of personal information by natural and juridical persons, including private entities and government institutions ([Data Privacy Act of 2012](#L1.30); [KAPIT v. City of Manila, et al., G.R. Nos. 261892, 262192, 263752, 2026](#J1.477)).
Public access may be relevant to determining the circumstances of collection and the reasonable expectations of users, but it does not by itself establish a lawful basis for commercial scraping or AI model training.
What Law Governs Social Media Scraping?
The principal statute is Republic Act No. 10173, or the Data Privacy Act of 2012. Its implementing rules require processing to comply with the principles of transparency, legitimate purpose, and proportionality ([IRR of Republic Act No. 10173](#L2.36)).
These principles apply whether the data is obtained directly from the individual, collected through an online platform, aggregated from multiple public sources, or extracted through automated tools. They also apply to the development, training, testing, and deployment of AI systems ([NPC Advisory No. 2024-04 (2024)](#I6.5)).
The NPC’s recent guidance specifically addresses organizations that scrape publicly available personal data and personal information controllers that host data capable of being scraped ([NPC Advisory No. 2026-01 (2026)](#I1.1)).
What Counts as Processing?
Processing is not limited to viewing information. It may include collecting, recording, organizing, storing, altering, retrieving, consulting, using, combining, disclosing, or deleting personal information.
Scraping may therefore constitute processing when a company uses automated tools to copy social media information, places it in a database, links it with other datasets, uses it to train or test a machine learning model, or makes the resulting dataset available to another entity.
The fact that a scraping operation is automated does not remove it from the Data Privacy Act. Nor does the fact that the resulting AI model does not visibly reproduce the original posts necessarily eliminate privacy concerns. The legality of the activity must be assessed by examining the full data lifecycle.
What Lawful Basis May Support AI Training?
Consent is only one possible lawful basis under the Data Privacy Act. Depending on the circumstances, processing may also be based on a legal obligation, a statutory or public-authority function, protection of lawful rights, or legitimate interests, subject to the statutory requirements and the rights of data subjects.
For ordinary personal information, legitimate interest may be considered under Section 12(f), but it is not automatic. The processing must not be otherwise prohibited by law, must serve a legitimate interest, and must not override the fundamental rights and freedoms of the data subject ([NPC 24-006 (2025)](#I3.30); [Eastwest Rural Bank v. Philippine National Police Anti-Cybercrime Group, et al., G.R. No. 273720, 2025](#J6.28)).
For sensitive personal information, the permitted grounds are narrower. Section 13(f) may allow processing necessary for the protection of lawful rights and interests in court proceedings, the establishment, exercise, or defense of legal claims, or when information is provided to a government or public authority under a constitutional or statutory mandate. That provision is generally not a blanket authorization for commercial AI training.
The Supreme Court recognized in Azarraga v. Jalbuna that processing information for the protection of lawful rights in litigation may be authorized under Section 13(f), provided that the processing is necessary and lawful. The Court also emphasized that the principles of transparency, legitimate purpose, and proportionality continue to govern the activity ([Azarraga v. Jalbuna, A.C. No. 13678, 2023](#J3.29)).
When May Legitimate Interest Apply?
A company relying on legitimate interest should be able to identify, document, and defend at least three matters:
- The legitimate interest: such as developing a service, improving cybersecurity, conducting research, or training a model for a defined business purpose.
- The necessity of the processing: whether scraping identifiable personal data is reasonably required, or whether less intrusive and anonymized data would suffice.
- The balancing of interests: whether the company’s interest outweighs the privacy rights, reasonable expectations, and potential risks to affected individuals.
Legitimate interest cannot be used to defeat another law that restricts access, use, reproduction, or disclosure. The NPC has stated that the lawful basis under Section 12(f) is not absolute and cannot authorize processing that is prohibited by another law ([NPC 24-006 (2025)]).
What Do Transparency and Notice Require?
Transparency requires the organization to explain, in a clear and accessible manner, what information is being collected, where it comes from, why it is being processed, how long it will be retained, with whom it will be shared, and what rights data subjects may exercise.
For large-scale scraping, individualized notice may be difficult, but difficulty does not automatically remove the transparency obligation. A company should consider layered notices, public privacy statements, platform-specific disclosures, accessible explanations of the AI training purpose, and mechanisms through which data subjects may raise objections or request applicable forms of relief.
The NPC’s data-scraping guidance states that publicly available information remains subject to privacy safeguards and that processing must have a lawful basis, a declared purpose, and appropriate notice where required ([NPC Advisory No. 2026-01 (2026)](#I1.0); [NPC Advisory No. 2026-01 (2026)](#I1.11)).
Why Purpose Limitation Matters
Information posted for one purpose should not automatically be repurposed for an unrelated commercial activity. A user may publish a comment to participate in a discussion, upload a photograph for friends, or maintain a professional profile for recruitment. Those circumstances do not necessarily imply agreement to have the information copied into a commercial AI training dataset.
Under the purpose-limitation principle, the organization should define the intended AI use before collection. If the company later intends to use the scraped data for a substantially different purpose, it should reassess the lawful basis, provide appropriate notice, conduct a privacy impact assessment, and comply with other requirements under the NPC guidance ([NPC Advisory No. 2026-01 (2026)](#I1.11)).
How Does Proportionality Affect Scraping?
Proportionality requires the processing to be adequate, relevant, and not excessive in relation to the declared purpose. A company training a language model may not need users’ telephone numbers, precise locations, private messages, government identifiers, or sensitive health information.
Proportionate controls may include:
- collecting only fields reasonably needed for the defined model purpose;
- excluding sensitive personal information and confidential communications;
- removing direct identifiers before training;
- using filtering tools to exclude minors’ data, credentials, and high-risk content;
- setting retention limits for raw scraped data; and
- restricting access to personnel and service providers with a legitimate need.
The greater the scale, sensitivity, identifiability, and commercial value of the dataset, the stronger the justification and safeguards should be.
Are Sensitive Data and Children’s Data Higher Risk?
Yes. Sensitive personal information includes data relating to a person’s race, ethnic origin, marital status, age, health, education, criminal record, government-issued information, and other categories identified by the Data Privacy Act. Scraping such information for AI training requires a specific and defensible legal basis under Section 13.
Organizations should also take particular care with information relating to children, location patterns, biometric identifiers, financial information, health information, political opinions, religious beliefs, and data that could facilitate identity fraud, harassment, profiling, or targeted cyberattacks.
The NPC’s data-scraping guidance identifies risks such as doxing, malicious disclosure, unauthorized profiling, identity fraud, surveillance, and collection of login credentials as prohibited or particularly harmful uses ([NPC Advisory No. 2026-01 (2026)](#I1.11)).
What Practices May Create Legal Exposure?
Commercial AI developers face heightened risk when they scrape data at scale without a documented lawful basis, collect information unrelated to the stated purpose, retain raw datasets indefinitely, ignore platform restrictions, process sensitive information indiscriminately, or disclose the dataset to third parties without appropriate safeguards.
Risk is also increased where the organization collects information that is technically public but reasonably expected to remain within a limited social or professional context. Public visibility is relevant, but it is not an unrestricted license for mass extraction and repurposing.
Unauthorized processing of personal information may carry criminal penalties under Section 25 of the Data Privacy Act. Unauthorized processing of personal information may be punished by imprisonment of one to three years and a fine of ₱500,000 to ₱2,000,000. Unauthorized processing of sensitive personal information may be punished by imprisonment of three to six years and a fine of ₱500,000 to ₱4,000,000 ([Data Privacy Act of 2012]).
Does Anonymization Remove Privacy Obligations?
Not necessarily. Properly anonymized information that can no longer be linked to an identified or identifiable individual may fall outside the concept of personal information. However, pseudonymization, hashing, tokenization, or removal of a username may not be enough if the person can still be reidentified by combining the dataset with other information.
Organizations should test whether reidentification is reasonably possible, considering the size of the dataset, available external information, model outputs, unique personal characteristics, and access by third parties. If reidentification remains reasonably possible, the data should continue to be handled as personal information.
What Compliance Process Should Companies Follow?
A responsible AI developer should complete the following steps before undertaking large-scale social media scraping:
- Define the use case. State the model’s purpose, expected users, commercial activity, and intended outputs.
- Map the data. Identify the types of content, data subjects, platforms, sources, jurisdictions, and recipients involved.
- Classify the information. Separate personal information, sensitive personal information, privileged information, confidential communications, and non-personal content.
- Select and document the lawful basis. Do not assume that public availability or business necessity is sufficient.
- Conduct a privacy impact assessment. Assess risks involving scale, reidentification, profiling, discrimination, children, sensitive data, and model memorization.
- Apply minimization and filtering. Exclude unnecessary identifiers, credentials, private communications, and high-risk categories.
- Prepare transparency measures. Publish clear notices and provide appropriate channels for data-subject requests and objections.
- Control vendors and recipients. Use written agreements, security measures, access controls, and deletion obligations.
- Test the model. Check for memorization, reproduction of personal information, harmful inference, and unauthorized disclosure through prompts or outputs.
- Maintain records. Keep documentation showing the purpose, legal basis, assessment, safeguards, retention period, and decision-making process.
What Should Data Subjects and Businesses Expect?
Data subjects may question how their information was collected, request clarification regarding its use, and exercise rights recognized under the Data Privacy Act, subject to statutory limitations and the circumstances of the processing.
Businesses should not wait for a complaint before examining their scraping activities. A dataset may be lawful at the time of collection but become problematic when reused for a new purpose, combined with other datasets, transferred to another entity, or incorporated into a model that reproduces identifiable information.
The Supreme Court has recognized that processing personal information from official records may be lawful when performed under a statutory mandate or for a legally protected purpose, but the processing must still comply with the Data Privacy Act and applicable limitations ([Zoleta v. Investigating Staff, et al., G.R. No. 258888, 2024](#J5.34)).
Conclusion
Scraping public social media data for commercial AI training is not automatically illegal, but neither is it automatically lawful. The decisive considerations are the type of data collected, the purpose and necessity of the processing, the lawful basis, the transparency provided, the proportionality of the collection, and the safeguards used throughout the data lifecycle.
Companies should treat public social media data as potentially protected personal information. Before collecting it at scale, they should conduct a documented privacy assessment, exclude unnecessary and sensitive data, establish a defensible lawful basis, provide appropriate notice, and implement controls against reidentification, profiling, harmful disclosure, and unauthorized reuse.
About Nicolas and De Vega Law Offices
Nicolas and de Vega Law Offices is a full-service law firm in the Philippines. You may visit us at the 16th Flr., Suite 1607 AIC Burgundy Empire Tower, ADB Ave., Ortigas Center, 1605 Pasig City, Metro Manila, Philippines. You may also call us at +632 84706126, +632 84706130, +632 84016392 or e-mail us at [email protected]. Visit our website https://ndvlaw.com.

