Scraping SOP: Ethics + Reliability
A repeatable 6-step procedure for running scraping jobs that stay inside platform rules, produce clean output, and do not put your accounts at risk.
Before You Start
45-90 min setup: output, prerequisites, and risk
Output: A repeatable scraping workflow that collects structured lead data, deduplicates against existing records, and pushes a clean export to your sequencer or CRM without triggering platform restrictions.
Time: 45-90 minutes for initial setup. Under 15 minutes per repeat run once configured and tested.
Active account in PhantomBuster, Apify, TexAu, or Captain Data. Defined target source (LinkedIn search URL, Sales Navigator list, Google Maps query, or website URL). CRM or sequencer ready to receive output. ToS status confirmed for the target platform.
LinkedIn's ToS explicitly prohibits scraping and unauthorized automation. Cookie-based scrapers without rate limiting risk permanent account restriction. This SOP reduces that risk but cannot eliminate it: read every ethics gate before running a LinkedIn job.
Workflow Overview
6-step scraping SOP: scope to clean export
| Step | Action | Output | Ethics gate? |
|---|---|---|---|
| 1 | Define scope, source, and legal boundary | Approved source list with ToS status confirmed | Yes |
| 2 | Configure tool, execution limits, and delay settings | Scraper config file with rate limits set | Yes |
| 3 | Set up proxy or cloud execution | Active proxy or cloud session connected to tool | No |
| 4 | Run test batch (50 rows max) | Sample output reviewed for field accuracy and completeness | No |
| 5 | Deduplicate against existing database | Net-new records only, ready for enrichment | No |
| 6 | Push clean output to CRM or sequencer | Verified, field-mapped records in downstream tool | No |
Step by Step
6 steps: source check to clean push
- Define your source and confirm its ToS status
Check the platform's ToS and robots.txt before opening any scraping tool. LinkedIn explicitly prohibits scraping and risks permanent account suspension. Public sites and business directories carry lower but variable risk.
- Set execution limits and per-action delay intervals
Override default speed settings: 3-8 seconds per action on LinkedIn, 1-3 seconds on other sites. Cap LinkedIn runs at 100-200 records per day as a conservative ceiling.
- Configure proxy or cloud execution
Use residential proxies for social network sources: datacenter IPs are trivially detected by LinkedIn and most anti-bot systems. Apify includes built-in residential proxy infrastructure; PhantomBuster uses your browser cookie session.
- Run a test batch of 50 rows and review every output field
Set the run limit to 50 records and verify every required field populates correctly before scaling. If a key field is absent in over 20% of rows, fix the scraper config before proceeding.
- Deduplicate against your existing database before full export
Match on LinkedIn URL or email as the deduplication key, not name or company alone. Most scraping tools don't deduplicate automatically: use Clay, a VLOOKUP, or a CRM rule to filter before any downstream push.
- Map output fields to your CRM or sequencer schema before pushing
Create a field mapping document before the first push: scraper column on the left, CRM field on the right. Mismatched fields create data hygiene problems that compound across every future campaign run.
Scraping a public page does not create a lawful basis for marketing EU-based contacts. Use legitimate interest as the lawful basis for B2B outreach and verify compliance for your stack: Apify is SOC 2 + GDPR-certified; TexAu is GDPR-compliant; PhantomBuster is self-declared.
When Things Break
4 failure modes and the fix for each
Tool Fit
PhantomBuster, Apify, TexAu, Captain Data: choose by use case




Does the platform's ToS permit automated collection? Are speed and daily caps below detection threshold? Does a 50-row test batch confirm all required fields populate? If any answer is no: stop.
FAQ
5 questions on scraping ethics, limits, and reliability
A scraping SOP is a repeatable set of steps covering pre-run checks, tool configuration, and output cleaning before records reach a CRM or sequencer. Without one, inconsistent settings produce duplicate records and account restrictions.
LinkedIn's ToS explicitly prohibits automated scraping regardless of jurisdiction. Even where scraping public profiles may be legally permissible, it violates ToS and risks permanent account restriction.
100-200 profile visits per day is the commonly cited safer range for cookie-based tools, though no threshold is confirmed by LinkedIn. Randomized delays between actions reduce detection risk but do not eliminate it.
Three checks prevent most failures: confirm the test batch fills all required fields before scaling, verify proxy is active before any production run, and deduplicate against your database before any downstream push. Skipping the test batch is the most common cause of wasted execution credits.
The framework (scope, rate limits, test batch, dedup, field mapping, push) transfers across PhantomBuster, Apify, TexAu, and Captain Data. Specific delay values and field names are tool-specific: consult each tool's documentation for exact configuration.
SOP in place? Connect the scraper output to a sending workflow next.
The Signal to Cold Email Sequence workflow shows how to route clean scraping output into a triggered sequence without manual steps between data collection and the first send.