AI Automation Β· Workflow

Scraping SOP: Ethics + Reliability

A repeatable 6-step procedure for running scraping jobs that stay inside platform rules, produce clean output, and do not put your accounts at risk.

Written for operators No vendor influence Practical, not theoretical

Before You Start

45-90 min setup: output, prerequisites, and risk

Output: A repeatable scraping workflow that collects structured lead data, deduplicates against existing records, and pushes a clean export to your sequencer or CRM without triggering platform restrictions.

Time: 45-90 minutes for initial setup. Under 15 minutes per repeat run once configured and tested.

πŸ“‹
Prerequisites

Active account in PhantomBuster, Apify, TexAu, or Captain Data. Defined target source (LinkedIn search URL, Sales Navigator list, Google Maps query, or website URL). CRM or sequencer ready to receive output. ToS status confirmed for the target platform.

🚨
LinkedIn prohibits scraping by default

LinkedIn's ToS explicitly prohibits scraping and unauthorized automation. Cookie-based scrapers without rate limiting risk permanent account restriction. This SOP reduces that risk but cannot eliminate it: read every ethics gate before running a LinkedIn job.

Workflow Overview

6-step scraping SOP: scope to clean export

StepActionOutputEthics gate?
1Define scope, source, and legal boundaryApproved source list with ToS status confirmedYes
2Configure tool, execution limits, and delay settingsScraper config file with rate limits setYes
3Set up proxy or cloud executionActive proxy or cloud session connected to toolNo
4Run test batch (50 rows max)Sample output reviewed for field accuracy and completenessNo
5Deduplicate against existing databaseNet-new records only, ready for enrichmentNo
6Push clean output to CRM or sequencerVerified, field-mapped records in downstream toolNo

Step by Step

6 steps: source check to clean push

  1. Define your source and confirm its ToS status

    Check the platform's ToS and robots.txt before opening any scraping tool. LinkedIn explicitly prohibits scraping and risks permanent account suspension. Public sites and business directories carry lower but variable risk.

  2. Set execution limits and per-action delay intervals

    Override default speed settings: 3-8 seconds per action on LinkedIn, 1-3 seconds on other sites. Cap LinkedIn runs at 100-200 records per day as a conservative ceiling.

  3. Configure proxy or cloud execution

    Use residential proxies for social network sources: datacenter IPs are trivially detected by LinkedIn and most anti-bot systems. Apify includes built-in residential proxy infrastructure; PhantomBuster uses your browser cookie session.

  4. Run a test batch of 50 rows and review every output field

    Set the run limit to 50 records and verify every required field populates correctly before scaling. If a key field is absent in over 20% of rows, fix the scraper config before proceeding.

  5. Deduplicate against your existing database before full export

    Match on LinkedIn URL or email as the deduplication key, not name or company alone. Most scraping tools don't deduplicate automatically: use Clay, a VLOOKUP, or a CRM rule to filter before any downstream push.

  6. Map output fields to your CRM or sequencer schema before pushing

    Create a field mapping document before the first push: scraper column on the left, CRM field on the right. Mismatched fields create data hygiene problems that compound across every future campaign run.

⚠️
GDPR applies to EU contacts

Scraping a public page does not create a lawful basis for marketing EU-based contacts. Use legitimate interest as the lawful basis for B2B outreach and verify compliance for your stack: Apify is SOC 2 + GDPR-certified; TexAu is GDPR-compliant; PhantomBuster is self-declared.

When Things Break

4 failure modes and the fix for each

If
LinkedIn account gets restricted mid-run
Stop immediately and pause the account for at least 24-48 hours. Reduce the daily cap by half, increase per-action delay, and consider a secondary account for future scraping runs.
If
Captcha blocks the scraper after a few pages
Switch to residential proxies and rotate user-agent strings between requests. Apify handles anti-blocking automatically on most Actors; PhantomBuster requires manual concurrency reduction and delay increases.
If
Output fields are empty or inconsistent across rows
The scraper is targeting the wrong DOM element or the page uses JavaScript rendering not accessible in page source. Apify's website content crawler handles JS-rendered pages natively.
If
Downstream tool receives duplicate or misformatted records
Deduplication was skipped or field mapping was not reviewed before the push. Use LinkedIn URL or verified email as the unique key and cross-check output column names against your CRM's exact field labels.

Tool Fit

PhantomBuster, Apify, TexAu, Captain Data: choose by use case

PhantomBuster
Multi-platform
130+ pre-built Phantoms for LinkedIn, Google Maps, Instagram, and 12+ other platforms. Cookie-based LinkedIn sessions carry ToS risk.
LinkedIn scraping Real-time leads CRM sync
Apify
Web data platform
19,000+ Actors with built-in anti-blocking, residential proxies, and API integration. Best for non-social-network sources and AI agent data pipelines.
Anti-blocking Proxy infra API native
TexAu
GTM enrichment
Waterfall enrichment across 150+ providers with built-in email verification, AI scoring, and CRM sync. GDPR-compliant, 256-bit encryption.
150+ providers Email verify CRM sync
Captain Data
B2B data API
Usage-based API with 7 endpoints for people search, enrichment, and employee discovery. Built for developers embedding B2B data into AI agents or SaaS features.
7 endpoints MCP ready SOC 2 certified
πŸ’‘
3 checks before every run

Does the platform's ToS permit automated collection? Are speed and daily caps below detection threshold? Does a 50-row test batch confirm all required fields populate? If any answer is no: stop.

FAQ

5 questions on scraping ethics, limits, and reliability

What is a scraping SOP and why does outbound need one?

A scraping SOP is a repeatable set of steps covering pre-run checks, tool configuration, and output cleaning before records reach a CRM or sequencer. Without one, inconsistent settings produce duplicate records and account restrictions.

Is scraping LinkedIn legal for B2B outreach?

LinkedIn's ToS explicitly prohibits automated scraping regardless of jurisdiction. Even where scraping public profiles may be legally permissible, it violates ToS and risks permanent account restriction.

How many records per day is safe on LinkedIn?

100-200 profile visits per day is the commonly cited safer range for cookie-based tools, though no threshold is confirmed by LinkedIn. Randomized delays between actions reduce detection risk but do not eliminate it.

What scraping SOP checklist items matter most for reliability?

Three checks prevent most failures: confirm the test batch fills all required fields before scaling, verify proxy is active before any production run, and deduplicate against your database before any downstream push. Skipping the test batch is the most common cause of wasted execution credits.

Can I apply this SOP across different scraping tools?

The framework (scope, rate limits, test batch, dedup, field mapping, push) transfers across PhantomBuster, Apify, TexAu, and Captain Data. Specific delay values and field names are tool-specific: consult each tool's documentation for exact configuration.

SOP in place? Connect the scraper output to a sending workflow next.

The Signal to Cold Email Sequence workflow shows how to route clean scraping output into a triggered sequence without manual steps between data collection and the first send.