
In Progress
Posted
I need a clean, well-commented Python script that starts from a single root URL, follows every internal link it finds (no keyword or structural filtering at all), and extracts only the visible text content on each page. The script should rely on mainstream libraries—requests plus BeautifulSoup is fine, but feel free to propose Scrapy or an async stack if it fits better. Please keep the code modular so I can later drop individual functions into a bigger application. Core expectations • Crawl every reachable link within the domain, respecting [login to view URL] and an adjustable polite delay. • Skip images, PDFs, or other binary assets; focus strictly on textual information. • Save each page’s URL alongside the extracted text in a single output file (CSV or JSON—whichever you prefer is acceptable). • Handle time-outs, redirects, and JavaScript-heavy pages gracefully so the run never crashes halfway through. • Include a short README that covers environment setup, command to launch the crawl, and tunable arguments such as depth and rate-limit. Acceptance test Running `python [login to view URL] [login to view URL]` on my machine must finish without uncaught exceptions and produce the output file containing every crawled URL plus its text. That’s the whole task—once the script meets the above criteria, the project is complete.
Project ID: 40495818
38 proposals
Remote project
Active 6 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
38 freelancers are bidding on average $20 USD/hour for this job

Hi, your "Python Text Scraper Development - 06/06/2026 11:40 EDT" project is right in my wheelhouse. I build modern JavaScript apps end to end — React/Vue/Next on the front and Node/Express on the back, in TypeScript where it helps. Working with c programming, javascript, python, translation, web scraping, software architecture, json, scrapy, data extraction, beautifulsoup, I focus on responsive, fast UIs, clean component structure, and reliable APIs — no page-builder shortcuts. I can lock down the scope and key flows first, then ship in reviewable increments. Can we hop on a quick chat to align on your requirements? ⭐ 5.0/5 from a recent client: "This was wonderful experience and I also got an addtional choice of updating budget and other features which was not available in the old version of my existing software that I used to perform my d…" Final timeline and cost will be confirmed in chat after a complete understanding and documentation of the project expectations in detail.
$20 USD in 1 day
6.8
6.8

Hi, this is a straightforward domain crawler on the surface, but the real engineering risk is making traversal exhaustive without letting redirects, duplicate URL variants, or render-heavy pages break the run or poison the output. I’ve built Python systems like this with the same emphasis on clean hand-off, modularity, and failure-safe execution. The closest examples in my background are Python Bug Localization Using Transformer Models (CodeBERT + TreeBERT), where I delivered a documented CLI pipeline with reproducible setup, and Custom Feature Development & Integration, where the work had to be modular enough to drop into a larger application cleanly. I usually structure this as separate URL discovery, fetch/parse, text extraction, and persistence layers so crawl behavior stays predictable and easy to extend. With requests plus BeautifulSoup, that covers most domains cleanly, and I’d isolate the fallback path for JS-heavy pages so it does not complicate the normal crawl path. I also recommend explicit normalization, retry rules, robots checks, content-type filtering, and checkpointed writes so the process finishes without uncaught exceptions and produces usable output even when parts of the site are messy. This would be built as a maintainable script, not a one-off scrape. If useful, I can sketch the crawl flow and output schema before implementation. Thanks, Hercules
$50 USD in 40 days
6.8
6.8

I have 7 years of experience in Beautiful soup, selenium, and worked on multiple scraping projects and have scraped texts from ecommerce like amazon, walmart, ebay.I can work on your project
$20 USD in 40 days
6.5
6.5

Hi there, You need a robust, fully‑commented Python crawler that starts from one URL, follows every internal link without filtering, and records the visible text of each page. The main challenge is reliably handling redirects, time‑outs and JavaScript‑generated content while keeping the process stable. I’ll build a modular script using requests, BeautifulSoup and httpx with asyncio for optional parallelism. A tiny robots‑txt parser will enforce crawl rules, and a configurable delay queue will manage politeness. The core functions—fetch_page, extract_text, discover_links, and write_output—will be isolated so you can drop them into any larger codebase. Output will be a JSON lines file (URL + text) with a simple CSV alternative switch. I’ll add retry logic, graceful handling of binary URLs, and a concise README covering setup, CLI arguments (depth, rate‑limit, user‑agent) and execution steps. If you prefer Scrapy, I can wrap the same logic into a spider with minimal changes. Do you have a preferred format for storing the extracted data (single JSON array vs. line‑delimited) or any existing logging framework you’d like integrated? Thanks, please get in touch – looking forward to collaborating!
$15 USD in 40 days
6.2
6.2

Hi, I have read the job description and now ready to start. I will provide 100% quality. Please send me a message for discussion. Let's discuss the job. Thanks
$15 USD in 40 days
5.1
5.1

Lets chat, a free consultation and no obligation. I understand you need a clean, professional, and user-friendly solution for your "Python Text Scraper Development - 06/06/2026 11:40 EDT" project. My skills in PHP, Java, JavaScript are a perfect fit for this project. While I am new to freelancer.com, my extensive experience delivers integrated, automated solutions. Regards, Jason McLachlan
$15 USD in 3 days
3.0
3.0

I can develop a clean, well-commented Python text scraper that starts from your root URL and follows every internal link as requested. I have solid experience building efficient web scrapers that handle complex site structures while ensuring data integrity. My approach involves using Python libraries like BeautifulSoup and Requests combined with robust link-following logic to cover all internal pages without redundancy. Quick question before I suggest an approach: Do you need the scraper to handle dynamic content or just static HTML pages?
$20 USD in 7 days
0.0
0.0

Hi there. I’ve recently completed a similar project involving web scraping, and I can deliver your project efficiently to the same standard. I specialize in Python web scraping using BeautifulSoup and requests libraries. I ensure clean, well-commented code and modular design for easy integration into larger applications. Regarding your requirement to handle JavaScript-heavy pages gracefully, have you considered implementing a headless browser like Selenium for such scenarios? Regards, Riyaaz
$19 USD in 30 days
0.0
0.0

I have extensive experience with database management and accurate data entry, ensuring every record is migrated field-by-field without errors or omissions. I will carefully maintain the original formatting and order, verify each entry before saving, and document any anomalies encountered during the process. Ready to start immediately and complete the migration efficiently with a detailed completion log upon delivery.
$20 USD in 20 days
0.0
0.0

Uhana, Sri Lanka
Payment method verified
Member since Jan 25, 2026
$750-1500 USD
₹100-400 INR / hour
$20 USD
$30-250 USD
€8-30 EUR
₹12500-37500 INR
$10-30 USD
$30-250 USD
$15-25 USD / hour
£20-250 GBP
$250-750 USD
$125-250 USD
$250-750 USD
$250-750 USD
₹750-1250 INR / hour
$2-8 USD / hour
$2-8 USD / hour
₹37500-75000 INR
₹400-750 INR / hour
$20 USD