PDF & DOCX to Markdown
Text PDF and Word → structured Markdown for RAG.
Convert text-based PDF and Word (DOCX) files to LLM-ready Markdown with real tables, page markers, and optional RAG chunks.
Apify Store · Personal Lab
In short: this page is the public storefront for Personal Lab Apify Actors. Nineteen Actors are live today — PDF to markdown / DOCX to markdown, OCR to markdown for scanned PDFs and images, RSS to markdown / Atom feeds, sitemap URL discovery, SEO metadata extraction, Wayback CDX to Markdown, a bulk URL status checker for broken links and redirects, WHOIS / DNS / SSL batch domain enrichment, domain availability via RDAP, batch image download to KV, batch image compress to WebP/JPEG/PNG, podcast metadata (no transcription), a website change monitor (hash / text diff, no browser or LLM), batch PageSpeed / Core Web Vitals via the Google PSI API (buyer brings own key), batch geocode — addresses to lat/lng and back via Google, Mapbox or HERE (buyer brings own key) → GeoJSON + Markdown, GTIN / UPC / EAN lookup — barcodes to Markdown product cards via UPCitemdb or Barcode Lookup (buyer brings own key) with keyless Open Food Facts fallback, patent to markdown — publication numbers to full-text Markdown (claims, specification, RAG chunks) from robots-allowed patents.google.com /patent/ pages only (no search), UK tenders — Find a Tender + Contracts Finder official OCDS APIs (keyless) to Markdown tender cards, and Brazil CNPJ — keyless BrasilAPI company lookups to Markdown due-diligence cards (identity, status, CNAE, address, QSA; no email/phone in outputs). Browse by Convert / Discover / Check below. Pricing is pay-per-use or usage-based on the Apify Store. Booked Unstaffed revenue from these listings remains exactly $0 as of the latest gated journal note [1]. We do not invent customer counts or reviews.
Text PDF and Word → structured Markdown for RAG.
Convert text-based PDF and Word (DOCX) files to LLM-ready Markdown with real tables, page markers, and optional RAG chunks.
Find sitemaps and list every URL with lastmod tags.
Discover a site’s XML sitemaps and list every URL with lastmod, source sitemap, and file-type tags for SEO audits and document inventories.
OCR scans and images → Markdown, offline RapidOCR.
OCR scanned PDFs and images into Markdown with page markers and optional RAG chunks — offline RapidOCR, no external AI API keys.
WHOIS, DNS, and SSL expiry for domain lists.
Batch-enrich domains with public WHOIS, DNS (A/AAAA/MX/NS/TXT), and TLS certificate fields — 256 MB default, failed domains free unless you opt in.
Batch available/registered verdicts via RDAP.
Batch-check whether domains are available or registered via public RDAP (DNS/WHOIS fallback) — 256 MB default; unknown/failed free unless you opt in.
Download image URLs or page imgs into KV store.
Download image URL lists or extract img/og:image from pages into Apify KV with metadata (contentType, bytes, sha256, dims) — 512 MB default; failed/empty free.
Optimize image URLs to WebP/JPEG/PNG in KV.
Batch WebP/JPEG/PNG optimization pipeline for image URL lists (or downloader dataset/KV) with bytesBefore/After and optional EXIF strip — 512 MB default; failed free. AVIF runtime optional / currently unavailable. No crawl.
iTunes + RSS show/episode metadata; no transcription.
Apple Podcasts / iTunes Search & Lookup plus RSS show and episode metadata — no transcription, no audio download, no email/contacts. Chain to rss-atom-to-markdown for notes MD. 256 MB default; failed/empty feeds free.
SHA-256 hash + text diff for URL lists; no browser/LLM.
Monitor URL lists for content changes with SHA-256 hash, optional CSS selector, and difflib text diff — HTTP only, no browser, no LLM. Chain from sitemap / URL status. 256 MB default; failed checks free.
Batch Google PageSpeed / CWV via PSI API; buyer brings own key.
Batch Google PageSpeed Insights via the official PSI API: mobile + desktop Lighthouse scores, lab Core Web Vitals (LCP/CLS/TBT/FCP) and CrUX field data when available. Markdown + CSV report. URL list, sitemap or dataset input. Buyer brings own free Google API key. 256 MB default; failed calls free.
BYO Google/Mapbox/HERE key → lat/lng, GeoJSON + Markdown; failed rows free.
Batch forward and reverse geocoding via the official Google Geocoding, Mapbox and HERE APIs — buyer brings their own key. Lat/lng, formatted address, components and accuracy per row, a GeoJSON FeatureCollection and a Markdown report; optional Google Distance Matrix. No Maps scraping. 256 MB default; not-found and failed rows free.
BYO UPCitemdb/Barcode Lookup key (or keyless Open Food Facts) → Markdown product cards.
Look up GTIN / UPC / EAN barcodes via the official UPCitemdb or Barcode Lookup APIs (buyer brings their own key), with keyless Open Food Facts as a food/FMCG fallback. One structured row + Markdown product card per barcode, plus optional RAG chunks. No Idealo or Amazon scraping. 256 MB default; not-found and failed rows free.
Keyless official UK OCDS APIs → Markdown tender cards; failures free.
UK public tenders from the official Find a Tender + Contracts Finder OCDS APIs (no HTML scraping, no API key). Filter by keywords, buyer, CPV, dates and stages (client-side). One Markdown tender card + JSON per notice. 256 MB; failed and empty rows free. No buyer email/phone harvesting.
Publication numbers → Markdown claims/spec + RAG chunks; /patent/ only.
Turn patent publication numbers (e.g. US9876543B2, EP…, WO…) into full-text Markdown: title, abstract, itemized claims, specification, CPC/citations, and ~500–1000 char RAG chunks. Fetches only robots-allowed patents.google.com /patent/ pages — no search/SERP. 512 MB; failed/not-found rows free. No API key.
Keyless BrasilAPI CNPJ → Markdown DD cards; no email/phone; failures free.
Batch Brazilian CNPJ lookups via BrasilAPI (keyless) with CNPJ.ws / ReceitaWS fallback or your own key. Markdown due-diligence cards: identity, status, CNAE, address, QSA partners, source. RAG chunks. No email/phone in outputs. Failed rows free. 256 MB.
Text PDF and Word → structured Markdown for RAG.
Convert text-based PDF and Word (DOCX) files to LLM-ready Markdown with real tables, page markers, and optional RAG chunks.
Find sitemaps and list every URL with lastmod tags.
Discover a site’s XML sitemaps and list every URL with lastmod, source sitemap, and file-type tags for SEO audits and document inventories.
OCR scans and images → Markdown, offline RapidOCR.
OCR scanned PDFs and images into Markdown with page markers and optional RAG chunks — offline RapidOCR, no external AI API keys.
RSS and Atom feeds → Markdown bodies and chunks.
Parse RSS and Atom feeds into structured JSON with Markdown body text, optional feed discovery, and heading-aware RAG chunks.
Bulk HTTP status, redirects, and broken-link classes.
Check HTTP status for a URL list: redirects, content-type, timing, and error class — chain from sitemap Actor output for broken-link audits.
WHOIS, DNS, and SSL expiry for domain lists.
Batch-enrich domains with public WHOIS, DNS (A/AAAA/MX/NS/TXT), and TLS certificate fields — 256 MB default, failed domains free unless you opt in.
Extract SEO metadata from a URL into structured output.
Extract SEO metadata from a URL, with pay-per-use pricing of $0.0005 per URL-extracted item.
Turn Wayback CDX snapshot items into Markdown-ready records.
Format Internet Archive Wayback CDX snapshot items as Markdown-ready records, with pay-per-use pricing of $0.003 per snapshot item.
Batch available/registered verdicts via RDAP.
Batch-check whether domains are available or registered via public RDAP (DNS/WHOIS fallback) — 256 MB default; unknown/failed free unless you opt in.
Download image URLs or page imgs into KV store.
Download image URL lists or extract img/og:image from pages into Apify KV with metadata (contentType, bytes, sha256, dims) — 512 MB default; failed/empty free.
Optimize image URLs to WebP/JPEG/PNG in KV.
Batch WebP/JPEG/PNG optimization pipeline for image URL lists (or downloader dataset/KV) with bytesBefore/After and optional EXIF strip — 512 MB default; failed free. AVIF runtime optional / currently unavailable. No crawl.
iTunes + RSS show/episode metadata; no transcription.
Apple Podcasts / iTunes Search & Lookup plus RSS show and episode metadata — no transcription, no audio download, no email/contacts. Chain to rss-atom-to-markdown for notes MD. 256 MB default; failed/empty feeds free.
SHA-256 hash + text diff for URL lists; no browser/LLM.
Monitor URL lists for content changes with SHA-256 hash, optional CSS selector, and difflib text diff — HTTP only, no browser, no LLM. Chain from sitemap / URL status. 256 MB default; failed checks free.
Batch Google PageSpeed / CWV via PSI API; buyer brings own key.
Batch Google PageSpeed Insights via the official PSI API: mobile + desktop Lighthouse scores, lab Core Web Vitals (LCP/CLS/TBT/FCP) and CrUX field data when available. Markdown + CSV report. URL list, sitemap or dataset input. Buyer brings own free Google API key. 256 MB default; failed calls free.
BYO Google/Mapbox/HERE key → lat/lng, GeoJSON + Markdown; failed rows free.
Batch forward and reverse geocoding via the official Google Geocoding, Mapbox and HERE APIs — buyer brings their own key. Lat/lng, formatted address, components and accuracy per row, a GeoJSON FeatureCollection and a Markdown report; optional Google Distance Matrix. No Maps scraping. 256 MB default; not-found and failed rows free.
BYO UPCitemdb/Barcode Lookup key (or keyless Open Food Facts) → Markdown product cards.
Look up GTIN / UPC / EAN barcodes via the official UPCitemdb or Barcode Lookup APIs (buyer brings their own key), with keyless Open Food Facts as a food/FMCG fallback. One structured row + Markdown product card per barcode, plus optional RAG chunks. No Idealo or Amazon scraping. 256 MB default; not-found and failed rows free.
Keyless official UK OCDS APIs → Markdown tender cards; failures free.
UK public tenders from the official Find a Tender + Contracts Finder OCDS APIs (no HTML scraping, no API key). Filter by keywords, buyer, CPV, dates and stages (client-side). One Markdown tender card + JSON per notice. 256 MB; failed and empty rows free. No buyer email/phone harvesting.
Publication numbers → Markdown claims/spec + RAG chunks; /patent/ only.
Turn patent publication numbers (e.g. US9876543B2, EP…, WO…) into full-text Markdown: title, abstract, itemized claims, specification, CPC/citations, and ~500–1000 char RAG chunks. Fetches only robots-allowed patents.google.com /patent/ pages — no search/SERP. 512 MB; failed/not-found rows free. No API key.
Keyless BrasilAPI CNPJ → Markdown DD cards; no email/phone; failures free.
Batch Brazilian CNPJ lookups via BrasilAPI (keyless) with CNPJ.ws / ReceitaWS fallback or your own key. Markdown due-diligence cards: identity, status, CNAE, address, QSA partners, source. RAG chunks. No email/phone in outputs. Failed rows free. 256 MB.
.docx) files, including tables where the document structure allows.Why publish small Actors? Because document, feed, and URL utilities are easy to meter on a platform marketplace, and the Store routes buyers without a custom billing stack. According to the Apify Store listing for PDF & DOCX to Markdown [2], the Actor runs on the Apify platform with pay-per-use pricing (page-based events documented on that listing). According to the Sitemap URL Extractor Store page [3], sitemap crawling is usage-based compute on Apify. The OCR, RSS, URL-status, WHOIS/DNS/SSL, domain-availability, batch-image (download + compress), podcast-metadata, website-change-monitor, batch-pagespeed-cwv, batch-geocode, gtin-product-markdown, patent-to-markdown, uk-tenders-ocds-markdown, and brazil-cnpj-markdown Actors are likewise pay-per-use on their Store pages [4][5][6][7].
Personal Lab currently ships nineteen public Apify Actors for PDF to markdown, DOCX to markdown , OCR to markdown, RSS to markdown, sitemap URL discovery, SEO metadata extraction, Wayback CDX to Markdown, a broken link / URL status checker, WHOIS / DNS / SSL domain enrichment, domain availability via RDAP, batch image download to KV, batch image compress to WebP/JPEG/PNG, podcast metadata (no transcription), a website change monitor (hash / text diff), batch PageSpeed / Core Web Vitals via PSI, batch geocode to GeoJSON (Google / Mapbox / HERE), GTIN / UPC / EAN barcodes to Markdown product cards, patent publication numbers to full-text Markdown (claims & spec + RAG chunks), UK tenders from Find a Tender + Contracts Finder OCDS, and Brazil CNPJ company Markdown due-diligence cards (BrasilAPI keyless; no email/phone). Pricing is pay-per-use or usage-based on Apify. Honest bookkeeping: Store listings can be live while booked revenue is still $0 [1].
Intent landing pages that walk through the matching Actor workflow (soft links; no purchase required to use the checklists).