Apify Store · Personal Lab

Tools — PDF to Markdown, OCR, PageSpeed CWV, Geocode, Patents, Brazil CNPJ & More

In short: this page is the public storefront for Personal Lab Apify Actors. Nineteen Actors are live today — PDF to markdown / DOCX to markdown, OCR to markdown for scanned PDFs and images, RSS to markdown / Atom feeds, sitemap URL discovery, SEO metadata extraction, Wayback CDX to Markdown, a bulk URL status checker for broken links and redirects, WHOIS / DNS / SSL batch domain enrichment, domain availability via RDAP, batch image download to KV, batch image compress to WebP/JPEG/PNG, podcast metadata (no transcription), a website change monitor (hash / text diff, no browser or LLM), batch PageSpeed / Core Web Vitals via the Google PSI API (buyer brings own key), batch geocode — addresses to lat/lng and back via Google, Mapbox or HERE (buyer brings own key) → GeoJSON + Markdown, GTIN / UPC / EAN lookup — barcodes to Markdown product cards via UPCitemdb or Barcode Lookup (buyer brings own key) with keyless Open Food Facts fallback, patent to markdown — publication numbers to full-text Markdown (claims, specification, RAG chunks) from robots-allowed patents.google.com /patent/ pages only (no search), UK tenders — Find a Tender + Contracts Finder official OCDS APIs (keyless) to Markdown tender cards, and Brazil CNPJ — keyless BrasilAPI company lookups to Markdown due-diligence cards (identity, status, CNAE, address, QSA; no email/phone in outputs). Browse by Convert / Discover / Check below. Pricing is pay-per-use or usage-based on the Apify Store. Booked Unstaffed revenue from these listings remains exactly $0 as of the latest gated journal note [1]. We do not invent customer counts or reviews.

  • Livepay-per-useConvertpdf-docx-to-markdown

    PDF & DOCX to Markdown

    Text PDF and Word → structured Markdown for RAG.

    Convert text-based PDF and Word (DOCX) files to LLM-ready Markdown with real tables, page markers, and optional RAG chunks.

  • Liveusage-basedDiscoversitemap-url-discovery

    Sitemap URL Extractor

    Find sitemaps and list every URL with lastmod tags.

    Discover a site’s XML sitemaps and list every URL with lastmod, source sitemap, and file-type tags for SEO audits and document inventories.

  • Livepay-per-useConvertscanned-ocr-to-markdown

    Scanned PDF/Image OCR to Markdown

    OCR scans and images → Markdown, offline RapidOCR.

    OCR scanned PDFs and images into Markdown with page markers and optional RAG chunks — offline RapidOCR, no external AI API keys.

  • Livepay-per-useConvertrss-atom-to-markdown

    RSS & Atom to Markdown — JSON + RAG Chunks

    RSS and Atom feeds → Markdown bodies and chunks.

    Parse RSS and Atom feeds into structured JSON with Markdown body text, optional feed discovery, and heading-aware RAG chunks.

  • Livepay-per-useCheckurl-status-checker

    Bulk URL Status Checker — Broken Links & Redirects

    Bulk HTTP status, redirects, and broken-link classes.

    Check HTTP status for a URL list: redirects, content-type, timing, and error class — chain from sitemap Actor output for broken-link audits.

  • Livepay-per-useCheckwhois-dns-ssl-lookup

    WHOIS DNS SSL Lookup — Batch Domain Enrichment

    WHOIS, DNS, and SSL expiry for domain lists.

    Batch-enrich domains with public WHOIS, DNS (A/AAAA/MX/NS/TXT), and TLS certificate fields — 256 MB default, failed domains free unless you opt in.

  • Livepay-per-useDiscoverseo-metadata-extract

    SEO Metadata Extractor

    Extract SEO metadata from a URL into structured output.

    Extract SEO metadata from a URL, with pay-per-use pricing of $0.0005 per URL-extracted item.

  • Livepay-per-useDiscoverwayback-cdx-markdown

    Wayback CDX Markdown

    Turn Wayback CDX snapshot items into Markdown-ready records.

    Format Internet Archive Wayback CDX snapshot items as Markdown-ready records, with pay-per-use pricing of $0.003 per snapshot item.

  • Livepay-per-useCheckdomain-availability-rdap

    Domain Availability Checker — RDAP Batch

    Batch available/registered verdicts via RDAP.

    Batch-check whether domains are available or registered via public RDAP (DNS/WHOIS fallback) — 256 MB default; unknown/failed free unless you opt in.

  • Livepay-per-useConvertbatch-image-downloader

    Batch Image Downloader — URLs & Pages to KV

    Download image URLs or page imgs into KV store.

    Download image URL lists or extract img/og:image from pages into Apify KV with metadata (contentType, bytes, sha256, dims) — 512 MB default; failed/empty free.

  • Livepay-per-useConvertbatch-image-compress

    Batch Image Compress & Convert — WebP/JPEG/PNG

    Optimize image URLs to WebP/JPEG/PNG in KV.

    Batch WebP/JPEG/PNG optimization pipeline for image URL lists (or downloader dataset/KV) with bytesBefore/After and optional EXIF strip — 512 MB default; failed free. AVIF runtime optional / currently unavailable. No crawl.

  • Livepay-per-useDiscoverpodcast-metadata

    Podcast Metadata — Shows & Episodes (No Transcription)

    iTunes + RSS show/episode metadata; no transcription.

    Apple Podcasts / iTunes Search & Lookup plus RSS show and episode metadata — no transcription, no audio download, no email/contacts. Chain to rss-atom-to-markdown for notes MD. 256 MB default; failed/empty feeds free.

  • Livepay-per-useCheckwebsite-change-monitor

    Website Change Monitor — Hash & Text Diff

    SHA-256 hash + text diff for URL lists; no browser/LLM.

    Monitor URL lists for content changes with SHA-256 hash, optional CSS selector, and difflib text diff — HTTP only, no browser, no LLM. Chain from sitemap / URL status. 256 MB default; failed checks free.

  • Livepay-per-useCheckbatch-pagespeed-cwv

    Batch PageSpeed & Core Web Vitals — Mobile+Desktop+CrUX

    Batch Google PageSpeed / CWV via PSI API; buyer brings own key.

    Batch Google PageSpeed Insights via the official PSI API: mobile + desktop Lighthouse scores, lab Core Web Vitals (LCP/CLS/TBT/FCP) and CrUX field data when available. Markdown + CSV report. URL list, sitemap or dataset input. Buyer brings own free Google API key. 256 MB default; failed calls free.

  • Livepay-per-useConvertbatch-geocode

    Batch Geocode — Google / Mapbox / HERE → GeoJSON

    BYO Google/Mapbox/HERE key → lat/lng, GeoJSON + Markdown; failed rows free.

    Batch forward and reverse geocoding via the official Google Geocoding, Mapbox and HERE APIs — buyer brings their own key. Lat/lng, formatted address, components and accuracy per row, a GeoJSON FeatureCollection and a Markdown report; optional Google Distance Matrix. No Maps scraping. 256 MB default; not-found and failed rows free.

  • Livepay-per-useConvertgtin-product-markdown

    GTIN / UPC / EAN → Product Markdown Card

    BYO UPCitemdb/Barcode Lookup key (or keyless Open Food Facts) → Markdown product cards.

    Look up GTIN / UPC / EAN barcodes via the official UPCitemdb or Barcode Lookup APIs (buyer brings their own key), with keyless Open Food Facts as a food/FMCG fallback. One structured row + Markdown product card per barcode, plus optional RAG chunks. No Idealo or Amazon scraping. 256 MB default; not-found and failed rows free.

  • Livepay-per-useDiscoveruk-tenders-ocds-markdown

    UK Tenders — Find a Tender + Contracts Finder OCDS

    Keyless official UK OCDS APIs → Markdown tender cards; failures free.

    UK public tenders from the official Find a Tender + Contracts Finder OCDS APIs (no HTML scraping, no API key). Filter by keywords, buyer, CPV, dates and stages (client-side). One Markdown tender card + JSON per notice. 256 MB; failed and empty rows free. No buyer email/phone harvesting.

  • Livepay-per-useConvertpatent-to-markdown

    Patent Number → Full-Text Markdown (Claims & Spec)

    Publication numbers → Markdown claims/spec + RAG chunks; /patent/ only.

    Turn patent publication numbers (e.g. US9876543B2, EP…, WO…) into full-text Markdown: title, abstract, itemized claims, specification, CPC/citations, and ~500–1000 char RAG chunks. Fetches only robots-allowed patents.google.com /patent/ pages — no search/SERP. 512 MB; failed/not-found rows free. No API key.

  • Livepay-per-useConvertbrazil-cnpj-markdown

    Brazil CNPJ → Company Markdown DD Card

    Keyless BrasilAPI CNPJ → Markdown DD cards; no email/phone; failures free.

    Batch Brazilian CNPJ lookups via BrasilAPI (keyless) with CNPJ.ws / ReceitaWS fallback or your own key. Markdown due-diligence cards: identity, status, CNAE, address, QSA partners, source. RAG chunks. No email/phone in outputs. Failed rows free. 256 MB.

Definitions

  • PDF to markdown is the conversion of a text-based PDF into structured Markdown — headings, lists, tables, and page markers — so RAG and LLM pipelines can read documents cleanly.
  • DOCX to markdown is the same idea for Microsoft Word (.docx) files, including tables where the document structure allows.
  • OCR to markdown means turning scanned pages or images into Markdown via optical character recognition (print scans; handwriting is unreliable).
  • RSS to markdown means converting RSS or Atom feed items into structured JSON with Markdown body text for archives and RAG pipelines.
  • Sitemap URL discovery means finding a website’s XML sitemaps (via robots.txt or common paths) and listing every URL with lastmod, source sitemap, and a file-type tag (for example PDF).
  • A broken link checker / URL status checker reports HTTP status, redirects, and error class for a list of URLs (optionally chained from sitemap output).
  • SEO metadata extraction means collecting page metadata from a URL into structured output for audits and inventories.
  • Wayback CDX to Markdown means formatting Internet Archive snapshot items as Markdown-ready records.
  • WHOIS / DNS / SSL lookup means batch-enriching a domain list with public WHOIS fields, DNS records (A/AAAA/MX/NS/TXT), and TLS certificate expiry — useful before RAG ingest or crawl.
  • Domain availability (RDAP) means batch-checking whether names are available or registered via public RDAP, with optional DNS/WHOIS fallback — inventory checks, not lead-gen.
  • Batch image download means fetching image URL lists (or extracting img/og:image from pages) into Apify key-value store with metadata for RAG and archives.
  • Batch image compress means optimizing image URL lists (or an upstream downloader dataset/KV) to WebP/JPEG/PNG with bytesBefore/After and optional EXIF strip — no page crawl. AVIF is runtime optional / currently unavailable.
  • Podcast metadata means fetching Apple Podcasts / iTunes and RSS show and episode fields without transcription, audio download, or contact scraping — then chaining to RSS & Atom to Markdown for show-notes bodies.
  • Website change monitor means watching URL lists for content changes via SHA-256 hash and text diff (optional CSS selector) — HTTP only, no browser, no LLM.
  • Batch PageSpeed / Core Web Vitals means calling Google's PageSpeed Insights API for URL lists (mobile + desktop), collecting Lighthouse scores, lab CWV, and CrUX field data when available, then writing Markdown + CSV — buyer brings their own Google API key; failed calls free.
  • Batch geocode means turning address lists into latitude / longitude (or coordinates into addresses) through the official Google Geocoding, Mapbox or HERE APIs with the buyer's own key, then writing a GeoJSON FeatureCollection + Markdown report; not-found and failed rows free.
  • GTIN / UPC / EAN to product card means looking up retail barcodes through the official UPCitemdb or Barcode Lookup APIs with the buyer's own key (or keyless Open Food Facts for food / FMCG) and writing one structured row + Markdown product card per barcode, with optional RAG chunks; not-found and failed rows free.
  • Patent to markdown means turning patent publication numbers into full-text Markdown (title, abstract, itemized claims, specification) plus ~500–1000 character RAG chunks and CPC/citation metadata — fetching only robots-allowed patents.google.com /patent/ pages (no search/SERP); failed and not-found rows free; no API key.
  • UK tenders (OCDS) means fetching UK public-sector notices from the official Find a Tender and Contracts Finder OCDS APIs (no API key, no HTML scraping), filtering client-side by keywords / buyer / CPV / stages, and writing one Markdown tender card + JSON per notice; failed and empty searches free; buyer email/phone redacted.

How these tools work

Why publish small Actors? Because document, feed, and URL utilities are easy to meter on a platform marketplace, and the Store routes buyers without a custom billing stack. According to the Apify Store listing for PDF & DOCX to Markdown [2], the Actor runs on the Apify platform with pay-per-use pricing (page-based events documented on that listing). According to the Sitemap URL Extractor Store page [3], sitemap crawling is usage-based compute on Apify. The OCR, RSS, URL-status, WHOIS/DNS/SSL, domain-availability, batch-image (download + compress), podcast-metadata, website-change-monitor, batch-pagespeed-cwv, batch-geocode, gtin-product-markdown, patent-to-markdown, uk-tenders-ocds-markdown, and brazil-cnpj-markdown Actors are likewise pay-per-use on their Store pages [4][5][6][7].

  1. Pick the Actor that matches the job (document conversion, OCR, feed parsing, sitemap discovery, SEO metadata, Wayback CDX, URL status, WHOIS/DNS/SSL, domain availability, batch image download, batch image compress, podcast metadata, website change monitor, batch PageSpeed / CWV, batch geocoding, GTIN / UPC / EAN product lookup, patent to markdown, UK tenders OCDS, or Brazil CNPJ).
  2. Open the Apify Store link and run with your own Apify account.
  3. Pay only for usage on the platform — this Lab page does not sell seats or invent “customers.”

In summary

Personal Lab currently ships nineteen public Apify Actors for PDF to markdown, DOCX to markdown , OCR to markdown, RSS to markdown, sitemap URL discovery, SEO metadata extraction, Wayback CDX to Markdown, a broken link / URL status checker, WHOIS / DNS / SSL domain enrichment, domain availability via RDAP, batch image download to KV, batch image compress to WebP/JPEG/PNG, podcast metadata (no transcription), a website change monitor (hash / text diff), batch PageSpeed / Core Web Vitals via PSI, batch geocode to GeoJSON (Google / Mapbox / HERE), GTIN / UPC / EAN barcodes to Markdown product cards, patent publication numbers to full-text Markdown (claims & spec + RAG chunks), UK tenders from Find a Tender + Contracts Finder OCDS, and Brazil CNPJ company Markdown due-diligence cards (BrasilAPI keyless; no email/phone). Pricing is pay-per-use or usage-based on Apify. Honest bookkeeping: Store listings can be live while booked revenue is still $0 [1].

Workflow guides

Intent landing pages that walk through the matching Actor workflow (soft links; no purchase required to use the checklists).

Sources

  1. D044 — Apify Actor went live; booked revenue still $0 — Personal Lab gated journal, 2026-09-29.
  2. PDF & DOCX to Markdown — Apify Store · AGPL source on GitHub
  3. Sitemap URL Extractor — Apify Store
  4. Scanned PDF/Image OCR to Markdown — Apify Store
  5. RSS & Atom to Markdown — Apify Store
  6. Bulk URL Status Checker — Apify Store
  7. WHOIS DNS SSL Lookup — Apify Store
  8. SEO Metadata Extractor — Apify Store · GitHub source
  9. Wayback CDX Markdown — Apify Store · GitHub source
  10. Domain Availability Checker — RDAP Batch — Apify Store · GitHub source
  11. Batch Image Downloader — Apify Store · GitHub source
  12. Batch Image Compress & Convert — Apify Store · GitHub source
  13. Podcast Metadata — Shows & Episodes (No Transcription) — Apify Store · GitHub source
  14. Website Change Monitor — Hash & Text Diff — Apify Store · GitHub source
  15. Batch PageSpeed & Core Web Vitals — Apify Store · GitHub source
  16. Batch Geocode — Google / Mapbox / HERE → GeoJSON — Apify Store · GitHub source
  17. GTIN / UPC / EAN → Product Markdown Card — Apify Store · GitHub source
  18. UK Tenders — Find a Tender + Contracts Finder OCDS — Apify Store · GitHub source
  19. Patent Number → Full-Text Markdown (Claims & Spec) — Apify Store · GitHub source
  20. Brazil CNPJ → Company Markdown DD Card — Apify Store · GitHub source