End of Life – Cloudflare protection

This topic is: resolved

 

Thank you for contacting me. Please note that I live in the GMT+3 time zone - responses might be delayed by this.

This topic has 1 reply, 2 voices, and was last updated 1 week, 5 days ago by Szabi – CodeRevolution.

Viewing 1 reply thread
  • Author
    Posts
    • #12811


      willbo987
      Participant
      Post count: 6

      Hi Szabi
      Been using Crawlomatic brilliantly over year or so.
      However a few of the main sites I scrape are now using Cloudflare protection.

      Despite now using Headless Browser (Puppeteer), it cannot get passed the protection.

      Do you have any other thoughts or do you think this may be the start of the end of API scrapers!

      One such site is https://www.commercialmotor.com/news/article/rh-commercial-vehicles-turns-to-ai-for-warranty-claims

    • #12812


      Szabi – CodeRevolution
      Keymaster
      Post count: 5110

      Hello,

      I am glad that the plugin was useful for you. The site you linked is probably using Cloudflare’s ‘Under Attack’ mode, this shows the captcha on almost all page request and is very aggressive against scrapers.

      I was not able to find a way around it yet. HeadlessBrowserAPI can still scrape sites behind Cloudflare, as long as they don’t enable this high level security feature.

      AI agents and search engine crawlers also struggle to read these highly protected pages, which on the long run will impact the website’s traffic, so I assume that on the long run, they will not keep these very restrictive security measures active.

      I will keep improving the scraping methods when possible, but I don’t want to promise a bypass that may stop working as soon as Cloudflare changes its protection again.

      Regards,
      Szabi – CodeRevolution.

Viewing 1 reply thread

You must be logged in to reply to this topic.