forumNew topic

We tried catching bot traffic with a honeypot — can we trust the results and what should we do next?

SSultan A***Member
Job title
QA Tester
Sector
Freight
Organization type
300-person organization
Joined
Aug 2025
Message
143

Doki · KVKK compliance consulting · 2023

#1

We run an e-commerce site offering a B2B spare parts catalog. We get around 4,000 unique visitors and roughly 25,000 pageviews a day. Over the last two months, our server costs skyrocketed from 4,500 TL to 11,000 TL. When we checked the logs, we noticed suspicious bots constantly scraping product prices and stock levels. As a fix, we set up honeypot traps using CSS-hidden fake links in our contact form and at the bottom of the page.

During a three-week test, we logged 1,400 distinct IP addresses that clicked these hidden links or filled out the form. Looking at the list, though, it seems some corporate proxy IPs, company security scanners, and even search engine cache bots fell into the trap too.

We're worried that if we trust these honeypot logs and block the IPs outright at the firewall, we might lose actual customers or hurt our SEO visibility. What is the proper way to filter this captured bot traffic, weed out the actual threats, and take safe action?

ÖÖzgür G***Member
Job title
Software developer
Sector
Cosmetics
Organization type
8-person team
Joined
Jun 2023
Message
16
Most Helpful#2

Short answer: permanently blocking every IP that hits a honeypot trap is the wrong approach. Corporate proxy servers, screen readers for visually impaired users, or legitimate search engine crawlers can easily trigger hidden page elements. Instead of banning captured requests immediately, you should run them through a tiered verification process.

Follow these steps for safe filtering: 1) First, run a reverse DNS lookup to check if the IP belongs to a verified, legitimate search engine. 2) Rather than flat-out banning an IP that triggers the trap, flag it as suspicious and serve a lightweight JavaScript challenge on subsequent requests from that address. 3) Because regular users' IPs change constantly on dynamic pools, any bans you issue should be temporary (15 minutes to an hour) rather than permanent. That way, you avoid locking out innocent users sharing the same IP pool.

The scrapers driving up your hosting costs usually follow an aggressive crawl rate. Instead of acting solely on a honeypot hit, rate-limit any address requesting more than 50 pages per minute. This protects server resources without risking real customers or your search engine indexing.

TTolga K***Member
Job title
IT manager
Sector
E-commerce
Organization type
a company within a holding
Joined
Jul 2024
Message
76
#3

If you hid elements using just display: none or visibility: hidden, modern bots simply parse the stylesheet, figure it out, and skip them. That leaves you catching only very primitive bots or accessibility tools. Pushing the hidden field off-screen (via text-indent or absolute positioning) and adding an aria-hidden attribute will significantly cut down on false positives.

EEsra A***Member
Job title
Technical service technician
Sector
Healthcare services
Organization type
120-person company
Joined
Dec 2025
Message
9
#4

We set up a similar honeypot last year. When we reviewed the first 800 captured IPs one by one, 120 turned out to be shared corporate egress IPs from our biggest wholesale clients. If we'd done a blanket ban that day, we would've completely cut off our top three ordering dealers.

edit: typed from phone, sorry for typos.

RRabia B***ExpertCommunity member
Joined
Mar 2025
Message
232
#5

Honeypots alone don't stop modern scrapers anymore. Commercial price-scraping tools run full headless browsers and strictly click elements that are genuinely visible on-screen. The 1,400 IPs you caught are likely just harmless background noise crawling the web; the bots actually harvesting your data probably never touched the trap.

AAylinMember
Job title
CRM and email
Organization type
300-person organization
Joined
Jun 2024
Message
118
#6

Instead of throwing an outright 403 Forbidden, serve a server-level JavaScript challenge to IPs triggering the trap. If it's a real person, their browser solves it in a fraction of a second and lets them in. A simple python- or curl-based scraper can't execute JS, hits a brick wall, and stops taxing your server.

PPerihan M***MemberCommunity member
Joined
Apr 2024
Message
83
#7

Don't panic and definitely avoid blanket IP bans. CGNAT is extremely common across Turkish ISPs; banning a single bot IP from a dynamic pool can instantly lock out hundreds of innocent home or office users. Score suspicious activity based on session behavior, not raw IPs.

Correction: I misremembered the figure, it was a bit lower.

EEsra K***MemberCommunity member
Joined
Feb 2023
Message
3
#8

search engine bots got caught in ours once and our indexes almost got wiped out, never write firewall ban rules without an rdns check first

LLale A***Member
Job title
Site Manager
Sector
IT services
Organization type
cooperative
Joined
Jul 2023
Message
86
#9

Did the hidden link at the bottom have a rel="nofollow" tag, and was that path disallowed in robots.txt? If it wasn't listed in robots.txt, it's completely normal for search engine crawlers to follow it.

YYasemin Y***MemberCommunity member
Joined
Oct 2023
Message
3
#10

Instead of just adding hidden fields to your honeypot, implement a timestamp check; any form submitted in under 2 seconds is almost guaranteed to be a bot.

RRıdvan K***MemberCommunity member
Joined
Oct 2024
Message
2
#11

I completely agree. When you try to change everything at once, nothing settles.

Good luck with that.

GGürkan Y***Member
Job title
Human Resources Specialist
Sector
Chemistry
Organization type
regional distributor
Joined
May 2023
Message
234
#12

Thanks, that was the answer I was looking for.

HHande B***Member
Job title
Operations manager
Sector
Sports and fitness
Organization type
boutique agency
Joined
Jun 2023
Message
353
#13

We experienced almost the exact same thing last year. Hasty decisions become decisions you have to fix six months later.

Don't rely on a single measure; go layer by layer. Good luck with that.

HHalil K***Member
Job title
Clinic manager
Sector
Seafood
Organization type
medium-sized business
Joined
May 2024
Message
208

Doki · Interface design · 2026

#14

You're right.

LLevent U***MemberCommunity member
Joined
Mar 2022
Message
270
#15

I agree, and I'd like to emphasize that. When making a decision, first look at what data you have on hand.

Correct me if I'm wrong.

DDilekNew member
Job title
Pastry Shop
Organization type
a company within a holding
Joined
Nov 2024
Message
19
#16

saved.

DDamla E***Veteran
Job title
Customer service representative
Sector
Freight
Organization type
boutique agency
Joined
Aug 2024
Message
414
#17

Great work. Having backups accessible on the same network and with the same identity makes them part of the target.

If the notification path is long, notifications don't arrive; missing notifications mean delayed incident detection. Correct me if I'm wrong.

RRabia K***MemberCommunity member
Joined
May 2022
Message
16
#18

I'm curious too.

HHakan A***ExpertCommunity member
Joined
Jan 2024
Message
3
#19

Do you think this works at any scale? Taking notes for two weeks yields better results than a six-month estimate.

Of course, it varies if your situation is different.

LLeyla P***MemberCommunity member
Joined
Oct 2022
Message
2
#20

Great work.

Reply