About The Watchtower

The Watchtower is this board, the list it publishes and the network of sensors behind it: today our own hosts, later any Linux box whose operator installs the Watchtower client and contributes what hits it, and pulls the list back into its firewall.

<!-- Source text for /threats/about. The page task (openspec/changes/threat-board/tasks.md phase 2, render_threats_about() in public_html/lib/pages/threats.php) renders this file; it does not write new prose. Policy numbers here are cited from docs/adr/0006-public-threat-listing-policy.md and openspec/changes/threat-board/specs/threat-board/spec.md and must never drift from them — if a number changes, change it there first, then here. -->

servbg.com runs its own small network of sensors — on this host, on our mail host, on the origin we use for a few other sites — that record traffic attacking them: credential stuffing, CVE exploitation attempts, brute-force logins, scanner probes. This page explains what we collect, how an address ends up named on the board, what we never publish, and how to get an address removed if we got it wrong.

What counts as an attack

We only store a request when it matches a known attack pattern — a request shaped like a known exploit or scanner tool, a login attempt against SSH, a mail server or an admin panel, or a request to a sensitive path that comes back 401/403/404. Ordinary traffic — a normal page view, a search engine crawler, a visitor browsing the site — is never stored by this pipeline at all, not stored and then ignored: it is never written to the database in the first place.

Every stored request is tagged with two things, from a fixed, published vocabulary (vocab-v1.json in the feeds below):

database, FTP, VPN, an IoT/router endpoint, or one of our own honeypots.

files, exploiting a known CVE, brute-forcing or spraying credentials, credential stuffing, injection, acting as a spam or open-proxy relay, a denial-of-service pattern, a known scanner tool, bot impersonation, an AI crawler ignoring robots rules, or unknown.

How an address gets named

Every address we see traffic from starts recorded — stored, scored, never shown on any public page. An address only becomes public once its behaviour clears a bar meant to rule out one bad request or one mistyped password:

One request, or ten requests in the same minute, is not enough — repetition over time is what moves an address from "recorded" to "listed".

seen from at least two of our separate vantage points, or one high-severity attack (an exploit attempt, not a login guess) repeated at least 10 times over at least 2 days.

reach "blocked" only by trying at least 5 different usernames against the same address (a spray), or by the ordinary 2-active-days rule above — never from one person mistyping their own password many times in a hurry.

Scores decay over time; a quiet address drifts back down and its public page reflects that, but its history and evidence are not deleted early (see Retention, below).

What we publish, and what we never do

Published, once an address is listed or blocked: the address itself, its reverse DNS, its network (ASN) and organisation name, its network type (hosting, cloud, residential, mobile, education, or unknown), its country, Tor/VPN/cloud flags, the request paths and user-agent strings it sent (with anything that could identify a real person already stripped before we ever write the row — see below), timestamps, the attack classes observed, and aggregate honeypot-credential statistics (the top 20 most common passwords submitted across every honeypot hit, shown as a simple counter with no address attached).

Never published, by design:

attacking our sensors and honeypots, never about who visits a site or what a site contains.

identify a real person — these are stripped before the row is ever written to the database, not filtered later when a page renders.

one place a password shows up in clear is the aggregate top-20 counter above, which names no address.

"this address tried a WordPress credential-stuffing tool against several sites" — never a label about who runs it. "Hosts a web server" is a fact we can check; "is probably a compromised home router" or "is likely run by a known group" is a guess about a third party, and we do not publish guesses.

Residential and mobile addresses are treated more carefully

Hosting, cloud and datacenter addresses are very rarely a real person's only connection — a hosting account has an operator, not usually a household. Residential, mobile, education addresses, and any address whose network type we cannot determine (the cautious default), are different: they can be a shared office connection, a home router that got compromised, or an address a provider reassigns to someone new tomorrow. For that reason:

offending, instead of a year.

wanted" ranking**, no matter how high their score climbs — they still get a page if listed or blocked, but it carries noindex and stays out of any ranked table.

address on a hosting network is overwhelmingly likely to still belong to the same operator a year later.

hosting, because cloud IP ranges get recycled between different customers within weeks — a year-long block increasingly punishes whoever rents the address next, not the original attacker.

AI crawlers get a facts-only table, never a score

The board's AI-crawler table (last 30 days) counts requests whose User-Agent claims to be one of the well-known AI-training or search crawlers (GPTBot, ClaudeBot, PerplexityBot, DeepSeekBot, Bytespider, Amazonbot, CCBot and a few others we've actually seen) against three facts, and nothing else:

treats as sensitive (.env, wp-admin, credentials/config filenames, and the rest of classify.php's own sensitive-path list), or somewhere ordinary?

live lookup made when you load this page) land in the domain that crawler's own User-Agent string names as its home (e.g. ClaudeBot's own UA names anthropic.com)? "Not confirmed" means exactly that: the one fact didn't check out. It does not mean the traffic is fake, malicious, or anything else — plenty of real infrastructure sits behind addresses whose rDNS was never set to anything informative.

the request before it ever reached our servers, or did it arrive at the origin (regardless of what HTTP status it then got)?

None of the three facts in this table — the path split, the rDNS check, the edge outcome — feed the scoring in "How an address gets named" above; they are computed only for this table. An address whose traffic happened to carry one of these User-Agents is scored by the exact same rules as any other address (volume, how many of our sensors saw it, auth-failure patterns) — the attack type "ai-crawler" gets no special weight, bonus or exemption in that scoring. The table exists so you can see what we see; it is not a verdict.

Blocking: what it actually stops

A block applies on every one of our hosts' own firewalls and, separately, at our Cloudflare edge. A host-level block only stops a direct connection to that host and stops web traffic to any site on that host that is not behind Cloudflare. It does not stop a blocked address from still reaching a Cloudflare-proxied site's web traffic — only the edge block does that. Our own management SSH port is exempted from every block, on every host, ahead of the drop rule, so a listed address can still reach it and we can never lock ourselves out by running this system.

Retention

DataKept for
Raw attack events90 days
An address's aggregate record (score, tier, history)365 days after its last activity
Up to 50 redacted evidence lines per listed/blocked addressthe aggregate's own lifetime above — these do not vanish at 90 days just because the raw row did
Abuse-report records we sent to hosters2 years
Dispute and removal tickets2 years
Honeypot usernames tied to a specific eventfollow that event's own retention above
Honeypot password hashesfollow that event's own retention above; never shown in clear per address
The top-20 clear-password counterkept indefinitely — it names no address and holds no personal data

Legal basis and your rights

We rely on GDPR Article 6(1)(f), legitimate interest — protecting our own hosts and the sites on them, and giving other small operators a free, signed feed of the same confirmed offenders. We do not profile people; we score addresses and their observed traffic. Because direct notice to an attacking address is not practical, this page serves as the Article 14 notice to anyone whose address appears here. Full legal reasoning, the balancing test, and the controller of record (СЕРВБГ ООД) are in docs/adr/0006-public-threat-listing-policy.md.

If you believe an address was listed in error, or you are the registrant of a listed address and want it removed or want to object to its processing (GDPR Articles 17 and 21), write to [email protected] with the address in the subject line. We review every request by hand and resolve it — by removing the address or by explaining why we did not — within 7 days. A dispute never auto-removes anything while it is open; it only guarantees a human looks at it inside that week.

(That mailbox is being stood up alongside its own sending domain, abuse.servbg.com, as a later build step. Until it is live, write to [email protected] with the same subject-line convention and we will handle it the same way.)

The feeds

We publish the blocked-address list as a set of plain files, updated every 10 minutes:

own blocklist mechanism and Cloudflare's IP-list import both already expect, so no separate CSF-specific or Cloudflare-specific file exists.

the dossier, and an expires field set to that address's own block window (90, 30 or 365 days from the table above — not one flat number for everyone).

which surface they were seen attacking, for an operator who only wants (say) the SSH list.

classes without reading our code.

Every file carries a header comment with its generation time, a stale_after: 7d figure, this page's URL, and the dispute contact above. Every file is signed with minisign; the public key is published at pubkey.minisig next to the feeds, so you can verify a file before trusting it.

Terms: the feeds are free, require no account, and carry no warranty. We sign them so you can verify where they came from; we do not promise they are complete, free of false positives, or fit for any particular firewall. If you consume them, you are responsible for your own firewall configuration and for deciding when a feed is too stale to trust (we recommend flushing your local block set if a feed is older than the stale_after value in its header). A small, dependency-free pull script is provided as a convenience for exactly this.

Running your own sensor

Phase 1 of this project is the estate's own three hosts. A later phase opens an ingest endpoint so another operator can run a small client and contribute events the same way our own collectors do — authenticated by a per-sensor token, rate-limited, versioned, with the client refusing to send anything outside the published vocabulary. That contract, and the first reference client, will be documented here once it ships; an address reported only by a third-party sensor, with no estate-owned vantage point ever confirming it, is never enough on its own to reach "blocked" — we always require at least one of our own sensors to agree.

What this page is not

This page states policy and method for a technical reader. It is not legal advice, and it does not replace docs/adr/0006-public-threat-listing-policy.md, which is the authoritative decision record — if the two ever disagree, the ADR is right and this page is stale and needs fixing.