Bot directory · Utility & developer tool

Java HTTP

Default UAs from Java's built-in HTTP clients, typically enterprise software that didn't identify itself.

Operator
CategoryUtility & developer tool
Respects robots.txtNo - documented as not applying
Executes JavaScriptNo - reads raw HTML
robots.txt tokenJava
User-agent stringJava/17.0.2

Official documentation: docs.oracle.com

See what Java HTTP sees on your site → try Peekabot, free

What Java HTTP does

The bare Java/x.y.z UA comes from Java's old HttpURLConnection; Java-http-client/x.y.z is the newer built-in HttpClient. Java runs a huge share of enterprise back ends, so these requests are often integration middleware, feed readers, or security scanners embedded in corporate software. Neither client knows anything about robots.txt or JavaScript.

Why it visits your site

Some Java application fetched your page and its developer never set a UA. The version number tells you which Java runtime, which is not much to go on.

The bigger picture: utility & developer tools

This category is the honest miscellany of a server log: monitoring services, platform callbacks, and the default user agents of HTTP libraries. The unifying rule is that the string identifies software, not intent. A curl request might be your host's health check or a vulnerability probe; a python-requests hit might be a researcher or a scraper; a WordPress UA is one site talking to another. Reading these names as a threat list misses what they actually are - a census of the scripted web touching your site, most of it mundane, some of it yours.

robots.txt barely applies here, and not out of rudeness: monitors fetch the exact URL a customer configured, feed fetchers act on a user's subscription, and bare libraries simply ship with no robots logic at all. Blocking a default library UA filters the honest and the lazy while the malicious change one line and continue, so the effective levers are different - rate limiting by IP, watching which paths get requested, and reserving blocks for named services you can verify and have decided you don't want.

Should you block it?

Old vulnerability scanners do love a bare Java UA, so some firewalls flag it - but so do legitimate enterprise feed fetchers and integrations. Blocking the string stops neither a determined scanner (one line to change) nor tells you who was knocking. Judge by the URLs requested; a Java client probing wp-login.php deserves an IP block, not a UA rule.

Block via robots.txt

Add these lines to the robots.txt at your site root to ask Java HTTP to stay away:

User-agent: Java
Disallow: /

Remember that robots.txt is a request, not a lock - compliance is voluntary, and enforcement happens at the server. Viz's AI Control pairs robots.txt rules with server-level 403 responses and then probes your site to verify the block is actually holding.

How to see Java HTTP traffic on your site

You have three windows onto it, each with a catch. Raw hosting access logs contain every Java HTTP request, but many managed WordPress hosts don't expose them, and when they do you are grepping text files by hand. A CDN dashboard sees the traffic at the edge, but it lives outside WordPress and speaks in totals, not in your site's terms. And JavaScript analytics - GA4 and friends - will never show it at all, because Java HTTP doesn't run scripts.

Viz's answer is the request log: filter to Java HTTP and read exactly which URLs it fetched and when, right inside wp-admin - then watch the same filter after any blocking decision to see whether the visits stopped, slowed, or kept coming.

Measure first

Is Java HTTP on your site right now? Find out in two minutes.