Most visits to a website aren't made by people. Search engines, AI companies, marketing tools, chat apps and security scanners all send programs (bots, crawlers, spiders) that read pages on their own. Most of them say who they are in the name they give with each visit (the “user agent”), and that name is how Live Traffic sorts them. Here is who they are.
// People
People
You, and anyone else who opens a page in a web browser.
Bots
Visits that looked like someone in a browser but were gone within 4 seconds. Mostly programs driving a real browser through ordinary home internet connections around the world, so they pass for visitors: they open one page, take what they came for and move on. Now and then it will be someone who changed their mind.
// Search engines
Googlebot reads pages so they can turn up in Google Search. Google has other crawlers too, for Google Images, shopping, ads and product research. Whether pages may be used for Google's AI (Gemini) is controlled by a separate robots.txt name, Google-Extended; it is a rule, not a separate visitor.
Worth knowing: Googlebot nowadays reads the web as a smartphone. Since Google's switch to “mobile-first indexing”, the phone version of a page is the one that counts in its search results.
Bing
Microsoft's crawler for Bing Search. Bing's index goes further than Bing itself: it supplies results to other search engines such as Yahoo, and in part DuckDuckGo and Ecosia, and it helps Microsoft's Copilot answer questions about the web.
Worth knowing: Bing and Yandex started the IndexNow system (2021), which lets sites tell search engines straight away when a page changes, instead of waiting for a crawler to notice.
Yandex
The crawler of Yandex, Russia's biggest search engine.
Worth knowing: in Russia Yandex is used for more searches than Google, one of the few countries where Google isn't the leader.
Baidu
The crawler of Baidu, China's largest search engine. Google Search isn't available in mainland China, so Baidu is the main way people there search the web.
Worth knowing: Baidu also makes its own AI chatbot, Ernie.
DuckDuckGo
The crawler of DuckDuckGo, the search engine that doesn't track its users. Much of what it shows comes from Bing, but it also crawls some pages itself, and fetches pages for its AI answers (DuckAssist).
Huawei (Petal)
The crawler of Petal Search, Huawei's search engine. Huawei built it after US restrictions meant its new phones could no longer come with Google's apps and services.
Seznam
The crawler of Seznam, a Czech search engine and web portal.
Worth knowing: Seznam is one of the very few home-grown search engines that has held its own against Google in its own country for decades.
Qwant
The crawler of Qwant, a French search engine that promises not to track or profile its users.
Mojeek
The crawler of Mojeek, a small independent search engine from Brighton, England, that has been building its own index of the web since 2004 rather than borrowing Google's or Bing's.
// AI companies
OpenAI
The company behind ChatGPT has three visitors. GPTBot collects pages that may be used to train its AI models. OAI-SearchBot indexes pages for ChatGPT's search. ChatGPT-User fetches a page when someone asks ChatGPT about it, so that visit is prompted by a person in a chat.
Worth knowing: OpenAI announced GPTBot in August 2023 along with a way to block it in robots.txt, and within weeks many large news sites had done so.
Anthropic
The company behind the Claude AI models. ClaudeBot collects pages that may be used to train its models; Claude-User fetches a page when someone asks Claude to look at it; Claude-SearchBot helps Claude search the web. Anthropic says its crawlers follow robots.txt.
Worth knowing: in 2024 the repair site iFixit said ClaudeBot had visited it about a million times in a single day, one of several complaints that year about how hard AI crawlers were hitting websites.
Meta
Facebook, Instagram and WhatsApp's parent company. Its most common visitor, facebookexternalhit (and WhatsApp's), fetches a page to build the little preview card when someone shares a link, so a visit from it usually means someone just posted or sent a link to this site. meta-externalagent is Meta's crawler for training its AI models (Llama and Meta AI).
Worth knowing: in March 2026 Meta bought Moltbook, a social network where only AI agents post, and brought its founders into its AI research lab, Meta Superintelligence Labs.
Apple
Apple's crawler, used for Siri and Spotlight suggestions and Safari's search suggestions. Whether pages may also be used to train Apple's AI (Apple Intelligence) is set with a separate robots.txt name, Applebot-Extended.
Amazon
Amazon's crawler, which Amazon says is used to improve its products and services, such as helping Alexa answer questions.
ByteDance
ByteDance is the Chinese company that owns TikTok. Its crawler, Bytespider, is widely thought to gather material for ByteDance's AI models (such as its Doubao chatbot) and search.
Worth knowing: in 2024 the web-security company Cloudflare reported Bytespider as the busiest AI crawler it saw across the sites it protects, and Fortune reported that it ignores robots.txt. Research by TollBit covering the first half of 2026 found it was one of three bots that reached blocked pages on close to half of the European sites that named them, alongside OpenAI's ChatGPT‑User and Youbot.
But: ByteDance publishes nothing about Bytespider, not even a way to check that a visit really is its crawler. So any program can call itself Bytespider, and studies that count by name may blame ByteDance for impostors. Some of the ByteDance visits here may not be ByteDance at all.
Perplexity
Perplexity is an AI “answer engine”: it searches the web and writes an answer with links. PerplexityBot indexes pages; Perplexity-User fetches a page when someone's question needs it.
Worth knowing: in 2024 and 2025 journalists and Cloudflare said Perplexity also read pages that had blocked it, using visitors that didn't give its name. Perplexity disputed this.
Common Crawl
A non-profit that has been crawling the web since 2008 and gives away the results: a free archive of billions of web pages, used by researchers, and by almost every company building AI language models.
Worth knowing: filtered Common Crawl data made up about 60% of the training mix for GPT-3, the model that came before ChatGPT, and a 2024 Mozilla Foundation study concluded that generative AI in its current form would probably not be possible without it. Common Crawl has taken donations from AI companies, OpenAI and Anthropic among them.
Disputed: in November 2025 an investigation by The Atlantic said the archive holds millions of paywalled articles from outlets such as the New York Times, the Economist and the Atlantic itself, because its crawler never runs the code that puts the paywall up, and that articles publishers asked to have removed were hidden from Common Crawl's search tool rather than taken out of the archive. Its director, Rich Skrenta, told the magazine “The robots are people too.” Common Crawl called the article's claims false and misleading.
Cohere
A Canadian AI company that makes language models for businesses.
Worth knowing: one of its founders, Aidan Gomez, co-wrote “Attention Is All You Need” (2017), the paper by a team at Google that introduced the transformer. It became the cornerstone of today's large language models: the “T” in GPT (as in ChatGPT and GPT-3.5) stands for transformer, and Claude, Gemini and Llama are built on it too.
// Marketing and search-ranking tools
Semrush
Semrush sells tools for online marketing. Its crawler maps which sites link to which, so that customers can study how other sites rank in search results, including their competitors'.
Worth knowing: since 28 April 2026 Semrush belongs to Adobe, which bought it for about $1.9 billion, so this crawler is now Adobe's.
Ahrefs
Ahrefs, based in Singapore, sells tools for studying links and search rankings, built on its own huge crawl of the web.
Worth knowing: Ahrefs says AhrefsBot is one of the busiest crawlers on the web after Googlebot, and it has used its index to start its own search engine, Yep.
Moz
Moz, from Seattle, makes search-ranking tools. DotBot builds its index of links between sites; rogerbot checks sites for Moz's customers.
Worth knowing: Moz invented “Domain Authority”, a score many marketers use to guess how well a site will rank, and rogerbot is named after Moz's robot mascot, Roger.
Majestic
Majestic, a British company, keeps a map of the links between websites for marketers.
Worth knowing: its crawler grew out of a distributed project, Majestic-12, in which volunteers ran the crawler on their own computers, so its visits can come from all over the world.
DataForSEO
DataForSEO sells search and link data to other marketing tools, which build their own products on top of it.
Serpstat
Serpstat is a search-marketing platform, started in Ukraine, whose crawler gathers link and ranking data for its customers.
Screaming Frog
A program made by a British search-marketing agency, which someone runs on their own computer to check every page of a site for problems. A visit from it usually means a person somewhere is looking closely at this site.
BLEXBot
The crawler of WebMeUp, which gathers data about links between sites for search-marketing tools.
Barkrowler
The crawler of Babbar, a French company that maps links between sites for search-marketing tools.
// Link previews
X (Twitter)
Fetches a page to make the preview card when someone posts a link on X. A visit usually means someone just shared this site there.
Fetches a page to make the preview when someone shares a link on LinkedIn.
Visits pages whose pictures have been pinned, to keep the pins' details up to date.
Slack
When someone pastes a link into a Slack chat, Slack fetches the page to show a preview (“unfurling”). A visit means the link was just shared in a workplace chat.
Discord
Fetches a page to show its preview when someone posts a link in a Discord chat.
Telegram
Fetches a page to show its preview when someone sends a link in Telegram.
// Security scanners
Security scanners
Programs that sweep the whole internet asking every site for files that are often left open by mistake on other kinds of sites: password files such as .env, WordPress login pages, admin panels. Some are run by researchers and security companies, some by people looking for weak sites to break into.
This site is plain pages with no login, database or admin area, so they find nothing here. Any visit asking for one of those addresses is counted here, whatever name it gives.
Censys
Censys, from Michigan in the US, keeps checking every address online and catalogues what it finds (devices, websites, security certificates) for security teams and researchers. It grew out of ZMap, free scanning tools built at the University of Michigan from 2013.
Worth knowing: ZMap can scan every public address on the internet (the whole of IPv4) in under 45 minutes from a single computer. Censys co-founder J. Alex Halderman compares it to Google Street View: “we just take a picture from the sidewalk. We don't peek in the door, we don't jiggle the locks.”
Palo Alto Networks
A large US security company. Its Cortex Xpanse service scans the internet to find everything an organisation has online, including forgotten or unsafe systems, so they can be fixed before someone else finds them.
// Programs without a company
curl
A free command-line tool for fetching web addresses, used by programmers and in countless scripts. A visit from it is a person's program rather than a browser: a developer checking the site, a script, or a scanner.
Worth knowing: curl was started in 1998 by Daniel Stenberg in Sweden, who still leads it, and it is built into cars, TVs, phones and games consoles: its makers estimate it is installed over twenty billion times.
A sign of the times: in January 2026 Stenberg shut down curl's bug bounty, which paid people for finding security flaws, because it was flooded with fake reports written by AI. Over almost seven years it had found 87 real vulnerabilities and paid out over $100,000, but by 2025 fewer than one report in twenty turned out to be real.
python-requests
Requests is a popular add-on for the Python programming language that fetches web pages. A visit from it is someone's script: a data project, a scraper or a bot someone wrote themselves.
Go-http-client
Go is a programming language made at Google and released in 2009. It is good at doing many network jobs at once, so it is widely used for servers and internet tools: Docker and Kubernetes, which run much of the world's cloud computing, are written in it.
“Go-http-client” is the name Go gives a program automatically when it fetches a web page. To give it any other name, the programmer has to set one on purpose. Many don't: a quick script someone runs once doesn't need one, and a program that visits millions of sites, such as a scanner, may prefer not to say whose it is. So a visit under this name is a program written in Go, and nothing more is known about it.
Wget
GNU Wget, a free tool from 1996 for downloading files from the command line. It comes ready installed on many Linux computers and servers, it can pick up a download where it broke off, and it can follow every link on a site to save a copy of the whole thing.
That makes it the usual choice for keeping copies of websites: someone saving a site to read offline or as a backup, a server script fetching a file every night, or archivists saving sites before they disappear.
Worth knowing: Archive Team, the volunteer group that rescues websites before they shut down (GeoCities in 2009 among them), has long used its own version of Wget to do it, sending much of what it saves to the Internet Archive.
HeadlessChrome
Google Chrome running without a window, driven by a program: used for automated testing, taking screenshots of pages, and scraping.
Scrapy
A free Python toolkit for building web scrapers, programs that collect information from websites.
Other named bots
Bots that give a name of their own but aren't from any company listed here. Live Traffic shows them by the name they give. Many are small crawlers for search or marketing start-ups, research projects, or website monitoring services.
Unnamed programs using a browser name
Visits that say they come from an ordinary web browser, like Chrome or Safari, but don't behave like one: they leave out things every browser sends, or open pages faster than anyone could read them. They are most likely programs pretending to be browsers so that sites treat them like people, for example scrapers, price checkers, or crawlers that would rather not be named. Now and then it may be an unusual or old browser.
Unnamed bots
Visits that admit to being a program but don't say whose, or that were counted before Live Traffic started noting bots' names.
Link previews
Other apps fetching a page to preview a shared link.
Worth knowing: on Mastodon every server fetches its own preview, so one post with a link can bring a sudden burst of visits from hundreds of servers at once.