Rendered at 19:43:00 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
arcfour 1 days ago [-]
> some of the bots won't respect robots.txt
The link you posted literally only lists the scanner related to security/safe browsing. Why would it? What kind of bot would that be if an attacker could just put up a robots.txt that causes it to ignore the site?! And many do...for obvious reasons.
It's misleading because this implies you're getting heavy scraper traffic from Google bots that don't respect robots.txt for reasons of greed rather than because it's necessary so they can proactively avoid surfacing malicious websites.
ArcHound 1 days ago [-]
I was thinking about this line a lot. Because yes, "not all bots" and "not in all cases".
But what gives them the right to do whatever with the sites if they claim it's for security purposes? They are not law enforcement. Moreover, other bots are missing from the docs which they clearly state on the same page. How many other bots of theirs are ignoring robots.txt?
I think we should hold Google to a higher standard than some random blogger (me).
arcfour 6 hours ago [-]
...Because law enforcement does not have the interest, reach, or time to investigate every single cybercrime, and if they do it won't be timely - it will be after the fact, punishing the criminals, and not preventing harm.
Google is in a unique situation to protect their users, and they also want to avoid serving search results that are malicious (yes, I know about their ads problem) - so they are proactive about scanning for sites that are malicious so they can avoid sending users to them.
ArcHound 2 days ago [-]
I keep stumbling over different Google bots in my logs. So today I'm looking at each and every Googly eye that takes a look at my site. Because when you stare at me, I'll stare back.
You have the right and the means to define what Google is allowed to do on your site, whether the Google traffic you face is extensive or you'd like to opt-out of AI training (check Google Other). Sadly, using robots.txt is not enough, Google admits that some of their bots will not respect the rules.
The link you posted literally only lists the scanner related to security/safe browsing. Why would it? What kind of bot would that be if an attacker could just put up a robots.txt that causes it to ignore the site?! And many do...for obvious reasons.
It's misleading because this implies you're getting heavy scraper traffic from Google bots that don't respect robots.txt for reasons of greed rather than because it's necessary so they can proactively avoid surfacing malicious websites.
But what gives them the right to do whatever with the sites if they claim it's for security purposes? They are not law enforcement. Moreover, other bots are missing from the docs which they clearly state on the same page. How many other bots of theirs are ignoring robots.txt?
I think we should hold Google to a higher standard than some random blogger (me).
Google is in a unique situation to protect their users, and they also want to avoid serving search results that are malicious (yes, I know about their ads problem) - so they are proactive about scanning for sites that are malicious so they can avoid sending users to them.
You have the right and the means to define what Google is allowed to do on your site, whether the Google traffic you face is extensive or you'd like to opt-out of AI training (check Google Other). Sadly, using robots.txt is not enough, Google admits that some of their bots will not respect the rules.