How to Block Some of the Bots

(nochan.net)

27 points | by Bender 1 hour ago

6 comments

  • fxtentacle 1 hour ago
    I like the idea of adding a fake cpanel subdomain for 169.254.169.254 so that script kiddies will start port-scanning their own hosting provider, which will likely get them flagged/banned.
    • Bender 1 hour ago
      When I first experimented with that I was not expecting anything to happen. Within a few days one person in Amazon EC2 in Germany started trying to do zone transfers for some of my domains likely to figure out which records to avoid and then they just excluded my domains entirely. All of the scanning stopped shortly thereafter. The scanning noise was literally all coming from one person despite the source IP's being all over the internet.
    • binaryturtle 1 hour ago
      410 and a "Sec-Fetch-Mode:" string in the response body. I guess it thinks I'm a bot? Thanks!

      Nothing to read, nothing to see, I move along. (Yikes, the modern web sucks!)

      • Bender 1 hour ago
        Real browsers send it. [1] Some of the reader apps do not. Most bots aside from those utilizing Chrome Headless do not.

        People can see a few headers here [2]

        [1] - https://caniuse.com/?search=sec-fetch-mode

        [2] - https://nochan.net/.env

        • gruez 53 minutes ago
          I didn't even make it that far, I got PR_END_OF_FILE_ERROR which indicates it couldn't even get past the tls handshake.
          • ErroneousBosh 45 minutes ago
            I mean that just might be the dying gasp of a puddle of scorched fibreglass and boiling copper where the server used to be.
            • Bender 43 minutes ago
              Mostly harmless(c)

                  load average: 0.00, 0.01, 0.00
        • ajsnigrutin 34 minutes ago
          There's a special place in hell for people who block curl and wget, especially on sites with downloadable files (eg source code tgz's, media, etc.), basically anything i might need to wget on a server.
          • Bender 26 minutes ago
            A guy gets sent to hell. The demons guide him into the lobby. Satan jumps out and lets out a big evil roar to no effect. Satan then glances down and sees the spot on the mans finger where there was a ring. "Oh... well you've been through worse. The break room is down the hallway to the left, mail room is upstairs to the right..." Satan just walks away.

            Every step is optional of course. There are ways to make curl work on my site but I choose to add friction as some botters abuse libcurl. I doubt anyone else will do what I do. They could spoof the user-agent but most botters seem to not know how to do that despite the myth that they all do. For what it's worth if I had a site specific to sharing code artifacts or archives I would not block curl and I would also enable native rsyncd.

          • iririririr 53 minutes ago
            most (all?) of those will 100% block valid traffic too
            • Bender 44 minutes ago
              I would be interested in some examples of valid traffic that they would block. For the purposes of the document I do not consider anything from a data-center to be valid traffic. People can of course skip any or all steps, experiment with one at a time on a test server as they should.
              • ferngodfather 29 minutes ago
                > Obvious proxy is obvious.

                > if ($http_x_forwarded_for) {....

                This may block schools and libraries that use content blockers. Often the internal client is left to make abuse tracking easier (or because the overworked admin didn't know they could turn it off).

                • Bender 22 minutes ago
                  This may block schools and libraries

                  Oh, well that is ok for me I suppose. I add RTA/adult headers that hopefully they also look for and block using parental controls as adult content should not be viewed in a school or library. I could add a note suggesting to skip that step if one wishes schools and libraries that may be using a proxy to view.

                  • ferngodfather 12 minutes ago
                    That's fair enough.

                    I think it's a great article to be fair. We need more of this cheap and quick bot blocking. The fact the solution to unwanted traffic is often "use Cloudflare" is _not_ great for the internet, and nobody really actually likes deploying or managing ModSecurity. Its a nice middleground.

            • Capricorn2481 26 minutes ago
              Am I the only one that exclusively gets attacks with spoofed user agents and rotating TLS signatures? I feel like every post I see about not needing a CDN has tips that could be overcome in under an hour of scripting.
              • receptopalak 3 minutes ago
                [flagged]