| Subject: Re: How to completely ignore 'robots.txt'? |
Author: Vladimir |
Date: 10/03/2010 02:01 |
| | > > Ok, I'm trying to mirror a site that tells engines
> > like httrack to not go down to certain directories.
> > Which version of httrack allows me to complete
> > ignore
> > these files and go down into this certain directory?>
> See options/spider/spider: robots.txt -> 'never'
>
> But also ensure that you set proper bandwidth limiter
> if you are crawling big files or a large number of
> generated pages (robots.txt are often used to avoid
> server overload)
>
The option has changed somewhat, over the years..:)
img843.imageshack.us/img843/8703/103201014518am.png | |
|
|
|
|
7
Created with FORUM 2.0.11