Search engines keep their listings current and show the most accurate results by performing an action known as a “crawl.” This means sending out a “bot” — sometimes called a “spider” — to move through the internet and look for new pages, updated pages, or pages the search engine did not previously know existed. After the crawl, the search engine results page can be updated to include the pages found during that process. In simple terms, it is a way for search engines to discover sites and pages online.
There may be times, however, when you have a page on your website that you do not want to appear in search engine results. For example, you might still be building a page and prefer that it not be listed until it is finished. In these cases, you can use a file called robots.txt to tell search engine bots to ignore specific pages on your website.
Robots.txt is essentially a way of telling a search engine, “don’t come in here, please.” When a bot finds a robots.txt file, it will read it and should ignore the URLs listed within it. As a result, those pages are not included in search results. It is not a failsafe, though. Robots.txt is a request for bots to ignore a page rather than a complete block, but most bots will follow the instructions in the file. When you are ready for the page to be included in search engines, you simply modify your robots.txt file and remove the URL of that page.
Speak Your Mind