Simple web crawler/spider

by 11 replies
15
Hi all,

I am just developing a very simple web spider/crawler. Here is the code:

PHP Code:
<?php

$seed 
= "http://www.akosblog.com";
$html = file_get_contents($seed);
echo 
"Page : " . $seed;
preg_match_all("/http:\/\/[^\"\s']+/", $html, $matches, PREG_SET_ORDER);

foreach (
$matches as $val) {
echo 
"<br><font color=red>links :</font> " . $val[0] . "\r\n";


}
?>
This code just gets all the links from the selected page. Now I want to move on, I want the spider to follow links and index another link and another.
So how could I do that?

Regards,
Akos
#programming #crawler or spider #simple #web
  • here is a link to a GPL search script that spiders your site for search terms.

    Take a look at how its doing the spider system.
    Orca PHP Scripts - Camouflaged PHP/MySQL Web Applications
  • Have a look at sphider => sphider.eu/about.php
  • [DELETED]
  • You need to put everything into a recursive function.
    • [1] reply
    • Storing the data in bulk is going to be more of a issue then retrieving it. Check out the NoSQL solutions like Mongo or Couchbase to store as a collection then batch it into a RBDMS is needed. Will give you decent performance without a write locked table
  • Dude I suggest you to use "PHP Simple HTML DOM Parser". It will make your job more easier. You can download and read the documentations from here: simplehtmldom.sourceforge.net
    • [1] reply
  • I also suggest you use an existing spider codebase instead of rolling your own, also ideally the code should be DOM-based instead of using regular expressions. Regex tends to be more brittle to website changes.

    Regarding your actual question about spidering sub-links, what you usually do is initialize a queue to store the URLs. Then you populate the queue with the seed URL. Then you create a loop that pops a URL from the queue, downloads the page, extracts the sub-links from it, and adds those sublinks to the queue (You probably want to add some more sophistication to it like only adding urls on the same domain, and up to a certain depth). You loop over the queue until it's empty. When adding the sublinks to the queue you usually add them to the end of the queue and pop from the front (creates a breadth-first search).
  • I think I'd look into using something like Nutch unless you're doing this for simply educational purposes.
  • You are making a search engine?
    You have to do some mathematics and algorithm study to bring a best solution here. In fact google too hires to mathematicians who develop a faster and economical internet search algorithm for them.
    Study two things -
    1. Mathematics Algorithms - it is a subject of engineering students, if you have a friend in engineering consult him/her.
    2. Data Structures - this is the concept of structure and deals with how to get onto different structures. Internet is also a structural based.

    Thanks.
    • [1] reply
    • python + BS will be much easier than PHP.
  • Generally you will have to decide how much you want to code yourself. Is this for a one-off solution it is probably best to use already well developed solutions created by others. However, if it is a product you want to sell, you may want to write a larger part yourself, so you are not dependent on anything 3rd party. (Of course, depending on if your project is commercial or not, you may have a fairly large amount of open source projects to pick from.)

    Also worth noting is that solutions that work well on 1000 page websites can fail on one million page websites. (And a whole slew of other potential website problems and issues.)

Next Topics on Trending Feed