Thursday, June 19, 2008

Optimizing for MSN

The Advent of MSN as a Search Engine

Microsoft had been napping for a long time and ignored the advancements in the field of Search Engine and Content Targeted Advertising. Although now dependent on Yahoo's Intokmi for their search results, Microsoft has made it very clear that they will compete with Yahoo and Google for their share in the Search Engine market. Given Microsoft's aggressive nature in fighting competition, it would be a grave mistake to underestimate them.

The recently re-launched MSN Search and future MSN Search integration with upcoming versions of Windows is about to make MSN one of the biggest and most important players in the world of searching. Thus, it is imperative to get good ranking in MSN if you want the share of traffic they can give to your Web page. Although MSN search spider does a fairly good job in crawling Web pages, you may benefit by submitting your website at http://search.msn.com/docs/submit.aspx .

Optimizing for MSN Search

With Microsoft sharing Yahoo's Inktomi search index to provide their search results, optimizing for yahoo meant optimizing for MSN. But things are changing at a rapid phase and with Microsoft getting active on the patent front, it is evident that they are working on their own search algorithm.

Luckily for us, the rules of Web page optimization that thought to be followed to please the MSN search algorithm aren't very different when compared to those already followed for other search Engines.

What They Lay Emphasis On?

As with most other Search Engines, MSN Search places heavy emphasis on content. They even allow higher keyword density than Google does. For MSN Search, it is best to keep your pages at least 200 words long and have phrases which searchers commonly use. Other than that, they lay importance in the following in the order they are listed.

• As MSN team declares in their blog that, they attach a lot of importance to the number and quality of sites that link to your pages.

• Clean coding is necessary with MSN Search. They even go to the extent of asking Webmasters to ensure that their pages are HTML validated. MSN's spider has a strong preference for well-written code. If a Website's coding is poorly written, it appears that MSN Search downgrades the site's search rankings heavily.

• A well-designed site map with good link text will help the MSN spider to crawl the site and ensure that all pages are indexed.

• Title tag should be less than 80 characters long and should be attractive enough to make a searcher click on the link.

• MSN Search doesn't rank based on Meta Keywords and Description, but it seems to place some importance on meta tags. So adding appropriate meta tags for each page might be beneficial as well.

• MSN Search recommends that an HTML page with no pictures should be under 150 KB. Therefore, ensure that you limit the size of your Web pages to a reasonable limit.

What MSN Doesn't Like?

MSN search lists the following as being search engine unfriendly due to the difficulty search engine robots have with this type of content:
• Frames
• Flash
• JavaScript navigation
• HTML Image Maps
• Dynamic URL's

Techniques not liked by MSN Search

MSN thinks the following to be unscrupulous SEO practices:
• Loading pages with irrelevant words in an attempt to increase a page's keyword density, this includes stuffing ALT tags that users are unlikely to view.
• Using hidden text or links. You should use only text and links that are visible to users.
• Using techniques to artificially increase the number of links to your page, such as link farms.

As you can see, these "rules" are no different from those mentioned by the rest of the industry. So avoid the above-mentioned techniques and the chances of your getting banned by any search engine are remote. On a related note, this is what MSN search's Program Manager, Eytan Seidman, has to say about spamming MSN.

"You crawled my site, so why can't I find it in your search index? This is one is a little bit easier. The reason that this is most likely happening is that we are detecting the page as spam when we analyze the page to build our index. How can you make sure that this does not happen? The best thing to do is to not spam us. On our site owners help, we talk about some of the things that we consider spam. In case you have not read it, here is a quick refresher: dirty javascript redirects, stuffing alt text, white on white links, off topic links etc. We take this stuff very seriously and we are continuously working to improve our spam detection."

Conclusion

With the increasing popularity of MSN Search and with Microsoft planning to make the search a part of their next windows release, your efforts to optimize your site for MSN are sure to pay off. For more details on optimization for MSN Search, read their help document and blog.

Hotel Internert Marketing by Gatesix

Optimization for Yahoo

Why Optimize for Yahoo?

According to a recent study, of all the searches done through search engines, around 25% of searches are done through Yahoo. That means, if your site is not coming up on Yahoo, you lose 25 percent of your potential visitors. Until February 2004, Yahoo used Google results. So, optimizing for Google was enough to get a top rank on Yahoo. As of 17th February 2004, Yahoo dropped Google results and instead showed search results using Inktomi algorithm. Yahoo's shift from Google to Inktomi made optimizing for Yahoo inevitable.

The New Yahoo Search Engine

Although Inktomi / Yahoo search algorithm doesn't differ too much from that of Google's, it is not exactly a clone of Google's algorithm. Based on the search results on Yahoo, it seems Yahoo's new algorithm gives much importance to keyword density in body text, Title tag, and META tags and to inbound links. Therefore, concentrating on these two items will definitely increase your site's ranking on Yahoo.

Keyword Density

The new Yahoo search engine gives more importance to keyword density. A Website with high keyword density may fare well in Yahoo. The average keyword counts that seem to work are as follows:

  • Title Tag - 15% to 20%: Yahoo displays the title tag content in its result page. Therefore, write the title as a readable sentence. A catchy title will attract the reader to come to your Website.
  • Body Text - 3%: Boldfacing the keywords sometimes boosts the page's ranking. But, be careful not to be awkward to the readers. Too much of boldfaced content irritates readers.
  • META Tags - 3%: In META description and Keyword tags, provide important keywords at the beginning. Do not use the keywords repeatedly in the Keyword tag, because Yahoo may consider it Spamming. Write the description tag as a readable sentence.
Inbound Links/Back Links

Yahoo considers inbound links highly important. Inbound links are the links from other sites pointing to your site. Having considerable links, with appropriate link texts, from quality sites increases your site's ranking in Yahoo.

Static Pages versus Dynamic Pages

Like most other search engines Yahoo prefers static pages to dynamic pages. Sometimes, Yahoo may fail to index dynamic pages. Therefore, consider the following tips to ensure that Yahoo indexes all your web pages:

  • Have static pages with keyword-rich content; it increases the rank of your site on Yahoo
  • If you have some important dynamic pages, prepare a site map or quick links section with links to all the Web pages. This would help the Yahoo spider to craw all your pages.

Frames

Most of the search engines, Yahoo in particular, hate frames. Avoid using frames on your site, because Yahoo spider finds it difficult to crawl them.

The sure-shot solution to rank high on yahoo is simply getting plenty of back links from quality sites, and then having copious keywords in the body text, title, META, and alt tags. As Yahoo holds the 25% share of the Internet searches, it is prudent to have your Web pages optimized for Yahoo. When your Web page ranks high on Yahoo, you get additional traffic that will convert into increased sales.

Vincent S Brown - 1st September 2005

Internet Marketing by Gatesix

Tuesday, June 17, 2008

Podcasting and SEO

Podcasting and SEO: How to SEO your podcasts

There has been plenty of discussion in the blogosphere about blogs and search engine optimization (SEO). Google in particular seems to love blogs. Blogs are rich in content, heavily linked, with links that tend to be contextual, and without much in the way of code bloat or gratuitous flash animation. In short, blogs are search engine friendly out-of-the-box.

But what about SEO’ing a podcast, the blog’s newest cousin?

Podcasting (where anyone can become an Internet radio talk show host or DJ) presents unique opportunities to the marketer/content producer that blogging does not. I expound on this a bit more in my recent MarketingProfs article but the benefits of podcasting from an SEO standpoint wouldn’t seem as obvious. Podcasts are usually audio content, so you don’t get all this rich textual content that the search engine spiders can snarf up. You also don’t get the rich inter-linking that happens with blogs because you can’t embed clickable URLs throughout your MP3 files.

Nonetheless, I believe you can SEO your podcasts. Here’s how:

  • Come up with a name for your podcast show that is rich with relevant heavily searched-on keywords.
  • Make sure your MP3 files have really good ID3 tags — rich with relevant keywords. ID3V2 even supports comment and URL fields. The major search engines may not pick up the ID3 tags now, but they will! And besides, there are specialty engines and software tools that already do.
  • Synopsize each podcast show in text and blog that. Put your most important keywords as high up in the blog post as possible but still keep it readable and interesting.
  • Encourage those who link directly to your MP3 file to also link to your blog post about the podcast.
  • Consider using a transcription service to transcribe your podcast or at least excerpts of it for use as search engine fodder. Break the transcript up into sections. Make sure each section is on a separate web page and each separate web page has a great keyword-rich title relating to that segment of the podcast. And, of course, link to the podcast MP3 from those web pages. There are many transcription services out there, where you can just email them the MP3 file or give them an URL and they send you back a Word document. Here’s a partial list of transcription services.
  • Submit your podcast site to podcast directories and search engines such as audio.weblogs.com.
  • Let people in your industry, such as bloggers and the media, know that you have a podcast because podcasting is quite new and novel. It will be more newsworthy and link worthy than just another blog in your industry.
  • Don’t just get up on your soapbox. Have conversations with others, in the form of recorded phone interviews, and podcast those as well. Pick people who have great reputations on the web and great Page Rank scores, and ask that they link to your site and to your podcast summary page.

This isn’t meant to be a comprehensive list of tactics. It is simply meant as a catalyst for creative thinking. SEO, in particular the link building aspect, isn’t about just following a set list of formulae. It is about creatively thinking outside the box and differentiating yourself in ways that make your site eminently more links worthy than your competitors.

Search Engine Optimization for Podcasts

By Grant Crowell | March 9, 2006 Podcasting is comparatively new, though there are already numerous podcast search engines and it's important to optimize your audio files if you want listeners to find your spoken content.

A special report from the Search Engine Strategies conference, December 5-8, 2005, Chicago, IL.

Podcasting—recording an audio or video file and uploading it to the web so that users with iPods or other media players can download the content—is a hot subject. Panelists on this session focused on how best to prepare and optimize podcasts for search engines.

Podcasting and search

"Podcasting is an interesting challenge from a search standpoint," said Joe Hayashi, Senior Director of Product Management at Yahoo "It is not only audio, it's also video. It's also a subset of audio—it is meant to be consumed in a particular way. Podcasts are a subset of multimedia, and the techniques to really find a podcast need to scale across multiple domains."

In some ways, podcast search engines are similar to traditional search engines except that podcast search engines crawl the Web constantly for rich media files. "If we come across things like podcasts or any other audio or video file," said Suranga Chandratillake, Co-Founder and CTO of Blinkx, "we ingest those into our index and allow people for people to search for that content on either our own site or thru various syndication partners."

"Most people are still wondering what a podcast is and have trouble not only finding it," he said. "So we put a lot of energy into not only the search (finding) aspects but the consumption aspects as well. We have done a variety of things—search, editorial, a browsing system and a tagging system for podcasting."

"We're really leveraging the community out there to provide great content to people," Hayashi continued. "We provide a lot of community tools: a tagging system, a ratings and review system—this lets us discover high quality content. The tagging system also influences search results."

Metadata and Podcasts.

In the past, many multimedia search engines relied heavily on metadata to determine relevancy. Now these search engines are able to utilize speech recognition to determine the content of an audio file.

"Podscope is the first podcast search engine that actually looks for and listens to every spoken word in a podcast," said David Ives, President and CEO of TVEyes. "We believe that speech recognition and actually cracking open the audio file is essential for finding relevant podcasters. We have a solution called 'pinpoint audio' which enables us to play an audio snippet to determine the relevancy of that term within a podcast."

"Metadata alone is not a sufficient indexing criteria to find relevant podcasts," he said.
"I also agree that just metadata is not enough, said Chandratillake.”The average podcast today, which is about 15-20 minutes long, only has 25-30 words describing it. There is no way that short description contains everything that is in the 'meat' of podcast. That is also why we use speech recognition to understand more completely what it’s about."

Podcast optimization tips and guidelines

Speakers offered the following tips and guidelines for optimizing podcasts:

  • Promote only one feed. "Many podcasters create a podcast, then move over to a different content management system, promote a new RSS feed, and wind up with all of these different feeds out there for every podcast," said Dick Costolo, CEO of Feedburner. "You want your content to be easily discovered. Promoting one feed makes it easy for search engines to know where your content is."
  • Optimize the audio file. A lot of people listen to a podcast on their computer as well as MP3 players.
  • Close the findability gap. "Optimize a landing page for each episode of your show, as well as your category page," said Amanda Watlington, owner of Searching for Profit. "Provide subscription information on the landing pages that's very visible."
  • Build correct and valid feeds. "Validate your feeds with feed validator tools," said Watlington. "Remember that iTunes does not redistribute. So you must build a separate feed for iTunes. I like to promote doing 3 separate feeds: a 2.0 feed, a media feed and an iTunes feed."
  • Include a transcript or summary. Whether or not it is a transcript or a summary will depend on the podcast's time span. "If you're giving just a little short tip, that's one thing," said Watlington. "Typically, a summary is all you need for your landing page, a nicely optimized page that covers the podcast's high points."
For marketers, more of your focus needs to be on the development and findability side, not gadget seduction. "Nobody is going to listen to the podcast no matter how elegant it may seem," Watlington concluded. "Focus on findability, focus on quality content and engaging the user. Focus on something people will want to listen to."

Page Rank Algorithm

Few Important Points to Note about PageRank

The first points we may notice from theses mathematical explanations about PageRank are so important that we prefer telling them here.

The PageRank of one page B only depend on 3 factors:

• the number of pages Ak linking to B,
• the PageRank of each page Ak,
• the number of outward links of each page Ak.

So it does not depend on the following criteria:

• the traffic of the sites linking to B,
• the number of clicks on the links to B within the pages Ak,
• the number of clicks on the links to B within the results pages in Google.

These points having been mentioned, let's go on with a question that should interest a lot of webmasters

The Page Rank Algorithm


The original Page Rank algorithm was described by Lawrence Page and Sergey Brin in several publications. It is given by

PR(A) = (1-d) + d (PR(T1)/C(T1) + ... + PR(Tn)/C(Tn))

where

PR(A) is the Page Rank of page A,
PR(Ti) is the Page Rank of pages Ti which link to page A,
C(Ti) is the number of outbound links on page Ti and
d is a damping factor which can be set between 0 and 1.

So, first of all, we see that Page Rank does not rank web sites as a whole, but is determined for each page individually. Further, the Page Rank of page A is recursively defined by the Page Ranks of those pages which link to page A.

The Page Rank of pages Ti which link to page A does not influence the Page Rank of page A uniformly. Within the Page Rank algorithm, the Page Rank of a page T is always weighted by the number of outbound links C(T) on page T. This means that the more outbound links a page T has, the less will page A benefit from a link to it on page T.

The weighted Page Rank of pages Ti is then added up. The outcome of this is that an additional inbound link for page A will always increase page A's Page Rank.

Finally, the sum of the weighted Page Ranks of all pages Ti is multiplied with a damping factor d which can be set between 0 and 1. Thereby, the extend of Page Rank benefit for a page by another page linking to it is reduced.

The Random Surfer Model

In their publications, Lawrence Page and Sergey Brin give a very simple intuitive justification for the Page Rank algorithm. They consider Page Rank as a model of user behaviour, where a surfer clicks on links at random with no regard towards content.

The random surfer visits a web page with a certain probability which derives from the page's Page Rank. The probability that the random surfer clicks on one link is solely given by the number of links on that page. This is why one page's Page Rank is not completely passed on to a page it links to, but is devided by the number of links on the page.

So, the probability for the random surfer reaching one page is the sum of probabilities for the random surfer following links to this page. Now, this probability is reduced by the damping factor d. The justification within the Random Surfer Model, therefore, is that the surfer does not click on an infinite number of links, but gets bored sometimes and jumps to another page at random.

The probability for the random surfer not stopping to click on links is given by the damping factor d, which is, depending on the degree of probability therefore, set between 0 and 1. The higher d is, the more likely will the random surfer keep clicking links. Since the surfer jumps to another page at random after he stopped clicking links, the probability therefore is implemented as a constant (1-d) into the algorithm. Regardless of inbound links, the probability for the random surfer jumping to a page is always (1-d), so a page has always a minimum Page Rank.

A Different Notation of the Page Rank Algorithm

Lawrence Page and Sergey Brin have published two different versions of their Page Rank algorithm in different papers. In the second version of the algorithm, the Page Rank of page A is given as

PR(A) = (1-d) / N + d (PR(T1)/C(T1) + ... + PR(Tn)/C(Tn))

where N is the total number of all pages on the web. The second version of the algorithm, indeed, does not differ fundamentally from the first one. Regarding the Random Surfer Model, the second version's Page Rank of a page is the actual probability for a surfer reaching that page after clicking on many links. The Page Ranks then form a probability distribution over web pages, so the sum of all pages' Page Ranks will be one.

Contrary, in the first version of the algorithm the probability for the random surfer reaching a page is weighted by the total number of web pages. So, in this version Page Rank is an expected value for the random surfer visiting a page, when he restarts this procedure as often as the web has pages. If the web had 100 pages and a page had a Page Rank value of 2, the random surfer would reach that page in an average twice if he restarts 100 times.

As mentioned above, the two versions of the algorithm do not differ fundamentally from each other. A Page Rank which has been calculated by using the second version of the algorithm has to be multiplied by the total number of web pages to get the according Page Rank that would have been caculated by using the first version. Even Page and Brin mixed up the two algorithm versions in their most popular paper "The Anatomy of a Large-Scale Hypertextual Web Search Engine", where they claim the first version of the algorithm to form a probability distribution over web pages with the sum of all pages' Page Ranks being one.

In the following, we will use the first version of the algorithm. The reason is that Page Rank calculations by means of this algorithm are easier to compute, because we can disregard the total number of web pages.

The Characteristics of Page Rank

The characteristics of Page Rank shall be illustrated by a small example.

We regard a small web consisting of three pages A, B and C, whereby page A links to the pages B and C, page B links to page C and page C links to page A. According to Page and Brin, the damping factor d is usually set to 0.85, but to keep the calculation simple we set it to 0.5.












The exact value of the damping factor d admittedly has effects on PageRank, but it does not influence the fundamental principles of PageRank. So, we get the following equations for the PageRank calculation:

PR(A) = 0.5 + 0.5 PR(C)
PR(B) = 0.5 + 0.5 (PR(A) / 2)
PR(C) = 0.5 + 0.5 (PR(A) / 2 + PR(B))

These equations can easily be solved. We get the following PageRank values for the single pages:

PR(A) = 14/13 = 1.07692308
PR(B) = 10/13 = 0.76923077
PR(C) = 15/13 = 1.15384615

It is obvious that the sum of all pages' PageRanks is 3 and thus equals the total number of web pages. As shown above this is not a specific result for our simple example.

For our simple three-page example it is easy to solve the according equation system to determine PageRank values. In practice, the web consists of billions of documents and it is not possible to find a solution by inspection.

The Iterative Computation of PageRank

Because of the size of the actual web, the Google search engine uses an approximative, iterative computation of PageRank values. This means that each page is assigned an initial starting value and the PageRanks of all pages are then calculated in several computation circles based on the equations determined by the PageRank algorithm. The iterative calculation shall again be illustrated by our three-page example, whereby each page is assigned a starting PageRank value of 1.

Iteration PR(A) PR(B) PR(C)
0 1 1 1
1 1 0.75 1.125
2 1.0625 0.765625 1.1484375
3 1.07421875 0.76855469 1.15283203
4 1.07641602 0.76910400 1.15365601
5 1.07682800 0.76920700 1.15381050
6 1.07690525 0.76922631 1.15383947
7 1.07691973 0.76922993 1.15384490
8 1.07692245 0.76923061 1.15384592
9 1.07692296 0.76923074 1.15384611
10 1.07692305 0.76923076 1.15384615
11 1.07692307 0.76923077 1.15384615
12 1.07692308 0.76923077 1.15384615

We see that we get a good approximation of the real PageRank values after only a few iterations. According to publications of Lawrence Page and Sergey Brin, about 100 iterations are necessary to get a good approximation of the PageRank values of the whole web.

Also, by means of the iterative calculation, the sum of all pages' PageRanks still converges to the total number of web pages. So the average PageRank of a web page is 1. The minimum PageRank of a page is given by (1-d). Therefore, there is a maximum PageRank for a page which is given by dN+(1-d), where N is total number of web pages. This maximum can theoretically occur, if all web pages solely link to one page, and this page also solely links to itself.

The Random Surfer Model


In their publications, Lawrence Page and Sergey Brin give a very simple intuitive justification for the PageRank algorithm. They consider PageRank as a model of user behaviour, where a surfer clicks on links at random with no regard towards content. The random surfer visits a web page with a certain probability which derives from the page's PageRank. The probability that the random surfer clicks on one link is solely given by the number of links on that page. This is why one page's PageRank is not completely passed on to a page it links to, but is devided by the number of links on the page.

So, the probability for the random surfer reaching one page is the sum of probabilities for the random surfer following links to this page. Now, this probability is reduced by the damping factor d. The justification within the Random Surfer Model, therefore, is that the surfer does not click on an infinite number of links, but gets bored sometimes and jumps to another page at random.
The probability for the random surfer not stopping to click on links is given by the damping factor d, which is, depending on the degree of probability therefore, set between 0 and 1. The higher d is, the more likely will the random surfer keep clicking links. Since the surfer jumps to another page at random after he stopped clicking links, the probability therefore is implemented as a constant (1-d) into the algorithm. Regardless of inbound links, the probability for the random surfer jumping to a page is always (1-d), so a page has always a minimum PageRank.


A Different Notation of the PageRank Algorithm

Lawrence Page and Sergey Brin have published two different versions of their PageRank algorithm in different papers. In the second version of the algorithm, the PageRank of page A is given as

PR(A) = (1-d) / N + d (PR(T1)/C(T1) + ... + PR(Tn)/C(Tn))

where N is the total number of all pages on the web. The second version of the algorithm, indeed, does not differ fundamentally from the first one. Regarding the Random Surfer Model, the second version's PageRank of a page is the actual probability for a surfer reaching that page after clicking on many links. The PageRanks then form a probability distribution over web pages, so the sum of all pages' PageRanks will be one. Contrary, in the first version of the algorithm the probability for the random surfer reaching a page is weighted by the total number of web pages. So, in this version PageRank is an expected value for the random surfer visiting a page, when he restarts this procedure as often as the web has pages. If the web had 100 pages and a page had a PageRank value of 2, the random surfer would reach that page in an average twice if he restarts 100 times.

As mentioned above, the two versions of the algorithm do not differ fundamentally from each other. A PageRank which has been calculated by using the second version of the algorithm has to be multiplied by the total number of web pages to get the according PageRank that would have been caculated by using the first version. Even Page and Brin mixed up the two algorithm versions in their most popular paper "The Anatomy of a Large-Scale Hypertextual Web Search Engine", where they claim the first version of the algorithm to form a probability distribution over web pages with the sum of all pages' PageRanks being one.

In the following, we will use the first version of the algorithm. The reason is that PageRank calculations by means of this algorithm are easier to compute, because we can disregard the total number of web pages.

Wednesday, June 11, 2008

Google Trend Update

Google Trends is a service that can be used to see how popular certain search terms are across geographic regions, cities, and languages. Google updated its Trends tool that allows users to see the popularity of a term and compare the level of interest in favorite topics--or people, such as those who of you who like to Google your name. It has updated its Trends tool, to include numbers.

With the new Google Trends, you can now view the numbers on the graph and can also download them to a spreadsheet.

Google Trends analyzes a portion of the search engine giant's Web searches to compute how many searches have been done for the terms entered relative to the total number done on Google over time. Based on that information, the Search Volume Index graph charts the results. Users can search up to five terms.

Previously the tool allowed people to view graphs showing trends including how frequently particular key terms were searched for across geographic regions, by different age groups or in different languages, however no numerical data could be inputted.

In the Google Blog, Google has explained the new updates to Google Trends in a very 'delicious' manner. The comparison in the example is between two prominent ice cream flavors, vanilla and chocolate. The aim is to learn as to how many searches are made for each flavor.

A subset of the tool is Google Hot Trends, which shows what people are searching for on the day of their search. Instead of showing the most popular searches overall, which would always be generic terms, Hot Trends highlights searches that experience sudden surges in popularity and updates that information hourly. Google's algorithm analyzes millions of Web searches performed on the search engine and displays the results that deviate the most from their historic traffic pattern. The algorithm also filters out spam and removes inappropriate material.

Now the file can be downloaded allowing users to analyse it along with the numbers involved, although it will remain scaled with relative results rather than actual ones.

Google Trends can be used for fun as well as for a more practical purpose, although users must first sign into their Google account. The search engine company has been busily expanding its brand of late. Earlier in the year it unveiled plans for the creation of Google Health, an online database that gives internet users access to their own medical histories.

Currently, Google Trends is only available in English and in Chinese. Hot Trends is only available in English. The company said it hopes to roll out Google Trends in other regions and languages in the future.

Tuesday, June 3, 2008

Presentation of Google

With nearly 50% of the traffic generated by the whole of the search engines and directories (in France), Google must not be neglected. From a PhD research subject for two American academics (Larry Page and Sergey Brin) Google became a company on the international scene.

The success of this engine comes on the one hand from the algorithm worked out by the 2 founders, and on the other hand of the application of an elementary principle: the simplest things are sometimes most effective. In this case, Google chose a very stripped interface, without advertisement, by concentrating its services on the search for Web pages and nothing else. The engine also enjoys a very great speed in the interrogation of its data base.

In addition to results of research considered to be relevant by many users, Google succeeded to index a very great number of pages: its "index" is from now on one of the largest in the world (if it is not the first), with approximately 2 billion pages. Recently, new types of documents were indexed, in addition to the traditional HTML: Word, Excel, Acrobat, PowerPoint, WordPad, etc.

The algorithm is based on two systems

1) a precise analysis of the contents of the indexed pages (keywords, occurrences, positions in the document, type of HTML tag, etc.)
2) a classification of the pages according to their popularity (PageRank), calculated from the topology of
the Web (i.e. the whole structure of the documents and the links between them).

Indexing of Webpages by Google

Google set up a crawler-type software, named Googlebot. It is a robot indexing Web pages (and now other types). Its principle is simple (but not its implementation!): when it reads a page, it adds to its list of pages to visit all those linked to the page in the current process.

Theoretically, it should thus be able to know the majority of the pages of the Web, i.e. all those which are not orphan (a page is known as orphan if no other links to it). The volume of data to be treated being important, this robot is a program distributed on hundreds of servers.

In addition to the knowledge of the greatest number of pages, Google also wants to index them regularly, because many the pages are updated from time to time. Moreover the frequency of visit of Googlebot on a Web page depends on its PageRank : the larger it is, the more it will often index it. From one passage to another, Googlebot can detect a page become non-existent ("error 404").

This colossal mass of information will be analyzed by Google in full details. Each word or sentence will be associated to a type, based on HTML tags. Thus a word contained in the title will be considered to be more significant than in the body text. These types may be classified according to their importance (title of the page , headings H1 to H6, bold, italic, etc). This preprocessing, associated with other criteria including the PageRank, makes it possible to provide the most relevant results in first.