Showing posts with label web search. Show all posts
Showing posts with label web search. Show all posts

Monday, June 29, 2009

Also, Consciousness vs. Memory

"Real time taps into consciousness, search taps into memory." Erick Schonfeld's recent TechCrunch post is illuminating. Beyond what Erick discussed of the real-time search, however, there is some more subtle distinction between the two that is worth of being thought. Hence here are my two cents.

Memory is the reserved and refined consciousness. Unlike consciousness that is often fuzzy and even pointless, memory generally has its definite theme and well-organized. We remember for a reason. The reason thus forms the backbone of memory. To search a particular piece of memory out of many, it is possible for us to figure out the algorithms disclosing the backbones. This is what the search engines have done and continue to improve.

Since real time message is indeed about the genuine consciousness in contrast to the refined consciousness (i.e., memory), the previous matrix of Web search no longer sustains. The genuine consciousness itself often lacks of reason. Most of the time it simply tells the undigested truth and the unconscious feeling. Therefore, it is impossible to "rank" the genuine consciousness the same way as if we have ranked (successfully?) the refined consciousness.

The thinking we have done points us to a new hint of effective real-time information search. In fact, we might not call it information search, but information refining. In contrast to search the real time messages, what the engine really needs to do is to refine the messages with respect to the pre-defined channels. By squeezing the water out, the refining process can eventually bring the high quality search results back to the information seekers.

The traditional/standard Web search delivers the direct access to the embodied, refined-already human consciousness. The real-time search then should direct the access to the procedures of refining the embodied, genuine human consciousness. This might be the intrinsic distinction between the two.

Friday, December 05, 2008

Some quick thoughts on Qi Lu's new appointment

Qi Lu, a former high profile Yahoo, joined Microsoft to help the traditional software giant fighting against the newly rising Google. I have a few quick thoughts of this news in the following. Readers who are interested in learning more of this news can read more reports from TechMeme. In particular, Kara Swisher at BoomTown had a unique interpretation of Ballmer's email.

(1) Congratulation to Qi! Great move, personally. Especially as a Chinese myself, I sincerely wish him the best being a great China-born technology leader.

(2) Will Live Search be reformed under the new leadership and be able to compete to Google? Cautiously, I doubt. Allow me be straight. If Qi has not led a compelling plan to compete Google under the visionary leadership of Jerry Yang, could he accomplish this hard assignment under the practical leadership of Steve Ballmer? Unquestionably Qi is a great Guru in the realm of search engine design and operation. However, it is impossible to defeat Google by following the Google way (i.e. the traditional thinking of Web search). Is the experiences that Microsoft lacks at this moment? I question it. Microsoft does not lack of experiences; and experiences can hardly help Microsoft fight Google. What Microsoft lacks is, ironically, the opposite, i.e., the less-experienced fresh soul of innovation. We have to be realistic that there are not many Steve Jobs in this world. Therefore, though I am pleased of Qi's appointment and believe in Qi's contribution to Live Search, this appointment is less than enough for Microsoft to threat Google.

(3) Microsoft seems having abandoned the plan of purchasing the Yahoo Search. It is good for Microsoft; and it is good for Yahoo. I am still looking forward to the newest experiments of BOSS and Search Monkey Yahoo is performing. The worst thing is that they might become another PowerSet.

Tuesday, August 12, 2008

The Age of Google (2): Open Sesame

(the previous installment: The Age of Google (1))

It is an old Arabic story. By accident, a poor Arabic young man named Ali Baba overheard a message spoken by a group of thieves---forty in total. From the message, he learned a tremendous amount of treasure hidden in a cave. The magic spell to open the cave was "Open Sesame." Ali Baba thus entered the cave with the secret spell and took some of the treasure home.

World Wide Web is the cave full of treasure. The search engine sites are the magic doors. The user-specified keywords are, however, the "open sesame."

Among all the magic doors, Google is the most marvelous one. Google has standardized the way of ranking online information. Before Google, we had diverse standards to rank the relevancy of information on the Web. Because of Google, the standards converged. This convergence is an important sign about the age of Google.

Subjective ranking vs. Objective ranking

In primary should ranking be subjective or objective? This was (and probably is still) a crucial debate in the realm of Web search. The issue is not only a technological argument but also a business decision.

Technologically, in the time before Google the performance of subjective ranking was generally incomparable to the performance of objective ranking. Yahoo was providing significantly better relevancy in its search results than the other search engines which performed objective ranking methods. Although due to the subjective policy Yahoo's execution expense was higher, its superior quality of search results made the cost worthy.

On the side of business execution, from the beginning Web search was coupled with the online advertising business model. Popularly, Web search engines were selling their first few related search results to certain advertisers; such a policy was once the basis of online advertisement. The subjective ranking search engines apparently executed the policy better than their objective competitors because the former ones selected relevancy subjectively anyway. There was nearly no extra cost for subjective search engines to embed the business model into their framework. On the contrary, the objective search engines might have to execute two policies in parallel to reach both the technological goal and the business goal. Therefore, the overhead of subjectively deploying advertisement reduced the advantage of objective ranking in its low cost of execution.

Google cleverly resolved the dilemma on the side of objective ranking. As the result, objective ranking declared the victory over subjective ranking, at least until the present. (Be note that now the subjective search strategy is striking back. Represented by Mahalo, the subjective search policy is reclaiming its momentum. I will analyze this phenomenon in the following installment of this series when discussing the challenge Google faces.) Another consequence of the resolution is what we all know: the victory clinched Google's championship on Web search over the previous leader Yahoo.

What Google did was actually on two folds. One fold was at the technological side on which Google implemented a brand new objective ranking policy that significantly improved the performance. We will discuss this fold a little bit later in this post.

The other fold was the revision of the online advertising business model. We have known that the previous online advertising business model favored subjective search. Since Google's technology belonged to objective search, the company was trying to discover a new business model that was more compatible to the objective search. Finally, Google invented AdSense and its sister program AdWords. There have been many discussions about the two programs. In fact, at the next installment I will discuss the two programs again. At here, however, the discussion is solely on the impact of the two programs to the combat between the subjective search and the objective search.

AdSense and AdWords reshaped the online advertising business model from the subjective judgments made by the search engines to the objective decisions made by the run-time mapping between search queries and advertising words. Please be note that this change did not indeed have improved the performance of advertising in the sense of technology. It, however, brought two critical improvements for online advertising business: (1) it decreases the unit cost of deploying an advertising word, hence by spending the same amount of money advertisers may now ask for more advertising words associated to their advertisement; and (2) it is implicitly empowered by the effect of the long tail since objective search automatically associates any unpopular deployment of the advertising words to the advertisement.

In order to thoroughly exploring the advantage of objective ranking, Google made another great decision. Google abandoned mixing the search results with advertisers' product links. Google displays all search results in order purely by their objective ranks. By contrast, advertisers' links are laid separately in such as the side bar of the page of the related search results. By this revision of advertisement deployment, Google won the name of integrity and objectiveness on Web search. This fame has been a crucial part of the foundation for Google's business success.

Cluster weighting vs. Link popularity

Another debate of Web search is between ranking by cluster weighting and ranking by link popularity.

In short, ranking by cluster weighting is to classify documents based on the measurement between typed keywords and predefined clusters of information. The implementation of the methodology may be machine learning or the mathematical analysis of vector computation. Ranking by link popularity, however, is to measure the relevancy of documents based on how popular the document is linked in the Web. More popularly linked ones are assigned greater value of relevancy.

Certainly, as many of us know, Google is a supporter of the latter policy. The PageRank algorithm is a famous representative of the thought.

In theory, however, we should expect that cluster weighting must be superior to link popularity on ranking search results. Essentially, cluster weighting methodology directly looks for the content relevancy, while link popularity just indirectly reveals the favorite of relevancy through human activities. In the other words, the truth itself is always more correct than people's vote of the truth. If we can directly reveal the truth, we do not need to rely on the secondary votes to guess the truth.

The problem is, however, whether we are truly able to compute the truth or even if we could, how much the computation would cost. This is where the problem of ranking by cluster weighting is.

The beauty of PageRank is to greatly avoid the complexity of vector computation between search keywords and content keywords. The mentioned computation is very costly in both of run-time execution and off-time optimization requirement. By investing the same amount of money to store and analyze the topological structure of World Wide Web, Google believes that it might gain more reward than investing on cluster weighting computation. Indeed, Google proves itself.

In essence, public voting is probably the most cost-efficient way to approach the truth when the truth is unknown. Public voting does not always reveals the truth. But if the truth is too expensive to be revealed, public voting is what most of the people accept and it generally reveals something that is close to the truth. This is the same philosophy of democracy. So actually Google was excising Web 2.0 implicitly even before Web 2.0. In person, I believe this is an important reason that Google eventually became an early leader of Web 2.0.

But ranking by cluster weighting is not dying. The policy is actually preparing its strong fightback now. Semantic search is the modern mutation of this traditional policy. As many people suggest and expect, semantic search (if it be realized) would certainly outperform the current search executed by Google. But how to reduce the execution cost remains to be a grand challenge. We will discuss more about semantic search in the following installments.

Open Sesame

In summary, Google's "open sesame" is to explicitly execute low-cost objective ranking over implicit, free, subjective ranking (link popularity) performed by regular Web users. This is a model putting "objectiveness" on top of "subjectiveness of the crowd."

However, the business success of this "open sesame" was still not enough for Google to be an age. Google might have been just another successful company if nothing else happened. There is another crucial reason that eventually pushed Google from a successful company to a legend. At the next installment we are going to discuss it.

(The next installment: Web 2.0)

Referenced resources:

Monday, July 28, 2008

Cuil Search

Beginning from the last night, Cuil (pronounced "cool") starts to hit the headline of many technological blogs.

New Design of Interface

Cuil is experiencing a new design of magazine-style search-result display interface other than the "standard" list-style display of Web search results. The following is a screen shot by typing my name into the Cuil search. This change of design potentially may mean much more than attracting eyeballs.



In an earlier post at Alt Search Engines, I shared that a critical but often overlooked issue in the current Web search is the production of link resources, i.e., how to better formulate the generated links according to the user search requests. To clarify the issue, let me explain it using a metaphor. If one has a brilliant pearl but present it inside a crude lunchbox, how good it might be known? A brilliant pearl needs to have a well-designed box (such as the right one) that matches its superior quality. It is exciting to watch a breakthrough on this issue by Cuil.

This new design provides Cuil lots of potential. For example, each of the short related story shown at the first page of result could be more than just Web links. By contrast, it may be an entry point to a related Web thread and each story is describing the theme of the thread. A combination of Web search and Techmeme-style stories may bring Cuil users very different experiences from using, such as, Google for search.

New ranking and privacy

Another exciting thing to watch is that apparently Cuil has performed a different algorithm on ranking its search results. After Google's success, ranking based on objective link popularity is generally accepted as the foundation of page rank on the Web. Cuil, however, seems trying to apply a new standard of ranking over its stored over 120 billion (as it claimed) Web pages. As the result, from the previous figure we can see that the rank of my related links is very different from the rank of links returned by Google. Discarding the performance until now (anyway, Cuil is launched barely for one day but Google has optimized its results for more than a decade), I strongly support Cuil's attempt. We want to have an alternative solution which does provide us DIFFERENCE. On the other hand, if a new search engine ranks the Web the same way as Google, how could we be convinced that it might do better than Google? Therefore, no matter whatever Cuil has chosen a correct path to walk and we wish it the best luck in the future.

Cuil also claimed some exciting news about its advanced privacy protection technology. Danny Sullivan on Search Engine Land uncovered that Cuil search engine would not log IP information. Be note that Google, Yahoo and Ask.com all perform this IP log in their search engines. Cuil's claim helps protect users' privacy on both of the publishing and surfing on the Web. I recommend this improvement especially to the places where information censorship is rigorous.

Still long way to go

Despite of all the improvements, Cuil still have a long way to go before it may indeed threaten Google. For example, it seems that search engine repeatedly looks for the same links and there must be some severe bugs about how to break a circles in graphs in its algorithm.

The following is the second page of the Cuil search results by input my name. Comparing to the first page in the former figure, we can see that the second page repeats quite a few links that have already shown in the first page. I have tested Cuil by the other queries and it seems that this is a bug generally occurred.



Certainly Cuil has more than this problem. But I am still looking forward to its future. Although Sullivan only showed "cautiously optimistic" to the future of Cuil, I think the value of Cuil would not be the precision of its search results. I agree to Sullivan that it is hard to believe that Cuil can do significantly better on bringing back more accurate results than Google, Yahoo, and Microsoft by employing the similar infrastructure of Web search. On the other hand, however, Cuil does show us that it may help build final link resources in better quality. Only if Cuil can continuously convince people that it can bring people alternatives (even though no better results) that they can hardly get from Google, Cuil would be a success at the end because we want to hear different voices.

Tuesday, July 22, 2008

Evolution of Web Search

I have a guest post about Web search evolution at Alt Search Engines.

In short, the main point of the article is as follows.

There are two basic types of quality Web search engines need to be aware.

1) There is a quality about how well the search results produced to match the user requests.

2) There is a quality about how a search engine has produced the link resources based on the search results so that the produced link resources are more feasible to be further manufactured.

I argue that it is the second quality that drives the evolution of Web search.

Thursday, July 10, 2008

Enough ants may bite an elephant to death

A traditional Chinese idiom says that enough ants may bite an elephant to death (Chinese: 蚁多咬死象). The idiom perfectly describes the strategy beneath Yahoo's newest Web search service---BOSS (Build your Own Search Service). Google is facing the most severe challenge ever in its history.

As we all know, until now Google is likely unbeatable by its dominating power on Web search. Neither Yahoo nor Microsoft, nor even Microhoo, may compete to Google head-to-head. Google is a huge elephant and nobody may even shake it.

Although even a leopard cannot fight an elephant, by raising a huge amount of ants they may bite the elephant to death. Yahoo thinks so, and I agree to it. Google can easily defeat any individual competitor, but Google cannot beat the united force of all competitors. It is the power of collectivity. Brilliant strategy, Yahoo!

On the other hand, however, this action is a double-edged sword. If the ants are so many numbered that they may bite an elephant to death, certainly they can kill a leopard without a question. Therefore, Yahoo search itself might be swapped out of the market even before the decaying of Google search.

A few people may be curious on why Yahoo search might be hurt. Isn't Yahoo that is the service provider of BOSS? If Amazon can make great deal of profit from its public Web service, how could Yahoo's public service eventually hurt itself? To answer this question, we must take a deep look at the difference between Yahoo' service product and Amazon's service product.

Amazon self-produces a large amount of data and makes itself be a huge data center. Amazon Web services allow users accessing and using Amazon's data for varied purposes, including business purposes. To the end, however, Amazon controls the source of data. Hence Amazon have the control over all of its service users.

Now let's turn to Yahoo. Surely Yahoo also produces data, but original data production is the secondary task in Yahoo. Primarily, Yahoo indexes the Web. In contrast to Amazon as a data-resource producer, Yahoo is primarily a link-resource producer. The problem of a link-resource producer is that it actually does not have the power of controlling the linked data. In short, Yahoo knows links, but Yahoo actually does not own the linked data. Because of this reason, any niche search engine may use Yahoo infrastructure to build up a self index of links in its niche domain during the steal mode. After the niche search engine comes to public, it can primarily works on its own index and just using Yahoo infrastructure to be a reference checker. The key point here is that links are normally not the end data users look for.

Does it mean that Yahoo has done a work only hurts others but not benefits itself? Not at all. The key point is that Yahoo must not try to charge the search flow through its opened infrastructure ever after, even if some search flow is for commercial purposes. Be not like Amazon Web services because they are totally different (thought they are similar on surface but completely varied in essence).

Yahoo should make itself be the biggest player over its own free and open search infrastructure in contrast to make itself be the leader of its not totally free and not totally open search infrastructure. Or in other words, Yahoo makes itself to be the largest ant instead of a leopard. If Yahoo can keep on its strategy in this way, Yahoo will have a chance to not only compete to Google, but also defeat Google at the end.

Keep on your great work! Jerry, your passion is truly respectful!


For readers who want to learn more details about BOSS, here are a few resources they can start for looking.

  • Read/WriteWeb: Search War: Yahoo! Opens Its Search Engine to Attack Google With An Army of Verticals
  • Yahoo! Search Blog: BOSS -- The Next Step in our Open Search Ecosystem
  • Between the Lines: Yahoo’s desperate search times call for open source
  • VentureBeat: Yahoo opens up its search platform to third parties, Me.dium takes the plunge
  • GigaOM: Yahoo, Now Offering Search as a Web Service
  • TechCrunch: Yahoo Radically Opens Web Search With BOSS
  • zooie’s blog (a BOSS team insider): Yahoo! Boss - An Insider View

Wednesday, May 21, 2008

Pay you to Live Search, brilliant?

Both TechCrunch and VentureBeat reported that Microsoft would announce a new search advertising model, which pays users who use the Live Search engine to search and eventually finish an online transaction. Is this a brilliant idea? Or not?

One week ago, I was at Redmond with the Live Search team. In Live Search, there is a group whose job is "to explore all the crazy ideas." The group picked me to interview and asked about my "crazy ideas". Unfortunately, however, I did not have any crazy idea except of a fairly novel vision suggested for Microsoft to compete with Google. In short, my suggestion was a totally un-Google Web search strategy that the Live Search crazy-idea people could not catch where my craziness was. As the result, they were not impressed by idea that did not sound crazy. Now I see what the Live-Search-craziness is.

The philosophy beneath this "crazy" idea is straightforward: when there are two sites from which we could buy the same product, we often choose the site that gives me more discount when checking out. Since Microsoft has lots of money, why not directly buy searchers from Google?

Is the idea "crazy"? Crazy, indeed.

I would like to quote a comment written by a reader (Tyler Wright) of TechCrunch. He has made a very cute analogy that points to the problem of this "crazy" idea---

"GM and Ford offer cash back to buyers, and they’re on their way out - and losing credibility daily.. Sounds kinda similar."

Ah-ha, this is the problem. GM and Ford often do more discount and promotion than Honda or Toyota does. But the discount and promotion seem do not really save their fate. Why does Microsoft believe that discount and promotion would save it from Google?

A deeper thought behind this strategy is that "online search" and "online transaction" are actually two varied phases. At the online-search phase, we look for a good search engine that can help us find what we want quickly and conveniently. During this phase, we also frequently look for advertisement provided by the search engine. By contrast, at the online-transaction phase, we simply want to finish the transaction as quick as possible. But at the same time, an extra bargain is alway appreciated. Very few people, however, will continue looking for product advertisement when they are checking out (because they are tired of long-time shopping).

By the former analysis, I cannot see why I should abandon Google for being my search engine, while at the same time I can use Live Search to check out. If many Web users adopt this simple strategy to maximize their benefits, I cannot see how much Live Search may gain by this scenario. To the end, Live Search could gain a few net-flow from Google due to the final check-out transaction. The problem is, however, this extra net-flow gained by Microsoft does little help for prompting advertisement in Live Search. This is thus the key of the entire issue.

Sorry, I am still too calm to be crazy.

Sunday, May 11, 2008

Return from Microsoft

The past weekend I attended a special event organized by the Microsoft Live Search team. The Live Search team invited 28 up-to-graduate PhDs from US and Canada to come to an on-site interview. The uniqueness of this event is that nearly none of the 28 candidates have been telephone interviewed by the team before and many candidates (including me) even have not submitted a resume to Microsoft before being invited. The Live Search team searched out these candidates (probably through the Live Search) and invited us. The theme over the event is straight and clear---beat Google!

About the Trip

In general, the trip is full of pleasant. The Microsoft recruiting team have organized a wonderful event. It is warm, joyful, and with exciting surprises. Thank you, Erin Bucholz (my direct contact recruiter), Jared Singer, Rondell Honcoop, Ben Mercer, and others (sorry I am not good at memorizing names).

Each of the 28 candidates is a selected one with a particular background that is related to Web search. For instance, there are candidates with the background of query optimization, distributed computing, image processing, data mining, natural language processing, and so on. My particular background is labeled Semantic Web. Moreover, I found that I was the only candidate invited to this event with the primary background as Semantic Web. It is exciting to be in a group of young scientists of varied disciplinary areas while at the same time with a focused general theme.

Microsoft has arranged most of the candidates (including me) staying at Westin Bellevue, a modern 5-star hotel that is adjacent to upscale shopping centers. The general environment of the hotel is superior. The guestrooms are luxuriously furnished in modern style---clean, simple, straight, and with modern decor.




At the very morning of the interview day, Microsoft sent a limousine (a 14-passenger Navigator model I believe) to pick up the candidates to the campus. It was a slight surprise to everybody and we were joking to attend a party rather than an interview event.

limousine exterior
limousine interior

The Interview Event

In this interview event, the Live Search team has scheduled three group discussions and four-round individual interview for every candidate. All the three group-discussion leaders are well-known Web researchers from Microsoft Research. They are Yi-Min Wang, Chris Burges, and Paul Viola. There are totally more than 10 individual interviewers. Each of them is a program leader with a particular focus on Live Search. Microsoft was indeed serious to this event.

Yi-Min Wang is a passionate speaker. He has a strong passion on competing Google. At the beginning of his session, Wang briefly introduced himself and described a general paradigm of Web search from the industrial point of view. In short, Web search is about finding a matching between billions of Web pages and billions of search queries, while at the same time the numbers of both pages and queries are increasing.

After the brief start, the rest of the session was focused on questioning and answering. In particular, when answering a question Wang described his experiences on fighting to fake Web pages, which was once reported by The New York Time. This topic also led to many discussions of a broader issue of the online advertisement business model and the risk of this popular business model.

I asked a question to Wang how he thought of the factor of humans in his described grand picture of Web search. His answer was primarily focused on the user interface issue. Nevertheless did I agree with his points, he had neglected mentioning the connections between Web resources (such as data, services, and links) and the humans who create or own these resources. I thought that the factor of these connections should at least be another critical issue with respect to my question. But time for question answering was limited and thus he might just not have enough time to expand his discussion.

The session with Burges was slightly different from the previous one. Burges began by asking every candidate why they came to this event. I answered by quoting myself a motto---"We may beat Google, by not by following the Google way." Microsoft is thinking of defeating Google on Web search. Hence I am very interested in coming to hear how Microsoft would approach this goal and I am willing to share with Microsoft how I think this goal could be approached. I would be glad to join them towards this goal together.

The addressing of "beating Google" does not mean at all, however, that Google is bad or evil. It is only about that we need competition to improve Web search better and better. Eventually Web users will be the biggest winners. With the same purpose, I have shared my viewpoints with the CEOs of Mechanical Zoo and Imindi (two ambitious startup companies but with great potential) in the last two weeks. Max Ventilla (Mechanical Zoo), Adam Lindemann (Imindi), and I have shared this common belief---by not following the Google way, we may approach alternative fascinating solutions for Web search and Web knowledge reorganization. Again, I shared this motto with Burges and he told me that it was also exactly what he believed.

Burges did a brief slide show for us about his new role at the Web search team and the Microsoft Live Search ranking algorithm. After his talk, I asked how he would compare the Microsoft ranking to the famous PageRank algorithm used by Google. He replied that the real ranking algorithm used by Google has already not been the original PageRank for long time. The very core of the Google ranking algorithm is a top secret. But Microsoft is catching up.

I agree with Burges. The PageRank algorithm is too raw for real-world products. To get high quality search results, Google must have done a significant revision of this general algorithm. The revision might have been so great that the eventual Google ranking used now may actually no longer be named the PageRank in its standard sense.

On the other hand, however, the PageRank algorithm reflects the Google's philosophical vision of World Wide Web. Google evaluates the Web to be a network of linked nodes where the importance of individual nodes is primarily determined by the linking popularity inside the entire network. This philosophy is the fundamental of "the Google way." Hence by "not following the Google way" it means that we need to think of the Web in a fairly different picture from the one Google thinks. I have such a different picture described. Adam Lindemann of Imindi shares a very similar vision as mine. But what is the picture that Microsoft thinks? Burges had not directly addressed an answer, and neither had Wang. Unfortunately, I have not gotten another chance to explore this issue deeper with another Microsoft developer or researcher in this trip.

In the third session, Paul Viola did a formal presentation to all the candidates in lunch. His talk was primarily on how to do a good research particularly in an industrial company rather than in an academic institute. He made several constructive suggestions for young scientists and engineers. The talk is informing and helpful. Due to the time constraint, however, we have no chances to ask him specific questions.

After the lunch, we came to the most fun part of the event---the swimsuit competition (i.e., individual interviews).

All the interviewers are excellent leading developers and researchers at Live Search. They are very well experienced and more important, they all have the passion on what they are doing (at least for the four who had interviewed me).

The four-round individual interview is scheduled into four main topics---research background discussion, coding, algorithm, and design issues. At each round, a candidate got to talk to an interviewer for about 45 minutes. In general, I have done a fair but not an excellent interview. I have described my thoughts to them, written a few lines of code, and solved some problems. At the same time, however, there was a problem I could not figure out the final answer to the end. I had tried to think of it from all angles except of the angle the interviewer looked for (and it is not a familiar territory of mine, what a pity).

In short, I had emphasized two points about the future of Web search: (1) the switch from the "God" role to the "Guru" role of search engines, and (2) the importance of "proactiveness" in the next-generation Web search. I emphasized to the last interviewer (a kind lady) that we were not just searching the Web. By contrast, we are searching the evolving Web; and this is a key when thinking of beating Google. If Live Search can think of the Web evolution a step further beyond Google, Live Search would get a better chance to beat Google.

Final Address

The final decision of this interview event will come to me in few days. No matter whatever, however, I appreciate this oppotunity and it gives me a chance to hear the first-hand opinions about the future of Web search from the most frontier industrial developers and researchers as well as from many peer PhD candidates all over the North America, let it alone that I had a wonderful journey at Seattle.

Tuesday, April 29, 2008

Microsoft Windows: more than operating system

(revised May 3, 2008)

Recent reports (such as this one) tell the drooping of Microsoft Windows in the market of operating system (OS). How should Microsoft react to this slide? There may be several options. For example, (1) to accelerate improving Windows Vista, (2) to speed up the migration towards Windows 7, (3) to buy Yahoo, or (4), as I am going to suggest in this post, to explore the potential of Windows that is more than operating system. In particular, I would rather devote this post to the Microsoft Live Search team. Live Search might become a central piece of the new Windows as I would suggest in the following.

Windows, less of momentum to keep on going?

The progress of Windows seems have been slowed down in recent years. I had once asked myself, should I upgrade the installed Windows XP in my laptop to Windows Vista? At last, I decided to wait for a while because I could not see any emergency of the upgrade. Now it seems that this decision might be smart. Many users do not like Vista and they would rather stay with the old Windows XP even for their new computers.

Besides this Vista story, another thread to Windows is from a competitor of Microsoft---Apple. The increasing influence of the Mac operating system has been gradually severe to Microsoft. When the technology of operating system on PCs is becoming more and more mature now, adopting Mac instead of Windows becomes a more and more persuasive excuse especially when the UI design of Mac is generally thought prettier than that of Windows.

Do Microsoft engineers suddenly lose their mind to create this unsatisfactory Vista product? Or are Apple engineers becoming much smarter to compete to Microsoft? Actually, none of the answers are true. The truth is that any further improvement of OS on PC has become so difficult and there are so few rooms to grow on advanced OS technologies on PC that the version-upgrading strategy executed by Microsoft for years may come to its end if no major change of action would been made. We are actually not complaining of how bad Vista is (in fact, Vista is not a bad product). The real issue is that Vista is not innovative enough to be a new version of Windows. Hence if someone really looks for experiencing a new operating system, switch to Mac is probably a better option than upgrade to Vista. This is thus the problem.

Windows, more than an operating system!

To save Windows from collapsing, buying Yahoo is not the only option as some experts have suggested. Although it becomes harder and harder for Windows to upgrade just as an operating system, Windows is actually more than an operating system.

The term "Windows" represents operating system when we talk about personal computers (PCs). But there is a question: will Windows mean something more when a PC is linked to World Wide Web? This question is crucial.

About Windows and World Wide Web, here are several statements.

1. Windows was born when the Web was infant.
2. Windows is not designed for the Web.
3. Windows is an operating system for PCs but it is not a Web operating system.
4. There are no Web operating systems so far and probably there should never been a Web operating system ever.

But all these statements still do not touch the question we just asked: what does Windows to users mean when a Windows-installed PC is connected to the Web?

To answer this question, we should first look at what a PC becomes when linked to the Web. By linking a PC to the Web, the PC automatically becomes part of the Web. Or more precisely, the storage space on the PC becomes a small portion of the gigantic space of the Web. When Windows manages the space on the PC, the Windows manages part of the Web. Hence in this local environment, Windows plays the role to be a web-resource operating mechanism (WebROM). This is a critical recognition of a new function of PC operating systems.

There is another effect caused by linking PCs to the Web. When PC users explore the Web through their PCs, the user activities on the Web result in a local topology of the Web on the PC. If we connect all the Web nodes navigated by the users of a PC, we can obtain a topology of a typical portion of the Web that reflects the interest of the respective PC users. Therefore, linking a PC to the Web causes not only PCs on the Web but also the Web mapped into PCs.

Based on the discussed we have done in the previous two paragraphs, a PC operating system (such as Windows) is indeed a WebROM of not only a particular Web space (the particular local PC) but also a typical topology of Web resources that are physically stored in other places. This is why Windows is more than an operating system by the traditional mean.

I must emphasize that the last statement just made does not, however, suggests Windows to be a Web OS. As I discussed in an earlier post, a WebROM is fundamentally different from a WebOS. What each Windows manages is a topology of a very small portion of the entire Web. Such a topology is consistently a closed world by contrast to that the entire Web is always an open world. I do not believe in generic Web OS.

Live Search, bring a new life for Windows

One Microsoft product will be crucial if Microsoft decides to explore the new Windows by expanding its ability of Web resource management. The product is Live Search.

According to Web resource operating, a central task is search. Unlike local resources stored in personal computers, PC users do not have the full control over Web resources in general. Many traditional OS issues on PCs such as I/O device management are not critical to the role of WebROM. By contrast, PC users do have the right to SEARCH all Web resources even though they cannot control most of them. The resource search operation thus becomes primary.

Up to the date, Live Search at Microsoft is following a route that Google and Yahoo have experienced successfully so far. Microsoft is developing a generic, centralized search engine that is supposed to have indexed the entire Web for users to search. Yahoo has succeeded with this strategy, and so has Google.

A problem to Microsoft by adopting this "successful experience" is that this strategy makes Microsoft forget its unique and powerful weapon that neither Yahoo nor Google has. The weapon is Windows. By giving up this weapon, Microsoft is just "a three-year-old kid comparing to the 12-year-old big boy Google" on Web search, once said by Mr. Ballmer. However, what might happen if the three-year-old kid decides to pick up a WMD (weapon of mass destruction) on hand? Google will be really afraid when Microsoft starts to assign Windows a new interpretation towards the new Web age.

In addition to the current strategy, Microsoft may think of an alternative strategy of Live Search to defeat Google. In my mind, this new search strategy should have a decentralized paradigm. Based on the network of registered Windows users, Microsoft can develop a novel social-search-style strategy by embedding Live Search into Windows (by contrast to add a Live Search link into Web browsers, I will explain the difference at the end of the post). This new strategy will put Live Search to the center of the new Microsoft Windows.

I will not discuss more details of this new search strategy in this post since it has already gotten be too long. I may start another post on this topic later depending on my time. But I will share the details of my thoughts with Microsoft Live Search scientists and engineers next week at Redmond.

More about the Yahoo Deal

As last, I want to say a little bit more about the Microsoft-Yahoo deal. In an earlier post, I briefly expressed that the deal may hurt Live Search in a long run. But certainly it is not because of Yahoo! Search that Microsoft wants to buy Yahoo.

Microsoft wants to jump into the market of online advertisement and buying Yahoo may be the fastest way to get into it and obtain a decent percentage of market share immediately. Unquestionably there are enough proper reasons for Microsoft to take this action. The problem is, however, that whether the action is really the best option, let it alone the only option, on the table, as some analysts have argued. I think it is not.

Merging with Yahoo could be a huge burden to Microsoft. This action would cost Microsoft both of the time and money to develop new-age Web-resource management strategy that is critical to the next-generation Web search. Anyway, Yahoo is constructed on top of its original successful Web search portal and Google's success on online advertisement is also on top of its successful Web search platform too. Designing a new-age Live Search is much more important for the future of Microsoft than merging Yahoo. Anyway, if Windows can keep on its strength on growing, buying Yahoo is actually much less critical to Yahoo, as the same analysts have implicitly suggested. Let's choose Live Search and new Windows instead of merging Yahoo!

(The binding of Live Search and Windows is not to enforce all the search flow to live.com by Windows. Otherwise Microsoft must get sued immediately by Google and all other search engines. The binding is actually about reinterpreting Microsoft to be a provider of WebROM and Windows is the product. By this reinterpretation, Web search is nothing but another basic function of new Windows such as the other ordinary Windows operations, e.g. creating a new file in PC. By this change, it does not matter who is the default search engine set in a computer. Even if Google is set to be the default search engine in a PC stalled by this new Windows, Microsoft may still gain a big (or probably the greatest) share on online advertisement because Windows is always the default search platform, in contrast to the particular search engine. Windows becomes the manager of the Web.)

Saturday, December 22, 2007

Thinking Space 2007 in 12 months

This post is the highlight of what was on Thinking Space in 2007 month-by-month. I am grateful to all the readers of Thinking Space and wish you merry Christmas and happy new year!

January 28, 2007, Web 2.0 panel on World Economic Forum

How would Web 2.0 and the emerging social networks affect world business? The annual World Economic Forum at Davos organized a panel with five outstanding Web business leaders addressing this issue at the beginning of 2007. The talks, however, showed that the executives from traditional big companies such as Microsoft and NIKE were less alerted to the new technologies than the executives from new-age companies such as YouTube and Flickr. In short, both Bill and Mark were talking in languages other than Web 2.0. By their viewpoints, the Web-2.0 phenomenon was certainly less important than their own imaginary vision towards the future. What web evolution really impacts world business was severely underestimated.

At the end, the speech by Viviane is worth of re-emphasizing. When the Web evolves to be more and more mature, who are going to govern the virtual world? This question may gradually become a severe issue when web evolution goes further. Will there be conflicts between the virtual world governments and the real world governments? I do not think that in 2008 we will immediately see this type of conflicts. But the traditional means of national boards do have started to diminish while the new means of digital boards are forming; these changes are slowly but inevitably.

February 18, 2007, The Two-Year Birthday of AJAX

Few technologies have affected the Web so much as AJAX has done. AJAX is more than a technology; it is a philosophy. What AJAX really does is to decompose Web content into smaller portable pieces that are feasible to be uploaded and updated independently. AJAX prompts the dynamic recomposition of pieces of Web content from varied resources. Hence it significantly improves the reuse of information on the Web.

The prevalence of AJAX causes the fragmentation of the Web. The reverse side of this phenomenon is, however, how we may defragment the small pieces of information and reorganize them from end-users' perspectives. This defragmentation issue is the next critical challenge for Web information management. Twine is an example that has started to address this issue. I expect to see more proposals to solve this defragmentation issue in 2008.

March 23, 2007, Will the Semantic Web fail? Or not?

Whether the Semantic Web is going to succeed is always debatable. There are many supporters of Semantic Web, and there are nearly as many as the opponents as well. Will Semantic Web become true? The answer partially depends on whether the Semantic Web researchers can humbly learn from the success of Web 2.0. The normal public might not welcome Semantic Web if its research is still kept inside the ivory tower. Practices such as Microformat are good examples that the Semantic Web research approaches normal web users. But there are still too few of this type of examples. For instance, will the new W3C RDFa proposal be too complicated again? We don't know yet. Hopefully this time W3C would focus more on simple solutions that are feasible to normal users rather than on sound and complete solutions that the academic researchers favor. In comparison, if our real human society is far less than being perfect in reasoning and inference, why must we have theoretically perfect plans to build a virtual world?

April 18, 2007, New web battle is announced

Google is expanding rapidly. Google had replaced Yahoo being the leading Web search engine. Google has already been the largest site that produces Web-2.0 products. Google is competing against Microsoft to be the leading online document editor. Google is fighting against Facebook to be the leading social network through the OpenSocial initiative. More recently, Google starts another battle against Wikipedia to be the leading online knowledge aggregator by the announcement of Google Knol. Can Google succeed simultaneously in all of these fields? Are Google's plans too ambitious to be successful?

The age of Google is about to pass; this is my prediction after watching all these ambitious plans issued by Google. Google has started losing its momentum on originality. By contrast, Google is now repeating a "successful" path of many traditional big companies, i.e., dominating the market by defeating the opponents not by new achievements on technologies but by its superior money resources. This strategy has been proved successfully in many fields. However, it is not a winning strategy on web industry. The reason is that World Wide Web itself is evolving. When the Web evolves, Web technologies evolves. Any company that stops evolving would be thrown away. The history once happened to Yahoo may happen to Google again in the future. The age of Google will be passed with the over of Web 2.0.

May 8, 2007, Web Search, is Google the ultimate monster?

Google is beatable, but Google is not going to be defeated by another Google-style solution. When I predict that the age of Google is about to pass, I mean new revolution on Web technologies. Google is thinking of itself as the God of World Wide Web; and indeed many Web users accept this interpretation (because we have no other better choices at present). But history has already told us that this type of fake gods like Google could not stay forever. In history, we humans abandoned most of the fake gods as soon as the public education system was prevailed. In similar, this history will repeat itself in the virtual world of the Web. The fake God of the virtual world (Google) will step down when the education on Web machine agents prevails. Hakia would not threaten Google if it continues following the Google strategy by addressing itself to be a more powerful fake God on the Web.

In addition to this short summary, I have a preliminary funding request. I will graduate next year and currently I am looking for an assistant professor position. If I'd get an offer, I would start a new research project on next-generation Web search that is beyond the current Google-style search strategy. In fact, I have already done the project proposal. For any reader, if you are responsible on looking for and funding new research projects that are full of potential in the future, I am far more than happy to discuss my project with you. I can be contacted through yihong.ding@gmail.com. The philosophy underneath my new web search strategy can be read at here.

June 29, 2007, Epistemological extension to ontologies: a key of realizing Semantic Web?

The application of epistemology into Semantic Web is less explored than it should have been. We need ontologies to enhance the collaboration and agreements. We also need epistemologies to emphasize the individuality and privacy. I expect more research on this topic in 2008.

July 31, 2007, What does tagging contribute to the web evolution? | An introduction of web thread

There are many ways to describe web evolution. One unique expression is the transformation from the node-driven web to the tread-driven web. Web thread is a new term proposed by myself. In short, a web thread is a connection that links multiple web nodes to a fixed inbound. I observed that the Web was not only syntactically connected by human-specified links, but also semantically connected by latent threads each of which expresses a fixed meaning. A straightforward evidence of the existence of web threads is Web-2.0 tags. On Web 2.0, resources are automatically mutual-connected when they are specified the same tag by individual human users. When weaving these tags together, we obtain an interconnected network of all web pages.

The existence of web threads is an interesting phenomenon that lacks of insightful research at present. From one side, web threads are part of the implicit web because they are generally latent at this moment. On the other side, by proactively revealing web threads and explicitly weaving them, we might produce more comprehensive social graphs for individual web users. This new concept thus may contribute significantly to the vision of Giant Global Graph. I will publish more research on this concept in 2008. By the way, a broader discussion of web links and web threads can be found at here.

August 24, 2007, Mapping between Web Evolution and Human Growth, A View of Web Evolution, series No. 4

World Wide Web is evolving. But why does the Web evolve and how does it evolve? Few answers have been given. The view of web evolution is the first systematic study in the world that directly addresses the answer to these questions based on a theoretic exploration.

This view of web evolution stands upon the analogical comparison between web evolution and human growth. I argue that the two progresses are not only similar to each other by their common evolutionary patterns, but also literally simulate each other from all the major aspects. At present, the simulation mainly happens in the uni-direction from the real world to the virtual world. In the future, however, we are going to see more evidences of simulation on the reversed direction, i.e. from the virtual world to the real world.

The virtual world represented by the Web is nothing but a reflection of our human society. Due to the limit of web technologies, however, we are not able to completely simulate our society from every aspect into this virtual world. In particular, we are not able to well simulate all the activities of individual humans on the Web. By contrast, we can simulate individuals at a certain level within any specific evolutionary stage. This continuous upgrade of simulation of individuals on the Web represents the main stream of web evolution.

This theory of web evolution has published for half a year and I have received many requests on discussing this vision. I hope this study would bring more attention to the fascinating web evolution research.

September 16, 2007, A Simple Picture of Web Evolution

The simple picture of web evolution expresses a straightforward timeline of web evolution. The Web is evolving from a read-or-write web to a read/write web, and eventually it may become a read/write/request web. The implementation of the "Request" operation would be a fundamental next-step towards the next generation Web.

October 7, 2007, What is Web 2.0? | The Path towards Next Generation, Series No.1

What is the next generation Web? This is a grand question to all Web researchers at this moment. We might see critical breakthrough on answering this question in 2008.

At present, the advance of Web 2.0 has already slowed down. The progress of web evolution has reached another stable quantitative expansion period after the exciting qualitative transition from 1.0 to 2.0. The seed of next transition is growing underground now.

In order to figure out the path towards the next generation Web, we need to know the present and where the present was coming from. In the first post of this series "towards the next generation", I summarized the various definitions of Web 2.0. In the following installments at this series, I will continue discussing my vision of the path towards Web 3.0. I feel sorry about the slow progress of this series. I will try to post this series more frequently in the coming year.

November 23, 2007, Multi-layer Abstractions: World Wide Web or Giant Global Graph or Others

Giant Global Graph is a new concept. Although Tim Berners-Lee proposed this concept intuitively for freely deploying personal social networks onto the Web, my view of the intent of this concept is beyond this intuition. In general, I believe that the proposal of this concept is the first sign of a great transition---the organization of web information is transforming from the publisher-oriented point of view to the viewer-oriented point of view.

The impact of this transformation could be greater than we may imagine. Most importantly, this transformation will show that the Web may automatically re-organize its information system without a human-controlled organization such as W3C or Google. World Wide Web is a self-organizing system. This observation is essential to the understanding of web evolution.

December 3, 2007, Collectivism on the Web

The implementation of collectivism has been the landmark of Web 2.0. But do we know how many types of collectivism we may implement onto the Web? This last selected article at December 2007 summarized a few typical implementations of collectivism on the Web. Some of them (such as collective intelligence) have been well known, while others (such as collective responsibility and collective identity) are less known by the public. I expect to watch more creative implementations of collectivism in 2008.

Tuesday, October 02, 2007

Yahoo updated its search

Believe or not, I like Yahoo though many of my discussions about Yahoo are negative so far. It was because of Yahoo that I were able to learn Unite States. I still remember those old days when I explored the Yahoo list of US universities and how much exciting I was. I do love Yahoo.

But I do blame Yahoo a lot too. Yahoo has been just stayed where it was for long time. The Yahoo site is still popular. But the reputation of Yahoo search is ruined indefinitely. I blog about and blame Yahoo because in personal feeling I still love it, though I am now using Google as my default search engine.

My Experiences on New Yahoo! Search

It has been long time not hearing Yahoo search declaring exciting news of its creation. Finally, it seems we got one. Yahoo updates it search engine and brings us some new hope for this classic site.

new Yahoo

By typing in my name "Yihong Ding", I got several suggestion of related concepts about this search request. Fortunately (or unfortunately), all of the Yahoo suggested keywords are actually about me. (Sorry, the other "Yihong Ding"s.) For instance, "semantic web", "ontologies", "web evolution", and "Web 2.0". Bingo! All of them are in my expertise areas.

Moreover, Yahoo also suggested "innsbruck, austria". Good enough, this is where I did my internship last summer. The next one is "w. embley", who is my PhD advisor. Weird! Where does his first name go? The rest ones are "semantic annotation" and "rdf". All of them are indeed related to me. Done!

Am I satisfied? Sure, indeed. But how about the other "Yihong Ding"s (not me), are they satisfied? Probably not. Very likely they, if any, will drop off Yahoo for another search engine immediately.

So the problem is that this suggested set is heavily biased. They are only about one particular "Yihong Ding" but not the others. This is an intrinsic problem of the current semantic understanding technologies.

If the semantic technology would succeed at the end, it must overcome the winner-take-all problem. In this world, the general public are unpopular ones and they do not want to see that their existence has been overlooked.

Discussion

To me, the really inspiring contribution this new Yahoo search brings is the philosophy beneath these keyword suggestions, i.e. the idea of "search assist". Presented by Yahoo researchers, search assist is an attempt of changing from what to do to what have done. By leveraging the search history collected by Yahoo servers, Yahoo tries to provide as many suggestions as possible to help users rapidly get their search assignment done. This is definitely a positive progress.

A curious question about this new Yahoo search assist is which direction it is going to pursue. I feel two paths, while one is dangerous and the other is adventurous. The dangerous path is to repeat the fallacy of Yahoo Directory. Eventually, the search assist becomes a tedious human-managed taxonomy. The adventurous path is, however, to upgrade "search assist" to "search assistant". That is, Yahoo should give up the control of search assist. By contrast, Yahoo hands the power of control to individual users and let them hire Yahoo search assistant to search. Yahoo changes its role from a central web search hub to a central search distribution hub. Digital search assistants become Yahoo's employees who work for individual web users.

In summary, new Yahoo Search does bring new hope. Will Yahoo start to go for a new path and avoid being trapped again into the same old problem? We don't know, but I wish Yahoo the best of its future.

Wednesday, June 13, 2007

Semantic search has two legs

The discussion of semantic search has gradually become popular. Just not long time ago, semantic search was thought to be barely a little bit more than a dream. At present, optimistic researchers have started to believe its possibility in the near future. Very recently at Read/WriteWeb, Dr. Riza C. Berkan, the CEO of Hakia (a company declared to perform "semantic search"), posted an article about semantic search that attracted much attention. Despite of agreeing with the post, here are more thoughts about semantic search.

Semantic search has two legs: semantic understanding and proactive collaboration. Until now, however, most semantic search articles only have focused on the first one. Including Hakia, an "ideal" semantic search engine is popularly thought to be alike a "semantics-enhanced Google." This is, however, a narrowed thought. The intension of semantic search is more beyond "semantics + search."

In order to better understand these two legs, we may watch a regular semantic search scenario in human society that is, however, often overlooked. We humans have daily practised a type of semantic search very successfully for centuries. We ask questions; everybody asks questions, from children to adults. We ask questions to look for answers. These question-and-answer behaviors are typical semantic search activities.

When we are young, we look for answers from parents, whose words are oracles to us. When we grow older, we look for answers from teachers, whose words are oracles to us. When we grow even older, we start to realize that there are indeed no oracles. We start to look for answers by ourselves. In particular, we make friends with various specialities. These friends become our sources of question answering when we get troubles in particular realms. At the meantime, we ourselves also become such a type of sources to our friends. These links constitute a delicate, complicate, and successfully executed network of semantic search in our human society.

If we take a closer look at this successful semantic search network, we can find two fundamental factors that support its execution. First, its success relies on the ability of semantic understanding at each but not some of its nodes. It is generally believed that the set of human knowledge is too rich and too complicated to be executed in a centralized way. For instance, Mor Naaman at Yahoo! Research very recently said that "there is no way that we can engage the masses in annotating media with 'semantic' labels" in a WWW2007 panel. Therefore, representations of global semantics are better to be distributed widely other than be accumulated onto only a few special nodes. In consequence, every node in this semantic search network has its ability to perform a certain level of search depending on its own capability of semantic understanding. This is the basis of a successful semantic search network.

Beyond the local semantic understanding on every node, a successful semantic search network also requires proactive collaborations. In a search network, some nodes (such as professors) may have much greater capability than others (such as first-grade students). But even the node with the greatest ability is still very much limited in its search capability when the search space is about the whole set of human knowledge. A successful semantic search network demands well collaboration among individual search nodes. Moreover, such a type of collaborations appeals to being proactive.

Proactiveness is a unique factor in the network of friendship. The network of friendship is not only a regular social network, but also a search network. When we get troubles, we used to get to our friends for help. Nevertheless, we often make friends on purpose, i.e. in contrast to randomly or aimlessly. A successful semantic search network in human society is priorly built on the joint or depending interest of individuals. For example, both John and Mary love music; so John actively make friend with Mary. Another example, John play piano and does not know to tune a piano; Kate, however, is good at tuning pianos. For the sake of his future requirement of piano-tuning, John proactively make friend with Kate. The third example, Rose is good at history literature; John, however, does not like to read history literature. In consequence, John inactively make friend with Rose. These examples show that the establishing of a search network very much depends on the proactiveness (which in turn decided by the semantic understanding of interests) of these nodes to make connections.

In summary, semantic search naturally contradicts to the centralized web search strategy. In order to activate semantic search to the practical level, we need a search network that is participated by all web users beyond the few independent and aggregated semantic search nodes such as Hakia. The entire web search strategy must experience some revolutionary change other than simple makeups. In the second part of my article about web evolution, I have more discussions about the collaborative search for the future semantic web.

The initial draft of this post is published at SemanticFocus.

Tuesday, May 08, 2007

Web Search, is Google the ultimate monster?

(Revised at Sept. 29, 2007)

If investing on web technologies is buying jewelry, web search technology is the most brilliant diamond on top of these jewelries. A recent post about Top 17 Search Innovations Outside Of Google at Read/WriteWeb attracted hundreds of click and dozens of comments again. Google or Post-Google, this question attracts eyeballs.

New companies may beat Google, but not by the Google way. At present, this claim represents not only innovation of search technologies, but also revolution of basic web search strategies. As I discussed in an earlier post, unlikely we can build another, more advanced "semantic Google" to beat the current Google. To defeat Google in real, the basic strategy of web search must experience a revolutionary change.

The current web search strategy can be summarized as the oracle-based web search. When web users search information, they look for oracles from the "Gods of the Web." These "Gods" are search engines, which are assumed knowing about the Web way more than us as normal persons know. Among these "Gods" the greatest one is Google. In the current WWW, Google is the "God." We beg for the oracles from it to access expected web resources. If there are no answers from Google, by convention we just believe that there are no answers to our question on the Web. This scenario typically shows how much we have trusted and relied on this "God." Although nearly all of us know that Google often makes mistakes and even Google can only search a comparatively small portion of all web resources, we indeed have barely better choices. To many companies, competing a better Google rank is a premier task. Isn't it a pity to make the life miserable because of some non-existing "God?"

This is not the first time in history we humans have this experience. In almost all the ancient countries, our ancestors had worshiped various non-existing Gods and begged for vague oracles from them. During the period of these dark days, formal education was not prevailed. Knowledge was primarily holden in the hand of few priests who were "servants" to the various Gods. These priests produced vague oracles using the names of these non-exiting Gods to control the mind of people. If even priests could not answer a question, normal people at the meantime had to believe that there were no answers to the question. Asking priests was the primary way to access knowledge of the unknown world. This scenario is exactly the reflection to the current stage of web search.

How did our ancestors get out of this darkness? The answer is one word --- education! The prevalence of formal education liberated humans out of the control of vague oracles. When knowledge education was prevailed, people found new ways to look for the unknown. Rather than to be new priests who generally know everything, the new knowledge education system educated normal persons to be specialists of various domains. Therefore, people no longer needed to look for vague oracles from priests. In contrast, they looked for much more clear and precise advices from particular domain specialists. They no longer laid the hopes on the shoulder of the man-made Gods. They started learning to help each other by sharing individual knowledge. This was the great Renaissance. Such a formal education and social system became the basis of our modern society.

The evolution of web search will follow a similar path. Sooner or later, the "Gods of the Web" will gradually step down from the stage. Though the algorithms about centralized web search are improving, the speed of knowledge accumulation is much faster than the improvement of the algorithms. This was the essential reason why in old days the priests could no longer pretended to be Gods. The set of knowledge was simply become too much to be holden by small groups. Educating everyone became the sole solution, and it did work.

What is an educated web? The educated web is the Semantic Web. To build this educated web, we need to prevail the cognition of eduction. On the Web, it means the cognition that we need to educate machines. The progress of constructing semantic web is not to build a system. No, we do not build a system; but we educate individuals and the network of these individuals automatically is an educated web.

The specialist-based search (i.e. a combination of vertical search and collaborative search) will gradually replace the oracle-based search. Engaged with more and more user-generated tags, the web is moving towards this direction. What is still lacking now is advocates of web education of machines. But we are certainly on the right track.

Final Address

Don't try to build another Tower of Babel. Our ancestors had tried once, and the attempt was failed. If the physical tower failed, could we succeed in building a mental Tower of Babel? I doubt it strongly. Constructing the Semantic Web is not to build a Tower of Babel. It is to simulate our modern human society. It is impossible to achieve "Semantic Google." But Semantic Web will still be available. At the day, however, there are no noble seats reserved for an oracle-announcer Google.

For readers who like to know more about this vision of web evolution, the Part 2 of the web evolution article is an old but coherent version of the view. A newer version is in this blog, the series of "A View of Web Evolution." (Follow the tag web evolution and you can get it.) Any comments are far more than welcomed.


Trackback:

Saravanan's post: Beat Google!

Friday, March 23, 2007

retrievr: an interesting progress on image search

retrievr is a new image search service that let users find flickr images by drawing rough sketches of them. It is not searched by keywords or meanings. In contrast, it is searched by visual effects. For example, when I upload my own photo to the site, it returns me a set of black-and-white pictures that has similar visual effect as my picture.



I am not sure how this technique can be used in real-world cases. But it is fun to play with it. If retrievr creates a Web-2.0 community and allow users voting and sharing their results, it may greatly help them improve their technique and bring up more creative ideas on how to use this cool technique on real world applications.

Thanks for JurijMLotman pointing this site and its mother site SystemOne to me. They are doing really cool stuffs. The web indeed becomes more and more interesting.