Thursday, June 21, 2007

Moving toward machine processing---the certain destiny of web evolution

This is indeed not new. But I had some strong feeling to say something after reading a new post at the Read/WriteWeb. In the post, Alex Iskold discussed a new, but common, phenomenon after the rise of Web 2.0---attention distraction.

Nevertheless Web 2.0 provides us powerful facilities to build virtual social network on the web, there is downside of this advancement. Unlike the previous ages, we become more and more easily "being caught alive" on the web. As a result, we are distracted regularly. We may often have to interrupt the normal work flow to handle exceptions, an inevitable negative side-effect of "being popular." This phenomenon is addressed as the problem of continuous partial attention.

This "continuous partial attention" is a very interesting issue, especially when we think of it by the view of web evolution. In this view of web evolution, we analogize the progress of World Wide Web to be the growth of humans. In particular, we have analogized Web 1.0 to be a society of newborns, Web 2.0 to be a society of pre-school kids, and the ideal Semantic Web to be a society of educated people. In fact, this analogy can also well explain the reason of "continuous partial attention" on the web and foresee how this issue could be solved gradually with the evolution of WWW.

We are seldom interrupted by newborns. In fact, though newborns may cry, we can ignore them if we want because they do not have the ability to interrupt our normal work flow. On Web 1.0, machines can deliver emails (a type of interruption) to us at any time. But we can choose ignore them at run-time and only choose to take care of them in our scheduled time. Our normal schedule is kept as usual in the environment of Web 1.0.

When children grow up, parents start feeling pain of "continuous partial attention" caused by their kids. Especially pre-school kids, they still do not have much ability to do things by themselves. But unfortunately (or fortunately), they have learned limited knowledge and started to request. They ask questions and require accompanies playing with them. Moreover, they deliver messages. Thought this is often thought positively, these messages are indeed irregular interruptions because these kids often ask for the highest priority to the handling of their delievered messages. This is what we have encountered at present, as in Alex's post, the issue of "continuous partial attention" on Web 2.0.

On Web 2.0, machines have been augmented by limited knowledge. They are equipped by various widgets and active functions. At run-time, we (as virtual parents of these machines) are often interrupted by the messages delievered by these kids from the other parents (other web users). The prevalence of Twitter only worsens the already disturbed schedule. We are often caught alive online; and we often have no choice but interrupt our regular work flow to handle these exceptions so that we can maintain a good relationship in the constructed online social network. This is a pain to have growing-up children; and this is a pain to all Web-2.0 dedicators.

How to solve this problem? Certainly we do not want our children going back to their newborn stage. As well, we certainly do not want to switch back to Web 1.0 or shut up ourselves from online only to avoid this "continuous partial attention" issue. In contrast, we want our children to grow up and start to be able to handle things, from simple to complicated, by themselves. For humans, this process is called education; and the people after this process is called the educated people. For the web, this process is called annotation (or adding semantics); and the web after this process is called the semantic web. We need to educate machines. Let them understand semantics, from simple to complicated. This is the certain direction of web evolution.

In summary, continuous partial attention is a certain side-effect in the process of web evolution. In this Web 2.0 stage, the severity of this problem will reach its peak. But this problem will be gradually solved during the process of web evolution when more and more machine-processable semantics are added to the web. Though it may not be solved totally (just like in our real life we cannot totally avoid being interrupted), it would not be a serious problem in the future web with rich machine-processable semantics.

Trackback list:

*** Continuous Partial Attention: Software & Solutions

*** Dealing with partial attention issues

*** Supernova 2005: Attention

Monday, June 18, 2007

Yahoo! had a new CEO. Can Jerry Yang lead the company to a new level?

A very recent news, Jerry Yang has replaced Terry Semel to be Yahoo's CEO. Though Yahoo! had survived from the dot-com bust under the lead of Semel, its influence declines conspicuously in recent years accompanied by the rise of Google. Can the crowning of Yang slow down or even reverse the decline of Yahoo? It is hard to tell at this moment. But one thing is for sure---Yang will have a long to-do list to accomplish. In a recent survey, up to now (people still can vote at this moment) 44% of people believed that Yang might not be the right choice for Yahoo CEO; in contrast to that only 22% voted yes.

Indeed, I don't think that the fate of the battle between Yahoo and Google will be any difference only if Yahoo has gotten a new CEO, even if this one is Jerry Yang. Personally, I have great respect to Yang for his great vision of founding such a great company (Yahoo) in history. But the once glory of Yahoo has gone with the rise of Google. Yahoo had once established a new model of web search. But Google perfected this model to its ultimate. Once upon a time, Google was a little follower of Yahoo. Except of the PageRank algorithm, Google followed everything Yahoo had invented. But now, many evidences show that Yahoo is following Google. Google has invented so many great applications by facilitating searched web resources. It is even difficult for Yahoo to follow up; let it alone beating the ambitious Google. At present, the opponent of Google is no longer Yahoo, it is Microsoft.

In the second part of my web evolution article, I have presented a brief study of the rise of Google, as well as the decline of Yahoo. Along with several of my previous posts, I believed that the future of Yahoo is lay on a complete new vision of web search. I doubt anybody could beat against Google any more underlying this traditional web search model. Google has executed it too well. And this traditional web search model allows the winner taking all the shares eventually. Yahoo must figure out a new solution, an alternative solution. Otherwise, Novell's present would be Yahoo's future. Novell was once a company leading the world in the network realm, and it had the power to decide the fate of others. But now Novell becomes barely more than a normal middle-size company that is struggling its own survival among the big brothers (once he was one of them).

Can Yang again lead Yahoo to a new route as he had done several years before? Can my vision of web evolution be realized by Yahoo? Best wishes to Yang and Yahoo!

Wednesday, June 13, 2007

Semantic search has two legs

The discussion of semantic search has gradually become popular. Just not long time ago, semantic search was thought to be barely a little bit more than a dream. At present, optimistic researchers have started to believe its possibility in the near future. Very recently at Read/WriteWeb, Dr. Riza C. Berkan, the CEO of Hakia (a company declared to perform "semantic search"), posted an article about semantic search that attracted much attention. Despite of agreeing with the post, here are more thoughts about semantic search.

Semantic search has two legs: semantic understanding and proactive collaboration. Until now, however, most semantic search articles only have focused on the first one. Including Hakia, an "ideal" semantic search engine is popularly thought to be alike a "semantics-enhanced Google." This is, however, a narrowed thought. The intension of semantic search is more beyond "semantics + search."

In order to better understand these two legs, we may watch a regular semantic search scenario in human society that is, however, often overlooked. We humans have daily practised a type of semantic search very successfully for centuries. We ask questions; everybody asks questions, from children to adults. We ask questions to look for answers. These question-and-answer behaviors are typical semantic search activities.

When we are young, we look for answers from parents, whose words are oracles to us. When we grow older, we look for answers from teachers, whose words are oracles to us. When we grow even older, we start to realize that there are indeed no oracles. We start to look for answers by ourselves. In particular, we make friends with various specialities. These friends become our sources of question answering when we get troubles in particular realms. At the meantime, we ourselves also become such a type of sources to our friends. These links constitute a delicate, complicate, and successfully executed network of semantic search in our human society.

If we take a closer look at this successful semantic search network, we can find two fundamental factors that support its execution. First, its success relies on the ability of semantic understanding at each but not some of its nodes. It is generally believed that the set of human knowledge is too rich and too complicated to be executed in a centralized way. For instance, Mor Naaman at Yahoo! Research very recently said that "there is no way that we can engage the masses in annotating media with 'semantic' labels" in a WWW2007 panel. Therefore, representations of global semantics are better to be distributed widely other than be accumulated onto only a few special nodes. In consequence, every node in this semantic search network has its ability to perform a certain level of search depending on its own capability of semantic understanding. This is the basis of a successful semantic search network.

Beyond the local semantic understanding on every node, a successful semantic search network also requires proactive collaborations. In a search network, some nodes (such as professors) may have much greater capability than others (such as first-grade students). But even the node with the greatest ability is still very much limited in its search capability when the search space is about the whole set of human knowledge. A successful semantic search network demands well collaboration among individual search nodes. Moreover, such a type of collaborations appeals to being proactive.

Proactiveness is a unique factor in the network of friendship. The network of friendship is not only a regular social network, but also a search network. When we get troubles, we used to get to our friends for help. Nevertheless, we often make friends on purpose, i.e. in contrast to randomly or aimlessly. A successful semantic search network in human society is priorly built on the joint or depending interest of individuals. For example, both John and Mary love music; so John actively make friend with Mary. Another example, John play piano and does not know to tune a piano; Kate, however, is good at tuning pianos. For the sake of his future requirement of piano-tuning, John proactively make friend with Kate. The third example, Rose is good at history literature; John, however, does not like to read history literature. In consequence, John inactively make friend with Rose. These examples show that the establishing of a search network very much depends on the proactiveness (which in turn decided by the semantic understanding of interests) of these nodes to make connections.

In summary, semantic search naturally contradicts to the centralized web search strategy. In order to activate semantic search to the practical level, we need a search network that is participated by all web users beyond the few independent and aggregated semantic search nodes such as Hakia. The entire web search strategy must experience some revolutionary change other than simple makeups. In the second part of my article about web evolution, I have more discussions about the collaborative search for the future semantic web.

The initial draft of this post is published at SemanticFocus.

Tuesday, May 08, 2007

Web Search, is Google the ultimate monster?

(Revised at Sept. 29, 2007)

If investing on web technologies is buying jewelry, web search technology is the most brilliant diamond on top of these jewelries. A recent post about Top 17 Search Innovations Outside Of Google at Read/WriteWeb attracted hundreds of click and dozens of comments again. Google or Post-Google, this question attracts eyeballs.

New companies may beat Google, but not by the Google way. At present, this claim represents not only innovation of search technologies, but also revolution of basic web search strategies. As I discussed in an earlier post, unlikely we can build another, more advanced "semantic Google" to beat the current Google. To defeat Google in real, the basic strategy of web search must experience a revolutionary change.

The current web search strategy can be summarized as the oracle-based web search. When web users search information, they look for oracles from the "Gods of the Web." These "Gods" are search engines, which are assumed knowing about the Web way more than us as normal persons know. Among these "Gods" the greatest one is Google. In the current WWW, Google is the "God." We beg for the oracles from it to access expected web resources. If there are no answers from Google, by convention we just believe that there are no answers to our question on the Web. This scenario typically shows how much we have trusted and relied on this "God." Although nearly all of us know that Google often makes mistakes and even Google can only search a comparatively small portion of all web resources, we indeed have barely better choices. To many companies, competing a better Google rank is a premier task. Isn't it a pity to make the life miserable because of some non-existing "God?"

This is not the first time in history we humans have this experience. In almost all the ancient countries, our ancestors had worshiped various non-existing Gods and begged for vague oracles from them. During the period of these dark days, formal education was not prevailed. Knowledge was primarily holden in the hand of few priests who were "servants" to the various Gods. These priests produced vague oracles using the names of these non-exiting Gods to control the mind of people. If even priests could not answer a question, normal people at the meantime had to believe that there were no answers to the question. Asking priests was the primary way to access knowledge of the unknown world. This scenario is exactly the reflection to the current stage of web search.

How did our ancestors get out of this darkness? The answer is one word --- education! The prevalence of formal education liberated humans out of the control of vague oracles. When knowledge education was prevailed, people found new ways to look for the unknown. Rather than to be new priests who generally know everything, the new knowledge education system educated normal persons to be specialists of various domains. Therefore, people no longer needed to look for vague oracles from priests. In contrast, they looked for much more clear and precise advices from particular domain specialists. They no longer laid the hopes on the shoulder of the man-made Gods. They started learning to help each other by sharing individual knowledge. This was the great Renaissance. Such a formal education and social system became the basis of our modern society.

The evolution of web search will follow a similar path. Sooner or later, the "Gods of the Web" will gradually step down from the stage. Though the algorithms about centralized web search are improving, the speed of knowledge accumulation is much faster than the improvement of the algorithms. This was the essential reason why in old days the priests could no longer pretended to be Gods. The set of knowledge was simply become too much to be holden by small groups. Educating everyone became the sole solution, and it did work.

What is an educated web? The educated web is the Semantic Web. To build this educated web, we need to prevail the cognition of eduction. On the Web, it means the cognition that we need to educate machines. The progress of constructing semantic web is not to build a system. No, we do not build a system; but we educate individuals and the network of these individuals automatically is an educated web.

The specialist-based search (i.e. a combination of vertical search and collaborative search) will gradually replace the oracle-based search. Engaged with more and more user-generated tags, the web is moving towards this direction. What is still lacking now is advocates of web education of machines. But we are certainly on the right track.

Final Address

Don't try to build another Tower of Babel. Our ancestors had tried once, and the attempt was failed. If the physical tower failed, could we succeed in building a mental Tower of Babel? I doubt it strongly. Constructing the Semantic Web is not to build a Tower of Babel. It is to simulate our modern human society. It is impossible to achieve "Semantic Google." But Semantic Web will still be available. At the day, however, there are no noble seats reserved for an oracle-announcer Google.

For readers who like to know more about this vision of web evolution, the Part 2 of the web evolution article is an old but coherent version of the view. A newer version is in this blog, the series of "A View of Web Evolution." (Follow the tag web evolution and you can get it.) Any comments are far more than welcomed.


Trackback:

Saravanan's post: Beat Google!

Sunday, May 06, 2007

Evolution of Web Links, another direction of thoughts

Danny Ayers had his most recent column at IEEE Internet Computing: Evolving the Link. In this article, Ayers summarized his vision of how web links evolve with the progress of World Wide Web. Agreeing with his vision, I, however, have some supplementary thoughts about the evolution of web links.

Before presenting my supplementary thoughts, I would like to briefly review what Danny presented about web link evolution. In general, Danny's vision followed a strict technical line. At the beginning, web link were anonymous connections that linked one web page to another. Using his example, "< href="http://creativecommons.org/licences/by/2.0/">cc by 2.0< /a>" produces an anonymous connection from the page that contains this specification to a particular web location: "http://creativecommons.org/licences/by/2.0/". Though each destination has its distinctive meaning, web links themselves mean nothing else except that the referenced web resources are ABOUT the local text. The meaning of ABOUT is, however, simply too rich to be properly distinguished.

To solve the problem, the "rel" attribute is designed so that it describes de relationship from the current document to the destination resource. For example, "< href="http://creativecommons.org/licences/by/2.0/" rel="license">cc by 2.0< /a>" shows the meaning of the link to be "license." Again, this solution still has its problem because the meaning of content inside the rel attribute is often not machine-processable.

In this column article, Danny presented that a potential solution to this last problem is to treat links as data. Thus, we map not only data but also links to proper RDF descriptions. With these mappings, machines can automatically interpret the meanings of not only resource nodes on the web, but also the links among these resource nodes. This vision of web links thus concluded the article.

Nevertheless I agree with Danny's vision, I feel something else also important to the evolution of web links but missed in discussion at this article. Besides semantic meanings, vulnerability is another essential aspect of web links. Traditionally handcrafted, anonymous web links are often vulnerable to individual prejudice. For example, I have produced a normal web link from this post to Danny's blog, which shows a relation of this post to Danny. Meanwhile, I can also subjectively remove this link (but not the content) so that there becomes no immediate link from this post to Danny, though indeed there should be such a link since the content is not changed. This simple example shows a basic problem about traditional links---they are vulnerable to individual prejudice.

Though to solve this vulnerability problem is not the intentional driving force of web link evolution, an important side effect of this evolution is the gradual invulnerability of web links. With the emergence of Web 2.0, more and more indirect web links are created based on common tags. For example, I have tagged this post with a keyword "Danny Ayers," while at the same time, Danny's blog is also tagged by the keyword "Danny Ayers." Therefore, even if I have removed a normal href-style direct web link in my post to Danny's blog, there is still an immediate link between this post to Danny's blog because we have shared a common tag. Due to the objectiveness of tagging, this type of web links becomes less vulnerable to individual prejudice than the previous purely handcrafted web links.

When the web evolve forward to the ideal semantic web, we can predict the creation of more and more objective links between web resources. These links are going to more and more based on common meanings (objectiveness) rather than individual preferences (subjectiveness). As Danny presented in his column, when links are shared as open data annotated by formal taxonomies, the existence of these links becomes less and less vulnerable to individual prejudice. It is the content itself that would decided the links.

This evolutionary aspect of web links is important to not only web links, but also the WWW in general. It means that the web is going to be weaved not only in more and more details, but also more and more objectively. Based on this conclusion, we can make several interesting predictions about the future web.


  1. Tags and annotations are going to be premier resources on web search. Due to their objectiveness (less vulnerable), the network composed by tags and annotations is going to be more stable than the network composed by handcrafted links. This fact may significantly help improve the efficiency of web search.

  2. The weight of tags and annotations are going to be more and more critical on ranking research results. The balance between these weights (the side of objectiveness) and link popularities (the side of subjectiveness) will be an essential issue on new web search ranking algorithms. Does it mean the dusk of the PageRank algorithm and thus the declining of Google? We are not sure yet. But at least this is not a positive news to Google.

  3. Tags and annotations are going to weave the web into close-related communities with distinct topics. As a result, vertical search engines may replace horizontal search engines becoming the basis of web search. At present, vertical search is relied on horizontal search and then provide more details on particular domains. In the future, horizontal search will be based on vertical search and then provide more details on cross-domain communication. This role switch on web search may significantly affect the structure of web industry, especially the web search businesses.


I have more discussions on web evolution in general in my article of Evolution of World Wide Web. The most recent post is the Part 2, Web Evolution Theory and The Next Stage. In this part, we studied several web evolution laws and composed them together to be a basis for predicting the evolutionary future of World Wide Web. Though these are only our viewpoints, we hope it brings some fresh air into the study of web evolution.

Wednesday, May 02, 2007

A fantastic map of a world of the web

Randall Munroe of xkcd.com has drawn a map of online communities as a world of the web. Nevertheless that I believe we are going to explore more land of unknown in this world of the web, this is a very creative artifact. I think I would like to have a T-Shirt about it.

Evolution of World Wide Web, Part 2, Theory and the Next Stage

I have uploaded the second part of the article---Evolution of World Wide Web on April 27. In this article we start with the discussion of several exciting web evolution laws. Then we apply these laws to predict the next stage on web evolution, which may be called the Web 3.0.

This article is still in progressing. Please leave your comments after you read it. I will update this post and the article regularly.