Showing posts with label web thread. Show all posts
Showing posts with label web thread. Show all posts

Sunday, August 09, 2009

The Link in Linked Web

Kingsley Idehen posted a thoughtful article on URI, URL, and linked data this weekend. In the style of Q&A, the post concisely answers some of the most confusing questions about linked data. It explains the subtle distinction between URI and URL when dealing with the linked data. Moreover, the post implies that "a new level of Link Abstraction on the Web" is likely needed for us in order to efficiently consume the linked data Web.

After I left a comment for the post, however, I feel the issue deserve a second thinking. Before approaching forward, I pasted the main section of my original comment to Kingsley's post in the following.

Another thought I have, however, is that we may have three, in contrast to two, fundamental definitions on describing the Web. The two well-known ones are data and service; or in RDF we define Class and Property respectively. Until now, we assert the third one---link---to be nothing but a special form of data. The reality is, however, that this special form is so special that we may consider to give it a little bit honor so that it becomes the third member of the fundamental building block of the Web. That is, a link is not a data, and nor is it a service, but a link. Or with respect to your post, a URI is not a data, but a form of link, PERIOD.

I believe that this distinction, once it is made, could be important as well as valuable. A trick thing here is that, following this distinction we can start to think of other forms of links that is beyond URI (which is just a binary model). By contrast, we may start to invent the links in higher order, such as the link of links (metalink) or the thread of links (group link). Be honest, if the Web is moving towards a web of linked data (and I believe so since the Web data is more and more interconnected), we must breakthrough this traditional thinking of the link model. The key is, however, from today we start to think link to be link but not a data.

World Wide Web: from the dualistic view to the ternaristic view

The thought that Web link is a fundamental element of the Web that is independent to data and service was originated when I wrote a model of Web evolution. (Actually, it could be traced back to January 2007 when I first started to think of how the Web evolves.) By observing the evolution of the Web, more and more I felt that data, service, and Web link are three equivalently fundamental elements of the Web. This interpretation of the Web is different from the classic dualistic view of the Web in which the Web is said to be built upon two fundamental first-class entities: data (which expresses the static description) and service (which expresses the dynamic action). In this classic dualistic model, Web link is a special second-class member that is partially static description and partially implied by dynamic action.

RDF is a typical design according to this dualism philosophy of the Web. In RDF, relation (RDF:Property) is a first-class entity along with the normal object entity (RDF:Class). While class expresses the static fact of the Web, relation expresses how the static facts are interacted to each other. Both the elements are equally fundamental. Two models with the identical classes may not necessarily be equivalent to each other since the properties that are among the classes could be different.

Now we need to start discussing a few subtle implication of this philosophical view of the Web.

By the dualism philosophy, a relation is first of all a service and secondary a link. For example, suppose there are two statements: (1) Mary is a teacher of John, and (2) Mary is a friend of John. Mary and John are two objects. "teacher-of" and "friend-of" are two relations. Philosophically, however, the primary meaning of each of the relations is a typical service defined in between the two objects. In the first relation, Mary provides a teaching service for John, by which a teacher-of relation is established. In the second relation, Mary provides a friendship service for John, by which a friend-of relation is established. Be note each of the links is a consequence of the respective service (and there could be other consequences as well) in contrast to a prerequisite of the service. The dualism philosophy tells that service implies link and every link must be a consequence of a service. Moreover, no link actually makes sense if no services imply the link. Every link has a reason, which is a known service, conceptually (means that the service is unnecessarily implemented, however).

There is also another side of Web link according to this dualism philosophy. Once a link is implied by a service, it becomes a data. Unlike service that always leads to an action or a production, link describes certain static fact, which by the dualism philosophy is a data.

Therefore, link, which inherits the features from both of the first-class entities, is a special secondary entity in the dualistic view of the Web.

The ternarism (3 fundamental elements to model the world) philosophy to which I prefer may describe the same Web but in a different picture. By this philosophy, a link is not the consequence of a service and neither is it an unique type of data. A link is a link, which in the ternaristic view of the Web exists without the need of being implied by a service or being stored in the form of data.

In the dualism world, wherever there is a link, it must exist a data that represents the link and a service (implemented or not) that implies the link.

In the ternaristic world, however, when there is a link, it may or may not exist a data that represents the link, and it may or may not exist a service (let it alone implemented) that implies the link. A link is nothing but a pure connection among (could be more than between) the things.

In the dualism world, that one thing is linked to another thing is always due to some reason. In the ternaristic world, however, link is a matter of natural connection that does not require a reason to be existed. A link is as fundamental as a data or a service.

As we know, the Web is a world of information. Following this ternarism philosophy, the Web we understand becomes different world from what we normally think by the dualistic view. It tells that in the world of information, data reveals the encapsulation of information, service reveals the action and production of information, and link reveals the transportation (in contrast to connection) of information. Under this view, any Thing in the information Web is composed by three fundamental elements---data, service, and link. The data elements contains the information, the service element enables the production of the information as well as the interaction of the information to the other things, and the link element determines whether or not the information being able to be passed to another Thing.

By the ternaristic view of the Web, when we say there is a link from Thing A to Thing B, it means that the information carried by Thing A can be directly transported to Thing B without the help of any other information carrier.

By the ternaristic view of the Web, when there are no links between Thing A and Thing B, it means that unless there are additional information carrier participated in the transaction, the information carried by A cannot be passed to B. Once properly the additional information carriers joins the protocol (possibly in both sides), a link in higher order can be established between A and B.

By the ternaristic view of the Web, there is always a link (i.e. a direct link in the classic mean) between any two things though the link is often in higher order, i.e., it is often not a binary link that involves only the two designated things.

The regular Thinking Space readers may have found that this ternaristic view of the Web is also influenced by the quantum theory. Unlike that in the dualistic presentation of the Web we often need to perform an expensive computation to discover a link (concatenated by several direct binary links) between two objects in the Web, in the ternaristic presentation of the Web any two objects are directly linked, but possibly linked in varied orders. Moreover, I realize that we may directly apply many classic quantum theories to the Web if we start to think of the Web in the ternaristic view, which I will share later in the other posts.

Does the ternaristic view actually reveal the more intrinsic fact of the Web? I do not know. But there is one thing I feel certain. Link is not a simple issue. By better understanding the nature of link in the information world, we may eventually release the tremendous power of computation that we might not even imagine now. For the companies that aim to monetize linked data (such as Kingsley's OpenLink Software), it would be even more valuable for them to rethink the nature of the links that they are working against every day.

Monday, July 28, 2008

Cuil Search

Beginning from the last night, Cuil (pronounced "cool") starts to hit the headline of many technological blogs.

New Design of Interface

Cuil is experiencing a new design of magazine-style search-result display interface other than the "standard" list-style display of Web search results. The following is a screen shot by typing my name into the Cuil search. This change of design potentially may mean much more than attracting eyeballs.



In an earlier post at Alt Search Engines, I shared that a critical but often overlooked issue in the current Web search is the production of link resources, i.e., how to better formulate the generated links according to the user search requests. To clarify the issue, let me explain it using a metaphor. If one has a brilliant pearl but present it inside a crude lunchbox, how good it might be known? A brilliant pearl needs to have a well-designed box (such as the right one) that matches its superior quality. It is exciting to watch a breakthrough on this issue by Cuil.

This new design provides Cuil lots of potential. For example, each of the short related story shown at the first page of result could be more than just Web links. By contrast, it may be an entry point to a related Web thread and each story is describing the theme of the thread. A combination of Web search and Techmeme-style stories may bring Cuil users very different experiences from using, such as, Google for search.

New ranking and privacy

Another exciting thing to watch is that apparently Cuil has performed a different algorithm on ranking its search results. After Google's success, ranking based on objective link popularity is generally accepted as the foundation of page rank on the Web. Cuil, however, seems trying to apply a new standard of ranking over its stored over 120 billion (as it claimed) Web pages. As the result, from the previous figure we can see that the rank of my related links is very different from the rank of links returned by Google. Discarding the performance until now (anyway, Cuil is launched barely for one day but Google has optimized its results for more than a decade), I strongly support Cuil's attempt. We want to have an alternative solution which does provide us DIFFERENCE. On the other hand, if a new search engine ranks the Web the same way as Google, how could we be convinced that it might do better than Google? Therefore, no matter whatever Cuil has chosen a correct path to walk and we wish it the best luck in the future.

Cuil also claimed some exciting news about its advanced privacy protection technology. Danny Sullivan on Search Engine Land uncovered that Cuil search engine would not log IP information. Be note that Google, Yahoo and Ask.com all perform this IP log in their search engines. Cuil's claim helps protect users' privacy on both of the publishing and surfing on the Web. I recommend this improvement especially to the places where information censorship is rigorous.

Still long way to go

Despite of all the improvements, Cuil still have a long way to go before it may indeed threaten Google. For example, it seems that search engine repeatedly looks for the same links and there must be some severe bugs about how to break a circles in graphs in its algorithm.

The following is the second page of the Cuil search results by input my name. Comparing to the first page in the former figure, we can see that the second page repeats quite a few links that have already shown in the first page. I have tested Cuil by the other queries and it seems that this is a bug generally occurred.



Certainly Cuil has more than this problem. But I am still looking forward to its future. Although Sullivan only showed "cautiously optimistic" to the future of Cuil, I think the value of Cuil would not be the precision of its search results. I agree to Sullivan that it is hard to believe that Cuil can do significantly better on bringing back more accurate results than Google, Yahoo, and Microsoft by employing the similar infrastructure of Web search. On the other hand, however, Cuil does show us that it may help build final link resources in better quality. Only if Cuil can continuously convince people that it can bring people alternatives (even though no better results) that they can hardly get from Google, Cuil would be a success at the end because we want to hear different voices.

Tuesday, July 08, 2008

Invariants on the Web

Invariant is something that does not change under a set of transformations. The picture on the right shows Pappus’s Invariant in geometry. The invariant tells that by following certain rules the three intersection points shown in the figure are always collinear no matter how people may draw the two lines and locate ABC and DEF in the lines respectively.

Invariant study is fundamental to any scientific research, especially when the research domain is as complex as World Wide Web. Invariants are supposed to be constant within the specified research scope. By well understanding the invariants we may effectively improve the knowledge over many complicated issues. Therefore, it is unsurprisingly for us to see the discussion of invariant study in the new Web Science Research Initiative.

In "A Framework for Web Science", the flag article of Web Science, Tim Berners-Lee and his colleagues have carefully studied several invariants on the Web. In particular, one invariant is outstanding among all the others. The one is URI (Uniform Resource Identifier). In the paper Berners-Lee et. al. had focused on discussing which invariant represents the binding of semantics with declared objects. There was no final best solution concluded in the paper, however, the one closest to the best was URI.

In varied programming languages we have widely used an invariant, which is declared name. In programming languages such as Java or C++, "each unique object (i.e. with distinct semantics) is declared with a distinct name in one program. By referencing a name, a program accesses the semantics behind the name." Hence declared name is taken to be invariant.

On the Web we are currently using another invariant. "Web researchers decide to use location binding to solve the problem, i.e. URIs and URLs. By default, identical URIs reference the same semantics. Identical URIs on web is the same as identical declared names in programs. However, the name of this URI is varied, i.e. name is no longer an invariant. In constrast, URI becomes a new invariant."

The authors, however, pointed out that indeed neither of the two was proper invariant on Semantic Web (or on the future Web). "The difference is, however, that the requirement of machine understanding," said by the authors. We actually have no ways to promise the consistency of the meaning to which a URI points. It is the same as we cannot enforce users to consistently bind the same name to any unique object on the Web.

Although with the problem, the authors did not provide a satisfactory answer to the problem in their paper. By contrast, they simply emphasized that "W3C suggests that do not transfer URI to another object. That is, whenever you create an object, giving it a unique URI. This requirement is thus the same as that whenever we create a new object in program, make sure we give it a unique name." In other words, please do not change the referred destination of any URI though anybody has the right to perform such a change. This passive resolution is not a satisfactory answer. Deprecated URI has gradually become a severe problem when more and more Web applications start to assume URI to be invariant on the Web. May we have an alternate, active answer to the question?


The figure above shows three basic components when we bind semantics with certain object. They are the declared name, the object itself, and a link connecting the two sides. So which one of them is truly invariant when they are presented on the Web?

As the paper has discussed, neither the declared name nor the link (i.e., uri) is true invariant. "Apple" may be fruit or a software company. We have no way to restrict a handpointing to a fixed destination.

The only exceptional one is the object itself. Although by nature an object can only be itself and it is automatically an invariant to itself, how can we present this invariant besides name and link? This is thus the problem.

We humans have so customized of binding semantics with declared names that we have almost forgotten some more intrinsic binding beneath the surface.

When we are binding the declared name "apple" with the object apple, we are actually making a semantic computation in our brain such as to determine whether it is a fruit with red or yellow or green skin and sweet to tart crisp whitish flesh. For people, a name is not just a name, but also a computational procedure in human brains. It is actually not the name that identifies an object, it is the procedure that identifies the object. The declared name is only a named shortcut referred to the particular procedure in brain. When we convert the procedure to machines, it is an epistemological process.


The picture above shows the new paradigm of semantic binding on the Web. The left side is changed to a particular epistemological procedure (which could be implemented in various ways such as the one we have suggested). Unlike names, these procedures are unique since they can unambiguously answer either yes or no for any identification request. Based on these epistemological procedures, Web links (such as URIs) are upgraded to be Web threads. The Web threads connect the same Web into a varied layer. Moreover, from the philosophical and economical aspects the construction of epistemological procedures and Web threads would be the basis for the production of mind asset.

In summary, epistemological procedure and Web thread are invariants on the Web. Through Imindi, we are going to demonstrate the world something extraordinary happening on the Web.

UPDATE: related reading about URI, "What do people have against URLs or URIs?" by Kingsley Idehen.

Saturday, December 22, 2007

Thinking Space 2007 in 12 months

This post is the highlight of what was on Thinking Space in 2007 month-by-month. I am grateful to all the readers of Thinking Space and wish you merry Christmas and happy new year!

January 28, 2007, Web 2.0 panel on World Economic Forum

How would Web 2.0 and the emerging social networks affect world business? The annual World Economic Forum at Davos organized a panel with five outstanding Web business leaders addressing this issue at the beginning of 2007. The talks, however, showed that the executives from traditional big companies such as Microsoft and NIKE were less alerted to the new technologies than the executives from new-age companies such as YouTube and Flickr. In short, both Bill and Mark were talking in languages other than Web 2.0. By their viewpoints, the Web-2.0 phenomenon was certainly less important than their own imaginary vision towards the future. What web evolution really impacts world business was severely underestimated.

At the end, the speech by Viviane is worth of re-emphasizing. When the Web evolves to be more and more mature, who are going to govern the virtual world? This question may gradually become a severe issue when web evolution goes further. Will there be conflicts between the virtual world governments and the real world governments? I do not think that in 2008 we will immediately see this type of conflicts. But the traditional means of national boards do have started to diminish while the new means of digital boards are forming; these changes are slowly but inevitably.

February 18, 2007, The Two-Year Birthday of AJAX

Few technologies have affected the Web so much as AJAX has done. AJAX is more than a technology; it is a philosophy. What AJAX really does is to decompose Web content into smaller portable pieces that are feasible to be uploaded and updated independently. AJAX prompts the dynamic recomposition of pieces of Web content from varied resources. Hence it significantly improves the reuse of information on the Web.

The prevalence of AJAX causes the fragmentation of the Web. The reverse side of this phenomenon is, however, how we may defragment the small pieces of information and reorganize them from end-users' perspectives. This defragmentation issue is the next critical challenge for Web information management. Twine is an example that has started to address this issue. I expect to see more proposals to solve this defragmentation issue in 2008.

March 23, 2007, Will the Semantic Web fail? Or not?

Whether the Semantic Web is going to succeed is always debatable. There are many supporters of Semantic Web, and there are nearly as many as the opponents as well. Will Semantic Web become true? The answer partially depends on whether the Semantic Web researchers can humbly learn from the success of Web 2.0. The normal public might not welcome Semantic Web if its research is still kept inside the ivory tower. Practices such as Microformat are good examples that the Semantic Web research approaches normal web users. But there are still too few of this type of examples. For instance, will the new W3C RDFa proposal be too complicated again? We don't know yet. Hopefully this time W3C would focus more on simple solutions that are feasible to normal users rather than on sound and complete solutions that the academic researchers favor. In comparison, if our real human society is far less than being perfect in reasoning and inference, why must we have theoretically perfect plans to build a virtual world?

April 18, 2007, New web battle is announced

Google is expanding rapidly. Google had replaced Yahoo being the leading Web search engine. Google has already been the largest site that produces Web-2.0 products. Google is competing against Microsoft to be the leading online document editor. Google is fighting against Facebook to be the leading social network through the OpenSocial initiative. More recently, Google starts another battle against Wikipedia to be the leading online knowledge aggregator by the announcement of Google Knol. Can Google succeed simultaneously in all of these fields? Are Google's plans too ambitious to be successful?

The age of Google is about to pass; this is my prediction after watching all these ambitious plans issued by Google. Google has started losing its momentum on originality. By contrast, Google is now repeating a "successful" path of many traditional big companies, i.e., dominating the market by defeating the opponents not by new achievements on technologies but by its superior money resources. This strategy has been proved successfully in many fields. However, it is not a winning strategy on web industry. The reason is that World Wide Web itself is evolving. When the Web evolves, Web technologies evolves. Any company that stops evolving would be thrown away. The history once happened to Yahoo may happen to Google again in the future. The age of Google will be passed with the over of Web 2.0.

May 8, 2007, Web Search, is Google the ultimate monster?

Google is beatable, but Google is not going to be defeated by another Google-style solution. When I predict that the age of Google is about to pass, I mean new revolution on Web technologies. Google is thinking of itself as the God of World Wide Web; and indeed many Web users accept this interpretation (because we have no other better choices at present). But history has already told us that this type of fake gods like Google could not stay forever. In history, we humans abandoned most of the fake gods as soon as the public education system was prevailed. In similar, this history will repeat itself in the virtual world of the Web. The fake God of the virtual world (Google) will step down when the education on Web machine agents prevails. Hakia would not threaten Google if it continues following the Google strategy by addressing itself to be a more powerful fake God on the Web.

In addition to this short summary, I have a preliminary funding request. I will graduate next year and currently I am looking for an assistant professor position. If I'd get an offer, I would start a new research project on next-generation Web search that is beyond the current Google-style search strategy. In fact, I have already done the project proposal. For any reader, if you are responsible on looking for and funding new research projects that are full of potential in the future, I am far more than happy to discuss my project with you. I can be contacted through yihong.ding@gmail.com. The philosophy underneath my new web search strategy can be read at here.

June 29, 2007, Epistemological extension to ontologies: a key of realizing Semantic Web?

The application of epistemology into Semantic Web is less explored than it should have been. We need ontologies to enhance the collaboration and agreements. We also need epistemologies to emphasize the individuality and privacy. I expect more research on this topic in 2008.

July 31, 2007, What does tagging contribute to the web evolution? | An introduction of web thread

There are many ways to describe web evolution. One unique expression is the transformation from the node-driven web to the tread-driven web. Web thread is a new term proposed by myself. In short, a web thread is a connection that links multiple web nodes to a fixed inbound. I observed that the Web was not only syntactically connected by human-specified links, but also semantically connected by latent threads each of which expresses a fixed meaning. A straightforward evidence of the existence of web threads is Web-2.0 tags. On Web 2.0, resources are automatically mutual-connected when they are specified the same tag by individual human users. When weaving these tags together, we obtain an interconnected network of all web pages.

The existence of web threads is an interesting phenomenon that lacks of insightful research at present. From one side, web threads are part of the implicit web because they are generally latent at this moment. On the other side, by proactively revealing web threads and explicitly weaving them, we might produce more comprehensive social graphs for individual web users. This new concept thus may contribute significantly to the vision of Giant Global Graph. I will publish more research on this concept in 2008. By the way, a broader discussion of web links and web threads can be found at here.

August 24, 2007, Mapping between Web Evolution and Human Growth, A View of Web Evolution, series No. 4

World Wide Web is evolving. But why does the Web evolve and how does it evolve? Few answers have been given. The view of web evolution is the first systematic study in the world that directly addresses the answer to these questions based on a theoretic exploration.

This view of web evolution stands upon the analogical comparison between web evolution and human growth. I argue that the two progresses are not only similar to each other by their common evolutionary patterns, but also literally simulate each other from all the major aspects. At present, the simulation mainly happens in the uni-direction from the real world to the virtual world. In the future, however, we are going to see more evidences of simulation on the reversed direction, i.e. from the virtual world to the real world.

The virtual world represented by the Web is nothing but a reflection of our human society. Due to the limit of web technologies, however, we are not able to completely simulate our society from every aspect into this virtual world. In particular, we are not able to well simulate all the activities of individual humans on the Web. By contrast, we can simulate individuals at a certain level within any specific evolutionary stage. This continuous upgrade of simulation of individuals on the Web represents the main stream of web evolution.

This theory of web evolution has published for half a year and I have received many requests on discussing this vision. I hope this study would bring more attention to the fascinating web evolution research.

September 16, 2007, A Simple Picture of Web Evolution

The simple picture of web evolution expresses a straightforward timeline of web evolution. The Web is evolving from a read-or-write web to a read/write web, and eventually it may become a read/write/request web. The implementation of the "Request" operation would be a fundamental next-step towards the next generation Web.

October 7, 2007, What is Web 2.0? | The Path towards Next Generation, Series No.1

What is the next generation Web? This is a grand question to all Web researchers at this moment. We might see critical breakthrough on answering this question in 2008.

At present, the advance of Web 2.0 has already slowed down. The progress of web evolution has reached another stable quantitative expansion period after the exciting qualitative transition from 1.0 to 2.0. The seed of next transition is growing underground now.

In order to figure out the path towards the next generation Web, we need to know the present and where the present was coming from. In the first post of this series "towards the next generation", I summarized the various definitions of Web 2.0. In the following installments at this series, I will continue discussing my vision of the path towards Web 3.0. I feel sorry about the slow progress of this series. I will try to post this series more frequently in the coming year.

November 23, 2007, Multi-layer Abstractions: World Wide Web or Giant Global Graph or Others

Giant Global Graph is a new concept. Although Tim Berners-Lee proposed this concept intuitively for freely deploying personal social networks onto the Web, my view of the intent of this concept is beyond this intuition. In general, I believe that the proposal of this concept is the first sign of a great transition---the organization of web information is transforming from the publisher-oriented point of view to the viewer-oriented point of view.

The impact of this transformation could be greater than we may imagine. Most importantly, this transformation will show that the Web may automatically re-organize its information system without a human-controlled organization such as W3C or Google. World Wide Web is a self-organizing system. This observation is essential to the understanding of web evolution.

December 3, 2007, Collectivism on the Web

The implementation of collectivism has been the landmark of Web 2.0. But do we know how many types of collectivism we may implement onto the Web? This last selected article at December 2007 summarized a few typical implementations of collectivism on the Web. Some of them (such as collective intelligence) have been well known, while others (such as collective responsibility and collective identity) are less known by the public. I expect to watch more creative implementations of collectivism in 2008.

Tuesday, October 23, 2007

Metadata or Hyperdata, Link or Thread, What is a Web of Data?

This is my most recent post at Semantic Focus. In this post I shared my view of a web of data. The following are selected quotes from the article.

A web of data is a network of data whose local characters are specified by metadata and global characters are specified by hyperdata.

A web thread is a reference to a named web location. Unlike a web link, a web thread connects arbitrary numbers of objects at the same time. In contrast to unidirectional, a web thread is omnidirectional. Data in a thread is automatically connected to all other data in the same thread. All the data connected by the same web thread mutually supplements each other in semantics.

With web threads, do we still need web links in a web of data? The answer is yes. Web threads cannot completely replace web links. Web links have their irreplaceable semantics.

But isn't "incorrectness" a synonym of creativeness? If we want to engage collective intelligence in a web of data, allowing and encouraging subjective (and biased) assignment of web links is fundamental to explore human creativity.

In summary, a web of agents is what ordinary users can see about the Semantic Web at the front end, while a web of data is what professional developers understand to be the essence of the Semantic Web at the back end. These two presentations tell a common story from two different sides.

If you are interested in my interpretation about "a web of data," check out the full story at Semantic Focus.

Tuesday, July 31, 2007

What does tagging contribute to the web evolution? | An introduction of web thread

(Revised October 15, 2007)

tagging Tagging becomes popular with the prevalence of Web 2.0. A tag is a keyword or term bound to a piece of information. By default, a tag is assumed to be a correct partial explanation of the bound object. This explanation is unnecessary to be complete, and it is also unnecessary to be about the key characters of the bound object. The purpose of tagging is primarily to associate web content to human conventions.

From the view of web evolution, however, we have another explanation of this Web-2.0 tagging. By this view, tagging is a process of drawing threads across the Web. The activities of tagging convert the traditional node-driven web to be a new thread-driven web; this is a prediction based on web evolution.

Node-driven Web versus Thread-driven Web

A node-driven web is a web whose structure is primarily driven by how web users link their authored web nodes to each other. Albert-Laszlo Barabasi had made great contributions on this field by studying the so-called scale-free network, which is typically node driven. Whenever a new node is added into a node-driven web, its owner decides how to connect this new node to existing ones. This type of decisions is, however, often not random. In contrast, people always tend to link a new node to the most popular node existed in the current web. The chance that a less popular node gets a new link decreases exponentially to its popularity. This theory is fundamental to the growth of node-driven webs.


Web thread is a new term. In my mind, a web thread is a connection that link various multiple web nodes to a fixed inbound. Unlike a standard web link that connects exactly one outbound node to one inbound node, a web thread may simultaneously connect more than two web nodes omnidirectionally. For example, a Web-2.0 tag is a web thread. Other than linking from one node to another, a Web-2.0 tag simultaneously connects arbitrary numbers of web nodes that share the same tag. This is why we call it a "thread" but not a "link" on the Web.

A thread-driven web is a web whose structure is primarily driven by the interconnections of web threads. When a new node is added into a thread-driven web, the web itself automatically decides how to locate this new node by assigning proper threads to it based on the content in this node. Although we do not prohibit user-specified links or threads, the popularity of a web node is determined by the richness of its content (i.e. how many threads are weaved through this node) instead of the number of human-specified links connected to this node.

The popularity of web nodes in a node-driven web is primarily decided by the votes of humans. The most popular hubs in a node-driven web may not necessarily contain rich information themselves. For example, the homepage of Google is just a simple interface without much information. But humans subjectively decide that these hubs are more valuable than many other web nodes.

In contrast, the popularity of web nodes in a thread-driven web is primarily decided by the richness of content in these nodes. The most popular hubs in a thread-driven web may not necessarily be favorite sites for most of the human users. But these nodes definitely contain richest information on the Web for machines to process.

With this comparison, we can see that a node-driven web is human-oriented while a thread-driven web is machine-oriented. If the future is Semantic Web, the Web certainly is evolving from a node-driven web to a thread-driven web. To machines, the popular human favorite sites such as the main page of Google is less valuable because it does not really provide much useful information. In contrast, machines look for nodes weaved by the most number of threads that concentrate on their searched topic. This switch of vision on World Wide Web may eventually change the methods of web search in the future.

Some Potential Impacts to the Future

This re-interpretation of tagging may bring us several positive impacts on developing next-generation web technologies.

(1) This re-interpretation brings us a new picture of World Wide Web. World Wide Web has gradually turning from a random network to be a more and more well-organized network. Before Web 2.0, the model of World Wide Web is widely known as a scale-free network formally introduced by Albert-Laszlo Barabasi. This new interpretation, however, states that underneath the scale-free architecture, World Wide Web actually maintains a fairly organized latent structure, whose backbones are web threads. A scale-free network is significantly dominated by few highly connected hubs. A well-organized thread-driven network, however, is dominated by highly adopted threads. This crucial difference between the two network architecture is a key of developing next-generation web technologies, especially the next-generation web search techniques.

(2) This re-interpretation brings us a new understanding of what web tags are. Web tags are not standard web links. When we produce a new web tag, we are producing a new thread of World Wide Web. The popularity of these threads, however, are determined by their acceptance among web users. On the other hand, the popularity of web thread may still follow the Yule-Simon distribution---a power law relationship, as what the scale-free model follows. These thoughts might be helpful for the further study of web tags and threads.

(3) This re-interpretation suggests us a new way of building a semantic web. Instead of creating semantic-web nodes (as we are creating normal web nodes), building a semantic web is creating web threads and throwing these threads across a network. By weaving these threads, we acquire semantic-web nodes by their intersections.

(4) This re-interpretation brings us a new vision to personalize World Wide Web. Based on this view of thread-driven web, a personalized web becomes a web weaved by personalized threads. By delicately mapping personalized threads to widely adopted threads, we can produce personalized web for individuals. In return, these personalized webs become latent chaos patterns of the entire World Wide Web.

Wednesday, July 25, 2007

Weaving the Thread-Driven Semantic Web

World Wide Web was designed to be a node-driven weaved web. The current web is a set of nodes connected by manually assigned links. In contrast, the ideal Semantic Web (if ever realized) must be a thread-driven weaved web. In such a Semantic Web, the machine-processable semantics are the objective threads that connect all the nodes. In such a Semantic Web, the importance of individual nodes become less and less. On the contrary, threads become fundamental.

I posted a new article at Semantic Focus in which I presented a new vision of Thread-Driven Semantic Web.