Showing posts with label collective intelligence. Show all posts
Showing posts with label collective intelligence. Show all posts

Friday, February 04, 2011

Turing Test and Web Evolution

Machines may have will, but it is not free will.

In 1950, Alan Turing challenged if machines could ever deceive a human judge from being identified non-human by answering questions indistinct to the real human responses. This proposed experiment is now well-known as the Turing Test, an essential concept in the philosophy of artificial intelligence. Now computer scientists generally agree that the Turing Test in a single standalone computer is likely unfeasible. With the rise of World Wide Web, however, people quickly revise the question by asking whether the collectivity of all computers in the world may eventually pass the Turing Test. May a network of computers do the trick? Or, a little bit more dramatically, may one day WWW evolve to be the Skynet?

To pass the Turing Test machines need to prove they may answer questions by free will, a unique character of the living creatures especially humans. In definition, free will is the putative ability of making choices discarding constraints. Constraints generally exist in any situation. The trick of free will is that independent to any other agent's prediction the agent may determine whether or not his decision is bound to any of the constraints. Free will means the ability of breaking rules with no limit and warning. Then a question follows: may we be able to assign this ability to the machines (discarding whether we will if we can)?

Here is my thought to the answer. If we continue building the computers based on the current computer architecture and programming them using the current type of the programming languages, there is no hope of inventing machine free will no matter of the standalone intelligence or the collective intelligence. The reason? The computer architecture and the programming languages themselves have already settled a set of rigid logic constraints. The machines built upon the architecture and programmed by the languages will not work correctly unless obeying these constraints. Please be aware that working incorrectly due to the technology problem is not a free will. Therefore, no true free will can be realized. In theory, we can always develop a set of questions that can cause trouble by the rigid set of logic to distinguish a machine from the real humans.

Now let's retreat half a step. May the machines have will despite it is not free? I believe, however, that the Web evolution is approaching this goal.

Recently in Wired Magazine, there is an article titled "Algorithms Take Control of Wall Street". In it the authors shared how the robots hired by the Wall Street traders instead of the traders themselves are controlling the stock prices. This scenario is a classic example on how the machines start having will, which is programmed by their human masters. In the article the authors also mentioned how some tragical events had happened due to the machines' will, which, the authors argued, was not really humans' expectation or intention. I think the authors may have underestimated the greedy nature of human beings. Behind the numerous rules there sits a word called "greed". Any single rule must be fair. The overall is, however, the consolidation of the greed. So it was not that the machines suddenly fell into greediness by their free will, the tragedies were nothing but the revealing of the overall greed in Wall Street. "There be no beasts." (quoted from Lois Lowry's Gathering Blue)

With the evolution the Web starts embodying the wills of not only the Wall Street traders but also many of us as the regular Web users. For example, the Google robots silently execute your will every time you press the search button. The Web starts being filled with the embodied wills and the machines claimed the ownership. That machines have will is already a fact.

With will, the machines now can do something fantastic. Also with will, the machines someday may do something horrible than we could imagine. The Wall Street tragedies mentioned by the Wired Magazine article are warnings. There is no any single rule that is particularly evil. It is the overall of all the rules/wills (established by the humans but) executed by the machines that causes the problem. If overall all of us are with good, positive intention, the embodiment of our wills overall will only lead our society and our world to becoming better and better. On the contrary, the tragedies will be inevitable. A little bit "harmless" greed by each of us can be accumulated by the machines that claim the embodied wills to become a great evil overall. Which future and which type of the Web we want it to be? The decision is in the hand of everyone of us. Now!

To the end, let's remember it. Machines may have will, but it is not free will. The future is in the hands of the free-thinking people, who are you and me. Let everyone of us uses the Web with good intention and thus the Web will be good to us. Otherwise, the tragedies will not just happen in the Wall Street.

Tuesday, September 16, 2008

Improve Human Intelligence: the opposite to Artificial Intelligence

My newest post at Internet Evolution has been online. Essentially, it is a rethinking of artificial intelligence, a debating topic for decades.

Thinking Space readers may have seen that I have blogged about Imindi, a new startup company, heavily in these few weeks. I am particularly interested in this company because it represents so many great new thoughts that I can hardly see from any other Web companies at this moment. Imindi service is built on a brand new vision of World Wide Web and its potential extension is broad. At this new post, I mainly focuses on one philosophy behind the Imindi service, i.e., to improve human beings because of computers instead of to improve computers because of human beings.

Be interested, you may watch the original article at here. Or do not have much time, watch the imindi graphs below then. ;-)

Saturday, August 23, 2008

Radar Themes, be aware

O'Reilly radar is a principal site exploring the very frontier of technologies. Recently, it publishes an index of themes that O'Reilly believes about the future. In contrast to review all the list, I want to share a few of my thoughts on the Web-related topics.

Collective Intelligence

Without a question, every Web 2.0 success shows the value of collective intelligence. But is collective intelligence panacea for launching new Web services?

(1) Collective intelligence is not necessarily the most crucial factor of collectivism, let it alone being the only one factor. There are many other factors about collectivism, for example, collective behavior and collective identity. The future of the Web will not be just about collective intelligence.

(2) Overemphasizing collectivism may cause prejudice of judgment. We should look for a balance between collectivism and individualism.

Open Beyond Source

The fundamental of open source is actually an alternate of collectivism. Through open source, different people can voluntarily contribute to varied aspects of a common project. Open source is actually more than project collaboration. For example, the linked data achievement (by the Semantic Web community) is another typical application about open source on data collaboration.

Besides open source and open data, can we also open "mind"?

Web Ops

Nat Torkington has quoted that "every 100ms of latency costs Amazon 1% of profit." If this is true, the data strongly supports to develop efficiently Web Ops technologies.

But there is a problem. World Wide Web is an open system in contrast to desktop computers are closed systems. Therefore, Web operating system is probably an inappropriate term since open systems are too unstable to be uniformly operated.

On the other hand, what may happen if we think of the entire Web to be an operating system? That is, instead of inventing certain external Web operation protocols, we let the Web operate itself. By this thought, every Web service is part of the Web Ops functions in contrast to existing extra specific Web Ops services to operate the Web.

Social Networking

Social networking becomes popular at Web 2.0. Unquestionably it represents a huge trend of demand from regular users. People want to be social and they expect to know more people through the Web.

A problem is, however, that why we have to firstly be "friends" in order to share information. In the other words, sharing common interest does not necessarily make two people be friend and vice versa. This misconception of friendship is a critical problem in the current implementation of social networking.

Web 2.0

We have already discussed this one too much. But the time of Web 2.0 is passing.

Money/Web

I am not sure what O'Reilly's view about Money/Web. To me, however, the topic is mainly about how the Web may produce wealth for mankind in general. Adam Lindemann and I have shared the thoughts of mind asset, which could be an interesting issue to explore under this cup.

Physical Web

World Wide Web is no longer just about linked computers. It is now about all kinds of devices linked global wide. Until now, however, devices other than computers are nothing but terminals of the Web. Will some day new devices invented to be central nodes of the Web besides being the terminals?

Neo-Geo

On the Web, we are actually rebuilding the real world into a virtual environment. Google Maps and Google Earth are in the vanguard. But we still have a long journey to go to really digitalize our real world into the virtual world known as the World Wide Web.

Clean Energy Tech

Though the topic seems unrelated to the Web, isn't mind another form of clean energy? If it is, then the Web is related.

New User Interfaces

Will the Web always be primarily linked through regular data connection? Are the regular hyperlinks necessarily be the primary structure of the Web? These are actually the question for new user interfaces. They are more crucial than switch the interface from desktop computers to mobile devices.

Overload

Web 2.0 solves information overload by engaging more human interaction. In consequence, however, the solution causes a new problem of identity overload, i.e., every Web user has multiple identities on the Web at varied Web 2.0 site. Solving this identity overload problem is going to be critical of the future Web.

Referenced resources:

Tuesday, January 01, 2008

Macroscopic regularity over microscopic Brownian motion | the secret beneath the wisdom of crowds

Happy New Year! 2008 will be an exciting new year for many reasons. The philosophy of Web 2.0 has been understood and accepted by more and more people and organizations. "Engaging the wisdom of crowds" has been a slogan of many new-age startups, as well as a few old corporations. In the year 2008, we will watch more deep implementation of this philosophy in the industrial realm. On the other hand, the new concept "Giant Global Graph" starts to be a bridge connecting Web 2.0 and Semantic Web. As the result, Semantic Web, after many years research in labs, will gradually approaches the public audience in 2008. At last, in person I will start a new career in this year. 2008 thus means especially different to me.

In this first post at 2008, I want to address an interesting and essential topic about Web 2.0---why is the wisdom of crowds often superior to the wisdom of individuals even though the individuals might be domain experts and the crowd is generally unprofessional?

The previous argument is the foundation of a best-selling book The Wisdom of Crowds written by James Surowiecki. In his book, James describes many evidences to show the superiority of crowd wisdom and he has also suggested several ways to approach the crowd wisdom while at the same time avoiding some regular traps of misusing this concept. Nevertheless is the book well written, there are a few important absences in the book. One Amazon book reviewer Aaron Swartz criticized that James's book was lack of thoughtful analysis on the intrinsic reasons beneath the described phenomenon of crowd wisdom. A similar critique was made also by another Amazon book critic, David J. Gannon, who wrote that "[James's] choices seem to be crafted to provide maximum support while eliminating any element of contraindication whatsoever." Despite of my sincere support to the basic concept in James's book, I have to confess that I incline to the arguments made by the two critics. These negative viewpoints, however, does not decrease the value the book (they only mean that the book could be even better). But they really pointed out at least one essential missed issue in the book, i.e., the question I rose at the beginning: what are the intrinsic reasons beneath the fact that the wisdom of crowds is often superior to the wisdom of individuals? With this question, I read through the book carefully and finally, I got an insight---the superiority of the wisdom of crowds is another example of a natural fact that there is often macroscopic regularity over any microscopic irregularity.

Brownian motionOne of the most famous irregular natural events is Brownian motion (click the picture on the left). Brownian motion is the random movement of particles suspended in a fluid or the mathematical model used to describe such random movements. In nature, Brownian motion exists everywhere. For example, in a body of water every individual H2O molecule moves towards random directions and with varied velocities at the microscopic level regardless, however, how the entire body of water actually flows at the macroscopic level. An interesting phenomenon is that despite of the pure random movement of each individual H2O molecule, a body of water always has its regular path of flow at the macroscopic level. Hence the irregularity of Brownian motion that may be supposed by many people to lead to random unpredictable flow actually generally cause regular predictable flow at the macroscopic level.

At the mathematical abstraction level, the phenomenon of Brownian motion and the phenomenon of crowd wisdom are indeed the same. In both cases, we have vectors that point to desultory directions and have random magnitudes in their respective directions. In Brownian motion these vectors represent the momentum of the particles and in crowd wisdom these vectors represent the decisions made by individual persons. By adding up these vectors, we may obtain a collective final result towards which the entire body moves. In Brownian motion the result is where the fluid body flows and in crowd wisdom it is what the collective decision is made by the crowd.

This mathematical model well explains several likely controversial claims in James's book. For example, in his book James observed that the collective decisions made by independent participants with highly diverse disciplinary areas is generally at least not worse than the collective decisions gathered by participants who are professional experts in the particular disciplinary area of the question. This observation is indeed surprising when we first see it because we often expect that to the same challenge the decision produced by a group of experts must be generally better than the decision made by a group of laymen. But based on the mathematical model of Brownian motion, we can see that although every individual particle acts purely random to each other, the sum of their momentum vectors always points to a fixed direction, i.e., the direction where the fluid body flows at the macroscopic level. In similar, although every individual person makes a decision solely upon his own profession that may be far away from the destinate disciplinary area, the sum of these decision vectors will point to a certain direction, i.e., the direction where the true answer sits (though nobody in the group really knows this direction). In this situation, it does not matter whether this group of people are domain experts or not. This is the myth and beauty of the wisdom of crowds.

The Brownian motion also explains why a group of laymen may even often outbeat a group of experts on producing a better collective decision. The figure on the left illustrate the idea. Above all, nobody truly knows the real direction of the goal. But experts often make decisions that are closer to the goal (this is why they are experts). By contrast, laymen often make decisions that are far away from the real goal. But if we sum up the decisions made by experts and the decisions made by laymen, the figure shows that it is not necessary that the collective decision of experts is better than the collective decision of laymen. Why? Experts are often biased in the same way, while laymen seldom have the type of biases experts have. Individual expert is certainly better than any individual layman on making professional decisions. But collectively experts often make biased decision---albeit the collective decision is also close to the real goal---since they are trained in similar ways. By contrast, the sum of layman's decisions might be more closer to the real goal since they do not have disciplinary bias in their mind. Though this vector addition diagram is simple, it shows why a group of laymen may beat a group of experts.

As James has emphasized in his book, independence is a crucial property of gathering better collective decisions, especially when these decisions are collected from a group of laymen. The new diagram at the left illustrates the reason. Unlike the previous one in which all experts and laymen make their decisions independently, in this new situation Expert 2 makes his decision after Expert 1 and so is Layman 2 after Layman 1. Humans are social creatures; so we often adjust our decisions to compromise the other persons in a group, even unconsciously. As the result, both Expert 2 and Layman 2 shift their decisions a little bit closer to the decision made by the respective former players. Immediately we see the consequence. The collective decision made by experts is still close to the goal since the decision made by Expert 1 is close to the goal. By contrast, the collective decision made by laymen starts to be away from the goal since the decision made by Layman 1 is away to the goal. This simple diagram shows why it is generally unconstructive (and often even destructive) to have a group of laymen communicate when they vote because most of the time they will unconsciously follow a direction that is far away from the real goal. By contrast, we may encourage the communication among experts since they are more likely to figure out a better solution after discussion. To the least, the worst expert decision may still be close to the goal.

Does the superiority of crowd wisdom suggest that we should not (or at least should not actively) hire experts on making decisions? The answer is no. There are two fundamental reasons why the existence of experts is actually crucial to the success of engaging the wisdom of crowds: (1) how to build a real diverse group of crowds and (2) how to aggregate the decisions made by the crowds.

As James has emphasized in his book, the property of diversity is critical to obtain high quality crowd-made decisions. In fact, this requirement of diversity is equivalent to the perspective of free of bias. When the voters in a group have diverse enough disciplinary backgrounds, disciplinary biases are eliminated to the least. Thus the aggregated collective decision would be closer to the real truth. But how to build a truly diverse set of participants is a problem. A random group of people invited from street might not necessary be a real diverse set to a particular question. To build a real diverse set requires highly professional experiences on the respective disciplinary field.

The two examples James told in the Introduction of his book are typical examples of why the construction of diverse groups is a highly professional job. In the first example, many people on the marketplace were beating on the weight of an ox. In his story, James addressed the participants as unprofessional normal people with respect to the issue of ox weight. Nevertheless was James right, these people were not so "unprofessional" as James had emphasized. We can safely assume that these people who had participated the beat regularly bought stuffs from the market. So they had basic knowledge on how to evaluate weight of varied things, even though they indeed were not professions on weighting oxes. This observation is important because it means that this group was really "diverse" with respect to the challenge. Think of repeating the same challenge among a group of first-grade elementary school students and we may see the difference between really "diverse" and fake "diverse" groups. By randomly calling up a group of first-grade elementary school students we may also have a diverse set, which is, however, not really "diverse" with respect to the demand of collecting crowd wisdom. The first-grade elementary school students are short of knowledge of weighting basic stuffs and thus the collective answer made by them is certainly "biased" by their short of knowledge.

Similar situation is for the second example, in which a group of professionals in varied fields was assembled by a naval official John Craven to guess the position of Scorpion, a US submarine disappeared in the North Atlantic. The story in the book was impressive; but would we be able to repeat this story by assembling a random group of professors at MIT? Certainly these professors must be brilliant and unquestionable experts in their disciplinary areas, but I bet they would certainly not able to guess the correct position of Scorpion, even collectively. Why? These professors generally have not been attended to the particular scenario and thus their decisions would be just little bit better than normal you and me in this case. On the contrary, the people John Craven had organized (as in the story) were the ones who were familiar to the submarine operation even though nobody had the complete knowledge of the particular case. So this group called by John Craven was a really diverse group and a random group of MIT professors is not.

Both the stories point out that assembling a "diverse" set is a highly professional request that demands the knowledge of domain experts.

Comparing to the demand of diversity, we may require more expertise on aggregating individual decisions made by a crowd to be a valuable collective decision. This type of aggregations is normally much more sophisticated than simply calculating the number of votes in different categories. As we show in the previous diagrams, the process of aggregating crowd wisdom is basically a procedure of vector addition (or more precisely a tensor addition since many times the number of dimensions would be more than three). How to divide a problem into varied dimensions and assign a measurement standard to it is a highly professional work that requires superior expertise on the disciplinary application area.

In summary, we now have a clear picture of how to take the benefit of crowd wisdom in real-world applications. First, we need a few experts. In contrast to rely on these experts to make decisions directly, however, we ask them to assemble a crowd that is truly diverse according to the problem. We let the crowd make their decisions independently. Then we ask the experts to aggregate the crowd decisions objectively based on the expertise of the experts. This is thus the procedure of engaging the wisdom of crowds.

Monday, December 03, 2007

Collectivism on the Web

Collectivism emphasizes on human interdependence and the importance of collective. As probably the greatest collective project of mankind in history, World Wide Web engages enormous practices of collectivism. In this article, we take a brief look at several typical examples of these engagements.

Collective Intelligence

Collective intelligence is the most well-known engagement of collectivism on World Wide Web. In particular, Web 2.0 advocates have declared "harnessing collective intelligence" to be the touchstone of the Web 2.0 revolution. By definition, collective intelligence is a form of intelligence that emerges from the collaboration and competition of many individuals. If someone feels a little bit puzzled of this definition, here is an alternative explanation that is imprecise but much easier to be understood. Informally, collective intelligence on the Web is the collections of user generated "intelligence".

A keen reader may immediately find an interesting comparison: are there any differences between user generated "intelligence" and user generated "content" (or user generated "data")? On Web 2.0, we have almost mentioned users generation content (UGC) as many times as collective intelligence. In many people's mind, UGC almost equals to the collective intelligence. But the actual meanings between "intelligence" and "content" or "data" are very much different. The intent of "intelligence" is much richer than "content/data". Tim O'Reilly also had briefly mentioned this distinction in one of his earlier post about harnessing collective intelligence.

Content/data is a type of intelligence but at the low end. Jean Piaget, a Swiss philosopher and pioneer of the constructivist epistemology, had a compact description about intelligence: "Intelligence is what you use when you don't know what to do." Content/data provides shallow and unrefined information for people to use. Content/data is often too crude to be efficiently used. Keeping the user generation intelligence at the level of content/data is not enough. This is a problem.

I foresee that the degree of complexity (as well as the degree of efficient usage) of the collective intelligence on the Web is going to evolve with the Web. For example, by tagging content with formal labels that are defined by ontologies, the user generated content/data would evolve to be the user generated knowledge. This is exactly what the vision of Semantic Web wants to bring to us. Moreover, by augmenting formally labeled content with external logic routines, the user generated knowledge would evolve to be the user generated wisdom. By encoding the mechanism of proactiveness into machine computation, the user generated wisdom might evolve to be the user generated creativity. By engaging user generated content/data, knowledge, wisdom, creativity together, we might eventually get the user generated personality, through which the human evolution reaches a new stage of being artificially immortal. Is this path a long way? Yes, there is a long way to go. Is this path an impossible dream? No, it is not. The practice of collective intelligence is converting our society into a virtual world simultaneously from the level of individuals and the level of collective groups.

Collective Behavior

Collective intelligence is not the only practice of collectivism on the Web. Another key practice of collectivism on the Web is the implementation of collective behavior.

Collective behavior is very much difference from collective intelligence. All types of collective intelligences are static and thus they can be easily presented in an explicit way. In comparison, collective behaviors are dynamic and it is difficult to present them in an explicit way. As the result, collective behaviors are much harder to be used than collective intelligences on the Web though in fact at the same time the amount of collective behaviors is much greater than the amount of collective intelligences. The reason of this amount difference is indeed trivial. Every piece of collective intelligence on the Web must be related to at least one human behavior (i.e. the one action that post this piece of information online). The majority of the time, any piece of collective intelligence must be associated with many human behaviors such as reading and writing. With such a large pool of collective behaviors, it is surprising to see that so few actions have been made so far to manage and utilize this large pool.

Fortunately, Web researchers have started to pay their attention to the collective behaviors. The recent proposal of the implicit web is a typical example. The implicit web is a network that defragments every piece of implicitness on the explicit web. The majority of the implicitness on the Web actually belongs to the collective behaviors.

Collective Responsibility

The collective intelligence is a popular concept. The discussion of collective behavior is also not rare. But the rest of practices of collectivism on the Web I am going to discuss are indeed uncommonly. Many readers may not even hear of them before. But all these practices are unexceptionally important and valuable for the evolution of World Wide Web. The first one I introduce is the collective responsibility.

Collective responsibility is a concept, or doctrine, according to which individuals are to be held responsible for other people's actions by tolerating, ignoring, or harboring them, without actively collaborating in these actions. This concept is particularly important to the study of Internet security.

On the age of Web 2.0 and afterwards, security is no longer a solo issue with the deeper and wider implementation of collectivism. As a result, being innocent may no longer be simply taken as an individual issue. We must start to consider collective responsibility, i.e., some people may have to be punished not due to their own guilty but because they have not actively prohibited the guilty happened regularly in their participated societies. This issue is going to be very much debatable and exciting.

Collective Identity

A collective identity refers to individuals' sense of belonging to a group.

Identity is a tough issue on the Web. Normally, a web user may have varied identities on different sites. These varied identities, however, cause serious problems when people try to organize their information of interest across the boundaries of web sites. To address this problem, web researchers have issued the project OpenID that allows users to use a single ID over the entire Web.

But OpenID, even if it would be a standard over the Web, is not the end of the Web identity issue. Similar to that individual persons have their particular roles in real life, individual identities on the Web must gain their particular social roles in virtual life. The identification of these roles is particularly important when we would start to manipulate human generated information on the Web, i.e. collective intelligence, collective behaviors, etc. Only until humans or machines may identify the social roles of the information producers or owners, these humans or machines may be able to properly manipulate the information. The research of collective identity will focus on the identification of social roles of individual identities.

The collective identities are identities of identities. The study of this issue is another exciting and unexplored field that may cause much attention in the future.

Collective Consciousness

Collective consciousness refers to the shared beliefs and moral attitudes which operate as a unifying force within society. In the other words, the collective consciousness is about machine morality because human consciousness on the Web is handled by machines. The machine morality is not a sci-fi term; this issue is indeed real. Machine morality is the reflection of human morality onto the virtual world.

The implementation of collective consciousness is very much related to all the previously mentioned collective factors. Human consciousnesses are materialized on the Web as static intelligence and dynamic behaviors. Moreover, the collective identities assign social roles for the materialized consciousnesses. The integrity of these materialized consciousnesses is closely related to the level of collective responsibility that is maintained at the meantime. The combination of all these issues compose the intent of the machine morality.

Collective Effervescence

Collective effervescence is a perceived energy formed by a gathering of people as might be experienced at a sporting event, a carnival, a rave, or a riot. This energy can cause people to act differently than in their everyday life.

Collective effervescence is the emotion web site owners want to bring. Collective effervescence represents one word---hype! Collective effervescence is the ultimate goal of implementing collectivism on the Web. At the same time, how much an implementation of collectivism successfully brings collective effervescence into a web site is the fundamental standard that we can measure the quality of the implementations of collectivism. This concept encloses the entire set of collective factors and upgrades the evaluation of collectivism into the computational realm.

Summary

We have discussed several examples of how we may engage practices of collectivism onto the Web. Certainly there could be many other possible practices that are beyond this article. But one thing is certain. Collectivism is a crucial phenomenon on the evolving Web. The study of collectivism on the Web is going to be a critical issue of the Web Science.

Wednesday, November 28, 2007

Blink: an embarrassment of collective intelligence

Blink is another best-selling book authored by Malcolm Gladwell after his influential The Tipping Point. The book Blink is about the unconsciousness of human being. In the book, Malcolm argues that a decision made by well-trained unconsciousness many times is better than an alternate decision made by through thoughts. Reason: well-trained unconsciousness (or the so-called "thin slicing") only catches the very core of the problem, while through thoughts often wander into unessential branches that lead to the burying of the core. This is thus "the power of thinking without thinking," as the subtitle of the book.

This observation of the importance about the "thin slicing" shows an embarrassing side of the collective intelligence: if there is a conflict between a decision made from a collective base and an alternate decision made by the instinct of few top experts, which one should we trust? The Web-2.0 experiences ask us to vote for the first decision, but Malcolm's book tells us that most of the time it is the second one that is more trustworthy. Which one would you pick in real then?

This is a vague question that may not have an absolute answer in general. But at least the question shows that collective intelligence is not a panacea. An opinion from a domain expert and another opinion from a layman certainly should be weighted differently when we apply both to make a decision. Some time, as what Blink tells, the instinct of very few experts is much more correct than a collective decision.

So is the YouBeTheVC competition a really serious event? Maybe it is just another American Idol show. Think of it, would Larry Page and Sergey Brin (or Mark Zuckerberg) attend this kind of idol show when they had the blueprint of Google (or Facebook) in mind? I doubt it. Distinctive idea is more often out of a blink in contrast to out of a collective vote.

Tuesday, November 06, 2007

The Implicit Web

(This article is cross-posted at ZDNet's Web 2.0 Explorer.)
(watch the article also in Chinese, translated by the author)

Implicit web is a new concept coined in 2007. Due to the first Defrag conference right now, discussion of this new term is timely.

Generally this concept implicit web intends to alert us a fact that besides all the explicit data, services, and links, the Web engages with much more implicit information such as which data users have browsed, which services users have invoked, and which links users have clicked. This type of information is often too boring and tedious to be human readable. So, inevitably, this type of information is only implicitly stored (if stored) on the Web. The implicit web intends to describe a network of this implicit information.

Implicitness Everywhere

Implicit information is everywhere. Implicit information on the Web is about things to which human web users have paid attention. For example, it is about which web pages are frequently read, how often they are read, and who read them. It is also about which services are frequently invoked, how often they are invoked, and who invoked them. Consider the number of web users and how many activities everybody has done daily on the Web, the amount of implicit information must be astonishing. The implicit information co-exists with every web page, every web service, and every web link. In short, great amount of implicitness co-exists with every little piece of explicitness on the Web.

Implicit does not mean insignificant or unimportant. By contrast, implicit web information is often valuable and even crucial in various situations. For example, implicit information of click rates can help editors decide which news are the most popular ones and thus they should put these news on the front page. In similar, the same type of implicit click rates can help salespeople decide which merchandises are among the greatest demanding and so they can arrange the next supply line.

Many companies have already started to collect implicit information and they take benefits from it. Alex Iskold had written a compact introduction on how some companies have utilized implicit information in their products. One well-known example is Amazon.com, which always lists related buyer recommendations with each of its online merchandise. "Customers Who Bought This Item Also Bought," many readers must be familiar to this label. And more importantly, many web users do care of the content underneath this label. This is a typical example of how implicit web information helps.

Amazon is not the only company that benefits from implicit information. Amazon is not one of the few companies that benefit from implicit information. In fact, nowadays almost every website that sells something, from baby toys to cars, has some back-end mechanism on analyzing the traffic (a typical implicit information) and adjust their sales plan based on the analysis. Implicitness is indeed everywhere.

Connect Implicitness

Implicitness is everywhere, but is fragmented everywhere. Implicit information on the Web is not connected. This is a problem.

Until now, implicit web information is generally separately stored, typically by individual companies. For example, both Gap.com and jcrew.com have their own stored visitor history but not shared to the other, although we may imagine that this information must be well connectible since both companies sell apparel and accessories. Someone may argue that Gap and J. Crew are competitors. So let us switch the pair to be Banana Republic and Victoria's Secret. The products of these two companies are well complement (in contrast to compete) to each other. But still the implicit information is isolated to itself, despite that both sides can benefit by connecting this independent implicit information. Readers can find many more this type of examples.

If sharing implicit information among big companies is still questionable (because these big boys hardly believe that they could get help from their little sisters), this type of sharing is much more critical to small websites. There are numerous individual sites that cannot utilize themselves well enough from their own implicit information because they are too small in size. At the same time, however, there are no effective way for them to share and find helpful implicit information, though everybody knows that there is plenty of this information on the Web.

All these discussions lead to one demand: we need the implicit web, which is not there yet. The goal of the implicit web is to defragment all the fragments of implicitness (where the name Defrag is gotten for the conference). But how can we indeed connect all the different types of implicitness on the Web to be a coherent implicit web? This is a grand challenge to the newly formed community of implicit web research. We do not have a clear answer yet.

No matter whatever, however, the solution to the question must be beyond web links. The implicit web engages with complex types of semantics. The amount of information on the implicit web is gigantic. The implicit Web is also very much dynamic. The traditional model of web link is too simple, too shallow, and too static to deal with all these challenges at the same time. We need big, creative thoughts to store and link all the implicitness.

The greatest potential problem to the implicit web is privacy. To companies, some implicit information may be too confidential to be shared. To individual persons, some implicit information may be too private to be public. We need innovative methods of privacy control on the implicit web.

Implicit Web in nutshell

In summary, I briefly list my beliefs about the implicit web.

1. The implicit web is a network that defragments every piece of implicitness on the explicit web, which is the generally known World Wide Web itself.

2. If the explicit web reveals the static side of human knowledge through posted data, services, and links, the implicit web reveals the dynamic side of human knowledge by recording how users access these data, services, and links.

3. The explicit web engages collective human intelligence. The implicit web engages collective human behaviors.

4. The implicit web is not part of the Semantic Web, but they are closely related. If the Semantic Web constructs a conceptual model of World Wide Web, the implicit web constructs a behavior model of World Wide Web.