The End of the Search Engine as We Know It?
|
Smart Searchers |
 |
Three new and innovative solutions for the efficient information retrieval on the Web.
Seven million pages are being added to the Web each day. It is painfully clear
that we lack solutions that can keep up with a system growing at such breakneck
pace. But as the problem gets worse, the solutions get better. This time we'll
take a look at three really different and innovative products that are still
in the early stages of development. Of course, the gap between hype and reality
is not new to the AI world, and some of them will not live up to the expectations.
On the other hand, some of them may become future industry standards, so it
is never too early to start exploring their possibilities.
OpenCOLA
OpenCOLA, a Toronto-based company with
a unique marketing strategy, just released an autonomous and collaborative open
agent framework under the same name - where COLA stands for Collaborative
Object Lookup Architecture. It is a true collaborative computing environment,
in which all instances of the program are peers. To avoid referring to "clients"
and "servers" in such a peer to peer environment, the OpenCola platforms are
called "clervers". Another peer-to-peer startup trying to promote yet another
way to share electronic documents and MP3 files? When combined with another
key concept, expert communities, it could actually offer an elegant solution
to the problem of overcrowded indexes of today's search engines. You will share
your hard drives with all other users who have similar interests. Your personal
agent (stored as XML file) will be trained to recognize and seek out only the
specific information. The whole idea relies on the network of users with equal
rights, essentially spidering and indexing small chunks of the Web in their
spare time.
Several applications should use this concept for achieving different tasks.
OpenCola Swarmcast will be a large file distribution system, allowing
content providers to serve large files significantly faster, breaking large
files (> 1 MB) into multiple, easily transferred chunks. These chunks will be
served simultaneously from multiple nodes in the OpenCola Swarmcast network.
OpenCola Streamfind is an extensible tool for announcing, distributing,
searching, filtering and viewing streamed media content. However, the most interesting
products from the AI standpoint will be called OpenCola Smart Folders.
Using a combination of machine intelligence, human decisions gathered from across
the OpenCola network, and direct user feedback to improve results it will continuously
and transparently search and retrieve the interesting information: The way
OpenCola Smart Folders works is simple: Name a folder. Drag a file into that
folder. The folder automatically fills up with other files that are similar.
If you take something out of the folder, it doesn't find stuff like that anymore.
You can create as many folders as you want by setting them up yourself or by
adopting them from other users.
Early adopters and experimenters can now download the OpenCOLA SDK. Documentation
is still incomplete, so the knowledge of C/C++, XML and HTTP is absolutely essential.
WDBC
LOTONtech Limited recently released
Web DataBase Connectivity, a Java2 adapter that lets you run SQL queries against
live Web pages as though you were querying relational database tables using
JDBC. It uses an enhanced version of SQL - called HTMSQL - that allows hierarchical
HTML data to be accessed according to the flat two-dimensional relational database
paradigm. I already dedicated an earlier feature article to the similar tools
for "SQL-style information
retrieval". WDBC shares a similar approach, and will be very attractive
for Java programmers building search engines, data mining tools, or knowledge-based
applications. Beta release applet can be tried from their Web pages. There are
still some bugs waiting to be fixed, but it offers an intuitive and easy way
to extract information from the semi-structured Web data.
DolphinSearch
DolphinSearch, Inc. was founded
in 1999 to exploit Dr. Herbert L. Roitblat's pioneering research into how dolphin
brains process echolocation signals. The research has led to breakthroughs in
using pattern recognition techniques to enable machines to recognize the meaning
of human language. They offer a family of technologies for automatically reading
documents and storing them in a searchable database that is very different from
the search engines as they are used today. The main concept behind this technology
relates relational pattern recognition mechanisms that model how the dolphin
recognizes objects to the text recognition techniques. "Just as a sound wave
changes its meaning depending on the context of other waves around it, so too
do words change their meaning - words have senses, nuances, and multiple meanings
- depending on the context of other words around them."
According to Dr. Roitblat, "knowledge management should be an inherent network
function". Following that principle, their main product, KnowledgeBox, is a
plug and play, hardware-based network appliance capable of finding any unstructured
information stored anywhere on a computer network. KnowledgeBox quickly scans
through the text on a company's network, regardless of format, understands its
meaning and remembers where it is stored. Pricing is $10/seat/month, with a
minimum commitment of $10,000 (US) in the first year. In subsequent years the
minimum annual commitment drops to $2000.