In an article in Computer World, Alfred Spector, Google’s VP of research, states that Google has developed technologies that enable the Google crawler to get content hidden behind passwords and user names.
This was in response to the following question:
Do you have plans to go after that huge body of information on the Internet that is not currently searched?
His answer:
“There is stuff on the Web, the so-called Deep Web, that is only “materialized” when a particular query is given by filling fields in a form. Since crawlers only follow HTML links, they cannot get to that “hidden” content. We have developed technologies to enable the Google crawler to get content behind forms and therefore expose it to our users. In general, this kind of Deep Web tends to be tabular in nature. It covers a very broad set of topics. It’s a challenge, but we’ve made progress.”
Several companies are currently developing technologies to access what is known as the Deep Web. We profiled Deep Peep recently. Google’s purchase of Transformics is helping them compete in what will be an important aspect of online search.
Read More!
Showing posts with label Deep Peep. Show all posts
Showing posts with label Deep Peep. Show all posts
Sunday, March 29, 2009
Sunday, March 8, 2009
A Deeper Peep into Search
Search engine Deep Peep (http://www.deeppeep.org) is one of the few search engines purporting to provide results from the deep web, which is a term associated with providing access to information contained behind passwords and user names. They describe themselves on the site:
"DeepPeep is a search engine specialized in Web forms. The current beta version tracks 13,000 forms across 7 domains. DeepPeep helps you discover the entry points to content in Deep Web (aka Hidden Web) sites, including online databases and Web services. This search engine is designed to cater to the needs of casual Web users in search of online databases (e.g., to search for forms related to used cars), as well as expert users whose goal is to build applications that access hidden-Web information (e.g., to obtain forms in job domain that contain salary, or discover common attribute names in a domain). The development of DeepPeep has been funded by National Science Foundation award #0713637 III-COR: Discovering and Organizing Hidden-Web Sources."
I haven’t seen much out of Google, Yahoo, and Microsoft on indexing content in the hidden web. Deep web search is difficult to implement because of the human element required in accessing and manipulating the content on the sites.
Read More!
"DeepPeep is a search engine specialized in Web forms. The current beta version tracks 13,000 forms across 7 domains. DeepPeep helps you discover the entry points to content in Deep Web (aka Hidden Web) sites, including online databases and Web services. This search engine is designed to cater to the needs of casual Web users in search of online databases (e.g., to search for forms related to used cars), as well as expert users whose goal is to build applications that access hidden-Web information (e.g., to obtain forms in job domain that contain salary, or discover common attribute names in a domain). The development of DeepPeep has been funded by National Science Foundation award #0713637 III-COR: Discovering and Organizing Hidden-Web Sources."
I haven’t seen much out of Google, Yahoo, and Microsoft on indexing content in the hidden web. Deep web search is difficult to implement because of the human element required in accessing and manipulating the content on the sites.
Read More!
Subscribe to:
Posts (Atom)
