Face recognition equipment and method
Patent Information
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Current Assignee / Owner
- GOOGLE LLC
- Publication Date
- 2008-04-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing document navigation techniques lack the ability to automatically generate and integrate relevant links within the document content based on user-specific information, limiting the reader's browsing experience.
A method that includes generating descriptive information based on the document content and user personal information, identifying relevant additional documents, and embedding these documents inline within the original document at relevant locations.
Enhances the reader's experience by providing contextually relevant additional documents inline with the original document, improving navigation and information retrieval efficiency.
Abstract
Description
"Methods for Enhancing Document Navigation and First Document and Respective System and Device" Descriptive Report Background of the Invention A - Field of the Invention The systems and methods described here generally relate to information retrieval and, more specifically, to techniques for navigating information. B - Description of the Related Technique The World Wide Web (“Web”) contains a vast amount of information. A very common use of the Web is to read documents, such as news articles or other publications. When reading a particular document, such as a news article, it is common practice to provide links to other documents that are somehow related to that document. For example, when a user selects a news document from a news search engine or an online news service, the website may provide links to other news articles or announcements that are related to the news document. Typically, these related documents are determined based on the content of the document being read and are displayed as additional links outside the document's content. By providing convenient links to related material, these additional documents can enhance the reader's browsing experience. It would be desirable to provide improved techniques for enhancing document navigation by providing automatically generated links to information relevant to the reader. Summary of the Invention According to one aspect, a method of improving document navigation includes receiving personal information about the user, generating descriptive information based on the content of an initial document and personal information, and identifying documents. Additional information can be generated based on the descriptive information. Furthermore, a second document can be generated that includes at least part of the content of the first document, modified in a way that includes references to the additional documents. In another aspect, one method involves locating a second document that is relevant to at least one first document and embedding the second document within the first document at a location in the first document to which the second document is relevant. In another aspect, one method includes receiving a request for a first document from the user, identifying a named entity in the first document, locating a second document that is relevant to the named entity, and presenting the user with a modified version of the first document in which a link to the second document is displayed inline in the first document at a location in the first document that approximates the named entity to which the second document is relevant. Brief Description of the Drawings The accompanying drawings, which are incorporated into and form part of this Descriptive Report, illustrate an embodiment of the invention and, together with the description, explain the invention. In the drawings, Figures 1A and 1B are diagrams that illustrate exemplary graphical interfaces that can be presented to the user; Figure 2 is an illustrative diagram of a network in which concepts consistent with the principles of the invention can be implemented; Figure 3 is an example diagram of a client or server shown in the network of Figure 2; Figure 4 is a block diagram illustrating conceptual elements of the document locator shown in Figure 2; Figure 5 is a diagram that illustrates an exemplary embodiment of the search component shown in Figure 4; Figure 6 is a flowchart illustrating example operations performed by the document locator shown in Figure 2; and Figure 7 is a diagram illustrating an example of a document locator in the context of a Content Service Network site. Detailed Description The detailed description of the invention below refers to the accompanying drawings. The detailed description does not limit the invention. Assessment As described here, additional documents relevant to an original document, such as a document being read by the user, are automatically located. These additional documents can be located based on their content and / or the user's personal information. They can then be displayed inline with the original document. This allows the user to be efficiently presented with additional information relevant to the document being read. Figures 1A and 1B are diagrams illustrating exemplary graphical interfaces that can be presented to the user. The graphical interfaces can be presented via a web browser 100 being used to navigate the web. The example document 105 shown in Figures 1A and 1B relates to the effort of a hiker (Bill Cross) to climb Mount Everest. A number of additional documents may be relevant to document 105. In Figure IA, for example, links to three additional articles 110, 112, and 114 are embedded in document 105. Link 110 may reference a document about Mount Everest, and link 112 may reference a document about the Novolog Peaks and Poles. Challenge and link 114 may reference a document about diabetes. Each of the links 110, 112, and 114 references content that is, in some way, related to the original document 105. In this example, links 110, 112, and 114 are displayed with brief summary text (e.g., “related content: Mount Everest”) that informs the reader of the underlying link's content. Additionally, the summary text is underlined, indicating that the summary text is associated with a link. Suppose the reader of document 105 in Figure IA is located in San Jose, California. An advertisement 115 for a hiking equipment retailer in San Jose might also be displayed. Furthermore, the documents referenced by links 110, 112, and 114 may be documents that are particularly appropriate for a reader in the San Jose area. Although not shown in Figure IA, other links may also be displayed, such as links even more directly personalized to the reader's personal information. For example, if the reader has previously entered search queries into a search engine, such as queries related to photography, the other links might be links to documents describing "Photographs of Everest." Document 105 in Figure 1B is identical to that in Figure 1A. Multiple links 120, 122, and 124 are included in document 105 of Figure 1B. In this example, links 120, 122, and 124, instead of being shown as linked summary text, are implemented by simply modifying the format or display associated with certain words or phrases in document 105. For example, link 120 is shown to the reader by underlining “Mount Everest,” thus illustrating to the reader that the link refers to a document relating to Mount Everest. Another link 126 is inserted inline in document 105 that includes summary text similar to links 110, 112, and 114. Let us assume, for this example, that the reader is from Seattle, rather than San Jose. Link 126, which can be generated based on this fact, references a document about Hiking on Mount Rainier - a mountain near Seattle. Network Assessment Example iva Figure 2 is an illustrative diagram of a network 200 in which concepts consistent with the principles of the invention can be implemented. The network 200 may include multiple clients 210 connected to a server 220 via a network 240. The network 240 may include a local area network (LAN), a wide area network (WAN), a telephone network such as the Public Switched Telephone Network (PSTN), an intranet, the Internet, or a combination of networks. Two clients 210 and one server 220 have been illustrated as connected to the network 240 for simplicity. In practice, there may be more clients and / or servers. Also, in some examples, a client may perform one or more functions of a server, and a server may perform one or more functions of a client. A client 210 may include a device such as a cordless phone, a personal computer, a personal digital assistant (PDA), a laptop, or other type of computing or communication device, a line or process running on these devices, and / or an object executable by one of these devices. The server 220 may include a server device that processes, retrieves, and / or maintains documents and images in a manner consistent with the principles of the invention. Clients 210 and the server 220 may connect to the Network 240 via wired, wireless, or optical connections. Server 220 may include the additional document locator component 225 (also simply referred to here as “document locator 225”). Document locator 225 may locate and append references to other documents related to an input document, such as the references added to document 105 (Figures 1A and 1B). A document, as the term is used here, should be broadly interpreted to include any machine-readable and machine-storable product of work. A document It can be an email, a web log (blog), a file, a combination of files, one or more files with embedded links to other files, a newsgroup post, etc. In the context of the Internet, a common document is a web page, such as an HTML web page. Web pages often include content and may include embedded information (such as metadata, hyperlinks, etc.) and / or embedded instructions (such as Javascript, etc.). The documents discussed here generally include embedded images. A “Zin / c”, as the term is used here, should be interpreted as broadly including any reference to / from a document to another document or another part of the same document. Client Example / Server Architecture Figure 3 is an example diagram of a client 210 or server 220. The client / server 210 / 220 may include a bus 310, a processor 320, a main memory 330, a read-only memory (ROM) 340, a storage device 350, an input device 360, an output device 370, and a communication interface 380. The bus 310 may include drivers that allow communication between the client / server 210 / 220 components. The 320 processor may include conventional processors, microprocessors, or processing logic that interprets and executes instructions. The 330 main memory may include random access memory (RAM) or another type of dynamic storage device that stores information and instructions for execution by the 320 processor. The 340 ROM may include a conventional ROM device or another type of static storage device that stores static information and instructions for use by the 320 processor. The 350 storage device may include a magnetic and / or optical recording medium and its corresponding drive. The 360 input device may include one or more mecha Conventional mechanisms that allow the user to input information to the client / server 210 / 220, such as a keyboard, mouse, pen, voice recognition, and / or biometric mechanisms, etc. The output device 370 may include one or more conventional mechanisms that output information to the user, including a monitor, printer, speaker, etc. The communication interface 380 may include any transceiver-like mechanism that enables the client / server 210 / 220 to communicate with other devices and / or systems. For example, the communication interface 380 may include mechanisms for communication with another device or system over a network, such as a network 240. Server 220, consistent with the principles of the invention, may implement an additional document locator 225. The additional document locator 225 may be stored on a computer-readable medium, such as memory 330. A computer-readable medium may be defined as one or more physical or logical memory devices and / or carrier waves. The software instructions that define the additional document locator 225 can be read into memory 330 from another computer-readable medium, such as the data storage device 350, or from another device via the communication interface 380. The software instructions contained in memory 330 can cause the processor 320 to execute processes that will be described later. Alternatively, wiring circuits or other logic can be used in place of or in combination with software instructions to implement processes consistent with the invention. Thus, embodiments consistent with the principles of the invention are not limited to any specific combination of hardware and software circuits. Document locator 225 Figure 4 is a block diagram illustrating conceptual elements of document locator 225. Document locator 225 may include a descriptive information generator 405, a The descriptive information generator 405 can generate descriptive information that describes the current document and based on the user's personal information. In one embodiment, the descriptive information may include a search query. The descriptive information generator 405 can generate descriptive information based on the user's personal information and / or the current input document (or information relating to the current document). The descriptive information output from the descriptive information generator 405 can be fed into the search component 410, which can use the descriptive information to generate additional documents. Links or other references to the additional documents can be inserted with the original document via the formatting component 415. The descriptive information generator 405, the search component 410, and the formatting component 415 will each be described in more detail below. Descriptive Information Generator 405 As mentioned, the 405 descriptive information generator can generate descriptive information, such as a search query. Descriptive information can generally be based on information relating to the document the user is currently viewing (or has requested to be viewed) as well as the user's personal information. Information relating to the current document may include information based on the text of the current document. The text can be processed to obtain, for example: (1) all terms that appear more than a predetermined number of times, (2) named entities that can be automatically extracted; (3) dates in the document; (4) the author and publication names; and / or keyword or category extraction. With regard to (1) above, it can be considered that the terms that appear more than a certain predetermined number of times These are important or particularly descriptive terms in the document and can be considered descriptive information for the document. The number of terms selected for inclusion in the descriptive information may, for example, be limited to a predetermined number of the most frequently occurring terms. In a possible variation of this concept, the number of times a term occurs may be considered in conjunction with the overall frequency with which the term appears in the document's language. Thus, terms that tend to occur relatively rarely in a language may be selected before a common term that occurs more often in the document. A list of predetermined named entities or other names can be stored using the descriptive information generator 405. For example, place names, celebrity names, well-known trade or consumer product names, and company names can be pre-generated manually (i.e., entered by a human operator) or by automated techniques. As mentioned above, the document text can be compared with these named entities, and matches are included in the descriptive information for the document. With reference to the examples in Figures IA and IB, the list of predetermined named entities may have included terms such as “Mt. Everest*” and “Novolog Peaks and Poles Challenge,” causing these terms to be included in the descriptive information for the document 105. The dates in the document (item (3) above), the document author, and the publication names (item (4)) can be included in the descriptive information. This information can often be determined automatically by pattern matching techniques applied to the document. The document date can be used to locate other contemporaneously published documents. Similarly, the publishing entity (e.g., website) and the document author can be used to locate documents from the same (or similar) publications or documents. written by the same author. The document date, author, and publication date can be particularly useful in the context of new episodes. In relation to (5), a document can be analyzed for its keywords, such as keywords extracted based on the term frequency or through named entity extraction. In addition to generating descriptive information based on the document, the 405 descriptive information generator can generate descriptive information based on user-specific information (“personal information”). For example, personal information may include the user’s geographic location (e.g., previously submitted search queries or selected links), personal information provided by the user when registering an account, personal information based on the user’s browsing history, personal information extracted from user-generated documents, or other sources of personal information. The user’s geographic location may be calculated based on the user’s IP address. Personal information may also include temporal information, such as the current date or season. Temporal information can be useful for correlating events with personal preferences or document content.For example, if a document being browsed is about Edinburgh and the current month is July or August, then documents related to the Edinburgh Arts Festival may be displayed. In one approach, personal information can be based on user profiles built from previous search queries submitted to a search engine. Category matching techniques can be used to infer user interests from search queries. For example, even if the user never actually entered the search term "photography," but instead queried the terms "Nikon," "aperture," and "f-stop," these terms can be used to deduce that the user is interested in photography. A technique for generating category mappings from The analysis of search queries is based on aggregating a large number of historical user search queries that are labeled based on the user's search sessions. The reasoning is that people searching for a term such as "Canon" are likely to also enter, in the same search session, other search queries such as "photography" or "f-stop," which are related to the same category. By analyzing many of these search query sessions, inferences about categories can be made (for example, if a person searches for "Nikon," they are likely interested in photography). The 405 descriptive information generator can format descriptive information as a search query. In one embodiment, the search query can be obtained by concatenating descriptive information (e.g., the user's personal information and descriptive information related to the document) to obtain the search query. As an example, consider document 105 in Figure IA. Based on an analysis of the document and the user's personal information, the 405 descriptive information generator can generate the descriptive information “Mt. Everest”, “Novolog Peaks and Poles Challenge”, “diabetes”, “San Jose”, and “photography.” These terms can be combined into a single search query: “Mount Everest Novolog Peaks Poles Challenge diabetes San Jose photography”.In other modes, multiple search queries can be generated, each including a subset of terms from the document and the user's personal information, such as the search queries: "Mount Everest San Jose", "hiking San Jose", "photography Mount Everest", etc. A person of ordinary technical ability will recognize that other techniques can be used to formulate search queries from the descriptive information generated. For example, additional information can be used in determining whether to include a term in the query, such as general frequencies of occurrence. The term in the language. Additionally, additional weights may be given to certain names, entities, or other predefined terms in determining whether they should be included in the query. Some terms, such as geographical names, may be weighted differently from other terms, such as product names. Product names may be automatically limited by adding their associated company name after the product name. Furthermore, descriptive information may be used with grouping or category matching techniques, such as those described above, to generate other terms that may be used in search queries. Search Component 410 Figure 5 is a diagram that illustrates in further detail an exemplary embodiment of a 410 search component. The 410 search component may include a 505 search engine and a 510 ranking component. The 505 search engine can receive descriptive information from the 405 descriptive information generator and, in response, locate one or more documents relevant to the descriptive information. The 505 search engine can be a query-based search engine that returns a sorted set of documents related to the input search query. The 505 search engine can be a general search engine, such as one based on all documents in a large collection, such as documents on the Web, or a more specialized search engine, such as a news search engine. The techniques for implementing search engines are generally known and, therefore, will not be described further here. The 510 classification component can operate to classify and / or prune the set of documents returned by the 505 search engine. In one embodiment, the 510 classification component can order the returned set of documents based on a query matching count, which defines how well each document in the returned set matches the query. The 510 ranking component can rank documents based on other relevance or quality measurements, such as a measurement of document quality based on links. The top N ranked documents (e.g., N=3) can be selected by the 510 ranking component for presentation to the user. Other techniques can be used to classify or prune the set of relevant documents by the 510 classification component. For example, documents that appear in multiple document sets that match multiple related search queries can be selected, documents that are very recent can be selected, or documents that are the most popular can be selected (e.g., based on the number of times the document link has been selected). As other examples, documents from commercial sites can be explicitly excluded (or included). In some modalities, multiple potential search queries that correspond to the descriptive information may be received, and the queries that return the "best" results may be used. The "best" results can be measured in various ways, such as targeting objective ranking values that correspond to documents returned from a search engine in response to the potential search query. Furthermore, multiple different search engines could be used, such as a news search engine, a product search engine, or a general web-based search engine. Formatting Component 415 The 415 formatting component can incorporate the doeu- Additional documents located by the 410 search component in the current document (that is, the document currently being viewed by the user) or in a new document that includes the current document. The additional documents can be embedded in a way that informs the user that the documents are available, without unduly interrupting the user's reading of the current document. In one embodiment, the formatting component 415 can insert links (e.g., hyperlinks) to additional documents inline with the text of the current document. When possible, the link to each additional document can be inserted in a section of the current document that is particularly relevant to the additional document. This concept is illustrated in Figures 1A and 1B where links to related content, such as a link to a document about Mount Everest, are inserted near the term “Mount Everest” in document 105. Although the links in Figures 1A and 1B are shown as including parenthetical summary information and as links that are identified by modifying the display of words in the current input document, other techniques can be used to graphically display the links. Different inline hyperlinking techniques can be used to embed additional documents within the current document. For example, "floating above" text can be used, which appears when the user positions the cursor over a particular word, image, or other object in the current document. Document Locator Operation 225 Figure 6 is a flowchart illustrating exemplary operations performed by document locator 225. Document locator 225 can start the operation in response to a user requesting a document, such as a request made from a website or a search engine. Document locator 225 may receive or locate the user's personal information (act 601). Personal information may include information such as, for example, the geographic location of the user. The document locator includes personal information provided by the user when registering an account (or at another time), personal information based on the user's browsing history, or personal information extracted from user-generated documents. The document locator also receives the actual input document that the user is requesting (act 602). Descriptive information relating to the input document can be generated (act 603). As previously discussed, this descriptive information can be generated using the descriptive information generator 405 and can include a search query containing terms related to the current input document and the user's personal information. This descriptive information can be used to locate additional relevant documents (act 604). As discussed, this can be done using the search component 410, which submits a search query to a search engine. One or more of the additional relevant documents may be embedded or otherwise associated with the current input document (Act 605). As shown in Figures 1A and 1B, the additional relevant documents may be embedded inline with the current input document. The modified version of the current input document, including the links to the additional relevant documents, may then be presented to the user (Act 606). Exemplary Modality iva from Document Locator 225 Figure 7 is a diagram illustrating an exemplary embodiment of an additional document locator 225 implemented in the context of a content service Network site, such as a Network site dedicated to articles about a particular hobby (e.g., automobiles). A person of ordinary ability in the art will observe that the document locator 225 could be implemented in a number of additional network environments, such as in the general context of a news search engine or a more general search engine. Multiple users 705 can connect to the content network site 710 on a network 715. Users can request specific documents from the content network site 710. Before returning the requested document to the user, the network site 710 can transmit the document (or the information that identifies the document), potentially along with the requesting user's personal information, to the document locator 225. The document locator 225 can return its modified version of the requested document, as previously discussed, to the network site 710, which can then send the document to the user. In this way, documents from the network site 710 can be self-augmented to enhance their appeal before being returned to the user. Many variations on this example are possible. For instance, instead of document locator 225 returning the augmented document to the Network 710 site, the Network 710 site could simply redirect the user's document request to document locator 225, which could then return the augmented document to the user. Conclusion Techniques for automatically locating additional documents relevant to an original document and / or to the user's personal information, such as a document being read by the user, have been described here. In one embodiment, the additional documents were located based on the user's personal information, as well as on content relevant to the document being read by the user. The additional documents can be presented inline with the document being read, such as via links inserted in locations within the document that are particularly relevant to the additional document. The user can thus be efficiently presented with additional information that is relevant to the document being read. It will be evident to a person of ordinary ability in The technique demonstrates that aspects of the invention, as described above, can be implemented in many different forms of software, firmware, and hardware in the embodiments illustrated in the Figures. The actual software code or specialized control hardware used to implement aspects consistent with the invention is not limiting to the invention. Thus, the operation and behavior of the aspects have been described without reference to specific software code – it being understood that a person of ordinary skill in the art could design software and control hardware to implement the aspects based on this description. The preceding description of preferred embodiments of the invention provides illustration and description, but it is not intended to be exhaustive, nor to limit the invention to the precise form disclosed. Modifications and variations are possible by taking into account prior learning or may be acquired from practice of the invention. For example, although many of the operations disclosed above have been described in a particular order, many of the operations are capable of being performed simultaneously or in different orders to achieve the same or equivalent results. No element, act, or instruction used in this Application should be construed as critical or essential to the invention unless explicitly described as such. Also, as used herein, the article “a” is intended to potentially allow for one or more items. Furthermore, the phrase “based on” is intended to mean “based, at least in part, on,” unless explicitly stated otherwise. stated otherwise.
Claims
"Methods for enhancing document navigation and the respective first document and system and device" Claims 1 - Method of intensifying document navigation, characterized by comprising: to receive personal information relating to a user; receive an initial document requested by the user; Generate descriptive information based on the content of the first document and personal information; including personal information such as the user's geographic location, information provided by the user when registering an account, or information based on the user's browsing history; Identify an additional document based on the descriptive information; and Generate a second document that includes at least some of the content of the first document and includes references to the additional document. 2 - A method for enhancing document navigation, according to Claim 1, characterized in that references to the additional document include an embedded link inline with the first document. 3 - A method for enhancing document navigation, as described in Claim 2, characterized in that the link includes text describing the references. 4 - A method for enhancing document navigation, according to Claim 2, characterized in that the link includes floating text. 5 - A method for enhancing document navigation, according to Claim 1, characterized in that the descriptive information includes named entities in the first document that match a list of predetermined named entities. 6 - Method for intensifying document navigation, according to with Claim 5, characterized in that the list of predetermined named entities includes names of consumer locations and products. 7 - A method for enhancing document navigation, according to Claim 1, characterized in that descriptive information based on the content of the first document includes terms that appear more than a predetermined number of times in the first document. 8 - A method for enhancing document navigation, according to Claim 1, characterized in that personal information includes information extracted from user-generated documents. 9 - A method for enhancing document navigation, according to Claim 1, characterized in that personal information includes temporal information. 10 - A method for enhancing document navigation, according to Claim 1, characterized in that it comprises: Format the descriptive information as a search query; and locate additional documents by submitting the search query to a search engine. 11 - System, characterized by comprising: means of receiving personal information relating to a user; means to receive an initial document requested by the user; means to generate descriptive information based on the content of the first document and personal information, including personal information such as the user's geographic location, information provided by the user when registering an account, or information based on the user's browsing history; means to locate additional documents based on the descriptive information; and means to generate a second document that includes the content Content of the first document including references to additional documents. 12 - System according to claim 11, characterized in that it further includes: methods for formatting descriptive information as a search query; and There are ways to locate additional documents by submitting the search query to a search engine. 13 - Method, characterized by comprising: Locate at least one second document that is relevant to a first document, where the relevance of the second document to the first document is based on content from the first document and on personal information of an intended reader of the first document, including personal information such as geographic location, information provided by the reader when registering an account, or information based on the intended reader's browsing history; and Embed a link to the second document within the first document at a location in the first document where the second document is relevant, in order to obtain a modified version of the first document. 14 - A method, according to Claim 13, characterized in that the content of the first document includes named entities matched from a list of named entities predetermined in the first document. 15 - Method, according to Claim 14, characterized in that the list of predetermined named entities includes names of consumer locations and products. 16 - Method, according to Claim 13, characterized in that the content of the first document includes terms that appear more than a predetermined number of times in the first document. 17 - Method, according to Claim 13, characterized in that Personal information includes information extracted from documents generated by the intended reader. 18 - A method, according to Claim 13, characterized in that the second document is embedded in the first document as a hyperlink associated with a named entity in the first document. 19 - Method, according to Claim 13, characterized in that embedding the second document within the first document at a location in the first document for which the second document is relevant includes: further: Insert a hyperlink that includes text describing the second document in the location. 20 - Device, characterized by comprising: a memory, which contains program instructions; and A processor coupled to memory and configured to execute program instructions for: to receive personal information relating to a user; receive an initial document requested by the user; generate descriptive information based on the content of the first document and personal information; including personal information such as the user's geographic location, information provided by the user when registering an account, or information based on the user's browsing history; Locate at least one additional document based on the descriptive information; and Generate a second document that includes the content of the first modified document, so as to include references to at least one additional document. 21 - Method of intensification of the first document, characterized by comprising: to receive personal information relating to a user, including personal information such as the user's geographic location, information provided by the user when registering an account, or information based in a user's browsing history; Generate descriptive information based on the content of the first document and personal information; Format the descriptive information as a search query; and to locate additional documents by submitting the search query to a search engine; and Generate a second document that includes the content of the first document, modified to include an inline embedded link to the first document that references at least one of the additional documents.