Methods and systems for determining a document's meaning to compare the document to the content

BRPI0413097AInactive Publication Date: 2006-10-03GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
BR · BR
Patent Type
Applications
Current Assignee / Owner
GOOGLE LLC
Publication Date
2006-10-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing document analysis methods fail to accurately identify relevant regions and concepts, leading to mismatched content that can dilute the overall meaning and reduce user engagement, particularly in web pages with irrelevant frames or sections.

Method used

A system and method to analyze documents by identifying and ranking local concepts within regions, eliminating irrelevant regions, and combining the remaining concepts to determine a source meaning, which is then matched with appropriate content items.

Benefits of technology

Enhances the relevance of content displayed on documents by ensuring that only relevant regions and concepts are used for matching, thereby improving user engagement and advertiser satisfaction.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

"METHODS AND SYSTEMS FOR DETERMINING A MEANING OF A DOCUMENT TO COMPARE THE DOCUMENT TO THE CONTENT". Systems and methods for determining a document's meaning for comparing the document to the content are described. In one aspect, a source article is accessed, a plurality of regions in the source article are identified, at least one local concept associated with each region is determined, the local concepts of each region are analyzed to identify any unrelated regions, the local concepts associated with any unrelated regions are eliminated to determine relevant concepts, the relevant concepts are analyzed to determine a source meaning for the source article, and the source meaning is compared to an item meaning associated with an item in a set of items.
Need to check novelty before this filing date? Find Prior Art

Description

Descriptive Report of the Invention Patent for: "METHODS "AND SYSTEMS FOR DETERMINING THE MEANING OF A DOCUMENT TO COMPARE THE DOCUMENT TO ITS CONTENT." Field of Invention This invention generally relates to documents. More particularly, the invention relates to methods and systems for determining the meaning of a document in order to match the document to its content. Background of the Invention Documents, such as web pages, can be combined with other content on the internet, for example. Documents include, for example, web pages of various formats, such as HTML, XML, XHTML; Portable Document Format (PDF) files; and word processing and application program document files. An example of combining documents with content is in internet advertising. For example, a website publisher might allow an advertisement for a fee on their web pages. When the publisher wants to display an advertisement on a web page to a user, a The facilitator can provide an announcement to the publisher for display on the web page. The facilitator can select the announcement due to a variety of factors, such as a Demographic information about the user, the category of the web page, for example, sports or entertainment, or the content of the web page. The facilitator can also combine the content of the web page with a piece of knowledge, such as a keyword, from a list of keywords. An advertisement associated with the combined keyword can then be displayed on the web page. A user can manipulate a mouse or other input device and "click" on the advertisement to see a web page on the advertisement's website that offers items or services for sale. In another example of internet advertising, the actual keywords combined are displayed on a publisher's web page in a Related Links section or similar. Similar to the example above, the web page content is combined with one or more keywords, which are then displayed in the Related Links section, for example. When a user clicks on a particular keyword, the user may be directed to a search results page, which may contain a mix of ads and regular search results. Advertisers bid for the keyword to have their ads appear on a search results page like this for the keyword. A user can Manipulate a mouse or other input device and "click" on the ad to view a webpage about the advertiser's website offering items and services for sale. Advertisers want the content of the web page to closely relate to the advertisement because a user viewing the web page is more likely to click on the advertisement and purchase the items or services being offered if they are highly relevant to what the user is reading on the web page. The web page publisher also wants the advertiser's content to match the web page content because the publisher often receives compensation if the user clicks on the advertisement, and a lack of matching could be offensive to the advertiser or the publisher in the case of sensitive content. Documents, such as web pages, can consist of various regions, such as frames in the case of web pages. Some of these regions may be irrelevant to the main content of the document. Therefore, the content of irrelevant regions can dilute the overall content of the document with irrelevant subject matter. Therefore, it is desirable to analyze a source document for its most relevant regions when determining the meaning of the source document, in order to match the document with... content. Summary The embodiments of the present invention comprise systems and methods that determine the meaning of documents by combining the document with content. One aspect of an embodiment of the present invention comprises accessing a source article, identifying a plurality of regions in the source article, determining at least one local concept associated with each region, analyzing the local concepts of each region to identify any unrelated regions, eliminating the local concepts associated with any unrelated regions to determine relevant concepts, analyzing the relevant concepts to determine a source meaning for the source article, and combining the source meaning with an item meaning associated with an item from a set of items. The item may be the content itself or may be associated with content. In one embodiment, the invention further comprises displaying the combined item in the source article.In another embodiment, the invention further comprises displaying the content associated with the item in the source article. Additional aspects of the present invention are directed to computer systems and media that can be read on a computer having such capabilities. relating to the preceding aspects. Brief Description of the Drawings These and other features, aspects and advantages of the present invention are best understood when the Detailed Description is read with reference to the accompanying drawings, where: Figure 1 illustrates a block diagram of a system according to an embodiment of the present invention; Figure 2 illustrates a flowchart of a method according to an embodiment of the present invention; and Figure 3 illustrates a flowchart of a subroutine of the method shown in Figure 2. Detailed Description of Specific Modalities The present invention comprises methods and systems for determining the meaning of a document in order to match the document with its content. Reference will now be made in detail to exemplary embodiments of the invention, as illustrated in the accompanying text and drawings. The same reference numbers are used by all drawings and in the following description for reference to the same parts or to equal parts. Various systems according to the present invention can be constructed. Figure 1 is a diagram illustrating an example system in which example embodiments of the invention are shown. The present invention can operate. The present invention can operate and be implemented in other systems as well. The system 100 shown in Figure 1 includes multiple client devices 102a and 104, server devices 104 and 140, and a network 106. The network 106 shown includes the Internet. In other embodiments, other networks, such as an intranet, may be used. Furthermore, the methods according to the present invention may operate on a single computer. Each of the client devices 102a and 108 shown includes a computer-readable medium, such as random access memory (RAM) 108, in the embodiment shown coupled to a processor 110. The processor 110 executes a set of computer executable program instructions stored in memory 108. These processors may include a microprocessor, an ASIC, and state machines.These processors include, or may communicate with, media, for example, computer-readable media, which store instructions that, when executed by the processor, cause the processor to perform the steps described herein. Types of computer-readable media include, but are not limited to, any electronic, optical, magnetic, or other storage or transmission device capable of providing a... A computer processor, such as a processor communicating with a touch-sensitive input device, transmits instructions that can be read into a computer. Other examples of suitable media include, but are not limited to, a floppy disk, CD-ROM, magnetic disk, memory chip, ROM, RAM, an ASIC, a configured processor, all optical media, all magnetic tape or other magnetic media, or any other media from which a computer processor can read instructions. Also, various other forms of computer-readable media can transmit or carry instructions to a computer, including a router, a private or public network, or another wired and wireless transmission device or channel. The instructions may comprise code in any computer programming language, including, for example, C, C++, C#, Visual Basic, Java, and JavaScript. Client devices 102a an can also include various external or internal devices, such as a mouse, a CD-ROM, a keyboard, a display, or other input or output devices. Examples of client devices 102a an are personal computers, digital assistants, personal digital assistants, telephones cell phones, telephones furniture, telephones Intelligent devices, radio calling equipment, digital tablet computers, laptop computers, processor-based devices, and similar types of systems and devices. In general, a client device can be any type of processor-based platform connected to a network and interacting with one or more application programs. The client devices shown include personal computers running a browser application program, such as Internet Explorer™, version 6.0 from Microsoft Corporation, Netscape Navigator™, version 7.1 from Netscape Communications Corporation, and Safari™, version 1.0 from Apple Computer. Through client devices, users can communicate with each other over the network and with other systems and devices connected to the network. As shown in Figure 1, server devices 104 and 140 are also connected to network 106. The document server device 104 shown includes a server running a document agent application program. The content server device 140 shown includes a server running a content agent application program. System 100 may also include multiple other server devices. Similarly to Client devices 102a and server devices 104, 140 each include a processor 116, 142 coupled to a memory that can be read by computer 118, 144. Each server device 104, 140 is described as a single computer system, but it can be implemented as a network of computer processors. Examples of server devices 104, 140 are servers, mainframe computers, networked computers, a processor-based device, and similar types of systems and devices. The client processors 110 and the server processors 116, 142 can be any of several well-known computer processors, such as processors from Intel Corporation of Santa Clara, California, and Motorola Corporation of Schaumburg, Illinois. Memory 118 of document server device 104 contains a document agent application program, also known as document agent 124. The document agent 124 determines a meaning for a source article and matches the source article with an item, such as another article or a knowledge item. The item can be the content itself or can be associated with the content. Source items can be received from other devices connected to the network 106. Items include Documents, for example, web pages of various formats such as HTML, XML, XHTML, Portable Document Format (PDF) files, and word processing document files, databases and application programs, audio, video, or any other information of any kind that is made available on a network (such as the Internet), a personal computer, or other computing or storage medium. The modalities described here are generally described in relation to documents, but the modalities can operate on any type of item. Knowledge items are generally anything physical or non-physical that can be represented through symbols and can be, for example, keywords, nodes, categories, people, concepts, products, phrases, documents, and other units of knowledge.Knowledge items can take any form, for example, a single word, a term, a short phrase, a document, or some other structured or unstructured information. The modalities described here are generally related to keywords, but the modalities can operate on any type of knowledge item. The document agent 124 shown includes a preprocessor 134, a meaning processor 136, and a combination processor 137. In the embodiment shown, each one comprises a computer code residing in memory. 118. The document agent 124 receives a content request to be placed in a source document. This request may be received from a device connected to the network 106. The content may include documents, such as web pages and advertisements, and knowledge items, such as keywords. The preprocessor 134 receives the source document and analyzes the source document to determine the concepts contained in the document and regions within the document. A concept may be defined using a grouping or a set of words or terms associated with it, where the words or terms may be, for example, synonyms. A concept may also be defined by various other information, such as, for example, relationships with related concepts, the strength of relationships with related concepts, parts of speech, common usage, frequency of use, the breadth of the concept, and other statistics on concept usage in language.The meaning processor 136 analyzes the concepts and regions to eliminate regions unrelated to the main concepts of the source document. The meaning processor 136 then determines a source meaning for the source document from the remaining regions. The processor... Combination 137 combines the meaning of the source document with the meaning of an item from a set of items. Memory 144 of the content server device 140 contains a content agent application program, also known as the content agent 146. In the embodiment shown, the content agent comprises computer code residing in memory 144. The content agent 146 receives the matched item from the document server device 104 and places the item or the content associated with the item in the source document. In one embodiment, the content agent 146 receives a matched keyword from the match processor 137 and associates a document, such as an advertisement, with it. The advertisement is then sent to a requester's website and placed in the source document, such as a frame on a web page, for example. The document server device 104 also provides access to other storage elements, such as a meaning storage element, in the example shown a meaning database 120. The meaning database can be used to store meanings associated with source documents. The content server device 140 also provides access to Other storage elements, such as a content storage element, in the example shown a content database 14 8. The content database can be used for storing items and content associated with the items, such as keywords and associated advertisements. Data storage elements can include any one or a combination of methods for storing data, including, without limitation, arrays, hash tables, lists, and pairs. Other similar types of data storage devices can be accessed by the server devices 104 and 140. It should be noted that the present invention may comprise systems having a different architecture from that shown in Figure 1. For example, in some systems according to the present invention, the preprocessor 134 and the meaning processor 136 may not be part of the document agent 124, and may perform other offline operations. In one embodiment, the meaning of a document is determined periodically as the document agent slowly advances through documents, such as web pages. In another embodiment, the meaning of a document is determined when a request for content to be placed in the document is received. The system 100 shown in Figure 1 is This is merely an example, and is used to explain the example methods shown in Figures 2 and 3. In the example shown in Figure 1, a user 112a can access a document on a device connected to network 106, such as a web page on a website. For example, user 112a can access a web page containing a story about salmon fishing with artificial lures in Washington on a news website. In this example, the web page contains four regions: a title region containing the story title, the author, and a one-sentence summary of the story; a main story section containing the story text and images; a banner ad related to car sales; and a link section containing links to other web pages on the website, such as national news, weather, and sports. The owner of the news website may wish to sell advertising space on the source web page and thus sends a request to document server 104 via network 106 for an item, such as an advertisement, to be displayed on the web page. In order to match the source web page with an item, the meaning of the source web page is first determined. The document agent 124 accesses the source web page and can receive the web page. The The original meaning of the web page may have been previously determined and may be stored in the meaning database 120. If the original meaning has been previously determined, then the document agent 124 retrieves the original meaning. If the original meaning of the web page has not been determined, preprocessor 134 first identifies the concepts contained in the web page and the regions contained within the web page. For example, the preprocessor might determine that the web page has four regions corresponding to the title region, the story region, the banner ad region, and the links region, and that the web page contains concepts related to salmon, lure fishing, Washington, automobiles, news, weather, and sports. The regions do not necessarily correspond to frames on a web page. The meaning agent then determines local concepts for each region and ranks all local concepts. A variety of weighting factors can be used for ranking the concepts, such as region importance, concept importance, concept frequency, the number of regions in which the concept appears, and concept breadth, for example. The meaning processor 136 then identifies Regions that are not related to most concepts are eliminated, and local concepts associated with them are removed. In the example, the band region and the link region do not contain concepts particularly relevant to the story, and thus, concepts related to these regions are eliminated. The meaning agent then determines an origin, based on the remaining concepts. The meaning could be a vector of weighted concepts. For example, the meaning could be salmon (40%), lure fishing (40%), and Washington (20%). This meaning can be combined with an item by the matching processor 137. Items can include documents, such as web pages and advertisements, and knowledge items, such as keywords, and can be received from the content server device 140. Items can be stored in the content database 148. For example, if the items are keywords, such as lure fishing, backpacker-type trips, CDs, and travel, the matching agent determines a match. Guiding factors, such as cost-per-click data associated with each keyword, can be used. For example, if the keyword match lure fishing is a closer match than the meaning of the keyword travel, but the If an ad that currently uses the keyword "travel" has a higher cost-per-click rate, the meaning agent can match the source meaning with the keyword "travel." Content filters can also be used to filter out any adult or sensitive content. The combined keyword can be received by the content server device 140. The content agent 146 associates an advertisement with the combined keyword and displays it on the source web page. For example, if the keyword "travel" were combined, the content agent would display the advertisement associated with the keyword "travel" on the source web page containing the story about fishing for salmon with artificial lures in Washington. If the user 112a points their input device at the advertisement and clicks on it, the user may be directed to a web page associated with the advertisement. Several methods according to the present invention can be implemented. One example method according to the present invention comprises accessing the source article, identifying a plurality of regions in the source article, determining at least one local concept associated with each region, analyzing the local concepts of each region, in order to identify any unrelated regions for the Determining relevant concepts involves analyzing the relevant concepts to determine a source meaning for the source article, and matching the source meaning with an item meaning associated with an item from a set of items. Guiding factors can be used to match the source meaning with an item meaning. The item meaning can be a vector of weighted concepts. In some modalities, the method still involves displaying the combined item within the source item. In these modalities, the source item may be a web page and the combined item may be a keyword. Alternatively, the source item may be a web page and the combined item may be an advertisement. In some modalities, the method also includes displaying content associated with the combined item in the source article. In these modalities, the source article may be a web page, the combined item may be a keyword, and the associated content may be an advertisement. Alternatively, the source article may be a web homepage, the combined item may be a second web page, and the associated content may be an advertisement. Associated content can be a link to a second web page. In some modalities, determining at least one local concept involves assigning a score to each local concept in each region. The local concepts in each region with the highest scores are the most relevant local concepts. Furthermore, identifying unrelated regions involves first determining a revised score for each local concept. Then, a ranked global list is determined containing all local concepts based on the revised scores. Local concepts whose combined revised score contributes less than a predetermined amount to the total score for the global list are removed to produce a resulting list. Then, uncorrelated regions with local concepts that are not the most relevant in the resulting list are determined. Local concepts associated with unrelated regions are then removed from the resulting list to produce a list of relevant concepts.Furthermore, a meaning of origin is determined by normalizing the revised scores for the relevant concepts. Another example method according to the present The invention comprises access one original article, Identify at least one first content region and a second content region in the source article, determine at least one first local concept associated with the second content region, match the first content region with a first item from a set of items based at least in part on the first local concept, and match the second content region with a second item from the set of items based at least in part on the second local concept. Figures 2 to 3 illustrate an example method 200 according to the present invention in detail. This example method is provided by way of example, as there are a variety of embodiments of the methods according to the present invention. The method 200 shown in Figure 2 can be performed or otherwise carried out by any of several systems. The method 200 is described below as carried out by the system 100 shown in Figure 1 by way of example, and several elements of the system 100 are referenced in the explanation of the example method in Figures 2 to 3. The method 200 shown provides a determination of the meaning of a source document for matching the source document with an item. Each block shown in Figures 2 to 3 represents one or more steps performed in Example Method 200. Referring to Figure 2, in block 202, example method 200 begins. Block 202 is followed by block 204, in which a document is accessed. The document can be accessed and received, for example, from a device on network 106 or from other sources. Block 204 is followed by block 206, in which a meaning for the source document is determined. In the embodiment shown, a meaning is determined for the source document by separating the document into regions, eliminating useless regions, and analyzing the concepts contained in the remaining regions of the document. For example, in the embodiment shown, preprocessor 134 initially determines the concepts contained in the source document and determines regions within the document. Meaning processor 136 classifies the concepts and removes regions and associated concepts unrelated to the majority of concepts. From the remaining concepts, meaning processor 136 determines a source meaning for the document. Figure 3 illustrates a subroutine 206 for performing method 200 shown in Figure 2. Subroutine 206 provides a meaning for the received source document. An example of the subroutine is as follows. The subroutine begins in block 300. In block 300, the The source document is pre-processed to determine the concepts contained within it. This can be done using natural language processing and text processing to decipher the document into words and then match the words with concepts. For example, cards corresponding to words are first determined by natural language processing and text processing and combined with cards contained in a semantic network of interconnected meanings. From the combined cards, the terms are then determined from the semantic network. Concepts for the determined terms are then assigned and given a probability of being related to the terms. Block 300 is followed by block 302, in which document regions are identified. Document regions can be determined, for example, based on certain heuristics, including information formatting. For example, for a source document that is a web page, which includes HTML labels, the labels can be used to help identify regions. For example, text between tags <title> ... < / title> This can be marked as text in a heading region. 0 text in a paragraph where more than seventy percent of the text is between tags. ... can be marked as in a region The structure of the text can also be used to help identify regions. For example, text in short paragraphs or columns in a table, without the structure of a sentence, such as, for example, without a verb, with too few words or no punctuation until the end of the sentence, can be marked as being in a list region. Text in long sentences, with verbs and punctuation, can be marked as part of a text region. When the region type changes, a new region can be created starting with the text marked with the new type. In one mode, if a text region takes up more than twenty percent of the document, it can be divided into smaller pieces. Block 302 is followed by block 304, in which the The most relevant concepts for each region are determined. In the modality shown, the meaning processor 136 processes the concepts identified for each region to suggest a smaller set of local concepts for each region. The relationships between concepts, the frequency of occurrence of the concept in the region, and the... The breadth of the concept can be used in determining local concepts. In one modality, for each region, every concept is placed in a list. The concepts are classified in The ranking is done by determining a score for each concept using a variety of factors. For example, if a first concept has a strong connection with other concepts, this is used to intensify the score of the first concept and its related concepts. This effect is diluted by the frequency of occurrence of the first concept and by the focus (or breadth) of the first concept, to decrease the score of very common concepts and concepts that are of broader significance. Concepts whose frequencies are above a certain threshold may be filtered. The perceived importance of the concept also impacts the concept score. The importance of a concept can be determined earlier in the processing, for example, by whether or not the words that caused the inclusion of the concept are marked in bold. After the concepts for each region are ranked, the least relevant concepts may be removed.This can be done by selecting a regulated number of the highest-ranking concepts or by removing concepts that have a ranking score below a certain score. Block 304 is followed by block 306, in which the local concepts for each region are combined and analyzed. In the modality shown, meaning processor 136 receives all the local concepts for each The system sorts the region and creates a globally ranked list of all local concepts, for example, by a score for each local concept. Guiding factors, such as the importance of each region, can be used to determine the score. The importance of each region can be determined by the type of region and the size of the region. For example, a title region might be considered more important than a links region, and concepts appearing in the title region might receive more weight than concepts in the links region. Additional weight can be given to concepts that appear in more than one region. For example, duplicate concepts can be merged and their scores added together. This global list can then be ranked, and final concepts contributing less than twenty percent, for example, of the sum of the scores, can be removed to produce a resulting global list of local concepts. Block 306 is followed by block 308, in which regions whose main concepts refer to unrelated concepts are eliminated. In the modality shown, meaning processor 136 determines unrelated regions, regions containing concepts unrelated to most concepts, and eliminates them. It should be understood what related unrelated They don't have to be determined using absolute criteria. "Related "Related" indicates a relatively high degree of relatedness and / or a predetermined degree of relatedness. "Unrelated" indicates a relatively low degree of relatedness and / or a predetermined degree of relatedness. By eliminating unrelated regions, associated unrelated concepts are eliminated. For example, if the source document for a web page consists of several frames, some of the frames will relate to advertisements or links to other pages on the website and thus be unrelated to the main meaning of the web page. In one modality, for example, the resulting global list determined in block 306 may be an approximation of the document's meaning and may be used to remove regions that are not related to the document's meaning. The meaning processor 136 may determine, for each region, whether the most representative local concepts for the region are not present in the resulting global list. If the most representative local concepts for a region are not in the list, the region may be marked as irrelevant. The most representative local concepts for a region may be the concepts with the highest scores for the region, as determined in block 304, for example. Block 308 is followed by block 310, in which the meaning of the source document is determined. In the mode shown, meaning processor 136 recalculates the representativeness of local concepts for the non-eliminated regions to create a relevant list of concepts. These local concepts in the relevant list can then be selected for a fixed number of concepts, for the provision of a meaning list, and then normalized to provide a source meaning. For example, a meaning list can be created using only concepts contained in relevant regions, and everything except the twenty-five highest-scoring concepts is removed from the new list. The scores of the highest-scoring concepts can be normalized to provide a source meaning. In this example, the The meaning of origin can be a weighted vector of relevant concepts. With reference, again, to Figure 2, block 206 is followed by block 208, in which a set of items is received. Items may be received, for example, by the combination processor 137 from the content server device 140. Items may include, for example, knowledge items such as keywords, and documents, such as advertisements and web pages. Each The received item may have a meaning associated with it. For keyword meanings, for example, these can be determined through the use of information associated with the keyword, as described in U.S. Patent Application Serial No. 10 / 690,328 (Legal Protocol No. 53051 / 288071) entitled "Methods and Systems for Understanding a Meaning of a Knowledge Item Using Information Associated with the Knowledge Item," which is thus incorporated by reference. The meaning of a document can be determined in the same way as described with respect to Figure 3, for example. Block 208 is followed by block 210, in which the source document is combined with an item. Guiding factors can be used in the matching process. In one modality, the source meaning is combined with a keyword meaning associated with a keyword from a set of keywords. The matching agent compares the source meaning with the keyword meanings and uses guiding factors, such as cost-per-click data, associated with the keywords to determine a match. This matched keyword can then be sent to the content server device 140. The content agent 146 can match the matched keyword with your ad. Associated and displaying the ad in the source document. Alternatively, the content agent can display the keyword itself in the source document. In another mode, the meanings for ads are combined with the source meaning. In this mode, content agent 146 can cause the combined ad to be displayed in the source document. In another mode, the meanings for web pages are combined with the source meaning. In this mode, content agent 146 can cause an ad associated with the web page to be displayed. Block 210 is followed by block 212, where the method ends. In one embodiment, after the source document is accessed, the source document is analyzed by preprocessor 134 to determine the content regions of the source document. Content regions can be regions containing a substantial amount of text, such as, for example, a text region or a link region, or they can be a region of relative importance, such as, for example, the title region. These regions can be determined through the use of heuristics, as described above. Preprocessor 134 can also identify concepts located in each content region, as described above. These concepts can be The meaning processor 136 is used to determine a meaning for each content region. The combination processor 137 can combine the meaning of each content region with a keyword. The content agent 146 can combine the combined keyword with its associated ad and display the ad in the source document. Alternatively, the content agent can display the keyword itself in the source document. In another mode, meanings for ads are combined with region meanings. In this mode, the content agent 146 can cause the combined ad to be displayed in the source document. In another mode, meanings for web pages are combined with region meanings. In this mode, the content agent 146 can cause an ad associated with the web page to be displayed. In one mode, ads or keywords are displayed in the content region with which they are combined. Although the above description contains many specifics, these specifics should not be construed as limitations on the scope of the invention, but merely as examples of the embodiments shown. Those skilled in the art will discern many other possible variations that are within the scope of the invention.

Claims

CLAIMS 1) A method characterized by the fact that it includes: accessing a source article; to identify a plurality of regions in the source article; to determine at least one local concept associated with each region; Analyze the local concepts of each region to identify any unrelated regions; eliminate local concepts associated with any regions in order to determine relevant concepts; analyze the relevant concepts to determine a source meaning for the source article; and Compare the source meaning with an item meaning associated with an item in a set of items. 2) Method, according to claim 1, characterized in that it further comprises displaying the compared item in the source article. 3) Method, according to claim 2, characterized in that the source article is a web page and the compared item is a keyword. 4) Method according to claim 2, characterized in that the source article is a web page and the compared item is an advertisement. 5) Method according to claim 1, characterized because it still includes the display of content associated with the compared item in the source article. 6) Method, according to claim 5, characterized in that the source article is a web page, the compared item is a keyword, and the associated content is an advertisement. 7) Method, according to claim 5, characterized in that the source article is a first web page, the compared item is a second web page, and the associated content is an advertisement. 8) Method according to claim 5, characterized in that the source article is a first web page, the compared item is a second web page, and the associated content is a link to the second web page. 9) Method, according to claim 1, characterized in that the comparison of the source meaning with an item meaning comprises the use of bias factors. 10) Method, according to claim 1, characterized in that the source meaning is a vector of weighted concepts. 11) Method, according to claim 1, characterized in that the determination of at least one local concept To understand the determination of a brand for each local concept, where the local concepts in each region with the highest brands are the most relevant local concepts. 12) Method according to claim 11, characterized in that the identification of unrelated regions comprises determining a revised mark for each local concept, determining a categorized global list of all local concepts based on the revised marks, removing local concepts whose combined revised marks contribute less than a predetermined amount of a total mark to the global list to produce a resulting list, determining unrelated regions with no more relevant local concepts in the resulting list, and removing local concepts associated with the unrelated regions from the resulting list in order to produce a list of relevant concepts. 13) Method, according to claim 12, characterized in that the determination of a source meaning comprises the standardization of revised marks for the relevant concepts. 14) Computer-readable media containing code a program characterized by the fact that it includes: Program code for accessing a source article; Program code to identify multiple regions in the source article; Program code to determine at least one local concept associated with each region; Program code to analyze the local concepts of each region to identify any unrelated regions; Program code to eliminate local concepts associated with any regions in order to determine relevant concepts; program code to analyze the relevant concepts to determine a source meaning for the source article; and Program code to compare the source meaning with an item meaning associated with an item in a set of items. 15) Computer-readable media according to claim 14, characterized in that it further comprises program code for displaying the item compared in the source article. 16) Computer-readable media according to claim 15, characterized in that the source article is a web page and the compared item is a word. key. 17) Computer-readable media, in accordance with with Claim 15, characterized by the fact that the source article is a web page and the item being compared is an advertisement. 18) Computer-readable media according to claim 14, characterized in that it further comprises program code for displaying the content associated with the item compared in the source article. 19) Computer-readable media according to claim 18, characterized in that the source article is a web page, the compared item is a keyword, and the associated content is an advertisement. 20) Computer-readable media according to claim 18, characterized in that the source article being a first web page, the compared item being a second web page, and the associated content being an advertisement. 21) Computer-readable media, <i.e., in accordance with the Claim 18, characterized by the fact that the source article is a first web page, the compared item is a second web page, and the associated content is a link to the second web page. 22) Computer-readable media according to claim 14, characterized in that the program code for comparing the meaning of the source with a Meaning of item: understand program code for the use of slope factors. 23) Computer-readable media according to claim 14, characterized in that the source meaning is a vector of weighted concepts. 24) Computer-readable media according to claim 14, characterized in that the program code for parsing relevant local concepts comprises program code for categorizing relevant local concepts. 25) Computer-readable media according to claim 14, characterized in that the program code for determining a mark for each local concept, wherein the local concept in each region with the highest marks are the most relevant local concepts. 26) Computer-readable media according to claim 25, characterized in that the program code for identifying unrelated regions comprises program code for determining a revised mark for each local concept, program code for determining a globally categorized list of all local concepts based on the revised marks, and program code for removing local concepts whose Combined revised marks contribute less than a predetermined amount of a total mark to the global list to produce a resulting list, program code to determine unrelated regions with no more relevant local concepts in the resulting list, and program code to remove local concepts associated with the unrelated regions from the resulting list in order to produce a list of relevant concepts. 27) Computer-readable media according to claim 26, characterized in that the program code for determining a source meaning comprises program code for standardizing revised marks for the relevant concepts. 28) Method characterized by the fact that it includes: access a source article; Identify at least one primary content region and one secondary content region in the source article; Determine at least one first local concept associated with the first content region and determine at least one second local concept associated with the second content region; to compare the first content region with the first item in a set of items based at least in part on the first local concept; and Compare the second content region with a second item from a set of items based at least in part on the second local concept. 29) Method according to claim 28, characterized in that it further comprises displaying compared items in the source article. 30) Method according to claim 29, characterized in that the first item is displayed in The first content region and the second item will be shown in the second content region. 31) Method according to claim 29, characterized in that the source article is a web page and the items compared are advertisements. 32) Method, according to claim 29, characterized in that the source article is a web page, and the items compared are keywords. 33) Method according to claim 28, characterized in that it further comprises displaying a first content associated with the first item and displaying a second content associated with the second first item in the source article. 34) Method according to claim 33, characterized in that the first content is shown in the first content region and the second content is shown in the second content region. 5 35) Method according to claim 33, characterized by the fact that the source article is a web page, the items compared are keywords, and the associated content is advertisements.