Webpage content tracing method, knowledge graph construction method, and related device
By constructing a knowledge graph to automatically trace the source of web page content, the problem of manually searching for the source of web page references is solved, and efficient web page content tracing is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2021-09-18
- Publication Date
- 2026-05-19
AI Technical Summary
When users manually search for the source of content referenced on a webpage, the process is cumbersome and inefficient, making it impossible to achieve efficient webpage content tracing.
By constructing a knowledge graph, the entities and relationships in the knowledge graph are used to automatically trace the content of web pages, including querying the web page entities of the web page to be traced in the knowledge graph, identifying the target entity, and displaying the tracing results.
It has enabled automated web page content tracing, improved tracing efficiency, simplified user operations, and enhanced user experience.
Smart Images

Figure CN115840863B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to methods for tracing web page content, methods for constructing knowledge graphs, and related devices. Background Technology
[0002] When a webpage on the internet references content from other webpages, the webpage will usually indicate the source of the content using words such as "reference" or "image source." When indicating the source, the webpage can include the name of the website containing the reference information, for example, "Data source: Tencent."
[0003] In practice, if a user visits a webpage containing cited content and wants to trace the source of that content based on the source indicated on the webpage to find the webpage that first published the cited content, the user can only manually search and filter on the internet using a search engine based on the source indicated on the webpage. This process is very cumbersome and inefficient. Summary of the Invention
[0004] In view of this, it is necessary to provide methods for tracing web page content, methods for constructing knowledge graphs, and related equipment, which can overcome the above problems, realize automated web page content tracing, eliminate the process of users manually searching for sources, and greatly improve the efficiency of web page content tracing.
[0005] In a first aspect, one embodiment of this application provides a webpage content tracing method applied to a server, the method comprising:
[0006] The process involves querying the first webpage entity corresponding to the webpage to be traced in the knowledge graph, which includes multiple entities and the relationships between them; determining at least one target entity based on the knowledge graph and the first webpage entity, wherein there is a direct or indirect relationship between the at least one target entity and the first webpage entity; and determining the tracing result of the webpage to be traced, which includes at least one webpage or website corresponding to at least one target entity and the relationships between each webpage or website.
[0007] By adopting this technical solution, knowledge graphs can be used to automatically trace the source of web pages, effectively improving the efficiency of web page content tracing.
[0008] In one possible implementation, the multiple entities include at least one website entity and at least one webpage entity, and the relationships between the entities include referencing relationships and / or attribution relationships, which are determined by the relationship attributes of the website entity or the relationship attributes of the webpage entity.
[0009] Among them, relationship attributes can include reference object attributes and belonging object attributes.
[0010] By adopting this technical solution, the first webpage entity corresponding to the webpage to be traced can be identified from multiple webpage entities and multiple website entities in the knowledge graph. Based on the attribution and reference relationships, the target entity with a direct or indirect relationship with the first webpage entity can be identified, thereby achieving automated webpage tracing and improving the efficiency of content tracing.
[0011] In one possible implementation, the webpage entity also includes a webpage address attribute. Querying the first webpage entity corresponding to the webpage to be traced in the knowledge graph includes: determining the first webpage entity corresponding to the webpage to be traced in the knowledge graph based on the webpage address of the source webpage and the webpage address attributes of all webpage entities in the knowledge graph.
[0012] By adopting this technical solution, the first webpage entity corresponding to the webpage to be traced in the knowledge graph can be accurately determined based on the attribute value (i.e., webpage address) of the webpage address attribute of each entity in the knowledge graph and the webpage address of the webpage to be traced.
[0013] In one possible implementation, the webpage entity also includes a webpage identifier attribute. Querying the first webpage entity corresponding to the webpage to be traced in the knowledge graph includes: generating a webpage identifier corresponding to the webpage to be traced based on the webpage address of the webpage to be traced; and determining the first webpage entity corresponding to the webpage to be traced in the knowledge graph based on the webpage identifier corresponding to the webpage to be traced and the webpage identifier attributes of all webpage entities in the knowledge graph.
[0014] By adopting this technical solution, the webpage identifier of the webpage to be traced can be generated from the webpage address of the webpage to be traced, and the first webpage entity corresponding to the webpage to be traced in the knowledge graph can be accurately determined by the attribute value (i.e., webpage identifier) of the webpage identifier attribute of each entity in the knowledge graph.
[0015] In one possible implementation, determining at least one target entity based on the knowledge graph and the first webpage entity includes: determining at least one candidate entity based on the knowledge graph and the first webpage entity; and determining at least one target entity from the at least one candidate entity based on preset attributes of each candidate entity and preset attributes of the first webpage entity.
[0016] By adopting this technical solution, one or more candidate entities corresponding to the first webpage entity can be obtained in the knowledge graph. Based on the preset attributes of the entity, the candidate entities are filtered to obtain at least one target entity. The preset attributes here may include one or more of the pre-determined attributes such as keyword attributes and summary attributes. By utilizing the attributes of entities within the knowledge graph, the target entity among the candidate entities can be determined efficiently and accurately.
[0017] In one possible implementation, before querying the first webpage entity corresponding to the webpage to be traced in the knowledge graph, the method further includes: obtaining the knowledge graph.
[0018] By adopting this technical solution, knowledge graphs can be obtained from other computer devices and used locally for web page content tracing. Knowledge graphs in different fields may be different, and the storage resources occupied by a single knowledge graph may be large. Obtaining knowledge graphs from other computer devices can effectively save local storage resources and provide web page content tracing services in more fields.
[0019] In one possible implementation, the method further includes: sending the tracing result to the terminal, causing the terminal to render and display a user interface based on the tracing result. The user interface includes an image of the webpage to be traced, an image of the website or webpage corresponding to at least one target entity, and a relationship identifier between the image of the webpage to be traced and the image of the website or webpage corresponding to at least one target entity. The relationship identifier is determined based on the relationship between the first webpage entity and at least one target entity.
[0020] Secondly, one embodiment of this application provides a method for tracing the source of web page content, applied to a terminal, the method comprising:
[0021] Based on the webpage address of the webpage to be traced input by the user, a traceability request is generated for the webpage to be traced; the traceability request is sent to the server so that the server can determine the traceability result of the webpage to be traced based on the webpage address contained in the traceability request in the knowledge graph; the traceability result returned by the server is received, and the webpage to be traced and the images of the webpages or websites referenced by the webpage to be traced are displayed on the user interface according to the traceability result.
[0022] By adopting this technical solution, users can simply enter the web address of the webpage to be traced to view the traceability results, eliminating the need for manual searching and greatly simplifying user operations and improving the user experience.
[0023] Thirdly, one embodiment of this application provides a knowledge graph construction method, the method comprising:
[0024] Identify multiple websites and their associated web pages for constructing a knowledge graph; identify the content of the web pages within the websites; construct a knowledge graph based on the content of the web pages and the relationship between the websites and the web pages, wherein the knowledge graph includes multiple entities and the relationships between these entities.
[0025] By adopting this technical solution, a knowledge graph belonging to a certain domain can be constructed based on websites on the Internet. This knowledge graph can be used to automate the tracing of web page content, thereby improving the efficiency of web page content tracing.
[0026] In one possible implementation, relationships include referencing relationships and attribution relationships. Constructing a knowledge graph based on the content of multiple internal web pages and the attribution relationships between multiple websites and these internal web pages includes:
[0027] Based on the identification results of the web page content of multiple internal web pages, at least one referencing entity that has a referencing relationship with the corresponding entity of each internal web page is identified, and the web page or website corresponding to the referencing entity is the web page or website that is referenced by the internal web page.
[0028] A knowledge graph is constructed based on the reference relationships between multiple entities corresponding to multiple internal web pages and at least one corresponding referencing entity, as well as the affiliation relationships between multiple entities corresponding to internal web pages and the corresponding entities of the websites to which they belong.
[0029] In one possible implementation, multiple entities include multiple attributes, each attribute including at least one attribute value, the entities include at least one website entity and at least one web page entity, and the relationships include reference relationships between website entities or web page entities, and attribution relationships between website entities.
[0030] Fourthly, an embodiment of this application also provides a computer device, which includes at least one processor, a memory, and a communication module; the at least one processor is connected to the memory and the communication module; the memory is used to store instructions, the processor is used to execute instructions, and the communication module is used to communicate with the device under the control of the at least one processor; when the instructions are executed by the at least one processor, the at least one processor performs the web page content tracing method of any possible implementation of the first aspect or the second aspect, or the knowledge graph construction method of any possible implementation of the third aspect.
[0031] Fifthly, an embodiment of this application also provides a computer-readable storage medium storing a program that causes a computer device to execute the web page content tracing method of any possible implementation of the first or second aspect, or the knowledge graph construction method of any possible implementation of the third aspect.
[0032] Sixthly, an embodiment of this application also provides a computer program product, the computer program product including computer execution instructions, the computer execution instructions being stored in a computer-readable storage medium; at least one processor of the computer device can read the computer execution instructions from the computer-readable storage medium, and the at least one processor executes the computer execution instructions causing the computer device to perform the web page content tracing method of any possible implementation of the first aspect or the second aspect, or the knowledge graph construction method of any possible implementation of the third aspect.
[0033] For a detailed description of aspects four through six and their various implementations in this application, please refer to the detailed descriptions in aspects one, two, three and their various implementations; and for a detailed analysis of the beneficial effects of aspects four through six and their various implementations, please refer to the beneficial effect analyses in aspects one, two, three and their various implementations, which will not be repeated here. Attached Figure Description
[0034] Figure 1 A schematic diagram illustrating a scenario for tracing web page content based on knowledge graphs, as provided in this application;
[0035] Figure 2 This is a diagram illustrating the execution system architecture of the webpage tracing method described in this application.
[0036] Figure 3 A flowchart illustrating the knowledge graph construction method provided in this application;
[0037] Figure 4 Example diagram of attributes of the web page entities provided in this application;
[0038] Figure 5 Example diagram of the attributes of the website entity provided in this application;
[0039] Figure 6 Example diagram of the knowledge graph provided in this application;
[0040] Figure 7 A flowchart illustrating the process of tracing web page content based on knowledge graphs provided in this application;
[0041] Figure 8 Example diagram of the user interface for displaying content tracing results provided in this application;
[0042] Figure 9 The overall execution flowchart for constructing knowledge graphs and tracing web page content provided in this application;
[0043] Figure 10 This is a schematic diagram of the structure of a possible computer device provided in this application. Detailed Implementation
[0044] It should be noted that in this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and drawings of this application are used to distinguish similar objects, not to describe a specific order or sequence.
[0045] The method in this application can be executed by at least one computer device, which may include a terminal, a server, etc. The terminal may include a laptop, smartphone, desktop computer, tablet computer, smart wearable device, smart TV, smart screen, etc., and the server may include a local server, cloud server, etc. The computer devices can be connected to each other via wired or wireless means.
[0046] For example, see Figure 1 The method of this application can be jointly executed by terminal 10 and server 20. Specifically, terminal 10 can receive the web address of the web page to be traced by the user and send the web address to server 20. Server 20 can trace the web address through a knowledge graph to obtain the content tracing result and send the content tracing result to terminal 10. Terminal 10 can display the tracing result of the web page to be traced to the user on the terminal page according to the received content tracing result.
[0047] See Figure 2 This application may include two systems during execution. The offline system can construct a knowledge graph, crawl web page data from the website to obtain multiple web pages, parse and process the web page content of the web pages to obtain the attribute information of the web pages, and combine the attribute information of the web pages / website in the manual database construction module to construct and save the knowledge graph.
[0048] The online system can perform content tracing on web pages to be traced. It can receive web page information (such as URL) input by the user, query the content tracing results of the web page to be traced in the knowledge graph, that is, trace the web page to be traced, and parse and display the content tracing results.
[0049] The website in this application can be a collection of web pages, which can contain content in the form of text, images, etc., for users to browse. For example, the website can be CCTV News, and the news web pages under CCTV News can be, for example, epidemic report web pages, weather forecast web pages, etc. The websites and web pages in this application can have different categories of relationships.
[0050] The relationship between web pages and websites can include referencing and attribution relationships. For example, CCTV News may include a webpage about the COVID-19 pandemic; this webpage has an attribution relationship with CCTV News. Similarly, if the text content of a COVID-19 pandemic webpage is quoted from Tencent.com, there is a referencing relationship between the two websites. The relationship between web pages can also include referencing relationships. For instance, if the image content of a COVID-19 pandemic webpage is taken from an information delivery webpage, there is a referencing relationship between the two websites.
[0051] In this embodiment, the knowledge graph construction method will be described in detail, see [link to relevant documentation]. Figure 3 , Figure 3 This is a flowchart illustrating the knowledge graph construction method provided in this embodiment. The method may include:
[0052] 101. Identify multiple websites for building the knowledge graph.
[0053] Specifically, multiple websites can be selected for constructing the knowledge graph based on its characteristics (such as the domain to which the content in the knowledge graph belongs) and the characteristics of the website (such as whether the website is an official website or a large website with high traffic).
[0054] 102. Determine the multiple internal web pages contained in each website, and obtain the attribute information of the corresponding website entity for each website, and the attribute information of the corresponding web page entity for each internal web page.
[0055] Attribute information can include information that reflects certain characteristics of the target website / webpages within the site. Attribute information can be recorded by attributes and attribute values. For example, the attributes of website A can include industry, nature, etc., and the attribute values corresponding to these attributes can be scientific research, official, etc. As another example, the attributes of webpage 1 can include keywords, URL, belonging object, referencing object, etc., and the attribute values corresponding to these attributes can be scientific research, URL 4, website A, website C, etc.
[0056] An attribute can have one or more attribute values. For example, the attribute values for the alias of website A can include "x bean" or "A website".
[0057] In some embodiments, a web crawler can be used to crawl data from a target website to determine all the web pages within the target website, and also to determine some attribute information of the web page entities corresponding to these web page entities, such as the attribute value of the object to which they belong. For example, a web crawler can be used to crawl data from website A to obtain 20 web pages within website A, and at the same time, it can be determined that the attribute value of the object to which the web page entity belongs for each web page is: website A.
[0058] In some embodiments, some attribute information of a website / webpage entity needs to be determined manually. For example, the attribute value of the alias of a website entity can be manually input. For instance, it can receive the attribute value of the alias of the website entity corresponding to website A: x bean, A network.
[0059] In some embodiments, the method for obtaining the attribute information of the web page entity corresponding to the web page within the site can be: to perform recognition processing on the web page content of the web page within the site to obtain the attribute information of the web page within the site, which is the attribute information of the web page entity corresponding to the web page within the site. Specifically, the technology used for recognition processing can be flexibly selected according to the form of the web page content. For example, the web page content can be in the form of images, videos, audio, text, etc., and can be recognized and processed by image recognition, video semantic recognition, audio recognition, text recognition, etc.
[0060] In some embodiments, the webpage content can be text content. In this case, the attribute values of certain attributes of the webpage can be obtained from the webpage content. Specifically, the text content can be identified. When a preset attribute character is identified in the text content, the attribute text that satisfies the first positional relationship with the preset attribute character is extracted from the text content, and the attribute text is determined to be the feature information of the webpage under the attribute features.
[0061] For example, if the webpage content is the text of an academic paper, the text will usually contain the words "Abstract" and "Supervisor," and the content information of the abstract and the name of the supervisor will be recorded in the adjacent positions of these words. Therefore, by identifying whether there are preset attribute characters (such as "Abstract," "Supervisor," etc.) in the text content, the attribute value corresponding to the preset attribute characters can be extracted from the text content. For example, when the text content of webpage 1 contains "Abstract" (i.e., the preset attribute character of the abstract), the attribute text adjacent to "Abstract" (i.e., satisfying the first positional relationship) can be extracted from the text content, and this attribute text is determined to be the attribute value of the abstract attribute of webpage 1.
[0062] The referencing object can record another webpage or website from which the content of a webpage originates. For example, if the referencing object of webpage 1 is webpage 2, it means that the content of webpage 1 is referenced from webpage 2. The attribution object can record the website to which a webpage belongs. For example, if the attribution object of webpage 1 is website A, it means that webpage 1 is a webpage on website A.
[0063] The default attribute characters for a webpage's referenced object can include: "reference", "image source", "excerpt", "reprinted from", "cr", "references", etc. Similarly, determining the attribute value of a referenced object can be achieved, for example, by recognizing the text content of webpage 1. When the default attribute characters for a referenced object are detected in the text content of webpage 1, the adjacent identifier text "CCTV News" can be extracted from the text content. Based on this, the attribute value of the referenced object on webpage 1 can be determined to be: CCTV News.
[0064] 103. Construct a knowledge graph based on the attribute information of each website entity and the attribute information of each webpage entity.
[0065] A knowledge graph can include a directed graph that reveals the relationships between web pages within a site and target websites. A knowledge graph can include multiple web page entities and website entities, which can be connected by directed lines. These directed lines represent the relationships between the two connected entities. Relationships can include reference relationships between web page entities or website entities, indicating that the content of a web page corresponding to a web page entity references another web page entity or a website entity. For example, a one-way relationship between web page entity 1 and web page entity 2 indicates that the content of the web page corresponding to web page entity 1 is referenced from the web page corresponding to web page entity 2. Relationships can also include attribution relationships between web page entities and website entities, indicating that the web page corresponding to a web page entity is a web page within the website entity corresponding to that website entity. For example, an attribution relationship between web page entity 1 and website entity 1 indicates that the web page corresponding to web page entity 1 belongs to the website entity corresponding to website entity 1, and so on.
[0066] Knowledge graphs can be constructed using different methods (such as top-down or bottom-up). The constructed knowledge graphs can be stored in a database (such as a graph database). The specific method can be chosen based on the actual data situation, and no restrictions are imposed here.
[0067] This embodiment can construct a knowledge graph that shows the relationship between web pages within a site and target websites. Then, the knowledge graph can be used to automatically trace the source of web pages on the Internet, eliminating the need for users to manually search and trace the source, and effectively improving the efficiency of web page content tracing.
[0068] The knowledge graph construction method will be described below in conjunction with specific application scenarios. One application scenario of this application is: constructing a knowledge graph in the field of health and wellness. The knowledge graph construction method in this application scenario can be implemented by computer equipment.
[0069] Specifically, the process of building a knowledge graph in the field of health and wellness may include: identifying the websites from which data will be collected, such as CCTV News, the State Council client, the National Health Commission, and Tencent.
[0070] Then, a web crawler service can be used to crawl data from each website, obtaining the multiple web pages contained in the website and the content data of each web page.
[0071] The web pages required to construct a knowledge graph in the field of health and wellness are those containing health and wellness information. However, the web page data obtained may not necessarily contain health and wellness information. For example, the web pages contained in a comprehensive website may also include those containing weather information, entertainment information, etc. Therefore, it is necessary to filter the obtained web pages and retain those that contain health and wellness information (for ease of description, web pages containing health and wellness information are referred to as health and wellness information web pages below).
[0072] The above steps yield a large number of health information web pages. Then, useless data (such as advertisements) can be filtered from the content data of these web pages to obtain the actual content. This content can then be analyzed to identify the presence of specific attribute characters. Attribute characters can be various, and identifying them determines the attribute values of the web page's attributes. For example, the attribute character for a summary is "summary," the attribute character for keywords is "keywords," and the attribute characters for cited objects include "citation," "image source," "excerpt," "from," and "source." After analyzing the web page content, attribute information for multiple web pages can be obtained.
[0073] For example, through data crawling and identification analysis, we can obtain the attribute tables of some health information web pages (Table 1).
[0074] Table 1. Characteristics of Health and Family Planning Information Webpages
[0075]
[0076] The characteristics of a website or webpage can also be improved by manually creating a database. Humans can perform operations such as data annotation, data processing, and data editing. For example, aliases for the website can be manually entered. For instance, the aliases for the website "National Health Commission" can be determined by manual input, including both "National Health Commission" and "Health Commission".
[0077] By crawling, identifying, and analyzing data from a website, and then processing it manually, we can obtain the website's attribute information.
[0078] For example, through data crawling, identification and analysis, and manual database construction, attribute tables of some websites can be obtained (Table 2).
[0079] Table 2 Website Characteristics
[0080]
[0081] Then a knowledge graph can be constructed, which can construct the web page entities corresponding to web pages. The attribute information of a web page is determined as the attribute information of its corresponding web page entity. For example, see... Figure 4 The webpage entity “COVID-19 Updates” can include four attributes: keywords, summary, attribution object, and reference object. The corresponding attribute values are “COVID-19”, “Text 1”, “Tencent.com”, and “CCTV News | National Health Commission”, respectively.
[0082] You can construct website entities corresponding to a website, for example, see [link to website entity]. Figure 5 The website entity "CCTV News" can include three attributes: alias, industry, and nature, with corresponding attribute values of "CCTV News App," "News," and "Official," respectively. Then, based on the referencing and attributing objects in the attributes of the webpage entities, a knowledge graph is constructed. The knowledge graph can include multiple entities, including website entities and webpage entities. Each entity can include multiple attributes, and each attribute corresponds to one or more attribute values. See, for example... Figure 6 The knowledge graph contains multiple entities, among which the website entity "CCTV News" includes three attributes: alias, industry, and nature, corresponding to the attribute values: CCTV News client, news, and official. The relationships related to the website entity "CCTV News" include: the webpage entity "COVID-19 Updates" has a reference relationship with the website entity "CCTV News", and the webpage entity "Summary of COVID-19 Risk Areas Nationwide" has an attribution relationship with the website entity "CCTV News".
[0083] In this application, multiple websites for constructing a knowledge graph can be identified, and then multiple internal web pages contained in each website can be identified. The attribute information of the website entity corresponding to each website and the attribute information of the web page entity corresponding to each internal web page can be obtained. Based on the attribute information of each website entity and the attribute information of each web page entity, a knowledge graph can be constructed. Then, web page content can be automatically traced based on the obtained knowledge graph, eliminating the need for users to manually search and effectively improving the efficiency of web page content tracing.
[0084] The following section will introduce the process of using knowledge graphs for web page content tracing.
[0085] In this embodiment, the webpage content tracing method will be described in detail, see [link to documentation]. Figure 7 , Figure 7 This is a flowchart illustrating the webpage content tracing method provided in this embodiment. The method may include:
[0086] 201. Receive a knowledge graph for tracing web page content. The knowledge graph includes multiple entities and the relationships between them.
[0087] Because knowledge graphs contain the relationships between web pages and websites on the Internet, computer devices can automatically trace the source of web page content through knowledge graphs, eliminating the need for manual search and querying, and effectively improving the efficiency and convenience of web page content tracing.
[0088] The internet contains a massive amount of web pages and information. Users, when acquiring information online, have a more pressing need to trace the origins of web pages and information in certain fields, such as policy and regulations, health and wellness, scientific research, and internet content copyright. On the other hand, the sheer number of web pages on the internet, coupled with the vast differences in visitor volume between different pages, means that some pages have low access value, resulting in low visitor volume. Therefore, in practice, this application can, based on actual needs, obtain a knowledge graph encompassing the relationships between several web pages / websites across several fields. This avoids unnecessary memory consumption caused by obtaining an excessively large knowledge graph, or poor results in tracing the origins of web page content due to obtaining an excessively small knowledge graph.
[0089] Specifically, there are various ways to determine a knowledge graph. For example, a knowledge graph can be constructed according to actual needs, or a knowledge graph tracing interface can be called, which corresponds to an already constructed knowledge graph, and so on.
[0090] 202. Query the webpage entity corresponding to the webpage to be traced in the knowledge graph.
[0091] The webpage to be traced can be a webpage on the Internet, such as webpage 1 containing information about the publication of paper A. Webpage entities can include entities in the knowledge graph that correspond to the webpage to be queried, such as webpage 1 (i.e., the webpage to be traced) corresponding to webpage entity 1 in the knowledge graph.
[0092] In some embodiments, to facilitate differentiation and labeling, a unique entity identifier can be assigned to each entity in the knowledge graph, and the entity identifier of each entity can be stored in the knowledge graph. The method for querying the webpage entity corresponding to the webpage to be traced in the knowledge graph includes: generating an entity identifier corresponding to the webpage to be traced based on features such as the webpage content or webpage address; querying the webpage entity corresponding to the entity identifier in the knowledge graph; and that webpage entity is the webpage entity corresponding to the webpage to be traced. For example, based on the webpage address of webpage 1 (i.e., the webpage to be traced), an entity identifier 1 corresponding to the webpage to be traced can be generated; querying the webpage entity 1 corresponding to entity identifier 1 in the knowledge graph confirms that the entity corresponding to webpage 1 is webpage entity 1.
[0093] In some embodiments, the attribute values of an entity are uniquely associated with that entity. For example, these may include the Uniform Resource Locator (URL) of a webpage entity and the ICP filing number of a website entity. These uniquely associated attribute values can be used to directly query the knowledge graph, efficiently and quickly determining the webpage entity corresponding to the webpage in the knowledge graph and the website entity corresponding to the website in the knowledge graph summary.
[0094] For example, the attribute value of the webpage address of webpage 1 (i.e. the webpage to be traced) is: URL 1. In the knowledge graph, all entities with URL addresses are identified, as well as the attribute values of these entities' URL addresses. These attribute values are compared with URL 1 in turn. When there is an attribute value that is the same as URL 1, the entity to which the attribute value belongs is determined to be webpage entity 1 corresponding to webpage 1.
[0095] 203. Identify at least one target entity in the knowledge graph corresponding to the webpage entity of the source webpage, where the target entity includes entities that are related to the webpage entity.
[0096] A knowledge graph can include relationships between entities. After determining the webpage entity corresponding to the webpage to be traced in the knowledge graph, one or more target entities corresponding to the webpage entity in the knowledge graph can be determined based on these relationships. Target entities can include entities that have relationships with webpage entities and / or have indirect relationships with network entities.
[0097] In some embodiments, the target entity may include an entity that is related to the web page entity. For example, the web page to be traced: web page 1 corresponds to web page entity 1. In the knowledge graph, web page entity 2 that is related to web page entity 1 is determined. This web page entity 2 is the target entity corresponding to web page entity 1. It can be seen that web page 1 references the content of the web page corresponding to web page entity 2.
[0098] In some embodiments, the target entity may include entities that are related to the web page entity and entities that are indirectly related to the target entity. For example, the web page to be traced: web page 1 corresponds to web page entity 1. In the knowledge graph, web page entity 2 is determined to be related to web page entity 1. This web page entity 2 is a target entity corresponding to web page entity 1. In the knowledge graph, web page entity 3 is determined to be related to web page entity 2. This web page entity 3 is another target entity corresponding to web page entity 1. The step of determining entities that are related to new target entities in the knowledge graph is repeated until no new target entity is related. Multiple target entities corresponding to the entity to be traced are obtained: web page entity 2, web page entity 3, and web page entity 4. It can be seen that web page 1 references the content of the web page corresponding to web page entity 2, the web page corresponding to web page entity 2 references the content of the web page corresponding to web page entity 3, and the web page corresponding to web page entity 3 references the content of the web page corresponding to web page entity 4.
[0099] In some embodiments, a webpage may identify multiple reference objects. For example, webpage A may identify webpage B, and webpage B may identify webpages C and D. However, in reality, the content of webpage A references webpage D, and the content of webpage A is unrelated to webpage C. Knowledge graphs can record the reference relationships between these webpage entities. However, if webpage content is traced solely based on these reference relationships, it is impossible to determine the target entity from the entities corresponding to webpage C and webpage D. To solve this problem, multiple candidate entities corresponding to webpage entities can be identified first, and then filtered using the attribute values of the webpage entities and candidate entities to determine the target entity from among the multiple candidate entities.
[0100] Candidate entities may include entities that are related to or indirectly related to web page entities.
[0101] The process of determining multiple candidate entities corresponding to a webpage entity in a knowledge graph may include: determining a candidate entity in the knowledge graph that has a relationship with the webpage entity; iterating through the steps of determining another candidate entity in the knowledge graph that has a relationship with the candidate entity until the candidate entities no longer have a relationship, at which point the loop ends, resulting in multiple candidate entities corresponding to the webpage entity.
[0102] For example, consider the following webpages to be traced: Webpage A corresponds to webpage entity A in the knowledge graph. Webpage entity B, which has a reference relationship with webpage entity A, is identified as a candidate entity. Webpage entities C and D, which also have reference relationships with webpage entity B, are identified. However, since neither webpage entity C nor webpage entity D has any reference relationships with other entities, the candidate entities corresponding to webpage entity A in the knowledge graph are: webpage entity B, webpage entity C, and webpage entity D. Then, based on the attribute information of the webpage entities and each candidate entity, at least one target entity can be determined from the candidate entities. For example, preset attributes required to determine the target entity from the candidate entities can be predetermined. Then, based on the attribute values of the preset attributes of the webpage entities and the attribute values of the preset attributes of each candidate entity, the target entity can be filtered from the candidate entities.
[0103] In some embodiments, if the preset attribute is a summary and the attribute value of the summary is a piece of text, the filtering method can be to perform semantic recognition on the attribute value of the summary attribute of the web page entity and the attribute value of the summary attribute of each candidate entity, and calculate the similarity between the semantic recognition result of the attribute value of the summary attribute of each candidate entity and the semantic recognition result of the attribute value of the summary attribute of the web page entity, and determine the candidate entity whose similarity is greater than a preset threshold as the target entity.
[0104] For example, if the default attribute is attribute 1, the target entities are selected from the three candidate entities: web entity B and web entity D based on the attribute value 1 of attribute 1 of web entity A, the attribute value 1 of attribute 1 of web entity B, the attribute value 2 of attribute 1 of web entity C, and the attribute value 3 of attribute 1 of web entity D.
[0105] In some embodiments, the referenced object indicated by a webpage may include the information source website, such as "Data source: Statistics Bureau of Province C" or "Image source: CCTV News" displayed on the webpage. The knowledge graph can record the reference relationship between the webpage entity corresponding to these webpages and the website entity corresponding to the website. However, if the specific webpage within the website referenced by the webpage cannot be determined based solely on the reference relationship, the attribution relationship between the webpage entity and the website entity in the knowledge graph, as well as the attribute values of the webpage entity's attributes, can be used to trace the source of the webpage that only indicates the information source website and determine the specific webpage it references.
[0106] For example, consider a webpage to be traced: Webpage 1, which identifies the referenced object as "Official Website A". Webpage 1 corresponds to Webpage Entity 1 in the knowledge graph. The knowledge graph then searches for candidate website entities that reference Webpage Entity 1: Website Entity A. Next, it searches for multiple candidate webpage entities that belong to Website Entity A: Webpage Entity 2, Webpage Entity 3, and Webpage Entity 4. Website Entity A is determined to be a target entity of Webpage Entity 1. Based on the attribute values of Webpage Entity 1 under preset attributes, and the attribute values of each candidate webpage entity under preset attributes, the target entity is selected from the three candidate webpage entities: Webpage Entity 2. Therefore, the content of Webpage 1 originates from the webpage corresponding to Webpage Entity 2, which belongs to the website corresponding to Website Entity A.
[0107] In some embodiments, the attribute values of preset attributes of web page entities can be compared with the attribute values of preset attributes of candidate entities. The matching criteria may include being the same, having a similarity greater than a preset threshold, having a numerical overlap rate greater than a preset value, or satisfying a preset correspondence, etc. The specific criteria can be flexibly selected in practice and will not be elaborated here.
[0108] For example, you can compare the attribute values of the candidate entity's preset attributes with the attribute values of the webpage entity. If they are the same, you can determine that the candidate entity is the target entity.
[0109] 204. Display the content tracing results of the webpage to be traced. The content tracing results are determined by at least one target entity and the relationship between the webpage entity and the target entity.
[0110] Specifically, the corresponding intermediate webpages / websites and source webpages can be determined based on the output target entity. The reference relationships or attribution relationships between the output webpage entity and the target entity, and between the target entity and the source webpage, can be determined based on the relationship between them.
[0111] When displaying the results of content tracing, you can directly display the source webpage corresponding to the webpage to be traced, or you can display the intermediate webpages / websites that are mutually referenced and belonged to during the tracing process, as well as the source webpage. You can also display the reference relationship or belonging relationship between the webpage to be traced, the intermediate webpages / websites, and the source webpage.
[0112] For example, if the target entity is webpage entity 2, determine the webpage 2 corresponding to webpage entity 2. Based on the relationship between the webpage entity 1 corresponding to the source webpage to be traced (webpage 1) and the target entity, determine the reference relationship between webpage 1 and webpage 2. Then, webpage 2 and the reference relationship between webpage 1 and webpage 2 can be displayed to the user.
[0113] For example, the first target entity is website entity A, and the second target entity is webpage entity 2. Based on the output website entity A, determine its corresponding website A; based on the output webpage entity 2, determine its corresponding webpage 2; based on the reference relationship between the output website entities: website entity 1 and website entity A, determine the reference relationship between webpage 1 and website A; based on the ownership relationship between the output website entity A and webpage entity 2, determine the ownership relationship between website A and webpage 2.
[0114] Show users the source of the webpage to be traced: the source of webpage 1 content: webpage 1, website A, webpage 2, and the reference relationship between webpage 1 and website A, and the attribution relationship between website A and webpage 2.
[0115] Content tracing results can be displayed on the page to show users the results. The page can display images of web pages, which may include part or all of the content of the web page. The web page may include the web page to be traced, the web page corresponding to the target entity, and the homepage of the website corresponding to the target entity. The page may also include reference relationship identifiers that represent the reference relationship between web page images, as well as attribution relationship identifiers.
[0116] Web page images can be used as an attribute of entities in a knowledge graph. The web page image attribute of a web page entity can be the image of the web page corresponding to that web page entity, and the web page image attribute of a website entity can be the image of the homepage of the website corresponding to that website entity. In this way, the web page images can be obtained from the knowledge graph.
[0117] A webpage image is captured from a webpage, and the webpage can be accessed via its URL. The webpage URL can be stored in the knowledge graph as the attribute value of the entity's URL attribute. For example, the first feature attribute can include the URL attribute. The attribute value of the target entity's URL attribute (the URL of the webpage corresponding to the target entity, or the URL of the homepage of the website corresponding to the target entity) can be obtained from the knowledge graph. By accessing this attribute value and capturing the webpage image, the target webpage image corresponding to the target entity can be obtained. Similarly, by accessing the URL of the webpage to be traced and capturing the webpage image, the initial webpage image of the webpage to be traced can be obtained.
[0118] This embodiment can automatically trace the source of web page content using a knowledge graph via computer devices, eliminating the need for manual searching and querying, and effectively improving the efficiency and convenience of tracing web page content.
[0119] This application can trace the source of health-related web pages through a constructed knowledge graph, either online or offline. The specific process may include:
[0120] It receives static web pages input by users and extracts the Uniform Resource Locator (URL) of the web page, or receives the URL of the web page directly input by the user, queries the knowledge graph for web page entities whose attribute value of the web page address is the URL, and then determines the web page / website entities that are related to the network entity in the knowledge graph.
[0121] For example, receiving the URL1 of a COVID-19 update webpage entered by the user, in Figure 6 The knowledge graph retrieved a web entity named "COVID-19 Updates" with the attribute value of URL1.
[0122] Search the knowledge graph for entities that have a reference relationship with the webpage entity "COVID-19 Updates". Figure 6 It can be seen that the entities that have a referencing relationship with the webpage entity "COVID-19 Updates" are the website entities "CCTV News" and "National Health Commission".
[0123] Depend on Figure 6 The knowledge graph shows that the keyword attribute value of the webpage entity "COVID-19 Updates" is "COVID-19". Therefore, all webpage entities with a relationship to the website entity "CCTV News" can be filtered to identify those with the keyword attribute value "COVID-19", namely the webpage entity "National COVID-19 Risk Areas Summary". This can be further analyzed based on... Figure 6The knowledge graph identified the website entity "State Council Client" as having a referencing relationship with the webpage entity "Summary of National Epidemic Risk Areas". All webpage entities with affiliation to the website entity "State Council Client" were then filtered to identify those with the keyword attribute value of "COVID-19", namely the webpage entity "Epidemic Risk Screening". The webpage entity "Epidemic Risk Screening" is... Figure 6 If there are no other reference relationships in the knowledge graph shown, then it can be determined that the webpage corresponding to the webpage entity "epidemic risk screening" is a source webpage.
[0124] It is possible Figure 6 The knowledge graph shown filters all web page entities that have a relationship with the website entity "National Health Commission" to identify those with the keyword attribute value of "COVID-19," namely the web page entity "Epidemic Report." The web page entity "Epidemic Report" is... Figure 6 If there are no other reference relationships in the knowledge graph, then the epidemic report webpage corresponding to the webpage entity "epidemic report" can be identified as a source webpage.
[0125] The system outputs the entities and relationships between them at each level of the tracing process, parses and renders the output information, and displays the tracing process on the page. For example, see... Figure 8 , Figure 8 The page displays the source tracing results of the COVID-19 pandemic update.
[0126] In some embodiments, the webpage content of the webpage to be traced can be identified to obtain information that can be used for tracing, such as the information of the referenced object "cited from website C", the information of the keyword "COVID-19", the information of the summary "text 1", etc. Then, the first website entity corresponding to website C can be queried in the knowledge graph, and then based on the first website entity and the information of the webpage to be traced, the webpage content of the webpage to be traced can be traced in the knowledge graph to obtain the tracing result of the webpage to be traced.
[0127] The execution process of this application can also be found in [reference needed]. Figure 9 The offline system can first crawl website content, then parse it, build a human-made knowledge base, and finally construct and save a knowledge graph. The online system can receive URLs input by users, initiate source tracing queries, and the webpage source tracing query module can use the knowledge graph to retrieve the URL's content source tracing results. The webpage source tracing display module can parse the query results and present them to the user.
[0128] In this application, when the webpage tracing module initiates a tracing query request to the knowledge graph module, the relevant code can be as follows, where webname can be the webpage name and sidename can be the website name to which the webpage name belongs.
[0129] {
[0130] "webname": "COVID-19 Updates",
[0131] "sidename": "Tencent.com",
[0132] "action": "link",
[0133] }
[0134] The code related to the process of a knowledge graph receiving a query request, querying the internal database storing the knowledge graph, and returning the query results can be as follows:
[0135] {
[0136] "webname": "COVID-19 Updates",
[0137] "sidename": "Tencent.com",
[0138] "links": [
[0139] {
[0140] "sidename": "National Health Commission",
[0141] "webname": "Epidemic Update"
[0142] },
[0143] {
[0144] "sidename": "CCTV News",
[0145] "webname": "Summary of COVID-19 Risk Areas Nationwide",
[0146] "links": [
[0147] {
[0148] "sidename": "State Council Client",
[0149] "webname": "COVID-19 Risk Inquiry"
[0150] } ]
[0152] } ]
[0154] }
[0155] As can be seen from the above, the source tracing results of the COVID-19 situation updates webpage belonging to Tencent.com can include: some webpage content of the COVID-19 situation updates webpage is quoted from the epidemic report webpage of the National Health Commission website; some webpage content of the COVID-19 situation updates webpage is quoted from the national epidemic risk area summary webpage of CCTV News website; and the webpage content of the national epidemic risk area summary webpage is quoted from the epidemic risk query webpage of the State Council client.
[0156] This embodiment can construct a knowledge graph that shows the relationship between web pages and websites, and automatically trace the source of web pages on the Internet through this knowledge graph, eliminating the need for users to manually search and trace the source, effectively improving the efficiency of web page content tracing. At the same time, it can display the intermediate websites / web pages between the web page to be traced and the source web page, making the entire tracing process clear and easy to understand.
[0157] refer to Figure 10 This is a schematic diagram of the hardware structure of the computer device 100 provided in an embodiment of this application. Figure 10 As shown, the computer device 100 may include a processor 1001, a memory 1002, a communication bus 1003, and a display screen 1004. The memory 1002 is used to store one or more computer programs 1005. The one or more computer programs 1005 are configured to be executed by the processor 1001. The one or more computer programs 1005 may include instructions that can be used to implement the aforementioned web page content tracing method and / or knowledge graph construction method in the computer device 100.
[0158] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the computer device 100. In other embodiments, the computer device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements.
[0159] Processor 1001 may include one or more processing units, such as application processors (APs), graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, DSPs, CPUs, baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.
[0160] The processor 1001 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 1001 is a cache memory. This memory can store instructions or data that the processor 1001 has just used or that are used repeatedly. If the processor 1001 needs to use the instruction or data again, it can retrieve it directly from this memory. This avoids repeated accesses, reduces the waiting time of the processor 1001, and thus improves the efficiency of the system.
[0161] In some embodiments, the processor 1001 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM interface, and / or a USB interface, etc.
[0162] In some embodiments, memory 1002 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0163] This embodiment also provides a computer storage medium storing computer instructions. When the computer instructions are executed on a computer device, the computer device performs the above-mentioned related method steps to implement the web page content tracing method and / or knowledge graph construction method in the above embodiment.
[0164] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement the web page content tracing method and / or knowledge graph construction method in the above embodiments.
[0165] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component, or module. The apparatus may include a connected processor and a memory; wherein the memory is used to store computer execution instructions. When the apparatus is running, the processor may execute the computer execution instructions stored in the memory to cause the chip to execute the web page content tracing method and / or knowledge graph construction method in the above-described method embodiments.
[0166] In this embodiment, the first computer device, computer storage medium, computer program product, or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0167] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0168] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0169] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0170] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0171] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0172] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered within the scope of protection of this application.
Claims
1. A method for tracing the source of web page content, applied to a server, characterized in that, The method includes: The process involves querying the first webpage entity corresponding to the webpage to be traced in a knowledge graph. The knowledge graph includes multiple entities and relationships between them. The multiple entities include at least one website entity and at least one webpage entity. The relationships between the entities include referencing relationships and attribution relationships. The referencing relationship or the attribution relationship is determined through the relationship attributes of the website entity or the webpage entity. Querying the first webpage entity corresponding to the webpage to be traced in the knowledge graph includes: generating a webpage identifier corresponding to the webpage to be traced based on its webpage address; and determining the first webpage entity corresponding to the webpage to be traced in the knowledge graph based on the webpage identifier corresponding to the webpage to be traced and the webpage identifier attributes of all webpage entities in the knowledge graph. Based on the knowledge graph and the first web page entity, at least one target entity is determined, and the at least one target entity has a direct or indirect relationship with the first web page entity. The tracing result of the webpage to be traced is determined, and the tracing result includes at least one webpage or website corresponding to the at least one target entity and the relationship between each webpage or website and webpages or websites corresponding to other target entities.
2. The method according to claim 1, characterized in that, The webpage entity also includes a webpage address attribute, and the first webpage entity corresponding to the webpage to be traced in the knowledge graph includes: Based on the webpage address of the source webpage and the webpage address attributes of all webpage entities in the knowledge graph, the first webpage entity corresponding to the webpage to be traced in the knowledge graph is determined.
3. The method according to claim 1, characterized in that, The step of determining at least one target entity based on the knowledge graph and the first webpage entity includes: Based on the knowledge graph and the first webpage entity, at least one candidate entity is determined; Based on the preset attributes of each candidate entity and the preset attributes of the first webpage entity, at least one target entity is determined from the at least one candidate entity.
4. The method according to claim 3, characterized in that, Before querying the first webpage entity corresponding to the webpage to be traced in the knowledge graph, the method further includes: Obtain the knowledge graph.
5. The method according to claim 4, characterized in that, The method further includes: The tracing result is sent to the terminal, causing the terminal to render a user interface based on the tracing result. The user interface includes an image of the webpage to be traced, an image of the website or webpage corresponding to the at least one target entity, and a relationship identifier between the image of the webpage to be traced and the image of the website or webpage corresponding to the at least one target entity. The relationship identifier is determined based on the relationship between the first webpage entity and the at least one target entity.
6. A method for tracing the source of web page content, applied to a terminal, characterized in that, The method includes: Based on the webpage address of the webpage to be traced input by the user, a traceability request is generated for the webpage to be traced. Send the tracing request to the server so that the server executes the web page content tracing method as described in any one of claims 1 to 5, so as to determine the tracing result of the web page to be traced in the knowledge graph based on the web page address contained in the tracing request; The system receives the tracing result returned by the server and displays an image of the webpage to be traced, as well as an image of the webpage or website referenced by the webpage to be traced, on the user interface based on the tracing result.
7. A computer device, characterized in that, The computer device includes at least one processor, memory, and communication module; The at least one processor is connected to the memory and the communication module; The memory is used to store instructions, the processor is used to execute the instructions, and the communication module is used to communicate with the device under the control of the at least one processor; When the instruction is executed by the at least one processor, it causes the at least one processor to perform the web page content tracing method as described in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that causes a computer device to perform the web page content tracing method as described in any one of claims 1 to 6.
9. A computer program product, characterized in that, The computer program product includes computer execution instructions stored in a computer-readable storage medium; at least one processor of the computer device can read the computer execution instructions from the computer-readable storage medium, and the at least one processor executes the computer execution instructions to cause the computer device to perform the web page content tracing method as described in any one of claims 1 to 6.