Encyclopedia corpus extraction method and device, electronic equipment, chip and medium

Through document object model tree parsing and structured data processing of encyclopedia corpora, the problem of inaccurate encyclopedia corpus acquisition is solved, high-quality corpus acquisition is achieved, and technical applications are applied in multiple fields.

CN120705326APending Publication Date: 2025-09-26BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410354815.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-26
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In the existing technology, the acquisition of encyclopedia corpus is inaccurate and of low quality, especially in the processing of unstructured and structured data, which makes it difficult to extract high-quality encyclopedia corpus.

Method used

By parsing the document object model tree based on encyclopedia resource web pages, we can filter the text content of unstructured data and construct key-value pairs of structured data, combine triples to build a knowledge graph, and improve data accuracy and quality.

Benefits of technology

It improves the accuracy and quality of encyclopedia corpus acquisition and is applicable to multiple fields such as natural language processing, machine translation, question-answering systems, knowledge graphs, educational technology, enterprise knowledge management, content creation, dialogue systems, healthcare, law and government services, financial analysis, and humanities and social sciences research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705326A_ABST
    Figure CN120705326A_ABST
Patent Text Reader

Abstract

The invention provides an encyclopedia corpus extraction method and device, electronic equipment, a chip and a medium, and relates to the technical field of knowledge extraction and representation. The encyclopedia corpus extraction method comprises the following steps: determining a target label through a target webpage source code; obtaining a corpus type of character information corresponding to the target tag, wherein the corpus type comprises unstructured data and structured data; if the corpus type is unstructured data, screening text content in the character information as a text of the unstructured data; and if the corpus type is structured data, extracting character information and encyclopedia information associated with the character information, and constructing a key value pair as the content of the structured data. Through the technical scheme provided by the invention, the problems of inaccuracy and low quality of the acquired encyclopedia corpus in the related technology are solved, and the accuracy and the quality of acquiring the encyclopedia corpus are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of knowledge extraction and representation, and in particular to an encyclopedia corpus extraction method, device, electronic device, chip, and medium. Background Art

[0002] With the rapid development of large language models (LLMs), the demand for high-quality encyclopedia corpora is growing. Large language models such as Generative Pre-trained Transformer 3 (GPT-3) and Bidirectional Encoder Representations from Transformers (BERT) have demonstrated outstanding performance in multiple fields such as text generation, natural language understanding, and machine translation. However, training these models requires a large amount of accurate, diverse, and high-quality data. As a knowledge-intensive resource, encyclopedia corpora have become one of the important sources of training data.

[0003] Encyclopedia corpora, such as Wikipedia, are collaboratively edited platforms that aggregate knowledge and information from users around the world. These corpora typically encompass a rich range of topics, covering a broad range of general and specialized knowledge. Due to their open nature, encyclopedia corpora are dynamically updated, quickly reflecting the latest state of knowledge.

[0004] However, it is difficult to obtain accurate and high-quality encyclopedia corpus information by directly extracting encyclopedia corpus from relevant encyclopedia corpora. Summary of the Invention

[0005] The present disclosure provides an encyclopedia corpus extraction method, device, electronic device, chip, and medium to address the inaccurate and low-quality encyclopedia corpus obtained in related technologies. By performing targeted parsing of the Document Object Model (DOM) tree of the encyclopedia resource webpage, the method extracts large sections of text as unstructured data and extracts attribute text of target objects corresponding to tags as key-value pairs, which serve as structured data. This method yields highly accurate and high-quality encyclopedia corpus information.

[0006] A first embodiment of the present disclosure provides an encyclopedia corpus extraction method, the method comprising:

[0007] Determine the target tag through the target web page source code;

[0008] Get the corpus type of the character information corresponding to the target label, which includes unstructured data and structured data;

[0009] If the corpus type is unstructured data, the text content in the character information is filtered as the text of the unstructured data;

[0010] If the corpus type is structured data, character information and encyclopedia information associated with the character information are extracted, and key-value pairs are constructed as the content of the structured data.

[0011] In one embodiment of the present disclosure, after extracting character information and encyclopedia information associated with the character information and constructing key-value pairs as structured data, the method further includes:

[0012] Convert structured data into triples to build a knowledge graph.

[0013] In one embodiment of the present disclosure, if the corpus type is unstructured data, filtering the text content in the character information as the text of the unstructured data includes:

[0014] Extracting the directory hierarchy structure of the document object model tree corresponding to the target web page source code;

[0015] Determining the density of text content in each directory of the directory hierarchy;

[0016] The text content with a density greater than a preset threshold is regarded as the text of unstructured data.

[0017] In one embodiment of the present disclosure, determining the density of text content in each directory in the directory hierarchy includes:

[0018] Based on the data type in the text content, obtain the weight factor corresponding to the data type;

[0019] Using the weight factor and the proportion of text of the data type in the text content, the density is obtained.

[0020] In one embodiment of the present disclosure, density is obtained using a weight factor and a ratio of text of a data type in text content, including:

[0021] If the data type is a string type, the density is obtained by multiplying the first weight factor by the proportion of text of the string type in the text content, where the first weight factor corresponds to the string type.

[0022] If the data type is a numerical type, the density is obtained by multiplying the proportion of numerical type text in the text content by a second weight factor, the second weight factor corresponds to the numerical type, and the first weight factor is different from the second weight factor.

[0023] In one embodiment of the present disclosure, determining a target tag through the source code of a target webpage includes:

[0024] Get the source code of the target web page;

[0025] Build a document object model tree based on the target web page source code;

[0026] Remove redundant tags in the document object model tree to preserve the target tags.

[0027] In one embodiment of the present disclosure, the method further includes:

[0028] Convert the target web page source code into a lightweight markup language text format using text formatting tools.

[0029] A second embodiment of the present disclosure provides an encyclopedia corpus extraction device, comprising:

[0030] A determination module is used to determine the target tag through the source code of the target web page;

[0031] An acquisition module is used to obtain the corpus type of the character information corresponding to the target tag, and the corpus type includes unstructured data and structured data;

[0032] A screening module is used to screen the text content in the character information as the text of the unstructured data if the corpus type is unstructured data;

[0033] The construction module is used to extract character information and encyclopedia information associated with the character information if the corpus type is structured data, and to construct key-value pairs as the content of the structured data.

[0034] The third aspect embodiment of the present disclosure proposes an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any one of the methods in the first aspect embodiment of the present disclosure.

[0035] The fourth aspect embodiment of the present disclosure proposes a non-transitory computer-readable storage medium storing computer instructions, characterized in that the computer instructions are used to enable a computer to execute the method in the first aspect embodiment of the present disclosure.

[0036] The fifth aspect of the present disclosure provides a computer program product, characterized in that it includes a computer program, and when the computer program is executed by a processor, it implements any one of the methods in the first aspect of the present disclosure.

[0037] The sixth aspect embodiment of the present disclosure proposes a chip, comprising at least one processor and a communication interface; the communication interface is used to receive signals input into the chip or signals output from the chip, the processor communicates with the communication interface and implements any one of the methods in the first aspect embodiment of the present disclosure through logic circuits or executing code instructions.

[0038] In summary, according to the encyclopedia corpus extraction method proposed in the present disclosure, the target tag is determined through the target web page source code, providing a data source for encyclopedia corpus extraction; the corpus type of the character information corresponding to the target tag is obtained, and the corpus type includes unstructured data and structured data. Different content extraction strategies are determined according to different types; if the corpus type is unstructured data, the text content in the character information is filtered as the text of the unstructured data, providing unstructured data for the encyclopedia corpus; if the corpus type is structured data, the character information and the encyclopedia information associated with the character information are extracted, and key-value pairs are constructed as the content of the structured data, providing the encyclopedia corpus with the content of the structured data. By extracting unstructured data and structured data by type, the accuracy of encyclopedia corpus acquisition is improved compared to the related art of obtaining encyclopedia corpus without distinguishing between types, especially for the acquisition of objective data such as structured data, the quality of encyclopedia corpus acquisition is improved.

[0039] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0041] Figure 1 This is a flow chart of an encyclopedia corpus extraction method according to an embodiment of the present disclosure;

[0042] Figure 2 A flowchart for constructing a knowledge graph according to an embodiment of the present disclosure;

[0043] Figure 3 A flowchart of filtering text content in character information as unstructured data according to an embodiment of the present disclosure;

[0044] Figure 4 This is a flow chart of determining the density of text content in each directory in a directory hierarchy according to an embodiment of the present disclosure;

[0045] Figure 5 A flow chart of obtaining density using a weight factor and a ratio of text of a data type in text content according to an embodiment of the present disclosure;

[0046] Figure 6 This is a flow chart of determining a target tag through the source code of a target web page according to an embodiment of the present disclosure;

[0047] Figure 7 This is a flow chart of a corpus extraction method according to an embodiment of the present disclosure;

[0048] Figure 8 This is a flow chart of an encyclopedia corpus extraction method according to an embodiment of the present disclosure;

[0049] Figure 9 This is a flow chart of an encyclopedia corpus extraction method according to an embodiment of the present disclosure;

[0050] Figure 10 Schematic diagram of the structure of an encyclopedia corpus extraction device according to an embodiment of the present disclosure;

[0051] Figure 11 is a block diagram of an electronic device for implementing the encyclopedia corpus extraction method disclosed herein, according to an exemplary embodiment;

[0052] Figure 12 It is a schematic structural diagram of a chip according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0053] The following describes in detail embodiments of the present disclosure, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout identify the same or similar components or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.

[0054] First, a brief introduction to the relevant terms in this disclosure is given:

[0055] Structured data: Data that is structured according to a predefined model or organized in a predefined way.

[0056] Unstructured data: Data that is neither structured according to a predefined data model nor organized in a predefined manner.

[0057] Regular expression: Use a simple string to describe and match all strings in the text that match the specified format.

[0058] Knowledge Graph: A knowledge graph is a structured semantic knowledge base used to quickly describe concepts and their relationships in the physical world. By effectively processing, handling, and integrating complex document data, it is transformed into simple, clear (entity, relationship, entity) and (entity, attribute, attribute value) triples, and finally aggregates a large amount of knowledge, thereby achieving rapid response and reasoning of knowledge.

[0059] DOM tree: Document Object Model tree. Through the DOM tree, you can not only intuitively see the overall structure of HTML, but also use some of the DOM tree's attributes to obtain information such as the child nodes and node names of an element.

[0060] In the era of large models, the accuracy and structure of training corpora are crucial. Currently, the most authoritative corpora for Chinese encyclopedia knowledge are Baidu Encyclopedia and Wikipedia. However, both are raw corpora that require multiple layers of cleaning to achieve high-quality results. Existing processing tools in the industry still have certain flaws. They extract unstructured information from Baidu Encyclopedia incompletely and contain a lot of redundant information. For structured information, they lack information on character relationships. For example, if character A's brother is character E, this data doesn't exist in Baidu Encyclopedia's existing structured information, requiring additional extraction of this area. For Chinese Wikipedia, all entries are officially packaged for download each month, resulting in a complex unstructured text format. Existing tools primarily rely on Python's Wikipedia Extractor and Gensim's WikiCorpus library. However, Wikipedia Extractor removes content marked with {{}}, and Gensim's WikiCorpus library presents even more serious issues, removing all punctuation. Furthermore, structured data on Chinese Wikipedia, such as the information box data on the right side of the page, is difficult to parse due to Wikipedia's unique parsing logic. For example, if person C's spouse is person D, the stored information is "spouse = {{marriage|[[person D]]|2015}}." This requires not only mapping spouse to spouse but also parsing {{marriage|[[person D]]|2015}} into (person D, 2015). With hundreds of such key-value pairs, processing is extremely challenging. This results in inaccurate and low-quality encyclopedia data.

[0061] The method proposed in this disclosure is applied to encyclopedia corpus extraction tasks, with a wide range of applications. It can clean and standardize existing online encyclopedia information to produce highly accurate and high-quality encyclopedia corpus. This method can be applied in natural language processing, for text classification, such as news classification and sentiment analysis; machine translation, improving translation accuracy and fluency; question-answering systems, providing more accurate answers and explanations; entity recognition and relationship extraction, better understanding entities and their relationships within text; and semantic search, providing more relevant search results. In the knowledge graph field, high-quality encyclopedia data can be used to construct and enrich knowledge graphs, providing support for search engines, recommendation systems, and more. In educational technology, it can be used to develop intelligent teaching assistance systems, such as AI-based tutoring robots. It can also be used to create interactive learning resources, such as intelligent exercise generators and personalized learning path recommendations. In enterprise knowledge management, it can be used to automatically categorize and organize internal knowledge. It can also be used for competitive intelligence analysis, extracting key information from large amounts of text. In content creation, it can be used to assist with writing, such as providing background information, drafting, or proofreading. It can also be used to automatically generate news articles and reports. In the field of dialogue systems, it is used to build more accurate and humane chatbots, and provide consistent and accurate information responses in customer service systems. In the field of medical health, it is used to extract and analyze medical literature to assist doctors in making diagnosis and treatment decisions. It is used to build a health knowledge base to provide users with reliable medical and health information. In the field of law and government services, it can be used to automatically analyze legal documents and provide legal consulting services. It provides public opinion analysis and policy-making support to government agencies. In the field of financial analysis, it analyzes financial reports and bulletins to predict market trends. Risk assessment and investment advice. In the field of humanities and social sciences research, it can be used for literary analysis, historical research, and data-driven research in the social sciences. In these applications, high-quality and high-accuracy Chinese encyclopedia corpus data can help improve system performance, user experience, and the reliability of the final results. There is no limitation on the application scenarios in the embodiments disclosed herein.

[0062] The encyclopedia corpus extraction method provided by the present disclosure is described in detail below with reference to the accompanying drawings.

[0063] Figure 1 This is a flow chart of an encyclopedia corpus extraction method according to an embodiment of the present disclosure. This method can be executed on a server, a mobile terminal, a personal computer, or other devices, such as Figure 1 In the embodiment shown, the encyclopedia corpus extraction method includes:

[0064] Step 101: determine the target tag through the source code of the target web page.

[0065] In this embodiment, the target web page source code refers to the source code of the web page of the encyclopedia website, such as the Hypertext Markup Language (HTML) code. The target tag refers to the different entry contents of the encyclopedia entry, which is located in the target web page source code. For example, in Baidu Encyclopedia, searching for "Person A" can obtain the encyclopedia information of Person A. The homepage of the encyclopedia information contains the encyclopedia directory and basic information about Person A. The encyclopedia directory contains information such as major events, early life, acting experience, personal life, major works, social activities, etc. The source code can be obtained on this page through a crawler tool. "Major works", "acting experience", etc. in the source code can be used as target tags.

[0066] Step 102: Obtain the corpus type of the character information corresponding to the target tag, where the corpus type includes unstructured data and structured data.

[0067] In this embodiment, character information refers to a string sequence consisting of letters and words. After determining the target tag of the target object in the source code of the target web page, the corpus type of the character information corresponding to the target tag is obtained, and the corpus type includes unstructured data and structured data. For example, the target tag is "Earth", and the character information can be "age" or "age". Optionally, unstructured data is usually a subjective description of a certain thing, and can have a variety of descriptive contents, such as character evaluation, music appreciation, movie review, etc., while structured data is usually objective data, such as objective information represented by attributes such as character age, gender, height, etc.

[0068] Step 103: If the corpus type is unstructured data, the text content in the character information is filtered as the text of the unstructured data.

[0069] In this embodiment, text content refers to the filtering of text content within character information if the corpus type is unstructured data, and this text content is treated as text for the unstructured data. For example, after obtaining the source code for the encyclopedia information about person A, the target tag "Basic Information" is not immutable objective data belonging to the target object "Person A" and can be treated as unstructured data. The text content within the character information within the target tag "Basic Information" is filtered and extracted as text for the unstructured data.

[0070] Step 104: If the corpus type is structured data, character information and encyclopedia information associated with the character information are extracted, and key-value pairs are constructed as the content of the structured data.

[0071] In this embodiment, encyclopedia information refers to a key-value pair, which is a data structure consisting of a key and a value. If the corpus type is structured data, the character information corresponding to the target tag of the target object and the encyclopedia information related to the character information are extracted to construct a key-value pair as the content of the structured data. For example, the target tag of the target object "Person A" is "wife", and the character information corresponding to the target tag is "Person B". In this case, "wife" is used as the key and "Person B" is used as the value. This key-value pair constitutes a key-value pair and can be used as the content of the structured data.

[0072] In summary, the encyclopedia corpus extraction method proposed in this disclosure determines the target tag through the target webpage source code, providing a data source for encyclopedia corpus extraction. The corpus type of the character information corresponding to the target tag is obtained, and the corpus types include unstructured data and structured data. Different content extraction strategies are determined based on the different types. If the corpus type is unstructured data, the text content in the character information is filtered as the text of the unstructured data, providing unstructured data for the encyclopedia corpus. If the corpus type is structured data, the character information and the encyclopedia information associated with the character information are extracted, and key-value pairs are constructed as the content of the structured data, providing the encyclopedia corpus with structured data content. This improves the accuracy and quality of encyclopedia corpus acquisition.

[0073] Figure 2 A flowchart for constructing a knowledge graph according to an embodiment of the present disclosure. Figure 2 Yes Figure 1 The instructions after step 104 are based on Figure 2 The embodiment shown includes the following steps:

[0074] Step 201: Convert structured data into triples to construct a knowledge graph.

[0075] In this embodiment, a triple refers to a data structure containing three elements. In the field of natural language processing, a triple usually refers to an SPO triple, namely Subject, Predicate, and Object. This is an information representation method used to describe the relationship or attributes between entities. For example, a triple describing a person's birthday may be (name, birthday, date). For example, the wife of "Person A" is Person B. These data are processed into triples such as "Person A--Acting Experience--XXX" and "Person A-Wife-Person B" to construct a knowledge graph. After obtaining the structured data, it is converted into triples to construct a knowledge graph to obtain high-quality corpus data. For the objective facts or attributes of the target topic, accurate output can be obtained through a large language model.

[0076] Figure 3The present invention is a flowchart of an embodiment of filtering text content in character information as unstructured data. Figure 3 Yes Figure 1 Further explanation of step 103 is based on Figure 3 The embodiment shown includes the following steps:

[0077] Step 301: extract the directory hierarchy structure of the document object model tree corresponding to the target web page source code.

[0078] In this embodiment, the directory hierarchy of the target web page is analyzed based on the DOM tree corresponding to the source code of the target web page. Optionally, the DOM tree of the page is directly accessed and manipulated through the DOM application program interface (API). Optionally, the DOM tree structure and directory hierarchy of the web page are displayed using the browser's element inspector. Optionally, the DOM tree structure and directory hierarchy of the web page are obtained based on keywords of the directory hierarchy using regular expressions.

[0079] Step 302: Determine the density of text content in each directory in the directory hierarchy.

[0080] In this embodiment, within a determined directory hierarchy, the density of text content within each directory is examined. For example, within a webpage structure, a directory may contain a large amount of text content, while the source code portion of the webpage structure is relatively small. This directory may contain a high density of text content. Alternatively, the total number of characters within the directory may be used as the denominator, and the number of text characters within the directory may be used as the numerator to determine the proportion of text content within the directory, i.e., its density.

[0081] Step 303: The text content with a density greater than a preset threshold is regarded as text of unstructured data.

[0082] In this embodiment, the preset threshold is a percentage value used to represent the density. After the density of the text content in the directory is determined, it is compared with the preset threshold. If the density is greater than the preset threshold, it indicates that the directory contains a large amount of text content rather than the source code of the web page structure. This text content can be treated as unstructured data, effectively alleviating the problem of a large amount of redundant information in the directory.

[0083] In one implementation of this embodiment, when extracting the contents of each directory, the text density is calculated, that is, the proportion of pure text in a tag is calculated, for example Character A is a famous singer and actor This text density is higher, and \\xe0 The density is lower. When calculating, we cannot only look at the Chinese character format, because it is also related to the sentence length. Normalization is required, and some rules need to be added to the digital types to give them weights. Finally, a threshold is set. If the density of the text block is greater than the preset threshold, the text is extracted as the final unstructured data content.

[0084] In this embodiment, the text content in the character information is filtered as the text of the unstructured data, redundant data is eliminated, and unstructured data is provided for the encyclopedia corpus.

[0085] Figure 4 This is a flowchart of determining the density of text content in each directory in a directory hierarchy structure according to an embodiment of the present disclosure. Figure 4 Yes Figure 3 The specific description of step 302 is based on Figure 4 The embodiment shown includes the following steps:

[0086] Step 401: Based on the data type in the text content, obtain a weight factor corresponding to the data type.

[0087] In this embodiment, the data type of the text content is determined. If the text content is a string type, a weight factor is set for it. If the text content is a numeric type, a different weight factor is set for it. This is used to express the importance of different types of text content.

[0088] Step 402 : Obtain density using the weight factor and the proportion of text of the data type in the text content.

[0089] In this embodiment, the density of the text content in the directory is calculated by multiplying the ratio of the text content in the directory and the weight factor corresponding to the data type of the text content.

[0090] In this embodiment, different weight factors are assigned to different text content types according to different importance levels, thereby determining the density of text content in each directory in the directory hierarchy and obtaining a balanced and comparable density value of the text content.

[0091] Figure 5 This is a flowchart of an embodiment of the present disclosure for obtaining density using a weight factor and a ratio of text of a data type in text content. Figure 5 Yes Figure 4 The specific description of step 402 is based on Figure 5 The embodiment shown includes the following steps:

[0092] Step 501: If the data type is a string type, density is obtained by multiplying a first weight factor by the proportion of text of the string type in the text content, where the first weight factor corresponds to the string type.

[0093] In this embodiment, the first weight factor is the weight value of the character information of the string type text in the directory. If the data type of the text content in the directory is a string type, the density is obtained by multiplying the first weight factor by the proportion of the string type text in the text content.

[0094] Step 502: If the data type is a numerical type, the density is obtained by multiplying the proportion of numerical type text in the text content by a second weight factor, where the second weight factor corresponds to the numerical type and the first weight factor is different from the second weight factor.

[0095] In this embodiment, the second weight factor is the weight value of the character information of the text of the numerical type in the directory. If the data type is a numerical type, the density is obtained by multiplying the second weight factor by the proportion of the text of the numerical type in the text content.

[0096] In this embodiment, normalization processing is performed on different types of text content, and density is obtained using weight factors and the proportion of text of different data types in the text content, so as to obtain higher quality encyclopedia corpus.

[0097] Figure 6 This is a flow chart of determining a target tag through the source code of a target web page according to an embodiment of the present disclosure. Figure 6 Yes Figure 1 The specific description of step 101 is based on Figure 1 The embodiment shown includes the following steps:

[0098] Step 601: Obtain the source code of the target web page.

[0099] In this embodiment, the target web page source code is obtained from the target object's encyclopedia page through a crawler tool or a browser's page source code conversion tool.

[0100] Step 602: construct a document object model tree based on the target web page source code.

[0101] In this embodiment, after the target webpage source code is obtained, a document object model tree is constructed using the hierarchical structure of the target webpage source code.

[0102] Step 603: Delete redundant tags in the document object model tree to retain the target tag.

[0103] In this embodiment, repeated and redundant tags in the document object model tree are deleted, thereby retaining the target tag. For example, messy punctuation is cleaned and only a single repeated tag is retained.

[0104] In this embodiment, the target tag is determined by the target web page source code, the web page source data of the encyclopedia corpus is cleaned, more valuable text content is obtained, and more accurate text content is retained as the encyclopedia corpus.

[0105] Figure 7 The present invention provides a flowchart of a corpus extraction method according to an embodiment of the present invention. Figure 7 Yes Figure 1 Further explanation is based on Figure 7 The embodiment shown includes the following steps:

[0106] Step 701: Convert the target webpage source code into a lightweight markup language text format using a text formatting tool.

[0107] In this embodiment, for corpus extraction, the target web page source code can also be converted into a lightweight markup language text format, for example, a complex table html source code is converted into a standard markdown format, so as to parse the text content in the table. Optionally, text formatting tools include pandoc and Turndown. For example, for the extraction of encyclopedia corpus, in terms of Wikipedia, many regular expressions have been added by observing and analyzing data to extract unstructured information that is not extracted by commonly used tools in the industry, and messy punctuation has been cleaned. At the same time, a new method of crawling Wikipedia pages to obtain structured information has been proposed. The two paths are combined to sort out the Wikipedia corpus. In addition, the existing encyclopedia-related tools do not handle tables well. Using table parsing tools, complex table html source codes can be converted into standard markdown format. This helps to extract high-quality encyclopedia corpus.

[0108] Figure 8 Flowchart of an encyclopedia corpus extraction method according to an embodiment of the present disclosure. Figure 8 As shown in the middle left figure, in Baidu Encyclopedia, we crawl the HTML page of Baidu Encyclopedia through the web page, and use the parsing tool to parse the unstructured data and structured data. Based on Baidu Encyclopedia, we crawl the Baidu Encyclopedia character relationship page using specific character relationships, and use the parsing tool to add related information to the structured data. Figure 8 As shown in the middle right image, on the Chinese Wikipedia, we crawled the HTML pages of the Chinese Wikipedia website and parsed them into structured data using the infobox parsing tool. For the Chinese Wikipedia, we can also download and decompress the source code and use the parsing tool to obtain unstructured data.

[0109] Figure 9 Flowchart of an encyclopedia corpus extraction method according to an embodiment of the present disclosure. Figure 9 As shown in the figure, for encyclopedia pages, the HTML source code of the web page is obtained through web page preprocessing. A DOM tree is constructed based on the HTML source code, and redundant tags in the DOM tree are deleted. Based on the cleaned web page source code, such as those that have been deleting redundant tags, the directory structure of the DOM tree is extracted, the text content density of each directory is calculated, and a threshold is set to filter the text content. Text content that meets the threshold is treated as unstructured data. In addition, based on the cleaned web page source code, such as those that have been deleting redundant tags, key-value pair information is extracted and directly used as structured data. This results in highly accurate and high-quality encyclopedia corpus.

[0110] The present disclosure provides an encyclopedia corpus extraction method. The method determines the target tag through the target webpage source code, providing a data source for encyclopedia corpus extraction. The method obtains the corpus type of the character information corresponding to the target tag, which includes unstructured data and structured data. Different content extraction strategies are determined based on the different types. If the corpus type is unstructured data, the text content in the character information is filtered as the text of the unstructured data, providing the encyclopedia corpus with unstructured data. If the corpus type is structured data, the character information and the encyclopedia information associated with the character information are extracted, and key-value pairs are constructed as the content of the structured data, providing the encyclopedia corpus with the content of the structured data. This method improves the accuracy and quality of encyclopedia corpus acquisition.

[0111] Corresponding to the methods provided in the above-mentioned embodiments, the present disclosure also provides an encyclopedia corpus extraction device. Since the device provided in the embodiment of the present disclosure corresponds to the methods provided in the above-mentioned embodiments, the implementation method is also applicable to the device provided in this embodiment and will not be described in detail in this embodiment.

[0112] Figure 10 FIG. 1 is a schematic diagram of the structure of an encyclopedia corpus extraction device 1000 according to an embodiment of the present disclosure. Figure 10 As shown, the encyclopedia corpus extraction device includes:

[0113] A determination module 1010 is configured to determine a target tag through the source code of a target web page;

[0114] An acquisition module 1020 is used to acquire the corpus type of the character information corresponding to the target tag, where the corpus type includes unstructured data and structured data;

[0115] A screening module 1030 is configured to screen text content in the character information as unstructured data if the corpus type is unstructured data;

[0116] The construction module 1040 is used to extract character information and encyclopedia information associated with the character information, and construct key-value pairs as the content of the structured data if the corpus type is structured data.

[0117] In some embodiments, after extracting character information and encyclopedia information associated with the character information and constructing key-value pairs as structured data, the construction module 1040 is further configured to:

[0118] Convert structured data into triples to build a knowledge graph.

[0119] In some embodiments, the screening module 1030 is configured to:

[0120] Extracting the directory hierarchy structure of the document object model tree corresponding to the target web page source code;

[0121] Determining the density of text content in each directory of the directory hierarchy;

[0122] The text content with a density greater than a preset threshold is regarded as the text of unstructured data.

[0123] In some embodiments, the screening module 1030 determines the density of text content in each directory in the directory hierarchy in the following manner:

[0124] Based on the data type in the text content, obtain the weight factor corresponding to the data type;

[0125] Using the weight factor and the proportion of text of the data type in the text content, the density is obtained.

[0126] In some embodiments, the screening module 1030 uses the weight factor and the ratio of text of the data type in the text content to obtain the density in the following manner:

[0127] If the data type is a string type, the density is obtained by multiplying the first weight factor by the proportion of text of the string type in the text content, where the first weight factor corresponds to the string type.

[0128] If the data type is a numerical type, the density is obtained by multiplying the proportion of numerical type text in the text content by a second weight factor, the second weight factor corresponds to the numerical type, and the first weight factor is different from the second weight factor.

[0129] In some embodiments, the determination module 1010 is configured to:

[0130] Get the source code of the target web page;

[0131] Build a document object model tree based on the target web page source code;

[0132] Remove redundant tags in the document object model tree to preserve the target tags.

[0133] In some embodiments, the encyclopedia corpus extraction device is further configured to:

[0134] Convert the target web page source code into a lightweight markup language text format using text formatting tools.

[0135] In summary, the encyclopedia corpus extraction device determines the target tag from the target webpage source code; obtains the corpus type of the character information corresponding to the target tag, which can be unstructured data or structured data; if the corpus type is unstructured data, filters the text content in the character information as the text of the unstructured data; if the corpus type is structured data, extracts the character information and the encyclopedia information associated with the character information, and constructs key-value pairs as the content of the structured data. This device solves the problem of inaccurate and low-quality encyclopedia corpus obtained in related technologies, improving the accuracy and quality of encyclopedia corpus acquisition.

[0136] The embodiments provided above in this disclosure describe the methods and devices provided in these embodiments. To implement the various functions in the methods provided in these embodiments, electronic devices may include hardware structures and software modules, and implement these functions in the form of hardware structures, software modules, or a combination of hardware structures and software modules. Certain of these functions may be implemented in the form of hardware structures, software modules, or a combination of hardware structures and software modules.

[0137] Figure 11 1 is a block diagram of an electronic device 1100 for implementing the above encyclopedia corpus extraction method according to an exemplary embodiment.

[0138] For example, the electronic device 1100 may be a mobile phone, a computer, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, and the like.

[0139] Reference Figure 11 , the electronic device 1100 may include one or more of the following components: a processing component 1102 , a memory 1104 , a power component 1106 , a multimedia component 1108 , an audio component 1110 , an input / output (I / O) interface 1112 , a sensor component 1114 , and a communication component 1116 .

[0140] The processing component 1102 generally controls the overall operation of the electronic device 1100, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 1102 may include one or more processors 1120 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 1102 may include one or more modules to facilitate interaction between the processing component 1102 and other components. For example, the processing component 1102 may include a multimedia module to facilitate interaction between the multimedia component 1108 and the processing component 1102.

[0141] The memory 1104 is configured to store various types of data to support operations on the electronic device 1100. Examples of such data include instructions for any application or method operating on the electronic device 1100, contact data, phone book data, messages, pictures, videos, etc. The memory 1104 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0142] The power supply component 1106 provides power to the various components of the electronic device 1100. The power supply component 1106 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 1100.

[0143] The multimedia component 1108 includes a screen that provides an output interface between the electronic device 1100 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 1108 includes a front camera and / or a rear camera. When the electronic device 1100 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0144] The audio component 1110 is configured to output and / or input audio signals. For example, the audio component 1110 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 1100 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 1104 or transmitted via the communication component 1116. In some embodiments, the audio component 1110 also includes a speaker for outputting audio signals.

[0145] I / O interface 1112 provides an interface between processing component 1102 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0146] The sensor assembly 1114 includes one or more sensors for providing various aspects of the status assessment of the electronic device 1100. For example, the sensor assembly 1114 can detect the open / closed state of the electronic device 1100, the relative positioning of components, such as the display and keypad of the electronic device 1100. The sensor assembly 1114 can also detect changes in the position of the electronic device 1100 or a component of the electronic device 1100, the presence or absence of user contact with the electronic device 1100, the orientation or acceleration / deceleration of the electronic device 1100, and changes in the temperature of the electronic device 1100. The sensor assembly 1114 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 1114 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 1114 can also include an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0147] The communication component 1116 is configured to facilitate wired or wireless communication between the electronic device 1100 and other devices. The electronic device 1100 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, 4G LTE, 5G NR (NewRadio) or a combination thereof. In an exemplary embodiment, the communication component 1116 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1116 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0148] In an exemplary embodiment, the electronic device 1100 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described methods.

[0149] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1104 including instructions, and the instructions can be executed by the processor 1120 of the electronic device 1100 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0150] The embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the encyclopedia corpus extraction method described in the above embodiments of the present disclosure.

[0151] An embodiment of the present disclosure further provides a computer program product, including a computer program, which, when executed by a processor, executes the encyclopedia corpus extraction method described in the above embodiment of the present disclosure.

[0152] Figure 12 1 is a schematic structural diagram of a chip 1200 for implementing the above-mentioned encyclopedia corpus extraction method according to an exemplary embodiment.

[0153] Reference Figure 12 The chip 1200 includes at least one communication interface 1201 and a processor 1202; the communication interface 1201 is used to receive signals input to the chip 1200 or signals output from the chip 1200, and the processor 1202 communicates with the communication interface 1201 and implements the encyclopedia corpus extraction method described in the above embodiment through logic circuits or execution code instructions.

[0154] An embodiment of the present disclosure further provides a vehicle, which includes the above-mentioned encyclopedia corpus extraction device.

[0155] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0156] Throughout this specification, references to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" indicate that a specific feature, structure, material, or characteristic described in conjunction with an embodiment or example is included in at least one embodiment or example of the present disclosure. In this specification, illustrative uses of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0157] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.

[0158] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processing module, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection having one or more wires (control method), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing it in a suitable manner if necessary, and then storing it in a computer memory.

[0159] It should be understood that the various parts of the embodiments of the present disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0160] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0161] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disk, etc.

[0162] Although the embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are exemplary and are not to be construed as limitations on the present disclosure. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present disclosure.

Claims

1. A method for extracting encyclopedia corpus, characterized in that: The method comprises: Determine the target tag through the target web page source code; Obtaining a corpus type of character information corresponding to the target tag, wherein the corpus type includes unstructured data and structured data; If the corpus type is the unstructured data, filtering the text content in the character information as the text of the unstructured data; If the corpus type is the structured data, the character information and the encyclopedia information associated with the character information are extracted, and a key-value pair is constructed as the content of the structured data.

2. The method according to claim 1, characterized in that After extracting the character information and the encyclopedia information associated with the character information and constructing a key-value pair as the content of the structured data, the method further includes: The structured data is converted into triples to construct a knowledge graph.

3. The method according to claim 1, characterized in that If the corpus type is the unstructured data, screening the text content in the character information as the text of the unstructured data includes: Extracting a directory hierarchy structure of a document object model tree corresponding to the target web page source code; determining the density of the text content in each directory of the directory hierarchy; The text content with a density greater than a preset threshold is regarded as the text of the unstructured data.

4. The method according to claim 3, characterized in that Determining the density of the text content in each directory of the directory hierarchy structure includes: Based on the data type in the text content, obtaining a weight factor corresponding to the data type; The density is obtained using the weight factor and the proportion of text of the data type in the text content.

5. The method according to claim 4, characterized in that The obtaining the density by using the weight factor and the ratio of text of the data type in the text content includes: If the data type is a string type, the density is obtained by multiplying a first weighting factor by a proportion of text of the string type in the text content, wherein the first weighting factor corresponds to the string type; If the data type is a numerical type, the density is obtained by multiplying a second weighting factor by the proportion of text of the numerical type in the text content, wherein the second weighting factor corresponds to the numerical type, and the first weighting factor is different from the second weighting factor.

6. The method according to claim 1, wherein Determining the target tag through the target webpage source code includes: Get the source code of the target web page; Constructing a document object model tree based on the target web page source code; Redundant tags in the document object model tree are deleted to retain the target tags.

7. The method according to claim 1, characterized in that The method further comprises: The target web page source code is converted into a lightweight markup language text format using a text format tool.

8. An encyclopedia corpus extraction device, characterized in that: The device comprises: A determination module is used to determine the target tag through the source code of the target web page; An acquisition module, configured to acquire a corpus type of the character information corresponding to the target tag, wherein the corpus type includes unstructured data and structured data; a screening module, configured to screen text content in the character information as text of the unstructured data if the corpus type is the unstructured data; A construction module is used to extract the character information and the encyclopedia information associated with the character information if the corpus type is the structured data, and to construct a key-value pair as the content of the structured data.

9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.

11. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 7.

12. A chip, characterized in that: The method comprises at least one processor and a communication interface; the communication interface is used to receive a signal input to the chip or a signal output from the chip, and the processor communicates with the communication interface and implements the method according to any one of claims 1 to 7 through a logic circuit or executing code instructions.