Heterogeneous document-oriented ontology entity high-precision identification and verification method and device
By converting heterogeneous documents into standardized text and combining high-dimensional semantic space mapping and triple verification mechanism, the problem of insufficient accuracy of ontology entity recognition in heterogeneous documents is solved, and high-precision entity recognition and verification are achieved.
Patent Information
- Application Number
- CN202510942933.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing entity recognition methods have problems such as insufficient accuracy and recognition results that do not conform to ontology constraints when processing heterogeneous documents, making it difficult to accurately obtain the data users want from the massive data on the Internet.
A multi-step collaborative processing approach is adopted, including converting heterogeneous documents into standardized text, mapping the candidate entity set into a high-dimensional semantic space using the local library, performing multi-granularity similarity matching through an improved hierarchical K-nearest neighbor algorithm, and filtering the candidate entity set through a triple verification mechanism to ensure compliance with ontology constraints.
It significantly improves the precision and accuracy of ontology entity recognition in heterogeneous documents, reduces noise interference, improves processing efficiency, supports cross-modal entity association analysis, and reduces false positive and false negative rates.
Smart Images

Figure CN120805909A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer science, natural semantic processing and ontology data standardization technology, and specifically relates to a high-precision identification and verification method for ontology entities for heterogeneous documents, which can be applied to smart cities, map services, data navigation and government address standardization scenarios. Background Art
[0002] With the massive amount of data on the Internet, how to obtain accurate data with high precision has become particularly important. However, in reality, the massive data on the Internet is not standardized and incomplete. Existing entity recognition methods have problems such as insufficient accuracy when processing heterogeneous documents and the recognition results do not meet the ontology constraints. How to accurately obtain the data users want from a large amount of data? Traditional computing methods are difficult to cope with complex situations.
[0003] In order to solve the above problems, there is an urgent need for a high-precision ontology entity recognition and verification method for heterogeneous documents, which can significantly improve the precision and accuracy of entity recognition based on Internet data, through standardized text, and multi-step collaborative processing. Summary of the Invention
[0004] The purpose of the present invention is to provide a high-precision ontology entity recognition and verification method for heterogeneous documents, so as to improve the precision and accuracy of entity recognition.
[0005] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention specifically comprises the following steps:
[0006] S1. Convert heterogeneous documents (PDF, DOC / DOCX, TXT, etc.) into standardized text. Let the heterogeneous document set be D = {d1, d2, ..., d n}, the standardized text set is T = {t1,t2,...,t n}, the conversion function is The conversion formula is: D is a heterogeneous document set, d1 represents the first heterogeneous document, d n represents the nth heterogeneous document, T represents the standard document set, t1 represents the first converted standard document, t n Represents the nth converted standard document, For the conversion method.
[0007] S2. Extract candidate entity set E from the standardized text set T i ={e1,e2,...,e n}, the extraction function is The extracted formula is: E iis the first entity, e n represents the nth entity, is the extraction method.
[0008] S3, the ontology library and the candidate entity set are mapped to a unified high-dimensional semantic space. The set of local libraries is O={o1, o2,..., o n}, and the candidate entity set is E i , and the function of joint embedding is The mapping formula is: Embed O is the vector of the ontology library, and Embed E is the space vector of the entity set E i , O is the local library, o1 is the first local library, and o n is the nth local library.
[0009] S4, multi-granularity similarity matching is performed through an improved hierarchical K-nearest neighbor algorithm (H-KNN) to generate a preliminary candidate entity set. The vectors in the joint embedding space are set to Embed O and Embed E , and the hierarchical K-nearest neighbor function is The mapping formula is: S i is the preliminary candidate entity set.
[0010] S5, the candidate entity set is filtered through a triple verification mechanism to ensure that it meets the ontology constraints. The preliminary candidate set result S i ={s1, s2,..., s p}, and the high-confidence candidate entity set after triple verification is H i ={h1, h2,..., h q}, and the verification functions are respectively The calculation formula is: where S i is the candidate result set, and H i is the calculated high-confidence candidate entity set. The triple verification mainly includes common sense reasoning verification semantic correlation degree filtering logical consistency verification The specific mapping formula is:
[0011]
[0012] s j is a subset of S i , and LLM is a language large model GTP-3.
[0013]
[0014] o k is a subset of the local library O, and τ is a threshold value of cosine similarity.
[0015]
[0016] LogicCheck is a logic reasoning tool Prolog.
[0017] S6, output the high-confidence candidate entity set that meets the ontology constraints, and the calculation formula is: h j is a subset of S i that passes the triple verification.
[0018] The application also provides a computer readable storage medium, which stores a computer program, and when the program is executed by a processor, the processor executes the ontology entity high-precision identification and verification method for heterogeneous documents.
[0019] The application also provides a computer system, which comprises:
[0020] a processor for executing instructions stored in a storage medium to implement the ontology entity high-precision identification and verification method for heterogeneous documents;
[0021] a memory for storing network texts and intermediate results generated in the calculation process of the ontology entity high-precision identification and verification method for heterogeneous documents;
[0022] an input / output interface for receiving specified texts or identification information and outputting identified ontology elements.
[0023] The application also provides an ontology entity high-precision identification and verification device for heterogeneous documents, which is configured to: acquire standardized text content; extract a candidate entity set; perform joint embedding space calculation; perform multi-granularity similarity matching; perform a triple verification mechanism: large language model reasoning verification, semantic correlation degree filtering, and logic consistency verification; and output a high-confidence candidate entity set.
[0024] The ontology entity high-precision identification and verification method and device for heterogeneous documents provided by the application have the following effects:
[0025] 1. The method eliminates differences in heterogeneous document formats, constructs a unified semantic baseline, accurately locates candidate entities, reduces noise interference, and improves subsequent processing efficiency.
[0026] 2. The method realizes semantic alignment of texts, images, tables and other modalities, supports cross-modal entity association analysis, dynamically adjusts matching weights through multi-granularity semantic comparison at the character level, word level and sentence level, and improves the robustness to synonyms, abbreviations and misspelled words.
[0027] 3. The method is based on a knowledge graph or ontology library, verifies the pre-defined semantic association between entities, ensures that the entity recognition result meets the requirements in terms of logic, semantics and domain rules, and reduces the false positive and false negative rates. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 The method is based on a knowledge graph or ontology library, verifies the pre-defined semantic association between entities, ensures that the entity recognition result meets the requirements in terms of logic, semantics and domain rules, and reduces the false positive and false negative rates.
[0029] Figure 2 The method is based on a knowledge graph or ontology library, verifies the pre-defined semantic association between entities, ensures that the entity recognition result meets the requirements in terms of logic, semantics and domain rules, and reduces the false positive and false negative rates. DETAILED DESCRIPTION
[0030] The specific embodiments of the present application are described below to facilitate understanding of the present application by those skilled in the art, but it should be clear that the present application is not limited to the scope of the specific embodiments, and that all applications utilizing the concept of the present application are within the scope of the appended claims, provided that various changes are obvious to those skilled in the art within the spirit and scope of the present application.
[0031] The specific implementation steps of the present application are described in detail below with the example of "identifying Wuyi Mountain range in the Mountain range ontology from multi-source documents".
[0032] 1. Obtain standardized text content:
[0033] Suppose there are the following multi-source documents:
[0034] A PDF document with the content: "Wuyi Mountain is located in Fujian Province, and is one of the famous mountain ranges in China."
[0035] A DOCX document with the content: "Wuyi Mountain in Fujian Province is a tourist destination."
[0036] A TXT document with the content: "Wuyi Mountain is famous for its unique natural scenery."
[0037] Use document parsing tools (such as PDFPlumber, Docx2txt, TxtParser) to convert the above documents into standardized text:
[0038] Text 1: Wuyi Mountain is located in Fujian Province, and is one of the famous mountain ranges in China.
[0039] Text 2: Wuyi Mountain in Fujian Province is a tourist destination.
[0040] Text 3: Wuyi Mountain is famous for its unique natural scenery.
[0041] 2. Use the pre-trained language model BERT to perform entity recognition on the standardized text and extract the candidate entity set:
[0042] Candidate entities for Text 1: {"Wuyi Mountains", "Fujian Province", "China", "Mountain Range"}
[0043] Candidate entities for Text 2: {"Fujian Province", "Wuyi Mountains", "Tourist Destination"}
[0044] Candidate entities for Text 3: {"Wuyi Mountains", "Natural Landscape"}
[0045] Merge all candidate entities to get the candidate entity set:
[0046] Candidate entity set: {"Wuyi Mountains", "Fujian Province", "China", "Mountain Range", "Tourist Destination", "Natural Landscape"}
[0047] 3. Joint Embedding Space Computation:
[0048] Let the ontology library (O = {"Wuyi Mountains Range", "Himalayan Range", "Mount Tai Range"}), and the candidate entity set (E = {"Wuyi Mountains", "Fujian Province", "China", "Mountain Range", "Tourist Destination", "Natural Landscape"}).
[0049] Use a contrastive learning algorithm (such as SimCLR) to train the embedding model, mapping the ontology library and candidate entities to a high-dimensional semantic space:
[0050] Ontology library embeddings: Wuyi Mountains Range: [[0.85, 0.12, -0.45,...]] Himalayan Range: [[0.92, 0.08, -0.38,...]] Mount Tai Range: [[0.88, 0.15, -0.42,...]]
[0051] Candidate entity embeddings: Wuyi Mountains: [[0.86, 0.11, -0.44,...]] Fujian Province: [[0.10, 0.75, -0.20,...]] China: [[0.15, 0.80, -0.25,...]] Mountain Range: [[0.81, 0.13, -0.43,...]] Tourist Destination: [[0.30, 0.70, -0.15,...]] Natural Landscape: [[0.35, 0.65, -0.18,...]]
[0052] 4. Multi-granularity Similarity Matching:
[0053] Use the H-KNN algorithm to calculate the similarity between candidate entities and the ontology library:
[0054] Wuyi Mountains: Similarity with "Wuyi Mountains Range" is 0.95
[0055] Fujian Province: Similarity with "Wuyi Mountains Range" is 0.20
[0056] China: Similarity to "Wuyi Mountain Range" is 0.25
[0057] Mountain Range: Similarity to "Wuyi Mountain Range" is 0.85
[0058] Tourist Destination: Similarity to "Wuyi Mountain Range" is 0.35
[0059] Natural Landscape: Similarity to "Wuyi Mountain Range" is 0.40
[0060] Select candidate entities with similarity higher than threshold (0.5), get preliminary candidate entity set:
[0061] Preliminary candidate entity set: {"Wuyi Mountain", "Mountain Range"}
[0062] 5. Triple verification mechanism:
[0063] Common sense reasoning verification: Use large language model GPT-3 to reason about "Wuyi Mountain" and "Mountain Range":
[0064] "Wuyi Mountain" is reasoned as a mountain range, consistent with common sense.
[0065] "Mountain Range" is reasoned as a geographical concept, consistent with common sense. Filter results: {"Wuyi Mountain", "Mountain Range"}
[0066] Semantic correlation filtering: Calculate the cosine similarity between candidate entities and ontology concepts, filter entities with low semantic correlation (threshold set to 0.8):
[0067] "Wuyi Mountain" has a similarity of 0.95 with "Wuyi Mountain Range"
[0068] "Mountain Range" has a similarity of 0.85 with "Wuyi Mountain Range" Filter results: {"Wuyi Mountain", "Mountain Range"}
[0069] Logical consistency verification: Use logical reasoning tools to verify the logical consistency of candidate entities and ontology:
[0070] "Wuyi Mountain" is logically consistent with "Wuyi Mountain Range".
[0071] "Mountain Range" is logically a superordinate concept of "Wuyi Mountain Range", but does not directly match. Filter results: {"Wuyi Mountain"}
[0072] 6. Output high-confidence candidate entity set:
[0073] High-confidence candidate entity set: {"Wuyi Mountain"}
[0074] As a result of analysis, through the above steps, the application successfully identifies the entity "Wuyi Mountain" from the multi-source document and accurately matches with "Wuyi Mountain Range" in the ontology library, verifying the effectiveness and high precision of the method.
[0075] In summary, the application aims to provide an ontology entity high-precision identification and verification method for heterogeneous documents. Through formal expression and accurate calculation, combined with pre-trained language models, contrastive learning algorithms, improved hierarchical K-nearest neighbor algorithm and triple verification mechanism, the precision and reliability of ontology entity identification in heterogeneous documents are significantly improved, providing reliable support for smart city, map service, data navigation and government address standardization scenarios.
Claims
1. A high-precision ontology entity recognition and verification method for heterogeneous documents, characterized by: The following steps are involved: S1: Obtain standardized text content; S2: Extract candidate entity set; S3: Joint embedding space computation; S4: Multi-granularity similarity matching; S5: Triple verification mechanism: large language model reasoning verification, semantic relevance filtering, and logical consistency verification; S6: Output a set of high-confidence candidate entities.
2. The method according to claim 1, characterized in that In step S1, the heterogeneous documents are converted into standardized texts. Let the heterogeneous document set be D = {d1, d2, ..., d n }, the standardized text set is T = {t1,t2,...,t n }, the conversion function is The conversion formula is: D is a heterogeneous document set, d1 represents the first heterogeneous document, d n represents the nth heterogeneous document, T represents the standard document set, t1 represents the first converted standard document, t n Represents the nth converted standard document, For the conversion method.
3. The method according to claim 1, characterized in that In step S2, the candidate entity set E is extracted from the standardized text set T i ={e1,e2,...,e n }, the extraction function is The extracted formula is: E i is a collection of entity classes, e1 is the first entity, e n represents the nth entity, For the extraction method.
4. The method according to claim 1, wherein In step S3, the ontology library and the candidate entity set need to be mapped to a unified high-dimensional semantic space. The set of local libraries is O = {o1, o2, ..., o n }, the candidate entity set is E i , the joint embedding function is The mapping formula is: Embed O is the vector of the ontology library, Embed E For the entity set E i The space vector of O is the local library, o1 is the first local library, o n For the nth local library.
5. The method according to claim 1, wherein In step S4, the improved hierarchical K-nearest neighbor algorithm is used to perform multi-granularity similarity matching to generate a preliminary candidate entity set. Set the vector in the joint embedding space as Embed O and Embed E , the hierarchical K nearest neighbor function is The mapping formula is: S i is the preliminary candidate entity set.
6. The method according to claim 1, wherein In step S5, the candidate entity set is filtered through a triple verification mechanism to ensure that it complies with the ontology constraints; the preliminary candidate set result S i ={s1,s2,...,s p }, the set of high confidence candidate entities after triple verification is H i ={h1,h2,...,h q }, the verification functions are The calculation formula is: Among them S i is the candidate result set, H i The high confidence candidate entity set for calculation, where the triple verification is mainly common sense reasoning verification Semantic relevance filtering Logical consistency verification The specific mapping formula is: s j For S i A subset of LLM, LLM is the language large model GTP-3; o k is a subset of the local library O, τ is the threshold of cosine similarity; LogicCheck is a logical reasoning tool for Prolog.
7. The method according to claim 1, characterized in that In step S6, a set of high-confidence candidate entities that meet the ontology constraints is output. The calculation formula is: h j For S that has passed triple verification i A subset of .
8. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to execute the high-precision ontology entity recognition and verification method for heterogeneous documents described in claims 1-7.
9. A computer system comprising: A processor, configured to execute instructions stored in a storage medium to implement the high-precision ontology entity recognition and verification method for heterogeneous documents as described in claims 1-7; A memory for storing network text and intermediate results generated during the calculation process of the ontology entity high-precision recognition and verification method for heterogeneous documents; The input / output interface is used to receive the specified text or recognition information and output the recognized ontology elements.
10. A high-precision ontology entity recognition and verification device for heterogeneous documents, characterized in that: The device is configured to: obtain standardized text content; extract candidate entity sets; jointly calculate embedding space; match multi-granularity similarity; and implement a triple verification mechanism: large language model reasoning verification, semantic relevance filtering, and logical consistency verification; and output a high-confidence candidate entity set.
Citation Information
Patent Citations
Contrast learning prediction method and system for heterogeneous knowledge graph
CN116401380A
Entity alignment method based on attribute and relation perception heterogeneous graph converter
CN116467395A
Semantic matching and retrieval of standardized entities
US20210303638A1