Vertical field-oriented text extraction and knowledge enhancement method

Through text optimization preprocessing, deep learning framework and vertical domain knowledge graph construction, the problems of low long text processing efficiency and lack of knowledge in vertical domains are solved, and efficient and accurate text extraction and knowledge enhancement are achieved.

CN120493923APending Publication Date: 2025-08-15INFORMATION & COMMUNICATION BRANCH STATE GRID JIBEI ELECTRIC POWER CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510654900.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art is inefficient when processing long and complex text data, making it difficult to accurately extract key information, and the existing knowledge-enhancing technology lacks rich background knowledge in the vertical field, resulting in incomplete extraction results or poor coherence.

Method used

Text optimization preprocessing, deep learning framework, parallel processing technology and vertical domain knowledge graph construction, combined with hyperbolic spatial learning and contrast learning, vertical domain knowledge graph is constructed through word segmentation, elimination of stop words, sentence segmentation, segmentation extraction and merging, and hyperbolic spatial learning and contrast learning technology is used to improve text extraction efficiency and knowledge enhancement effect.

Benefits of technology

It significantly improves the efficiency of text extraction and the completeness and coherence of results, effectively solves the problem of small scale and sparse structure of knowledge graphs in vertical fields, realizes automatic identification and efficient extraction of key information, and enhances the depth and accuracy of knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493923A_ABST
    Figure CN120493923A_ABST
Patent Text Reader

Abstract

The invention provides a vertical field-oriented text extraction and knowledge enhancement method, and relates to the technical field of natural language processing, the vertical field-oriented text extraction and knowledge enhancement method comprises the steps of text optimization preprocessing, adoption of a text extraction algorithm, segmented extraction and result combination, and introduction of a parallel processing technology; the knowledge enhancement method comprises the following steps: constructing a vertical domain knowledge graph, introducing hyperbolic space learning, and implementing comparative learning; through a series of steps of text optimization preprocessing, word segmentation, stop word elimination, sentence segmentation and the like, interference of noise information is effectively reduced, a long text or a complex-structure text is disassembled into units easier to process, key information can be automatically recognized and extracted in combination with a text extraction model trained by a deep learning framework, and the method is high in practicability and easy to popularize. And the text extraction efficiency and the result integrity and continuity are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a text extraction and knowledge enhancement method for vertical fields. Background Art

[0002] With the explosive growth of text data, text extraction and knowledge augmentation technologies have played an important role in information extraction, natural language processing, and other fields. However, existing technologies have significant limitations when processing long and complex text data. When faced with lengthy and complex documents, existing extraction technologies are often inefficient and have difficulty accurately capturing and extracting key information. This is mainly because these texts contain a large amount of details and layers, which exceeds the processing capabilities of traditional methods, resulting in incomplete or poorly coherent extraction results. On the other hand, existing knowledge augmentation technologies mainly rely on large-scale general domain knowledge graphs. Although this provides good support for a wide range of applications, it is insufficient for applications in specific vertical fields. The knowledge graphs in vertical fields are small in scale and sparse in structure, and cannot provide sufficiently rich background knowledge and detailed information, which directly affects the effectiveness and depth of knowledge augmentation. Therefore, in vertical fields, existing methods find it difficult to achieve effective information augmentation and supplementation through existing knowledge graphs. In summary, the present invention aims to provide a text extraction and knowledge enhancement method for vertical fields to address the above technical deficiencies. Summary of the Invention

[0003] The purpose of the present invention is to address the shortcomings of the existing technology, effectively improve the efficiency of text extraction and the completeness and consistency of the results, and how to use fewer and more specific knowledge graphs for effective knowledge enhancement.

[0004] To achieve the above objectives, the present invention adopts the following technical solutions: a text extraction and knowledge enhancement method for vertical fields, wherein the text extraction method includes text optimization preprocessing, using natural language processing technology to segment long texts or complex structured texts, and breaking them down into more fine-grained language units; Eliminate stop words to reduce the interference of noise information; Implement sentence segmentation to cut the text into independent sentences or paragraph units; Using a text extraction algorithm and based on a deep learning framework, we train a text extraction model. The model learns and masters the rules for extracting key text information through a large amount of annotated data. Segment extraction and result merging: split long text or complex structured text into multiple paragraphs or sentence units and extract them separately; then merge and organize the extraction results; Introducing parallel processing technology, using multi-threading or distributed computing frameworks to distribute text processing tasks to multiple processors or nodes, achieving parallel processing to improve efficiency; The knowledge enhancement method includes constructing a vertical field knowledge graph, collecting professional knowledge resources in the vertical field, including but not limited to academic papers, patents or industry reports; systematically organizing, classifying and integrating the collected knowledge, and then constructing a knowledge graph; Hyperbolic space learning is introduced to embed vertical domain knowledge graphs into hyperbolic space to learn the hierarchical semantic information of graph data; Implement contrastive learning and use the contrastive learning framework to construct positive and negative samples of different graph structure difficulties; further train the model to enhance the ability to handle semantic sparsity problems.

[0005] Preferably, the text optimization preprocessing uses a word segmentation algorithm and a stop word list in natural language processing (NLP) combined with regular expressions to perform sentence segmentation.

[0006] Preferably, the text optimization preprocessing introduces a domain-specific stop word library according to the corresponding domain.

[0007] Preferably, the deep learning framework uses BERT or Transformer.

[0008] Preferably, the deep learning model automatically identifies and extracts key information from the text, and uses fine-tuning technology to adapt the model to different vertical fields.

[0009] Preferably, the segmentation extraction and result merging process adopts a text segmentation algorithm to cut the text into independent processing units; After the extraction is completed, the text splicing algorithm is used to merge the results; In the merging stage, semantic similarity calculation technology is introduced to sort and optimize the extraction results.

[0010] Preferably, the parallel processing technology uses the Message Passing Interface (MPI) or Apache Spark parallel computing framework to achieve distributed processing of the text extraction task; Combined with load balancing strategies, it ensures balanced task distribution among processors or nodes.

[0011] Preferably, the knowledge graph construction technology uses entity recognition, relationship extraction and graph fusion; it can also be combined with crowdsourcing and expert annotation technology.

[0012] Preferably, the hyperbolic space learning is introduced to map the graph data into the hyperbolic space using a hyperbolic embedding algorithm, and is combined with technologies such as graph neural network (GNN) to improve the hyperbolic embedding effect.

[0013] Preferably, the contrastive learning constructs positive and negative sample pairs so that the model can learn the semantic similarity in the atlas data, and uses the contrastive loss function to train the model to improve the generalization ability of the model.

[0014] Compared with the prior art, the present invention has the following beneficial effects: Through a series of steps including text optimization preprocessing, word segmentation, stop word removal, and sentence segmentation, the present invention effectively reduces the interference of noise information and breaks down long or complex text structures into more easily processable units. Combined with a text extraction model trained in a deep learning framework, it can automatically identify and extract key information, significantly improving the efficiency of text extraction and the completeness and coherence of the results. This invention constructs a specialized knowledge graph based on the characteristics of vertical fields, and introduces hyperbolic space learning and contrastive learning techniques, effectively solving the problems of small scale and sparse structure of vertical field knowledge graphs. Through these technologies, the model can better learn and master the semantic information of vertical fields, and achieve effective knowledge enhancement and supplementation. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 Flowchart of the vertical field-oriented text extraction and knowledge enhancement method provided by the present invention. DETAILED DESCRIPTION

[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0017] To facilitate understanding of the present invention, the present invention will be described more comprehensively below with reference to relevant references, and several embodiments of the present invention are given. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0018] It should be noted that when an element is referred to as being "fixed on" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used in this article are for illustrative purposes only.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention pertains. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items. Example

[0020] like Figure 1 As shown, the present invention provides a technical solution: a text extraction and knowledge enhancement method for vertical fields, the text extraction method includes text optimization preprocessing, Step 1) Use natural language processing (NLP) to segment long or complex texts into smaller, more granular language units, such as words and phrases, for easier processing. This step uses NLP's word segmentation algorithm, combined with a domain-specific stop word library, to improve preprocessing accuracy and reduce the complexity of subsequent processing. Step 2: Remove stop words. These stop words usually include prepositions, conjunctions, and auxiliary words. They do not contribute much to the core information of the text, but they increase the processing noise. By removing these words, the text content can be further streamlined. Step 3) Implement sentence segmentation and use regular expressions or sentence segmentation algorithms in NLP to cut the text into independent sentences or paragraph units to facilitate subsequent text extraction processing.

[0021] Deep learning model training and text extraction: Use deep learning frameworks such as BERT or Transformer to train text extraction models. This model learns and masters the rules for extracting key information from texts through a large amount of annotated data, and can automatically identify and extract key information from texts. During the training process, fine-tuning technology is used to adapt the model to different vertical fields and improve the accuracy and pertinence of extraction; It should be added that the combination of attention mechanism enables the model to focus on the key parts of the text more accurately; Segment extraction and result merging: split long text or complex structure text into multiple paragraphs or sentence units and extract them separately; After the extraction is completed, the text splicing algorithm is used to merge and organize the extraction results to form a complete extraction result; In the merging stage, semantic similarity calculation technology is introduced to sort and optimize the extraction results to ensure the consistency and accuracy of the extraction results.

[0022] Parallel processing technology: By introducing parallel processing technology, text processing tasks can be distributed to multiple processors or nodes with the help of multi-threaded or distributed computing frameworks (such as the message passing interface MPI or Apache Spark), achieving parallel processing to improve efficiency; Furthermore, combined with the load balancing strategy, the balanced distribution of tasks among processors or nodes is ensured to avoid resource idleness or overload. With the help of multi-threaded or distributed computing framework, the present invention realizes the parallel processing of text processing tasks, improves processing efficiency, and further enhances overall performance. Example

[0023] like Figure 1 As shown in the figure, the knowledge enhancement method includes constructing a vertical field knowledge graph, collecting professional knowledge resources in the vertical field, including academic papers, patents, industry reports, etc.; systematically organizing, classifying and integrating the collected knowledge to construct a knowledge graph; in the construction process, technical means such as entity recognition, relationship extraction and graph fusion can be selected, and combined with crowdsourcing, expert annotation and other means to improve the quality and accuracy of the knowledge graph; Then, we introduce hyperbolic space learning and use hyperbolic embedding algorithms (such as Poincaré embedding) to map the graph data into the hyperbolic space to learn the hierarchical semantic information of the graph data. Combined with technologies such as graph neural networks (GNNs), this further improves the hyperbolic embedding effect and enhances the model's ability to understand complex semantic relationships; Finally, we implement contrastive learning, using the contrastive learning framework to construct positive and negative sample pairs with different graph structure difficulties. Through contrastive learning, the model can learn the semantic similarity in the graph data and improve the generalization ability of the model. The model is trained using the contrastive loss function and the model parameters are continuously optimized to make it more robust and accurate when dealing with semantic sparsity problems.

[0024] To sum up, the implementation process of the present invention is as follows: first, the input long text or complex structure text is preprocessed for text optimization, including steps such as word segmentation, removal of stop words and sentence segmentation, and then the preprocessed text is segmented and extracted using a trained deep learning model. After the extraction is completed, the extraction results of each segment are merged and sorted to form a complete extraction result; at the same time, the extraction results are enhanced using the constructed vertical field knowledge graph, and technical means such as hyperbolic space learning and contrastive learning are introduced to improve the effect of knowledge enhancement, and finally the enhanced text extraction results are output for subsequent applications.

[0025] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A vertical field-oriented text extraction and knowledge enhancement method, characterized by: The text extraction method includes text optimization preprocessing, (1) using natural language processing technology to segment long texts or complex structured texts and break them down into more fine-grained language units; (2) Eliminate stop words to reduce the interference of noise information; (3) Implement sentence segmentation to cut the text into independent sentences or paragraph units; Using a text extraction algorithm and based on a deep learning framework, we train a text extraction model. The model learns and masters the rules for extracting key text information through a large amount of annotated data. Segment extraction and result merging: split long text or complex structured text into multiple paragraphs or sentence units and extract them separately; then merge and organize the extraction results; Introducing parallel processing technology, using multi-threading or distributed computing frameworks to distribute text processing tasks to multiple processors or nodes, achieving parallel processing to improve efficiency; The knowledge enhancement method includes constructing a vertical field knowledge graph, collecting professional knowledge resources in the vertical field, including but not limited to academic papers, patents or industry reports; systematically organizing, classifying and integrating the collected knowledge, and then constructing a knowledge graph; Hyperbolic space learning is introduced to embed vertical domain knowledge graphs into hyperbolic space to learn the hierarchical semantic information of graph data; Implement contrastive learning and use the contrastive learning framework to construct positive and negative samples of different graph structure difficulties; further train the model to enhance the ability to handle semantic sparsity problems.

2. The vertical field-oriented text extraction and knowledge enhancement method according to claim 1 is characterized by: The text optimization preprocessing adopts the word segmentation algorithm and stop word list in natural language processing (NLP) combined with regular expressions to perform sentence segmentation.

3. The vertical field-oriented text extraction and knowledge enhancement method according to claim 1 is characterized by: The text optimization preprocessing introduces a domain-specific stop word library according to the corresponding domain.

4. The vertical field-oriented text extraction and knowledge enhancement method according to claim 1 is characterized by: The deep learning framework uses BERT or Transformer.

5. The vertical field-oriented text extraction and knowledge enhancement method according to claim 1 is characterized by: The deep learning model automatically identifies and extracts key information from text, and uses fine-tuning technology to adapt the model to different vertical fields.

6. The vertical field-oriented text extraction and knowledge enhancement method according to claim 1 is characterized by: The segmentation extraction and result merging process uses a text segmentation algorithm to cut the text into independent processing units; After the extraction is completed, the text splicing algorithm is used to merge the results; In the merging stage, semantic similarity calculation technology is introduced to sort and optimize the extraction results.

7. The vertical field-oriented text extraction and knowledge enhancement method according to claim 1 is characterized by: The parallel processing technology uses the Message Passing Interface (MPI) or Apache Spark parallel computing framework to achieve distributed processing of text extraction tasks; Combined with load balancing strategies, it ensures balanced task distribution among processors or nodes.

8. The vertical field-oriented text extraction and knowledge enhancement method according to claim 1 is characterized by: The knowledge graph construction technology uses entity recognition, relationship extraction and graph fusion; it can also be combined with crowdsourcing and expert labeling technical means.

9. The vertical field-oriented text extraction and knowledge enhancement method according to claim 1 is characterized by: The introduction of hyperbolic space learning utilizes a hyperbolic embedding algorithm to map graph data into a hyperbolic space, combined with technologies such as graph neural networks (GNNs).

10. The vertical field-oriented text extraction and knowledge enhancement method according to claim 1 is characterized by: The contrastive learning constructs positive and negative sample pairs so that the model can learn the semantic similarity in the atlas data and trains the model using the contrastive loss function.