Information extraction method fusing semantic analysis, rule matching and deep learning

By integrating information extraction methods with semantic analysis, rule matching and deep learning, the problems of language diversity, context dependence and ambiguity in text extraction are solved, and more efficient and accurate text information extraction is achieved.

CN119962539APending Publication Date: 2025-05-09STATE GRID NINGXIA ELECTRIC POWER CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510022464.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to deal with the diversity, context dependence and ambiguity of language in text extraction, resulting in inaccurate information extracted.

Method used

An information extraction method that integrates semantic analysis, rule matching and deep learning is adopted. Through paragraph slicing, word vector model and deep learning-based information processing module, combined with ambiguity rule matching module, the efficiency and accuracy of text extraction are improved.

Benefits of technology

It realizes more efficient and accurate text extraction, and can automatically generate rules to adapt to changing text content, reduce the cumbersome process of rule maintenance and update, and provide more accurate information matching and emotional color understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962539A_ABST
    Figure CN119962539A_ABST
Patent Text Reader

Abstract

The invention relates to an information extraction method fusing semantic analysis, rule matching and deep learning, and the method comprises the steps: carrying out the information extraction of a text through a paragraph segmentation module, a word vector model information extraction module and an information processing module based on deep learning; the method comprises the following specific steps: a paragraph segmentation module defines a paragraph segmentation rule according to rules and features of symbols used in text editing, performs matching by using a regular expression, and traverses each row in a text to complete paragraph segmentation; the word vector model information extraction module constructs a data set by adopting a sliding window mode, and converts an input text into word vectors; and inputting the word vectors into the information processing module based on deep learning for information extraction, and obtaining text information. The method has the advantages that multi-level information cleaning, screening, extraction and processing are performed on the text information by fusing semantic analysis, a specific rule and a deep learning technology, so that the finally extracted text information is more accurate and smoother.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and in particular to an information extraction method integrating semantic analysis, rule matching and deep learning. Background Art

[0002] The importance of text extraction in the information age is self-evident. It can help people improve the efficiency of information acquisition and save time and energy. The current research status of semantic analysis methods in the field of text extraction is very active. Natural language processing is an important branch of artificial intelligence. The semantic analysis technology in it aims to extract semantic information from text so that computers can better understand and process language. With the advent of the big data era, the development of semantic analysis technology has important practical significance, which helps to extract valuable information hidden in large amounts of text data and provide a basis for human decision-making.

[0003] As a key technology in natural language processing, semantic analysis can accurately extract valuable information through in-depth understanding of text. For example, using statistical language models, we can build a conditional probability model between words based on the frequency of words in different contexts, thereby predicting the probability of future word appearance and providing a basis for text extraction. Hidden Markov models can model the relationship between words, perform semantic analysis, and help identify key information in text extraction. Support vector machines perform well in tasks such as part-of-speech tagging and named entity recognition, providing support for information screening in text extraction. Deep learning, with its advantage of automatically learning semantic information, greatly improves the accuracy of semantic analysis, thereby improving the effect of text extraction.

[0004] However, it is not easy to achieve efficient and accurate text extraction based on semantic analysis. The diversity of languages ​​makes it difficult for computers to understand the rules of different languages, which increases the difficulty of text extraction. Context dependence requires us to fully consider the context when extracting text, otherwise the extracted information may be inaccurate. Ambiguity problems also often plague text extraction. Words in the same sentence may have multiple meanings, and sophisticated semantic analysis is required to determine their true meaning in a specific context. In addition, some words or expressions may be difficult for computers to understand, forming semantic barriers. Summary of the invention

[0005] In order to solve the above problems, the purpose of the present invention is to provide an information extraction method that integrates semantic analysis, rule matching and deep learning, which improves the efficiency and accuracy of text extraction.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] The invention discloses an information extraction method integrating semantic analysis, rule matching and deep learning. The method extracts information from a text through a paragraph segmentation module, a word vector model information extraction module and an information processing module based on deep learning. The specific steps are as follows: the paragraph segmentation module defines paragraph segmentation rules according to the rules and features of symbols used in text editing, uses regular expressions for matching, traverses each line in the text, and completes paragraph segmentation; the word vector model information extraction module constructs a data set in a sliding window manner and converts the input text into a word vector; the word vector is input into the information processing module based on deep learning for information extraction to obtain text information.

[0008] Preferably, after the text is processed by the paragraph segmentation module, paragraphs are obtained, and each paragraph is processed by the ambiguity rule matching module and then input into the word vector model information extraction module; the execution steps of the ambiguity rule matching module are as follows: formulating lexical ambiguity rules: collecting common polysemous words and homonyms, establishing a meaning list for each word, and then determining the part with lexical ambiguity by analyzing the keywords, part of speech, and semantic relationship of the context; formulating syntactic ambiguity rules: determining the part with syntactic ambiguity and the omitted part by analyzing the grammatical structure and semantic relationship; rule matching: traversing each word in the text, matching it through the lexical ambiguity rules, checking whether there are polysemous words or homonyms, if so, determining its specific meaning according to the context analysis rules; traversing each sentence in the text, analyzing the structure of the sentence through the syntactic ambiguity rules, checking whether there is structural ambiguity or omission ambiguity, if so, determining the correct interpretation according to the syntactic analysis rules; result processing: outputting the disambiguated paragraph as the final result for subsequent natural language processing tasks.

[0009] Preferably, the word vector model information extraction module adopts Word2Vec.

[0010] More preferably, the deep learning-based information processing module adopts a long short-term memory network.

[0011] The present invention has the following beneficial effects:

[0012] The present invention combines semantic analysis, rule-based matching and deep learning information extraction methods. In terms of efficiency, the deep learning model can automatically generate rules to adapt to the ever-changing text content, avoiding the cumbersome process of rule maintenance and updating in traditional rule-based methods, and improving the speed and accuracy of text extraction. At the same time, the present invention provides more accurate information matching for different fields by establishing information matching rules and ambiguous understanding rules in professional fields, which can better understand the emotional color and theme tendency of the text and achieve more accurate data preprocessing for the input text.

[0013] The present invention performs multi-level information cleaning, screening, extraction and processing on text information, so that the ultimately extracted text information is more accurate and smoother. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION

[0015] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0016] See also Figure 1 , an information extraction method integrating semantic analysis, rule matching and deep learning, extracts information from the text through a paragraph segmentation module, an ambiguous rule matching module, a word vector model information extraction module and an information processing module based on deep learning. The specific steps are as follows:

[0017] Step 10, the paragraph segmentation module defines paragraph segmentation rules according to the rules and features of the symbols used in text editing, uses regular expressions for matching, traverses each line in the text, and completes paragraph segmentation. A paragraph is usually composed of a group of related sentences, expressing a relatively complete theme or idea. When performing paragraph segmentation, some specific rules and features can be used, such as: ① Blank line: two or more consecutive blank lines often mark the separation of paragraphs. ② Indentation: In some text formats, indentation can indicate the beginning of a new paragraph. ③ Specific punctuation marks: such as a period, question mark, exclamation mark followed by one or more line breaks, may mean the beginning of a new paragraph. According to the above rules, traverse each line in the text and make a judgment according to the above paragraph segmentation rules. If a blank line, indentation, or a combination of specific punctuation marks and line breaks is encountered, it is marked as a paragraph separation point. Regular expressions can be used to match specific patterns, such as \n\n represents two consecutive line breaks, or ^\t represents indentation. Once paragraph segmentation is completed, each paragraph can also be processed separately. For example, you can perform text cleaning to remove extra spaces, special characters, etc.; perform natural language processing tasks such as part-of-speech tagging and named entity recognition to extract richer information; or perform keyword extraction to quickly understand the topic of a paragraph.

[0018] Step 20: After the text is processed by the paragraph segmentation module, each paragraph is processed by the ambiguous rule matching module to improve the accuracy of the text representation of each paragraph. The execution steps of the ambiguous rule matching module are as follows:

[0019] Step 21. Develop rules for lexical ambiguity: 1) Build a vocabulary list: Collect common polysemous words and homonyms, and build a list of meanings for each word. 2) Context analysis: Determine the specific meaning of a word based on the context in which it is located. For example, if in "I went to the fruit store to buy some apples", "apple" obviously refers to fruit; while in "I took a photo with an iPhone", "apple" refers to a technology company. Develop rules by analyzing the context's keywords, parts of speech, semantic relationships, and other factors.

[0020] Step 22, formulate syntactic ambiguity rules: 1) Structural analysis: Analyze the structure of the sentence to determine the parts that may be ambiguous. For example, for a sentence like "a prime minister who loves the people", the correct interpretation can be determined by analyzing the grammatical structure and semantic relationship. If there is a clear modification relationship or qualifier in the sentence, it can help eliminate ambiguity. 2) Omission completion: For sentences with omission ambiguity, you can try to complete the omitted components according to the context. For example, in "He likes to eat apples and bananas, and I like them too", if the above text clearly mentions "he likes to eat apples and bananas", then "I" probably also like apples and bananas.

[0021] Step 23, rule matching: 1) Lexical ambiguity matching: traverse each word in the text to check whether there are polysemous words or homophones. If so, determine its specific meaning based on the context analysis rules. Word vector technology, semantic similarity calculation and other methods can be used to assist in judgment. 2) Syntactic ambiguity matching: Analyze the structure of the sentence to check whether there is structural ambiguity or omission ambiguity. If so, determine the correct interpretation based on the syntactic analysis rules. You can use syntactic analysis tools such as Stanford Syntactic Analyzer, Harbin Institute of Technology Language Technology Platform, etc.

[0022] Step 24, result processing: 1) Disambiguation: Disambiguate the text according to the result of rule matching. If a word has multiple meanings, the meaning that best fits the context can be selected; if there is structural ambiguity in the sentence, the correct interpretation can be determined based on grammatical rules and semantic relationships; 2) Output results: The disambiguated text is output as the final result, which can be used for subsequent natural language processing tasks, such as text classification, sentiment analysis, machine translation, etc.

[0023] Step 30, the word vector model information extraction module uses a sliding window to construct a data set and convert the input text into a word vector. The word vector model information extraction module can use Word2Vec or GloVe. The word vector model plays a key basic role in semantic analysis. Word2Vec can represent words in the text with numerical vectors through training, and can measure the similarity of two words well in the vector space. For example, during the training process, Word2Vec not only updates the weight matrix of the network, but also updates the input mapping matrix, so that words with similar semantics such as man and boy are closer in the word vector space, while words with larger semantic differences such as water are farther away. Word2Vec training is an unsupervised process, and any text can be used as a data set. The data set is usually constructed in a sliding window manner, and the first few words are selected as input, and the last word is predicted as output. Its commonly used training schemes are CBOW and Skip-gram. CBOW predicts the central word based on the context, while Skip-gram predicts the context based on the central word. In order to simplify the calculation, optimization methods such as negative sample sampling can also be used.

[0024] Step 40, input the word vector into the information processing module based on deep learning to extract information and obtain text information. The deep learning model plays an important role in semantic analysis by virtue of its powerful automatic learning ability and ability to process complex data. Recurrent neural network (RNN) is also widely used in semantic analysis. RNN is a neural network specially used to process sequence data, which can mine the time series information and semantic information in the data. Its network structure makes the hidden layer vector at each moment not only determined by the input at the current moment, but also depends on the hidden layer vector at the previous moment, which enables RNN to remember the past time series information. Long short-term memory network (LSTM) is an improved model of RNN. LSTM solves the gradient disappearance and gradient explosion problems in long sequences, thereby realizing efficient processing of sequence data. The core structure of LSTM includes three gates (input, output, forget) and a memory cell (cell). The input gate is used to control the input information flow in and out, the output gate is used to control the output information flow in and out, and the forget gate is used to control the information flow of the memory cell. For example, in natural language processing tasks, LSTM can better process long texts, remember important information, and ignore irrelevant information. In tasks such as text generation and machine translation, LSTM can generate coherent subsequent text based on the previous text content, improving the accuracy and fluency of translation.

[0025] The above description is only a specific implementation mode of the present invention, and does not limit the patent scope of the present invention. Any equivalent structural transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. An information extraction method integrating semantic analysis, rule matching and deep learning, characterized in that: The text is extracted through the paragraph segmentation module, the word vector model information extraction module and the information processing module based on deep learning. The specific steps are as follows: The paragraph segmentation module defines paragraph segmentation rules according to the rules and features of symbols used in text editing, uses regular expressions for matching, traverses each line in the text, and completes paragraph segmentation; The word vector model information extraction module constructs a data set in a sliding window manner to convert the input text into a word vector; The word vector is input into the deep learning-based information processing module to extract information and obtain text information.

2. The information extraction method integrating semantic analysis, rule matching and deep learning according to claim 1, characterized in that: After the text is processed by the paragraph segmentation module, paragraphs are obtained, and each paragraph is processed by the ambiguous rule matching module and then input into the word vector model information extraction module; the execution steps of the ambiguous rule matching module are as follows: Formulate lexical ambiguity rules: Collect common polysemous words and homonyms, create a meaning list for each word, and then determine the parts with lexical ambiguity by analyzing the keywords, parts of speech, and semantic relationships in the context; Formulate syntactic ambiguity rules: determine the ambiguous and omitted parts by analyzing the grammatical structure and semantic relations; Rule matching: traverse each word in the text, match it according to the lexical ambiguity rules, check whether there are polysemous words or homonyms, and if so, determine its specific meaning according to the context analysis rules; Traverse each sentence in the text, analyze the structure of the sentence through the syntactic ambiguity rules, check whether there is structural ambiguity or omission ambiguity, and if so, determine the correct interpretation according to the syntactic analysis rules; Result processing: The disambiguated paragraph is output as the final result for subsequent natural language processing tasks.

3. The information extraction method integrating semantic analysis, rule matching and deep learning according to claim 1, characterized in that: The word vector model information extraction module adopts Word2Vec.

4. The information extraction method integrating semantic analysis, rule matching and deep learning according to claim 1, characterized in that: The deep learning-based information processing module adopts a long short-term memory network.