Information extraction method and system based on natural language text
By adopting a variety of artificial intelligence models and algorithms in the natural language text information extraction system, the problems of slow processing speed, low accuracy, poor flexibility and insufficient semantic understanding in the prior art are solved, and efficient, accurate and flexible information extraction effects are achieved.
Patent Information
- Application Number
- CN202510311445.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-17
AI Technical Summary
When processing large-scale real-time natural language texts, the prior art has slow processing speed, low accuracy, poor flexibility and insufficient semantic understanding, making it difficult to meet the needs of real-time and in-depth information extraction.
The text analysis model based on CNN-MLP, the strategy generation model of ISSA-MOGRPO and the information extraction model of BERT-LSTM-Attention-CRF are used to pre-process natural language text, text analysis, strategy generation and information extraction, and the processing strategy and model parameters are dynamically adjusted to improve extraction efficiency and accuracy.
It significantly shortens the processing time of information extraction, improves the accuracy and stability of information extraction, enhances the deep semantic understanding and flexibility of natural language text, and meets the real-time and efficient information extraction needs.
Smart Images

Figure CN120106080A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text information extraction, and in particular relates to an information extraction method and system based on natural language text. Background Art
[0002] Natural language text refers to text composed of human natural language (such as Chinese, English, etc.). It is one of the main ways for humans to communicate, express ideas and transmit information. Natural language text can include various forms, such as articles, conversations, books, web page content, emails, social media posts, etc. With the rapid development of the Internet, a large amount of natural language text is constantly generated. How to quickly and accurately extract key information from these unstructured natural language texts has become an important technical challenge.
[0003] Existing information extraction technology has the following defects:
[0004] 1) Slow processing speed: Existing methods often require a long processing time when processing large-scale real-time text data, which cannot meet real-time requirements. Complex algorithms and model structures lead to large consumption of computing resources, making it difficult to achieve efficient information extraction.
[0005] 2) Low accuracy: When facing complex and diverse natural language texts, the existing technology has a low accuracy rate in information extraction, which is prone to false positives and false negatives. It lacks effective text analysis and strategy generation mechanisms, resulting in unstable information extraction results.
[0006] 3) Poor flexibility: It relies on manually formulated rules, which are difficult to cover all the complex situations of natural language. Moreover, as the language develops and new scenarios emerge, the rules need to be constantly updated, resulting in high maintenance costs.
[0007] 4) Insufficient semantic understanding: Existing methods have deficiencies in semantic understanding, making it difficult to accurately capture the deep semantic information and association relationships in the text. The lack of semantic understanding causes the information extraction results to often remain at the surface level, making it difficult to mine valuable information. Summary of the invention
[0008] In order to solve the problems of slow processing speed, low accuracy, poor flexibility and insufficient semantic understanding in the prior art, the present invention aims to provide an information extraction method and system based on natural language text.
[0009] The technical solution adopted by the present invention is:
[0010] A method for extracting information based on natural language text comprises the following steps:
[0011] Collecting real-time natural language text, and preprocessing the real-time natural language text to obtain corresponding preprocessed real-time natural language text;
[0012] Use the text analysis model to perform text analysis on the preprocessed real-time natural language text to obtain the corresponding real-time text analysis results;
[0013] According to the real-time text analysis results, the strategy generation model is used to generate strategies for the pre-processed real-time natural language text to obtain the corresponding real-time information extraction strategy;
[0014] According to the real-time information extraction strategy, the information extraction model is used to extract information from the preprocessed real-time natural language text to obtain the corresponding real-time key information.
[0015] Furthermore, real-time natural language text is collected and preprocessed to obtain corresponding preprocessed real-time natural language text, including the following steps:
[0016] Collecting real-time natural language text, and performing format conversion on the real-time natural language text to obtain the real-time natural language text after format conversion;
[0017] The real-time natural language text after the format conversion is subjected to sentence segmentation processing and word segmentation processing in sequence to obtain a number of real-time text segments after word segmentation processing;
[0018] Perform part-of-speech tagging on each real-time text segment after word segmentation processing to obtain a number of real-time text segments after part-of-speech tagging;
[0019] Integrate several real-time text segments after part-of-speech tagging to obtain the corresponding pre-processed real-time natural language text.
[0020] Furthermore, the text analysis model is constructed based on the CNN-MLP algorithm, and the text analysis model includes a text feature extraction module constructed based on the CNN algorithm and a text analysis module constructed based on the MLP algorithm, which are connected in sequence.
[0021] Further, using the text analysis model, performing text analysis on the preprocessed real-time natural language text to obtain the corresponding real-time text analysis results includes the following steps:
[0022] Using the text feature extraction module of the text analysis model, extracting real-time text features of the preprocessed real-time natural language text;
[0023] According to the real-time text features, the text analysis module of the text analysis model is used to perform text analysis to obtain corresponding real-time text analysis results.
[0024] Furthermore, the strategy generation model is constructed based on the ISSA-MOGRPO algorithm, and the strategy generation model includes a network parameter optimization module constructed based on the ISSA algorithm and a strategy generation module constructed based on the MOGRPO algorithm. The strategy generation module includes an objective function set, an experience replay pool, a strategy network and an intelligent agent. The intelligent agent is respectively connected to the objective function set, the experience replay pool and the strategy network, and the network parameter optimization module is connected to the strategy network.
[0025] Furthermore, according to the real-time text analysis results, a strategy generation model is used to generate a strategy for the pre-processed real-time natural language text to obtain a corresponding real-time information extraction strategy, including the following steps:
[0026] Parse the real-time text analysis results to obtain the real-time mapping factor, and update the optimization target of the network parameter optimization module of the strategy generation model according to the real-time mapping factor to obtain an updated optimization target;
[0027] According to the updated optimization target, the policy network of the policy generation module of the policy generation model is optimized using the network parameter optimization module to obtain an optimized policy network;
[0028] A strategy generation module provided with an optimized strategy network is used to generate strategies for preprocessed real-time natural language text to obtain corresponding real-time information extraction strategies.
[0029] Furthermore, the real-time information extraction strategy includes real-time text processing decisions, real-time word embedding adjustment decisions, real-time attention weight adjustment decisions, and real-time sequence labeling adjustment decisions.
[0030] Furthermore, the information extraction model is constructed based on the BERT-LSTM-Attention-CRF algorithm, and the information extraction model includes a word embedding module constructed based on the BERT algorithm, a semantic feature extraction module constructed based on the LSTM algorithm, an attention weight module constructed based on the Attention mechanism, and an information extraction module constructed based on the CRF algorithm, which are connected in sequence.
[0031] Further, according to the real-time information extraction strategy, the information extraction model is used to extract information from the pre-processed real-time natural language text to obtain the corresponding real-time key information, including the following steps:
[0032] According to the real-time text processing decision in the real-time information extraction strategy, text processing is performed on the pre-processed real-time natural language text to obtain the processed real-time natural language text;
[0033] According to the real-time word embedding adjustment decision in the real-time information extraction strategy, the word embedding module of the information extraction model is adjusted to obtain an adjusted word embedding module;
[0034] According to the real-time attention weight adjustment decision in the real-time information extraction strategy, the attention weight module of the information extraction model is adjusted to obtain an adjusted attention weight module;
[0035] According to the real-time sequence labeling adjustment decision in the real-time information extraction strategy, the information extraction module of the information extraction model is adjusted to obtain an adjusted information extraction module;
[0036] Use the adjusted word embedding module of the information extraction model to embed the real-time natural language text after text processing to obtain a real-time word embedding vector;
[0037] Use the semantic feature extraction module of the information extraction model to extract real-time semantic features of the real-time word embedding vector;
[0038] Using the adjusted attention weight module of the information extraction model, several real-time feature components of the real-time semantic feature are weightedly fused to obtain a real-time weighted fused feature;
[0039] According to the real-time weighted fusion features, the adjusted information extraction module of the information extraction model is used to perform sequence annotation to obtain the corresponding real-time key information.
[0040] A natural language text-based information extraction system is used to implement an information extraction method. The system comprises a preprocessing unit, a text analysis unit, a strategy generation unit and an information extraction unit which are connected in sequence.
[0041] The beneficial effects of the present invention are:
[0042] The present invention provides a natural language text-based information extraction method and system, which preprocesses the natural language text and uses an artificial intelligence model to perform automated data processing, thereby significantly shortening the processing time of information extraction, meeting real-time requirements, effectively reducing computing resource consumption, and achieving efficient information extraction; dynamically generating information extraction strategies according to the text state of the natural language text, and using the information extraction model to extract information, thereby improving the accuracy of information extraction, reducing false positives and false negatives, and ensuring that the information extraction results are stable and reliable; using the artificial intelligence model to perform automated text analysis and information extraction, thereby reducing reliance on manual rules, thereby improving flexibility and scalability; and through the deep structure of the text analysis model, strengthening the semantic understanding ability, accurately capturing the deep semantic information and association relationships in the text, mining valuable information, and improving the depth and breadth of information extraction.
[0043] Other beneficial effects of the present invention will be further described in the specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flowchart of the information extraction method based on natural language text in the present invention.
[0045] Figure 2 It is a structural block diagram of the information extraction system based on natural language text in the present invention. DETAILED DESCRIPTION
[0046] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments.
[0047] Embodiment 1:
[0048] like Figure 1 As shown, this embodiment provides an information extraction method based on natural language text, comprising the following steps:
[0049] S1: collecting real-time natural language text and preprocessing the real-time natural language text to obtain corresponding preprocessed real-time natural language text, including the following steps:
[0050] S1-1: collecting real-time natural language text, and performing format conversion on the real-time natural language text to obtain the real-time natural language text after format conversion;
[0051] In this embodiment, the real-time natural language text is a news text: "Recently, the Chang'e-5 probe independently developed by my country was successfully launched, marking a major breakthrough in my country's lunar exploration. The probe is equipped with a variety of scientific instruments to collect soil and rock samples on the lunar surface and provide important data for scientific research. The successful implementation of this mission has not only enhanced my country's position in the international aerospace field, but also made important contributions to global lunar exploration research.";
[0052] S1-2: performing sentence segmentation and word segmentation processing on the real-time natural language text after format conversion in sequence to obtain a number of real-time text segments after word segmentation processing;
[0053] Clauses:
[0054] Sentence 1: "Recently, my country's independently developed Chang'e-5 probe was successfully launched, marking a major breakthrough in my country's lunar exploration.";
[0055] Sentence 2: "The probe carries a variety of scientific instruments designed to collect soil and rock samples from the lunar surface and provide important data for scientific research.";
[0056] Sentence 3: "The successful implementation of this mission not only enhanced my country's position in the international aerospace field, but also made important contributions to global lunar exploration research.";
[0057] Participle:
[0058] Sentence 1 word segmentation results: [“recently”, “our country”, “independently developed”, “of”, “Chang’e 5”, “probe”, “successfully”, “launched”, “,”, “marks that”, “our country”, “in”, “moon”, “exploration”, “field”, “achieved”, “significant”, “breakthrough”, “.”];
[0059] Sentence 2 word segmentation results: [“probe”, “carrying”, “carried”, “a variety of”, “scientific”, “instruments”, “,”, “aimed at”, “collecting”, “lunar”, “surface”, “soil”, “and”, “rock”, “samples”, “,”, “providing”, “important”, “data”, “for”, “scientific”, “research”, “.”];
[0060] Sentence 3 word segmentation results: [“this”, “mission”, “successfully”, “implemented”, “not only”, “enhanced”, “our”, “status”, “in”, “international”, “aerospace”, “field”, “but also”, “made”, “important”, “contributions”, “to”, “global”, “lunar”, “exploration”, “research”, “.”];
[0061] S1-3: performing part-of-speech tagging on each real-time text segment after word segmentation processing to obtain a number of real-time text segments after part-of-speech tagging;
[0062] Part-of-speech tagging:
[0063] Sentence 1 part-of-speech tagging results: [("recently","NT"),("our country","NR"),("independently developed","VV"),("of","DEG"),("Chang'e 5","NR"),("probe","NN"),("successfully","VV"),("launched","VV"),("","PU"),("marks","VV"),("our country","NR"),("in","P"),("moon","NN"),("detection","NN"),("field","NN"),("achieved","VV"),("had","AS"),("significant","JJ"),("breakthrough","NN"),("."","PU")];
[0064] Sentence 2 part-of-speech tagging results: [("probe","NN"),("carrying","VV"),("had","AS"),("various","CD"),("science","NN"),("instrument","NN"),("",","PU"),("aims to","VV"),("collect","VV"),("moon","NN"),("surface","NN"),("of","DEG"),("soil","NN"),("and","CC"),("rock","NN"),("sample","NN"),("",","PU"),("for","P"),("science","NN"),("research","NN"),("provide","VV"),("important","JJ"),("data","NN"),("."","PU")];
[0065] Sentence 3 part-of-speech tagging results: [("this time", "DT"), ("task", "NN"), ("of", "DEG"), ("successful", "NN"), ("implemented", "VV"), (",","PU"), ("not only", "AD"), ("improved", "VV"), ("had", "AS"), ("our country", "NR"), ("in", "P"), ("international", "NN"), ("aerospace", "NN"), ("field", "NN"),("of","DEG"),("status","NN"),("",","PU"),("also","AD"),("for","P"),("global","NN"),("moon","NN"),("exploration","NN"),("research","NN"),("made","VV"),("had","AS"),("important","JJ"),("contribution","NN"),("."","PU")];
[0066] S1-4: Integrate several real-time text segments after part-of-speech tagging to obtain corresponding real-time natural language text after preprocessing;
[0067] S2: Use the text analysis model to perform text analysis on the preprocessed real-time natural language text to obtain the corresponding real-time text analysis results;
[0068] The text analysis model is constructed based on a convolutional neural network (CNN)-multi-layer perceptron (MLP) algorithm, and the text analysis model includes a text feature extraction module constructed based on the CNN algorithm and a text analysis module constructed based on the MLP algorithm, which are connected in sequence;
[0069] The text feature extraction module uses the convolution layer of the convolutional neural network to extract local features of the preprocessed real-time natural language text. Through convolution kernels of different sizes, it captures local features of different scales such as characters, words, and phrases in the text. Through the pooling layer, the convolutional features are reduced in dimension to reduce the number of features while retaining the most important feature information, reducing the computational complexity, and mapping the extracted local features to a high-dimensional space to provide rich feature representations for the subsequent text analysis module. The text analysis module fuses the local features extracted by the CNN module to form a complete text feature representation. Through the nonlinear activation function of the multi-layer perceptron, the fused features are nonlinearly transformed to enhance the nonlinear modeling capability of the model. According to the specific task requirements, the output layer of the MLP is used to perform text classification or regression analysis to obtain the final text analysis results.
[0070] The method for constructing a text analysis model includes the following steps:
[0071] A-1: Collect a number of historical natural language texts, preprocess and add labels to the historical natural language texts, and obtain a number of preprocessed historical natural language texts with real text analysis results;
[0072] A-2: Divide a number of pre-processed historical natural language texts with real text analysis results into a first model training set and a first model test set in a ratio of 7:3;
[0073] A-3: Use the CNN-MLP algorithm to build an initial text analysis model, input the first model training set to optimize the initial text analysis model, obtain an optimized text analysis model, and generate several historical text analysis results;
[0074] A-4: Input the first model test set, perform model testing on the optimized text analysis model, obtain several historical text analysis results, and compare and count them with the corresponding real text analysis results to obtain the first model test accuracy;
[0075] A-5: If the test accuracy of the first model is greater than the first accuracy threshold, the final text analysis model is output; otherwise, optimization training continues;
[0076] Use the text analysis model to perform text analysis on the preprocessed real-time natural language text to obtain the corresponding real-time text analysis results, including the following steps:
[0077] S2-1: Use the text feature extraction module of the text analysis model to extract real-time text features of the preprocessed real-time natural language text;
[0078] S2-2: Perform text analysis using a text analysis module of a text analysis model based on real-time text features to obtain corresponding real-time text analysis results;
[0079] The real-time text analysis results include the real-time text type analysis results, real-time text defect recognition results, real-time text topic recognition results, and real-time text volume analysis results, etc.; among them, the real-time text type analysis results are the language types and fields of natural language texts, including Chinese prose, English poetry, Chinese news, etc., the real-time text defect recognition results are the format defects and content defects of natural language texts, etc., the real-time text topic recognition results are the topics of natural language texts, including life, technology, etc., and the real-time text volume analysis results are the volume of natural language texts;
[0080] S3: Based on the real-time text analysis results, the strategy generation model is used to generate strategies for the pre-processed real-time natural language text to obtain the corresponding real-time information extraction strategy;
[0081] The strategy generation model is constructed based on the Improved Sparrow Search Algorithm (ISSA)-Multi-Objective Group Relative Policy Optimization (MOGRPO) algorithm, and the strategy generation model includes a network parameter optimization module constructed based on the ISSA algorithm and a strategy generation module constructed based on the MOGRPO algorithm. The strategy generation module includes an objective function set, an experience replay pool, a strategy network and an intelligent agent. The intelligent agent is connected to the objective function set, the experience replay pool and the strategy network respectively, and the network parameter optimization module is connected to the strategy network.
[0082] The network parameter optimization module optimizes the policy network parameters under different optimization objectives through the ISSA algorithm to obtain the optimal initial network parameters of the policy network, thereby improving the accuracy of the policy generation model. The objective function set of the policy generation module can handle multiple conflicting optimization objectives, such as minimizing the prediction error value, minimizing the prediction accuracy, maximizing the prediction efficiency, maximizing the computing resource utilization, and minimizing the prediction cost, etc., and generate strategies that balance these objectives. The intelligent agent learns historical strategies through the experience replay pool and continuously optimizes its own policy generation capabilities. The intelligent agent controls the policy network based on the learned experience to generate more effective strategies. The design of the experience replay pool and the intelligent agent enables the model to continuously learn and optimize, thereby improving the quality of policy generation. The policy generation module adopts a group exploration method, which can avoid falling into the local optimal solution to a certain extent. The policy network outputs the distribution probability of actions under a given state. The policy generation module directly updates the policy network through the gradient, eliminating the Critic model in traditional reinforcement learning, making the algorithm structure more concise.
[0083] The method for constructing a strategy generation model includes the following steps:
[0084] B-1: Use the ISSA-MOGRPO algorithm to build an initial policy generation model; the initial policy generation model includes an initial network parameter optimization module and an initial policy generation module;
[0085] B-2: Based on several historical analysis results, the initial network parameter optimization module is optimized and trained to obtain the final network parameter optimization module;
[0086] B-3: Use the final network parameter optimization module to initialize the policy network of the initial policy generation module to obtain the initial policy network;
[0087] B-4: Set the objective function set and experience replay pool for the initial policy network, and set the action space and state space for the agent of the initial policy generation module;
[0088] B-5: The information extraction strategy generation problem is used as a simulation environment, and an optimized strategy generation module is obtained based on the initial strategy network and the intelligent agent with action space and state space.
[0089] B-6: Traverse all the objective functions in the objective function set, optimize and train the optimized strategy generation module according to several pre-processed historical natural language texts, obtain the final strategy generation module, and generate several historical strategy generation experiences;
[0090] B-7: Integrate the final network parameter optimization module and the final strategy generation module to obtain the final strategy generation model, and store several historical strategy generation experiences into the experience playback pool;
[0091] According to the real-time text analysis results, the strategy generation model is used to generate strategies for the pre-processed real-time natural language text to obtain the corresponding real-time information extraction strategy, including the following steps:
[0092] S3-1: parsing the real-time text analysis results to obtain the real-time mapping factor, and updating the optimization target of the network parameter optimization module of the strategy generation model according to the real-time mapping factor to obtain an updated optimization target;
[0093] The real-time mapping factors include prediction error value, prediction accuracy, prediction efficiency, computing resource utilization rate, and prediction cost, etc. When the real-time text type analysis result is a complex Chinese news type, the real-time mapping factor includes the prediction error value or the prediction accuracy rate. When the real-time text defect recognition result is a defect, the real-time mapping factor includes the prediction error value, the prediction accuracy rate, and the prediction efficiency. When the real-time text volume analysis result is a large volume, the real-time mapping factor includes the prediction efficiency, computing resource utilization rate, and the prediction cost. In this embodiment, the prediction error value is taken as an example to illustrate the workflow of network parameter optimization;
[0094] S3-2: According to the updated optimization target, the policy network of the policy generation module of the policy generation model is optimized using the network parameter optimization module to obtain an optimized policy network, including the following steps:
[0095] S3-2-1: According to the updated optimization target, set the fitness function of the network parameter optimization module, and set the algorithm parameters and maximum number of iterations of the ISSA algorithm;
[0096] The formula is:
[0097] f(X)=minMSN
[0098] Where f(X) is the fitness value of ISSA individual X; minMSN is the prediction error value of ISSA individual X; X is the ISSA individual reference parameter;
[0099] S3-2-2: Encode the initial network parameters of the strategy network of the strategy generation module of the strategy generation model into the solution vector of the ISSA individual, and initialize it using the Circle chaotic mapping sequence according to the solution vector to generate several initial solutions;
[0100] The formula is:
[0101]
[0102] Where X' c is the initial ISSA individual of the Circle chaos map; X c * is the randomly generated initial ISSA individual, i.e., the initial solution; mod(*) is the remainder function;
[0103] S3-2-3: Use the fitness function to obtain the initial fitness values of all initial ISSA populations, and sort the initial ISSA individuals according to the initial fitness values to obtain the initial discoverers, initial joiners, and initial predators;
[0104] S3-2-4: Use the network parameter optimization module to update the initial ISSA population to obtain an updated ISSA population; the updated ISSA population includes an updated discoverer, an updated joiner, and an updated predator;
[0105] The update formula of the discoverer is:
[0106]
[0107] In the formula, are the cth discoverer ISSA individuals of the t+1th and tth iterations respectively; iter max is the maximum number of iterations; ξ is a random number between 0 and 1; Q is a normally distributed random number; L is a 1×D matrix whose elements are all 1; R 2 is the warning value; ST is the safety threshold; c is the ISSA individual indicator; i is the update parameter;
[0108] The update formula for the joiner is:
[0109]
[0110] In the formula, are the cth joiner ISSA individuals in the t+1th and tth iterations respectively; The best position for the exposed person to occupy; is the current worst position; ξ is a random number between 0 and 1; L is a 1×D matrix whose elements are all 1 or -1; A + is the position update parameter; h is the total number of ISSA individuals;
[0111] The update formula of the predator is:
[0112]
[0113] In the formula, are the cth predator ISSA individuals of the t+1th and tth iterations respectively; δ is the step-size control parameter, and δ=a"·γ", a" is the convergence factor, γ" is a non-zero positive real number for step-size control; is the current best position; f c 、f g 、f w are the current, best and worst fitness of the ISSA individual respectively; γ is the minimum constant to prevent the denominator from being 0;
[0114]
[0115] In the formula, a" is the convergence factor; tanh(.) is the hyperbolic tangent function; t is the iteration indicator; t max is the maximum number of iterations; a max 、a min are the maximum and minimum values of the convergence factor, respectively; λ is the decreasing rate parameter, k" is the decreasing period parameter, λ = -2π, k" = π;
[0116] S3-2-5: Use the dynamic reverse learning algorithm to perform dynamic reverse learning on the updated ISSA population to generate a dynamic reverse ISSA population;
[0117] The formula is:
[0118]
[0119] In the formula, is the ISSA individual with dynamic reverse; γ* is the decreasing inertia coefficient; ub is the upper limit of the search space; lb is the lower limit of the search space; For the updated ISSA entity;
[0120] S3-2-6: Use the fitness function to obtain the updated fitness values of all ISSA individuals in the updated ISSA population and the dynamically reversed ISSA population, and take the ISSA individual with the smallest updated fitness value as the optimal solution;
[0121] S3-2-7: Decode the solution vector corresponding to the optimal solution to obtain the optimal initial network parameters of the policy network, and optimize the policy network of the policy generation module of the policy generation model according to the optimal initial network parameters to obtain an optimized policy network;
[0122] S3-3: using a strategy generation module provided with an optimized strategy network, performing strategy generation on the preprocessed real-time natural language text to obtain a corresponding real-time information extraction strategy;
[0123] The real-time information extraction strategy includes real-time text processing decisions, real-time word embedding adjustment decisions, real-time attention weight adjustment decisions, and real-time sequence labeling adjustment decisions;
[0124] Real-time text processing decisions include a series of "select", "skip", and "merge", including "select" ("Chang'e 5", "NR"), ("probe", "NN"), perform "merge", and obtain ("Chang'e 5 probe", "NR"), "skip" ("recently", "NT"), ("our country", "NR");
[0125] Real-time word embedding adjustment decision makes word embedding adjustments for keywords such as "Chang'e 5 probe" and "lunar surface";
[0126] Real-time attention weight adjustment decisions give higher weights to words or phrases such as "successful launch", "major breakthrough", and "scientific instrument";
[0127] The decision to adjust noun annotations marked "Chang'e 5" as a proper noun NN, and other nouns were annotated according to the context;
[0128] S4: According to the real-time information extraction strategy, the information extraction model is used to extract information from the pre-processed real-time natural language text to obtain the corresponding real-time key information;
[0129] The information extraction model is constructed based on the Bidirectional Encoder Representations from Transformers (BERT)-Long Short-Term Memory (LSTM)-Attention-Conditional Random Fields (CRF) algorithm, and the information extraction model includes a word embedding module constructed based on the BERT algorithm, a semantic feature extraction module constructed based on the LSTM algorithm, an attention weight module constructed based on the Attention mechanism, and an information extraction module constructed based on the CRF algorithm, which are connected in sequence;
[0130] The context-aware feature of the word embedding module enables word embedding to better reflect the contextual meaning of words, improve the model's understanding of complex semantics, and use pre-trained models to reduce the necessity of manual feature engineering and simplify the model building process. The semantic feature extraction module effectively captures long-term dependencies in the text, improves the ability to understand complex sentence structures, and generates richer semantic representations by integrating contextual information to provide more valuable information for subsequent modules. The attention weight module uses the attention mechanism to enable the model to focus more on key information, improve the accuracy of information extraction, and dynamically adjust feature representations so that the model can adapt to the needs of different texts and tasks. The information extraction module can effectively model the dependencies between labels, improve the accuracy of sequence annotation, and improve the overall performance and effect of the information extraction model through accurate sequence annotation.
[0131] The method for constructing an information extraction model includes the following steps:
[0132] C-1: Labeling a number of preprocessed historical natural language texts to obtain a number of preprocessed historical natural language texts with real key information;
[0133] C-2: Divide a number of pre-processed historical natural language texts with real key information into a second model training set and a second model test set in a ratio of 7:3;
[0134] C-3: Use the BERT-LSTM-Attention-CRF algorithm to build an initial information extraction model, input the second model training set to optimize the initial information extraction model, obtain the optimized information extraction model, and generate some historical key information;
[0135] C-4: Input the second model test set, perform model testing on the optimized information extraction model, obtain some historical key information, and compare and count the corresponding real key information to obtain the second model test accuracy;
[0136] C-5: If the test accuracy of the second model is greater than the second accuracy threshold, the final information extraction model is output, otherwise, optimization training continues;
[0137] According to the real-time information extraction strategy, the information extraction model is used to extract information from the pre-processed real-time natural language text to obtain the corresponding real-time key information, including the following steps:
[0138] S4-1: performing text processing on the preprocessed real-time natural language text according to the real-time text processing decision in the real-time information extraction strategy to obtain the processed real-time natural language text;
[0139] S4-2: According to the real-time word embedding adjustment decision in the real-time information extraction strategy, the word embedding module of the information extraction model is adjusted to obtain an adjusted word embedding module;
[0140] S4-3: According to the real-time attention weight adjustment decision in the real-time information extraction strategy, the attention weight module of the information extraction model is adjusted to obtain an adjusted attention weight module;
[0141] S4-4: adjusting the information extraction module of the information extraction model according to the real-time sequence labeling adjustment decision in the real-time information extraction strategy to obtain an adjusted information extraction module;
[0142] S4-5: Use the adjusted word embedding module of the information extraction model to embed the real-time natural language text after text processing to obtain a real-time word embedding vector;
[0143] S4-6: Use the semantic feature extraction module of the information extraction model to extract real-time semantic features of the real-time word embedding vector;
[0144] S4-7: Using the adjusted attention weight module of the information extraction model, weighted fusion is performed on several real-time feature components of the real-time semantic feature to obtain a real-time weighted fusion feature;
[0145] S4-8: According to the real-time weighted fusion features, the adjusted information extraction module of the information extraction model is used to perform sequence annotation to obtain the corresponding real-time key information;
[0146] The real-time key information is "Chang'e-5 probe", "successful launch", "lunar exploration field", "major breakthrough", "scientific instruments", "lunar surface", "soil and rock samples", "scientific research", "international space field", and "important contribution".
[0147] Embodiment 2:
[0148] like Figure 2 As shown, this embodiment provides an information extraction system based on natural language text, which is used to implement the information extraction method. The system includes a preprocessing unit, a text analysis unit, a strategy generation unit and an information extraction unit connected in sequence;
[0149] A preprocessing unit, used for collecting real-time natural language text and preprocessing the real-time natural language text to obtain corresponding preprocessed real-time natural language text;
[0150] A text analysis unit, used to use a text analysis model to perform text analysis on the preprocessed real-time natural language text to obtain corresponding real-time text analysis results;
[0151] A strategy generation unit, used to generate a strategy for the pre-processed real-time natural language text using a strategy generation model according to the real-time text analysis result, and obtain a corresponding real-time information extraction strategy;
[0152] The information extraction unit is used to extract information from the pre-processed real-time natural language text using an information extraction model according to a real-time information extraction strategy to obtain corresponding real-time key information.
[0153] The present invention provides a natural language text-based information extraction method and system, which preprocesses the natural language text and uses an artificial intelligence model to perform automated data processing, thereby significantly shortening the processing time of information extraction, meeting real-time requirements, effectively reducing computing resource consumption, and achieving efficient information extraction; dynamically generating information extraction strategies according to the text state of the natural language text, and using the information extraction model to extract information, thereby improving the accuracy of information extraction, reducing false positives and false negatives, and ensuring that the information extraction results are stable and reliable; using the artificial intelligence model to perform automated text analysis and information extraction, thereby reducing reliance on manual rules, thereby improving flexibility and scalability; and through the deep structure of the text analysis model, strengthening the semantic understanding ability, accurately capturing the deep semantic information and association relationships in the text, mining valuable information, and improving the depth and breadth of information extraction.
[0154] The present invention is not limited to the above optional implementations, and anyone can derive other various forms of products under the enlightenment of the present invention. The above specific implementations should not be understood as limiting the scope of protection of the present invention. The scope of protection of the present invention should be based on the definition in the claims, and the description can be used to interpret the claims.
Claims
1. A method for extracting information based on natural language text, characterized in that: The steps include: Collecting real-time natural language text, and preprocessing the real-time natural language text to obtain corresponding preprocessed real-time natural language text; Use the text analysis model to perform text analysis on the preprocessed real-time natural language text to obtain the corresponding real-time text analysis results; According to the real-time text analysis results, the strategy generation model is used to generate strategies for the pre-processed real-time natural language text to obtain the corresponding real-time information extraction strategy; According to the real-time information extraction strategy, the information extraction model is used to extract information from the preprocessed real-time natural language text to obtain the corresponding real-time key information.
2. The method for extracting information based on natural language text according to claim 1, characterized in that: Collecting real-time natural language text and preprocessing the real-time natural language text to obtain corresponding preprocessed real-time natural language text includes the following steps: Collecting real-time natural language text, and performing format conversion on the real-time natural language text to obtain the real-time natural language text after format conversion; The real-time natural language text after the format conversion is subjected to sentence segmentation processing and word segmentation processing in sequence to obtain a number of real-time text segments after word segmentation processing; Perform part-of-speech tagging on each real-time text segment after word segmentation processing to obtain a number of real-time text segments after part-of-speech tagging; Integrate several real-time text segments after part-of-speech tagging to obtain the corresponding pre-processed real-time natural language text.
3. The method for extracting information based on natural language text according to claim 2, characterized in that: The text analysis model is constructed based on the CNN-MLP algorithm, and the text analysis model includes a text feature extraction module constructed based on the CNN algorithm and a text analysis module constructed based on the MLP algorithm, which are connected in sequence.
4. The method for extracting information based on natural language text according to claim 3, characterized in that: Use the text analysis model to perform text analysis on the preprocessed real-time natural language text to obtain the corresponding real-time text analysis results, including the following steps: Using the text feature extraction module of the text analysis model, extracting real-time text features of the preprocessed real-time natural language text; According to the real-time text features, the text analysis module of the text analysis model is used to perform text analysis to obtain corresponding real-time text analysis results.
5. The method for extracting information based on natural language text according to claim 4, characterized in that: The strategy generation model is constructed based on the ISSA-MOGRPO algorithm, and the strategy generation model includes a network parameter optimization module constructed based on the ISSA algorithm and a strategy generation module constructed based on the MOGRPO algorithm. The strategy generation module includes an objective function set, an experience replay pool, a strategy network and an intelligent agent. The intelligent agent is respectively connected to the objective function set, the experience replay pool and the strategy network, and the network parameter optimization module is connected to the strategy network.
6. The method for extracting information based on natural language text according to claim 5, characterized in that: According to the real-time text analysis results, the strategy generation model is used to generate strategies for the pre-processed real-time natural language text to obtain the corresponding real-time information extraction strategy, including the following steps: Parse the real-time text analysis results to obtain the real-time mapping factor, and update the optimization target of the network parameter optimization module of the strategy generation model according to the real-time mapping factor to obtain an updated optimization target; According to the updated optimization target, the policy network of the policy generation module of the policy generation model is optimized using the network parameter optimization module to obtain an optimized policy network; A strategy generation module provided with an optimized strategy network is used to generate strategies for preprocessed real-time natural language text to obtain corresponding real-time information extraction strategies.
7. The method for extracting information based on natural language text according to claim 6, characterized in that: The real-time information extraction strategy includes real-time text processing decisions, real-time word embedding adjustment decisions, real-time attention weight adjustment decisions, and real-time sequence labeling adjustment decisions.
8. The method for extracting information based on natural language text according to claim 7, characterized in that: The information extraction model is constructed based on the BERT-LSTM-Attention-CRF algorithm, and the information extraction model includes a word embedding module constructed based on the BERT algorithm, a semantic feature extraction module constructed based on the LSTM algorithm, an attention weight module constructed based on the Attention mechanism, and an information extraction module constructed based on the CRF algorithm, which are connected in sequence.
9. The method for extracting information based on natural language text according to claim 8, characterized in that: According to the real-time information extraction strategy, the information extraction model is used to extract information from the pre-processed real-time natural language text to obtain the corresponding real-time key information, including the following steps: According to the real-time text processing decision in the real-time information extraction strategy, text processing is performed on the pre-processed real-time natural language text to obtain the processed real-time natural language text; According to the real-time word embedding adjustment decision in the real-time information extraction strategy, the word embedding module of the information extraction model is adjusted to obtain an adjusted word embedding module; According to the real-time attention weight adjustment decision in the real-time information extraction strategy, the attention weight module of the information extraction model is adjusted to obtain an adjusted attention weight module; According to the real-time sequence labeling adjustment decision in the real-time information extraction strategy, the information extraction module of the information extraction model is adjusted to obtain an adjusted information extraction module; Use the adjusted word embedding module of the information extraction model to embed the real-time natural language text after text processing to obtain a real-time word embedding vector; Use the semantic feature extraction module of the information extraction model to extract real-time semantic features of the real-time word embedding vector; Using the adjusted attention weight module of the information extraction model, several real-time feature components of the real-time semantic feature are weightedly fused to obtain a real-time weighted fused feature; According to the real-time weighted fusion features, the adjusted information extraction module of the information extraction model is used to perform sequence annotation to obtain the corresponding real-time key information.
10. An information extraction system based on natural language text, used to implement the information extraction method according to any one of claims 1 to 9, characterized in that: The system comprises a pre-processing unit, a text analysis unit, a strategy generation unit and an information extraction unit which are connected in sequence.
Citation Information
Patent Citations
Natural language processing method and system based on machine learning
CN117493491A
Internet of Things risk intrusion identification method based on artificial intelligence
CN118133146A
Hydropower remote centralized control safety anti-error method and system
CN118396260A
News interpretation method and system based on natural language processing
CN119025670A
Method and system for synthesizing reasoning data
CN119476479A