Data extraction method and system for entity evaluation

Through the large language model based on the Transformer model and the instruction fine-tuning of the LoRA module, combined with the Doc2Vec model, the problems of low redundant information processing efficiency and low short text matching accuracy in big data are solved, and efficient and accurate data extraction is achieved, which is especially suitable for entity evaluation tasks.

WO2025148607A1PCT designated stage expired Publication Date: 2025-07-17YINGTOU INFORMATION & TECH SHANGHAI CO LTD

Patent Information

Application Number
PCT/CN2024/138864
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2024-12-12
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The prior art has problems such as low redundant information processing efficiency, low short text matching accuracy and difficulty in extracting unstructured text information when processing big data, especially in the lack of high accuracy and fast data extraction methods in entity evaluation.

Method used

The large language model based on the Transformer model architecture is used for fine-tuning instructions, combining the LoRA module and the Doc2Vec model, through cosine similarity and text length judgment, the text of the same event is merged to achieve efficient and accurate data extraction.

Benefits of technology

Improves the accuracy and efficiency of data extraction, especially showing better migration capabilities on short text processing and untrained metric points, while reducing training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024138864_17072025_PF_FP_ABST
    Figure CN2024138864_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a data extraction method for entity evaluation. The method is used for extracting target data on the basis of a predefined indicator, and comprises: an underlying data processing step: at least comprising using a first language model to identify identical events in underlying data; and an indicator data extraction step: at least comprising, on the basis of the predefined indicator, using the first language model to extract target data from among the underlying data, wherein a base large model of the first language model comprises a large language model based on a Transformer model architecture, and the first language model is obtained by training on the basis of the base large model. The technology of the present invention can significantly improve the accuracy of data extraction and reduce training costs.
Need to check novelty before this filing date? Find Prior Art

Description

A data extraction method and system for entity evaluation Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a method and system for data extraction. Background Art

[0002] In the era of big data, the speed and frequency of data generation and collection are increasing significantly, making the need for big data processing an inevitable trend. Although different industries involve different types of underlying data, extracting relevant information from this underlying data is considered a key step in subsequent analysis and processing to achieve different business goals.

[0003] Entity evaluation generally refers to the assessment, scoring, or evaluation of a subject / entity in specific aspects or dimensions. Although the subject / entity, evaluation criteria, and evaluation methods can be arbitrary, no matter how the evaluation is conducted, it involves the step of extracting relevant information from the underlying data. For example, an ESG (environmental, social, and corporate governance) evaluation requires assessing an entity's performance in three aspects: environmental, social, and corporate governance, and potential going concern risks. It is easy to understand that the large amount of underlying data involved in this evaluation process contains unstructured characteristics and irrelevant information, which needs to be extracted and processed before it can be used for actual ESG evaluation and analysis.

[0004] However, when using big data technology, those skilled in the art usually face the following major problems.

[0005] First, because raw data often contains a large amount of repeated information, this information is redundant for data analysis. This includes information with highly similar content and content that references each other. Eliminating this repeated information can significantly improve the efficiency of data processing. One of the traditional methods is to identify similar parts by comparing the longest identical substrings between two news articles. However, the time complexity of this method is quadratic time complexity, that is, O(n 2 ), therefore, the efficiency is low when processing large amounts of text data. In addition, due to the presence of irrelevant information, such as advertisements, etc., it will affect the accuracy of text (especially short text) processing, and there is no better solution to this in the prior art.

[0006] Currently, pre-trained models dominate various natural language processing tasks. These models first require pre-training on large-scale text data to form a language model, and then are fine-tuned on data from specific downstream tasks to meet specific requirements. BERT (Bidirectional Encoder Representations from Transformers) is currently the most popular and best-performing pre-trained model, but the large datasets required for training result in high costs for training data processing. Existing short text semantic matching methods are primarily based on deep semantic matching models using deep neural networks. These models abstractly represent short texts as high-dimensional vectors and use specific matching algorithms to calculate similarities between short texts. While they can capture the deep semantic features of short texts, they perform poorly for shallow features, such as word-level features and shallow semantic features within sentences. Furthermore, since short texts typically have short word lengths and contain irrelevant textual content (such as advertisements), neural network models are prone to overfitting when extracting features, resulting in low short text matching accuracy.

[0007] Secondly, since the original data usually contains a large amount of unstructured text information, extracting data related to the task objectives is an important and complex task. Traditional machine reading comprehension-based methods have the limitation of extracting the beginning and end, and perform poorly on untrained task data. The machine reading comprehension task or algorithm MRC (Machine Reading Comprehension) is a task (or algorithm) that requires the model to read a text and then answer questions. The current mainstream MRC method is that after given the question and content text, the MRC model extracts the answer by predicting the start and end character positions of the answer in the content text. The defects of this algorithm are: (1) the training of the MRC model requires a large amount of data like other data-driven algorithms; (2) this method is likely to give an answer even for content text that has no answer, and it does not handle content text without an answer well.

[0008] Therefore, how to address the shortcomings of the above-mentioned data extraction technologies and provide a technology for extracting structured data with high accuracy and speed is an urgent problem to be solved in this field. Summary of the Invention

[0009] The present invention provides a data extraction method and system, the purpose of which is to improve the accuracy and speed of data extraction technology so that it can better adapt to and complete entity evaluation tasks.

[0010] To achieve the above-mentioned objectives, the present invention adopts a technical solution: a data extraction method for entity evaluation, the method being used to extract target data based on predefined indicators, comprising: an underlying data processing step, at least comprising using a first language model to identify identical events in the underlying data; an indicator data extraction step, at least comprising using the first language model to extract target data from the underlying data based on the predefined indicators; wherein the base large model of the first language model comprises a large language model based on the Transformer model architecture, and the first language model is trained based on the base large model; wherein the training comprises at least two rounds of optimization fine-tuning based on a first training dataset; the first training dataset comprises first task data and second task data, the first task data comprising 40%-50% NER task data, 40%-50% classification task data, and / or 0-20% other task data; and the second task data comprises at least 200 indicator points of different types, and the number of samples of each type of indicator point is at least between 10 and 100.

[0011] In a preferred embodiment, the first training data set includes third task data, and the third task data includes at least 300 to 500 task samples for identifying the same event and / or at least 2,000 to 4,000 task samples for extracting quantitative data; wherein, the ratio of the number of samples of the same event and different events in the task samples for identifying the same event is approximately 1:1 to 3:1.

[0012] In a preferred embodiment, the first task data includes at least 20,000 to 40,000 task samples; wherein the 0-20% other task data includes relationship extraction tasks, semantic role labeling tasks and / or event extraction tasks.

[0013] In a preferred embodiment, the optimization fine-tuning includes using a parameter adjustment method to gradually adjust the parameters of the base model according to the gradient of the loss function; wherein, the parameter adjustment method includes a back propagation method, the loss function includes a cross entropy loss function, and the parameters include creation using the LoRA method.

[0014] In a preferred embodiment, the method of using LoRA includes adding a LoRA module to the key components of all attention mechanism layers of the base model; wherein the key components include query (q), key (k), value (v) and projection linear layer (projection); the rank (r) of the low-rank matrix decomposition of the LoRA module is set to 4 to 64, and the regularization parameter (alpha) is set according to the rank (r).

[0015] In a preferred embodiment, the use of the first language model to identify identical events in the underlying data includes: cyclically determining whether any two texts contain the same events, and merging texts containing the same events; wherein, determining whether any two texts contain the same events includes: a) vectorizing the any two texts and calculating the cosine similarity of the any two texts; b) if the cosine similarity is less than a first threshold, determining that the events contained in the any two texts are different; if the cosine similarity is greater than or equal to the first threshold, and the length of at least one text is greater than or equal to a second threshold, determining that the events contained in the any two texts are the same; otherwise, executing step c); c) using the first language model to determine whether the events contained in the any two texts are the same; wherein, when the judgment results of any of the above steps a)-c) are the same, executing the step of merging texts containing the same events; the first threshold is set to 0.6-0.8, and the second threshold is set to 200-400 characters.

[0016] In a preferred embodiment, the vectorization includes using a Doc2Vec model, which is obtained through the following training method: performing Chinese word segmentation processing on the vectorized training data set; setting the parameters of the Doc2Vec model, including setting the vector dimension to 100-300, the window length to 2-10, and the dictionary size; and performing at least 10 rounds of iterative training based on the vectorized training data set.

[0017] In a preferred embodiment, the step of merging texts containing the same event includes: creating a set for each of the texts using a disjoint set; merging the sets of texts containing the same event through a Union operation based on the judgment result; and when the number of texts in a set is greater than 1, selecting one of the texts as the representative text of the set.

[0018] In a preferred embodiment, the Transformer-based model architecture is selected from the Qwen-7B model or the Llama model.

[0019] In order to achieve the above-mentioned purpose, another technical solution adopted by the present invention is: a data extraction system for entity evaluation, wherein the system is used to implement any of the methods described above.

[0020] Compared with the existing technology, the advantages of the present invention are: (1) it can efficiently and accurately detect highly similar news in a large number of news, thereby improving the accuracy and efficiency of information screening; (2) the present invention has high accuracy and robustness when processing short texts; (3) the present invention has higher accuracy for indicator point extraction tasks, especially showing better migration ability on untrained indicator points; (4) under the premise of unchanged or better technical effects, the implementation / training cost of the present invention is lower; (5) the present invention can simultaneously complete the extraction of various types of data such as qualitative and quantitative data. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] FIG1 is a schematic diagram of a data extraction process according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The following is a clear and complete description of the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0024] As used herein, "at least one" means one or more than one.

[0025] In this document, terms such as "include", "comprising", "containing" and "having" are open-ended and indicate the presence of the described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0026] In this document, the terms "first," "second," "third," "fourth," etc. (if any) are used to distinguish between similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances such that the embodiments described herein can be practiced in orders other than those illustrated or described herein.

[0027] In this article, "indicator" refers to a quantitative or qualitative standard used to describe a measurement, evaluation or metric of a certain state, performance or result.

[0028] In this article, "entity" refers to an individual, thing, organization or concept with specific attributes that exists in the real world or in a conceptual space. They may be objects with unique identifiers, attributes and characteristics that can be identified, described, evaluated and classified; specifically, the scope of entities may include an individual, a company, an organization, an investment portfolio or other economic entities.

[0029] The present invention provides a data extraction technology that can be used to complete various types of downstream tasks. The introduction of this technology will be described in detail below, and its potential benefits in various application scenarios will be demonstrated. Data extraction technology is a key tool that aims to extract key information from raw data and convert it into structured or semi-structured data for subsequent analysis and application. The data extraction technology of the present invention is not only highly flexible, but also can adapt to different data types and downstream task requirements. In particular, the present invention has excellent results for completing entity evaluation type tasks.

[0030] Specifically, the present invention provides a data extraction method for entity evaluation, which is used to extract data based on predefined indicators, including: an underlying data processing step, which at least includes using a first language model to merge identical events in the underlying data; an indicator data extraction step, which at least includes using a first language model to extract indicator data from the underlying data based on the predefined indicators; wherein the first language model is obtained by fine-tuning instructions based on a large language model based on the Transformer model architecture.

[0031] On the one hand, the method of the present invention can enable the language model to extract data according to predefined indicators, thereby meeting the data requirements of different entity evaluation tasks. On the other hand, the algorithm of the present invention has undergone targeted fine-tuning to improve the accuracy of data extraction, which helps to improve the quality and speed of data analysis. In particular, the present invention unexpectedly discovered that the language model based on the large language model based on the Transformer model architecture, especially the language model based on the Llama model and Qwen-7B, has unique advantages in short text processing and information extraction.

[0032] FIG1 shows a schematic diagram of a data extraction process according to an embodiment of the present invention.

[0033] As shown in the figure, this embodiment provides a data extraction method for entity evaluation, which is used to extract data based on predefined indicators, including: an underlying data processing step, which at least includes using a first language model to merge identical events in the underlying data; an indicator data extraction step, which at least includes using a first language model to extract indicator data from the underlying data based on the predefined indicators; wherein the first language model is obtained by fine-tuning instructions based on a large language model based on the Transformer model architecture.

[0034] Specifically, the method provided in this embodiment includes an underlying data processing step, which is the basis for data extraction. In an optional embodiment, the Transformer model architecture of this step adopts the Qwen-7B model or the Llama model. The key task of the underlying data processing step is to merge the same events in the underlying data, thereby reducing data redundancy and making subsequent processing more efficient. In addition, the method provided in this embodiment also involves a step of extracting indicator data using a first language model, which is a key link in extracting key indicator data from the underlying data. This step can accurately extract the required indicator information from the underlying data based on predefined indicators.

[0035] More specifically, the first language model of this embodiment is obtained by optimizing and fine-tuning the base large model based on at least two rounds of the first training data set. The first training data set may be a training data set containing natural language processing (NLP) tasks. In this embodiment, the task data of the first training data set includes first task data and second task data, the first task data including 40%-50% NER (Named Entity Recognition) task data, 40%-50% classification task data, and / or 0-20% other task data; the second task data includes at least 200 different types of indicator points, and the number of samples of each type of indicator point is at least between 10-100. In an optional embodiment, the first task data includes 20,000 to 40,000 task samples; wherein the 0-20% other task data includes relationship extraction tasks, semantic role labeling tasks, and / or event extraction tasks. Such a training data set setting can minimize the number of samples in the training data set while ensuring the output effect of the large model.

[0036] The present invention sets a larger proportion for NER task data and classification task data because the NER task focuses on identifying specific entities, such as names of people, places, etc., while the classification task focuses on classifying the entire event or situation. On the one hand, NER tasks and classification tasks are usually more relevant to the goal of identifying the same event. Using these two types of tasks as the main training data can improve the performance of the model on the target task; and compared with a single training data, the diversity of training tasks can enable the model to better understand contextual information and improve its generalization ability. On the other hand, in some cases, in order to improve the diversity and robustness of the model and prevent the model from overfitting, it is possible to consider adding other tasks, but this depends on the correlation between the tasks and the scope of knowledge that the model needs to cover. For example, in this embodiment, 0-20% of other tasks include but are not limited to relationship extraction tasks, semantic role labeling tasks, event extraction tasks, etc.

[0037] It should be understood that the various types of task data mentioned above or below are not mutually exclusive. They can be independent, overlapping, or related. For example, for the sample text "The new coronavirus is rampant, with more than X million people infected worldwide. The epidemic has seriously affected various industries, such as the aviation industry, tourism, and healthcare", in some cases, this task data can belong to both NER task data (such as the identification of epidemics, industry names, etc.) and classification task data (such as the classification of the impact of the epidemic), and the task data also includes indicator points (such as the extraction of industry indicators).

[0038] This embodiment performs at least two rounds of optimization and fine-tuning based on the first training data set. An epoch represents a complete forward propagation and backpropagation of the entire training data set. In an optional embodiment, optimization and fine-tuning can be performed for multiple rounds. In this embodiment, the first round is used to adjust the initial weights, that is, when the large model starts to use a new data set for training, the weights learned before are fine-tuned according to the characteristics of the new data in order to adapt to the new task or data set; the first round and subsequent rounds are used for gradual optimization of the large model. It can be understood that as the training progresses, the model gradually understands and adapts to the new data patterns and features. Each round allows the model to better fit the data, improve performance, and reduce possible overfitting or underfitting.

[0039] In an optional embodiment, the first training data set also includes third task data, and the third task data includes at least 300 to 500 task samples for identifying the same event and / or at least 2,000 to 4,000 task samples for extracting quantitative data; wherein the ratio of the number of samples of the same event and different events in the task samples for identifying the same event is approximately 1:1 to 3:1. Specifically, the task samples for identifying the same event are used to improve the effect of the underlying data processing steps; the task samples for extracting quantitative data can make the large model more accurate in extracting quantitative data, so that the model can accurately complete the extraction of qualitative and quantitative data. It should be understood that, generally, the more samples of task data there are, the higher the accuracy of the model in performing specific tasks, but the training cost will also increase accordingly; the present invention does not impose an upper limit on the number of samples of task data, and the lower limit of the number of samples is the minimum requirement to ensure the effect of the present invention. In addition, considering that the quality of task data may fluctuate, the lower limit of the number of samples is set to a range accordingly.

[0040] Furthermore, the optimization and fine-tuning of this embodiment includes using a parameter adjustment method to gradually adjust the parameters of the base model based on the gradient of the loss function; wherein the parameter adjustment method includes a backpropagation method, the loss function includes a cross-entropy loss function, and the parameters are created using the LoRA (Low-Rank Adaptation) method. In an optional embodiment, the LoRA method includes adding a LoRA module to the key components of all attention mechanism layers of the base model; wherein the key components include query (q), key (k), value (v), and projection linear layer (projection); the rank (r) of the low-rank matrix factorization of the LoRA module is set to 4 to 64, and the regularization parameter (alpha) is set based on the rank (r). The rank (r) is generally related to the regularization effect, and a smaller rank (r) value generally results in a stronger regularization effect, which helps to avoid model overfitting; the setting of the regularization parameter (alpha) requires balancing the complexity and generalization ability of the model. Increasing the regularization parameter (alpha) can help reduce the risk of overfitting, but a trade-off needs to be made based on the specific situation to avoid model underfitting. In this embodiment, the regularization parameter (alpha) is set to 0.1 to 16 according to the rank (r). It should be understood that the setting of these two parameters is very important for model training, and the setting values ​​will vary for different models and training data.

[0041] An embodiment of the present invention uses a first language model to merge identical events in the underlying data. This embodiment cyclically determines whether any two texts in the underlying data involve identical events until all texts containing identical events are recognized.

[0042] As shown in the figure, this embodiment determines whether the events involved in any two texts are the same through the following process, including:

[0043] a) vectorizing the two arbitrary texts and calculating the cosine similarity of the two arbitrary texts;

[0044] b) if the cosine similarity is less than a first threshold, then the arbitrary two texts are determined to be different; if the cosine similarity is greater than or equal to the first threshold, and the length of at least one text is greater than or equal to a second threshold, then the arbitrary two texts are determined to contain the same event; otherwise, executing step c);

[0045] c) using the first language model to determine whether the events contained in the arbitrary two texts are the same.

[0046] In an optional embodiment, Doc2Vec technology can be optionally used to convert any two texts into vector representations. Specifically, the semantic information of the text is encoded as vector features, thereby achieving a quantitative representation of the text. In an optional embodiment, the Doc2Vec model is obtained by the following training method: performing Chinese word segmentation processing on the vectorized training dataset; setting the parameters of the Doc2Vec model, including setting the vector dimension to 100-300, the window length to 2-10, and setting the dictionary size; and performing at least 10 rounds of iterative training based on the vectorized training dataset. It can be understood that the vector dimension determines the dimensionality of the embedding space of each document or word, affecting the expressiveness and feature dimension of the model; the window length defines how many words are considered in the context. A larger window can capture a wider range of contextual information, but may also introduce more noise; the dictionary size defines the number of different words seen by the model during training, and the appropriate dictionary size is usually selected based on the lexical richness of the dataset. In this embodiment, the dictionary size will be set according to the vectorized training dataset of the Doc2Vec model. In other optional embodiments, other text vectorization methods may be used, including but not limited to Bag of Words (BoW), TF-IDF (Term Frequency-Inverse Document Frequency), Sentence-BERT, etc.

[0047] This embodiment uses cosine similarity to preliminarily determine the similarity between two texts. The cosine similarity is calculated as follows: Cosine Similarity(A, B)=A·B / ||A||||B||.

[0048] Where A and B represent the vectors of two texts, · represents the dot product of the vectors, and ||A|| and ||B|| represent the norms of the two vectors. Cosine similarity can effectively measure the semantic similarity between texts and is particularly effective for long texts. In an optional embodiment, the Faiss Similarity Search tool can be used to quickly calculate the cosine similarity between a large number of news vectors. In an optional embodiment, the first threshold for determining cosine similarity can be set to between 0.6 and 0.8.

[0049] Furthermore, this embodiment divides similarity determination into two scenarios based on text length. The aforementioned cosine similarity is used to determine the similarity of long texts. If cosine similarity determines similarity but the text length does not meet the set standard, further similarity determination will be performed using the first language model. In an optional embodiment, a second threshold is used as the criterion for determining text length and can be optionally set between 200 and 400 characters. As can be seen, this embodiment combines cosine similarity with the first language model to simultaneously improve the accuracy and computational efficiency of text similarity calculations. On the one hand, cosine similarity is particularly suitable for long texts and does not require excessive computing resources. On the other hand, the contextual information of short texts may not be sufficient for accurate similarity calculations. Incorporating the first language model can help capture more in-depth semantic features, thereby improving the accuracy of short text similarity. For most entity evaluation tasks, the underlying data often involves both long and short texts. In this case, combining cosine similarity with the first language model can simultaneously meet text processing requirements without sacrificing accuracy. Furthermore, this embodiment eliminates the need for introducing additional large-scale models and achieves high accuracy with minimal resource consumption.

[0050] One embodiment of the present invention uses a loop-based traversal approach to determine similarity between texts in the underlying data. Specifically, when two texts are detected to involve the same event, the system merges the texts containing the same event and continues iterating until all texts containing the same event are identified.

[0051] In an optional embodiment, the step of merging texts containing the same event includes merging texts related to the same event using a disjoint set data structure, specifically including: creating a set for each of the texts using a disjoint set; merging the sets of texts containing the same event through a Union operation based on a similarity judgment result; and when the number of texts in a set is greater than 1, selecting one of the texts as the representative text of the set.

[0052] In an optional embodiment, each text set is determined using the process described in this application for determining whether any two texts involve the same event. The text set is traversed, and texts involving the same event are merged into the same set. In an optional embodiment, this merging can be achieved through a Union operation in the data structure. When two texts are merged, they will belong to the same set, and one of the texts will become the representative. In an optional embodiment, for each set, a text can be randomly selected as the representative, or the text that was added to the set earliest can be selected as the representative, or other methods can be used to select the representative. When all texts have been traversed and merged, different texts will be distributed in different sets, each with a representative. At this point, the underlying data has been divided into non-overlapping subsets based on similarity. The events of the texts in each subset are the same, and the representative text represents the entire subset. In an optional embodiment, when it is necessary to check whether the events of a certain text are identical, the text set can be searched to see if there are other text sets in the text set. If so, it means that the data set already contains text with the same event as the text; otherwise, the event of the text is new.

[0053] One embodiment of the present invention provides a data extraction system for entity evaluation, which can implement the method described in any embodiment of the present invention.

[0054] The above embodiments describe the present invention or a certain aspect of the present invention, and these embodiments can be combined arbitrarily. On this basis, in order to more clearly reflect the process of the present invention, the following embodiments are described in conjunction with specific cases.

[0055] Example 1

[0056] This example uses the assessment of corporate risk based on news information as an example. As can be seen, in this example, the entity is the company to be evaluated, the underlying data is news data, and indicators can be set based on the actual risk focus. For example, for the indicator "pollutant emission reduction measures," the indicator definition can be set as: "Measures taken by the company to reduce pollutant emissions."

[0057] In this regard, the indicator data expected to be obtained in this embodiment include news that meets the above definition. For example, the news may be "In accordance with relevant laws and regulations and the requirements of the Ecological Protection Bureau, XX Company conducts regular self-inspections and prepares self-monitoring reports, which are announced on the Chongqing Environmental Protection Bureau website", or "The XX Ecological Environment Bureau monitoring station conducts supervisory monitoring of XX Company's pollutant emissions every six months. The supervisory monitoring results in XX year all met the emission standards", or "XX Company signed a solid waste transfer agreement with a qualified third-party agency and obtained a solid waste transfer license from the XX Ecological Environment Bureau, transferring hazardous waste to a qualified third-party agency for disposal. A total of XX tons of hazardous waste were disposed of in XX year", etc.

[0058] Before conducting actual evaluation, this embodiment requires training of the Doc2Vec technology and the first language model to ensure the accuracy of the results.

[0059] Doc2Vec model training

[0060] The Doc2Vec technology in this embodiment vectorizes the collected data. Its advantage is that it can calculate relatively long texts and requires relatively short resources and computing time. The training steps include:

[0061] Data Preparation and Preprocessing: This example uses over 10 million publicly available news items from 2015 to 2022. Before training, we perform Chinese word segmentation using the Jieba tool. This step helps break the text data into words or lexical units, providing input for text embedding modeling.

[0062] Training the Model: This example uses the Python gensim package to train the Doc2Vec model. The training process uses 10 epochs. Each epoch represents a complete pass through the entire dataset. In each epoch, the model learns how to generate document vectors that capture the semantic information of the document.

[0063] The parameters used in training Doc2Vec in this embodiment include: vector dimension: 256; window length: 8; dictionary size: 1,000,000.

[0064] First language model training

[0065] The first language model of this embodiment is obtained by fine-tuning the instruction using the Qwen-7B model as the base model. The training steps include:

[0066] During the instruction fine-tuning phase, this embodiment uses a training data set to train the base large model for two rounds. The training data set contains approximately 55,000 samples of natural language processing (NLP) tasks, including: approximately 30,000 first task data, which includes 50% NER task data samples and 50% classification task data samples; approximately 20,000 second task data, which includes approximately 200 different types of indicator points, and the number of samples of each type of indicator point category is approximately 100; 450 samples that identify the same event, and 4,000 samples of quantitative data extraction. In this embodiment, the number of samples of each type of task data is calculated independently, and their content can be independent or related.

[0067] For samples identifying identical events, the ratio of text pairs containing identical events to different events is 2:1. This dataset partitioning method uses sampling ratios that reflect the actual event distribution (i.e., the probability that different news articles report the same event). This helps improve the model's performance and accuracy on real-world data because it better reflects actual application scenarios.

[0068] Specifically, each data item in the training dataset of this embodiment is structured as follows: the input is a prompt constructed from a pair of text events. The prompt is a text used to guide the model to generate an output, which typically includes a question, instruction, or request to tell the model how to process the text pair; the output is the model's response to the prompt, such as a summary of the content of the text event pair and a judgment on whether the text events are identical. Specifically, for training data related to identifying identical events, the input is a prompt constructed based on a given text pair and the task of determining whether the text pair describes the same event, and the output is the judgment result of whether the text pair describes the same event. For training data related to the quantitative data extraction task, taking the extraction result of the "paid leave" indicator point as an example, the extracted indicator point is "Our employees enjoy 5 days of paid leave each year." The additional input prompt constructed based on this result is: "Please tell me the specific number of paid leave days: Our employees enjoy 5 days of paid leave each year." The correct output result should be "5 days."

[0069] As mentioned above, during the instruction fine-tuning process, the model underwent fine-tuning for 2 instruction fine-tuning rounds (Epochs). Each round represents a complete traversal of the fine-tuning dataset. These two rounds allow the model to learn multiple times on the fine-tuning data to better adapt to the requirements of the task. Specifically, this embodiment uses the LoRA method to create model training parameters to reduce the model training cost and improve the performance and generalization ability of the model, including: adding LoRA modules to the key components of all attention mechanism layers of the first language model, wherein the key components include at least query (q), key (k), value (v) and projection linear layer (projection). The hyperparameters of all LoRA models are r=64, alpha=16. More specifically, the loss function used for optimization fine-tuning is the cross-entropy loss function (Cross-Entropy Loss), and the parameter adjustment method is backpropagation (Backpropagation), wherein the learning rate used for parameter adjustment in this embodiment is 0.0002, and the Adam optimizer is used to optimize and update the parameters.

[0070] For the target task of identifying identical events, this example uses 112 news stories isolated from the training dataset to test the similarity detection performance of the first language model. The use of this isolated test set helps to independently verify the model's performance on new data without being affected by the training data.

[0071] For the target task of identifying the same event, it can be described as text classification or text clustering. At the framework level, the input of the first language model is a text library containing a large amount of news text, and the output is a news library that has been classified or clustered. This processing process aims to classify news with similar themes or content into the same category, so that news within the same category exhibits highly similar characteristics, while news in different categories do not have this high similarity. For the indicator data extraction task, it can be described as extracting content related to a specific topic from the text. At the framework level, the input of the first language model is the definition of the indicator and the text to be extracted, and the output is the content related to the indicator in the text.

[0072] This example uses the Qwen-7B model as the base model and performs two rounds of instruction fine-tuning on the above training dataset. This process helps the model learn the feature representations of various NLP tasks, enabling it to handle multiple tasks simultaneously. The final model performed well in performance evaluation, with the first language model achieving an EM (Exact Match) index of 0.4 and an F1 score of 0.7. These indicators reflect the performance of the model in tasks such as MRC (Machine Reading Comprehension), with high exact matching and comprehensive scores.

[0073] Indicator data extraction

[0074] This embodiment uses the trained Doc2Vec model and the first language model to extract the index data. The news data used in this embodiment can be news collected from any source or any existing news database.

[0075] Low-level data processing steps

[0076] This embodiment loops through the news library until each piece of news is similarly judged / categorized. Specifically, the underlying data processing steps for any two pieces of news are as follows:

[0077] Step 1: Use the Doc2Vec model to convert the two news articles into text vectors.

[0078] Step 2: Calculate the cosine similarity of the two news articles. If the cosine similarity is less than 0.75, output "the two news articles are not similar"; if the cosine similarity is greater than or equal to 0.75, and at least one of the news articles is greater than or equal to 300 characters, output "the two news articles are similar". Otherwise (cosine similarity is greater than or equal to 0.75, and the length of both news articles is less than 300 characters), execute step 3.

[0079] Step 3: Use the first language model to determine the similarity between the two news articles and output "the two news articles are not similar" or "the two news articles are similar".

[0080] Through the above steps, this embodiment completes the similarity determination between any two news articles. After looping through the underlying database, all news data can be classified by similarity. In this embodiment, news in the same category is described as exhibiting highly similar characteristics, which generally refers to their similarity in terms of theme, content, etc.

[0081] Indicator data extraction steps

[0082] After the underlying data processing steps are completed, that is, all news data have been classified, the present embodiment performs the index data extraction step. Specifically, the index data extraction steps for a certain news article are as follows.

[0083] Each predetermined indicator point and news text are input as questions into the first language model. The model will perform question understanding and generate a textual answer associated with the news text. A predefined template can be constructed for each indicator point definition and news text and provided to the first language model. If the first language model outputs "no," it indicates that the news text does not contain content related to the indicator point. If the first language model's answer is not "no," the first language model's answer is considered to be the content of the news text related to the indicator point.

[0084] Comparative Example 1

[0085] Comparison Model 1: Doc2vec Model

[0086] Comparison Model 2: Fine-tuned BERT-large MRC Model: The top softmax layer of the original BERT-large model is replaced with a linear layer with two output neurons, which correspond to the logit values ​​at the beginning and end of the answer, respectively.

[0087] Comparison model 3: Baichuan-13B.

[0088] Comparative model 4: Qwen-7B without training of the present invention.

[0089] Comparison model 5: ChatGPT-4.

[0090] This comparative example uses Example 1 of the present invention and comparative models 1-5 that are more common on the market to compare performance.

[0091] Regarding the underlying data processing steps, this embodiment considers using precision, recall, and accuracy to evaluate model performance, where precision represents the model's ability to correctly identify news as duplicates; recall represents the model's ability to successfully find all actual duplicate news; and accuracy represents the ratio of the number of correctly predicted samples to the total number of samples, which measures the proportion of correct predictions made by the model among all predictions.

[0092] EM (Exact Match) and F1 (F1 Score) are commonly used metrics for evaluating the performance of extraction tasks. The following compares the performance of various models based on the metric data extraction steps.

[0093] As can be seen, the embodiments of the present invention significantly outperform other comparison models in the underlying data processing steps. Furthermore, it is noteworthy that despite the relatively small amount of training data used in the indicator data extraction step, the performance of the first language model of the present invention is comparable to that of the BERT-large MRC model fine-tuned using 180,000 pieces of data. This highlights the resource efficiency advantage of the present invention, while also achieving excellent performance across multiple tasks.

[0094] The present invention has been introduced in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The above implementation description is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. Changes and improvements to the present invention will be possible without exceeding the concept and scope specified in the appended claims. In summary, the contents of this specification should not be understood as limiting the present invention.

Claims

1. A data extraction method for entity evaluation, characterized in that The method is used to extract target data according to predefined metrics, including: A bottom - layer data processing step, at least including using a first language model to identify the same events in the bottom - layer data; A metric data extraction step, at least including using a first language model to extract target data in the bottom - layer data according to the predefined metrics; Wherein, the base large model of the first language model includes a large language model based on the Transformer model architecture, and the first language model is trained based on the base large model; Wherein, the training at least includes at least 2 rounds of optimization fine - tuning based on a first training dataset; The first training dataset includes first - task data and second - task data. The first - task data includes 40% - 50% NER task data, 40% - 50% classification task data, and / or 0 - 20% other task data; The second - task data includes at least 200 different types of metric points, and the sample quantity of each type of metric point is at least between 10 - 100; 2. The method according to claim 1, wherein The first training dataset includes third - task data. The third - task data includes at least 300 - 500 task samples for identifying the same events and / or at least 2000 - 4000 task samples for quantitative data extraction; wherein, the ratio of the sample quantity of the same events to the different events in the task samples for identifying the same events is about 1:1 - 3:

1.

3. The method according to claim 1 or 2, characterized in that, The first - task data includes at least 20,000 - 40,000 task samples; wherein, the 0 - 20% other task data includes relation extraction tasks, semantic role annotation tasks, and / or event extraction tasks.

4. The method according to claim 3, wherein The optimization fine - tuning includes using a parameter adjustment method to gradually adjust the parameters of the base model according to the gradient of the loss function; Wherein, the parameter adjustment method includes the back - propagation method, the loss function includes the cross - entropy loss function, and the parameters are created using the LoRA method.

5. The method according to claim 4, wherein The method of using LoRA includes adding LoRA modules to the key components of all attention mechanism layers of the base model; Wherein, the key components include query (q), key (k), value (v), and projection linear layer (projection); the rank (r) of the low - rank matrix decomposition of the LoRA module is set to 4 - 64, and the regularization parameter (alpha) is set according to the rank (r).

6. The method according to any one of claims 1-5, characterized in that, Using the first language model to identify the same events in the bottom - layer data includes: circularly judging whether the events included in any two texts are the same, and merging the texts containing the same events; wherein, judging whether the events included in any two texts are the same includes: a) Vectorizing the any two texts and calculating the cosine similarity of the any two texts; b) If the cosine similarity is less than the first threshold, it is determined that the events included in the any two texts are not the same; if the cosine similarity is greater than or equal to the first threshold, and the length of at least one text is greater than or equal to the second threshold, it is determined that the events included in the any two texts are the same; otherwise, execute step c); c) Use the first language model to determine whether the events contained in any two texts are the same; Wherein, when the judgment result of any one of the above steps a)-c) is the same, execute the step of merging the texts containing the same event; the first threshold is set to 0.6-0.8, and the second threshold is set to 200-400 characters.

7. The method according to claim 6, wherein The vectorization includes using the Doc2Vec model, and the Doc2Vec model is obtained through the following training method: Perform Chinese word segmentation on the vectorization training data set; Set the parameters of the Doc2Vec model, including setting the vector dimension to 100-300, the window length to 2-10, and setting the dictionary size; and Perform at least 10 rounds of iterative training based on the vectorization training data set.

8. The method according to claim 6, wherein The step of executing the merging of texts containing the same event includes: Create a set for each text using the disjoint set; According to the judgment result, merge the sets of texts containing the same event through the Union operation; When the number of texts in a set is greater than 1, select one of the texts as the representative text of the set.

9. The method according to any one of claims 1-8, characterized in that, The Transformer model architecture is selected from the Qwen-7B model or the Llama model.

10. A data extraction system for entity evaluation, characterized in that The system is used to implement the method described in any one of claims 1-9.

Citation Information

Patent Citations

  • Event entity joint extraction method and device, computer equipment and storage medium

    CN112052682A

  • Multi-task interaction enhanced electronic text event extraction method

    CN112069811A

  • Joint extraction method for named entities and relationships in judicial domain

    CN113221567A

  • Information extraction model training method, information extraction method and device

    CN116150613A

  • Generative cross-language event extraction method enhanced by using large language model

    CN116956922A

Cited By

  • Rapid classification and grading method and system for flow data and medium

    CN120850050A

  • A method, system, and medium for fast classification of traffic data

    CN120850050B

  • Text reasoning method, product, equipment and storage medium

    CN120851223A

  • Big language model-based pest control level evaluation method and system

    CN122452951A