Text extraction method and device, computer equipment, storage medium and computer program product

By combining a multi-head hybrid expert model and a dynamic prompt template with the BERT model, the problems of high model complexity and inaccurate boundary positioning in named entity recognition are solved, and efficient and accurate named entity extraction is achieved.

CN120706426APending Publication Date: 2025-09-26JIANGNAN INST OF COMPUTING TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510684595.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing named entity recognition methods have high model complexity when facing multi-category named entities, and it is difficult to accurately locate the boundaries of named entities, resulting in poor recognition results and redundancy.

Method used

A multi-headed hybrid expert model is used to identify the text language and target data category. Combined with the text annotation vector library and the right-or-wrong judgment model, the accuracy of language recognition and extraction type is achieved through a multi-task learning framework and dynamic prompt templates. The BERT model and classification head are used for final judgment.

Benefits of technology

It improves the accuracy and efficiency of named entity recognition, reduces model complexity, reduces redundant recognition, and ensures the reliability and accuracy of output data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706426A_ABST
    Figure CN120706426A_ABST
Patent Text Reader

Abstract

The invention relates to a text extraction method and device, computer equipment, a storage medium and a computer program product. The method comprises the steps that a to-be-recognized text is obtained, the to-be-recognized text comprises a target text and target data associated with the target text, and the language of the to-be-recognized text and the category of the target data are recognized through a multi-head hybrid expert model; obtaining a text extraction cue word corresponding to the language and the category from the text labeling vector library, calling a text extraction model corresponding to the text extraction cue word, extracting a target text and target data from the to-be-recognized text through the text extraction model, and obtaining structured text data based on the target text and the target data; outputting a probability value of the structured text data belonging to a preset target category through a correctness judgment model; and outputting the structured text data under the condition that the probability value is greater than a preset threshold value. By adopting the method, the accuracy of the extraction position can be ensured, and meanwhile, the complexity of the extraction model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of text extraction, and in particular to a text extraction method, apparatus, computer equipment, storage medium, and computer program product. Background Art

[0002] Named entity recognition (NER) is a crucial task in natural language processing (NLP). It identifies meaningful entities within text and categorizes them into predefined categories, such as names of people, places, organizations, times, dates, or numbers. It can extract specific entities from large amounts of text data, enabling rapid acquisition and integration of important information. It effectively mines the potential knowledge within text, reduces the time and effort required to manually screen information, and significantly improves the efficiency and quality of information processing.

[0003] Combining a pre-trained model with a conditional random field (CRF) is a common method for named entity recognition. During the prediction phase, the text to be recognized is fed into a trained model. The pre-trained model first extracts features. Then, the CRF layer predicts the named entity label corresponding to each word or character based on the learned label dependencies and feature information, completing the named entity recognition task. This method has strong recognition capabilities, but when there are many named entity categories, the number of labeled categories and the model's classification head also increase, increasing model complexity and requiring more computing resources. Furthermore, the majority of the text contains content that does not belong to named entities. Consequently, the recognition process suffers from issues such as inaccurately locating named entities and redundant recognized content. Summary of the Invention

[0004] Based on this, it is necessary to provide a text extraction method, device, computer equipment, computer-readable storage medium and computer program product that can ensure the accuracy of the extraction position while reducing the complexity of the extraction model to address the above technical problems.

[0005] In a first aspect, the present application provides a text extraction method, the method comprising:

[0006] Acquire a text to be recognized, wherein the text to be recognized includes a target text and target data associated with the target text, and identify the language of the text to be recognized and the category to which the target data belongs using a multi-headed mixed expert model;

[0007] Obtaining a text extraction prompt word corresponding to the language and the category from a text annotation vector library, calling a text extraction model corresponding to the text extraction prompt word, extracting the target text and target data from the text to be recognized using the text extraction model, and obtaining structured text data based on the target text and the target data;

[0008] The structured text data is input into a true / false judgment model, and the true / false judgment model outputs a probability value that the structured text data belongs to a preset target category; when the probability value is greater than a preset threshold, the structured text data is output.

[0009] In one embodiment, the method further comprises:

[0010] Acquire a multilingual text sample and a marking target, and mark the multilingual text sample according to the marking target to obtain a marked multilingual text sample;

[0011] Performing vectorization processing on the annotated multilingual text samples to obtain multilingual text vectors;

[0012] The annotated multilingual text samples and corresponding multilingual text vectors are stored in a preset data structure to obtain the text annotation vector library.

[0013] In one embodiment, obtaining structured text data based on the target text and the target data includes:

[0014] The target text and the target data are structurally parsed to obtain structured text data.

[0015] In one embodiment, the multi-head hybrid expert model includes a text language recognition expert model and a data category recognition expert model, and the method further includes:

[0016] Establishing a text language identification expert model and a data category identification expert model respectively, and defining a gating network for the text language identification expert model and the data category identification expert model, wherein the gating network includes a feature extraction layer, a nonlinear activation layer, and a normalized output layer;

[0017] Training the text language recognition expert model and the data category recognition expert model, assigning weights to the text language recognition expert model and the data category recognition expert model through the gating network, and performing weighted summation of the outputs of the text language recognition expert model and the data category recognition expert model according to the weights to obtain an output result of the multi-head hybrid expert model;

[0018] According to the difference between the output result and the language and data category of the marked preset text category, the parameters of the text language identification expert model and the data category identification expert model are adjusted until the difference meets the requirements, thereby obtaining a multi-head hybrid expert model.

[0019] In one embodiment, the method further comprises:

[0020] Obtaining a text extraction sample, inserting a preset category label, placeholder, and separator into the text extraction sample to obtain a text extraction correct and incorrect sample set; wherein the category label and the separator are used to mark the boundary of structured information in the text extraction sample, and the placeholder is used to mark the target text and target data in the text extraction sample;

[0021] A BERT-based correctness / error judgment model is trained using a correctness / error sample set of extracted text to obtain an extracted correctness / error judgment model; wherein the BERT-based extracted correctness / error judgment model is composed of a pre-trained BERT model and a classification head.

[0022] In one embodiment, obtaining a text extraction prompt word corresponding to the language and the category from a label vector library includes:

[0023] Generating a query vector according to the language and the category using the same model and method as used to construct the annotation vector library;

[0024] Calculate the similarity between the query vector and each vector in the annotation vector library,

[0025] The vectors in the annotation vector library are sorted according to the similarity, and the text extraction prompt word corresponding to the vector with the highest similarity is obtained.

[0026] In a second aspect, the present application further provides a text extraction device, characterized in that the device comprises:

[0027] A recognition module is configured to obtain a text to be recognized, wherein the text to be recognized includes a target text and target data associated with the target text, and identify the language of the text to be recognized and the category to which the target data belongs using a multi-headed mixed expert model;

[0028] an extraction module configured to obtain, from a text annotation vector library, a text extraction prompt word corresponding to the language and the category, invoke a text extraction model corresponding to the text extraction prompt word, extract the target text and target data from the text to be recognized using the text extraction model, and obtain structured text data based on the target text and the target data;

[0029] The judgment module is used to input the structured text data into a true or false judgment model, and output a probability value of the structured text data belonging to a preset target category through the true or false judgment model; when the probability value is greater than a preset threshold, the structured text data is output.

[0030] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the steps of the method described in the first aspect or any embodiment of the first aspect when executing the computer program.

[0031] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the method described in the first aspect or any embodiment of the first aspect are implemented.

[0032] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, characterized in that when the computer program is executed by a processor, the steps of the method described in the first aspect or any embodiment of the first aspect are implemented.

[0033] The above-mentioned text extraction method, apparatus, computer equipment, storage medium, and computer program product can realize multi-dimensional information recognition of text by identifying the language of the text to be identified and the category to which the target data belongs through a multi-head hybrid expert model. Accurate identification of the language helps to carry out subsequent processing based on different language characteristics. Identification of the category to which the target data belongs can clarify the nature and purpose of the data, facilitate subsequent more targeted text extraction operations, and improve the efficiency and accuracy of extraction. Text extraction prompt words corresponding to the language and category are obtained from the text annotation vector library, and the corresponding text extraction model is called for extraction. This can more accurately guide the text extraction model to operate and improve the accuracy of extraction. This method can select the most appropriate prompt words and models according to different languages ​​and categories. Structured text data is obtained based on the target text and target data, which is more convenient for subsequent analysis, storage, and use. The structured text data is input into the right and wrong judgment model, and the probability value of belonging to the preset target category is output, which can perform quality assessment on the extracted text data. The structured text data is output when the probability value is greater than a preset threshold, ensuring that the output data has high reliability and accuracy. The above text extraction method, apparatus, computer device, storage medium and computer program product can ensure the accuracy of the extraction position while reducing the complexity of the extraction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0035] Figure 1 A diagram showing an application environment of a text extraction method in one embodiment;

[0036] Figure 2 Schematic diagram of a flow chart of a text extraction method in one embodiment;

[0037] Figure 3 Schematic diagram of a flow chart of a method for extracting epidemic report text in one embodiment;

[0038] Figure 4 is a structural block diagram of a text extraction device in one embodiment;

[0039] Figure 5 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0041] The text extraction method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented as an independent server or a server cluster consisting of multiple servers.

[0042] In an exemplary embodiment, Figure 2 As shown, a text extraction method is provided, which is applied to Figure 1 The application environment in FIG. 1 is taken as an example to illustrate the method, which includes the following steps 202 to 206. Among them:

[0043] Step 202: Obtain a text to be recognized, wherein the text to be recognized includes a target text and target data associated with the target text, and identify the language of the text to be recognized and the category to which the target data belongs through a multi-headed mixed expert model.

[0044] The "text to be identified" refers to the input object of the text extraction method and can be text content obtained from sources such as web pages, documents, and social media. These texts may contain target text and target data. The target text refers to the portion of text within the text to be identified that has specific meaning or value. It may be about a specific event, topic, or technical details. The specific definition of the target text can be determined based on the purpose and application scenario of the text extraction. For example, if the text to be identified is a report related to an infectious disease epidemic, the target text may include terms such as "confirmed cases," "local cases," or "recovered cases." The target data refers to the specific data or statistical results associated with the target text. A multi-headed mixture of expert models is a model architecture composed of multiple expert models, each responsible for processing a specific type of task or data. The category of the target data refers to the category to which the target data belongs within a preset classification system. For example, if the text to be identified is a report related to an infectious disease epidemic and the target text is a confirmed case, the target data may include the number of newly confirmed cases or the cumulative number of confirmed cases.

[0045] For example, the text content in the web page can be extracted through a URL request, text data can be read from an existing database, or text data can be obtained from a local text file as the text to be recognized. After obtaining the text to be recognized, preprocessing steps such as text cleaning or word segmentation can be performed, and the expert model can be selected according to the specific needs of text processing. For example, for language recognition tasks, a language model based on deep learning, such as BERT, GPT, etc., can be selected as the expert model. For the task of identifying the target data category, a decision tree or support vector machine can be used as the expert model. The multi-head hybrid expert model is trained with a labeled data set, and the parameters are continuously adjusted during the training process to ensure that the recognition accuracy of the model meets the actual needs.

[0046] Step 204: Obtain text extraction prompt words corresponding to the language and the category from a text annotation vector library, call a text extraction model corresponding to the text extraction prompt words, extract the target text and target data from the text to be recognized through the text extraction model, and obtain structured text data based on the target text and the target data.

[0047] The text annotation vector library refers to a database that stores multilingual text samples and their annotation information. The text vectors stored in it contain the semantic information and annotation information of the text. Text extraction prompt words refer to keywords or phrases corresponding to specific languages ​​and categories obtained from the text annotation vector library. Prompt words can help the text extraction model more accurately locate and extract the required information, improving the efficiency and accuracy of text extraction. The text extraction model refers to the model used to extract target text and target data from the text to be identified. Structured text data refers to the data obtained by organizing the extracted target text and target data according to a certain structure.

[0048] Exemplarily, text samples of different languages ​​and categories can be collected, and information such as language labels, data categories, key texts to be extracted, and structured data can be processed. The text can be converted into a vector representation through a multilingual training model, and the language, category and other annotation information can be encoded into a vector, which can be spliced ​​or fused with the text vector. The text vector and the annotation information can be associated and stored to form a vector library. The category of the target text and target data can be converted into a vector representation and a query vector can be generated. The similarity between the query vector and the vectors in the vector library can be calculated, and the text samples corresponding to the first several vectors with the highest similarity can be selected. Keywords or phrases that are considered to be relevant to the target can be extracted as prompt words. The corresponding text extraction model can be loaded according to the mapping relationship between the prompt words and the model. The text to be recognized is input into the text extraction model to extract the target text and target data, and then structured processing is performed to obtain structured text data.

[0049] In step 206, the structured text data is input into a true / false judgment model, and the true / false judgment model outputs a probability value of the structured text data belonging to a preset target category; when the probability value is greater than a preset threshold, the structured text data is output.

[0050] Among them, the true or false judgment model refers to a model used to analyze the input structured text data and judge its degree of matching with the preset target category. Its input is structured text data, and its output is the probability value that the structured text data belongs to the preset target category. The accuracy of text extraction can be improved through verification of the true or false judgment model.

[0051] For example, the model architecture can be selected according to actual needs, samples can be collected and data can be labeled, appropriate hyperparameters such as learning rate, number of iterations and number of hidden layer neurons can be set according to the model architecture and data characteristics, and the model can be trained using optimization algorithms and loss functions. The model parameters can be continuously adjusted to gradually improve the prediction accuracy of the model on the training data set, and the performance of the model can be judged by indicators such as accuracy or recall rate. The model whose performance meets expectations can be used to determine whether the structured text data belongs to the preset category, and the structured text data belonging to the preset target category can be output as the extraction target of the text extraction task.

[0052] The above text extraction method uses a multi-head hybrid expert model to identify the language of the text to be identified and the category to which the target data belongs, which can realize multi-dimensional information recognition of the text. Accurate identification of the language helps to carry out subsequent processing based on different language characteristics. Identification of the category to which the target data belongs can clarify the nature and purpose of the data, facilitate more targeted text extraction operations in the future, and improve the efficiency and accuracy of extraction. Text extraction prompt words corresponding to the language and category are obtained from the text annotation vector library, and the corresponding text extraction model is called for extraction. This can more accurately guide the text extraction model to operate and improve the accuracy of extraction. This method can select the most appropriate prompt words and models according to different languages ​​and categories. Structured text data is obtained based on the target text and target data, which is more convenient for subsequent analysis, storage and use. The structured text data is input into the right and wrong judgment model, and the probability value of belonging to the preset target category is output, which can perform quality assessment on the extracted text data. When the probability value is greater than the preset threshold, the structured text data is output, ensuring that the output data has high reliability and accuracy. The above text extraction method can ensure the accuracy of the extraction position while reducing the complexity of the extraction model.

[0053] In one embodiment, the method further includes: obtaining a multilingual text sample and a labeling target, labeling the multilingual sample text according to the labeling target to obtain a labeled multilingual text sample; performing vectorization processing on the labeled multilingual text sample to obtain a multilingual text vector; and storing the labeled multilingual text sample and the corresponding multilingual text vector in a preset data structure to obtain the text labeling vector library.

[0054] Among them, multilingual text refers to text data containing two or more languages, annotation targets refer to the preset information types that need to be extracted or classified from the text, vectorization processing refers to the process of converting text into numerical vectors, and multilingual text vectors refer to text representations that have similar semantics in different languages ​​and are close in distance in the vector space after vectorization processing, and different language expressions of the same concept are mapped to similar areas.

[0055] For example, tools such as Prodigy, BRAT, and spaCy can be used to annotate text, a multilingual word segmenter can be used to process mixed text, vectors can be generated through a pre-trained model, and the text and the generated vectors can be stored in a relational database to obtain a text annotation vector library.

[0056] In one embodiment, obtaining structured text data based on the target text and the target data includes: performing structured parsing on the target text and the target data to obtain structured text data.

[0057] Among them, structured parsing refers to converting unstructured or semi-structured text into machine-understandable structured data with clear semantic relationships.

[0058] For example, semantic relationships between entities in a text may be extracted through rule-based methods or machine learning methods, structured associations may be constructed, and the extracted entities and relationships may be organized into a predefined data structure, such as a JSON data structure.

[0059] In one embodiment, the multi-head hybrid expert model includes a text language recognition expert model and a data category recognition expert model, and the method further includes: establishing the text language recognition expert model and the data category recognition expert model respectively, and defining a gating network for the text language recognition expert model and the data category recognition expert model, wherein the gating network includes a feature extraction layer, a nonlinear activation layer, and a normalized output layer; training the text language recognition expert model and the data category recognition expert model, assigning weights to the text language recognition expert model and the data category recognition expert model through the gating network, and performing weighted summation of the outputs of the text language recognition expert model and the data category recognition expert model according to the weights to obtain an output result of the multi-head hybrid expert model; and adjusting the parameters of the text language recognition expert model and the data category recognition expert model according to the difference between the output result and the language and data category of the marked preset text category until the difference meets the requirements, thereby obtaining the multi-head hybrid expert model.

[0060] Among them, the text language recognition expert model is a model that is only used to identify the language to which the text belongs. Its input can be text content, and the output is the language to which the text content belongs, such as Chinese, English or Japanese. The data category recognition expert model refers to a model used to classify text content. The gated network refers to a neural network that controls the weight distribution of the elder expert model, which can determine the contribution of each expert model to the output.

[0061] Exemplarily, the input text can be passed to two expert models and the gating network at the same time, the language recognition expert model is trained using language annotation data, and the category recognition expert model is trained using category annotation data. The gating network generates a weight vector, and the language recognition results and category recognition results are weighted and summed to obtain the final output of the multi-head hybrid expert model. The optimization goal is to minimize the cross-entropy loss function, and the parameters of the expert model and the gating network are simultaneously updated through backpropagation.

[0062] In one embodiment, the method further includes: obtaining a text extraction sample, inserting preset category labels, placeholders and separators into the text extraction sample to obtain a text extraction correct or incorrect sample set; wherein the category labels and the separators are used to mark the boundaries of structured information in the text extraction sample, and the placeholders are used to mark the target text and target data in the text extraction sample; using the extracted text correct or incorrect sample set to train a BERT-based correct or incorrect judgment model to obtain an extraction correct or incorrect judgment model; wherein the BERT-based extraction correct or incorrect judgment model is composed of a pre-trained BERT model and a classification head.

[0063] Among them, text extraction samples refer to text data used to train the extraction correctness judgment model, category labels refer to identifiers used to represent specific types of structured information, which are used to clarify the boundaries and semantics of different types of entities or information in the text, placeholders refer to symbols used to mark the target text or data to be extracted, which can indicate the key content that the model focuses on, separators refer to symbols used to divide different structured information fragments, structured information refers to data with clear semantics and format extracted from unstructured text, such as name, date, amount or relationship, etc., BERT model refers to the bidirectional encoder representation based on Transformer, and classification head refers to the neural network component connected after the output layer of the BERT model, which is used to map the semantic features extracted by BERT to specific classification tasks to realize the prediction function of the model.

[0064] For example, the structured information to be extracted from the text extraction samples can be manually or through specific tools, and the structured information can be marked with preset labels and symbols. Samples with labels and placeholders inserted in the correct positions are used as positive samples. Negative samples are generated by replacing, inserting, or deleting labels. The positive and negative samples are combined to form a text extraction correct and incorrect sample set to train a BERT-based correct and incorrect judgment model. The extraction correct and incorrect judgment model is then obtained by optimizing the model.

[0065] In one embodiment, obtaining a text extraction prompt word corresponding to the language and the category from a label vector library includes: generating a query vector based on the language and the category using the same model and method as used to construct the label vector library; calculating the similarity between the query vector and each vector in the label vector library, sorting the vectors in the label vector library according to the similarity, and obtaining a text extraction prompt word corresponding to the vector with the highest similarity.

[0066] The query vector is a vector generated based on a given language and category using the same model and method used to construct the annotation vector library. It serves as the basis for searching within the annotation vector library and contains feature information related to the currently input language and category, so as to find matching vectors within the annotation vector library. Similarity is an indicator used to measure the degree of similarity between the query vector and each vector in the annotation vector library.

[0067] For example, features such as language and character strings can be encoded using the same model used to construct the annotated vector library. The processed features are converted into a fixed-dimensional query vector. The cosine similarity, Euclidean distance, or dot product similarity between the query vector and the vectors in the vector library is calculated. The vector with the highest similarity is selected to obtain the corresponding prompt word pre-associated with the vector. This serves as the prompt word for invoking the text extraction model.

[0068] In the process of epidemic prevention and control, extracting the number of people related to the epidemic (newly confirmed cases, new deaths, cumulative confirmed cases, and cumulative deaths) from a large number of text reports to form structured data can clearly show the development and changes of the epidemic, provide effective basis and guidance for the analysis and command of relevant departments in epidemic prevention and control, and is of great significance to the monitoring, analysis and decision-making of the epidemic.

[0069] The task of extracting meaningful entities from text is called Named Entity Recognition (NER). Currently, the main method used is a pre-trained model (mainly the Bert model) combined with a Conditional Random Field (CRF). The annotation method used is BIO (B stands for Begin, the first character of the named entity; I stands for Inside, which represents all characters in the named entity except the first character; and O stands for Outside, which does not belong to the named entity). Although this model has strong recognition capabilities, it still has the following problems:

[0070] 1. There are errors in entity boundary positioning. For a piece of text, it is impossible to accurately locate the start and end positions of the named entity.

[0071] 2. When there are many categories of named entities, the number of labeled categories and the classification head of the model (the size of the state transition matrix of the CRF) must also be increased accordingly, increasing the complexity of the model.

[0072] 3. The majority of a text is not a named entity, which results in a lot of redundancy in the recognition process and restricts the recognition effect of the model to a certain extent.

[0073] 4. In practical applications, this model requires more training data.

[0074] For the task of extracting the number of people related to the epidemic, the accuracy of the extraction position is guaranteed, the complexity of the model is reduced, and the data annotation work is reduced. In an exemplary embodiment, Figure 3 As shown in the figure, a method for extracting case numbers from epidemic reports is proposed. This method first establishes a library of annotated vectors for multilingual epidemic data. Then, based on different epidemic inputs, a multi-head hybrid expert model in a multi-task learning framework is used to simultaneously implement language recognition and extraction type identification. A dynamically adjusted epidemic extraction prompt template is then designed to initially extract case numbers and convert epidemic numbers. Based on the results obtained in the above steps, incorrect extraction results may exist in practice, and the correct extraction results need to be screened out. Finally, a set of correct and incorrect samples of epidemic extraction is constructed, and a classification model is trained to obtain the final extraction results.

[0075] The present invention implements the task of extracting the number of epidemic cases from multilingual epidemic report texts, including the establishment of a labeled vector library for multilingual epidemic data, language and extraction type identification based on a multi-head hybrid expert model, the implementation of a dynamic epidemic extraction prompt template, and the judgment of the correctness of epidemic extraction.

[0076] 1. Establishment of a label vector library for multilingual epidemic data

[0077] For epidemic-related reports in different countries, this method establishes a unified annotation management library and performs vector conversion at the same time to realize vectorized processing of epidemic report texts in different countries for subsequent retrieval enhancement prompt.

[0078] 2. Language and extraction type recognition based on a multi-headed mixture of experts model

[0079] Based on epidemic reports in different languages ​​and the categories of epidemic cases involved, this method uses a multi-headed mixed expert model to simultaneously determine the text language category and the epidemic case category that needs to be extracted. The model involves text language recognition experts and epidemic case category recognition experts. The gating network of each expert is implemented using DNN+ReLU+softmax.

[0080] 3. Implementation of dynamic epidemic extraction prompt template

[0081] The dynamic epidemic extraction prompt first determines the language of the input epidemic report and the type of epidemic case to be extracted based on the previous step, and then uses the retrieval enhancement method to obtain the closest prompt example from the annotation vector library.

[0082] 4. Judgment of the correctness of the epidemic extraction

[0083] Regarding the results obtained in the previous step, there is a possibility of extraction errors in practice. This step aims to construct a set of correct and incorrect samples for epidemic extraction, and use the BERT pre-training model plus a neural network model with a classification head for training and prediction to ensure the accuracy of the extraction results.

[0084] First, the design of each sample is: [cls]{key}number of people{value}[sep]report text.

[0085] [cls]: This is the class label in classification tasks. It indicates the beginning of a sample and indicates that the subsequent content is related to the class label. The label is set to 1 or 0 depending on the positive or negative sample.

[0086] {key}: is a placeholder representing a specific epidemic-related keyword. The description after the keyword indicates that the following value is the number of people, such as "newly confirmed cases", "cumulative confirmed cases", "new deaths" or "cumulative deaths".

[0087] {value}: It is also a placeholder, representing the specific value associated with {key}, that is, the number of people in the epidemic data (pure number).

[0088] [sep]: A delimiter that distinguishes structured information from the subsequent unstructured text (i.e., the specific content of the epidemic report). In sequence annotation, delimiters are used to mark the end of structured information and the beginning of unstructured text.

[0089] For example, a specific sample: [cls] 120 new confirmed cases [sep] According to today's report, 120 new confirmed cases were reported in a certain area...

[0090] [cls] and [sep] serve as markers to help the model identify the boundaries of structured information, while {key} and {value} provide specific epidemic data points. This design enables the model to more accurately extract and understand key information from text.

[0091] Then, for the above training data, the present invention adopts a neural network model consisting of a BERT pre-training model and a classification head; finally, for the input epidemic reports and the obtained extraction results, the trained model is input in sequence for each case type, and finally the correct extraction result of the number of epidemic cases is obtained.

[0092] 1. Establishment of a label vector library for multilingual epidemic data

[0093] The designed fields include text ID, text title, report language, epidemic text content, epidemic extraction and annotation results, and epidemic text vectors. The epidemic extraction and annotation results are formatted in a structured JSON format. Keywords in the JSON field include, but are not limited to, "new confirmed cases," "cumulative confirmed cases," "new deaths," and "cumulative deaths." The epidemic text vector field uses a unified multilingual vector model to transform the "epidemic text content" field.

[0094] 2. Language and extraction type recognition based on a multi-headed mixture of experts model

[0095] Specifically, for the input epidemic text, we need to determine the language type of the text and the type of cases that can be extracted. This multi-head hybrid expert model involves text language recognition experts and epidemic case category recognition experts. Each expert k has the formula:

[0096]

[0097] in

[0098]

[0099] The gated network is implemented using DNN+ReLU+softmax:

[0100]

[0101] in is a trainable matrix, n is the number of expert networks, where n is 2, and d is the network dimension.

[0102] 3. Implementation of dynamic epidemic extraction prompt template

[0103] According to different languages ​​of the epidemic, the prompt language can be switched in real time;

[0104] Automatically change the type of entity extraction based on the case categories included in the epidemic;

[0105] Using retrieval enhancement technology, the most similar extracted examples are retrieved from the annotated database to guide model extraction;

[0106] 4. Input the large model to obtain the output structure and perform JSON parsing to obtain preliminary epidemic extraction results.

[0107] The specific prompt implementation is as follows:

[0108] You are an excellent expert in extracting the number of epidemic cases. Please identify the {{case type list}} based on the {{language type}} epidemic reports I entered. Please strictly output in JSON format. No additional content is allowed.

[0109] Example epidemic report: {{Epidemic reports in the annotation library found by vector search of new epidemic reports}}

[0110] Sample output: {{Annotation content corresponding to the sample epidemic report}}

[0111] New outbreak report: {{Input outbreak report}}

[0112] New output:

[0113] 4. Judgment of the correctness of the epidemic extraction

[0114] The positive and negative sample sets for the epidemic are designed as follows: {"content":"[cls]{key}number of people{value}[sep]report text","label" :"{label_value}"}; key refers to one of "newly confirmed cases", "cumulative confirmed cases", "new deaths", or "cumulative deaths", and value refers to the actual number of people infected with the epidemic corresponding to the key (pure number). Label_value is set to 1 or 0 according to the positive and negative samples.

[0115] Epidemic judgment model training: The model architecture is a BERT pre-trained model coupled with a neural network classification head. Specifically, after the data passes through the BERT model, the high-dimensional vector at the [cls] position is converted to a value in the range of 0-1 through a feed-forward network. A value greater than 0.5 is considered a positive sample, and otherwise a negative sample.

[0116] For each number of people in the epidemic extraction results, input it into the trained Bert to obtain four [0, 1] prediction values. After sorting according to the output values ​​of the model, if the predicted maximum value is greater than 0.5, the case number type corresponding to the number is obtained and imported into the database table. Otherwise, the number of people in the epidemic is not the final extraction result.

[0117] Compared with the existing technology, the beneficial effects of adopting the above technical solution are:

[0118] 1. Innovative prompt template design: Design dynamically generated prompt templates, intelligently select the most appropriate template based on text content and contextual information, and enhance the accuracy of prompt template matching through retrieval;

[0119] 2. Introducing unsupervised and semi-supervised learning to reduce dependence on large amounts of labeled data;

[0120] 3. Use a multi-task learning framework to simultaneously train the model's capabilities in named entity recognition and semantic understanding;

[0121] 4. With cross-language learning capabilities, the model can handle epidemic reports in multiple languages;

[0122] 5. With a real-time update mechanism, the model can adjust the extraction strategy and recognition mode according to the latest epidemic developments;

[0123] 6. It has the ability to migrate scenes and can migrate and extract key information from similar public event reports.

[0124] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0125] Based on the same inventive concept, embodiments of the present application also provide a text extraction device for implementing the aforementioned text extraction method. The implementation solution provided by this device is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations of one or more text extraction device embodiments provided below can be found in the above-mentioned limitations of the text extraction method and will not be repeated here.

[0126] In an exemplary embodiment, Figure 4 As shown, a text extraction device is provided, comprising: a recognition module, an extraction module and a judgment module, wherein:

[0127] A recognition module is configured to obtain a text to be recognized, wherein the text to be recognized includes a target text and target data associated with the target text, and identify the language of the text to be recognized and the category to which the target data belongs using a multi-headed mixed expert model;

[0128] an extraction module configured to obtain, from a text annotation vector library, a text extraction prompt word corresponding to the language and the category, invoke a text extraction model corresponding to the text extraction prompt word, extract the target text and target data from the text to be recognized using the text extraction model, and obtain structured text data based on the target text and the target data;

[0129] The judgment module is used to input the structured text data into a true or false judgment model, and output a probability value of the structured text data belonging to a preset target category through the true or false judgment model; when the probability value is greater than a preset threshold, the structured text data is output.

[0130] In an exemplary embodiment, the extraction module is further used to: obtain a multilingual text sample and an annotation target, annotate the multilingual sample text according to the annotation target to obtain an annotated multilingual text sample; perform vectorization processing on the annotated multilingual text sample to obtain a multilingual text vector; store the annotated multilingual text sample and the corresponding multilingual text vector in a preset data structure to obtain the text annotation vector library.

[0131] In an exemplary embodiment, the extraction module is further configured to perform structured parsing on the target text and the target data to obtain structured text data.

[0132] In an exemplary embodiment, the multi-head hybrid expert model includes a text language recognition expert model and a data category recognition expert model, and the recognition module is further used to: establish the text language recognition expert model and the data category recognition expert model respectively, and define a gating network for the text language recognition expert model and the data category recognition expert model, wherein the gating network includes a feature extraction layer, a nonlinear activation layer, and a normalized output layer; train the text language recognition expert model and the data category recognition expert model, assign weights to the text language recognition expert model and the data category recognition expert model through the gating network, and perform weighted summation of the outputs of the text language recognition expert model and the data category recognition expert model according to the weights to obtain an output result of the multi-head hybrid expert model; adjust the parameters of the text language recognition expert model and the data category recognition expert model according to the difference between the output result and the language and data category of the marked preset text category until the difference meets the requirements, thereby obtaining the multi-head hybrid expert model.

[0133] In an exemplary embodiment, the judgment module is also used to: obtain text extraction samples, insert preset category labels, placeholders and separators into the text extraction samples, and obtain a text extraction correct and incorrect sample set; wherein the category labels and the separators are used to mark the boundaries of structured information in the text extraction samples, and the placeholders are used to mark the target text and target data in the text extraction samples; use the extracted text correct and incorrect sample set to train a BERT-based correct and incorrect judgment model to obtain an extraction correct and incorrect judgment model; wherein the BERT-based extraction correct and incorrect judgment model is composed of a pre-trained BERT model and a classification head.

[0134] In an exemplary embodiment, the extraction module is further used to: generate a query vector based on the language and the category using the same model and method as used to construct the annotation vector library; calculate the similarity between the query vector and each vector in the annotation vector library, sort the vectors in the annotation vector library according to the similarity, and obtain the text extraction prompt word corresponding to the vector with the highest similarity.

[0135] Each module in the above-mentioned text extraction device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0136] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 5 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store various data involved in the text extraction process, such as epidemic news report data, model data and training data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a text extraction method is implemented.

[0137] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 5As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a text extraction method is implemented.

[0138] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0139] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0140] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0141] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0142] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0143] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0144] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0145] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A text extraction method, characterized in that: The method comprises: Acquire a text to be recognized, wherein the text to be recognized includes a target text and target data associated with the target text, and identify the language of the text to be recognized and the category to which the target data belongs using a multi-headed mixed expert model; Obtaining a text extraction prompt word corresponding to the language and the category from a text annotation vector library, calling a text extraction model corresponding to the text extraction prompt word, extracting the target text and target data from the text to be recognized using the text extraction model, and obtaining structured text data based on the target text and the target data; The structured text data is input into a true / false judgment model, and the true / false judgment model outputs a probability value that the structured text data belongs to a preset target category; when the probability value is greater than a preset threshold, the structured text data is output.

2. The method according to claim 1, characterized in that The method further comprises: Acquire a multilingual text sample and a marking target, and mark the multilingual text sample according to the marking target to obtain a marked multilingual text sample; Performing vectorization processing on the annotated multilingual text samples to obtain multilingual text vectors; The annotated multilingual text samples and corresponding multilingual text vectors are stored in a preset data structure to obtain the text annotation vector library.

3. The method according to claim 1, characterized in that The obtaining of structured text data based on the target text and the target data includes: The target text and the target data are structurally parsed to obtain structured text data.

4. The method according to claim 2, characterized in that The multi-head hybrid expert model includes a text language recognition expert model and a data category recognition expert model, and the method further includes: Establishing a text language identification expert model and a data category identification expert model respectively, and defining a gating network for the text language identification expert model and the data category identification expert model, wherein the gating network includes a feature extraction layer, a nonlinear activation layer, and a normalized output layer; Training the text language recognition expert model and the data category recognition expert model, assigning weights to the text language recognition expert model and the data category recognition expert model through the gating network, and performing weighted summation of the outputs of the text language recognition expert model and the data category recognition expert model according to the weights to obtain an output result of the multi-head hybrid expert model; According to the difference between the output result and the language and data category of the marked preset text category, the parameters of the text language identification expert model and the data category identification expert model are adjusted until the difference meets the requirements, thereby obtaining a multi-head hybrid expert model.

5. The method according to claim 1, wherein The method further comprises: Obtaining a text extraction sample, inserting a preset category label, placeholder, and separator into the text extraction sample to obtain a text extraction correct and incorrect sample set; wherein the category label and the separator are used to mark the boundary of structured information in the text extraction sample, and the placeholder is used to mark the target text and target data in the text extraction sample; A BERT-based correctness / error judgment model is trained using a correctness / error sample set of extracted text to obtain an extracted correctness / error judgment model; wherein the BERT-based extracted correctness / error judgment model is composed of a pre-trained BERT model and a classification head.

6. The method according to claim 1, characterized in that Obtaining text extraction prompt words corresponding to the language and the category from the annotation vector library, including: Generating a query vector according to the language and the category using the same model and method as used to construct the annotation vector library; Calculate the similarity between the query vector and each vector in the annotation vector library, The vectors in the annotation vector library are sorted according to the similarity, and the text extraction prompt word corresponding to the vector with the highest similarity is obtained.

7. A text extraction device, characterized in that: The device comprises: A recognition module is configured to obtain a text to be recognized, wherein the text to be recognized includes a target text and target data associated with the target text, and identify the language of the text to be recognized and the category to which the target data belongs using a multi-headed mixed expert model; an extraction module configured to obtain, from a text annotation vector library, a text extraction prompt word corresponding to the language and the category, invoke a text extraction model corresponding to the text extraction prompt word, extract the target text and target data from the text to be recognized using the text extraction model, and obtain structured text data based on the target text and the target data; The judgment module is used to input the structured text data into a true or false judgment model, and output a probability value of the structured text data belonging to a preset target category through the true or false judgment model; when the probability value is greater than a preset threshold, the structured text data is output.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.