Model construction method and device, equipment and storage medium
Patent Information
- Application Number
- CN202211078361.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-05
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2042-09-05
AI Technical Summary
[0002]目前,随着计算机技术的持续发展,实体识别模型已被广泛应用于各种系统(如问答系统、对话系统以及搜索系统等);具体的,在获取到目标对象输入的文本后,可采用实体识别模型中的语言模型对获取到的文本中的各个字符进行特征提取,从而基于各个字符的字符特征,识别出文本中的实体以响应相应的文本;但医疗等专业名词较为不常见的场景下,目标对象往往会输入错误,使输入的错误文本中包括错误实体(即拼写错误的实体),导致语言模型提取到的字符特征的准确性较低,从而使得实体识别模型进行实体识别的准确性较低,在此种情况下,难以识别出错误文本中的错误实体
[0029]本申请实施例可获取用于优化语言模型的纠错数据集,该纠错数据集包括目标场景下的多个纠错文本,且每个纠错文本包括至少一个错误实体;然后,可调用语言模型对各个纠错文本中的各个字符进行特征提取,得到各个纠错文本中的各个字符的字符特征,从而基于各个纠错文本中各字符的字符特征对语言模型进行优化,得到优化后的语言模型,以提高优化后的语言模型的模型性能,并得到更加精确的实体识别模型,基于此,可提升通过实体识别模型中的语言模型所提取到的字符特征的准确性,进而提升实体识别的准确性。另外,由于字符特征也可反映上下文关系等,那么通过纠错数据集优化后的语言模型能够更加准确地提取出包含错误实体的文本中的各个字符的字符特征,以更加精确地反映包含错误实体的文本所涉及的上下文关系等,进而使得实体识别模型支持识别出文本中的参考实体(即拼写正确的实体)和错误实体。
Smart Images

Figure CN117034928B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a model building method, apparatus, device and storage medium. Background Technology
[0002] Currently, with the continuous development of computer technology, entity recognition models have been widely applied in various systems (such as question-answering systems, dialogue systems, and search systems). Specifically, after obtaining the text input by the target object, the language model in the entity recognition model can be used to extract features from each character in the obtained text. Based on the character features of each character, the entity in the text can be identified to respond to the corresponding text. However, in scenarios where medical and other professional terms are not common, the target object often makes input errors, resulting in the input erroneous text containing incorrect entities (i.e., misspelled entities). This leads to lower accuracy of the character features extracted by the language model, thus reducing the accuracy of entity recognition models in entity recognition. In such cases, it is difficult to identify erroneous entities in erroneous text. Therefore, how to improve the accuracy of character features, and thus improve the accuracy of entity recognition, has become a research hotspot. Summary of the Invention
[0003] This application provides a model building method, apparatus, device, and storage medium that can optimize a language model using an error correction dataset, thereby constructing an entity recognition model using the optimized language model. This improves the accuracy of character features obtained by feature extraction from the language model in the entity recognition model, thereby improving the accuracy of entity recognition and enabling the entity recognition model to identify erroneous entities in the text.
[0004] On the one hand, embodiments of this application provide a model building method, the method comprising:
[0005] Obtain an error correction dataset for optimizing a language model. The error correction dataset includes multiple error correction texts in the target scene, and each error correction text includes at least one error entity, and an error entity includes at least one character.
[0006] For any one of the multiple error-corrected texts, the language model is invoked to extract features from each character in the error-corrected text, thereby obtaining the character features of each character.
[0007] For any character feature in any error-corrected text, obtain the probability difference between the probability of identifying the character as any of the various characters based on the character feature of the character and the reference probability of the character being the various characters.
[0008] Optimize the feature extraction parameters in the language model in the direction of reducing the probability difference to obtain an optimized language model;
[0009] Based on the optimized language model and entity recognition network, an entity recognition model for the target scene is constructed, which is used to perform entity recognition on text in the target scene.
[0010] On the other hand, embodiments of this application provide a model building apparatus, the apparatus comprising:
[0011] The acquisition unit is used to acquire an error correction dataset for optimizing the language model. The error correction dataset includes multiple error correction texts in the target scene, and each error correction text includes at least one error entity, and an error entity includes at least one character.
[0012] The processing unit is used to call the language model for any one of the plurality of error-corrected texts, and extract the features of each character in the error-corrected text to obtain the character features of each character.
[0013] The processing unit is further configured to, for any character feature in any error-corrected text, obtain the probability difference between the probability of identifying any character as each of the characters based on the character feature of any character and the reference probability of any character being each of the characters.
[0014] The processing unit is further configured to optimize the feature extraction parameters in the language model in the direction of reducing the probability difference, so as to obtain an optimized language model;
[0015] The processing unit is further configured to construct an entity recognition model for the target scene based on the optimized language model and entity recognition network, wherein the entity recognition model is used to perform entity recognition on the text in the target scene.
[0016] In another aspect, embodiments of this application provide a computer device, the computer device including a processor and a memory, wherein the memory is used to store a computer program, and the computer program, when executed by the processor, performs the following steps:
[0017] Obtain an error correction dataset for optimizing a language model. The error correction dataset includes multiple error correction texts in the target scene, and each error correction text includes at least one error entity, and an error entity includes at least one character.
[0018] For any one of the multiple error-corrected texts, the language model is invoked to extract features from each character in the error-corrected text, thereby obtaining the character features of each character.
[0019] For any character feature in any error-corrected text, obtain the probability difference between the probability of identifying the character as any of the various characters based on the character feature of the character and the reference probability of the character being the various characters.
[0020] Optimize the feature extraction parameters in the language model in the direction of reducing the probability difference to obtain an optimized language model;
[0021] Based on the optimized language model and entity recognition network, an entity recognition model for the target scene is constructed, which is used to perform entity recognition on text in the target scene.
[0022] In another aspect, embodiments of this application provide a computer storage medium storing a computer program adapted for loading by a processor and executing the following steps:
[0023] Obtain an error correction dataset for optimizing a language model. The error correction dataset includes multiple error correction texts in the target scene, and each error correction text includes at least one error entity, and an error entity includes at least one character.
[0024] For any one of the multiple error-corrected texts, the language model is invoked to extract features from each character in the error-corrected text, thereby obtaining the character features of each character.
[0025] For any character feature in any error-corrected text, obtain the probability difference between the probability of identifying the character as any of the various characters based on the character feature of the character and the reference probability of the character being the various characters.
[0026] Optimize the feature extraction parameters in the language model in the direction of reducing the probability difference to obtain an optimized language model;
[0027] Based on the optimized language model and entity recognition network, an entity recognition model for the target scene is constructed, which is used to perform entity recognition on text in the target scene.
[0028] In another aspect, embodiments of this application provide a computer program product, which includes a computer program that, when executed by a processor, implements the aforementioned model building method.
[0029] This application embodiment can obtain an error correction dataset for optimizing a language model. This dataset includes multiple error-corrected texts from a target scene, and each text includes at least one erroneous entity. Then, the language model can be invoked to extract features from each character in each error-corrected text, obtaining character features for each character. Based on these character features, the language model is optimized to obtain an optimized language model, improving its performance and resulting in a more accurate entity recognition model. This improves the accuracy of character features extracted by the language model in the entity recognition model, thereby enhancing entity recognition accuracy. Furthermore, since character features can also reflect contextual relationships, the language model optimized using the error correction dataset can more accurately extract character features from texts containing erroneous entities, more precisely reflecting the contextual relationships involved. This enables the entity recognition model to identify both reference entities (i.e., correctly spelled entities) and erroneous entities in the text. Attached Figure Description
[0030] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1a This is a flowchart illustrating a model building scheme provided in an embodiment of this application;
[0032] Figure 1b This is a schematic diagram illustrating the interaction between a terminal and a server, provided in an embodiment of this application.
[0033] Figure 2 This is a flowchart illustrating a model building method provided in an embodiment of this application;
[0034] Figure 3 This is a schematic diagram of a labeled sequence provided in an embodiment of this application;
[0035] Figure 4 This is a flowchart illustrating another model building method provided in an embodiment of this application;
[0036] Figure 5a This is a schematic diagram of a BK tree provided in an embodiment of this application;
[0037] Figure 5b This is a flowchart illustrating another model construction method provided in an embodiment of this application;
[0038] Figure 6 This is a schematic diagram of the structure of a model building device provided in an embodiment of this application;
[0039] Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0040] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0041] With the continuous development of internet technology, artificial intelligence (AI) technology has also seen significant advancements. Artificial intelligence refers to the theories, methods, technologies, and application systems that utilize digital computers or computers-controlled machines to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines capable of reacting in a manner similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.
[0042] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0043] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Deep learning, on the other hand, is a technique that utilizes deep neural network systems for machine learning. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0044] Based on machine learning / deep learning techniques in AI, this application proposes a model building scheme to improve the accuracy of character features, thereby enhancing the accuracy of entity recognition and enabling the entity recognition model to identify both reference and erroneous entities in text. It should be noted that this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.
[0045] See Figure 1a As shown, the general principle of the model construction scheme proposed in this application embodiment is as follows: First, an error correction dataset for optimizing the language model can be obtained. The error correction dataset includes multiple error correction texts in the target scene, and each error correction text includes at least one erroneous entity. Then, each error correction text in the error correction dataset can be used to optimize the feature extraction parameters in the language model to obtain an optimized language model. Based on the optimized language model and the entity recognition network, an entity recognition model in the target scene is constructed. This entity recognition model is used to perform entity recognition on the text in the target scene.
[0046] Practice has shown that the model construction scheme proposed in this application can have at least the following beneficial effects: ① It can improve the accuracy of character features, that is, improve the accuracy of feature extraction through the language model in the entity recognition model; ② It can obtain more accurate entity recognition results based on more accurate character features, that is, improve the accuracy of entity recognition of text; ③ It can enable the entity recognition model to not only recognize reference entities in the text, but also support the recognition of erroneous entities in the text.
[0047] In practical implementation, the model construction scheme mentioned above can be executed by a computer device, which can be a terminal or a server. The terminal mentioned here can include, but is not limited to, smartphones, tablets, laptops, desktop computers, smartwatches, smart voice interaction devices, smart home appliances, in-vehicle terminals, and aircraft. Various clients (apps) can run within the terminal, such as video playback clients, social media clients, browser clients, news feed clients, educational clients, and so on. The server mentioned here can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, etc. Cloud computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. Furthermore, the computer device mentioned in the embodiments of this application can be located outside or inside the blockchain network, and there is no limitation on this. The so-called blockchain network is a network composed of a peer-to-peer network (P2P network) and a blockchain. The blockchain refers to a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. In essence, it is a decentralized database, which is a series of data blocks (or blocks) linked together using cryptographic methods.
[0048] Alternatively, in other embodiments, the model building scheme mentioned above can also be jointly executed by the server and the terminal; the terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited herein. For example, the terminal can be responsible for obtaining the error correction dataset for optimizing the language model and sending the error correction dataset to the server, so that the server can use each error correction text in the error correction dataset to optimize the language model, obtain the optimized language model, and send the optimized language model to the terminal. Then, the terminal constructs an entity recognition model based on the optimized language model and the entity recognition network, such as... Figure 1bAs shown. For example, the terminal can acquire the error correction dataset and send it to the server. The server can then use each corrected text in the dataset to optimize the language model, obtaining an optimized language model. Based on this optimized language model and the entity recognition network, an entity recognition model can be constructed, and so on. It should be understood that this is merely an illustrative example of two scenarios where the terminal and server jointly execute the above model construction scheme, and is not an exhaustive list.
[0049] Based on the above description of the model building scheme, this application proposes a model building method, which can be executed by the aforementioned computer device (terminal or server); or, the model building method can be executed jointly by the terminal and the server. For ease of explanation, the following description will use the execution of the model building method by a computer device as an example; please refer to [link to relevant documentation]. Figure 2 The model construction method may include the following steps S201-S205:
[0050] S201, Obtain the error correction dataset for optimizing the language model. The error correction dataset includes multiple error correction texts in the target scene, and each error correction text includes at least one error entity, and an error entity includes at least one character.
[0051] The language model can be a pre-trained language model with bidirectional feature representations, i.e., the BERT (Bidirectional Encoder Representation from Transformers) series of models. BERT emphasizes that it no longer uses traditional unidirectional language models or shallow concatenation of two unidirectional language models for pre-training, but instead uses a new MLM (Masked Language Model) to generate deep bidirectional language representations. Optionally, the above language model can also be a unidirectional language model or an autoregressive pre-trained language model with bidirectional feature representations, etc.; this application does not limit this. It should be noted that, for ease of explanation, the following descriptions will use the BERT series of models as an example.
[0052] In this application embodiment, the aforementioned erroneous entity refers to an entity with a spelling error, and the types of errors of the erroneous entity include, but are not limited to, phonetic (i.e., pronunciation), visual (i.e., character shape), sequence confusion (i.e., disordered word order), repetition (i.e., extra characters), and omission (i.e., missing characters), etc.; this application does not limit the specific type of erroneous entity. Correspondingly, the aforementioned target scenarios include, but are not limited to, medical scenarios, product recommendation scenarios, and shopping scenarios, etc.; this application does not limit these; for ease of explanation, the following description will use the medical scenario as an example.
[0053] For example, when the target scenario is a medical scenario, the error correction text, including the error entities under each error type, can be as shown in Table 1:
[0054] Table 1
[0055]
[0056]
[0057] In this embodiment, a correction text can be a sentence, and all erroneous entities in the correction text are misspelled. It should be noted that the correction dataset is a large-scale expert-annotated dataset proposed in this application embodiment; optionally, this correction dataset can contain approximately 200,000 correction sample pairs (i.e., correction text and corresponding correction information), as shown in Table 1, where each row of data contains a sentence and the corrected content, which can serve as a correction sample pair. Furthermore, compared to existing open-domain correction datasets, when the target scenario is a medical scenario, the correction dataset proposed in this application embodiment can contain a large number of medical queries collected from the target medical dictionary, and the correction dataset in this application embodiment mainly targets the correction of medical entities.
[0058] In specific implementations, the methods for obtaining the aforementioned error correction dataset include, but are not limited to, the following:
[0059] The first method of acquisition: The computer device can first obtain the data download link of the error correction dataset, and then download the error correction dataset according to the data download link to obtain the error correction dataset.
[0060] The second method of acquisition: The computer device itself stores error correction text in its storage space. The computer device can then select multiple error correction texts from the stored error correction texts and use the selected error correction texts as the error correction texts in the error correction dataset, and so on.
[0061] S202: For any one of the multiple error-correcting texts, call the language model to extract features from each character in the error-correcting text and obtain the character features of each character.
[0062] It should be noted that the computer device can call the language model to extract features from each character (token) in each error-corrected text, thereby obtaining the character features of each character in each error-corrected text.
[0063] In the embodiments of this application, the character features of a character may include the character's own features (i.e., internal features), contextual features, part-of-speech features, etc.; this application does not limit the information contained in the character features.
[0064] S203, for any character feature in any error-corrected text, obtain the probability difference between the probability of recognizing any character as various characters based on the character feature of any character and the reference probability of any character being various characters.
[0065] The reference probability of any character being one of the other characters refers to the actual probability that any given character is one of the other characters, used to indicate the actual character represented by that given character. For example, when any of the above-mentioned error-correction texts is "the process of removing wisdom teeth", and any given character is "tooth", the reference probability of that given character being one of the other characters can be (0, 0, 1, 0, 0, 0). Based on this, the computer device can calculate cross-entropy loss or negative log-likelihood loss, etc., based on the probability of recognizing any given character as one of the other characters and the reference probability of any given character being one of the other characters, to obtain the probability difference between the probability of recognizing any given character as one of the other characters based on the character features of any given character and the reference probability of any given character being one of the other characters; that is, the computer device can calculate the corresponding loss value and use the calculated loss value as the above-mentioned probability difference.
[0066] S204. Optimize the feature extraction parameters in the language model in the direction of reducing probability differences to obtain the optimized language model.
[0067] S205, based on the optimized language model and entity recognition network, constructs an entity recognition model for the target scene. The entity recognition model is used to perform entity recognition on the text in the target scene.
[0068] In this application embodiment, when the target scenario is a medical scenario, the aforementioned entity recognition model can also be called Medical Entity Recognition (NER) or an error-based medical entity recognition model, i.e., named entity recognition in the medical field, which refers to extracting important medical entities (such as diseases, symptoms, etc.) from medical text. The result (i.e., the entity recognition result) is the foundation for subsequent medical tasks such as relation extraction. Correspondingly, in medical scenarios (i.e., the clinical medical field, etc.), the ability to accurately identify named entities in electronic medical data is of great significance for building a comprehensive medical knowledge base, accurate object profiling, and intelligent medical decision support. Similarly, in the field of medical entity error correction, Medical Entity Recognition (NER) is also one of the important basic models. The entity recognition in this application refers to named entity recognition.
[0069] Specifically, in medical AI applications, many scenarios, such as question-and-answer, dialogue, and search, require text input. Due to the uncommon nature of medical terminology, users often input incorrect text, leading to suboptimal answers and feedback from the system. Therefore, building more accurate entity recognition models to precisely identify entities within text and pinpoint erroneous entities is of great significance. Furthermore, in medical scenarios, computer devices can use entity recognition models to identify erroneous entities in text and correct input from users in medical settings, enabling the system to return accurate answers and feedback based on the corrected entities.
[0070] It should be noted that named entity recognition belongs to the sequence labeling task in natural language processing. In the medical context, it refers to identifying specific named medical entities from text, such as disease names, drug names, and symptoms. In this case, a natural language sequence (i.e., text containing at least one character) can be input into the entity recognition model, and the computer device can then provide the corresponding label sequence through the entity recognition model. It should be noted that the embodiments of this application can use BIO annotation for parsing, or BIOSE annotation for parsing, etc.; this application does not limit this. Optionally, when using BIO annotation for parsing, START and END can be added simultaneously to make the transition matrix more robust, where START represents the beginning of a sentence and END represents the end of a sentence. In this case, there are a total of 5 annotation labels: [B, I, O, START, END], where B can be used to represent the beginning of an entity, I can be used to represent the middle or end of an entity, and O can be used to represent entities that do not belong to an entity.
[0071] In specific implementations, the entity recognition model may include a language model and an entity recognition network. The entity recognition network may include a CRF (conditional random fields) layer, or a recurrent neural network, etc., which is not limited in this application. For example, taking an entity recognition network including a CRF layer as an example, when performing entity recognition on any text, the computer device can use the Emission_score, output based on the character features of each character in that text, as input to the CRF layer. Correspondingly, it can output the most likely predicted label sequence that meets the label transfer constraints, and further obtain the labels of each character in that text, such as... Figure 3 As shown; in this case, any text may include 5 characters, and the labels of each character in the text may be B, I, O, O and B in sequence.
[0072] This application embodiment can obtain an error correction dataset for optimizing a language model. This dataset includes multiple error-corrected texts from a target scene, and each text includes at least one erroneous entity. Then, the language model can be invoked to extract features from each character in each error-corrected text, obtaining character features for each character. Based on these character features, the language model is optimized to obtain an optimized language model, improving its performance and resulting in a more accurate entity recognition model. This improves the accuracy of character features extracted by the language model in the entity recognition model, thereby enhancing entity recognition accuracy. Furthermore, since character features can also reflect contextual relationships, the language model optimized using the error correction dataset can more accurately extract character features from texts containing erroneous entities, more precisely reflecting the contextual relationships involved. This enables the entity recognition model to identify both reference entities (i.e., correctly spelled entities) and erroneous entities in the text.
[0073] Please see Figure 4 This is a flowchart illustrating another model building method provided in this application. This model building method can be executed by the computer device (terminal or server) mentioned above; or, it can be executed jointly by the terminal and the server. For ease of explanation, the following description will use the execution of this model building method by a computer device as an example; please refer to [link to relevant documentation]. Figure 4 The model construction method may include the following steps S401-S409:
[0074] S401, Obtain an error correction dataset for optimizing the language model. The error correction dataset includes multiple error correction texts in the target scene, and each error correction text includes at least one error entity, and an error entity includes at least one character.
[0075] In this embodiment, the computer device can acquire a training dataset for optimizing an initial language model. This training dataset includes multiple training texts from a target scene, and each training text includes at least one reference entity. Each reference entity includes at least one training character. Based on this, the initial language model can be optimized using multiple training texts from the training dataset to obtain the aforementioned language model. A reference entity refers to a correctly spelled entity, i.e., an entity with no spelling errors, and the reference entity can also be called a correct entity.
[0076] Specifically, when optimizing the initial language model using multiple training texts from the training dataset to obtain the aforementioned language model, for any training text among the multiple training texts, the computer device can call the initial language model to extract features from each training character in that training text, obtaining the character features of each training character. Then, for the character features of any training character in that training text, the probability of recognizing any training character based on its character features as any training character can be obtained, along with the reference probability that any training character is any training character. Based on this, the feature extraction parameters in the initial language model can be optimized in the direction of reducing the training probability difference to obtain the language model.
[0077] It should be noted that when the language model is a BERT series model and the target scenario is a medical scenario, directly using the original BERT model as the language model will result in the named entity recognition results not meeting expectations. Therefore, in this embodiment, a pre-trained Chinese medical language model, MedBert (a BERT model pre-trained using a training dataset in a medical scenario), can be used as the language model to be optimized using the aforementioned error correction dataset. In this case, the BERT model can be used as the initial language model, and after pre-training the initial language model using the training dataset (i.e., normal medical data), the language model (i.e., MedBert) is obtained. Then, the labeled error correction dataset (i.e., medical error correction data) can be used to further fine-tune the language model to obtain the optimized language model.
[0078] At the model level, the pre-trained Chinese medical language model MedBERT (the aforementioned language model) and the ordinary BERT (the original BERT) model share the same aspect: both employ the encoder part of a transformer (a type of network model). However, the difference lies in the masking method used for the training data and the fact that the training data in this embodiment originates from medical scenarios. Based on this, a computer device can mask the characters in any of the aforementioned training texts to update the training text, obtaining an updated training text. This allows for the extraction of character features from each training character in the updated training text, thereby optimizing the initial language model. The specific implementation process for masking characters in any training text is the same as that for masking characters in any error-corrected text, as detailed below, and will not be elaborated upon here.
[0079] S402: For any one of the multiple error-correcting texts, call the language model to extract features from each character in the error-correcting text and obtain the character features of each character.
[0080] It should be noted that when the language model mentioned above is a BERT series model (such as a Chinese medical language model), the computer device can obtain the first mask position in any of the above-mentioned error-corrected texts according to a preset mask probability; and determine the second mask position in any of the error-corrected texts based on the target masking method and the first mask position; wherein, the target masking method is used to indicate the selection method of the mask position in any of the error-corrected texts, and the target masking method includes at least one of the following: character masking method, whole-word masking method, and entity masking method; based on this, the computer device can use mask characters to mask the characters at the second mask position to update any of the above-mentioned error-corrected texts, wherein each character in the updated error-corrected text is used for feature extraction. In this case, the computer device can first mask the characters in any of the above-mentioned error-corrected texts to obtain the updated error-corrected text; and call the language model to extract features from each character in the updated error-corrected text to obtain the character features of each character in the updated error-corrected text.
[0081] The masking probability can be set empirically or according to actual needs, and this application does not limit it; for example, the masking probability can be 15% or 20%, etc. Furthermore, the masking character can be [mask], a randomly selected character, or the character to be masked itself, and this application does not limit it. It should be noted that, for any corrected text, when there are multiple second masking positions in that corrected text, the masking characters used to mask the characters at the second masking positions in that corrected text can be the same or different, and this application does not limit it.
[0082] In the embodiments of this application, the method for determining the second mask position in any of the above-mentioned error-corrected texts includes: using the first mask position as the second mask position in any of the error-corrected texts; or, determining the complete word composed of the characters at the first mask position and using the character position of the complete word as the second mask position in any of the error-corrected texts; or, determining the target entity composed of the characters at the first mask position and using the character position of the target entity as the second mask position in any of the error-corrected texts.
[0083] Specifically, if the target masking method includes a character masking method, the computer device can use the first mask position as the second mask position in any error-correcting text; if the target masking method includes a whole-word masking method, the computer device can determine the complete word formed by the characters at the first mask position and use the character position of the complete word as the second mask position in any error-correcting text; if the target masking method includes an entity masking method, the computer device can determine the target entity formed by the characters at the first mask position and use the character position of the target entity as the second mask position in any error-correcting text. Furthermore, if the target masking method includes a mixed masking method of character masking, whole-word masking, and entity masking, the computer device can sequentially select one masking method to determine the second mask position corresponding to a first mask position, or it can randomly select a masking method to determine the second mask position corresponding to a first mask position; this application does not limit this.
[0084] It should be noted that computer devices can segment a complete word into several sub-words (i.e., characters) based on WordPiece (a word segmentation method). When generating training samples (i.e., masking the characters in any error-correcting text or training text to obtain the updated text), in the character masking method, these separated sub-words will be randomly masked; in the whole-word masking method and the entity masking method, if some sub-words of a complete word or entity are masked, then other sub-words belonging to the same word or corresponding entity will also be masked.
[0085] For example, when the target masking method includes entity masking, suppose any error-correcting text contains the symptom entity "stomach hurts", and "stomach hurts" can be segmented into "stomach" and "aches", or into "stomach", "child", "very" and "aches", then when the character position of one of the sub-words is used as the first mask position, the computer device can use the character position of the entity as the second mask position, so that as long as one sub-word of the entity is masked, the rest of the entity will be masked.
[0086] For example, when the target masking method includes the whole word masking method, assuming that any error-correcting text contains the complete word "useful", then when the character position of one of the sub-words in the word is used as the first mask position, the computer device can use the character position of the word as the second mask position, so that when one sub-word in the word is masked, the rest of the word will also be masked.
[0087] S403, for any character feature in any error-corrected text, obtain the probability difference between the probability of recognizing any character as various characters based on the character feature of any character and the reference probability of any character being various characters.
[0088] It should be noted that after the computer device performs masking processing on the characters in any of the above-mentioned error-corrected texts to obtain the character features of each character in the updated error-corrected texts, the computer device can obtain the corresponding probability difference for the character features of any character in the updated error-corrected texts in order to optimize the feature extraction parameters in the language model.
[0089] S404. Optimize the feature extraction parameters in the language model in the direction of reducing probability differences to obtain the optimized language model.
[0090] In this embodiment of the application, when optimizing the language model, the Next Sentence Prediction task can be removed, and only the Masked Language Model task can be retained. In other words, the computer device can only optimize the language model in the entity recognition model.
[0091] It should be noted that when using the error correction dataset to optimize the language model (i.e., during fine-tuning), the computer device can be set to a low training learning rate, such as 3*10^-5 or 4*10^-5, etc., and this application does not limit this.
[0092] S405, based on the optimized language model and entity recognition network, constructs an entity recognition model for the target scene. The entity recognition model is used to perform entity recognition on text in the target scene.
[0093] In this embodiment of the application, after fine-tuning the language model using all the error-corrected texts in the error-corrected dataset, the computer device can use the test error-corrected dataset to test and verify the effect of the entity recognition model constructed in this embodiment of the application on entity recognition.
[0094] S406, Obtain the obfuscation set under the target scenario. The obfuscation set includes multiple error correction character pairs under the target scenario, and each error correction character pair includes at least one erroneous entity and the corresponding entity after the error has been corrected.
[0095] Specifically, the aforementioned error correction dataset may also include error correction information corresponding to the erroneous entities in each error correction text. The target scenario refers to a medical scenario, and each error correction text includes at least one erroneous medical entity. Based on this, when obtaining the obfuscation set under the target scenario, the computer device can traverse multiple error correction texts, determine the corrected target medical entity based on the target erroneous medical entity and the corresponding error correction information in the currently traversed error correction text, and based on the target erroneous medical entity and the target medical entity, determine the erroneous entity and the corresponding corrected entity in the target error correction character pair, and save the target error correction character pair to the obfuscation set. After traversing multiple error correction texts, the obfuscation set under the target scenario is obtained. The error correction information corresponding to an erroneous entity can be the corrected character of the erroneous character in the corresponding erroneous entity, or it can be the corrected entity of the corresponding erroneous entity. This application does not limit this, as shown in Table 1.
[0096] In this embodiment, when a computer device determines the erroneous entity and its corresponding corrected entity in a target error-correction character pair based on the target erroneous medical entity and the target medical entity, if the obfuscation set already includes the target medical entity (i.e., the obfuscation set already includes the target error-correction character pair), the computer device can update the target error-correction character pair using the target erroneous medical entity, so that the target error-correction character pair includes the target erroneous medical entity. If the obfuscation set does not include the target medical entity (i.e., the obfuscation set does not include the target error-correction character pair), the computer device can generate the target error-correction character pair using the target erroneous medical entity and the target medical entity. Correspondingly, when the computer device saves the target error-correction character pair to the obfuscation set, if the obfuscation set already includes the target error-correction character pair, the computer device can update the target error-correction character pair in the obfuscation set; if the obfuscation set does not include the target error-correction character pair, the computer device can add the target error-correction character pair to the obfuscation set.
[0097] It should be noted that there are several common confusion sets in open-domain error correction tasks, but these confusion sets cannot be well applied to the medical field. Therefore, based on the error correction dataset (i.e., the medical error correction dataset), this application embodiment constructs a dictionary-style confusion set (i.e., the medical confusion set). This confusion set may include a large number of correct-incorrect character pairs (i.e., error correction character pairs) involved in all spelling errors in the error correction dataset. Based on this, given a feature (i.e., an entity) that is prone to errors in the medical field, the common incorrect entities corresponding to the corresponding entity can be easily found according to the confusion set.
[0098] Furthermore, through statistical analysis of the confusion set, the confusion set proposed in this application embodiment may include 2623 different misspelled characters, and most Chinese characters have one to twenty corresponding error characters; for example, the misspelled characters in the confusion set are shown in Table 2:
[0099] Table 2: Seven examples of medical confusion
[0100]
[0101] Accordingly, this application embodiment analyzed all 81,020 high-frequency medical entities appearing in the confusion set, which means that these entities are more likely to be misspelled in medical scenarios; for example, the top 5 high-frequency entities found from the confusion set and the corresponding set of incorrect entities in this application embodiment are shown in Table 3:
[0102] Table 3: Top 5 High-Frequency Entities
[0103]
[0104] S407, call the language model in the entity recognition model to extract features from each character in the text to be processed, and obtain the character features of each character in the text to be processed.
[0105] In the entity recognition model, the language model refers to the optimized language model. In other words, the computer device can call the optimized language model to extract features from each character in the text to be processed, and obtain the character features of each character in the text to be processed.
[0106] S408, the entity recognition network in the entity recognition model is invoked to perform entity recognition on the text to be processed based on the character features of each character in the text to be processed, and the recognized entities of the text to be processed are obtained.
[0107] S409 If a target erroneous entity matching the identified entity is found in the confusion set, the entity corresponding to the target erroneous entity is used to replace the identified entity in order to achieve error correction processing of the text to be processed.
[0108] In this embodiment, the module that uses the obfuscation set to perform error correction on the text to be processed can be called the obfuscation set replacement module, and when the target scenario is a medical scenario, it can also be called the medical obfuscation set replacement module. Based on this, the computer device can first identify the recognition entities in the text to be processed through the entity recognition model, and then check whether the recognition entities are common erroneous entities in the obfuscation set through the obfuscation set replacement module. If so, the corresponding entities can be directly replaced.
[0109] In one embodiment, the computer device may use the erroneous entity that is the same as the recognized entity in the confusion set as the target erroneous entity matching the recognized entity. For example, assuming that the confusion set comprises the error-corrected character pair "common cold-(ganchang, ganmao, ganmao)", and the recognized entity is "ganmao", the computer device can find the target erroneous entity "ganmao" matching the recognized entity in the confusion set, thereby replacing the recognized entity with the entity "common cold" corresponding to the target erroneous entity.
[0110] In another embodiment, if there is no erroneous entity identical to the recognized entity in the confusion set, and there is no reference entity identical to the recognized entity in the confusion set (that is, the entity obtained after error correction for any erroneous entity), the computer device may use any erroneous entity in the confusion set that has the same pinyin as the recognized entity as the target erroneous entity matching the recognized entity, etc. For example, assuming that the confusion set comprises the error-corrected character pair "common cold-(ganchang, ganmao, ganmao)", and the recognized entity is "ganmao", the computer device can use the erroneous entity "ganmao" or the erroneous entity "ganmao" in the confusion set as the target erroneous entity, thereby replacing the recognized entity with the entity "common cold" corresponding to the target erroneous entity.
[0111] Optionally, in other embodiments, the computer device may not perform steps S408 and S409, but instead adopts a hard (direct) mode to perform error correction processing on the text to be processed; that is, after extracting features for each character in the text to be processed by using the language model in the entity recognition model (that is, the optimized language model), the computer device can directly use the keyword matching method to search whether any erroneous word that has appeared in the confusion set exists in the text to be processed based on the character features of each character in the text to be processed. If there is an erroneous word that has appeared in the confusion set, the erroneous word is directly replaced with the correct word.
[0112] It should be noted that in other embodiments, if the target erroneous entity is not found in the obfuscation set, the computer device may not execute step S409. Further, in other embodiments, if the target erroneous entity is not found in the obfuscation set and the identified entity is not found in the target reference entity set, at least one candidate replacement entity corresponding to the identified entity is selected from the dictionary. The target reference entity set includes multiple reference entities. Then, the distance between each candidate replacement entity and the identified entity is determined, and the closest candidate replacement entity is used to replace the identified entity, thereby achieving error correction processing of the text to be processed. Optionally, the computer device may select at least one candidate replacement entity matching the pinyin (i.e., pronunciation) of the identified entity from the dictionary; it may also select at least one candidate replacement entity matching the shape of the identified entity from the dictionary; it may also select at least one candidate replacement entity matching the pinyin and shape of the identified entity from the dictionary, etc.; this application does not limit this.
[0113] The target reference entity set can also be called a common vocabulary or a common correct entity set. In this case, the computer device can first check whether the identified entity exists in both the target reference entity set and the confusion set. If it exists in the target reference entity set, it remains unchanged; if it exists in the confusion set, it is directly replaced and corrected; if it does not exist in either set, i.e., it does not exist in either the target reference entity set or the confusion set, then a candidate replacement entity is determined to replace the identified entity. Optionally, the implementation of replacing the identified entity with a candidate replacement entity can be completed by an entity correction module outside the confusion set.
[0114] In practical implementations, computer devices can utilize tree structures (such as Burkhard-Keller trees, a tree-based data structure) for recall to obtain the nearest candidate replacement entity (i.e., the candidate replacement entity closest to the identified entity). In this case, the entities in the dictionary are stored in the form of a tree structure, which includes multiple nodes and at least one edge. A node represents an entity in the dictionary (i.e., the number of nodes in the tree structure is the same as the number of entities in the dictionary). An edge connects two nodes and includes an integer weight indicating the edit distance between the corresponding entities of the two connected nodes. For example, assuming an edge from node u to node v has some edge weight w (i.e., the edge connecting node u to node v includes weight w), then w is the edit distance required to convert entity (i.e., string) u into entity v. The Burkhard-Keller tree is used for spell checking based on the edit distance concept and also for approximate string matching; based on this data structure, various automatic correction features in many software programs can be implemented.
[0115] The edit distance between two entities (i.e., strings) refers to the minimum number of steps required to transform one entity into another using only insertion, deletion, and replacement operations. Similarly, the edit distance from string A to string B refers to the minimum number of steps required to transform A into B using only insertion, deletion, and replacement operations. For example, transforming FAME into GATE requires two steps (two replacements), while transforming GAME into ACM requires three steps (deleting G and E and then adding C).
[0116] It should be noted that the core idea of BK trees is: let d(x, y) represent the edit distance from entity x to y, then d(x, y) = 0 if and only if x = y (edit distance is 0, i.e. entities are equal), d(x, y) = d(y, x) (the minimum number of steps from x to y is equal to the minimum number of steps from y to x), and d(x, y) + d(y, z) >= d(x, z) (the number of steps required to change from x to z will not exceed the number of steps from x to y and then to z). This property is called the triangle inequality, which means that, just like a triangle, the sum of any two sides must be greater than the third side.
[0117] Accordingly, for any candidate replacement entity among at least one candidate replacement entity, the computer device can determine the node corresponding to any candidate replacement entity in the structure tree, and the target edge connecting it to the node corresponding to the identified entity; and use the edit distance indicated by the target edge as the distance between any candidate replacement entity and the identified entity. In this case, after obtaining at least one candidate replacement entity, the nearest candidate replacement entity can be selected using the structure tree as the final entity replacement word to replace the identified entity.
[0118] For example, such as Figure 5a As shown, assuming that at least one candidate replacement entity includes entity A, entity B and entity C, the identified entity is entity D, and each edge in the structure tree indicates the edit distance between each entity and the identified entity, then the computer device can determine that the distances between entity A, entity B and entity C and the identified entity are 5, 8 and 3 respectively, and thus replace the identified entity with the candidate replacement entity (i.e. entity C) that is closest to it.
[0119] It should be noted that when the target scenario is a medical scenario, the overall framework of the model construction method proposed in this application mainly includes a medical named entity recognition model (i.e., entity recognition model) based on medical error correction data (i.e., error correction dataset) and a large-scale medical language model, as well as a post-processing model. The post-processing model includes a medical entity detection model (i.e., a medical confusion set replacement module and an entity error correction module outside the confusion set), such as... Figure 5bAs shown, this application utilizes an error correction dataset to generate an confusion set, and then uses a medical pre-trained language model (i.e., the aforementioned language model) and a medical entity discovery algorithm fine-tuned for medical error correction (i.e., an entity recognition model) to identify erroneous medical entities in the sentence. Then, it replaces erroneous medical entities in the medical confusion set and uses a structure tree for recall and replacement to further process the sentence (i.e., the text to be processed), finally outputting the corrected sentence. In other words, for object input error correction in medical scenarios, this application proposes a method based on a medical pre-trained language model and entity recognition augmented with confusion sets, combined with statistical rule post-processing, to detect and correct object input errors.
[0120] Optionally, if the target erroneous entity is not found in the confusion set and the identified entity is not found in the target reference entity set, the computer device may also use the recall method of the language model to replace the identified entity, or calculate the distance between each candidate replacement entity and the identified entity based on the entity features of each entity, thereby replacing the identified entity with the candidate replacement entity with the closest one, etc.; this application does not limit this.
[0121] It should be noted that when the target scenario is a medical scenario, the model construction method proposed in this application can be applied to various medical functions or products to correct erroneous entities in the input text, thereby achieving more accurate response results, such as automatically and effectively classifying short texts based on the corrected entities. For example, it can be applied to health assistants. When users input consultation searches or pre-diagnosis, medical entities, such as drug names, disease names, and symptoms, are often mistyped due to their technical nature. If the system does not correct these entities, it is very likely to return an incorrect answer or miss the correct one. Furthermore, it can be applied to medical question-and-answer dialogues. The questions input by users may contain incorrect medical entities, requiring the use of a correction engine to first correct the input, thereby expanding the correct recall of the question-and-answer or dialogue, and so on.
[0122] In this application, to better illustrate the performance of the entity recognition model, when the target scenario is a medical scenario, this application can use a total of 30,000 medical entity error correction data generated by rules to experimentally verify the entity recognition model, and this application can use the false correction rate and recall rate as hard indicators. The false correction rate refers to the ratio of correct sentences to corrected errors, and a large false correction rate will have a negative impact on the system and user experience; the recall rate refers to the ratio of all incorrect sentences to corrected errors. Based on this, the goal of this application's embodiments is to ensure that the number of corrected sentences is far greater than the number of incorrectly corrected sentences, i.e., K*RECALL >> (1-K)*FAR; where K is the sentence error probability, RECALL refers to the recall rate, and FAR refers to the false correction rate.
[0123] Based on this, this application also conducted experimental analysis on the aforementioned medical entity error correction data using common basic models and rule-based and common confusion set replacement models, respectively, to compare the recall and false positive rates with those of the entity recognition model in this application. The common basic model refers to using only a general language model for error correction verification, while the rule-based and confusion set replacement model refers to using only a common confusion set and an entity recognition model (including the general language model) for error correction. It is evident that this application achieves the best error correction effect after fine-tuning the error correction dataset and replacing the confusion set, making it a relatively effective method for correcting errors in medical scenarios. Specific comparison results are shown in Table 4.
[0124] Table 4
[0125] Common basic models 78% 78% Rule-based and confusion set replacement model 40% 15% The model proposed in this application 85% 12%
[0126] This application embodiment can train and predict entity recognition in the target scene using a language model and an error correction dataset. In other words, the language model can be fine-tuned using the error correction dataset, allowing the optimized language model to better adapt to erroneous entities (i.e., error-corrected entities). An entity recognition model is then constructed based on this optimized language model. This enables the identification of misspelled entities during entity recognition (i.e., entity extraction), preparing for the subsequent replacement of incorrect entities with correct entities using a confusion set. Furthermore, this application embodiment can obtain a confusion set for erroneous entity replacement and utilize methods such as BK trees to recall uncommon erroneous entities, thereby achieving error correction processing of the text to be processed and enabling more accurate results to be retrieved based on the corrected text.
[0127] Based on the description of the relevant embodiments of the above model building method, this application also proposes a model building apparatus, which can be a computer program (including program code) running on a computer device. This model building apparatus can execute... Figure 2 or Figure 4 The model construction method shown; please refer to [link / reference]. Figure 6 The model building apparatus can operate the following units:
[0128] The acquisition unit 601 is used to acquire an error correction dataset for optimizing the language model. The error correction dataset includes multiple error correction texts in the target scene, and each error correction text includes at least one error entity, and an error entity includes at least one character.
[0129] The processing unit 602 is used to call the language model for any one of the plurality of error-corrected texts, and extract the features of each character in the error-corrected text to obtain the character features of each character.
[0130] The processing unit 602 is further configured to, for any character feature in any error-corrected text, obtain the probability difference between the probability of identifying any character as each of the characters based on the character feature of any character and the reference probability of any character being each of the characters.
[0131] The processing unit 602 is further configured to optimize the feature extraction parameters in the language model in the direction of reducing the probability difference, so as to obtain an optimized language model;
[0132] The processing unit 602 is further configured to construct an entity recognition model for the target scene based on the optimized language model and entity recognition network, wherein the entity recognition model is used to perform entity recognition on the text in the target scene.
[0133] In one embodiment, the acquisition unit 601 can also be used to: acquire a training dataset for optimizing the initial language model, the training dataset including multiple training texts in the target scene, and each training text including at least one reference entity, and a reference entity including at least one training character.
[0134] The processing unit 602 can also be used to: for any training text among the plurality of training texts, call the initial language model to extract features from each training character in the training text to obtain the character features of each training character;
[0135] For any training character in any training text, the probability of identifying the training character based on the character features of the training character as each training character is obtained, and the training probability difference between the probability that the training character is the reference probability of each training character.
[0136] The feature extraction parameters in the initial language model are optimized in the direction of reducing the difference in training probabilities to obtain the language model.
[0137] In another embodiment, the processing unit 602 can also be used to: obtain the first mask position in any of the error-correcting texts according to a preset mask probability;
[0138] Based on the target masking method and the first masking position, a second masking position in any corrected text is determined; wherein, the target masking method is used to indicate the selection method of the masking position in any corrected text, and the target masking method includes at least one of the following: character masking method, whole word masking method, and entity masking method;
[0139] The characters at the second mask position are masked using a mask character to update any of the error-correcting texts, wherein each character in the updated error-correcting text is used for feature extraction.
[0140] In another implementation, the method for determining the position of the second mask in any error-corrected text includes:
[0141] The first mask position is used as the second mask position in any of the error-corrected texts;
[0142] Alternatively, determine the complete word formed by the characters at the first mask position, and use the character position of the complete word as the second mask position in any of the error-corrected texts;
[0143] Alternatively, the target entity composed of the characters at the first mask position can be determined, and the character position where the target entity is located can be used as the second mask position in any of the error-correcting texts.
[0144] In another embodiment, the acquisition unit 601 can also be used to: acquire the obfuscation set under the target scene, the obfuscation set including multiple error correction character pairs under the target scene, and an error correction character pair including at least one erroneous entity and the corresponding entity after the at least one erroneous entity has been corrected;
[0145] The processing unit 602 can also be used to: call the language model in the entity recognition model, extract features from each character in the text to be processed, and obtain the character features of each character in the text to be processed;
[0146] The entity recognition network in the entity recognition model is invoked to perform entity recognition on the text to be processed based on the character features of each character in the text to be processed, so as to obtain the recognized entities of the text to be processed.
[0147] If a target erroneous entity matching the identified entity is found in the obfuscation set, the identified entity is replaced with the entity corresponding to the target erroneous entity to achieve error correction processing of the text to be processed.
[0148] In another embodiment, the processing unit 602 can also be used to: if the target erroneous entity is not found in the confusion set and the identified entity is not found in the target reference entity set, then select at least one candidate replacement entity corresponding to the identified entity from the dictionary, wherein the target reference entity set includes multiple reference entities;
[0149] The distance between each candidate replacement entity in the at least one candidate replacement entity and the identified entity is determined, and the identified entity is replaced with the candidate replacement entity with the closest distance, so as to realize the error correction processing of the text to be processed.
[0150] In another embodiment, the error correction dataset further includes error correction information corresponding to the erroneous entities in each error correction text, the target scenario refers to a medical scenario, and each error correction text includes at least one erroneous medical entity; when acquiring the obfuscation set under the target scenario, the acquisition unit 601 may specifically be used for:
[0151] Traverse the multiple error correction texts, and based on the target erroneous medical entity in the currently traversed error correction text and the error correction information corresponding to the target erroneous medical entity, determine the target medical entity after the target erroneous medical entity has been corrected.
[0152] Based on the target erroneous medical entity and the target medical entity, determine the erroneous entity and the corresponding corrected entity in the target error correction character pair, and save the target error correction character pair to the obfuscation set;
[0153] After traversing the multiple error-correcting texts, the obfuscation set for the target scenario is obtained.
[0154] According to one embodiment of this application, Figure 2 or Figure 4 Each step involved in the method shown can be derived from... Figure 6 The model building apparatus shown is executed by the individual units within it. For example, Figure 2 Step S201 shown can be performed by Figure 6 The acquisition unit 601 shown is executed, and steps S202-S205 can all be performed by... Figure 6 The processing unit 602 shown executes this. For example, Figure 4 Steps S401 and S406 shown can both be performed by Figure 6 The acquisition unit 601 shown is executed, and steps S402-S405 and steps S407-S409 can be performed by... Figure 6 The processing unit 602 shown executes, etc.
[0155] According to another embodiment of this application, Figure 6The various units in the model building apparatus shown can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The above-mentioned units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the model building apparatus may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.
[0156] According to another embodiment of this application, the following can be achieved by running on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM), a device capable of performing operations such as... Figure 2 or Figure 4 The computer program (including program code) for each step involved in the corresponding method shown, to construct such... Figure 6 The model building apparatus shown herein, and the model building method for implementing the embodiments of this application, are described. The computer program may be recorded on, for example, a computer storage medium, loaded onto the aforementioned computing device via the computer storage medium, and run therein.
[0157] This application embodiment can obtain an error correction dataset for optimizing a language model. This dataset includes multiple error-corrected texts from a target scene, and each text includes at least one erroneous entity. Then, the language model can be invoked to extract features from each character in each error-corrected text, obtaining character features for each character. Based on these character features, the language model is optimized to obtain an optimized language model, improving its performance and resulting in a more accurate entity recognition model. This improves the accuracy of character features extracted by the language model in the entity recognition model, thereby enhancing entity recognition accuracy. Furthermore, since character features can also reflect contextual relationships, the language model optimized using the error correction dataset can more accurately extract character features from texts containing erroneous entities, more precisely reflecting the contextual relationships involved. This enables the entity recognition model to identify both reference entities (i.e., correctly spelled entities) and erroneous entities in the text.
[0158] Based on the description of the above method and apparatus embodiments, this application also provides a computer device. Please refer to... Figure 7The computer device includes at least a processor 701, an input interface 702, an output interface 703, and a computer storage medium 704. The processor 701, input interface 702, output interface 703, and computer storage medium 704 within the computer device can be connected via a bus or other means.
[0159] Computer storage medium 704 can be stored in the memory of a computer device. The computer storage medium 704 is used to store computer programs, which include program instructions. The processor 701 is used to execute the program instructions stored in the computer storage medium 704. The processor 701 (or CPU (Central Processing Unit)) is the computing and control core of the computer device, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve a corresponding method flow or corresponding function. In one embodiment, the processor 701 described in this application embodiment can be used to perform a series of model constructions, specifically including: acquiring an error correction dataset for optimizing a language model, the error correction dataset including multiple error correction texts in a target scene, and each error correction text including at least one error entity, an error entity including at least one character; for any one of the multiple error correction texts, calling the language model to optimize the language model. Feature extraction is performed on each character in any corrected text to obtain the character features of each character; for any character feature in any corrected text, the probability difference between the probability of recognizing the character as one of the various characters based on the character features of the character and the reference probability of the character being one of the various characters is obtained; the feature extraction parameters in the language model are optimized in the direction of reducing the probability difference to obtain an optimized language model; based on the optimized language model and the entity recognition network, an entity recognition model for the target scene is constructed, and the entity recognition model is used to perform entity recognition on the text in the target scene, etc.
[0160] This application embodiment also provides a computer storage medium (memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer storage medium provides storage space that stores the operating system of the computer device. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device; optionally, it can also be at least one computer storage medium located remotely from the aforementioned processor. In one embodiment, the processor can load and execute one or more instructions stored in the computer storage medium to implement the above-mentioned... Figure 2 or Figure 4 The various method steps in the embodiment of the model building method shown.
[0161] This application embodiment can obtain an error correction dataset for optimizing a language model. This dataset includes multiple error-corrected texts from a target scene, and each text includes at least one erroneous entity. Then, the language model can be invoked to extract features from each character in each error-corrected text, obtaining character features for each character. Based on these character features, the language model is optimized to obtain an optimized language model, improving its performance and resulting in a more accurate entity recognition model. This improves the accuracy of character features extracted by the language model in the entity recognition model, thereby enhancing entity recognition accuracy. Furthermore, since character features can also reflect contextual relationships, the language model optimized using the error correction dataset can more accurately extract character features from texts containing erroneous entities, more precisely reflecting the contextual relationships involved. This enables the entity recognition model to identify both reference entities (i.e., correctly spelled entities) and erroneous entities in the text.
[0162] It should be noted that, according to one aspect of this application, a computer program product or computer program is also provided, which includes computer instructions stored in a computer storage medium. The processor of a computer device reads the computer instructions from the computer storage medium, executes the computer instructions, and causes the computer device to perform the aforementioned actions. Figure 2 or Figure 4The methods provided are among the various alternative approaches to the model construction method embodiments shown.
[0163] Furthermore, it should be understood that the above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application are still within the scope of this application.
Claims
1. A model construction method, characterized in that, include: Obtain an error correction dataset for optimizing a language model. The error correction dataset includes multiple error correction texts in the target scene and error correction information corresponding to the error entities in each error correction text. Each error correction text includes at least one error entity, and an error entity includes at least one character. For any one of the multiple error-corrected texts, the language model is invoked to extract features from each character in the error-corrected text, thereby obtaining the character features of each character. For any character feature in any error-corrected text, obtain the probability difference between the probability of recognizing the character as any of the various characters based on the character feature of the character and the reference probability of the character being any of the various characters; the reference probability of any character being any of the various characters refers to the actual probability of the character being any of the various characters, used to indicate the actual character represented by the character. Optimize the feature extraction parameters in the language model in the direction of reducing the probability difference to obtain an optimized language model; Based on the optimized language model and entity recognition network, an entity recognition model for the target scene is constructed. The entity recognition model is used to perform entity recognition on text in the target scene. The entity recognition model supports the identification of correct and incorrect entities in the text. The error correction information is used to correct incorrect entities.
2. The method according to claim 1, characterized in that, The method further includes: Obtain a training dataset for optimizing the initial language model. The training dataset includes multiple training texts in the target scene, and each training text includes at least one reference entity. A reference entity includes at least one training character. For any training text among the plurality of training texts, the initial language model is invoked to extract features from each training character in the training text, thereby obtaining the character features of each training character. For any training character in any training text, the probability of identifying the training character based on the character features of the training character as each training character is obtained, and the training probability difference between the probability that the training character is the reference probability of each training character. The feature extraction parameters in the initial language model are optimized in the direction of reducing the difference in training probabilities to obtain the language model.
3. The method according to claim 1 or 2, characterized in that, The method further includes: According to the preset mask probability, obtain the first mask position in any of the error-corrected texts; Based on the target masking method and the first masking position, a second masking position in any corrected text is determined; wherein, the target masking method is used to indicate the selection method of the masking position in any corrected text, and the target masking method includes at least one of the following: character masking method, whole word masking method, and entity masking method; The characters at the second mask position are masked using a mask character to update any of the error-correcting texts, wherein each character in the updated error-correcting text is used for feature extraction.
4. The method according to claim 3, characterized in that, The method for determining the position of the second mask in any of the error-corrected text includes: The first mask position is used as the second mask position in any of the error-corrected texts; Alternatively, determine the complete word formed by the characters at the first mask position, and use the character position of the complete word as the second mask position in any of the error-corrected texts; Alternatively, the target entity composed of the characters at the first mask position can be determined, and the character position where the target entity is located can be used as the second mask position in any of the error-correcting texts.
5. The method according to claim 1 or 2, characterized in that, The method further includes: Obtain the obfuscation set under the target scenario, the obfuscation set includes multiple error correction character pairs under the target scenario, and each error correction character pair includes at least one erroneous entity and the corresponding entity after the at least one erroneous entity has been corrected; The language model in the entity recognition model is invoked to extract features from each character in the text to be processed, thereby obtaining the character features of each character in the text to be processed. The entity recognition network in the entity recognition model is invoked to perform entity recognition on the text to be processed based on the character features of each character in the text to be processed, so as to obtain the recognized entities of the text to be processed. If a target erroneous entity matching the identified entity is found in the obfuscation set, the identified entity is replaced with the entity corresponding to the target erroneous entity to achieve error correction processing of the text to be processed.
6. The method according to claim 5, characterized in that, The method further includes: If the target erroneous entity is not found in the confusion set and the identified entity is not found in the target reference entity set, then at least one candidate replacement entity corresponding to the identified entity is selected from the dictionary. The target reference entity set includes multiple reference entities. The distance between each candidate replacement entity in the at least one candidate replacement entity and the identified entity is determined, and the identified entity is replaced with the candidate replacement entity with the closest distance, so as to realize the error correction processing of the text to be processed.
7. The method according to claim 5, characterized in that, The error correction dataset also includes error correction information corresponding to the erroneous entities in each error correction text. The target scenario refers to a medical scenario, and each error correction text includes at least one erroneous medical entity. Obtaining the obfuscation set under the target scenario includes: Traverse the multiple error correction texts, and based on the target erroneous medical entity in the currently traversed error correction text and the error correction information corresponding to the target erroneous medical entity, determine the target medical entity after the target erroneous medical entity has been corrected. Based on the target erroneous medical entity and the target medical entity, determine the erroneous entity and the corresponding corrected entity in the target error correction character pair, and save the target error correction character pair to the obfuscation set; After traversing the multiple error-correcting texts, the obfuscation set for the target scenario is obtained.
8. A model building apparatus, characterized in that, include: The acquisition unit is used to acquire an error correction dataset for optimizing the language model. The error correction dataset includes multiple error correction texts in the target scene and error correction information corresponding to the error entities in each error correction text. Each error correction text includes at least one error entity, and an error entity includes at least one character. The processing unit is used to call the language model for any one of the plurality of error-corrected texts, and extract the features of each character in the error-corrected text to obtain the character features of each character. The processing unit is further configured to, for any character feature in any error-corrected text, obtain the probability difference between the probability of recognizing any character as any of the various characters based on the character feature of any character and the reference probability of any character being any of the various characters; the reference probability of any character being any of the various characters refers to the actual probability of any character being any of the various characters, used to indicate the actual character represented by any character; The processing unit is further configured to optimize the feature extraction parameters in the language model in the direction of reducing the probability difference, so as to obtain an optimized language model; The processing unit is further configured to construct an entity recognition model for the target scene based on the optimized language model and entity recognition network. The entity recognition model is used to perform entity recognition on the text in the target scene. The entity recognition model supports the identification of correct and incorrect entities in the text. The error correction information is used to correct incorrect entities.
9. A computer device, characterized in that, It includes a processor and a memory, wherein the memory is used to store a computer program that, when executed by the processor, implements the method as described in any one of claims 1-7.
10. A computer storage medium, characterized in that, The computer storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-7.
11. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-7.