Corpus Processing Model Training Method, Device, Storage Medium and Electronic Device

By splicing labels for sample corpus and performing slicing processing, using feature extraction and entity recognition networks to share advanced semantic information, the cascade error problem of named entity recognition and corpus classification models is solved, which improves training effect and accuracy, and reduces deployment difficulty.

CN113010647BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110356549.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-01
Publication Date
2025-07-18
Estimated Expiration
2041-04-01

AI Technical Summary

Technical Problem

In the prior art, cascade training of named entity recognition and corpus classification models leads to amplification of cascade errors and the inability to share advanced semantic information, affecting the training effect.

Method used

After splicing labels for sample corpus, slicing processing is performed, feature extraction networks and entity recognition networks are used for feature extraction and entity recognition, high-level semantic information is shared, and network parameters are adjusted to optimize the model.

Benefits of technology

The joint training of named entity recognition and corpus classification is realized, which improves the model's named entity recognition accuracy and corpus classification accuracy, avoids cascade errors and rules stacking, and reduces deployment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113010647B_ABST
    Figure CN113010647B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method, apparatus, storage medium, and electronic device for training a corpus processing model. The method includes splicing labels for sample corpus to obtain a splicing result; performing segmentation processing on the splicing result to obtain an object sequence; extracting features from the object sequence through a feature extraction network to obtain a feature information sequence; performing entity recognition on the feature information sequence through an entity recognition network to obtain an entity recognition result sequence, where each entity recognition result in the entity recognition result sequence represents the probability distribution of the corresponding object belonging to a preset category, and the preset category includes a named entity category and a classification category; determining an annotation path; adjusting the parameters of the feature extraction network and the entity recognition network according to the entity recognition result sequence and the annotation path; and obtaining a corpus processing model according to the adjustment result. Embodiments of the present application can jointly perform named entity recognition training and classification training on the model and achieve good training effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of artificial intelligence, and in particular, to a method, device, storage medium, and electronic device for training a corpus processing model. Background Art

[0002] Named entity recognition and corpus classification are two basic tasks in the field of natural language processing. To complete these two tasks, in practical application scenarios, it is often necessary to train a named entity recognition model and a corpus classification model separately. The named entity recognition model and the corpus classification model can be trained in a cascaded form. This cascaded training will amplify the cascaded error and may to a certain extent rely on the stacking of named entity recognition rules and corpus classification rules. The named entity recognition model and the corpus classification model can also be trained separately on the premise of sharing an embedding layer, but this separate training cannot share the high-level semantic information related to named entity recognition and corpus classification recognition, which affects the training effect. Summary of the Invention

[0003] In order to avoid the stacking of rules in the process of performing named entity recognition tasks and corpus classification tasks, share the high-level semantic information related to named entity recognition and corpus classification, and avoid the drawback of amplified cascaded error caused by cascaded training, embodiments of the present application provide a method, device, storage medium, and electronic device for training a corpus processing model.

[0004] On the one hand, embodiments of the present application provide a method for training a corpus processing model, the method comprising:

[0005] Concatenating a label to a sample corpus to obtain a concatenation result;

[0006] Performing a splitting process on the concatenation result such that each corpus unit in the sample corpus corresponds to an object of a named entity to be recognized and each of the labels corresponds to an object to be classified, and obtaining an object sequence according to the splitting result;

[0007] Performing feature extraction on the object sequence through a feature extraction network to obtain a feature information sequence;

[0008] Performing entity recognition on the feature information sequence through an entity recognition network to obtain an entity recognition result sequence, wherein each entity recognition result in the entity recognition result sequence represents the probability distribution of the corresponding object belonging to a preset category, and the preset category includes a named entity category and a classification category;

[0009] Determining the annotation category corresponding to the object in the object sequence to obtain an annotation path;

[0010] Adjusting the parameters of the feature extraction network and the entity recognition network according to the entity recognition result sequence and the annotation path;

[0011] Based on the adjusted feature extraction network and entity recognition network, the corpus processing model is obtained.

[0012] On the other hand, an embodiment of the present application provides a device for training a corpus processing model, the device includes:

[0013] A splicing module, configured to splice labels for a sample corpus to obtain a splicing result;

[0014] A segmentation module, configured to perform segmentation processing on the splicing result, so that each corpus unit in the sample corpus corresponds to an object of a named entity to be recognized and each label corresponds to an object to be classified, and an object sequence is obtained according to the segmentation processing result;

[0015] A feature extraction module, configured to extract features from the object sequence through a feature extraction network to obtain a feature information sequence;

[0016] An entity recognition module, configured to perform entity recognition on the feature information sequence through an entity recognition network to obtain an entity recognition result sequence, and each entity recognition result in the entity recognition result sequence represents the probability distribution of the corresponding object belonging to a preset category, and the preset category includes a named entity category and a classification category;

[0017] A labeled path determination module, configured to determine the labeled category corresponding to the object in the object sequence to obtain a labeled path;

[0018] A training module, configured to adjust the parameters of the feature extraction network and the entity recognition network according to the entity recognition result sequence and the labeled path;

[0019] A corpus processing model determination module, configured to obtain the corpus processing model according to the adjusted feature extraction network and entity recognition network.

[0020] On the other hand, an embodiment of the present application provides a computer-readable storage medium, characterized in that at least one instruction or at least one program is stored in the computer-readable storage medium, and the at least one instruction or at least one program is loaded and executed by a processor to implement the above-mentioned method for training a corpus processing model.

[0021] On the other hand, an embodiment of the present application provides an electronic device, characterized in that it includes at least one processor and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the at least one processor implements the above-mentioned method for training a corpus processing model by executing the instructions stored in the memory.

[0022] The embodiments of the present application provide a method, apparatus, storage medium, and device for training a corpus processing model. The embodiments of the present application can perform joint training on a model for named entity recognition and classification. During the training process, the named entity recognition task and the classification task can share high-level semantic features and assist each other for joint optimization, so that the trained corpus processing model can not only perform named entity recognition and corpus classification on the corpus, but also have good named entity recognition accuracy and corpus classification accuracy. Compared with the cascade training in the related art, the cascade error and rule stacking are avoided, and compared with the separate training in the related art, the named entity recognition accuracy and corpus classification accuracy are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the related art, the following will briefly introduce the drawings required for the description of the embodiments or the related art. Obviously, the following drawings are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 is a schematic flowchart of a method for training a corpus processing model provided by an embodiment of the present application;

[0025] Figure 2 is a schematic structural diagram of a neural network provided by an embodiment of the present application;

[0026] Figure 3 is a schematic basic structural diagram of a BERT model provided by an embodiment of the present application;

[0027] Figure 4 is a schematic basic structural diagram of a conversion model in a BERT model provided by an embodiment of the present application;

[0028] Figure 5 is a schematic basic structural diagram of an LSTM layer provided by an embodiment of the present application;

[0029] Figure 6 is a schematic diagram of the named entity recognition result in the named entity recognition result sequence provided by an embodiment of the present application;

[0030] Figure 7 is a schematic diagram of the named entity recognition result after adding tags provided by an embodiment of the present application;

[0031] Figure 8 is a schematic flowchart of adjusting the parameters of the above-mentioned feature extraction network and the above-mentioned entity recognition network provided by an embodiment of the present application;

[0032] Figure 9 is a schematic flowchart of adjusting the neural network parameters provided by an embodiment of the present application;

[0033] Figure 10 It is a block diagram of a corpus processing model training device provided by an embodiment of the present application;

[0034] Figure 11 It is a schematic diagram of the hardware structure of a device for implementing the method provided by an embodiment of the present application. Detailed implementation manners

[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the embodiments of the present application.

[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the embodiments of the present application and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0037] In order to make the purpose, technical solution and advantages of the disclosure of the embodiments of the present application clearer and more understandable, the embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain the embodiments of the present application, and are not used to limit the embodiments of the present application.

[0038] Hereinafter, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise stated, the meaning of "a plurality" is two or more. To facilitate the understanding of the above technical solutions of the embodiments of the present application and the technical effects produced thereby, the embodiments of the present application first explain the relevant professional terms:

[0039] Artificial Intelligence (AI): It is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning and decision-making.

[0040] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0041] Named Entities Recognition (NER): It is a basic task in Natural Language Processing (NLP). The purpose of named entity recognition is to identify named entities such as person names, place names, and organization names in the corpus, such as identifying person names, place names, organization names, time, dates, etc. from sentences.

[0042] Bidirectional Encoder Representation from Transformers (BERT): It is a model for pre-training language representations. A general "language understanding" model is trained based on the text corpus. Based on the BERT model, it can assist in performing Natural Language Processing (NLP) tasks.

[0043] Long Short-Term Memory (LSTM): It is a recurrent neural network suitable for capturing the position information before and after the sequence and making predictions on the sequence.

[0044] Conditional Random Fields (CRF): It is a probabilistic graphical model, often used for annotating or analyzing sequential data, such as natural language texts or biological sequences. A conditional random field is a conditional probability distribution model P(Y|X), representing a Markov random field of another set of output random variables Y given a set of input random variables X. That is to say, the characteristic of CRF is that it assumes the output random variables form a Markov random field. A conditional random field can be regarded as a generalization of the maximum entropy Markov model in the annotation problem.

[0045] In order to reduce the accumulation of rules in the process of performing named entity recognition tasks and corpus classification tasks, share high-level semantic information related to named entity recognition and corpus classification, and avoid the drawback of cascade error amplification caused by cascade training, the embodiments of the present application provide a method for training a corpus processing model.

[0046] The method provided by the embodiments of the present application may be related to the field of cloud technology. For example, it may be related to the field of big data. The method provided by the embodiments of the present application can perform corpus mining based on big data and train a corpus processing model according to the mined corpus. Big data refers to a collection of data that cannot be captured, managed, and processed by conventional software tools within a certain time range. It is a massive, high-growth-rate, and diverse information asset that requires new processing models to have stronger decision-making power, insight discovery ability, and process optimization ability. With the advent of the cloud era, big data has attracted more and more attention. Big data requires special technologies to effectively process a large amount of data tolerated over time. Technologies applicable to big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems.

[0047] The method provided by the embodiments of the present application may also be related to blockchain. That is, the method provided by the embodiments of the present application can be implemented based on blockchain, or the data involved in the method provided by the embodiments of the present application can be stored based on blockchain, or the execution entity of the method provided in the embodiments of the present application can be located in blockchain. Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.

[0048] The underlying blockchain platform may include processing modules such as user management, basic services, smart contracts, and operation monitoring. Among them, the user management module is responsible for the identity information management of all blockchain participants, including maintaining the generation of public and private keys (account management), key management, and the maintenance of the correspondence between the real identity of the user and the blockchain address (permission management). And under authorized circumstances, it supervises and audits the transaction situations of certain real identities, and provides the rule configuration for risk control (risk control audit); the basic service module is deployed on all blockchain node devices, used to verify the validity of business requests, and after reaching a consensus on valid requests, record them on the storage. For a new business request, the basic service first performs interface adaptation parsing and authentication processing (interface adaptation), then encrypts the business information through a consensus algorithm (consensus management), and after encryption, transmits it to the shared ledger completely and consistently (network communication), and performs record storage; the smart contract module is responsible for the registration and issuance of contracts, as well as contract triggering and contract execution. Developers can define contract logic through a certain programming language, publish it to the blockchain (contract registration), and according to the logic of the contract terms, call keys or other events to trigger execution, complete the contract logic, and at the same time also provide functions for contract upgrade and cancellation; the operation monitoring module is mainly responsible for the deployment, configuration modification, contract setting, cloud adaptation during the product release process, and the visual output of the real-time state during product operation, such as: alarming, monitoring network conditions, monitoring the health status of node devices, etc.

[0049] The platform product service layer provides the basic capabilities and implementation frameworks of typical applications. Developers can build on these basic capabilities and overlay the characteristics of the business to complete the blockchain implementation of the business logic. The application service layer provides application services based on the blockchain solution for business participants to use.

[0050] The embodiments of the present application can be applied to data processing devices. The data processing device can be a terminal device. The terminal device can be, for example, a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The data processing device can also be a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Of course, the data processing device can be a terminal device and a server, that is, the two cooperate to execute. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here.

[0051] The following introduces a method for training a corpus processing model according to an embodiment of the present application. Figure 1The figure shows a schematic flowchart of a method for training a corpus processing model provided by an embodiment of the present application. The embodiment of the present application provides the method operation steps as described in the embodiment or flowchart, but based on routine or non-creative labor, it may include more or fewer operation steps. The step order listed in the embodiment is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual system or server product executes, it may be executed in the order shown in the embodiment or the drawing or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). The above method may include:

[0052] S101. Concatenate labels to the sample corpus to obtain a concatenation result.

[0053] In the embodiment of the present application, the sample corpus may be a sentence for training a corpus processing model. Exemplarily, it may be a text sentence. Taking a text sentence in Chinese as an example, it is composed of multiple characters. Exemplarily, the sample corpus "Changjiang Automobile enters the bankruptcy liquidation process" is a sample corpus including 12 characters. Taking a text sentence in English as an example, it may be composed of multiple words. Exemplarily, the sample corpus "david is a cute boy" is a sample corpus including five words. For the convenience of subsequent processing, in a feasible embodiment, the obtained text content may be segmented to obtain a sample corpus with a length less than a preset length threshold.

[0054] On the one hand, the corpus processing model trained in the embodiment of the present application can output corresponding named entity recognition results for each corpus unit of the corpus to be processed. The corpus unit in the embodiment of the present application is the smallest unit processed by the corpus model. For a corpus in Chinese, the corpus unit may be a character, and for a corpus in English, the corpus unit may be a word. On the other hand, it can also classify the corpus to be processed and output a classification result, and the classification result may be output with the above label as the carrier.

[0055] In order to train such a corpus processing model, first, labels need to be concatenated to the sample corpus to obtain a concatenation result. Based on this concatenation result, a preset neural network is trained, and the trained neural network is used as the above corpus processing model. Taking "Changjiang Automobile enters the bankruptcy liquidation process" as an example, by concatenating the label [RL], the concatenation result "Changjiang Automobile enters the bankruptcy liquidation process[RL]" is obtained. Inputting this concatenation result into the above neural network can obtain the named entity recognition result and classification result corresponding to each character, and the classification result corresponds to the label [RL].

[0056] In a feasible embodiment, a reasonable number of tags can be set according to actual needs, and the number of tags is not limited in the embodiments of the present application. Exemplarily, if binary classification of the corpus is required, for example, if it is only necessary to classify whether the corpus is a negative corpus, one tag can be set. If multi-classification of the corpus is required, multiple corresponding tags need to be set. For example, if it is necessary to output both the first-level category of the corpus (such as sports, news, finance, entertainment) and the second-level category of the corpus under the first-level category. Taking sports as an example, the second-level categories can be basketball, baseball, and diving, and two tags can be set.

[0057] S102. Perform a segmentation process on the above splicing result, so that each corpus unit in the above sample corpus corresponds to an object of a named entity to be recognized and each of the above tags corresponds to an object to be classified, and obtain an object sequence according to the segmentation result.

[0058] In the embodiments of the present application, each object in the object sequence is processed by the above neural network without distinction to obtain the entity recognition result corresponding to each object. For the object corresponding to the corpus unit, its final training goal is to output the named entity to which it belongs. For the object corresponding to the tag, its final training goal is to output a certain classification result of the corpus.

[0059] The embodiments of the present application can segment the above splicing result based on the corpus unit and the tag. By segmenting each corpus unit in the splicing result, an object of a named entity to be recognized corresponding to each of the above corpus units is obtained. By segmenting each tag in the splicing result, an object to be classified corresponding to each of the above tags is obtained, and an object sequence is obtained according to the object of the named entity to be recognized and the object to be classified. Exemplarily, taking the splicing result "Changjiang Automobile enters the bankruptcy liquidation process [RL]" described above as an example, it includes 12 corpus units and 1 tag. By segmenting this splicing result, 12 objects of named entities to be recognized and 1 object to be classified can be correspondingly obtained, and finally an object sequence including 13 objects is obtained. Named entity recognition and classification both belong to subordinate concepts of entity recognition. The neural network can perform entity recognition on the object sequence including 13 objects without distinction to obtain the named entity recognition result corresponding to the first 12 objects and the classification result corresponding to the 13th object.

[0060] S103. Extract features from the above object sequence through a feature extraction network to obtain a feature information sequence.

[0061] In the embodiments of the present application, the feature extraction network can be formed by BERT, RoBERTa (A Robustly Optimized BERT Pretraining Approach), ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately). The embodiments of the present application do not limit the specific structure of the feature extraction network.

[0062] Please refer to Figure 2 , which shows a schematic structural diagram of a neural network. By training this neural network, a corpus processing model can be obtained. Figure 2 In [reference], the BERT model is used as the feature extraction network, where Trm represents the transformation model (Transformer) on which the BERT model depends. Ei represents each object, and Ti represents the corresponding feature extraction result, where i is an integer less than or equal to N, and N is the length of the object sequence.

[0063] Please refer to Figure 3 , which shows a schematic diagram of the basic structure of the BERT model. In combination with Figure 3 A brief introduction to the BERT model is given. The BERT model can perform word segmentation on the input corpus. That is to say, the BERT model can process the word segmentation results as units to obtain the corresponding feature extraction results for each word segmentation result. During the feature extraction process, the BERT model can add additional markers to the word segmentation results. Exemplarily, [CLS] represents the classification marker, and [SEP] represents the short sentence marker. Based on the obtained word segmentation results, word feature extraction (Token Embeddings), short sentence feature extraction (Segment Embeddings), and position feature extraction (Position Embeddings) can be performed. Taking "mydogiscute helikes playing" as an example of the input corpus, it is segmented word by word to obtain multiple "words", and Token Embeddings, Segment Embeddings, and Position Embeddings are performed on each "word", and finally the feature extraction results are obtained. In the embodiments of the present application, the object sequence is obtained by segmentation, and the BERT model can perform feature extraction on the objects in the object sequence to obtain the corresponding feature extraction results.

[0064] The core structure of the BERT model is the transformation model (Transformer). Please refer to Figure 4, which shows a schematic diagram of the basic structure of the conversion model in the BERT model. Transformer is a new architecture proposed in May 2018, which can replace traditional recurrent neural networks and convolutional neural networks for machine learning. The structure of Transformer is divided into the left encoder and the right decoder. It not only adds multi-head attention, but also adds self-attention and add & norm inside. Finally, it passes through linearization and the activation layer. The activation layer uses Softmax as the activation function. Transformer learns different features from different dimensions and adds position information through positional encoding. The conversion model can extract high-order semantic features of the above input corpus.

[0065] In one embodiment, the above feature extraction network extracts features from the above object sequence to obtain a feature information sequence, including: extracting lexical features from each object in the above object sequence to obtain a lexical feature sequence; extracting syntactic features from the above lexical feature sequence to obtain a target feature sequence; performing bidirectional semantic feature extraction on the above target feature sequence to obtain the above feature information sequence. Specifically, lexical feature extraction, syntactic feature extraction, and bidirectional semantic feature extraction can all be implemented based on the BERT model, so as to obtain better feature extraction results, which will not be elaborated in the embodiments of this application.

[0066] S104. Use the entity recognition network to perform entity recognition on the above feature information sequence to obtain an entity recognition result sequence. Each entity recognition result in the above entity recognition result sequence represents the probability distribution of the corresponding object belonging to a preset category, and the above preset category includes named entity categories and classification categories.

[0067] Please refer to Figure 2, an entity recognition network can be formed by an LSTM layer and a CRF layer. The LSTM is used to further supplement the front and back sequence position information on the basis of extracting feature information sequences from the above object sequences, so as to optimize the entity recognition result. That is, the LSTM extracts the sequence position information from the above feature information sequences, and the CRF performs conditional random field analysis on the above feature information sequences according to the extraction results to obtain the above entity recognition result sequences. In other feasible embodiments, the conditional random field analysis can also be directly performed on the above feature information sequences based on the CRF to obtain the above entity recognition result sequences. The LSTM layer can already output the probability distribution of the preliminary preset categories of each object. By adding transfer constraints through the CRF layer, the entity recognition result sequences output by the entity recognition network can be finally obtained. The CRF layer can learn the transfer constraints from the training corpus by itself to ensure the effectiveness of the finally predicted sequences, and this part of the content will not be elaborated here.

[0068] Please refer to Figure 5 , which shows a schematic diagram of the basic structure of the LSTM layer. The squares represent the neuron network layers, and the circles represent a certain fusion operation. The two squares with different names indicate that different activation functions are used. Figure 5 The basic structures in [reference] form the forget gate, input gate and output gate of the LSTM, and this will not be elaborated in the embodiments of the present application.

[0069] Please refer to Figure 6 , which shows a schematic diagram of the entity recognition result in the entity recognition result sequence. Combining Figure 2It can be known that for each object (Xi) in the object sequence, an entity recognition result (Ci) can be correspondingly output. Each entity recognition result represents the probability distribution of the corresponding object belonging to a preset category. The above preset categories include named entity categories and classification categories. Taking C1 as an example, it shows the probabilities that the object belongs to B-com, I-com... TRUE, where B-com and I-com belong to the named entity categories. Specifically, B-com represents the first character of the company name, and I-com represents the non-first character of the company name. "TRUE" belongs to the classification category, that is, the probability that the sample corpus belongs to a certain type. Exemplarily, in a binary classification scenario, the probability of "TRUE" can represent the probability that the sample corpus belongs to a negative corpus. Obviously, for each object, its corresponding entity recognition result includes the probabilities that the object belongs to each named entity category and the probabilities that the sample corpus corresponding to the object belongs to each classification category. Still taking C1 as an example, the corresponding entity recognition result represents that the probability that the object is the first character of the company name is 1.5, the probability that the object is the middle character of the company name is 0.9, and the probability that the sample corpus corresponding to the object is a negative corpus is 0.05. The corpus processing model trained in this scenario can be applied to the field of risk warning. The determination result of the negative corpus can prompt a certain risk, such as financial crime risk or investment risk.

[0070] S105. Determine the annotation category corresponding to the object in the above object sequence to obtain an annotation path.

[0071] Each object in the splicing result obtained from the sample corpus has its corresponding annotation category. The annotation category is used as a true value for training the neural network. Taking "Changjiang Automobile enters the bankruptcy liquidation process [RL]" as an example, the corresponding object sequence obtained is "Chang", "Jiang", "Qi", "Che", "Jin", "Ru", "Po", "Chan", "Qing", "Suan", "Cheng", "Xu", "RL", where each object has its annotation category as the true value. For example, "Chang" is the first character of the company name, and its annotation category is B-com. "Jiang" is the middle character of the company, and its annotation category is I-com. This sample corpus is a negative corpus, and the annotation category corresponding to "RL" is TRUE. The path formed by the annotation categories corresponding to a preset number of objects can be an annotation path.

[0072] S106. Adjust the parameters of the above feature extraction network and the above entity recognition network according to the above entity recognition result sequence and the above annotation path.

[0073] In one embodiment, the entity recognition result sequence can be used as a whole to determine the training target, and the parameters of the above feature extraction network and the above entity recognition network are adjusted according to the training target.

[0074] Specifically, the annotation category corresponding to each object in the above object sequence can be determined to obtain the above annotation path. Among them, the annotation category of the object corresponding to the corpus unit is the named entity category, and the annotation category of the object corresponding to the label is the classification category. Exemplarily, taking the object sequence "Chang", "Jiang", "Automobile", "Enter", "Bankruptcy", "Liquidation", "Procedure", "RL" as an example, this object sequence includes 13 objects, and an annotation path including 13 annotation categories is formed.

[0075] Correspondingly, the parameters of the above feature extraction network and the above entity recognition network can be adjusted with the highest ratio of the probability of the above annotation path to the total probability of the first path as the training objective. Among them, the total probability of the first path represents the sum of the probabilities of all possible paths obtained based on the above entity recognition result sequence. The probability of the above annotation path can be calculated according to the entity recognition result sequence.

[0076] Exemplarily, if there are 8 named entity categories and 1 classification category in the preset categories, then there are a total of 9 preset categories. Still taking the object sequence "Chang", "Jiang", "Automobile", "Enter", "Bankruptcy", "Liquidation", "Procedure", "RL" as an example, this object sequence includes 13 objects, and the entity recognition result of each object represents the probability distribution of the above 9 preset categories. That is to say, each object corresponds to 9 predicted probabilities, so there are 9 to the 13th power of possible paths formed by the entity recognition results of these 13 objects. Based on the above entity recognition results, the probability corresponding to each possible path and the probability corresponding to the annotation path are both computable, and this calculation process will not be elaborated here.

[0077] For tasks of sequence prediction, the highest probability of the annotation path can be used as the training objective. Exemplarily, the parameters of the above feature extraction network and the above entity recognition network can be adjusted with as the training objective, where P RealPath represents the annotation path, and ∑p i represents the sum of the probabilities of all possible paths obtained based on the above entity recognition result sequence.

[0078] In the embodiments of the present application, the label corresponds to the classification category, and each corpus unit corresponds to the named entity category. In some possible implementation scenarios, it may be that due to the number of labels being less than the number of corpus units in the sample corpus, the neural network pays too much attention to named entity recognition during the training process and ignores the classification of the sample corpus. Exemplarily, Table 1 shows the schematic data of the model training effect in the scenario where the number of labels is less than the number of corpus units. Among them, the F1 value is a parameter in the field of model effect evaluation, and its meaning is F1 value = accuracy * recall * 2 / (accuracy + recall).

[0079] Table 1

[0080] Preset category Accuracy Recall F1 score First classification category 54.30% 86.90% 66.84 Second classification category 57.72% 19.65% 29.32 First named entity category 98.93% 99.11% 99.02 Second named entity category 98.29% 96.02% 97.14 Third named entity category 97.45% 92.85% 95.09 Fourth named entity category 98.75% 97.48% 98.11 Fifth named entity category 100.00% 98.85% 99.42 Sixth named entity category 99.07% 92.41% 95.62 Seventh named entity category 100.00% 91.34% 95.48

[0081] It can be clearly seen from the results in Table 1 that the model has a poor effect on classification but a good effect on named entity recognition. This is because when performing entity recognition on the object sequence, the proportion of objects corresponding to the labels is too small. Correspondingly, it can be improved by increasing the form of labels.

[0082] In one embodiment, the number of labels can be adjusted to be equal to the number of corpus units in the sample corpus. Then, when the model performs entity recognition, it can equally focus on the two tasks of named entity recognition and classification, so that both named entity recognition and classification can achieve good results. Specifically, in the step of splicing labels for the sample corpus to obtain the splicing result mentioned above, the number of corpus units in the above sample corpus can be obtained; splice the above number of corpus units of labels for the above sample corpus to obtain the splicing result. Training based on this splicing result can obtain good results. Still taking the sample corpus "Changjiang Automobile enters the bankruptcy liquidation process" as an example, it includes 12 corpus units, so 12 labels are added to it to form the following splicing result "Changjiang Automobile enters the bankruptcy liquidation process

[0083] [RL][RL][RL][RL][RL][RL][RL][RL][RL][RL][RL][RL]". Please refer to Figure 7 , which shows a schematic diagram of the entity recognition result after adding labels. Obviously, the 12 corpus units in "Changjiang Automobile enters the bankruptcy liquidation process

[0084] [RL][RL][RL][RL][RL][RL][RL][RL][RL][RL][RL][RL]" output 12 named entity recognition results, and the 12 labels can correspondingly output 12 classification results.

[0085] In one embodiment, the annotation category corresponding to the object in the object sequence corresponding to the corpus unit can be determined to obtain a first annotation path; the annotation category corresponding to the object in the object sequence corresponding to the label can be determined to obtain a second annotation path. Both the first annotation path and the second annotation path are the acquisition results of the standard path in step S105, that is, the standard path in step S105 can include the first standard path and the second standard path. Taking the object sequence "Chang", "Jiang", "Automobile", "Enter", "Bankruptcy", "Liquidation", "Procedure", "RL" as an example, the annotation categories corresponding to the first 12 objects form the first annotation path, and the annotation category of the last object forms the second annotation path. Taking the object sequence "Chang", "Jiang", "Automobile", "Enter", "Bankruptcy", "Liquidation", "Procedure", "RL", "RL", "RL", "RL", "RL", "RL", "RL", "RL", "RL", "RL", "RL", "RL", "RL", "RL", "RL" as an example, the annotation categories corresponding to the first 12 objects form the first annotation path, and the annotation categories corresponding to the last 12 objects form the second annotation path.

[0086] Correspondingly, according to the entity recognition result sequence and the annotation path, the parameters of the feature extraction network and the entity recognition network are adjusted, as Figure 8 shown, including:

[0087] S1061. Extract the entity recognition result corresponding to the object corresponding to the corpus unit in the entity recognition result sequence to obtain a named entity prediction sequence.

[0088] Taking the object sequence "Chang", "Jiang", "Automobile", "Enter", "Bankruptcy", "Liquidation", "Procedure", "RL" as an example, the first 12 objects all correspond to the corpus unit, so the entity recognition results of the first 12 objects form a named entity prediction sequence.

[0089] S1063. Determine the named entity recognition loss according to the ratio of the probability of the first annotation path to the total probability of the second path, where the total probability of the second path represents the sum of the probabilities of all possible paths obtained based on the named entity prediction sequence.

[0090] Referring to the above, taking the object sequence "Chang", "Jiang", "Automobile", "Enter", "Bankruptcy", "Liquidation", "Procedure", "RL" as an example, it can be known that the named entity prediction sequence includes the entity recognition results corresponding to 12 objects. If there are 8 named entity categories and 1 classification category in the preset categories, then there are a total of 9 preset categories, and the possible paths formed by the entity recognition results of these 12 objects are 9 to the 12th power. Based on the above-named entity prediction sequence, the probability corresponding to each possible path and the probability corresponding to the labeled path are both calculable, and this calculation process will not be elaborated here.

[0091] S1065. For each labeled category in the above second labeled path, based on the entity recognition result of the object corresponding to the above labeled category and the above labeled category, determine the classification loss of the object corresponding to the above labeled category.

[0092] Taking the object sequence "Chang", "Jiang", "Automobile", "Enter", "Bankruptcy", "Liquidation", "Procedure", "RL" as an example, the second labeled path only includes the labeled category corresponding to the last object "RL". The classification loss corresponding to this object can be determined according to the difference between this labeled category and the entity recognition result of "RL". By the same token, taking the object sequence "Chang", "Jiang", "Automobile", "Enter", "Bankruptcy", "Liquidation", "Procedure", "RL", "RL", "RL", "RL", "RL", "RL", "RL", "RL", "RL", "RL", "RL", "RL", "RL" as an example, the second labeled path includes the labeled categories corresponding to 12 "RL" objects. For each "RL" object, the classification loss corresponding to this object can be determined according to the difference between the labeled category and the entity recognition result corresponding to this "RL".

[0093] S1067. According to the above-named entity recognition loss and each of the above classification losses, adjust the parameters of the above feature extraction network and the above entity recognition network.

[0094] Specifically, the weighted values of the above-named entity recognition loss and each of the above classification losses can be determined as the total loss value, and the parameters of the above feature extraction network and the above entity recognition network can be adjusted according to the above total loss value. In the embodiments of the present disclosure, the weights can be set according to the actual situation, and the setting method and specific values of the weights are not limited. Exemplarily, if the proportion of the object corresponding to the label in the object sequence is relatively low, the weight of the classification loss can be increased accordingly.

[0095] Specifically, please refer to Figure 9 , which shows a schematic diagram of the neural network parameter adjustment process. The above adjustment of the parameters of the above feature extraction network and the above entity recognition network according to the above-named entity recognition loss and each of the above classification losses includes:

[0096] S1. Determine a first quantity of objects corresponding to the corpus units.

[0097] S2. Determine a second quantity of objects corresponding to the labels.

[0098] S3. Determine a first weight corresponding to the named entity recognition loss and a second weight corresponding to the classification loss according to the first quantity and the second quantity.

[0099] By setting the weights, it is possible to pay attention to named entity recognition and corpus classification more evenly during model training, without overemphasizing one over the other, so that the trained model can perform well in both named entity recognition and corpus classification. Exemplarily, if the first quantity is higher than the second quantity, the weight corresponding to the classification loss can be appropriately increased.

[0100] S4. Determine the total loss according to the named entity recognition loss, the first weight, the classification loss, and the second weight.

[0101] S5. Adjust the parameters of the feature extraction network and the entity recognition network according to the total loss.

[0102] S107. Obtain the corpus processing model according to the adjusted feature extraction network and entity recognition network.

[0103] In the embodiments of the present application, the trained feature extraction network and entity recognition network can be used as a corpus processing model. The corpus processing model can process the corpus, thereby outputting the entity recognition results in the corpus and classifying the corpus, and can be widely applied to various application scenarios that require corpus processing. Named entity recognition and corpus classification are two basic tasks in the field of corpus processing. The corpus processing model trained in the embodiments of the present application can output both the named entity recognition results and the corpus classification results, which undoubtedly improves the corpus processing ability of the model and reduces the application difficulty. In scenarios that require corpus processing, only the corpus processing model in the embodiments of the present application needs to be deployed to complete both named entity recognition and corpus classification. Taking a certain news message as an example, inputting the news message into the corpus processing model can not only obtain the named entities included in the news message, but also obtain the classification result of the news message, which is convenient for further processing of the news message. The application scenarios of the corpus processing model in the embodiments of the present application are not limited. Exemplarily, it can be applied in recommendation scenarios, human-computer interaction scenarios, risk warning scenarios, and big data analysis scenarios.

[0104] The corpus processing model training method provided by the embodiments of the present application can simultaneously perform named entity recognition training and classification training on the model. During the training process, the named entity recognition task and the classification task can share high-level semantic features and assist each other to jointly optimize, so that the trained corpus processing model can not only perform named entity recognition and corpus classification on the corpus at the same time, but also have better named entity recognition accuracy and corpus classification accuracy. Compared with the cascaded training method in the related art, it avoids cascaded errors and rule stacking, and compared with the separate training in the related art, it improves the named entity recognition accuracy and corpus classification accuracy. According to the above description, it can be known that the corpus processing model at least includes a trained feature extraction network and an entity recognition network, the model structure is relatively simple, and the deployment difficulty is correspondingly low. Compared with the need to deploy a named entity recognition model and a corpus classification model in the related art, the deployment cost of the embodiments of the present application is low and it is easy to promote and apply.

[0105] The embodiments of the present application also disclose a corpus processing model training device, as Figure 10 shown, the above device includes:

[0106] The splicing module 10 is used to splice labels for the sample corpus to obtain a splicing result.

[0107] The splitting module 20 is used to perform splitting processing on the above splicing result, so that each corpus unit in the above sample corpus corresponds to an object of a named entity to be recognized and each of the above labels corresponds to an object to be classified, and an object sequence is obtained according to the splitting processing result.

[0108] The feature extraction module 30 is used to extract features from the above object sequence through a feature extraction network to obtain a feature information sequence.

[0109] The entity recognition module 40 is used to perform entity recognition on the above feature information sequence through an entity recognition network to obtain an entity recognition result sequence. Each entity recognition result in the above entity recognition result sequence represents the probability distribution that the corresponding object belongs to a preset category, and the above preset category includes a named entity category and a classification category.

[0110] The annotation path determination module 50 is used to determine the annotation category corresponding to the object in the above object sequence to obtain an annotation path.

[0111] The training module 60 is used to adjust the parameters of the above feature extraction network and the above entity recognition network according to the above entity recognition result sequence and the above annotation path.

[0112] The corpus processing model determination module 70 is used to obtain the above corpus processing model according to the adjusted above feature extraction network and the above entity recognition network.

[0113] Specifically, the training device for the corpus processing model disclosed in the embodiments of the present application and the corresponding method embodiments are all based on the same inventive concept. For details, please refer to the method embodiments and will not be elaborated herein.

[0114] The embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned method for training a corpus processing model.

[0115] The embodiments of the present application also provide a computer-readable storage medium. The above computer-readable storage medium can store multiple instructions. The above instructions can be adapted to be loaded and executed by the processor to execute the above-mentioned method for training a corpus processing model in the embodiments of the present application.

[0116] Furthermore, Figure 11 The hardware structure diagram of a device for implementing the method provided in the embodiments of the present application is shown. The above device can participate in forming or include the device or system provided in the embodiments of the present application. As Figure 11 shown, the device 10 may include one or more (shown as 102a, 102b,..., 102n in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 11 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the device 10 may further include more or fewer components than those Figure 11 shown, or have a different configuration from that Figure 11 shown.

[0117] It should be noted that the above one or more processors 102 and / or other data processing circuits are generally referred to as "data processing circuits" herein. The data processing circuit may be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be wholly or partially incorporated into any one of the other elements in the device 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0118] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the methods in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned method for training a corpus processing model. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the device 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0119] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the device 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0120] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the device 10 (or a mobile device).

[0121] It should be noted that: the above-mentioned sequence of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above-mentioned specific embodiments of the present application have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0122] The various embodiments in the embodiments of the present application are all described in a progressive manner. For the same or similar parts between the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and server embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

[0123] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The above program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a disk, an optical disc, etc.

[0124] The above are only the preferred embodiments of the embodiments of the present application, and are not intended to limit the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the embodiments of the present application shall be included in the protection scope of the embodiments of the present application.

Claims

1. A method for training a corpus processing model, characterized in that, The method includes: Appending labels to the sample corpus to obtain an appended result; Performing a segmentation process on the appended result such that each corpus unit in the sample corpus corresponds to an object of a named entity to be recognized and each of the labels corresponds to an object to be classified, and obtaining an object sequence according to the segmentation result; Performing feature extraction on the object sequence through a feature extraction network to obtain a feature information sequence; Performing entity recognition on the feature information sequence through an entity recognition network to obtain an entity recognition result sequence, where each entity recognition result in the entity recognition result sequence represents the probability distribution of the corresponding object belonging to a preset category, and the preset category includes a named entity category and a classification category; Determining the annotation category corresponding to each object in the object sequence to obtain the annotation path, where the annotation category of the object corresponding to the corpus unit is the named entity category, and the annotation category of the object corresponding to the label is the classification category; Taking the ratio of the probability of the annotation path to the total probability of the first path being the highest as the training objective, and adjusting the parameters of the feature extraction network and the entity recognition network, where the total probability of the first path represents the sum of the probabilities of all possible paths obtained based on the entity recognition result sequence; Obtaining the corpus processing model according to the adjusted feature extraction network and entity recognition network.

2. The method according to claim 1, wherein The annotation path includes a first annotation path and a second annotation path. The determining the annotation category corresponding to each object in the object sequence to obtain the annotation path includes: Determining the annotation category corresponding to the object corresponding to the corpus unit in the object sequence to obtain the first annotation path; Determining the annotation category corresponding to the object corresponding to the label in the object sequence to obtain the second annotation path; The taking the ratio of the probability of the annotation path to the total probability of the first path being the highest as the training objective and adjusting the parameters of the feature extraction network and the entity recognition network includes: Extracting the entity recognition results corresponding to the objects corresponding to the corpus unit in the entity recognition result sequence to obtain a named entity prediction sequence; Determining the named entity recognition loss according to the ratio of the probability of the first annotation path to the total probability of the second path, where the total probability of the second path represents the sum of the probabilities of all possible paths obtained based on the named entity prediction sequence; For each annotation category in the second annotation path, determining the classification loss of the object corresponding to the annotation category based on the entity recognition result of the object corresponding to the annotation category and the annotation category; Adjusting the parameters of the feature extraction network and the entity recognition network according to the named entity recognition loss and each of the classification losses.

3. The method according to claim 2, wherein The adjusting the parameters of the feature extraction network and the entity recognition network according to the named entity recognition loss and each of the classification losses includes: Determining a first quantity corresponding to the object corresponding to the corpus unit; Determining a second quantity corresponding to the object corresponding to the label; Determining a first weight corresponding to the named entity recognition loss and a second weight corresponding to the classification loss according to the first quantity and the second quantity; Determine the total loss according to the named entity recognition loss, the first weight, the classification loss, and the second weight; Adjust the parameters of the feature extraction network and the entity recognition network according to the total loss.

4. The method according to any one of claims 1 to 3, characterized in that, Splicing labels for the sample corpus to obtain a splicing result, including: Obtain the number of corpus units in the sample corpus; Splice the number of labels equal to the number of corpus units for the sample corpus to obtain the splicing result.

5. The method according to any one of claims 1, characterized in that, Performing feature extraction on the object sequence through a feature extraction network to obtain a feature information sequence, including: Performing lexical feature extraction on each object in the object sequence to obtain a lexical feature sequence; Performing syntactic feature extraction on the lexical feature sequence to obtain a target feature sequence; Performing bidirectional semantic feature extraction on the target feature sequence to obtain the feature information sequence.

6. The method according to any one of claims 1, characterized in that Performing entity recognition on the feature information sequence through an entity recognition network to obtain an entity recognition result sequence, including: Performing conditional random field analysis on the feature information sequence in the entity recognition network to obtain the entity recognition result sequence; Or, Performing sequence position information extraction on the feature information sequence in the entity recognition network; and performing conditional random field analysis on the feature information sequence according to the extraction result to obtain the entity recognition result sequence.

7. A corpus processing model training device, characterized in that, The device includes: A splicing module for splicing labels for the sample corpus to obtain a splicing result; A splitting module for splitting the splicing result so that each corpus unit in the sample corpus corresponds to an object of a named entity to be recognized and each label corresponds to an object to be classified, and obtaining an object sequence according to the splitting result; A feature extraction module for performing feature extraction on the object sequence through a feature extraction network to obtain a feature information sequence; An entity recognition module for performing entity recognition on the feature information sequence through an entity recognition network to obtain an entity recognition result sequence, where each entity recognition result in the entity recognition result sequence represents the probability distribution that the corresponding object belongs to a preset category, and the preset category includes a named entity category and a classification category; A labeled path determination module for determining the labeled category corresponding to each object in the object sequence to obtain the labeled path, where the labeled category of the object corresponding to the corpus unit is the named entity category, and the labeled category of the object corresponding to the label is the classification category; A training module for adjusting the parameters of the feature extraction network and the entity recognition network with the highest ratio of the probability of the labeled path to the total probability of the first path as the training target, where the total probability of the first path represents the sum of the probabilities of all possible paths obtained based on the entity recognition result sequence; A corpus processing model determination module for obtaining the corpus processing model according to the adjusted feature extraction network and entity recognition network.

8. A computer-readable storage medium, characterized in that, At least one instruction or at least one program is stored in the computer-readable storage medium, and the at least one instruction or at least one program is loaded and executed by a processor to implement a corpus processing model training method according to any one of claims 1 to 6.

9. An electronic device, characterized in that, Comprising at least one processor and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the at least one processor implements a corpus processing model training method according to any one of claims 1 to 6 by executing the instructions stored in the memory.

Citation Information

Patent Citations

  • Named entity recognition model training method and named entity recognition method and device

    CN110705294A