Model training method, recognition method, device, electronic device and computer storage medium

Through the shared combination structure of encoder and decoder, combined with random domain label-free training data for deep feature extraction, the problem of poor migration training effect in the source target field is solved, and efficient recognition of the named entity recognition model in the target field is achieved.

CN113553849BActive Publication Date: 2025-07-11ALIBABA GROUP HOLDING LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010340132.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-26
Publication Date
2025-07-11
Estimated Expiration
2040-04-26

AI Technical Summary

Technical Problem

现有技术在源领域和目标领域相距较远时,迁移训练方法无法有效实现知识迁移,反而可能损害目标领域的命名实体识别效果。

Method used

The combined structure of common encoder, source domain decoder, target domain decoder and random domain decoder is adopted, and the unmarked data of the input source domain and random domain are screened, combined with the target domain data for training, and the domain migration training sample includes random domain label-free training data for unsupervised training, forming an autoencoder structure for deep feature extraction.

Benefits of technology

The effectiveness of domain migration training of named entity recognition models is improved, ensuring effective recognition effect in the target domain, while reducing manual labeling workload and improving training accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113553849B_ABST
    Figure CN113553849B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a model training method, an identification method, a device, an electronic device, and a computer storage medium. The model training method is applied to a named entity recognition model, and includes: inputting training data in a source domain and unlabeled data in a random domain into a training model to obtain a first training result, where the training model includes a shared encoder, a source domain decoder, a target domain decoder, and a random domain decoder; performing data screening to obtain the screened training data in the source domain and the screened unlabeled data in the random domain; inputting the above-screened data and the training data in the target domain into the training model to obtain a second training result. Since the domain transfer training samples include unlabeled training data in the random domain, in domain transfer training, a wider range of training samples other than the source domain training samples and the target domain training samples are applied, thereby improving the effectiveness of domain transfer training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of communication technologies, and in particular, to a model training method, an identification method, a device, an electronic device, and a computer storage medium. Background Art

[0002] Named Entity Recognition (NER), also known as "Proper Name Recognition", refers to automatically identifying entities with specific meanings in text through automated methods, mainly including person names, place names, organization names, product names, proper nouns, etc. Named Entity Recognition plays an important role in the field of text processing, such as question answering systems, translation, etc.

[0003] When constructing various knowledge bases or various top-level services, a named entity recognition model is usually trained using training data. However, when the source domain where the existing training data is located is far from the target domain, the existing transfer training methods cannot achieve knowledge transfer, and sometimes it will even damage the effect of the original target domain. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a control method, a device, a mobile terminal, an electronic device, and a computer storage medium, which can effectively improve the effect of transfer training.

[0005] According to a first aspect of the embodiments of the present invention, a model training method is provided. The model training method is applied to a named entity recognition model and includes: inputting training data in the source domain and unlabeled data in a random domain into a training model to obtain a first training result. The training model includes a shared encoder, a source domain decoder, a target domain decoder, and a random domain decoder; screening the training data in the source domain and the unlabeled data in the random domain according to the training result to obtain the screened training data in the source domain and the screened unlabeled data in the random domain; inputting the screened training data in the source domain, the screened unlabeled data in the random domain, and training data in the target domain into the training model to obtain a second training result.

[0006] According to a second aspect of the embodiments of the present invention, an identification method is provided, including: inputting data to be predicted in the target domain into a training model, and the training model is trained by the method described in the first aspect; the training model encodes the data to be predicted by using the shared encoder; the training model decodes the encoded data to be predicted by using the target domain decoder to obtain a prediction result.

[0007] According to a third aspect of an embodiment of the present invention, there is provided a model training method applied to a named entity recognition model, including: determining an encoder-decoder structure, where the encoder-decoder structure includes a shared encoder and a source domain decoder, a target domain decoder, and a random domain decoder connected to the shared encoder. Among them, the shared encoder and the random domain decoder form an autoencoder structure, and the shared encoder and the target domain decoder form a target domain encoder-decoder structure; based on initial training samples, initially train at least part of the encoder-decoder structure, where the initial training samples include source domain labeled training data; based on domain transfer training samples, perform domain transfer training on at least part of the encoder-decoder structure, where the domain transfer training samples include target domain labeled training data and random domain unlabeled training data, and the random domain unlabeled training data is used for unsupervised training of the autoencoder structure; in at least part of the encoder-decoder structure that has undergone domain transfer training, determine the target domain encoder-decoder structure as the named entity recognition model.

[0008] According to a fourth aspect of an embodiment of the present invention, there is provided a model training device applied to a named entity recognition model, including: a first input module that inputs training data in the source domain and unannotated data in the random domain into a training model to obtain a first training result, where the training model includes a shared encoder, a source domain decoder, a target domain decoder, and a random domain decoder; a screening module that screens the training data in the source domain and the unannotated data in the random domain according to the training result to obtain the screened training data in the source domain and the screened unannotated data in the random domain; a second input module that inputs the screened training data in the source domain, the screened unannotated data in the random domain, and training data in the target domain into the training model to obtain a second training result.

[0009] According to a fifth aspect of an embodiment of the present invention, there is provided an identification device, including: an input module that inputs data to be predicted in the target domain into a training model, where the training model is trained by the above model training method; an encoding module that the training model encodes the data to be predicted using the shared encoder; a decoding module that the training model decodes the encoded data to be predicted using the target domain decoder to obtain a prediction result; an output module that outputs the prediction result.

[0010] According to a sixth aspect of an embodiment of the present invention, there is provided a model training apparatus applied to a named entity recognition model, including: an encoder-decoder structure determination module that determines an encoder-decoder structure, where the encoder-decoder structure includes a shared encoder and a source domain decoder, a target domain decoder, and a random domain decoder connected to the shared encoder. Among them, the shared encoder and the random domain decoder form an autoencoder structure, and the shared encoder and the target domain decoder form a target domain encoder-decoder structure; an initial training module that performs initial training on at least part of the encoder-decoder structure based on initial training samples, where the initial training samples include source domain labeled training data; a domain transfer training module that performs domain transfer training on at least part of the encoder-decoder structure based on domain transfer training samples, where the domain transfer training samples include target domain labeled training data and random domain unlabeled training data, and the random domain unlabeled training data is used for unsupervised training of the autoencoder structure; a model determination module that determines the target domain encoder-decoder structure as the named entity recognition model among at least part of the encoder-decoder structure that has undergone domain transfer training.

[0011] According to a seventh aspect of an embodiment of the present invention, there is provided an electronic device, where the device includes: one or more processors; a computer-readable medium configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the model training method according to any one of the first aspect to the third aspect, or the recognition method according to the fourth aspect.

[0012] According to an eighth aspect of an embodiment of the present invention, there is provided a computer-readable medium having a computer program stored thereon, and when the program is executed by a processor, it implements the model training method according to any one of the first aspect to the third aspect, or the recognition method according to the fourth aspect.

[0013] Since the domain transfer training samples include random domain unlabeled training data, in domain transfer training, a wider range of training samples other than the source domain training samples and the target domain training samples are applied, thus improving the effectiveness of domain transfer training. Description of the Drawings

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0015] Figure 1ASchematic flowchart of a model training method for an example of Embodiment 1 of the present invention;

[0016] Figure 1B Schematic diagram of training based on labeled training data in the source domain for Embodiment 1 of the present invention;

[0017] Figure 1C Schematic diagram of training based on labeled training data in the target domain for Embodiment 1 of the present invention;

[0018] Figure 1D Schematic diagram of training based on unlabeled training data in a random domain for Embodiment 1 of the present invention;

[0019] Figure 1E Schematic diagram of the model training method for Embodiment 1 of the present invention;

[0020] Figure 1F Schematic flowchart of a model training method for another example of Embodiment 1 of the present invention;

[0021] Figure 2A Schematic diagram of the recognition method for Embodiment 2 of the present invention;

[0022] Figure 2B Schematic block diagram of a recognition method for an example of Embodiment 2 of the present invention;

[0023] Figure 2C Schematic block diagram of a recognition method for another example of Embodiment 2 of the present invention;

[0024] Figure 3 Schematic flowchart of the model training method for Embodiment 3 of the present invention;

[0025] Figure 4 Schematic flowchart of the model training method for Embodiment 4 of the present invention;

[0026] Figure 5 Schematic block diagram of the model training device for Embodiment 5 of the present invention;

[0027] Figure 6A Schematic block diagram of the recognition device for Embodiment 6 of the present invention;

[0028] Figure 6B Schematic block diagram of the recognition device for Embodiment 6 of the present invention;

[0029] Figure 7 Schematic structural diagram of the electronic device for Embodiment 7 of the present invention;

[0030] Figure 8 Hardware structure of the electronic device for Embodiment 8 of the present invention. Detailed implementation manners

[0031] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art shall fall within the scope of protection of the embodiments of the present invention.

[0032] The following further illustrates the specific implementation of the embodiments of the present invention in conjunction with the accompanying drawings of the embodiments of the present invention.

[0033] Figure 1A It is a schematic flowchart of the model training method for the first embodiment of the present invention; Figure 1A The model training method is applied to a named entity model. The method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as mobile phones, PADs, etc.) and PCs, etc. A named entity in the text refers to a word with an actual meaning, such as a person's name, an organization name, a place name, and all other entities identified by a name. For example, named entity recognition is used to convert the entity recognition task into a sequence labeling task, that is, to assign a label to each character in the input sentence. For example, the label consists of a prefix + type. For example, the prefix B indicates that this character is the start of an entity, the prefix I indicates that this character is inside an entity, the prefix E indicates that this character is the end of an entity, and the prefix S indicates that this entity is a single-word entity. The type is defined by different data sets and can be time (TIME), location (LOC), person name (PER), organization (ORG), or other custom types. Each type set must contain other types (O) to represent the characters that are not classified into the predefined entity type set. For example, "China" is converted into the label sequence "B-LOC E-LOC". According to the first embodiment of the present invention Figure 1A The model training method includes:

[0034] 110: Input the training data in the source domain and the unlabeled data in the random domain into the training model to obtain the first training result. The training model includes a shared encoder, a source domain decoder, a target domain decoder, and a random domain decoder;

[0035] 120: Screen the training data in the source domain and the unlabeled data in the random domain according to the training result to obtain the screened training data in the source domain and the screened unlabeled data in the random domain;

[0036] 130: Input the screened training data in the source domain, the screened unlabeled data in the random domain, and the training data in the target domain into the training model to obtain the second training result.

[0037] It should be understood that, for example, the shared encoder and the source domain decoder can also form a source domain codec structure. For example, the auto-encoder structure is a Variational Auto-Encoder (VAE) structure, or it can also be other auto-encoder structures, which are not limited in the embodiments of the present invention. VAE is a neural network designed to copy its input to the output. For example, it works by compressing the input into a hidden space representation and then reconstructing the output from this representation. For example, the source domain decoder and the target domain decoder can be any type of decoder, including but not limited to probabilistic statistical model decoders. For example, probabilistic model decoders include but are not limited to conditional random field (CRF) decoders, Hidden Markov Model (HMM) decoders, etc. For example, the source domain decoder and the target domain decoder can be different probabilistic statistical model decoders.

[0038] It should also be understood that the shared encoder can use any type of encoder, such as a probabilistic statistical model encoder, any neural network, etc. For example, neural networks such as recurrent neural networks. The above encoding layer can adopt a probabilistic statistical model, such as a conditional random field or a hidden Markov model. For example, the above target domain and source domain are unrelated or weakly related domains, so the above domain transfer training can be remote transfer training. For example, the unrelated or weakly related domains can be any two of the news domain, medical domain, judicial domain, e-commerce domain, engineering domain, etc.

[0039] Since the domain transfer training samples include random domain unlabeled training data, in domain transfer training, a wider range of training samples other than the source domain training samples and the target domain training samples are applied, thus improving the effectiveness of domain transfer training.

[0040] In another implementation manner of the present invention, inputting the training data of the source domain and the unlabeled data of the random domain into the training model to obtain a first training result, including: inputting the first sentence in the training data of the source domain into the training model; the training model encodes the first sentence using the shared encoder; the training model decodes the encoded first sentence using the source domain decoder to obtain the training data of the first sentence; the training model decodes the encoded first sentence using the random domain decoder to obtain the loss function of the first sentence.

[0041] In another implementation manner of the present invention, the training data in the source domain and the unlabeled data in the random domain are input into the training model to obtain a first training result, including: inputting a second sentence in the unlabeled data in the random domain into the training model; the training model encodes the second sentence by using a shared encoder; the training model decodes the encoded second sentence by using a random domain decoder to obtain a loss function of the second sentence.

[0042] In another implementation manner of the present invention, the training data in the source domain and the unlabeled data in the random domain are screened according to the training result to obtain the screened training data in the source domain and the screened unlabeled data in the random domain, including: calculating a score of a first sentence according to a loss function of the first sentence, and using the first sentence with the score within a first preset range as the screened training data in the source domain.

[0043] In another implementation manner of the present invention, the training data in the source domain and the unlabeled data in the random domain are screened according to the training result to obtain the screened training data in the source domain and the screened unlabeled data in the random domain, including: calculating a score of a second sentence according to a loss function of the second sentence, and using the second sentence with the score within a second preset range as the screened unlabeled data in the random domain.

[0044] In another implementation manner of the present invention, before inputting the training data in the source domain and the unlabeled data in the random domain into the training model to obtain a first training result, the method further includes: inputting the training data in the source domain and the training data in the target domain into the training model to obtain an initial training result.

[0045] In another implementation manner of the present invention, the data volume of the training data in the target domain is smaller than the data volume of the training data in the source domain.

[0046] In another implementation manner of the present invention, the shared encoder and the random domain decoder form a variational autoencoder structure.

[0047] In another implementation manner of the present invention, the source domain decoder and the target domain decoder are conditional random field decoders.

[0048] Figure 1F It is a schematic flowchart of a model training method for another example of Embodiment 1 of the present invention. Figure 1F The method includes:

[0049] 160: Determine an encoder-decoder structure, where the encoder-decoder structure includes a shared encoder and a source domain decoder, a target domain decoder, and a random domain decoder connected to the shared encoder. Among them, the shared encoder and the random domain decoder form an autoencoder structure, and the shared encoder and the target domain decoder form a target domain encoder-decoder structure.

[0050] 170: Based on the initial training samples, perform initial training on at least part of the codec structure, where the initial training samples include labeled training data in the source domain.

[0051] It should be understood that the initial training samples may also include labeled training data in the target domain. Figure 1B This is a schematic diagram of training based on labeled training data in the source domain in Embodiment 1 of the present invention. For example, for training based on labeled training data in the source domain, supervised training can be performed on the source domain codec structure, as Figure 1B shown on the left, thereby reducing the computational amount. For example, for training based on labeled training data in the source domain, semi-supervised training can be performed on the source domain codec structure and the self-codec structure, as Figure 1B shown on the right, thereby improving the training accuracy. For example, unsupervised training can be performed on the self-codec structure based on the labeled training data in the source domain.

[0052] 180: Based on the domain transfer training samples, perform domain transfer training on at least part of the codec structure, where the domain transfer training samples include labeled training data in the target domain and unlabeled training data in a random domain, and the unlabeled training data in the random domain is used for unsupervised training of the autoencoder structure.

[0053] For example, Figure 1C This is a schematic diagram of training based on labeled training data in the target domain in Embodiment 1 of the present invention. For example, for training based on labeled training data in the target domain, supervised training can be performed on the target domain codec, as Figure 1C shown on the left, thereby reducing the computational amount. For example, for training based on labeled training data in the target domain, semi-supervised training can be performed on the target domain codec structure and the self-codec structure, as Figure 1C shown on the right, thereby improving the training accuracy. For example, Figure 1D This is a schematic diagram of training based on unlabeled training data in a random domain in Embodiment 1 of the present invention. For example, as Figure 1D shown, for unsupervised training based on unlabeled training data in a random domain. For example, unsupervised training can be performed on the self-codec structure based on the labeled training data in the target domain.

[0054] For example, for performing domain transfer training on at least part of the codec structure based on the domain transfer training samples, multiple trainings can be performed, or one training can be performed. For example, in multiple trainings, the training samples for each training may overlap or may not overlap.

[0055] 190: In at least some of the codec structures trained through domain transfer, determine the target domain codec structure as the named entity recognition model.

[0056] For example, the named entity recognition model can perform named entity recognition on the task to be recognized. For example, perform named entity recognition on the text to be recognized in the above-mentioned domain that is the target domain. For example, for each input sentence in the task to be recognized, perform time series modeling on it. For example, model the sentence into a time series structure, and then use a recurrent neural network to encode the time series. Through the recurrent neural network, the vector representation of each word (phrase) in the sentence can be obtained. For example, input this vector into the decoder to obtain the corresponding output.

[0057] For example, this model training method can be used for the recognition of person names (Per), abbreviated person names (Aper), location names (Loc), abbreviated location names (Aloc), organization names (Org), time words (Tim), and quantity words (Num). Movie reviews analyze and comment on the director, actors, shooting techniques, plot, clues, environment, color, etc. of a movie. With the development and popularization of the Internet, people's discussions about movies have moved from offline to online, resulting in a large amount of online movie reviews.

[0058] Person name recognition is one of the tasks of named entity recognition (NER). In the person name recognition scenario, the embodiments of the present invention can perform recognition on movie reviews. Extracting relevant director and actor names from a large number of movie reviews can provide information support for various upper-layer applications such as movie marketing, star marketing, and person sentiment analysis. Due to the diversity and complexity of the composition of person names, the model training method of the present invention can effectively achieve person name recognition.

[0059] Since the domain transfer training samples include random domain unlabeled training data, in domain transfer training, a wider range of training samples other than the source domain training samples and the target domain training samples are applied, thus improving the effectiveness of domain transfer training. In addition, the random domain unlabeled training data is used for unsupervised training of the autoencoder structure. The autoencoder structure can perform deep feature extraction, enabling the shared encoder to learn deep features, thereby achieving effective domain transfer training for the named entity recognition model.

[0060] In other words, when the source domain and the target domain are far apart, unlabeled random domain training data is used to enable the model to share an encoder between the source domain and the target domain. Since the unlabeled random domain training data serves as an auxiliary bridge for data from the source domain to the target domain, a domain-agnostic encoder is learned during remote migration. The present invention proposes using an autoencoder to select intermediate data to assist in the learning of the encoder, thereby achieving remote knowledge transfer.

[0061] In another implementation manner of the present invention, for example, when the amount of labeled training data in the target domain is less than the amount of labeled training data in the source domain, the solution of this embodiment can be used to perform training based on unlabeled random domain training data when the amount of labeled training data in the target domain is less than that in the source domain, thereby ensuring the accuracy of domain transfer training.

[0062] For example, the amount of unlabeled random domain training data can be greater than the amount of labeled training data in the target domain. Thus, by using the unlabeled random domain training data, the autoencoder structure can learn the common features of data from different domains and then learn domain-agnostic representations.

[0063] For example, the amount of unlabeled random domain training data can be greater than the amount of labeled training data in the target domain. Since the unlabeled random domain training data does not need to be labeled, the workload of manual labeling is reduced while ensuring the training accuracy.

[0064] In another implementation manner of the present invention, the domain transfer training samples include first domain transfer training samples and second domain transfer training samples. The first domain transfer training samples include labeled training data in the target domain, and the second domain transfer training samples include unlabeled random domain training data. The second domain transfer training samples further include labeled training data in the source domain. Among them, based on the domain transfer training samples, domain transfer training is performed on at least part of the codec structure, including: performing first domain transfer training on at least part of the codec structure based on the first domain transfer training samples, and performing second domain transfer training on at least part of the codec structure based on the second domain transfer training samples.

[0065] For example, at least part of the codec structure can be trained multiple times. For each training, the current training samples are determined by migrating training samples from the first domain. For example, the amount of the current training samples can be determined based on the relative value between the amount of labeled training data in the source domain and the amount of labeled training data in the target domain. For example, the larger the ratio of the amount of labeled training data in the source domain to the amount of labeled training data in the target domain, the larger the proportion of the labeled training data in the target domain in the current training samples compared to the migrated training samples from the first domain. For example, the ratio of the amount of labeled training data in the source domain to the amount of labeled training data in the target domain is positively correlated with the proportion of the labeled training data in the target domain in the current training samples compared to the migrated training samples from the first domain.

[0066] In another implementation of the present invention, based on the migrated training samples from the second domain, second domain transfer training is performed on at least part of the codec structure, including: training at least part of the codec structure multiple times. For each training, based on the autoencoder structure, the current training samples are determined from the migrated training samples from the second domain; and based on the current training samples, current training is performed on at least part of the codec structure.

[0067] In another implementation of the present invention, determining the current training samples from the migrated training samples from the second domain based on the autoencoder structure includes: inputting the migrated training samples from the second domain into the autoencoder structure; and determining the current training samples from the migrated training samples from the second domain by calculating the current loss function value using the current parameters of the autoencoder structure.

[0068] For example, Figure 1E is a schematic diagram of the model training method according to the first embodiment of the present invention. As Figure 1E shown, it is divided into ① initial training, ② sample screening, and ③ domain transfer training. Finally, the codec structure in the target domain (the bold box in the figure) is used as the named entity recognition model. It should be understood that for the initial training and the domain transfer training, the training methods based on different data are as Figure 1B - 1D shown, Figure 1E which only shows an example, and the embodiments of the present invention are not limited thereto. ② Sample screening and ③ domain transfer training can be regarded as a whole. For example, one or more sample screenings can be performed. Before each transfer training, sample screening is carried out. For example, sample screening can also be performed on the training data in the target domain, so as to improve the training efficiency, that is, reduce the calculation amount while ensuring the training accuracy.

[0069] For example, from the transferred training samples in the second domain, the current training samples are screened. For example, based on a preset threshold, at least one of multiple data is screened as the current training sample. For example, the preset threshold can be a relative preset threshold or other preset thresholds. For example, the relative preset threshold can be the proportion of the screened data volume. For example, the current training samples include a first current training sample and a second current training sample, and the second current training sample is the screened training sample among the current training samples. For example, by using the current parameters of the autoencoder structure, multiple current loss function values corresponding to multiple data in the transferred training samples in the second domain are determined. For example, the loss function value of the second current training sample is less than the loss function value of the first current training sample.

[0070] In another implementation manner of the present invention, by using the current parameters of the autoencoder structure to calculate the current loss function value, to determine the current training samples from the transferred training samples in the second domain, includes: by using the current parameters of the autoencoder structure, determining multiple current loss function values corresponding to multiple data in the transferred training samples in the second domain; based on a preset threshold, determining at least one of the multiple data as the current training sample, wherein each of the at least one current loss function values corresponding to the at least one data is less than the preset threshold.

[0071] For example, the loss function of the autoencoder structure can be any loss function. For example, the loss function includes but is not limited to square loss function, logarithmic loss function, mean square error, root mean square error, average error, cross-entropy cost function, maximum entropy loss function, etc.

[0072] In another implementation manner of the present invention, a shared encoder and a source domain decoder form a source domain encoder-decoder structure, and based on the current training samples, perform current training on at least part of the encoder-decoder structure, includes: determining the source domain labeled training data in the current training samples; based on the source domain labeled training data in the current training samples, performing supervised training on the source domain encoder-decoder structure.

[0073] In another implementation manner of the present invention, a shared encoder and a source domain decoder form a source domain encoder-decoder structure, and based on the current training samples, perform current training on at least part of the encoder-decoder structure, includes: determining the source domain labeled training data in the current training samples; based on the source domain labeled training data in the current training samples, performing semi-supervised training on the source domain encoder-decoder structure and the autoencoder structure.

[0074] For example, the source domain has labeled training data including first source domain labeled training data and second source domain labeled training data. Supervised training is performed using the first source domain labeled training data, and semi-supervised training is performed using the second source domain labeled training data. For example, the first source domain labeled training data and the second source domain labeled training data are determined by the loss function of the autoencoder structure. For example, the loss function value of the first source domain labeled training data is greater than the loss function value of the second source domain labeled training data. Since data with a smaller loss function value is used for semi-supervised training, and data with a larger loss function value is used for supervised training, the autoencoder structure performs deep feature extraction, reducing the computational amount while ensuring the training accuracy.

[0075] In another implementation manner of the present invention, semi-supervised training is performed on the source domain codec structure and the autoencoder structure based on the source domain labeled training data in the current training sample, including: calculating the loss function value of the source domain labeled training data in the current training sample based on the loss function of the source domain codec structure and the loss function of the autoencoder structure; performing semi-supervised training on the source domain codec structure and the autoencoder structure based on the loss function value.

[0076] For example, the loss functions of the source domain codec structure and the autoencoder structure can be summed to obtain a loss function sum, and semi-supervised training is performed using the obtained loss function sum. For example, backpropagation is performed based on the loss function sum for semi-supervised training. For example, the loss functions of the source domain codec structure and the autoencoder structure can be weighted and summed to obtain a weighted sum of the two loss functions, and semi-supervised training is performed using the weighted sum. For example, a function with the loss functions of the source domain codec structure and the autoencoder structure as binary independent variables can be constructed, and semi-supervised learning training is performed through the binary function.

[0077] In another implementation manner of the present invention, first domain transfer training is performed on at least part of the codec structure based on the first domain transfer training sample, including: performing semi-supervised training on the target domain codec structure and the autoencoder structure based on the target domain labeled training data.

[0078] For example, the labeled training data in the target domain includes the labeled training data in the first target domain and the labeled training data in the second target domain. Supervised training is performed using the labeled training data in the first target domain, and semi-supervised training is performed using the labeled training data in the second target domain. For example, the labeled training data in the first target domain and the labeled training data in the second target domain are determined by the loss function of the autoencoder structure. For example, the loss function value of the labeled training data in the first target domain is greater than the loss function value of the labeled training data in the second target domain. Since data with a smaller loss function value is used for semi-supervised training and data with a larger loss function value is used for supervised training, the autoencoder structure performs deep feature extraction, reducing the computational amount while ensuring the training accuracy.

[0079] In another implementation manner of the present invention, based on the labeled training data in the target domain, semi-supervised training is performed on the target domain codec structure and the autoencoder structure, including: calculating the loss function value of the labeled training data in the target domain through the loss function of the target domain codec structure and the loss function of the autoencoder structure; performing semi-supervised training on the target domain codec structure and the autoencoder structure through the loss function value.

[0080] For example, the loss functions of the target domain codec structure and the autoencoder structure can be summed to obtain the sum of the loss functions, and semi-supervised training is performed using the sum of the loss functions shown. For example, based on the sum of the loss functions, backpropagation is performed for semi-supervised training. For example, the loss functions of the target domain codec structure and the autoencoder structure can be weighted and summed to obtain the weighted sum of the two loss functions, and semi-supervised training is performed using the weighted sum. For example, a function with the loss functions of the target domain codec structure and the autoencoder structure as binary independent variables can be constructed, and semi-supervised learning training is performed through the binary function.

[0081] The loss functions of the target domain codec structure and the source domain codec structure can be any loss functions. For example, the loss functions include but are not limited to the square loss function, the logarithmic loss function, the mean square error, the root mean square error, the cross-entropy cost function, the maximum entropy loss function, etc. For example, the loss functions of the target domain codec structure, the source domain codec structure, and the autoencoder structure can be the same loss function or different loss functions. For example, for any one of the target domain codec structure, the source domain codec structure, and the autoencoder structure, different loss functions can be used in the initial training and the domain transfer training. For example, in multiple trainings during the domain transfer training, the same structure can use different loss functions, and the embodiments of the present invention do not limit this.

[0082] Figure 2A Schematic diagram of the recognition method according to the second embodiment of the present invention. Figure 2A The recognition method can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as mobile phones, PADs, etc.) and PCs, etc. As Figure 2A shown, the terminal device A sends a named entity recognition request for the text to be recognized to a network such as the mobile Internet. A server such as a cloud server receives the text to be recognized and the recognition request from the network, and in response to the recognition request, feeds back the named entity recognition result of the text to be recognized to the network. The terminal device A receives the recognition result from the network for subsequent operations by the user.

[0083] Figure 2B Schematic block diagram of the recognition method according to an example of the second embodiment of the present invention. Figure 2B The recognition method includes:

[0084] 210: Obtain the task to be recognized in the target domain;

[0085] 220: Use the named entity recognition model to perform named entity recognition on the task to be recognized in the target domain, where the named entity recognition model is trained by the model training method as in Embodiment 1.

[0086] For this named entity recognition model, since the domain transfer training samples include random domain unlabeled training data, in the domain transfer training, a wider range of training samples other than the source domain training samples and the target domain training samples are applied, thus improving the effectiveness of the domain transfer training. In addition, the random domain unlabeled training data is used for unsupervised training of the autoencoder structure, and the autoencoder structure can perform deep feature extraction, enabling the shared encoder to learn deep features, thereby achieving effective domain transfer training for the named entity recognition model. Therefore, effective recognition is performed using this named entity recognition model.

[0087] Figure 2C Schematic block diagram of the recognition method according to another example of the second embodiment of the present invention. Figure 2C The recognition method includes:

[0088] 230: Input the data to be predicted in the target domain into the training model, and the training model is trained by the above method;

[0089] 240: The training model encodes the data to be predicted using the shared encoder;

[0090] 250: The training model decodes the encoded data to be predicted using the target domain decoder to obtain the prediction result.

[0091] Figure 3Schematic flowchart of the model training method according to Embodiment 3 of the present invention; Figure 3 The model training method is applied to a named entity recognition model and includes:

[0092] 310: Determine the codec structure, which includes a shared encoder and an e-commerce domain decoder, a non-e-commerce domain decoder, and a random domain decoder connected to the shared encoder. Among them, the shared encoder and the random domain decoder form an autoencoder structure, and the shared encoder and the non-e-commerce domain decoder form a non-e-commerce domain codec structure;

[0093] 320: Based on the initial training samples, perform initial training on at least part of the codec structure, where the initial training samples include labeled training data in the e-commerce domain;

[0094] 330: Based on the domain transfer training samples, perform domain transfer training on at least part of the codec structure, where the domain transfer training samples include labeled training data in the non-e-commerce domain and unlabeled training data in the random domain, and the unlabeled training data in the random domain is used for unsupervised training of the autoencoder structure;

[0095] 340: Among at least part of the codec structures that have undergone domain transfer training, determine the non-e-commerce domain codec structure as the named entity recognition model.

[0096] It should be understood that the non-e-commerce domain includes but is not limited to any domain such as the news domain, the medical domain, the judicial domain, the engineering domain, etc.

[0097] Since the domain transfer training samples include unlabeled training data in the random domain, a wider range of training samples are applied in the domain transfer training, thus improving the effectiveness of the domain transfer training. In addition, the unlabeled training data in the random domain is used for unsupervised training of the autoencoder structure, and the autoencoder structure can perform deep feature extraction, enabling the shared encoder to learn deep features, thereby achieving effective domain transfer training for the named entity recognition model.

[0098] Figure 4 Schematic flowchart of the model training method according to Embodiment 4 of the present invention; Figure 4 The model training method is applied to a named entity recognition model and includes:

[0099] 410: Determine the codec structure, which includes a shared encoder and a non-e-commerce domain decoder, an e-commerce domain decoder, and a random domain decoder connected to the shared encoder. Among them, the shared encoder and the random domain decoder form an autoencoder structure, and the shared encoder and the e-commerce domain decoder form an e-commerce domain codec structure;

[0100] 420: Based on the initial training samples, perform initial training on at least part of the codec structure, where the initial training samples include labeled training data in non-e-commerce domains;

[0101] 430: Based on the domain transfer training samples, perform domain transfer training on at least part of the codec structure, where the domain transfer training samples include labeled training data in the e-commerce domain and unlabeled training data in random domains, and the unlabeled training data in random domains is used for unsupervised training of the autoencoder structure;

[0102] 440: In at least part of the codec structure that has undergone domain transfer training, determine the e-commerce domain codec structure as the named entity recognition model.

[0103] The non-e-commerce domains include, but are not limited to, any of the domains such as the news domain, medical domain, judicial domain, engineering domain, etc.

[0104] Since the domain transfer training samples include unlabeled training data in random domains, a more extensive set of training samples is applied in the domain transfer training, thus improving the effectiveness of the domain transfer training. In addition, the unlabeled training data in random domains is used for unsupervised training of the autoencoder structure, and the autoencoder structure can perform deep feature extraction, enabling the shared encoder to learn deep features, thereby achieving effective domain transfer training for the named entity recognition model.

[0105] Figure 5 It is a schematic block diagram of the model training device according to Embodiment 5 of the present invention; Figure 3 The model training device includes:

[0106] A codec structure determination module 510 that determines the codec structure, where the codec structure includes a shared encoder and a source domain decoder, a target domain decoder, and a random domain decoder connected to the shared encoder. Among them, the shared encoder and the random domain decoder form an autoencoder structure, and the shared encoder and the target domain decoder form a target domain codec structure;

[0107] An initial training module 520 that performs initial training on at least part of the codec structure based on the initial training samples, where the initial training samples include labeled training data in the source domain;

[0108] A domain transfer training module 530 that performs domain transfer training on at least part of the codec structure based on the domain transfer training samples, where the domain transfer training samples include labeled training data in the target domain and unlabeled training data in random domains, and the unlabeled training data in random domains is used for unsupervised training of the autoencoder structure;

[0109] The model determination module 540 determines the target domain codec structure as a named entity recognition model in at least a part of the codec structure that has undergone domain adaptation training.

[0110] Since the domain adaptation training samples include random domain unlabeled training data, a wider range of training samples other than the source domain training samples and the target domain training samples are applied in the domain adaptation training, thus improving the effectiveness of the domain adaptation training. In addition, the random domain unlabeled training data is used for unsupervised training of the autoencoder structure, and the autoencoder structure can perform deep feature extraction, enabling the shared encoder to learn deep features, thereby achieving effective domain adaptation training for the named entity recognition model.

[0111] In another implementation of the present invention, the domain adaptation training samples include first domain adaptation training samples and second domain adaptation training samples. The first domain adaptation training samples include target domain labeled training data, and the second domain adaptation training samples include random domain unlabeled training data. The second domain adaptation training samples further include source domain labeled training data.

[0112] The domain adaptation training module is specifically configured to: perform first domain adaptation training on at least a part of the codec structure based on the first domain adaptation training samples, and perform second domain adaptation training on at least a part of the codec structure based on the second domain adaptation training samples.

[0113] In another implementation of the present invention, the domain adaptation training module is specifically configured to: perform multiple trainings on at least a part of the codec structure. For each training, based on the autoencoder structure, determine the current training samples from the second domain adaptation training samples; and perform the current training on at least a part of the codec structure based on the current training samples.

[0114] In another implementation of the present invention, the domain adaptation training module is specifically configured to: input the second domain adaptation training samples into the autoencoder structure; and determine the current training samples from the second domain adaptation training samples by calculating the current loss function value using the current parameters of the autoencoder structure.

[0115] In another implementation of the present invention, the domain adaptation training module is specifically configured to: determine multiple current loss function values corresponding to multiple data in the second domain adaptation training samples by using the current parameters of the autoencoder structure; and determine at least one of the multiple data as the current training samples based on a preset threshold, where each loss function value in the at least one current loss function value corresponding to the at least one data is less than the preset threshold.

[0116] In another implementation of the present invention, the shared encoder and the source domain decoder form a source domain codec structure. The domain transfer training module is specifically configured to: determine the source domain labeled training data in the current training sample; and perform supervised training on the source domain codec structure based on the source domain labeled training data in the current training sample. In another implementation of the present invention, the shared encoder and the source domain decoder form a source domain codec structure. The domain transfer training module is specifically configured to: determine the source domain labeled training data in the current training sample; and perform semi-supervised training on the source domain codec structure and the autoencoder structure based on the source domain labeled training data in the current training sample.

[0117] In another implementation of the present invention, the domain transfer training module is specifically configured to: calculate the loss function value of the source domain labeled training data in the current training sample based on the loss function of the source domain codec structure and the loss function of the autoencoder structure; and perform semi-supervised training on the source domain codec structure and the autoencoder structure based on the loss function value.

[0118] In another implementation of the present invention, the initial training module is specifically configured to: perform semi-supervised training on the target domain codec structure and the autoencoder structure based on the target domain labeled training data.

[0119] In another implementation of the present invention, the initial training module is specifically configured to: calculate the loss function value of the target domain labeled training data through the loss function of the target domain codec structure and the loss function of the autoencoder structure; and perform semi-supervised training on the target domain codec structure and the autoencoder structure through the loss function value.

[0120] In another implementation of the present invention, the initial training sample further includes the target domain labeled training data.

[0121] In another implementation of the present invention, the data volume of the target domain labeled training data is smaller than that of the source domain labeled training data.

[0122] In another implementation of the present invention, the autoencoder structure is a variational autoencoder structure.

[0123] In another implementation of the present invention, the source domain decoder and the target domain decoder are conditional random field decoders.

[0124] The device of this embodiment is used to implement the corresponding methods in the foregoing multiple method embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here. In addition, the function implementation of each module in the device of this embodiment can refer to the description of the corresponding part in the foregoing method embodiments, which will not be elaborated here either.

[0125] Figure 6A Schematic block diagram of the recognition device according to the sixth embodiment of the present invention; Figure 6A The named entity recognition device includes:

[0126] An acquisition module 610, which acquires the task to be recognized in the target domain;

[0127] A recognition module 620, which performs named entity recognition on the task to be recognized in the target domain by using a named entity recognition model, wherein the named entity recognition model is trained by the model training method as in Embodiment 1.

[0128] For this named entity recognition model, since the domain transfer training samples include random domain unlabeled training data, in the domain transfer training, a wider range of training samples other than the source domain training samples and the target domain training samples are applied, so the effectiveness of the domain transfer training is improved. In addition, the random domain unlabeled training data is used for unsupervised training of the autoencoder structure, and the autoencoder structure can perform deep feature extraction, so that the shared encoder can learn deep features, thereby realizing effective domain transfer training for the named entity recognition model. Therefore, effective recognition is performed by using this named entity recognition model.

[0129] Figure 6B Schematic block diagram of the recognition device according to the sixth embodiment of the present invention. Figure 6B The recognition device includes:

[0130] An input module 630, which inputs the data to be predicted in the target domain into the training model, and the training model is trained by the above method;

[0131] An encoding module 640, and the training model encodes the data to be predicted by using the shared encoder;

[0132] A decoding module 650, and the training model decodes the encoded data to be predicted by using the target domain decoder to obtain a prediction result;

[0133] An output module 660, which outputs the prediction result.

[0134] The device in this embodiment is used to implement the corresponding methods in the foregoing multiple method embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated herein. In addition, the function implementation of each module in the device in this embodiment can refer to the description of the corresponding part in the foregoing method embodiments, which will not be elaborated herein either.

[0135] Figure 7 Schematic structural diagram of an electronic device according to the seventh embodiment of the present application; the electronic device may include:

[0136] One or more processors 701;

[0137] A computer-readable medium 702, which can be configured to store one or more programs,

[0138] When the one or more programs are executed by the one or more processors, the one or more processors implement the model training method as described in the first embodiment above.

[0139] The device of this embodiment is used to implement the corresponding methods in the foregoing multiple method embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here. In addition, the function implementation of each module in the device of this embodiment can be referred to the description of the corresponding part in the foregoing method embodiments, which will not be elaborated here either.

[0140] Figure 8 This is the hardware structure of the electronic device in the eighth embodiment of this application; as Figure 8 shown, the hardware structure of the electronic device may include: a processor 801, a communication interface 802, a computer-readable medium 803, and a communication bus 804;

[0141] Among them, the processor 801, the communication interface 802, and the computer-readable medium 803 complete mutual communication through the communication bus 804;

[0142] Optionally, the communication interface 802 may be an interface of a communication module;

[0143] Among them, the processor 801 can be specifically configured to: determine a codec structure, the codec structure includes a shared encoder and a source domain decoder, a target domain decoder, and a random domain decoder connected to the shared encoder, wherein, the shared encoder and the random domain decoder form an autoencoder structure, and the shared encoder and the target domain decoder form a target domain codec structure;

[0144] Based on the initial training samples, perform initial training on at least part of the codec structure, wherein, the initial training samples include source domain labeled training data;

[0145] Based on the domain transfer training samples, perform domain transfer training on at least part of the codec structure, wherein, the domain transfer training samples include target domain labeled training data and random domain unlabeled training data, and the random domain unlabeled training data is used to perform unsupervised training on the autoencoder structure;

[0146] In at least part of the codec structure that has undergone domain transfer training, determine the target domain codec structure as a named entity recognition model, or

[0147] Obtain a target domain task to be recognized;

[0148] Use the named entity recognition model to perform named entity recognition on the task to be recognized in the target field, where the named entity recognition model is trained by the model training method as described above.

[0149] The processor 801 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0150] The computer-readable medium 803 can be, but is not limited to, a random access storage medium (RAM), a read-only storage medium (ROM), a programmable read-only storage medium (PROM), an erasable programmable read-only storage medium (EPROM), an electrically erasable programmable read-only storage medium (EEPROM), etc.

[0151] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code configured to execute the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed structurally through the communication part and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed. It should be noted that the computer-readable medium described in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable medium can, for example, but is not limited to, be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access storage medium (RAM), a read-only storage medium (ROM), an erasable programmable read-only storage medium (EPROM or flash memory), an optical fiber, a portable compact disk read-only storage medium (CD-ROM), an optical storage medium, a magnetic storage medium, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program configured to be used by or combined with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0152] Computer program code configured to perform the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0153] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions configured to implement the specified logical function. There are specific sequential relationships in the above specific embodiments, but these sequential relationships are only exemplary. In specific implementations, these steps may be fewer, more, or the execution order may be adjusted. That is, in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0154] The modules described in the embodiments of this application can be implemented in software or in hardware. The names of these modules do not, in some cases, constitute a limitation on the module itself.

[0155] As another aspect, this application also provides a computer-readable medium having a computer program stored thereon, which when executed by a processor implements the method described in Embodiment 1 or Embodiment 2 above.

[0156] As another aspect, the present application also provides a computer-readable medium, which may be included in the device described in the above embodiments; or may exist alone without being assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the device, the device is caused to: determine a codec structure, the codec structure including a shared encoder and a source domain decoder, a target domain decoder, and a random domain decoder connected to the shared encoder, wherein the shared encoder and the random domain decoder form an autoencoder structure, and the shared encoder and the target domain decoder form a target domain codec structure;

[0157] Based on initial training samples, perform initial training on at least part of the codec structure, wherein the initial training samples include source domain labeled training data;

[0158] Based on domain transfer training samples, perform domain transfer training on at least part of the codec structure, wherein the domain transfer training samples include target domain labeled training data and random domain unlabeled training data, and the random domain unlabeled training data is used for unsupervised training of the autoencoder structure;

[0159] In at least part of the codec structure that has undergone domain transfer training, determine the target domain codec structure as a named entity recognition model, or

[0160] Obtain a target domain task to be recognized;

[0161] Use the named entity recognition model to perform named entity recognition on the target domain task to be recognized, wherein the named entity recognition model is trained by the model training method as described above.

[0162] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to:

[0163] (1) Mobile communication devices: Such devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones (such as iPhone), multimedia phones, functional phones, and low-end phones, etc.

[0164] (2) Ultra-mobile personal computer devices: Such devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc., such as iPad.

[0165] (3) Portable entertainment devices: Such devices can display and play multimedia content. This type of device includes: audio and video players (such as iPods), handheld game consoles, e-books, as well as smart toys and portable in-vehicle navigation devices.

[0166] (4) Servers: Devices that provide computing services. The composition of a server includes a processor, hard disk, memory, system bus, etc. Servers are similar to general computer architectures, but due to the need to provide highly reliable services, they have higher requirements in terms of processing power, stability, reliability, security, scalability, manageability, etc.

[0167] (5) Other electronic devices with data interaction functions.

[0168] So far, specific embodiments of this subject matter have been described. Other embodiments are within the scope of the above model training method. In some cases, the actions recited in the above model training method can be executed in a different order and still achieve the desired result. Additionally, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing can be advantageous.

[0169] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many improvements to method flows today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing some logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0170] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an Application Specific Integrated Circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, ASICs, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0171] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0172] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0173] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0174] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.

[0175] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.

[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.

[0177] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0178] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0179] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0180] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0181] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0182] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific transactions or implement specific abstract data types. The present application may also be practiced in distributed computing environments where transactions are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0183] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0184] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the above model training method of the present application.

Claims

1. A model training method, applied to a named entity recognition model, includes: Inputting the training data in the source domain and the unlabeled data in the random domain into the training model to obtain a first training result. The training model includes a shared encoder, a source domain decoder, a target domain decoder, and a random domain decoder. The first training result includes a loss function of a first sentence and a loss function of a second sentence. The loss function of the first sentence is obtained by decoding the encoded first sentence using the random domain decoder. The encoded first sentence is obtained by encoding the first sentence in the training data in the source domain using the shared encoder. The loss function of the second sentence is obtained by decoding the encoded second sentence using the random domain decoder. The encoded second sentence is obtained by encoding the second sentence in the unlabeled data in the random domain using the shared encoder. Screening the training data in the source domain and the unlabeled data in the random domain according to the first training result to obtain the screened training data in the source domain and the screened unlabeled data in the random domain. The screened training data in the source domain is the first sentence with a score within a first preset range. The score of the first sentence is determined according to the loss function of the first sentence. The screened unlabeled data in the random domain is the second sentence with a score within a second preset range. The score of the second sentence is determined according to the loss function of the second sentence. Inputting the screened training data in the source domain, the screened unlabeled data in the random domain, and the training data in the target domain into the training model to obtain a second training result. The second training result is the trained training model. The screened training data in the source domain includes first source domain labeled training data and second source domain labeled training data. The loss function value of the first source domain labeled training data is greater than that of the second source domain labeled training data. The first source domain labeled training data is used for supervised training of the source domain encoder-decoder structure. The second source domain labeled training data is used for semi-supervised training of the source domain encoder-decoder structure and the autoencoder structure. The source domain encoder-decoder structure is formed by the shared encoder and the source domain decoder. The autoencoder structure is formed by the shared encoder and the random domain decoder.

2. The method according to claim 1, wherein Inputting the training data in the source domain into the training model to obtain the first training result includes: Inputting the first sentence in the training data in the source domain into the training model. The training model encodes the first sentence using the shared encoder. The training model decodes the encoded first sentence using the source domain decoder to obtain the training data of the first sentence. The training model decodes the encoded first sentence using the random domain decoder to obtain the loss function of the first sentence.

3. The method according to claim 2, wherein, Inputting the unlabeled data in the random domain into the training model to obtain the first training result includes: Input the second sentence in the unlabeled data of the random domain into the training model; The training model encodes the second sentence by using the shared encoder; The training model decodes the encoded second sentence by using the random domain decoder to obtain the loss function of the second sentence.

4. The method according to claim 2, wherein, Screen the training data of the source domain and the unlabeled data of the random domain according to the training result to obtain the screened training data of the source domain, including: Calculate the score of the first sentence according to the loss function of the first sentence, and use the first sentence with the score within the first preset range as the screened training data of the source domain.

5. The method according to claim 3, wherein Screen the training data of the source domain and the unlabeled data of the random domain according to the training result to obtain the screened unlabeled data of the random domain, including: Calculate the score of the second sentence according to the loss function of the second sentence, and use the second sentence with the score within the second preset range as the screened unlabeled data of the random domain.

6. The method according to any one of claims 1 to 5, wherein Before inputting the training data of the source domain and the unlabeled data of the random domain into the training model to obtain the first training result, the method further includes: Input the training data of the source domain and the training data of the target domain into the training model to obtain the initial training result.

7. The method according to claim 1, wherein, The data volume of the training data of the target domain is smaller than the data volume of the training data of the source domain.

8. The method according to claim 1, wherein The shared encoder and the random domain decoder form a variational autoencoder structure.

9. The method according to claim 1, wherein The source domain decoder and the target domain decoder are conditional random field decoders.

10. An identification method, including: Input the data to be predicted in the target domain into the training model, and the training model is trained by the method according to any one of claims 1-9; The training model encodes the data to be predicted by using the shared encoder; The training model decodes the encoded data to be predicted by using the target domain decoder to obtain the prediction result.

11. A model training method, applied to a named entity recognition model, including: Determine the codec structure, where the codec structure includes a shared encoder and a source domain decoder, a target domain decoder, and a random domain decoder connected to the shared encoder. Among them, the shared encoder and the random domain decoder form an autoencoder structure, and the shared encoder and the target domain decoder form a target domain codec structure; Based on the initial training samples, perform initial training on at least part of the codec structure, where the initial training samples include the labeled training data of the source domain; Based on the domain transfer training samples and the training data of the source domain, perform domain transfer training on at least part of the codec structure, where the domain transfer training samples include labeled training data in the target domain and unlabeled training data in a random domain, the unlabeled training data in the random domain is used for unsupervised training of the autoencoder structure, the training data of the source domain includes labeled training data in the first source domain and labeled training data in the second source domain, the loss function value of the labeled training data in the first source domain is greater than the loss function value of the labeled training data in the second source domain, the labeled training data in the first source domain is used for supervised training of the source domain codec structure, the labeled training data in the second source domain is used for semi-supervised training of the source domain codec structure and the autoencoder structure, the source domain codec structure is formed by the shared encoder and the source domain decoder, and the autoencoder structure is formed by the shared encoder and the random domain decoder; In at least part of the codec structure that has undergone domain transfer training, determine the target domain codec structure as the named entity recognition model.

12. A model training device, applied to a named entity recognition model, includes: A first input module that inputs the training data of the source domain and the unannotated data of the random domain into the training model to obtain a first training result, where the training model includes a shared encoder, a source domain decoder, a target domain decoder, and a random domain decoder, the first training result includes the loss function of the first sentence and the loss function of the second sentence, the loss function of the first sentence is obtained by decoding the encoded first sentence using the random domain decoder, the encoded first sentence is obtained by encoding the first sentence in the training data of the source domain using the shared encoder, the loss function of the second sentence is obtained by decoding the encoded second sentence using the random domain decoder, and the encoded second sentence is obtained by encoding the second sentence in the unannotated data of the random domain using the shared encoder; A screening module that screens the training data of the source domain and the unannotated data of the random domain according to the first training result to obtain the screened training data of the source domain and the screened unannotated data of the random domain, where the screened training data of the source domain is the first sentence with a score within a first preset range, the score of the first sentence is determined according to the loss function of the first sentence, and the screened unannotated data of the random domain is the second sentence with a score within a second preset range, and the score of the second sentence is determined according to the loss function of the second sentence; A second input module that inputs the filtered training data of the source domain, the filtered unlabeled data of the random domain, and the training data of the target domain into the training model to obtain a second training result, where the second training result is the trained training model. The filtered training data of the source domain includes first source domain labeled training data and second source domain labeled training data. The loss function value of the first source domain labeled training data is greater than that of the second source domain labeled training data. The first source domain labeled training data is used for supervised training of the source domain codec structure, and the second source domain labeled training data is used for semi-supervised training of the source domain codec structure and the autoencoder structure. The source domain codec structure is formed by the shared encoder and the source domain decoder, and the autoencoder structure is formed by the shared encoder and the random domain decoder.

13. An identification device, comprising: An input module that inputs the data to be predicted in the target domain into the training model, where the training model is obtained by training according to any one of claims 1-9; An encoding module, where the training model encodes the data to be predicted by using the shared encoder; A decoding module, where the training model decodes the encoded data to be predicted by using the target domain decoder, to obtain a prediction result; An output module that outputs the prediction result.

14. A model training device applied to a named entity recognition model, comprising: A codec structure determination module that determines a codec structure, where the codec structure includes a shared encoder and a source domain decoder, a target domain decoder, and a random domain decoder connected to the shared encoder. The shared encoder and the random domain decoder form an autoencoder structure, and the shared encoder and the target domain decoder form a target domain codec structure; An initial training module that performs initial training on at least part of the codec structure based on an initial training sample, where the initial training sample includes source domain labeled training data; The domain transfer training module performs domain transfer training on at least part of the codec structure based on domain transfer training samples and the training data of the source domain. Among them, the domain transfer training samples include labeled training data in the target domain and unlabeled training data in a random domain. The unlabeled training data in the random domain is used for unsupervised training of the autoencoder structure. The training data of the source domain includes labeled training data in the first source domain and labeled training data in the second source domain. The loss function value of the labeled training data in the first source domain is greater than that of the labeled training data in the second source domain. The labeled training data in the first source domain is used for supervised training of the source domain codec structure. The labeled training data in the second source domain is used for semi-supervised training of the source domain codec structure and the autoencoder structure. The source domain codec structure is formed by the shared encoder and the source domain decoder. The autoencoder structure is formed by the shared encoder and the random domain decoder; The model determination module determines the target domain codec structure as the named entity recognition model among at least part of the codec structures that have undergone domain transfer training.

15. An electronic device, the device includes: One or more processors; A computer-readable medium configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-11.

16. A computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • A method and apparatus for text translation

    CN109190134A