Training Method, Device, Equipment and Storage Medium of Retrieval Matching Model

By using the training corpus pairs of existing fields to train the generative model, generating the correlation training corpus pairs of new fields, and training the search matching model, the problem of lack of training data in the search scenario of new fields is solved, and the search service adapted to the new fields is quickly deployed.

CN112579870BActive Publication Date: 2025-06-27BEIJING SANKUAI ONLINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011529224.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-22
Publication Date
2025-06-27
Estimated Expiration
2040-12-22

AI Technical Summary

Technical Problem

In the search scenario of new fields, due to the lack of correlation training corpus pairs, the search matching model cannot be effectively trained, resulting in the rapid deployment of search services adapted to new fields.

Method used

By obtaining the generative model obtained by training the first correlation training corpus pair of the existing field, input the searched text of the target field to generate the second correlation training corpus pair, and input it to the initialization model for training, obtaining a search matching model that is suitable for the target field.

Benefits of technology

It realizes the rapid deployment of retrieval matching models suitable for new fields without manual labeling of corpus and user behavior data, and solves the problem of lack of training data in new fields search scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112579870B_ABST
    Figure CN112579870B_ABST
Patent Text Reader

Abstract

The present application discloses a training method, device, equipment and storage medium for a retrieval matching model, belonging to the field of machine learning. The method includes: obtaining a generation model, where the generation model is trained based on the first relevance training corpus in the existing domain; inputting the text to be retrieved in the target domain into the generation model to obtain the second relevance training corpus in the target domain, where the second relevance training corpus includes the corresponding relationship between the text to be retrieved and the query term; and inputting the second relevance training corpus into an initialized model for training to obtain a retrieval matching model adapted to the target domain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of machine learning, and particularly to a training method, device, equipment and storage medium for a retrieval matching model. Background Art

[0002] The essence of search is to satisfy the supply-demand matching relationship between users and information (such as merchant products). The retrieval matching model plays a fundamental and important role in the search process.

[0003] A large number of retrieval matching models are constructed and learned using neural networks. However, neural networks rely heavily on a large amount of manually labeled corpus for training. For example, the manually labeled corpus includes the correlation level between the document to be retrieved (doc) and the query. Generally, the more samples in the manually labeled corpus, the better the performance of the trained retrieval matching model.

[0004] When a new search field appears, due to the lack of manually labeled corpus for the new search field, a retrieval matching model for the new search field cannot be trained in time. Summary of the Invention

[0005] The present application provides a training method, device, equipment and storage medium for a retrieval matching model. The technical solution is as follows:

[0006] According to one aspect of the present application, a training method for a retrieval matching model is provided. The method includes:

[0007] Obtain a generation model, where the generation model is trained using a first correlation training corpus in an existing field;

[0008] Input the text to be retrieved in the target field into the generation model to obtain a second correlation training corpus pair in the target field, where the second correlation training corpus pair includes the corresponding relationship between the text to be retrieved and the query;

[0009] Input the second correlation training corpus pair into an initialization model for training to obtain a retrieval matching model adapted to the target field.

[0010] According to one aspect of the present application, a training device for a retrieval matching model is provided. The device includes:

[0011] An obtaining module, configured to obtain a generation model, where the generation model is trained using a first correlation training corpus in an existing field;

[0012] An input module for inputting the text to be retrieved in the target field into the generation model to obtain a second relevance training corpus pair in the target field, where the second relevance training corpus pair includes the correspondence between the text to be retrieved and the query term;

[0013] A training module for inputting the second relevance training corpus pair into an initialized model for training to obtain a retrieval matching model adapted to the target field.

[0014] According to another aspect of the present application, there is provided a computer device, which includes: a processor and a memory, where the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the training method of the retrieval matching model as described above.

[0015] According to another aspect of the present application, there is provided a computer-readable storage medium storing a computer program, and the computer program is loaded and executed by a processor to implement the training method of the retrieval matching model as described above.

[0016] According to another aspect of the present application, there is provided a computer program product storing a computer program, and the computer program is loaded and executed by the processor to implement the training method of the retrieval matching model as described above.

[0017] The beneficial effects brought by the technical solution provided by the embodiments of the present application at least include:

[0018] By training a generation model with the first relevance training corpus pair in the existing field and calling the generation model to process the text to be retrieved in the new field to obtain a second relevance training corpus pair in the new field, the problem that there is no relevance training corpus pair in the search scenario for the new field and the retrieval matching model cannot be trained is solved. It enables the retrieval matching model to be quickly deployed in the search scenario for the new field without the need for manual annotation of the corpus and user behavior data. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 Shows a structural block diagram of a search system provided by an exemplary embodiment of the present application;

[0021] Figure 2Shows a flowchart of a method for training a retrieval matching model provided by another exemplary embodiment of the present application;

[0022] Figure 3 Shows a schematic diagram of the training of a retrieval matching model provided by another exemplary embodiment of the present application;

[0023] Figure 4 Shows a flowchart of a method for training a retrieval matching model provided by an exemplary embodiment of the present application;

[0024] Figure 5 Shows a schematic diagram of the training of a generation model provided by another exemplary embodiment of the present application;

[0025] Figure 6 Shows a flowchart of a method for training a retrieval matching model provided by an exemplary embodiment of the present application;

[0026] Figure 7 Shows a model architecture diagram of a retrieval matching model provided by another exemplary embodiment of the present application;

[0027] Figure 8 Shows a model architecture diagram of a retrieval matching model provided by another exemplary embodiment of the present application;

[0028] Figure 9 Shows a block diagram of a training device for a retrieval matching model provided by an exemplary embodiment of the present application;

[0029] Figure 10 Shows a block diagram of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners

[0030] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0031] Figure 1 Shows a block diagram of a retrieval system 100 provided by an exemplary embodiment of the present application. The retrieval system 100 includes: a user terminal 120, a search server 140, and a development terminal 160.

[0032] The user terminal 120 is the terminal used by the user. The user terminal 120 may be at least one of a desktop computer, a laptop computer, a tablet computer, an e-book, an MP3, and an MP4. An application program or a web client is run on the user terminal 120. The application program or the web client provides a search service.

[0033] The search server 140 is a backend server that provides search services. A retrieval matching model is stored in the search server 140, and the retrieval matching model is a neural network-based model. The input of the retrieval matching model is the search term sent by the user terminal 120, and the output of the retrieval matching model is the search text (result). The retrieval matching model can be multiple models for different fields.

[0034] The development terminal 160 is a terminal used by developers. The development terminal 160 is used to train the retrieval matching model.

[0035] The retrieval matching model can be trained by a computer device, which can be the search server 140, or a development terminal 160 different from the search server 140, or other computer devices.

[0036] Figure 2 The flowchart of the training method of the retrieval matching model provided by an exemplary embodiment of the present application is shown. In this embodiment, the method is exemplified by being applied to Figure 1 the computer device shown. The method includes:

[0037] Step 202: Obtain a generation model;

[0038] The generation model is a neural network model with the ability to predict query terms for the input text. The generation model is trained based on the first relevance training corpus of the existing field.

[0039] In the present application, the field refers to the division of different types of search scenarios. Taking the takeaway scenario as an example, in the historical search scenario, most of the query terms of users are in the catering fields such as food, desserts, and beverages. However, with the change of users' cognition, more and more people will search for goods in non-catering fields, such as fresh fruits and vegetables, books, mobile phones, etc. Here, the catering field is the existing field or the original field, while the new flash sale categories such as books and mobile phones are the new fields.

[0040] Schematically, the existing field is a general field, or one or more used search fields. The first relevance training corpus pair includes: multiple first corpus pairs of the existing field, and each first corpus pair includes: search term (query) - text to be retrieved (doc), and a relevance level. Among them, the relevance level also has other names such as relevance score and relevance gear. For example, the relevance level includes: strongly relevant, weakly relevant, and irrelevant.

[0041] Step 204: Input the text to be retrieved in the target field into the generation model to obtain the second relevance training corpus pair in the target field, and the second relevance training corpus pair includes the corresponding relationship between the text to be retrieved and the query term;

[0042] The target domain is a new domain relative to the existing domain, or the target domain is a sub-domain relative to the existing domain. Taking the target domain being a new domain as an example, a new domain is a domain where there is text to be retrieved but no query terms, or only a small number of query terms exist.

[0043] As Figure 3 shown, the text to be retrieved in the new domain is input into the generation model 10 for prediction. The second relevant training corpus pair for the target domain is output by the generation model. The second relevant training corpus pair includes: multiple second corpus pairs for the target domain, and each second corpus pair includes: a search term (query) - text to be retrieved (doc), and a relevance level. Among them, the relevance level also has other names such as relevance score, relevance level, etc.

[0044] Step 206: Input the second relevant training corpus pair into the initialization model for training to obtain a retrieval matching model adapted to the target domain.

[0045] The initialization model can be a basic model that has not been trained yet, or a model after initializing the model parameters using a pre-trained language model.

[0046] As Figure 3 shown, each second corpus pair in the second relevant training corpus pair is used as a training sample and input into the initialization model 20 for training to obtain a retrieval matching model 30 adapted to the target domain.

[0047] In summary, the method provided in this embodiment trains a generation model through the first relevant training corpus pair in the existing domain, and calls the generation model to process the text to be retrieved in the new domain to obtain the second relevant training corpus pair for the new domain, thereby solving the problem that in the search scenario for the new domain, there is no relevant training corpus pair and the retrieval matching model cannot be trained. It enables the effect of quickly deploying the retrieval matching model in the search scenario for the new domain without the need for manual annotation of the corpus and user behavior data.

[0048] Figure 4 shows a flowchart of a method for training a retrieval matching model provided by another exemplary embodiment of the present application. This embodiment takes the method being applied to Figure 1 the computer device shown as an example. The method includes:

[0049] Phase 1: The training phase of the generation model;

[0050] Step 402: Obtain the first relevant training corpus pair in the existing domain. Each corpus pair in the first relevant training corpus pair includes: a sample text, a sample query term, and a sample relevance level. The relevance level is used to indicate the degree of relevance between the sample text and the sample query term;

[0051] The first relevance training corpus pair is constructed based on the existing domain and can be denoted as Corpus 1.

[0052] The first relevance training corpus pair includes: sample query words (query), sample texts (doc), and sample relevance levels (label). Among them, the sample texts vary in form according to the uses of specific retrieval systems. For example, in the traditional web search scenario, the sample texts are mainly medium and long articles. In the fields of e-commerce and food delivery, they are mainly stores and shops. Among them, stores can be represented by Points of Information (POI), and products can be represented by Standard Product Units (SPU). Schematically, in addition to stores and shops, other information can be additionally added to the text to be retrieved, and this embodiment does not limit this.

[0053] The relevance level includes more than two levels. In one example, the relevance level includes relevant and irrelevant. In another example, the relevance level includes: strongly relevant, weakly relevant, and irrelevant. The source methods of the relevance level include but are not limited to at least one of the following:

[0054] · Manual annotation method;

[0055] Relying on personnel such as outsourcers, R & D personnel, and product managers to manually annotate between the query words and the text to be retrieved.

[0056] · Unsupervised generation method;

[0057] Automatically generated relying on the user's historical click behavior and historical order placement behavior. For example, if the user searches for "Word A" and places an order for "Product B" among multiple search results, then Word A and Product B are established as a corpus pair with strong relevance.

[0058] When obtaining the annotation results in the existing domain, the cost is relatively low and most of the existing data has been accumulated.

[0059] Step 404: Segment the sample text to obtain text segments; segment the sample query words to obtain query segments;

[0060] The generation model is a neural network model capable of predicting query words for the input text. Exemplarily, the generation model selects an encoder-decoder model structure based on the currently widely used Bidirectional Encoder Representations from Transformers (BERT) model in the industry.

[0061] It is necessary to first perform word segmentation on the sample text to obtain text word segments. After each text word segment is vectorized, it is denoted as doc_tokens; perform word segmentation on the sample query term to obtain query word segments. After each query word segment is vectorized, it is denoted as query_tokens. That is, the smallest unit after word segmentation is denoted as a token.

[0062] During the word segmentation process, the word segmentation module of the BERT model can be used for word segmentation (or character segmentation). Suppose the generative model is the ALBERT (A Lite BERT for Self-supervised Learning of Language Representations) model for self-supervised learning of language representations, then the word segmentation module selects the word segmentation module in the ALBERT model.

[0063] Optionally, during the process of constructing the generative model using the encoder and decoder of the ALBERT model, the model parameters of the pre-trained language model obtained by training with general corpus are used to initialize the model parameters of the encoder and decoder of the generative model. Among them, the encoder and decoder share model parameters during the training process.

[0064] Step 406: Input the text word segments into the encoder to obtain an encoded output;

[0065] Exemplarily, Figure 5 The structural schematic diagram of the generative model 10 is shown. The generative model 10 includes an encoder 12 and a decoder 14. During the training process, the computer device sequentially inputs each text word segment X in the text to be retrieved into the encoder 12 to obtain an encoded output.

[0066] Optionally, at the i-th encoding moment, the encoder 12 encodes the i-th text word segment in the text to be retrieved to obtain the encoded output corresponding to the i-th encoding moment. After all the text word segments in the text to be retrieved are encoded, the encoded output of the text to be retrieved is obtained.

[0067] Step 408: Input the encoded output and the relevance level into the decoder to obtain predicted query word segments;

[0068] In order to reflect the influence of the relevance level on the text to be retrieved, the computer device inputs the encoded output and the relevance level into the decoder to obtain predicted query word segments.

[0069] Schematically, an attention matrix is also designed between the encoder and the decoder. The computer device embeds the sample correlation level to obtain a correlation level vector. The correlation level vector and the encoded output are subjected to attention weighting through the attention matrix to obtain a weighted vector; the weighted vector is input into the decoder to decode and obtain the predicted query segmentation.

[0070] In one example, the weighted vector is only input into the decoder at the first decoding moment. At each subsequent decoding moment, the predicted query segmentation output by the decoder at the historical decoding moment is input into the decoder to obtain the predicted query segmentation of the decoder at the next decoding moment.

[0071] In another example, after the weighted vector is input into the decoder at the first decoding moment, at each subsequent decoding moment, the weighted vector and the predicted query segmentation output by the decoder at the historical decoding moment are input into the decoder to obtain the predicted query segmentation of the decoder at the next decoding moment.

[0072] Schematically, as Figure 5 shown, at the first decoding moment, the weighted vector is input into the decoder 14 to obtain the predicted query segmentation Y1 corresponding to the first decoding moment;

[0073] At the second decoding moment, the weighted vector and the predicted query segmentation Y1 output by the decoder at the historical decoding moment are input into the decoder to obtain the predicted query segmentation of the decoder at the second decoding moment;

[0074] At the third decoding moment, the weighted vector and the predicted query segmentations Y1 to Y2 output by the decoder at the historical decoding moment are input into the decoder to obtain the predicted query segmentation Y3 of the decoder at the third decoding moment, and so on, which will not be elaborated here.

[0075] Step 410: Update the model parameters of the encoder and the decoder according to the error between the predicted query segmentation and the query segmentation.

[0076] Optionally, the error between the predicted query segmentation and the query segmentation is represented by the standard cross entropy.

[0077] Phase II: The training corpus generation phase for the new domain;

[0078] Step 412: Obtain the generation model;

[0079] The computer device obtains the already trained generation model.

[0080] Step 414: Input the text to be retrieved in the target domain into the generation model to obtain the second correlation training corpus pair in the target domain, and the second correlation training corpus pair includes the corresponding relationship between the text to be retrieved and the query word;

[0081] The target domain is a new domain relative to the existing domain, or the target domain is a sub-domain relative to the existing domain. Taking the target domain as a new domain as an example, the new domain is a domain where there is retrieved text to be retrieved, but there is no query term, or only a small number of query terms exist.

[0082] The computer device inputs the retrieved text of the target domain into the generation model, and the generation model predicts the query term and the relevance level, so as to obtain the second relevance training corpus pair of the target domain. The second relevance training corpus pair includes the corresponding relationship between the retrieved text and the query term.

[0083] In one example, since the retrieved text of the target domain is small, the quantity of the retrieved text needs to be enhanced. As Figure 6 shown, the following steps may be included:

[0084] Step 61: Train a second pre-trained language model using the retrieved text of the target domain;

[0085] The goal of a language model is to describe the probability of a word in a sentence. A language model is a model trained from multiple corpus information to "learn" the probability of a certain word in the corpus domain.

[0086] The domain knowledge under the new domain can be trained by the second pre-trained language model. Schematically, a language model represented by the BERT model is still selected for pre-training to obtain the pre-trained language model. The "pre-trained language model" here and the "generation model" are two different models, but the model architectures used can be the same or different.

[0087] During the process of training the second pre-trained language model, first use the tokenization module to tokenize the retrieved text to obtain text tokens. After each text token is vectorized, it is denoted as doc_tokens. It should be noted that the tokenization modules used by the second pre-trained language model and the generation model are the same or consistent.

[0088] Schematically, taking the open-source model checkpoint model trained using a general corpus as the base model, use the text tokens of the retrieved text under the new domain for pre-training to learn the domain knowledge under the new domain. Finally, the second pre-trained language model under the new domain is obtained.

[0089] The pre-trained language model can be trained based on the Masked Language Model Task (MLM). MLM means that during the training process, a certain word in the retrieved text will be randomly blocked, and the second pre-trained language model will predict the currently blocked word according to the context information of the word (similar to cloze test), and the original text order and structure will not be changed after the prediction.

[0090] Step 62: Input the text to be retrieved in the target domain into the second pre-trained language model to obtain the enhanced text to be retrieved.

[0091] Since the number of texts to be retrieved in the new domain is small, input the text to be retrieved in the target domain into the pre-trained language model to obtain the enhanced text to be retrieved.

[0092] In the schematic enhancement process, the computer device randomly masks the positions of words in the text to be retrieved in the target domain, predicts the positions of words through the pre-trained language model to obtain predicted words; substitute the predicted words into the positions of words to obtain the enhanced text to be retrieved.

[0093] Optionally, at least one word in the text to be retrieved in the target domain is masked each time. That is, one word or multiple words in the text to be retrieved in the target domain are masked each time. Where n is a preset value.

[0094] For example, if the text to be retrieved is "stewed beef with tomatoes", when masking "tomatoes", the text to be retrieved predicted by the pre-trained model is "stewed beef with tomatoes"; when masking "beef", the text to be retrieved predicted by the pre-trained model is "stewed tomatoes with brisket". When masking "tomatoes" and "stewed", the text to be retrieved predicted by the pre-trained model is "stewed beef with potatoes".

[0095] Determine the collection of the original text to be retrieved in the new domain and the text to be retrieved predicted by the pre-trained language model as the enhanced text to be retrieved. The text content of the enhanced text to be retrieved is more than the text content of the text to be retrieved.

[0096] Step 63: Input the enhanced text to be retrieved into the generation model to obtain the second relevance training corpus pair in the target domain.

[0097] The second relevance training corpus pair includes the correspondence between the text to be retrieved and the query word;

[0098] Input the text to be retrieved in the new domain into the generation model for prediction. The generation model outputs the second relevance training corpus pair in the target domain. The second relevance training corpus pair includes: multiple second corpus pairs in the target domain, and each second corpus pair includes: search word (query) - text to be retrieved (doc), and the relevance level. The relevance level is also called other names such as relevance score and relevance gear.

[0099] There are n relevance levels. For illustrative purposes, taking three relevance levels as an example, the relevance levels include: strong correlation, weak correlation, and no correlation. The label 0 is used to represent strong correlation, the label 1 is used to represent weak correlation, and the label 3 is used to represent no correlation. The text to be retrieved in each new field and the three labels are input into the generation model, and the generation model generates query words (queries) under different relevance levels. Finally, for each text to be retrieved, three "query-doc" pairs are generated, corresponding to the three labels respectively.

[0100] Phase 3: The training phase of the retrieval matching model;

[0101] Step 416: Input the second relevance training corpus pair into the initialization model for training to obtain a retrieval matching model adapted to the target field;

[0102] The initialization model can be a basic model that has not been trained yet, or a model after initializing the model parameters using a pre-trained language model.

[0103] Use each second corpus pair in the second relevance training corpus pair as a training sample, input it into the initialization model for training, and obtain a retrieval matching model adapted to the target field. The retrieval matching model is a neural network model that has the ability to output the corresponding text to be retrieved for the input query word. The input of the retrieval matching model is the query word, and the output is the text to be retrieved.

[0104] For illustrative purposes, the initialization model includes a second encoder. The model parameters of the second encoder in the initialization model can be initialized using the model parameters of the encoder in the second pre-trained language model. Among them, the second pre-trained language model is trained based on the text to be retrieved in the new field.

[0105] In one example, the initialization model uses a classic twin two-tower semantic matching model represented by the Deep Structured Semantic Models (DSSM). The characteristic of the DSSM model is that the query and doc are encoded in the same vector space, and the corresponding vector results can be retained offline, which helps to improve the online performance significantly.

[0106] Such as Figure 7 As shown, the DSSM model includes: an input layer 71, a representation layer 72, and a matching layer 73.

[0107] In the input layer 71, CLS represents the beginning and end of the first sentence, Tok1 represents the first word of the sentence, and Tokn represents the nth word of the sentence. SEP is used to separate two sentences. That is, the computer device uses the above-mentioned word segmentation module to perform word segmentation on the query and doc, and then inputs them into the representation layer 72 for encoding.

[0108] In the presentation layer 72, it includes two cascaded BERT encoders (the second encoder) and an average pooling layer. Both BERT encoders are initialized with the model parameters of the second pre-trained language model. One set of the BERT encoder and the average pooling layer corresponds to the input of the query terms, and the output is the first feature representation of the query terms; the other set of the BERT encoder and the average pooling layer corresponds to the input of the store name and the product name in the text to be queried, and the output is the second feature representation of the store name and the store name. Taking the query terms including N query tokens as an example, E [CLS] is the input representation of the CLS by the input layer 71, E1 is the input representation of the first query token Tok1 by the input layer 71, E N is the input representation of the Nth query token TokN by the input layer 71. C is the semantic representation vector of E [CLS] by the BERT encoder, T1 is the input representation of E1 by the BERT encoder, T N is the semantic representation vector of E N by the BERT encoder. And so on, which will not be elaborated here.

[0109] In the matching layer 73, the cosine similarity between the first feature representation and the second feature representation is calculated. The cosine similarity is used to determine whether two vectors point in the same direction. When two vectors have the same direction, the value of the cosine similarity is 1; when the included angle between two vectors is 90 degrees, the value of the cosine similarity is 0. Among them, the matching layer 73 is also called the softmax layer.

[0110] In actual model training, the model structure of the Deep Structured Semantic Models (DSSM) model can be simplified according to the hardware conditions and time requirements. For example, the model parameters (or called network weights) of the first n layers in the BERT encoder can be fixed, and only the last few layers are trained, etc.

[0111] In another example, the initialized model adopts an improved interactive deep semantic matching model. As Figure 8 shown, this improved interactive deep semantic matching model is divided into left and right network structures. The left network structure mainly includes: an input layer 81, an interaction layer 82, and a fully connected layer 83.

[0112] The input layer 81 includes two second encoders, and both second encoders are initialized with the model parameters of the second pre-trained language model. One of the second encoders is used to encode the text to be retrieved, and the output is the semantic representation vector doc-vec of the text to be retrieved. The other second encoder is used to encode the query terms, and the output is the semantic representation vector query-vec of the query terms. The two second encoders will share weights.

[0113] The interaction layer 82 includes an average pooling layer, a max pooling layer, and a normalization (Norm) layer. The semantic representation vectors of the query and the doc are used to calculate a similarity vector based on the vector results obtained through the max pooling layer and the average pooling layer respectively. There are various ways to calculate the similarity vector, such as cosine, jaccard, dot-product, etc. The normalization layer is used to perform normalization processing on the similarity vector.

[0114] The fully connected layer 83 is used to concatenate the two similarity vectors output by the interaction layer 82 and then input them into the upper fully connected layer.

[0115] The right network architecture is a multi-layer perceptron model (Multi-Layer Perception, MLP), and its specific structure will not be elaborated here. Some additional features will be used in the right network architecture, which depend on the work of feature engineering. Schematically, the additional features can include the literal text features of the query and the doc, such as the number of characters, the number of words, the number of co-occurring characters / words, the co-occurring positions, etc.; or, the additional features can include text similarity, such as BM25, Term Frequency-Inverse Document Frequency (TF-IDF), edit distance, etc.; or, the additional features can include vector similarity, such as the results of word vectors like BERT, word2vec, fast-text, etc.; or, the additional features include category similarity, such as text classification labels, merchant categories, product categories, etc., and features in other dimensions.

[0116] Using artificial feature engineering can improve the extensibility of the initialized model, lay a foundation for subsequent iterations, and effectively improve the prediction accuracy of the initialized model. Finally, the vectors output by the network results on both the left and right sides are concatenated together, and the final result is output after passing through the fully connected layer 83, the dense layer, and the output (softmax) layer.

[0117] Phase Four: The usage phase of the retrieval matching model;

[0118] Step 418: Use the trained retrieval matching model to provide retrieval services in the target domain.

[0119] Schematically, developers deploy the trained retrieval matching model to the search server, and the search server uses the trained retrieval matching model to provide retrieval services in the target domain. For example, if the new domain is the book domain, the search server uses the retrieval matching model to provide retrieval services in the book domain; or, if the new domain is the mobile phone domain, the search server uses the retrieval matching model to provide retrieval services in the mobile phone domain.

[0120] In summary, the method provided in this embodiment does not require manual annotation of sample training data in the new domain, and does not require prior accumulation of user behavior in the new domain. It can automatically generate relevant corpora for learning the retrieval matching model, and at the same time, by using the pre-trained language model, it can well adapt to the new domain.

[0121] In addition, the method provided in this embodiment has strong adaptability and excellent scalability. It is not only more handy when facing more new domains during the high-speed expansion of the business, but also can provide a large amount of candidate annotation data before the retrieval matching model is launched. After there are manual annotation data or a large amount of user behavior accumulation in the later stage, the retrieval matching model trained by this method can still be used for continuous learning, greatly improving the iteration efficiency of the retrieval matching model.

[0122] In addition, the method provided in this embodiment has high accuracy. On the one hand, the pre-trained language model itself contains domain knowledge. On the other hand, since the text used in training the retrieval matching model is the text to be retrieved in the new domain, the prediction accuracy of the retrieval matching model will be more accurate than the relevance obtained by traditional methods such as only using the original domain knowledge, or simple literal matching, or using semantic similarity, improving the overall user experience when searching in the new domain, and thus enhancing the user's trust in the brand.

[0123] Figure 9 The block diagram of the training device of the retrieval matching model provided by an exemplary embodiment of the present application is shown. The device includes:

[0124] An acquisition module 920, configured to acquire a generation model, where the generation model is trained according to the first relevance training corpus of the existing domain;

[0125] An input module 940, configured to input the text to be retrieved in the target domain into the generation model to obtain a second relevance training corpus pair in the target domain, where the second relevance training corpus pair includes the correspondence between the text to be retrieved and the query term;

[0126] A training module 960, configured to input the second relevance training corpus pair into an initialization model for training to obtain a retrieval matching model adapted to the target domain.

[0127] In an alternative design of the present application, the generation model includes a first encoder and a first decoder; the device further includes: a word segmentation module 980;

[0128] The obtaining module 920 is further configured to obtain a first relevance training corpus pair in the existing field, and each corpus pair in the first relevance training corpus pair includes: a sample text, a sample query term, and a sample relevance level, where the relevance level is used to indicate the degree of relevance between the sample text and the sample query term;

[0129] The word segmentation module 980 is further configured to perform word segmentation on the sample text to obtain text word segments; perform word segmentation on the sample query term to obtain query word segments;

[0130] The input module 940 is further configured to input the text word segments into the encoder to obtain an encoded output; input the encoded output and the relevance level into the decoder to obtain a predicted query word segment;

[0131] The training module 960 is configured to update the model parameters of the encoder and the decoder according to the error between the predicted query word segment and the query word segment.

[0132] In an alternative design of the present application, the input module 940 is further configured to perform embedding processing on the sample relevance level to obtain a relevance level vector; perform attention weighting on the relevance level vector and the encoded output through an attention matrix to obtain a weighted vector; input the weighted vector into the decoder to decode and obtain the predicted query word segment.

[0133] In an alternative design of the present application, the training module 960 is further configured to initialize the model parameters of the encoder and the decoder using the model parameters of a first pre-trained language model; the first pre-trained language model is trained using a general corpus;

[0134] Wherein, the encoder and the decoder share the model parameters during the training process.

[0135] In an alternative design of the present application, the input module 940 is further configured to input the text to be retrieved in the target field into a second pre-trained language model to obtain an enhanced text to be retrieved; input the enhanced text to be retrieved into the generation model to obtain a second relevance training corpus pair in the target field;

[0136] Wherein, the text content of the enhanced text to be retrieved is more than the text content of the text to be retrieved, and the second pre-trained language model is trained based on the text to be retrieved in the target field.

[0137] In an alternative design of the present application, the input module 940 is further configured to randomly mask the positions of words in the text to be retrieved in the target field, predict the positions of the words through the first pre-trained language model to obtain predicted words; and substitute the predicted words into the positions of the words to obtain the enhanced text to be retrieved.

[0138] In an alternative design of the present application, the initialization model includes a second encoder, and the training module 960 is further configured to initialize the model parameters of the second encoder in the initialization model by using the model parameters of the encoder in the second pre-trained language model;

[0139] wherein, the second pre-trained language model is trained based on the text to be retrieved in the target field.

[0140] Figure 10 FIG. shows the structural framework diagram of a computer device 1000 provided by an embodiment of the present application. Specifically: The computer device 1000 includes a central processing unit (CPU) 1001, a system memory 1004 including a random access memory (RAM) 1002 and a read-only memory (ROM) 1003, and a system bus 1005 connecting the system memory 1004 and the central processing unit 1001. The computer device 1000 further includes a basic input / output system (I / O system) 1006 for facilitating the transfer of information between various components within the computer, and a mass storage device 1007 for storing an operating system 1013, application programs 1014, and other program modules 1015.

[0141] The basic input / output system 1006 includes a display 1008 for displaying information and an input device 1009 such as a mouse, keyboard, etc. for user input of information. Wherein both the display 1008 and the input device 1009 are connected to the central processing unit 1001 through an input / output controller 1100 connected to the system bus 1005. The basic input / output system 1006 may further include an input / output controller 1010 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1010 also provides output to a display screen, printer, or other types of output devices.

[0142] The mass storage device 1007 is connected to the central processing unit 1001 through a mass storage controller (not shown) connected to the system bus 1005. The mass storage device 1007 and its associated computer-readable medium provide non-volatile storage for the computer device 1000. That is to say, the mass storage device 1007 may include a computer-readable medium (not shown) such as a hard disk or a CD-ROM drive.

[0143] Without loss of generality, the computer-readable medium may include a computer storage medium and a communication medium. The computer storage medium includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The computer storage medium includes RAM, ROM, EPROM, EEPROM, flash memory or other solid-state storage technologies, CD-ROM, DVD or other optical storage, magnetic tape cartridges, magnetic tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage medium is not limited to the above several types. The above-mentioned system memory 1004 and mass storage device 1007 can be collectively referred to as memory.

[0144] According to various embodiments of the present application, the computer device 1000 can also run on a remote computer on the network connected through a network such as the Internet. That is, the computer device 1000 can be connected to the network 1012 through the network interface unit 1011 connected to the system bus 1005, or rather, the network interface unit 1011 can also be used to connect to other types of networks or remote computer systems (not shown).

[0145] The memory further includes one or more programs, and the one or more programs are stored in the memory. The one or more programs include a method for training the retrieval matching model provided in the embodiments of the present application.

[0146] The present application also provides a computer-readable storage medium, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the method for training the retrieval matching model provided in the above method embodiments.

[0147] Optionally, the present application also provides a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the method for training the retrieval matching model described in the above aspects.

[0148] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0149] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disc, etc.

[0150] The above are only alternative embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.

Claims

1. A training method for a retrieval matching model, characterized in that The method includes: Obtaining a generation model, which is trained based on a first relevance training corpus in an existing domain; Inputting a text to be retrieved in the target domain into the generation model to obtain a second relevance training corpus in the target domain, where the second relevance training corpus includes the correspondence between the text to be retrieved and a query term; Inputting the second relevance training corpus into an initialized model for training to obtain a retrieval matching model adapted to the target domain; The generation model includes a first encoder and a first decoder, and the method further includes: Obtaining the first relevance training corpus in the existing domain, where each corpus pair in the first relevance training corpus includes: a sample text, a sample query term, and a sample relevance level, and the relevance level is used to indicate the degree of relevance between the sample text and the sample query term; Performing word segmentation on the sample text to obtain text word segments; performing word segmentation on the sample query term to obtain query word segments; Inputting the text word segments into the encoder to obtain an encoded output; Inputting the encoded output and the relevance level into the decoder to obtain a predicted query word segment; Updating the model parameters of the encoder and the decoder according to the error between the predicted query word segment and the query word segment.

2. The method according to claim 1, wherein The generation model further includes an attention matrix, and the step of inputting the encoded output and the relevance level into the decoder to obtain a predicted query word segment includes: Performing embedding processing on the sample relevance level to obtain a relevance level vector; Performing attention weighting on the relevance level vector and the encoded output through the attention matrix to obtain a weighted vector; Inputting the weighted vector into the decoder to decode and obtain the predicted query word segment.

3. The method according to claim 1, wherein The method further includes: Initializing the model parameters of the encoder and the decoder using the model parameters of a first pre-trained language model; the first pre-trained language model is trained using a general corpus; Wherein, the encoder and the decoder share the model parameters during training.

4. The method according to any one of claims 1 to 3, characterized in that, The step of inputting the text to be retrieved in the target domain into the generation model to obtain the second relevance training corpus in the target domain includes: Inputting the text to be retrieved in the target domain into a second pre-trained language model to obtain an enhanced text to be retrieved; Inputting the enhanced text to be retrieved into the generation model to obtain the second relevance training corpus in the target domain; Wherein, the text content of the enhanced text to be retrieved is more than the text content of the text to be retrieved, and the second pre-trained language model is trained based on the text to be retrieved in the target domain.

5. The method according to claim 4, characterized in that The step of inputting the text to be retrieved in the target domain into the first pre-trained language model to obtain an enhanced text to be retrieved includes: randomly masking the word positions in the text to be retrieved in the target domain, predicting the word positions through the first pre-trained language model to obtain a predicted word; Substituting the predicted word into the word position to obtain the enhanced text to be retrieved.

6. The method according to any one of claims 1 to 3, characterized in that The initialization model includes a second encoder, and the method further includes: Initializing the model parameters of the second encoder in the initialization model by using the model parameters of the encoder in the second pre-trained language model; wherein, the second pre-trained language model is trained based on the text to be retrieved in the target domain.

7. A training device for a retrieval matching model, characterized in that The device includes: An acquisition module, configured to acquire a generation model, the generation model is trained based on the first relevance training corpus in the existing domain, the generation model includes a first encoder and a first decoder, and the device further includes: Acquiring the first relevance training corpus pair in the existing domain, each corpus pair in the first relevance training corpus pair includes: a sample text, a sample query term, and a sample relevance level, and the relevance level is used to indicate the relevance degree between the sample text and the sample query term; Performing word segmentation on the sample text to obtain text word segments; performing word segmentation on the sample query term to obtain query word segments; Inputting the text word segments into the encoder to obtain an encoded output; Inputting the encoded output and the relevance level into the decoder to obtain a predicted query word segment; Updating the model parameters of the encoder and the decoder according to the error between the predicted query word segment and the query word segment; An input module, configured to input the text to be retrieved in the target domain into the generation model to obtain a second relevance training corpus pair in the target domain, and the second relevance training corpus pair includes the correspondence between the text to be retrieved and the query term; A training module, configured to input the second relevance training corpus pair into the initialization model for training to obtain a retrieval matching model adapted to the target domain.

8. A computer device, characterized in that, The computer device includes: a processor and a memory, the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the training method of the retrieval matching model according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the training method of the retrieval matching model according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Sequencing system based on machine learning

    CN103530321A

  • Probability prediction model training method, probability prediction method and device

    CN111782676A