Interactive Retrieval Method, Apparatus, Device, and Storage Medium Based on a Dual-Tower Model

By building a dual-tower model based on a pre-trained language model, combining multi-channel computing and attention mechanisms, the problem of difficulty in matching efficiency and effectiveness of the dual-tower model is solved, and efficient matching of complex long text and multi-channel data is achieved.

CN115309865BActive Publication Date: 2025-05-27PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210962907.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-11
Publication Date
2025-05-27
Estimated Expiration
2042-08-11

AI Technical Summary

Technical Problem

The prior art is difficult to take into account the characteristics of high matching efficiency of double tower models and good matching effects of interactive models, especially when processing complex long text matching and multi-channel and multi-domain data.

Method used

By building a dual-tower model based on a pre-trained language model, the dual-tower model is used to perform multi-channel calculation and splicing of the target sample set, and vector calculation and optimization are performed in combination with the attention mechanism to achieve interactive matching of search entries and samples.

Benefits of technology

It improves the matching efficiency and effect of the double tower model, enhances the matching accuracy of complex long text and multi-channel data, and meets the performance requirements of the industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115309865B_ABST
    Figure CN115309865B_ABST
Patent Text Reader

Abstract

The present invention relates to artificial intelligence technology, and discloses an interactive retrieval method, device, equipment and medium based on a dual-tower model. The method includes: obtaining a target retrieval entry and a target sample set according to each retrieval entry in the retrieval log and its corresponding sample set; respectively performing model calculations on the target samples in the target sample set and the target retrieval entry by using a pre-trained dual-tower model to obtain a fused sample vector and an attention vector; optimizing the dual-tower model into a target dual-tower model according to the similarity between the attention vector and the fused sample vector; offline, calculating a fused matching vector of data to be matched by using the target dual-tower model; online, calculating the attention vector of the retrieval entry to be retrieved by using the dual-tower model and performing average pooling to obtain an average attention vector; and obtaining a retrieval result according to the similarity between the fused matching vector and the average attention vector. The present invention can improve the efficiency and accuracy of retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an interactive retrieval method, device, electronic device and computer-readable storage medium based on a double-tower model. Background Art

[0002] Information retrieval is an important area in the field of natural language processing (NLP). Information retrieval mainly involves the storage, indexing and retrieval of massive unstructured or semi-structured data. Its purpose is to help users efficiently obtain the information they want from massive data. Usually, in industrial application scenarios, it is necessary to screen a certain number of candidate subsets that meet user needs from billions of massive candidate data based on user search terms, so not only the effect of the solution is required, but also the efficiency is very strict.

[0003] At present, with the rapid development of deep neural networks, many deep semantic matching schemes have emerged in the field of information retrieval. These schemes can be generally divided into two categories, namely representational and interactive. The representational scheme is generally based on two structurally similar semantic models to encode the search terms and candidate data respectively, and then recall the candidate subset according to the similarity of the encoded vectors, which is also called the double tower model. The advantage is that a large number of candidate vectors can be calculated offline in advance, and online only needs to encode the vector of the search term and calculate the similarity with the pre-calculated candidate vector; this scheme is fast, but due to the lack of interaction between the search terms and candidate data in the training stage, the model does not learn enough about their relevance, and it is often not effective for complex long text matching. The interactive model splices the search terms and candidate data in the initial stage, and inputs them as a whole into a more complex neural network structure, which can learn the deeper relevance between the search terms and candidate data; this scheme can often achieve higher semantic matching accuracy; but because the interactive model needs to calculate a large number of spliced ​​vectors online, it is difficult to meet the performance requirements of the industry.

[0004] In addition, if the traditional semantic model is used to encode the candidate data, the candidate data needs to be integrated as a whole as input; in many application scenarios in the industry, candidate data often has multi-channel and multi-domain characteristics, rather than a simple long text; for example, in the Taobao product search scenario, each product has text information such as title, advertisement, subtitle, label, etc., and the weight of these text information in the search process is also different. If all text fields are modeled separately, the entire search solution will become complicated and inefficient.

[0005] In summary, the existing technology has the problem of being difficult to take into account both the high efficiency of the dual-tower model matching and the good effect of the interactive model matching. Summary of the invention

[0006] The present invention provides an interactive retrieval method, device, electronic device and computer-readable storage medium based on a dual tower model, and its main purpose is to solve the problem of difficult to balance the two characteristics of high matching efficiency of the dual tower model and good matching effect of the interactive model.

[0007] To achieve the above object, an interactive retrieval method based on a dual tower model provided by the present invention includes:

[0008] Construct a dual tower model according to a pre-trained standard language model, obtain retrieval logs, construct a corresponding sample set according to each retrieval entry in the retrieval logs, select a target retrieval entry from the retrieval entries, and select samples that meet preset conditions from the sample set corresponding to the target retrieval entry as the target sample set;

[0009] Perform multi-channel calculations on each target sample in the target sample set using the dual tower model to obtain target sample vectors, and perform vector calculations on the target retrieval entry using the dual tower model to obtain target retrieval vectors;

[0010] Perform splicing and network calculations on the target sample vectors to obtain fused sample vectors, initialize multiple parameter vectors, and perform attention calculations according to the parameter vectors and the target retrieval vectors to obtain attention vectors;

[0011] Calculate the similarity between the attention vector and the fused sample vector, determine the matching degree between the target sample and the target retrieval entry according to the similarity, and optimize the dual tower model according to the matching degree between the target sample set and the target retrieval entry to obtain a target dual tower model;

[0012] When a preset retrieval server is offline, obtain preset data to be matched, and calculate a fused matching vector of the data to be matched using the target dual tower model;

[0013] When the retrieval server is online, obtain a retrieval entry to be retrieved, calculate an attention vector of the retrieval entry to be retrieved using the dual tower model, and perform average pooling on the attention vector of the retrieval entry to be retrieved to obtain an average attention vector;

[0014] Select the data to be matched corresponding to the retrieval entry to be retrieved according to the similarity between the fused matching vector and the average attention vector, and use the data to be matched corresponding to the retrieval entry to be retrieved as the retrieval result.

[0015] Optionally, the constructing a dual tower model according to a pre-trained standard language model includes:

[0016] Obtain general training corpus data, and train a preset language model in a horizontal domain according to the general training corpus data to obtain a preliminary language model;

[0017] Obtain vertical training corpus data, and train the preliminary language model according to the vertical training corpus data to obtain the standard language model;

[0018] Construct the two-tower model according to the standard language model, a preset attention mechanism, and a network calculation layer, where the two-tower model includes a retrieval end model and a matching end model.

[0019] Optionally, the multi-channel calculation of each target sample in the target sample set by using the two-tower model to obtain a target sample vector includes:

[0020] Perform multi-channel recognition on the target sample to obtain multi-channel data corresponding to the target sample;

[0021] Use the matching end model in the two-tower model to perform vector calculation on the multi-channel data to obtain a language vector for each channel, and use the language vector as the target sample vector corresponding to the target sample.

[0022] Optionally, the splicing and network calculation of the target sample vector to obtain a fused sample vector includes:

[0023] Splice the target sample vectors, and perform full connection and activation on the spliced language vectors to obtain the channel weights corresponding to each channel of the target sample;

[0024] Perform weighted summation according to the channel weights of each channel and the target sample vector to obtain the fused sample vector corresponding to the target sample.

[0025] Optionally, the performing weighted summation according to the channel weights of each channel and the target sample vector to obtain the fused sample vector corresponding to the target sample includes:

[0026] Perform weighted summation on the channel weights of each channel and the target sample vector through the following formula:

[0027]

[0028] where f(d) is the fused sample vector corresponding to the target sample d; N is the total number of channels corresponding to the target sample; w i is the channel weight corresponding to the i-th channel of the target sample; is is the target sample vector corresponding to the i-th channel of the target sample.

[0029] Optionally, calculating an attention vector based on the parameter vector and the target retrieval vector includes:

[0030] Generating a set of vector sequences according to the parameter vector and the retrieval vector, and successively selecting a target vector sequence from the set of target vector sequences;

[0031] Performing a first weight calculation on the target vector sequence to obtain an updated sequence corresponding to the target vector sequence;

[0032] Performing a second weight calculation on the updated sequence to obtain a plurality of representation vectors corresponding to the updated sequence;

[0033] Performing an attention operation on the representation vectors corresponding to all updated sequences to obtain an initial attention vector;

[0034] Performing an activation calculation on the initial attention vector using a preset activation function, and performing a dot product summation on the result of the activation calculation and the representation vectors corresponding to the updated sequence to obtain an attention vector.

[0035] Optionally, determining the matching degree between the target sample and the target retrieval entry according to the similarity includes:

[0036] Determining the matching degree between the target sample and the target retrieval entry through the following formula:

[0037]

[0038] where h(q, d) is the matching degree between the target sample d and the target retrieval entry q; (P j ·f(d)) similarity is the similarity between the j-th attention vector P j in the target retrieval entry and the fused sample vector f(d) corresponding to the target sample d.

[0039] To solve the above problems, the present invention further provides an interactive retrieval device based on a two-tower model, and the device includes:

[0040] A target sample set generation module, configured to obtain a retrieval log, construct a corresponding sample set according to each retrieval entry in the retrieval log, select a target retrieval entry from the retrieval entries, and select samples meeting preset conditions from the sample set corresponding to the target retrieval entry as a target sample set;

[0041] A dual - tower model calculation module, which is used to construct a dual - tower model according to a pre - trained standard language model, perform multi - channel calculations on each target sample in the target sample set by using the dual - tower model to obtain target sample vectors, and perform vector calculations on the target retrieval entry by using the dual - tower model to obtain a target retrieval vector;

[0042] A target dual - tower model generation module, which is used to splice and perform network calculations on the target sample vectors to obtain fused sample vectors, initialize multiple parameter vectors, perform attention calculations according to the parameter vectors and the target retrieval vector to obtain an attention vector; calculate the similarity between the attention vector and the fused sample vector, determine the matching degree between the target sample and the target retrieval entry according to the similarity, and optimize the dual - tower model according to the matching degree between the target sample set and the target retrieval entry to obtain a target dual - tower model;

[0043] An offline calculation module, which is used to obtain preset data to be matched when a preset retrieval server is offline, and calculate the fused matching vector of the data to be matched by using the target dual - tower model;

[0044] An online calculation module, which is used to obtain a retrieval entry to be retrieved when the retrieval server is online, calculate the attention vector of the retrieval entry to be retrieved by using the dual - tower model, perform average pooling on the attention vector of the retrieval entry to be retrieved to obtain an average attention vector; select the data to be matched corresponding to the retrieval entry to be retrieved according to the similarity between the fused matching vector and the average attention vector, and use the data to be matched corresponding to the retrieval entry to be retrieved as the retrieval result.

[0045] To solve the above problems, the present invention also provides an electronic device, which includes:

[0046] At least one processor; and,

[0047] A memory communicatively connected to the at least one processor; wherein,

[0048] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above - mentioned interactive retrieval method based on a dual - tower model.

[0049] To solve the above problems, the present invention also provides a computer - readable storage medium, in which at least one computer program is stored, and the at least one computer program is executed by a processor in an electronic device to implement the above - mentioned interactive retrieval method based on a dual - tower model.

[0050] In the embodiments of the present invention, each retrieval entry and its corresponding sample set are constructed by retrieving logs, and a two-tower model is constructed according to a language model. Then, the similarity between the sample set and the vectors corresponding to the retrieval entries is calculated according to the two-tower model. According to the similarity results, the two-tower model can be optimized, and the interaction between the retrieval entries and the samples in the training stage is realized, solving the problem of insufficient learning of the correlation between the retrieval entries and the samples by the two-tower model, and improving the matching accuracy of the subsequent use of the two-tower model. By initializing multiple parameter vectors and calculating the attention of the retrieval vectors according to the parameter vectors, the vector representation of the retrieval entries is improved. When the retrieval server is online, by performing average pooling on the attention vectors of the retrieval entries to be retrieved, certain vector representations are retained and the online matching speed of the model is improved. The two-tower model is used to perform multi-channel calculation, splicing, and network calculation on the target samples respectively to obtain fused sample vectors, improving the accuracy of generating sample vectors for samples with multi-channel and multi-domain characteristics. Therefore, the interactive retrieval method, device, electronic device, and computer-readable storage medium based on the two-tower model proposed by the present invention can solve the problem of being difficult to balance the two characteristics of high matching efficiency of the two-tower model and good matching effect of the interactive model. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 FIG. is a schematic flow chart of an interactive retrieval method based on a two-tower model provided by an embodiment of the present invention;

[0052] Figure 2 FIG. is a schematic flow chart of constructing a two-tower model according to a pre-trained standard language model provided by an embodiment of the present invention;

[0053] Figure 3 FIG. is a schematic flow chart of splicing and network calculation on the target sample vectors provided by an embodiment of the present invention;

[0054] Figure 4 FIG. is a functional module diagram of an interactive retrieval device based on a two-tower model provided by an embodiment of the present invention;

[0055] Figure 5 FIG. is a schematic structural diagram of an electronic device for implementing the interactive retrieval method based on the two-tower model provided by an embodiment of the present invention.

[0056] The implementation, functional features, and advantages of the objectives of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0058] The embodiments of the present application provide an interactive retrieval method based on a dual - tower model. The execution subject of the interactive retrieval method based on the dual - tower model includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiments of the present application. In other words, the interactive retrieval method based on the dual - tower model can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0059] Refer to Figure 1 As shown, it is a schematic flowchart of the interactive retrieval method based on the dual - tower model provided by an embodiment of the present invention. In this embodiment, the interactive retrieval method based on the dual - tower model includes:

[0060] S1. Construct a dual - tower model according to a pre - trained standard language model, obtain retrieval logs, construct a corresponding sample set according to each retrieval entry in the retrieval logs, select a target retrieval entry from the retrieval entries, and select samples that meet preset conditions from the sample set corresponding to the target retrieval entry as the target sample set.

[0061] In the embodiments of the present invention, the dual - tower model can be composed of two Transformer - based language models, an attention mechanism, a neural network layer, etc.

[0062] Please refer to Figure 2 As shown, in the embodiments of the present invention, the constructing of the dual - tower model according to the pre - trained standard language model includes:

[0063] S21. Obtain general training corpus data, perform horizontal - domain training on a preset language model according to the general training corpus data to obtain a preliminary language model;

[0064] S22. Obtain vertical training corpus data, perform training on the preliminary language model according to the vertical training corpus data to obtain the standard language model;

[0065] S23. Construct the dual - tower model according to the standard language model, a preset attention mechanism, and a network calculation layer, where the dual - tower model includes a retrieval - end model and a matching - end model.

[0066] In the embodiments of the present invention, the general training corpus data is based on the generally publicly available general corpus, and a Transformer-based language model is pre-trained in a semi-supervised or unsupervised manner. This model has the general semantic expression and information extraction capabilities in the general field; the vertical training corpus data is based on the vertical domain corpus in the retrieval scenario, and the Transformer model trained based on the general training corpus data can also be fine-tuned in a semi-supervised or unsupervised manner, so that the model has the language representation ability containing specific domain information.

[0067] In the embodiments of the present invention, the two-tower model includes a retrieval end model and a matching end model. The retrieval end model and the matching end model can be constituted by the standard language model, and then attention calculation is performed on the output of the retrieval end model by using the attention mechanism, and network calculation is performed on the output of the matching end model by using pooling, fully connected and activation; further, the retrieval result of the two-tower model can be output by means of similarity calculation of the result of the attention calculation and the result of the network calculation.

[0068] In the embodiments of the present invention, the retrieval log refers to the log generated by the user during the historical retrieval process according to the retrieval entry and the subsequent click or browsing record. Among them, the data selected and clicked by the user according to the retrieval entry is the positive sample; in the data generated by the retrieval entry, the data other than the data browsed by the user is used as the negative sample; or the data generated by all retrieval entries other than the retrieval entry is used as the negative sample.

[0069] For example, assume that there are retrieval entries A and B in the retrieval log. Among them, retrieval entry A generates data x (browsed data) and data y (unbrowsed data), and retrieval entry B generates data u (browsed data) and data v (unbrowsed data). Therefore, for retrieval entry A, there are positive samples: data x, and negative samples: data y, data u, data v; for retrieval entry B, there are positive samples: data u, and negative samples: data x, data y, data v.

[0070] In the embodiments of the present invention, selecting samples that meet the preset conditions from the sample set corresponding to the target retrieval entry can be taking the positive samples and all negative samples corresponding to the target retrieval entry as the target sample set. Further, the positive samples corresponding to the target retrieval entry and the negative samples second only to the positive samples can be taken as the target sample set; among them, the negative samples second only to the positive samples can be obtained by calculating the similarity of the target retrieval entry and its corresponding positive samples and all negative samples by the two-tower model, and the negative sample with the similarity calculation result closest to the similarity calculation result corresponding to the positive sample is selected from the negative samples as the negative sample second only to the positive sample.

[0071] S2. Use the two-tower model to perform multi-channel calculations on each target sample in the target sample set to obtain target sample vectors, and use the two-tower model to perform vector calculations on the target retrieval entry to obtain a target retrieval vector.

[0072] In the embodiment of the present invention, the step of using the two-tower model to perform multi-channel calculations on each target sample in the target sample set to obtain target sample vectors includes:

[0073] Perform multi-channel recognition on the target sample to obtain multi-channel data corresponding to the target sample;

[0074] Use the matching end model in the two-tower model to perform vector calculations on the multi-channel data to obtain language vectors for each channel, and use the language vectors as the target sample vectors corresponding to the target sample.

[0075] In the embodiment of the present invention, assume that the target sample is an online commodity, and each commodity has text information such as a title, advertisement, subtitle, tags, etc., which constitute multi-channel data. The weights of these text information in the subsequent retrieval process are also different. Therefore, multi-channel recognition can be performed on the target sample to obtain language vectors for each channel. Subsequently, by performing vector calculations on the language vectors for each channel, the weights for each channel can be obtained, making the finally generated fused sample vector more accurate.

[0076] In the embodiment of the present invention, the retrieval end model (a Transformer-based language model) in the two-tower model performs vector calculations on the target retrieval entry to generate a target retrieval vector.

[0077] S3. Concatenate the target sample vectors and perform network calculations to obtain a fused sample vector, and initialize multiple parameter vectors. Perform attention calculations based on the parameter vectors and the target retrieval vector to obtain an attention vector.

[0078] Please refer to Figure 3 As shown, in the embodiment of the present invention, the step of concatenating the target sample vectors and performing network calculations to obtain a fused sample vector includes:

[0079] S31. Concatenate the target sample vectors, and perform full connection and activation on the concatenated language vectors to obtain the channel weights corresponding to each channel of the target sample;

[0080] S32. Perform weighted summation based on the channel weights for each channel and the target sample vectors to obtain the fused sample vector corresponding to the target sample.

[0081] Specifically, in the embodiments of the present invention, the channel weights of each channel and the target sample vector can be weighted and summed by the following formula:

[0082]

[0083] where h(q, d) is the matching degree between the target sample d and the target retrieval entry q; (P j ·f(d)) similarity is the similarity between the j-th attention vector P j in the target retrieval entry and the fused sample vector f(d) corresponding to the target sample d.

[0084] In the embodiments of the present invention, the calculating the attention vector according to the parameter vector and the target retrieval vector includes:

[0085] generating a vector sequence set according to the parameter vector and the retrieval vector, and successively selecting a target vector sequence from the target vector sequence set;

[0086] performing a first weight calculation on the target vector sequence to obtain an updated sequence corresponding to the target vector sequence;

[0087] performing a second weight calculation on the updated sequence to obtain a plurality of representation vectors corresponding to the updated sequence;

[0088] performing an attention operation on the representation vectors corresponding to all updated sequences to obtain an initial attention vector;

[0089] performing activation calculation on the initial attention vector by using a preset activation function, and performing dot product summation on the result of the activation calculation and the representation vectors corresponding to the updated sequence to obtain the attention vector.

[0090] In the embodiments of the present invention, the attention operation is the operation of attention, and the operation of attention can be implemented by scaled dot product or the like.

[0091] In the embodiments of the present invention, a first weight calculation can be performed by using the weight coefficients defined by the network to obtain a new vector sequence; a second weight calculation can be performed on the updated sequence by using three different weight matrices defined by the network to obtain three representation vectors corresponding to the updated sequence.

[0092] S4. Calculate the similarity between the attention vector and the fused sample vector, determine the matching degree between the target sample and the target retrieval entry according to the similarity, and optimize the two-tower model according to the matching degree between the target sample set and the target retrieval entry to obtain a target two-tower model.

[0093] In the embodiments of the present invention, the similarity between the attention vector and the fused sample vector can be calculated using Euclidean distance, cosine similarity, Pearson coefficient, etc.

[0094] In the embodiments of the present invention, the matching degree between the target sample and the target retrieval entry can be determined by the following formula:

[0095]

[0096] where h(q, d) is the matching degree between the target sample d and the target retrieval entry q; P j is the j-th attention vector in the target retrieval entry; f(d) is the fused sample vector corresponding to the target sample d; and similarity is the similarity calculation.

[0097] In the embodiments of the present invention, if the matching degree between the positive sample in the target sample set and the target retrieval entry is less than the matching degree between the negative sample in the target sample set and the target retrieval entry, it indicates that there is a deviation in model training. The model can be updated by adding a penalty mechanism to update the network parameters in the twin tower model, thereby optimizing the model.

[0098] Further, in the embodiments of the present invention, after updating the network parameters in the twin tower model, the updated model can be used to recalculate the matching degree between the target sample set and the target retrieval entry, and it is determined whether to update the twin tower model again according to the calculated result meeting the preset matching degree condition. For example, when the matching degree between the positive sample in the target sample set and the target retrieval entry is greater than the matching degree between the negative sample in the target sample set and the target retrieval entry, the update of the twin tower model can be stopped.

[0099] In another alternative embodiment of the present invention, when there are multiple negative samples in the target sample set, multi-classification task training is performed on the samples in the target sample set, that is, positive samples and multiple negative samples are found from multiple samples in the target sample set; when there is only one negative sample in the target sample set (the matching degree calculation result obtained through multi-classification task training), this negative sample is of a certain difficulty in discrimination, then advanced binary classification task training is performed on the samples in the target sample set, that is, positive samples and multiple negative samples are found from multiple samples in the target sample set, that is, positive samples are found from two samples. Through this model training method, the model can have high accuracy and improve the ability to distinguish difficult negative samples.

[0100] S5. When the preset retrieval server is offline, obtain the preset data to be matched, and calculate the fused matching vector of the data to be matched using the target twin tower model.

[0101] In an embodiment of the present invention, the data to be matched is data that needs to perform multi-channel calculations and network fusion calculations, and is used to match with the entries to be retrieved.

[0102] In an embodiment of the present invention, by performing offline calculations on the data to be matched offline, the processing pressure for calculating the data to be matched online is reduced, and the efficiency of entry retrieval is improved.

[0103] In an embodiment of the present invention, the process of calculating the fusion matching vector of the data to be matched using the target dual tower model is similar to the process of performing multi-channel calculations on each target sample in the target sample set using the dual tower model in step S2 above to obtain a target sample vector, and the process of splicing and network calculation on the target sample vector in S3 to obtain a fusion sample vector, and will not be elaborated here.

[0104] S6. When the retrieval server is online, obtain the entry to be retrieved, calculate the attention vector of the entry to be retrieved using the dual tower model, and perform average pooling on the attention vector of the entry to be retrieved to obtain an average attention vector.

[0105] In an embodiment of the present invention, the average pooling of the attention vector of the entry to be retrieved can be performed using the following formula, including:

[0106]

[0107] where, is the average attention vector; P k is the k-th attention vector of the entry to be retrieved; M is the total number of attention vectors.

[0108] In an embodiment of the present invention, the average attention vector is generated based on multiple attentions, expanding the information content contained in the vector representation, and increasing the vector representation for searching the entry to be retrieved.

[0109] S7. Select the data to be matched corresponding to the entry to be retrieved according to the similarity between the fusion matching vector and the average attention vector, and use the data to be matched corresponding to the entry to be retrieved as the retrieval result.

[0110] In an embodiment of the present invention, the fusion matching vector corresponding to the maximum similarity value between the fusion matching vector and the average attention vector can be selected, and the data to be matched corresponding to this fusion matching vector is used as the retrieval result.

[0111] In the embodiments of the present invention, by directly calculating the similarity between the fusion matching vector and the average attention vector, the efficiency of vector matching is improved; and by pre-calculating the fusion matching vectors of the data to be matched offline, the time for online calculation of the fusion matching vectors is reduced, and the retrieval efficiency is improved.

[0112] In the embodiments of the present invention, each retrieval entry and its corresponding sample set are constructed through retrieval logs, and a two-tower model is constructed according to a language model. Then, the similarity between the vectors corresponding to the sample set and the retrieval entry is calculated according to the two-tower model. According to the similarity result, the two-tower model can be optimized, and the interaction between the retrieval entry and the sample in the training stage is realized, solving the problem of insufficient learning of the relevance between the retrieval entry and the sample by the two-tower model, and improving the matching accuracy of the subsequent use of the two-tower model; by initializing multiple parameter vectors and calculating the attention of the retrieval vector according to the parameter vectors, the vector representation of the retrieval entry is improved. When the retrieval server is online, by performing average pooling on the attention vectors of the retrieval entries to be retrieved, certain vector representations are retained and the online matching speed of the model is improved; the two-tower model is used to perform multi-channel calculation, splicing, and network calculation on the target sample respectively to obtain a fused sample vector, improving the accuracy of generating sample vectors for samples with multi-channel and multi-domain characteristics. Therefore, the interactive retrieval method based on the two-tower model proposed by the present invention can solve the problem of being difficult to balance the two characteristics of high matching efficiency of the two-tower model and good matching effect of the interactive model.

[0113] As Figure 4 shown, it is a functional module diagram of an interactive retrieval device based on a two-tower model provided by an embodiment of the present invention.

[0114] The interactive retrieval device 100 based on the two-tower model of the present invention can be installed in an electronic device. According to the functions realized, the interactive retrieval device 100 based on the two-tower model can include a target sample set generation module 101, a two-tower model calculation module 102, a target two-tower model generation module 103, an offline calculation module 104, and an online calculation module 105. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0115] In this embodiment, the functions of each module / unit are as follows:

[0116] The target sample set generation module 101 is configured to obtain retrieval logs, construct a corresponding sample set according to each retrieval entry in the retrieval logs, select a target retrieval entry from the retrieval entries, and select samples that meet preset conditions from the sample set corresponding to the target retrieval entry as the target sample set;

[0117] The dual - tower model calculation module 102 is configured to construct a dual - tower model according to a pre - trained standard language model, perform multi - channel calculations on each target sample in the target sample set by using the dual - tower model to obtain target sample vectors, and perform vector calculations on the target retrieval entry by using the dual - tower model to obtain a target retrieval vector;

[0118] The target dual - tower model generation module 103 is configured to splice and perform network calculations on the target sample vectors to obtain fused sample vectors, initialize multiple parameter vectors, perform attention calculations according to the parameter vectors and the target retrieval vector to obtain an attention vector; calculate the similarity between the attention vector and the fused sample vector, determine the matching degree between the target sample and the target retrieval entry according to the similarity, and optimize the dual - tower model according to the matching degree between the target sample set and the target retrieval entry to obtain a target dual - tower model;

[0119] The offline calculation module 104 is configured to obtain preset data to be matched when a preset retrieval server is offline, and calculate the fused matching vector of the data to be matched by using the target dual - tower model;

[0120] The online calculation module 105 is configured to obtain a retrieval entry to be retrieved when the retrieval server is online, calculate the attention vector of the retrieval entry to be retrieved by using the dual - tower model, perform average pooling on the attention vector of the retrieval entry to be retrieved to obtain an average attention vector; select the data to be matched corresponding to the retrieval entry to be retrieved according to the similarity between the fused matching vector and the average attention vector, and use the data to be matched corresponding to the retrieval entry to be retrieved as the retrieval result.

[0121] Specifically, each module in the interactive retrieval device 100 based on the dual - tower model in the embodiments of the present invention adopts the same technical means as the interactive retrieval method based on the dual - tower model described in the accompanying drawings when in use, and can produce the same technical effects, which will not be elaborated here.

[0122] As Figure 5 shown, it is a schematic structural diagram of an electronic device for implementing an interactive retrieval method based on a dual - tower model provided by an embodiment of the present invention.

[0123] The electronic device 1 may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may further include a computer program stored in the memory 11 and executable on the processor 10, such as an interactive retrieval program based on a dual - tower model.

[0124] Among them, in some embodiments, the processor 10 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and circuits, and by running or executing programs or modules stored in the memory 11 (such as executing an interactive retrieval program based on a dual-tower model, etc.), and calling the data stored in the memory 11, to perform various functions of the electronic device and process data.

[0125] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device, such as the mobile hard disk of the electronic device. In some other embodiments, the memory 11 may also be an external storage device of the electronic device, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device. The memory 11 can not only be used to store application software installed on the electronic device and various types of data, such as the code of an interactive retrieval program based on a dual-tower model, etc., but can also be used to temporarily store data that has been output or will be output.

[0126] The communication bus 12 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable connection communication between the memory 11 and at least one processor 10, etc.

[0127] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device and to display a visual user interface.

[0128] Figure 5 Only the electronic device with components is shown, and those skilled in the art can understand that Figure 5 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0129] For example, although not shown, the electronic device may further include a power source (such as a battery) for powering each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or an inverter, and a power status indicator. The electronic device may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0130] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0131] The interactive retrieval program based on the dual-tower model stored in the memory 11 in the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can implement:

[0132] Construct a dual-tower model according to a pre-trained standard language model, obtain retrieval logs, construct a corresponding sample set according to each retrieval entry in the retrieval logs, select a target retrieval entry from the retrieval entries, and select samples that meet preset conditions from the sample set corresponding to the target retrieval entry as the target sample set;

[0133] Perform multi-channel calculations on each target sample in the target sample set using the twin tower model to obtain target sample vectors, and perform vector calculations on the target retrieval entry using the twin tower model to obtain target retrieval vectors;

[0134] Perform splicing and network calculations on the target sample vectors to obtain fused sample vectors, and initialize multiple parameter vectors. Perform attention calculations based on the parameter vectors and the target retrieval vectors to obtain attention vectors;

[0135] Calculate the similarity between the attention vector and the fused sample vector, determine the matching degree between the target sample and the target retrieval entry according to the similarity, and optimize the twin tower model according to the matching degree between the target sample set and the target retrieval entry to obtain a target twin tower model;

[0136] When the preset retrieval server is offline, obtain the preset data to be matched, and calculate the fused matching vector of the data to be matched using the target twin tower model;

[0137] When the retrieval server is online, obtain the retrieval entry to be retrieved, calculate the attention vector of the retrieval entry to be retrieved using the twin tower model, and perform average pooling on the attention vector of the retrieval entry to be retrieved to obtain an average attention vector;

[0138] Select the data to be matched corresponding to the retrieval entry to be retrieved according to the similarity between the fused matching vector and the average attention vector, and use the data to be matched corresponding to the retrieval entry to be retrieved as the retrieval result.

[0139] Specifically, the specific implementation method of the above instructions by the processor 10 can refer to the description of the relevant steps in the corresponding embodiments of the accompanying drawings, which will not be elaborated here.

[0140] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory).

[0141] The present invention also provides a computer-readable storage medium, and the readable storage medium stores a computer program. When the computer program is executed by the processor of the electronic device, it can implement:

[0142] Construct a two - tower model based on a pre - trained standard language model, obtain retrieval logs, construct a corresponding sample set according to each retrieval entry in the retrieval logs, select a target retrieval entry from the retrieval entries, and select samples that meet the preset conditions from the sample set corresponding to the target retrieval entry as the target sample set;

[0143] Perform multi - channel calculations on each target sample in the target sample set using the two - tower model to obtain target sample vectors, and perform vector calculations on the target retrieval entry using the two - tower model to obtain target retrieval vectors;

[0144] Perform splicing and network calculations on the target sample vectors to obtain fused sample vectors, and initialize multiple parameter vectors. Perform attention calculations based on the parameter vectors and the target retrieval vectors to obtain attention vectors;

[0145] Calculate the similarity between the attention vector and the fused sample vector, determine the matching degree between the target sample and the target retrieval entry according to the similarity, and optimize the two - tower model according to the matching degree between the target sample set and the target retrieval entry to obtain a target two - tower model;

[0146] When the preset retrieval server is offline, obtain the preset data to be matched, and calculate the fused matching vector of the data to be matched using the target two - tower model;

[0147] When the retrieval server is online, obtain the retrieval entry to be retrieved, calculate the attention vector of the retrieval entry to be retrieved using the two - tower model, and perform average pooling on the attention vector of the retrieval entry to be retrieved to obtain an average attention vector;

[0148] Select the data to be matched corresponding to the retrieval entry to be retrieved according to the similarity between the fused matching vector and the average attention vector, and use the data to be matched corresponding to the retrieval entry to be retrieved as the retrieval result.

[0149] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0150] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0151] In addition, in each embodiment of the present invention, each functional module can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.

[0152] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0153] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed by the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

[0154] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, an application service layer, etc.

[0155] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results.

[0156] In addition, obviously, the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Words such as first, second, etc. are used to represent names and do not indicate any specific order.

[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. An interactive retrieval method based on a dual - tower model, characterized in that, the method includes: Construct a dual - tower model according to a pre - trained standard language model, obtain retrieval logs, construct a corresponding sample set according to each retrieval entry in the retrieval logs, select a target retrieval entry from the retrieval entries, and select samples that meet preset conditions from the sample set corresponding to the target retrieval entry as the target sample set; Perform multi - channel calculations on each target sample in the target sample set using the dual - tower model to obtain target sample vectors, and perform vector calculations on the target retrieval entry using the dual - tower model to obtain a target retrieval vector; Perform splicing and network calculations on the target sample vectors to obtain fused sample vectors, and initialize multiple parameter vectors, and perform attention calculations according to the parameter vectors and the target retrieval vector to obtain an attention vector; Calculate the similarity between the attention vector and the fused sample vector, determine the matching degree between the target sample and the target retrieval entry according to the similarity, and optimize the dual - tower model according to the matching degree between the target sample set and the target retrieval entry to obtain a target dual - tower model; When a preset retrieval server is offline, obtain preset data to be matched, and calculate a fused matching vector of the data to be matched using the target dual - tower model; When the retrieval server is online, obtain a retrieval entry to be retrieved, calculate an attention vector of the retrieval entry to be retrieved using the dual - tower model, and perform average pooling on the attention vector of the retrieval entry to be retrieved to obtain an average attention vector; Select the data to be matched corresponding to the retrieval entry to be retrieved according to the similarity between the fused matching vector and the average attention vector, and use the data to be matched corresponding to the retrieval entry to be retrieved as the retrieval result; Among them, the performing attention calculations according to the parameter vector and the target retrieval vector to obtain an attention vector includes: generating a vector sequence set according to the parameter vector and the retrieval vector, and successively selecting a target vector sequence from the vector sequence set; performing a first weight calculation on the target vector sequence to obtain an updated sequence corresponding to the target vector sequence; performing a second weight calculation on the updated sequence to obtain a plurality of representation vectors corresponding to the updated sequence; performing an attention operation on the representation vectors corresponding to all updated sequences to obtain an initial attention vector; performing activation calculations on the initial attention vector using a preset activation function, and performing dot - product summation on the result of the activation calculation and the representation vectors corresponding to the updated sequence to obtain an attention vector.

2. The interactive retrieval method based on a dual - tower model according to claim 1, characterized in that, the constructing a dual - tower model according to a pre - trained standard language model includes: Obtain general training corpus data, and perform training in a horizontal domain on a preset language model according to the general training corpus data to obtain a preliminary language model; Obtain vertical training corpus data, and perform training on the preliminary language model according to the vertical training corpus data to obtain the standard language model; Construct the two-tower model according to the standard language model, the preset attention mechanism, and the network calculation layer, where the two-tower model includes a retrieval-end model and a matching-end model.

3. The interactive retrieval method based on the two-tower model according to claim 1, characterized in that the multi-channel calculation of each target sample in the target sample set by using the two-tower model to obtain a target sample vector includes: performing multi-channel recognition on the target sample to obtain multi-channel data corresponding to the target sample; using the matching-end model in the two-tower model to perform vector calculation on the multi-channel data to obtain a language vector for each channel, and using the language vector as the target sample vector corresponding to the target sample.

4. The interactive retrieval method based on the two-tower model according to claim 1, characterized in that the splicing and network calculation of the target sample vector to obtain a fused sample vector includes: splicing the target sample vectors, and performing full connection and activation on the spliced language vectors to obtain the channel weight corresponding to each channel of the target sample; performing weighted summation according to the channel weight of each channel and the target sample vector to obtain the fused sample vector corresponding to the target sample.

5. The interactive retrieval method based on the two-tower model according to claim 4, characterized in that the performing weighted summation according to the channel weight of each channel and the target sample vector to obtain the fused sample vector corresponding to the target sample includes: performing weighted summation on the channel weight of each channel and the target sample vector through the following formula: Among them, is the target sample corresponding fused sample vector; is the total number of channels corresponding to the target sample; is the channel weight of the th channel corresponding to the target sample; is the target sample vector of the th channel corresponding to the target sample.

6. The interactive retrieval method based on the two-tower model according to any one of claims 1 to 5, characterized in that the determining the matching degree between the target sample and the target retrieval entry according to the similarity includes: determining the matching degree between the target sample and the target retrieval entry through the following formula: Among them, is the target sample and the matching degree with the target retrieval entry ; is the th attention vector in the target retrieval entry and the similarity with the fusion sample vector corresponding to the target sample.

7. An interactive retrieval device based on the two-tower model, which is used to implement the interactive retrieval method based on the two-tower model according to any one of claims 1 to 6, characterized in that the device includes: a target sample set generation module, configured to obtain a retrieval log, construct a corresponding sample set according to each retrieval entry in the retrieval log, select a target retrieval entry from the retrieval entries, and select samples that meet preset conditions from the sample set corresponding to the target retrieval entry as the target sample set; a two-tower model calculation module, configured to construct a two-tower model according to a pre-trained standard language model, perform multi-channel calculation on each target sample in the target sample set by using the two-tower model to obtain a target sample vector, and perform vector calculation on the target retrieval entry by using the two-tower model to obtain a target retrieval vector; A target two-tower model generation module is configured to splice and perform network calculations on the target sample vectors to obtain fused sample vectors, initialize multiple parameter vectors, perform attention calculations based on the parameter vectors and the target retrieval vectors to obtain attention vectors; calculate the similarity between the attention vectors and the fused sample vectors, determine the matching degree between the target samples and the target retrieval entries according to the similarity, and optimize the two-tower model according to the matching degree between the target sample set and the target retrieval entries to obtain a target two-tower model; An offline calculation module is configured to obtain preset data to be matched when a preset retrieval server is offline, and calculate the fused matching vector of the data to be matched by using the target two-tower model; An online calculation module is configured to obtain retrieval entries to be retrieved when the retrieval server is online, calculate the attention vectors of the retrieval entries to be retrieved by using the two-tower model, perform average pooling on the attention vectors of the retrieval entries to be retrieved to obtain average attention vectors; select the data to be matched corresponding to the retrieval entries to be retrieved according to the similarity between the fused matching vector and the average attention vector, and use the data to be matched corresponding to the retrieval entries to be retrieved as retrieval results.

8. An electronic device, wherein, the electronic device includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the interactive retrieval method based on a two-tower model according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, wherein, the computer program, when executed by a processor, implements the interactive retrieval method based on a two-tower model according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • FAQ method, question and answer retrieval system, electronic equipment and storage medium

    CN111198940A

  • Commodity information recall method, device and equipment and computer storage medium

    CN114820134A