Model Training Method, Retrieval Method and Related Devices

By integrating text and visual features in a dual tower model to generate multi-modal representations, the method addresses the lack of rich semantic expression in existing models, enhancing search accuracy in e-commerce platforms.

CN116503127BActive Publication Date: 2025-07-15BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310317346.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-27
Publication Date
2025-07-15
Estimated Expiration
2043-03-27

AI Technical Summary

Technical Problem

In the prior art, there is only single modal information on one side of the item tower, resulting in poor retrieval matching accuracy of the trained double tower model during the item search process.

Method used

The target text representation of the first training text for retrieval is generated by the first tower model of the two tower model, and the second tower model is used to fuse the multimodal representation of the training item, including feature information of text and visual information related to the training item, and train the double tower model to determine the target item from the candidate item.

Benefits of technology

It improves the accuracy of searches, can match the search text more accurately, and improves the service quality and user experience of the search system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116503127B_ABST
    Figure CN116503127B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a model training method, a retrieval method, related devices, an electronic device, and a computer-readable storage medium, and relates to the field of computer technologies. The model training method includes: using a first tower model of a two-tower model to generate a target text representation of a first training text for retrieval; according to a second training text related to a training item and training visual information, using a second tower model of the two-tower model to generate a target multi-modal representation of the training item, wherein the target multi-modal representation integrates feature information of the second training text and the training visual information, and there is an association relationship between the training item and the first training text; training the two-tower model according to the target text representation of the first training text and the target multi-modal representation of the training item, wherein the trained two-tower model is configured to determine a target item from candidate items as a retrieval result. According to the present disclosure, the retrieval text can be more accurately matched, and the retrieval accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a model training method, a retrieval method, related devices, an electronic device, and a computer-readable storage medium. Background Art

[0002] In the e-commerce scenario, the main task of searching or retrieving is to quickly retrieve a small number (e.g., at the ten-thousand level) of items related to the user's search intent from a massive item library (e.g., at the billion level).

[0003] In the related art, during the process of training a two-tower model, the user tower and the item tower of the two-tower model are respectively used to generate a text representation of the retrieval text for user retrieval and a representation of single-modal information related to the retrieved items, and the two-tower model is trained based on the text representation and the representation of single-modal information. Summary of the Invention

[0004] In the related art, there is only single-modal information of the item on the item tower side, which cannot meet the requirement of enriching the semantic expression of the item tower. During the process of using the trained two-tower model for item retrieval, the accuracy of retrieval matching is relatively poor.

[0005] In view of the above technical problems, the present disclosure proposes a solution, which can more accurately match the retrieval text and improve the accuracy of retrieval.

[0006] According to a first aspect of the present disclosure, there is provided a model training method, including: using a first tower model of a two-tower model to generate a target text representation of a first training text for retrieval; according to a second training text related to a training item and training visual information, using a second tower model of the two-tower model to generate a target multi-modal representation of the training item, where the target multi-modal representation integrates feature information of the second training text and the training visual information, and the training item is associated with the first training text; training the two-tower model according to the target text representation of the first training text and the target multi-modal representation of the training item, where the trained two-tower model is configured to determine a target item from candidate items as a retrieval result.

[0007] According to a second aspect of the present disclosure, a retrieval method is provided, including: obtaining a current text for retrieval; generating a target text representation of the current text by using a first tower model of a dual-tower model; and determining a target item from the candidate items as a retrieval result according to the target multi-modal representation of the candidate items and the target text representation of the current text, wherein the target multi-modal representation of the candidate items is generated by using a second tower model of the dual-tower model according to candidate text and candidate visual information related to the candidate items, and the target multi-modal representation fuses the feature information of the candidate text and the candidate visual information.

[0008] According to a third aspect of the present disclosure, a model training device is provided, including: a first generation module configured to generate a target text representation of a first training text for retrieval by using a first tower model of a dual-tower model; a second generation module configured to generate a target multi-modal representation of a training item by using a second tower model of the dual-tower model according to a second training text and training visual information related to the training item, wherein the target multi-modal representation fuses the feature information of the second training text and the training visual information, and the training item is associated with the first training text; and a training module configured to train the dual-tower model according to the target text representation of the first training text and the target multi-modal representation of the training item, wherein the trained dual-tower model is configured to determine a target item from candidate items as a retrieval result.

[0009] According to a fourth aspect of the present disclosure, a retrieval device is provided, including: an obtaining module configured to obtain a current text for retrieval; a generating module configured to generate a target text representation of the current text by using a first tower model of a dual-tower model; and a determining module configured to determine a target item from the candidate items as a retrieval result according to the target multi-modal representation of the candidate items and the target text representation of the current text, wherein the target multi-modal representation of the candidate items is generated by using a second tower model of the dual-tower model according to candidate text and candidate visual information related to the candidate items, and the target multi-modal representation fuses the feature information of the candidate text and the candidate visual information.

[0010] According to a fifth aspect of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, the processor being configured to execute the model training method or the retrieval method according to any one of the above embodiments based on instructions stored in the memory.

[0011] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the instructions are executed by a processor, the model training method or the retrieval method according to any one of the above embodiments is implemented.

[0012] In the above embodiments, the retrieval text can be more accurately matched, improving the accuracy of retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.

[0014] With reference to the accompanying drawings, the present disclosure can be more clearly understood from the following detailed description, wherein:

[0015] Figure 1 is a flowchart showing a model training method according to some embodiments of the present disclosure;

[0016] Figure 2 is a schematic diagram showing the generation of a target text representation, a target visual representation, and a target multi-modal representation using a two-tower model according to some embodiments of the present disclosure;

[0017] Figure 3 is a flowchart showing a retrieval method according to some embodiments of the present disclosure;

[0018] Figure 4 is a block diagram showing a model training apparatus according to some embodiments of the present disclosure;

[0019] Figure 5 is a block diagram showing a retrieval apparatus according to some embodiments of the present disclosure;

[0020] Figure 6 is a block diagram showing an electronic device according to some embodiments of the present disclosure

[0021] Figure 7 is a block diagram showing a computer system for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present disclosure.

[0023] At the same time, it should be understood that, for the sake of convenience of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationships.

[0024] The following description of at least one exemplary embodiment is merely illustrative in nature and in no way serves as a limitation on the present disclosure or its application or use.

[0025] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be considered as part of the specification.

[0026] In all examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.

[0027] It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, further discussion thereof is not required in subsequent figures.

[0028] Figure 1 is a flowchart showing a model training method according to some embodiments of the present disclosure.

[0029] As Figure 1 shown, the model training method includes: step S110, generating a target text representation of a first training text for retrieval by using a first tower model of a two-tower model; step S120, generating a target multi-modal representation of a training item by using a second tower model of the two-tower model according to a second training text related to the training item and training visual information, wherein the target multi-modal representation fuses the feature information of the second training text and the training visual information; and step S130, training the two-tower model according to the target text representation of the first training text and the target multi-modal representation of the training item. In some embodiments, the model training method is executed by a model training device. For example, the visual information includes but is not limited to at least one of images and videos.

[0030] In the above embodiment, in the process of training the second tower model of the two-tower model, a multi-modal representation that fuses the text and visual feature information related to the item is introduced, enriching the semantic expression of the item on the second tower model side, and the semantic relationship between the retrieval text and the visual and text information of the item can be learned from a large amount of user retrieval data. As a result, when the trained two-tower model is used for retrieving items, it can more accurately match the retrieval text input by the user, improving the accuracy of the retrieval.

[0031] In the above embodiments, the design of the dual - tower structure is mainly due to the huge number of candidate items. In an e - commerce search system, for a given retrieval text, relevant items need to be retrieved from millions or even hundreds of millions of candidate items. Therefore, considering the computational cost and online efficiency, it is impossible for the model to perform online scoring. By separating the first - tower model and the second - tower model, the feature representations of relevant information can be calculated separately so as to quickly establish an index online, and then quickly retrieve using Faiss (Facebook AI Similarity Search) online. A high - quality semantic understanding model can help the retrieval platform better understand the user's needs, recall a rich variety of items that are more relevant to the user's request and more matched with the retrieval text, and can significantly improve the service quality and user experience of the platform.

[0032] In step S110, the first - tower model of the dual - tower model is used to generate the target text representation of the first training text for retrieval. In some embodiments, the first - tower model can be referred to as the Query tower model or the user tower model.

[0033] In some embodiments, the output of the first - tower model can be expressed as where k represents the dimension of the representation, and q identifies the first training text. For example, the value of k includes but is not limited to 128.

[0034] In some embodiments, the training data for training the dual - tower model includes the first training samples and the second training samples. The first training text in the first training samples is generated based on the description information of the training items (at this time, the first training text can be referred to as the synthetic retrieval text), and the first training text in the second training samples includes the text used in the actual retrieval process.

[0035] In some embodiments, the first training text in the first training samples is a random substring with a specified length intercepted from the description information of the training items. For example, the description information of the training items includes but is not limited to the title information of the training items.

[0036] In some embodiments, the text used in the actual retrieval process includes the retrieval text input by the user, and the first training text in the second training samples further includes the user's feature information. For example, the user's feature information includes but is not limited to personalized feature information such as the user's gender and age. By introducing the user's feature information during the training process, personalized retrieval results can be provided for the user, further improving the accuracy and precision of the retrieval.

[0037] In step S120, according to the second training text and training visual information related to the training item, the second tower model of the two-tower model is used to generate the target multi-modal representation of the training item, where the target multi-modal representation fuses the feature information of the second training text and the training visual information. The training item is associated with the first training text. In some embodiments, the second tower model may be referred to as an item tower model.

[0038] In some embodiments, taking the training data for training the two-tower model including the first training sample and the second training sample as an example, the training items corresponding to the second training sample include the items that are subjected to specified operations during the actual retrieval process. For example, the specified operation includes at least one of a click operation, an operation of browsing for a specified duration, and a sharing operation.

[0039] In some embodiments, in the first training sample, the second training text and training visual information related to the training item corresponding to the description information for generating the first training text are used as positive samples, and the second training text and training visual information related to other training items are used as negative samples; in the second training sample, the second training text and training visual information related to the training item that is subjected to the specified operation during the actual retrieval process are used as positive samples, and the second training text and training visual information related to other training items during the actual retrieval process are used as negative samples.

[0040] For example, the negative samples can be obtained by using a batch negative sampling strategy, that is, when the batch size is B, each item serves as a positive example for the corresponding first training text once, and the other items corresponding to the other first training texts in this batch are its negative examples, and the number of negative examples is B - 1.

[0041] The training set can be expressed as where represents a training sample, including the synthetic retrieval text q i , the corresponding positive example item the negative example item set N i . The positive example relationship is represented by , and the negative example relationship unrelated to the synthetic retrieval text is represented by .

[0042] Training the two-tower model by combining positive and negative samples can improve the learning ability of model training, thereby improving the accuracy of the multi-modal representation of the candidate items generated by the model, and further improving the accuracy of retrieval.

[0043] In some embodiments, the output of the second tower model can be expressed as where k represents the dimension of the representation, and s can identify the item. For example, the value of k includes but is not limited to 128.

[0044] In some embodiments, the second training text includes at least one of the identity document (ID) of the training item, the title of the training item, the product word to which the training item belongs, the category to which the training item belongs (e.g., at least one of the first-level category, the second-level category, and the third-level category), the color information of the training item (e.g., color words), the model information of the training item (e.g., model words), the store information to which the training item belongs (e.g., at least one of the store name and the store logo), and at least one of the product name and the product logo corresponding to the training item. The training visual information includes at least one of a picture (such as a main picture) and a video (such as a main video) of the training item. The second training text can be increased or decreased according to the scenario.

[0045] In step S130, a dual-tower model is trained according to the target text representation of the first training text and the target multi-modal representation of the training item.

[0046] The trained dual-tower model is configured to determine a target item from the candidate items as the retrieval result. In some embodiments, the trained first tower model is configured to generate the target text representation of the current text for retrieval; the trained second tower model is configured to generate the target multi-modal representation of the candidate item according to the candidate text and the candidate visual information related to the candidate item; the target text representation of the current text and the target multi-modal representation of the candidate item are used to determine the target item from the candidate items as the retrieval result.

[0047] In some embodiments, taking the training data for training the dual-tower model including the above first training samples and second training samples as an example, the dual-tower model can be trained through the following steps 1) to 2).

[0048] In step 1), the dual-tower model is first trained according to the target text representation of the first training text in the first training sample and the target multi-modal representation of the training item corresponding to the first training sample.

[0049] In step 2), after the first training of the dual-tower model, the dual-tower model is second trained according to the target text representation of the first training text in the second training sample and the target multi-modal representation of the training item corresponding to the second training sample.

[0050] In the above embodiments, by synthesizing the first training text based on the description information of the training items, a huge number of first training samples can be obtained, ensuring the diversity and richness of the training instances. The dual-tower model is first trained using a larger number of first training samples, and then the dual-tower model is second-trained using the second training samples that are more in line with the actual retrieval scenario. By adopting the method of pre-training combined with fine-tuning, the dual-tower model learns deeper modal alignment and fusion in the first training stage, fully excavates the semantic expressions of the retrieval text and the items, and learns the accurate matching process between the retrieval text and the items in the second training stage. The user retrieval data is used to fit the real data distribution, making the retrieved items more in line with the retrieval text and fully excavating the semantic correlation between the retrieval text and the items. This method can further improve the accuracy of the retrieval matching during the retrieval process of the trained dual-tower model, thereby further improving the retrieval accuracy.

[0051] In some embodiments, the above step 1) can be implemented in the following manner.

[0052] First, according to the second training text related to the training items corresponding to the first training samples, the target text representation of the second training text is generated using the second tower model.

[0053] Then, according to the training visual information related to the training items corresponding to the first training samples, the target visual representation of the training visual information is generated using the second tower model.

[0054] Finally, based on the target text representation of the first training text in the first training samples, the target text representation of the second training text related to the training items corresponding to the first training samples, the target visual representation of the training visual information related to the training items corresponding to the first training samples, and the target multi-modal representation of the training items corresponding to the first training samples, the dual-tower model is first trained.

[0055] In the above embodiments, during the first training process, based on the target text representation of the first training text and the target multi-modal representation of the training items, the target text representation of the second training text and the target visual representation of the training visual information are further introduced, realizing the process of model training from easy to difficult and gradually improving the model learning ability, improving the accuracy of model training, so that the model can more accurately fuse the feature information of the second training text and the training visual information, improve the accuracy of generating the target multi-modal representation, and further improve the retrieval accuracy.

[0056] In some embodiments, the first training of the two-tower model based on the target text representation of the first training text in the first training sample, the target text representation of the second training text related to the training item corresponding to the first training sample, the target visual representation of the training visual information related to the training item corresponding to the first training sample, and the target multi-modal representation of the training item corresponding to the first training sample includes the following steps.

[0057] First, determine the similarity between the target text representation of the second training text related to the training item corresponding to the first training sample and the target text representation of the first training text in the first training sample as the first similarity.

[0058] Second, determine the similarity between the target visual representation of the training visual information related to the training item corresponding to the first training sample and the target text representation of the first training text in the first training sample as the second similarity.

[0059] Then, determine the similarity between the target multi-modal representation of the training item corresponding to the first training sample and the target text representation of the first training text in the first training sample as the third similarity.

[0060] Finally, perform the first training on the two-tower model according to the first similarity, the second similarity, and the third similarity.

[0061] In the above embodiments, during the first training process, the model learns multiple different tasks, such as the semantic mapping between the retrieved text (the first training text) and the text features related to the item, the semantic alignment between the retrieved text and the visual features related to the item, the semantic fusion of the text features and visual features related to the item, and the asymmetric semantic alignment between the retrieved text and the multi-modal features of the item. By means of multi-task learning, combined with the fact that in the second training process, the model only learns the asymmetric modal alignment task, the complexity of the model in the second training stage can be reduced, the accuracy of feature generation and modal fusion by the model can be improved, and thus the accuracy of retrieval can be further improved.

[0062] In some embodiments, the first training of the two-tower model according to the first similarity, the second similarity, and the third similarity can be implemented in the following manner.

[0063] First, determine the first loss value, the second loss value, and the third loss value according to the first similarity, the second similarity, and the third similarity respectively. In some embodiments, the cross-entropy loss function can be used to determine the first loss value, the second loss value, and the third loss value.

[0064] For example, the similarity is calculated for the relevant representations of the first training text and the training item in each training sample (the target text representation of the second training text, the target visual representation of the training visual information, and the target multimodal representation), and the softmax function is used to normalize the similarity so that it belongs to the interval (0, 1). In some embodiments, the normalized similarity can be expressed as

[0065] For example, the logarithm of the normalized similarity is calculated and multiplied by the value representing the positive or negative example relationship between the first training text and the training item, and then the products of all training samples are summed to obtain the loss value. In some embodiments, the loss value can be expressed as

[0066] Combined with the above embodiments, for the semantic alignment task, the scoring function for calculating the first similarity is f t (q, s), and the first loss value is Similarly, for the cross-modal alignment task, the scoring function for calculating the second similarity is f i (q, s), and the second loss value is For the asymmetric modality alignment-fusion task, the scoring function for calculating the third similarity is f m (q, s), and the third loss value is

[0067] Then, according to the first weight value, the second weight value, and the third weight value, the first loss value, the second loss value, and the third loss value are weighted to obtain the total loss value. The first weight value, the second weight value, and the third weight value are configured to adjust the contribution degrees of different modality information of the training item.

[0068] In some embodiments, considering the different contributions of the text information and the visual information of the item, different weights are multiplied to each loss term, so that the final loss function of the pre-trained model is: where α, β, and γ are the weights representing different loss terms, that is, the first weight value, the second weight value, and the third weight value.

[0069] Finally, according to the total loss value, the two-tower model is first trained.

[0070] In the above embodiments, by using different weights to adjust the contribution degrees of different modality information to the multimodal retrieval task, the complementarity and redundancy elimination between different modality information can be achieved, so as to achieve the best retrieval effect and improve the retrieval accuracy. Complementarity and redundancy refer to the ability of different modalities to promote mutual understanding after adding multimodal information, compared with only using the single information of a certain modality.

[0071] For example, for a certain dress, there is a case of keyword stuffing in its item title: "Floral Skirt Dress A-line Skirt". Just looking at the title, it is impossible to tell which type of skirt the item is. If the visual information of the item is added, then the model can understand that the skirt is a dress, rather than other types of skirts. In this way, the representation of the item and the representation of query = "dress" are more matched, thus achieving redundancy removal. Similarly, if the item title of a blue dress is very simple, "dress", just looking at the title cannot know the specific color and other details of this skirt. If the visual information of the item is combined, the model can understand that the skirt is blue. Then the representation of the item and the representation of query = "blue dress" are more matched, thus achieving complementarity.

[0072] In some embodiments, the above step 2) can be implemented in the following manner, that is, the dual-tower model is second-trained according to the target text representation of the first training text in the second training sample and the target multi-modal representation of the training item corresponding to the second training sample.

[0073] First, determine the similarity between the target multi-modal representation of the training item corresponding to the second training sample and the target text representation of the first training text in the second training sample as the fourth similarity.

[0074] Then, second-train the dual-tower model according to the fourth similarity. In some embodiments, according to the fourth similarity, determine the fourth loss value. Similar to the first training process, the fourth loss value can be expressed as

[0075] Finally, second-train the dual-tower model according to the fourth loss value.

[0076] Next, the process of generating the target text representation of the first training text and the second training text, the target visual representation of the training visual information, and the target multi-modal representation of the training item will be described in conjunction with Figure 2 the following.

[0077] Figure 2 FIG. is a schematic diagram showing the generation of the target text representation, the target visual representation, and the target multi-modal representation using a dual-tower model according to some embodiments of the present disclosure.

[0078] As Figure 2 shown, the dual-tower model includes a first tower model and a second tower model.

[0079] In some embodiments, referring to Figure 2 the first tower model includes a feature expression layer and an encoding layer. In this case, the target text representation of the first training text for retrieval can be generated in the following manner.

[0080] First, use the feature representation layer of the first tower model to process the first training text to obtain the initial text representation of the first training text.

[0081] In some embodiments, taking the first training text denoted as q = {t1, …, t n} as an example, the first training text is embedded as a k-dimensional vector in the feature representation layer of the first tower model. t i is a token after unigram tokenization, and the dictionary of tokens is V. Among them, represents the learnable parameter one-hot encoding matrix, m is the size of the dictionary V; e i represents the one-hot encoding of t i ; stack represents embedding connection along the first dimension; represents the text serialization embedding vector of the retrieval text with length n. For example, for the first training text "floral dress", t i includes "broken", "flower", "connected", "clothes", "skirt", collectively called tokens, and the set of tokens is called the dictionary V. A long text such as a sentence or a paragraph is decomposed into a sequence with tokens as the smallest unit, and a token is each item after tokenization.

[0082] Then, use the encoding layer of the first tower model to process the initial text representation of the first training text to obtain the intermediate text representation of the first training text. Among them, the intermediate text representation incorporates the weight relationship between different sub-texts of the first training text relative to the initial text representation. For example, the sub-texts in the first training text can include at least one of a word, a phrase, and a short sentence in the first training text, and the weight relationship between different sub-texts in the first training text can include at least one of the weight relationship between different words, the weight relationship between different phrases, the weight relationship between different short sentences, the weight relationship between a word and a phrase, the weight relationship between a word and a short sentence, and the weight relationship between a phrase and a short sentence in the first training text.

[0083] In some embodiments, taking the encoding layer of the first tower model as a Transformer encoder as an example, the intermediate text representation output by this encoding layer can be expressed as h q = cls(H q ), where H q = transformer q (X q ), cls() means inserting the [CLS] token at the beginning of the retrieval word sequence and using this to represent Hq For example, the input Q (query, query vector), K (key, key vector) and V (value, value vector) in the Transformer encoder of the first tower model are the same, and the self-attention mechanism is adopted.

[0084] Finally, a target text representation of the first training text is generated according to the intermediate text representation of the first training text.

[0085] In some embodiments, reference Figure 2 , the first tower model also includes a normalization layer. In this case, the normalization layer of the first tower model can be used to normalize the intermediate text features of the first training text to obtain the target text representation of the first training text.

[0086] In some embodiments, L2 regularization can be used to normalize the intermediate text features of the first training text. For example, the target text representation of the first training text can be expressed as Q(q)=norm(h q ), where norm() represents the normalization operation.

[0087] In some embodiments, reference Figure 2 , the second tower model includes a feature expression layer and an encoding layer. In this case, the target text representation of the second training text can be generated in the following way.

[0088] First, based on the second training text, an initial text representation of the second training text is generated using the feature expression layer of the second tower model. Figure 2 , taking the feature expression layer in the second tower model including the first feature expression module as an example, the first feature expression module can be used to generate an initial text representation of the second training text.

[0089] In some embodiments, the initial text representation output by the first feature expression module can be expressed as in, Indicates the title information of the item. Represents other text information of the item, such as category, brand, store, etc. and X q The same one-hot encoding matrix is used to achieve the interaction between the search text (first training text) and the text information of the item (second training text). If there is common text in the search text and the item title, then under the premise of using the same one-hot encoding matrix, the one-hot encoding of the common text is the same, which can shorten the distance between the search text and the item information, improve the accuracy of model training, and thus further improve the accuracy of retrieval.

[0090] Then, according to the initial text representation of the second training text, the encoding layer of the second tower model is used to generate the intermediate text representation of the second training text, where the intermediate text representation incorporates the weight relationships between different sub-texts of the second training text relative to the initial text representation. In some embodiments, referring to Figure 2 , taking the encoding layer in the second tower model including the first encoding module as an example, the first encoding module can be used to generate the intermediate text representation of the second training text.

[0091] For example, the sub-texts in the second training text may include at least one of a word, a phrase, and a short sentence in the second training text, and the weight relationships between different sub-texts in the second training text may include at least one of the weight relationships between different words, different phrases, different short sentences, the weight relationship between a word and a phrase, the weight relationship between a word and a short sentence, and the weight relationship between a phrase and a short sentence in the second training text.

[0092] In some embodiments, taking the first encoding module as a Transformer encoder as an example, the intermediate text representation output by the first encoding module can be expressed as h t = cls(H t ), where H t = transformer t (X t ), cls() represents inserting the [CLS] token at the beginning of the second training text sequence and using this to represent the output of H t . For example, the input Q (query), K (key), and V (value) in the first encoding module (Transformer encoder) are the same, and the self-attention mechanism is adopted.

[0093] Finally, according to the intermediate text representation of the second training text, the target text representation of the second training text is generated. In some embodiments, referring to Figure 2 , taking the second tower model further including a normalization layer as an example, the normalization layer of the second tower model can be used to perform normalization processing on the intermediate text representation of the second training text to obtain the target text representation of the second training text. For example, referring to Figure 2 , the normalization layer of the second tower model includes a first normalization module, and the first normalization module can be used to perform normalization processing on the intermediate text representation of the second training text to obtain the target text representation of the second training text.

[0094] In some embodiments, the target text representation output by the first normalization module can be expressed as S t(s) = norm(h t ).

[0095] For example, in the above embodiment, the first feature expression module, the first encoding module, and the first normalization module in the second tower model together constitute the item text tower model in the second tower model.

[0096] In some embodiments, taking the second tower model including a feature expression layer and an encoding layer as an example, the target visual representation for training visual information can be generated in the following manner.

[0097] First, according to the training visual information, the initial visual representation of the training visual information is generated using the feature expression layer of the second tower model. In some embodiments, referring to Figure 2 , taking the feature expression layer in the second tower model including a second feature expression module as an example, the initial visual representation of the training visual information can be generated using the second feature expression module.

[0098] In some embodiments, taking the visual information as an image as an example, the initial visual representation output by the second feature expression module can be expressed as For example, the second feature expression module can adopt a VIT (Vision Transformer) model based on CLIP (Contrastive Language-Image Pre-training) as a feature extractor.

[0099] Then, according to the initial visual representation of the training visual information, the intermediate visual representation of the training visual information is generated using the encoding layer of the second tower model, where the intermediate visual representation incorporates the weight relationship between different pixel regions of the training visual information relative to the initial visual representation. In some embodiments, referring to Figure 2 , taking the encoding layer in the second tower model including a second encoding module as an example, the intermediate visual representation of the training visual information can be generated using the second encoding module.

[0100] In some embodiments, taking the second encoding module as a Transformer encoder as an example, the intermediate text representation output by the second encoding module can be expressed as h i = cls(H i ), where H i = transformer i (X i ), cls() represents inserting the [CLS] token at the beginning of the pixel sequence of the visual information and using this to represent H iOutput. For example, the Q (query), K (key), and V (value) vectors input in the second encoding module (Transformer encoder) are the same, and the self-attention mechanism is adopted.

[0101] Finally, according to the intermediate visual representation of the training visual information, the target visual representation of the training visual information is generated. In some embodiments, referring to Figure 2 , taking the second tower model further including a normalization layer as an example, the normalization layer of the second tower model can be used to normalize the intermediate visual representation of the training visual information to obtain the target visual representation of the training visual information. For example, referring to Figure 2 , the normalization layer of the second tower model includes a second normalization module, and the second normalization module can be used to normalize the intermediate visual representation of the training visual information to obtain the target visual representation of the training visual information.

[0102] In some embodiments, the target text representation output by the second normalization module can be expressed as S i (s) = norm(h i ).

[0103] For example, in the above embodiments, the second feature expression module, the second encoding module, and the second normalization module in the second tower model together constitute the item visual tower model in the second tower model. In some embodiments, the above target text representation can also be obtained by performing feature extraction operations through a multi-layer perceptron, an RNN (Recurrent Neural Network) model, an LSTM (Long short-term memory) model, etc.

[0104] In some embodiments, referring to Figure 2 , taking the encoding layer of the second tower model including a first encoding module, a second encoding module, and a cross-attention module as an example, the target multi-modal representation of the training item can be generated in the following manner.

[0105] First, according to the intermediate text representation of the second training text and the intermediate visual representation of the training visual information, using the cross-attention module, the intermediate multi-modal representation of the training item is generated, where the intermediate multi-modal representation incorporates the weight relationship between different modalities. The generation of the intermediate text representation of the second training text and the intermediate visual representation of the training visual information can refer to the foregoing embodiments and will not be elaborated here.

[0106] In some embodiments, taking the cross-attention module as a Transformer-based cross-attention model as an example, the intermediate multi-modal representation output by the cross-attention module can be expressed as hm = cls(H m ), where H m = transformer m (H t , H i , H i ). The cls() indicates inserting the [CLS] token at the beginning of the sequence of the combined visual information and text, and using this to represent the output of H m . For example, in the cross-attention module, the input Q (query, query vector), K (key, key vector), and V (value, value vector) are not the same, and the cross-attention mechanism is adopted. By using the cross-attention mechanism, the text features and visual features of the item can be combined to achieve multi-modal fusion between the text modality and the visual modality.

[0107] Then, according to the intermediate multi-modal representation, the target multi-modal representation of the training item is generated. In some embodiments, referring to Figure 2 , taking the example that the second tower model further includes a normalization layer, the normalization layer of the second tower model can be used to perform normalization processing on the intermediate multi-modal representation to obtain the target multi-modal representation of the training item. For example, referring to Figure 2 , the normalization layer of the second tower model includes a third normalization module, and the third normalization module can be used to perform normalization processing on the intermediate multi-modal representation to obtain the target multi-modal representation of the training item.

[0108] In some embodiments, the target multi-modal representation output by the third normalization module can be expressed as S m (s) = norm(h m ).

[0109] For example, in the above embodiment, the first feature expression module, the second feature expression module, the first encoding module, the second encoding module, the cross-attention module, and the third normalization module in the second tower model together constitute the item multi-modal tower model in the second tower model.

[0110] In some embodiments, in combination with Figure 2 , the first similarity can be expressed as f t (q, s) = F(Q(q), S t (s)), where f t (q, s) represents the output of the scoring function F corresponding to the first tower model, and is used to calculate the first similarity. For example, the scoring function F uses the inner product. Similarly, the second similarity can be expressed as f i (q, s) = F(Q(q), S i (s)), and the third similarity can be expressed as f m(q, s) = F(Q(q), S m (s)).

[0111] In the above embodiments, each tower model in the two - tower model respectively includes an input layer, a hidden - layer network, and an output layer. The feature - expression layer belongs to the input layer, the encoding layer belongs to the hidden - layer network, and the normalization layer belongs to the output layer. The input layer completes the feature expression of the visual and text information of the retrieval text and the item. The hidden - layer network uses a deep neural network (for example, the Transformer model) for deep expression. The output layer completes the semantic interaction between the query tower (or retrieval - term tower) and the item tower. The output layer outputs the final representation of the visual and text information of the retrieval text and the item.

[0112] The present disclosure uses the recall rate, precision rate of the retrieval system, and the harmonic mean of the recall rate and precision rate as the measurement indicators of the retrieval performance, and compares the retrieval performance of the two - tower model after training according to the present disclosure with the two - tower model in the related art in all categories, fashion categories, and non - fashion categories of items. For example, the recall - precision rate, that is, the ratio of the number of query - related items retrieved in the top - K results to the number of all related items in the inventory, is used to measure the recall rate of the retrieval system, and the higher the value of this indicator, the better. Another example is to use the ratio of the number of query - related items retrieved in the top - K results to the total number of retrieved items to measure the precision rate of the retrieval system, and the higher the value of this indicator, the better. The higher the harmonic mean, the better the performance of the retrieval system.

[0113] Through the comparative experiment, it is found that the retrieval system based on the two - tower model after training according to the present disclosure has improved or maintained in terms of the recall rate, precision rate, and harmonic - mean performance compared with the two - tower model in the related art, verifying the effectiveness of the two - tower model after training according to the present disclosure and being able to effectively improve the accuracy of retrieval. For the embodiments of multi - task learning, different learning tasks not only help the model learn semantic mapping, but also can ensure the performance of fusion and alignment between different modalities, and can further effectively improve the accuracy of retrieval.

[0114] Taking the embodiment of multi - task learning as an example, Table 1 shows a set of comparative - experiment data of the retrieval performance of the present disclosure. The experimental data in Table 1 are obtained by removing the semantic - alignment task, cross - modal alignment task, or asymmetric - modality alignment - fusion task and comparing with the embodiment in which all three tasks of the present disclosure exist in terms of retrieval performance.

[0115] As shown in Table 1, the retrieval system in the case of the multi-task multi-modal retrieval model of the present disclosure has improved or maintained in terms of recall rate, precision rate, and harmonic mean performance compared to other cases, verifying the effectiveness of the dual-tower model trained by multi-task learning of the present disclosure. Different learning tasks not only help the model learn semantic mapping but also ensure the performance of fusion and alignment between different modalities, effectively improving the accuracy of retrieval.

[0116] Table 1

[0117]

[0118] In addition, by testing the impact of the dual-tower model trained based on the embodiments of the present disclosure after going online on the Gross Merchandise Volume (GMV) and User Conversion Rate (UCR), it reflects from the side that the present disclosure can improve the accuracy of retrieval. Table 2 shows the impact of the dual-tower model trained based on the embodiments of the present disclosure after going online on the gross merchandise volume and user conversion rate.

[0119] Table 2

[0120] GMV UCVR All Categories +0.285% +0.174% Fashion Categories +1.112% +0.437

[0121] As shown in Table 2, the trained dual-tower model of the present disclosure has improved in terms of gross merchandise volume and user conversion rate for all categories and fashion categories.

[0122] In addition, a visualization test was also conducted on the trained dual-tower model of the present disclosure. Table 3 shows the retrieval results in three cases: pure visual information retrieval (only using visual information on the item side), pure text retrieval (only using text information on the item side), and multi-modal retrieval (using multi-modal information that fuses visual information and text information on the item side).

[0123] Table 3

[0124]

[0125] As shown in Table 3, for the retrieval terms "nude high heels" and "solid color long dress" in the fashion category, there are mismatches in the recall results of the single-modal retrieval model (pure visual information retrieval and pure text retrieval). For example, "silver" instead of "nude" high heels is recalled, and "floral" instead of "solid color" dress is recalled. For the retrieval term "tooth cleaner" in the regular category, both single-modal models recall categories that do not match the retrieval term, such as "toothbrush" and "electric toothbrush head". It can be seen that the multi-modal retrieval model of the present disclosure performs better in both regular categories and fashion categories, and the retrieval is more accurate and precise.

[0126] Figure 3 It is a flowchart showing a retrieval method according to some embodiments of the present disclosure.

[0127] As Figure 3 shown, the retrieval method includes step S310 to step S330. In some embodiments, the retrieval method is executed by a retrieval device.

[0128] In step S310, obtain the current text for retrieval.

[0129] In some embodiments, obtaining the current text for retrieval can be achieved in the following manner.

[0130] First, receive the retrieval text input by the user.

[0131] Then, in response to receiving the retrieval text input by the user, obtain the user's characteristic information. For example, the user's characteristic information includes, but is not limited to, personalized characteristic information such as the user's gender, age, etc. By introducing the user's characteristic information, personalized retrieval results can be provided for the user, further improving the accuracy and precision of the retrieval.

[0132] Finally, determine the retrieval text input by the user and the user's characteristic information as the current text.

[0133] In step S320, use the first tower model of the two - tower model to generate the target text representation of the current text.

[0134] In step S330, according to the target multimodal representation of the candidate item and the target text representation of the current text, determine the target item from the candidate items as the result of the retrieval. The target multimodal representation of the candidate item is generated using the second tower model of the two - tower model based on the candidate text and candidate visual information related to the candidate item, and the target multimodal representation integrates the characteristic information of the candidate text and candidate visual information.

[0135] In some embodiments, determining the target item from the candidate items includes: determining the similarity between the target multimodal representation of the candidate item and the target text representation of the current text; determining, from the candidate items, the candidate items with a similarity greater than the similarity threshold as the target items.

[0136] In the above - mentioned embodiments, retrieval is performed based on the multimodal representation that integrates the candidate text and candidate visual information, enriching the semantic expression of the candidate items and improving the accuracy of the retrieval.

[0137] In some embodiments, taking the first tower model including a feature expression layer and an encoding layer as an example, generating the target text representation of the current text includes the following steps.

[0138] First, use the feature representation layer of the first tower model to process the current text and obtain the initial text representation of the current text.

[0139] Then, use the encoding layer of the first tower model to process the initial text representation of the current text and obtain the intermediate text representation of the current text, where the intermediate text representation incorporates the weight relationships between different sub-texts of the current text with respect to the initial text representation.

[0140] For example, the sub-texts in the current text may include at least one of a word, a phrase, and a short sentence in the current text, and the weight relationships between different sub-texts in the current text may include at least one of the weight relationships between different words, between different phrases, between different short sentences, between a word and a phrase, between a word and a short sentence, and between a phrase and a short sentence in the current text.

[0141] Finally, generate the target text representation of the current text based on the intermediate text representation of the current text.

[0142] In some embodiments, the first tower model further includes a normalization layer, and generating the target text representation of the current text based on the intermediate text representation of the current text includes: using the normalization layer of the first tower model to perform normalization processing on the intermediate text features of the current text to obtain the target text representation of the current text.

[0143] In some embodiments, the second tower model includes a feature representation layer, an encoding layer, and an output layer, and the encoding layer includes a first encoding module, a second encoding module, and a cross-attention module.

[0144] The feature representation layer of the second tower model is configured to generate the initial text representation of the candidate text and the initial visual representation of the candidate visual information based on the candidate text and the candidate visual information respectively using the feature representation layer of the second tower model.

[0145] The first encoding module of the second tower model is configured to generate the intermediate text representation of the candidate text based on the initial text representation of the candidate text, where the intermediate text representation incorporates the weight relationships between different sub-texts of the candidate text with respect to the initial text representation.

[0146] The second encoding module of the second tower model is configured to generate the intermediate visual representation of the candidate visual information based on the candidate visual information, where the intermediate visual representation incorporates the weight relationships between different pixel regions of the candidate visual information with respect to the initial visual representation.

[0147] The cross-attention layer is configured to generate an intermediate multi-modal representation of the candidate item based on the intermediate text representation of the candidate text and the intermediate visual representation of the candidate visual information, wherein the intermediate multi-modal representation incorporates the weight relationship between different modalities.

[0148] The output layer of the second tower model is configured to generate a target multi-modal representation of the candidate item based on the intermediate multi-modal representation of the candidate item.

[0149] In some embodiments, the output layer of the second tower model is a normalization layer. The normalization layer of the second tower model is configured to perform normalization processing on the intermediate multi-modal representation to obtain the target multi-modal representation of the candidate item.

[0150] For the descriptions of the corresponding parts in the retrieval method and the model training method, reference can be made to the various embodiments of the model training method, which will not be elaborated here.

[0151] Figure 4 is a block diagram showing a model training apparatus according to some embodiments of the present disclosure.

[0152] As Figure 4 shown, the model training apparatus 4 includes a first generation module 41, a second generation module 42, and a training module 43.

[0153] The first generation module 41 is configured to use the first tower model of the two-tower model to generate a target text representation of the first training text for retrieval, for example, performing the steps as Figure 1 shown in step S110.

[0154] In some embodiments, the first tower model includes a feature expression layer and an encoding layer. In this case, the first generation module 41 is configured to use the feature expression layer of the first tower model to process the first training text to obtain an initial text representation of the first training text; use the encoding layer of the first tower model to process the initial text representation of the first training text to obtain an intermediate text representation of the first training text, wherein the intermediate text representation incorporates the weight relationship between different sub-texts of the first training text relative to the initial text representation; and generate a target text representation of the first training text based on the intermediate text representation of the first training text.

[0155] In some embodiments, the first tower model further includes a normalization layer, and the first generation module 41 is configured to use the normalization layer of the first tower model to perform normalization processing on the intermediate text features of the first training text to obtain the target text representation of the first training text.

[0156] The second generation module 42 is configured to generate a target multi-modal representation of the training item according to the second training text and training visual information related to the training item, using the second tower model of the two-tower model, wherein the target multi-modal representation integrates the feature information of the second training text and the training visual information, and there is an association relationship between the training item and the first training text, for example, performing steps such as Figure 1 shown in step S120.

[0157] In some embodiments, the second tower model includes a feature expression layer and an encoding layer, and the encoding layer includes a first encoding module, a second encoding module, and a cross-attention module.

[0158] In this case, the second generation module 42 is configured to generate an initial text representation of the second training text and an initial visual representation of the training visual information respectively according to the second training text and the training visual information, using the feature expression layer of the second tower model; generate an intermediate text representation of the second training text according to the initial text representation of the second training text, using the first encoding module of the second tower model, wherein the intermediate text representation incorporates the weight relationship between different sub-texts of the second training text compared to the initial text representation; generate an intermediate visual representation of the training visual information according to the training visual information, using the second encoding module of the second tower model, wherein the intermediate visual representation incorporates the weight relationship between different pixel regions of the training visual information compared to the initial visual representation; generate an intermediate multi-modal representation of the training item according to the intermediate text representation of the second training text and the intermediate visual representation of the training visual information, using the cross-attention module, wherein the intermediate multi-modal representation incorporates the weight relationship between different modalities; generate a target multi-modal representation of the training item according to the intermediate multi-modal representation.

[0159] In some embodiments, the second tower model further includes a normalization layer, and the second generation module 42 is configured to use the normalization layer of the second tower model to perform normalization processing on the intermediate multi-modal representation to obtain the target multi-modal representation of the training item.

[0160] The training module 43 is configured to train the two-tower model according to the target text representation of the first training text and the target multi-modal representation of the training item, wherein the trained two-tower model is configured to determine a target item from the candidate items as the retrieval result, for example, performing steps such as Figure 1 shown in step S130.

[0161] In some embodiments, the trained first tower model is configured to generate a target text representation of the current text for retrieval; the trained second tower model is configured to generate a target multi-modal representation of a candidate item based on candidate text and candidate visual information related to the candidate item; the target text representation of the current text and the target multi-modal representation of the candidate item are used to determine a target item from the candidate items as the retrieval result.

[0162] In some embodiments, the training data for training the two-tower model includes first training samples and second training samples. The first training text in the first training samples is generated based on the description information of the training items. The first training text in the second training samples includes the text used in the actual retrieval process, and the training items corresponding to the second training samples include the items that have performed specified operations in the actual retrieval process.

[0163] In this case, the training module 43 includes a first training unit and a second training unit. The first training unit is configured to perform first training on the two-tower model according to the target text representation of the first training text in the first training samples and the target multi-modal representation of the training items corresponding to the first training samples. The second training unit is configured to perform second training on the two-tower model according to the target text representation of the first training text in the second training samples and the target multi-modal representation of the training items corresponding to the second training samples after the first training on the two-tower model.

[0164] In some embodiments, the first training unit is configured to generate a target text representation of the second training text by using the second tower model according to the second training text related to the training items corresponding to the first training samples; generate a target visual representation of the training visual information by using the second tower model according to the training visual information related to the training items corresponding to the first training samples; and perform first training on the two-tower model according to the target text representation of the first training text in the first training samples, the target text representation of the second training text related to the training items corresponding to the first training samples, the target visual representation of the training visual information related to the training items corresponding to the first training samples, and the target multi-modal representation of the training items corresponding to the first training samples.

[0165] In some embodiments, the first training unit is configured to determine the similarity between the target text representation of the second training text related to the training item corresponding to the first training sample and the target text representation of the first training text in the first training sample as the first similarity; determine the similarity between the target visual representation of the training visual information related to the training item corresponding to the first training sample and the target text representation of the first training text in the first training sample as the second similarity; determine the similarity between the target multi-modal representation of the training item corresponding to the first training sample and the target text representation of the first training text in the first training sample as the third similarity; and perform the first training on the two-tower model according to the first similarity, the second similarity, and the third similarity.

[0166] In some embodiments, the first training unit is configured to determine a first loss value, a second loss value, and a third loss value respectively according to the first similarity, the second similarity, and the third similarity; perform a weighting operation on the first loss value, the second loss value, and the third loss value according to a first weight value, a second weight value, and a third weight value to obtain a total loss value, where the first weight value, the second weight value, and the third weight value are configured to adjust the contribution degrees of different modal information of the training item; and perform the first training on the two-tower model according to the total loss value.

[0167] In some embodiments, the second tower model includes a feature expression layer and an encoding layer. In this case, the first training unit is configured to generate an initial text representation of the second training text and an initial visual representation of the training visual information by using the feature expression layer of the second tower model according to the second training text and the training visual information; respectively generate an intermediate text representation of the second training text and an intermediate visual representation of the training visual information by using the encoding layer of the second tower model according to the initial text representation of the second training text and the initial visual representation of the training visual information, where the intermediate text representation incorporates the weight relationship between different sub-texts of the second training text relative to the initial text representation, and the intermediate visual representation incorporates the weight relationship between different pixel regions of the training visual information relative to the initial visual representation; and generate a target text representation of the second training text and a target visual representation of the training visual information respectively according to the intermediate text representation of the second training text and the intermediate visual representation of the training visual information.

[0168] In some embodiments, the second tower model further includes a normalization layer, and the first training unit is configured to perform normalization processing on the intermediate text representation of the second training text and the intermediate visual representation of the training visual information respectively by using the normalization layer of the second tower model to obtain a target text representation of the second training text and a target visual representation of the training visual information.

[0169] In some embodiments, the second training unit is configured to determine the similarity between the target multi-modal representation of the training item corresponding to the second training sample and the target text representation of the first training text in the second training sample as the fourth similarity; and perform the second training on the two-tower model according to the fourth similarity.

[0170] In some embodiments, in the first training sample, the second training text and the training visual information related to the training item corresponding to the description information for generating the first training text are used as positive samples, and the second training text and the training visual information related to other training items are used as negative samples; in the second training sample, the second training text and the training visual information related to the training item that has performed the specified operation during the actual retrieval process are used as positive samples, and the second training text and the training visual information related to other training items during the actual retrieval process are used as negative samples.

[0171] Figure 5 FIG. is a block diagram showing a retrieval device according to some embodiments of the present disclosure.

[0172] As Figure 5 shown, the retrieval device 5 includes an acquisition module 51, a generation module 52, and a determination module 53.

[0173] The acquisition module 51 is configured to acquire the current text for retrieval, for example, perform step S310 as Figure 3 shown.

[0174] The generation module 52 is configured to use the first tower model of the two-tower model to generate the target text representation of the current text, for example, perform step S320 as Figure 3 shown.

[0175] In some embodiments, the first tower model includes a feature expression layer and an encoding layer. The generation module 52 is configured to use the feature expression layer of the first tower model to process the current text to obtain the initial text representation of the current text; use the encoding layer of the first tower model to process the initial text representation of the current text to obtain the intermediate text representation of the current text, where the intermediate text representation incorporates the weight relationship between different sub-texts of the current text relative to the initial text representation; and generate the target text representation of the current text according to the intermediate text representation of the current text.

[0176] In some embodiments, the first tower model further includes a normalization layer. The generation module 52 is configured to use the normalization layer of the first tower model to perform normalization processing on the intermediate text features of the current text to obtain the target text representation of the current text.

[0177] The determination module 53 is configured to determine a target item from candidate items according to the target multi-modal representation of the candidate items and the target text representation of the current text as a retrieval result, where the target multi-modal representation of the candidate items is generated by using the second tower model of the two-tower model according to candidate text and candidate visual information related to the candidate items, and the target multi-modal representation integrates the feature information of the candidate text and the candidate visual information, for example, performing steps such as Figure 3 shown in step S330.

[0178] In some embodiments, the determination module 53 is configured to determine the similarity between the target multi-modal representation of the candidate items and the target text representation of the current text; and determine, from the candidate items, the candidate items with a similarity greater than the similarity threshold as the target items.

[0179] In some embodiments, the second tower model includes a feature expression layer, an encoding layer, and an output layer, and the encoding layer includes a first encoding module, a second encoding module, and a cross-attention module.

[0180] The feature expression layer of the second tower model is configured to generate an initial text representation of the candidate text and an initial visual representation of the candidate visual information respectively according to the candidate text and the candidate visual information by using the feature expression layer of the second tower model.

[0181] The first encoding module of the second tower model is configured to generate an intermediate text representation of the candidate text according to the initial text representation of the candidate text, where the intermediate text representation incorporates the weight relationship between different sub-texts of the candidate text relative to the initial text representation.

[0182] The second encoding module of the second tower model is configured to generate an intermediate visual representation of the candidate visual information according to the candidate visual information, where the intermediate visual representation incorporates the weight relationship between different pixel regions of the candidate visual information relative to the initial visual representation.

[0183] The cross-attention layer is configured to generate an intermediate multi-modal representation of the candidate items according to the intermediate text representation of the candidate text and the intermediate visual representation of the candidate visual information, where the intermediate multi-modal representation incorporates the weight relationship between different modalities.

[0184] The output layer of the second tower model is configured to generate a target multi-modal representation of the candidate items according to the intermediate multi-modal representation of the candidate items.

[0185] In some embodiments, the output layer of the second tower model is a normalization layer, and the normalization layer of the second tower model is configured to perform normalization processing on the intermediate multi-modal representation to obtain the target multi-modal representation of the candidate items.

[0186] Figure 6is a block diagram showing an electronic device according to some embodiments of the present disclosure.

[0187] As Figure 6 shown, the electronic device 6 includes a memory 61; and a processor 62 coupled to the memory 61. The memory 61 is used to store instructions for implementing corresponding embodiments of the model training method or the retrieval method. The processor 62 is configured to execute the model training method or the retrieval method in any of the embodiments of the present disclosure based on the instructions stored in the memory 61.

[0188] Figure 7 is a block diagram showing a computer system for implementing some embodiments of the present disclosure.

[0189] As Figure 7 shown, the computer system 70 may be embodied in the form of a general-purpose computing device. The computer system 70 includes a memory 710, a processor 720, and a bus 700 connecting different system components.

[0190] The memory 710 may include, for example, a system memory, a non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs. The system memory may include a volatile storage medium, such as a random access memory (RAM) and / or a cache memory. The non-volatile storage medium stores, for example, instructions for implementing corresponding embodiments of at least one of the model training method and the retrieval method. The non-volatile storage medium includes, but is not limited to, a disk memory, an optical memory, a flash memory, etc.

[0191] The processor 720 may be implemented in the form of a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, or discrete hardware components such as transistors. Correspondingly, each module such as a judgment module and a determination module may be implemented by a central processing unit (CPU) running instructions for performing corresponding steps in the memory, or may be implemented by a dedicated circuit for performing the corresponding steps.

[0192] The bus 700 may use any of a variety of bus structures. For example, the bus structure includes, but is not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus.

[0193] The computer system 70 may also include an input / output interface 730, a network interface 740, a storage interface 750, etc. These interfaces 730, 740, 750, the memory 710, and the processor 720 may be connected through a bus 700. The input / output interface 730 may provide a connection interface for input / output devices such as a display, a mouse, and a keyboard. The network interface 740 provides a connection interface for various networking devices. The storage interface 750 provides a connection interface for external storage devices such as a floppy disk, a USB flash drive, and an SD card.

[0194] Here, various aspects of the present disclosure have been described with reference to the flowcharts and / or block diagrams of methods, apparatuses, and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of the blocks, can be implemented by computer-readable program instructions.

[0195] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable devices to generate a machine, such that the device implemented by executing the instructions by the processor realizes the functions specified in one or more blocks in the flowcharts and / or block diagrams.

[0196] These computer-readable program instructions can also be stored in a computer-readable memory, and these instructions cause the computer to work in a specific manner, thereby generating a manufactured article including instructions for realizing the functions specified in one or more blocks in the flowcharts and / or block diagrams.

[0197] The present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.

[0198] It should be noted that in the technical solution of the present disclosure, aspects such as the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken for user personal information to prevent illegal access to user personal information data, and to safeguard the security of user personal information, network security, and national security.

[0199] Through the model training method, retrieval method, related apparatuses, electronic devices, and computer-readable storage media in the above embodiments, the retrieval text can be more accurately matched, and the accuracy of retrieval can be improved.

[0200] So far, the model training method, retrieval method, related apparatuses, electronic devices, and computer-readable storage media according to the present disclosure have been described in detail. To avoid obscuring the concept of the present disclosure, some details well known in the art have not been described. Those skilled in the art can clearly understand how to implement the technical solution disclosed herein based on the above description.

Claims

1. A model training method, comprising: Generating a target text representation of a first training text for retrieval by using a first tower model of a two-tower model; Generating a target multi-modal representation of the training item by using a second tower model of the two-tower model according to a second training text related to the training item and training visual information, wherein the target multi-modal representation fuses feature information of the second training text and the training visual information, and there is an association relationship between the training item and the first training text; Training the two-tower model according to the target text representation of the first training text and the target multi-modal representation of the training item, wherein the trained two-tower model is configured to determine a target item from candidate items as a retrieval result.

2. The model training method according to claim 1, wherein, The training data for training the two-tower model includes first training samples and second training samples. The first training text in the first training samples is generated based on the description information of the training item. The first training text in the second training samples includes the text used in the actual retrieval process. The training items corresponding to the second training samples include the items that have performed specified operations in the actual retrieval process. Training the two-tower model includes: Performing a first training on the two-tower model according to the target text representation of the first training text in the first training samples and the target multi-modal representation of the training item corresponding to the first training samples; After performing the first training on the two-tower model, performing a second training on the two-tower model according to the target text representation of the first training text in the second training samples and the target multi-modal representation of the training item corresponding to the second training samples.

3. The model training method according to claim 2, wherein, Performing a first training on the two-tower model according to the target text representation of the first training text in the first training samples and the target multi-modal representation of the training item corresponding to the first training samples includes: Generating a target text representation of the second training text by using the second tower model according to the second training text related to the training item corresponding to the first training samples; Generating a target visual representation of the training visual information by using the second tower model according to the training visual information related to the training item corresponding to the first training samples; Performing a first training on the two-tower model according to the target text representation of the first training text in the first training samples, the target text representation of the second training text related to the training item corresponding to the first training samples, the target visual representation of the training visual information related to the training item corresponding to the first training samples, and the target multi-modal representation of the training item corresponding to the first training samples.

4. The model training method according to claim 3, wherein, Performing a first training on the two-tower model according to the target text representation of the first training text in the first training samples, the target text representation of the second training text related to the training item corresponding to the first training samples, the target visual representation of the training visual information related to the training item corresponding to the first training samples, and the target multi-modal representation of the training item corresponding to the first training samples includes: Determine the similarity between the target text representation of the second training text related to the training item corresponding to the first training sample and the target text representation of the first training text in the first training sample as the first similarity; Determine the similarity between the target visual representation of the training visual information related to the training item corresponding to the first training sample and the target text representation of the first training text in the first training sample as the second similarity; Determine the similarity between the target multi-modal representation of the training item corresponding to the first training sample and the target text representation of the first training text in the first training sample as the third similarity; Perform the first training on the two-tower model according to the first similarity, the second similarity, and the third similarity.

5. The model training method according to claim 4, wherein, Performing the first training on the two-tower model according to the first similarity, the second similarity, and the third similarity includes: Determine a first loss value, a second loss value, and a third loss value according to the first similarity, the second similarity, and the third similarity respectively; Perform a weighted operation on the first loss value, the second loss value, and the third loss value according to a first weight value, a second weight value, and a third weight value to obtain a total loss value, where the first weight value, the second weight value, and the third weight value are configured to adjust the contribution degrees of different modal information of the training item; Perform the first training on the two-tower model according to the total loss value.

6. The model training method according to claim 3, wherein, The second tower model includes a feature expression layer and an encoding layer, where Respectively according to the second training text and the training visual information, use the feature expression layer of the second tower model to generate an initial text representation of the second training text and an initial visual representation of the training visual information; Respectively according to the initial text representation of the second training text and the initial visual representation of the training visual information, use the encoding layer of the second tower model to generate an intermediate text representation of the second training text and an intermediate visual representation of the training visual information, where the intermediate text representation incorporates the weight relationship between different sub-texts of the second training text relative to the initial text representation, and the intermediate visual representation incorporates the weight relationship between different pixel regions of the training visual information relative to the initial visual representation; Respectively according to the intermediate text representation of the second training text and the intermediate visual representation of the training visual information, generate a target text representation of the second training text and a target visual representation of the training visual information.

7. The model training method according to claim 6, wherein, The second tower model further includes a normalization layer, and generating the target text representation of the second training text and the target visual representation of the training visual information includes: Use the normalization layer of the second tower model to perform normalization processing on the intermediate text representation of the second training text and the intermediate visual representation of the training visual information respectively to obtain the target text representation of the second training text and the target visual representation of the training visual information.

8. The model training method according to claim 2, wherein, The second training of the two-tower model based on the target text representation of the first training text in the second training sample and the target multi-modal representation of the training item corresponding to the second training sample includes: Determining the similarity between the target multi-modal representation of the training item corresponding to the second training sample and the target text representation of the first training text in the second training sample as the fourth similarity; Performing second training on the two-tower model according to the fourth similarity.

9. The model training method according to claim 2, wherein In the first training sample, the second training text and training visual information related to the training item corresponding to the description information for generating the first training text are used as positive samples, and the second training text and training visual information related to other training items are used as negative samples; In the second training sample, the second training text and training visual information related to the training item that performs a specified operation during the actual retrieval process are used as positive samples, and the second training text and training visual information related to other training items during the actual retrieval process are used as negative samples.

10. The model training method according to any one of claims 1-9, wherein, The second tower model includes a feature expression layer and an encoding layer. The encoding layer includes a first encoding module, a second encoding module, and a cross-attention module. Generating the target multi-modal representation of the training item includes: Respectively according to the second training text and the training visual information, using the feature expression layer of the second tower model to generate an initial text representation of the second training text and an initial visual representation of the training visual information; According to the initial text representation of the second training text, using the first encoding module of the second tower model to generate an intermediate text representation of the second training text, wherein the intermediate text representation incorporates the weight relationship between different sub-texts of the second training text relative to the initial text representation; According to the training visual information, using the second encoding module of the second tower model to generate an intermediate visual representation of the training visual information, wherein the intermediate visual representation incorporates the weight relationship between different pixel regions of the training visual information relative to the initial visual representation; According to the intermediate text representation of the second training text and the intermediate visual representation of the training visual information, using the cross-attention module to generate an intermediate multi-modal representation of the training item, wherein the intermediate multi-modal representation incorporates the weight relationship between different modalities; Generating the target multi-modal representation of the training item according to the intermediate multi-modal representation.

11. The model training method according to claim 10, wherein, The second tower model further includes a normalization layer. Generating the target multi-modal representation of the training item according to the intermediate multi-modal representation includes: Using the normalization layer of the second tower model to perform normalization processing on the intermediate multi-modal representation to obtain the target multi-modal representation of the training item.

12. The model training method according to any one of claims 1-9, wherein, The first tower model includes a feature expression layer and an encoding layer. Generating the target text representation of the first training text for retrieval includes: Using the feature expression layer of the first tower model to process the first training text to obtain an initial text representation of the first training text; Using the encoding layer of the first tower model, process the initial text representation of the first training text to obtain the intermediate text representation of the first training text, where the intermediate text representation incorporates the weight relationship between different sub-texts of the first training text with respect to the initial text representation; Generate the target text representation of the first training text according to the intermediate text representation of the first training text.

13. The model training method according to claim 12, wherein, The first tower model further includes a normalization layer. Generating the target text representation of the first training text according to the intermediate text representation of the first training text includes: Using the normalization layer of the first tower model, perform normalization processing on the intermediate text features of the first training text to obtain the target text representation of the first training text.

14. The model training method according to any one of claims 1-9, wherein, The trained first tower model is configured to generate the target text representation of the current text for retrieval; the trained second tower model is configured to generate the target multi-modal representation of the candidate item according to the candidate text and candidate visual information related to the candidate item; the target text representation of the current text and the target multi-modal representation of the candidate item are used to determine the target item from the candidate items as the result of the retrieval.

15. A retrieval method, including: Obtain the current text for retrieval; Use the first tower model of the two-tower model to generate the target text representation of the current text; According to the target multi-modal representation of the candidate item and the target text representation of the current text, determine the target item from the candidate items as the result of the retrieval, where the target multi-modal representation of the candidate item is generated using the second tower model of the two-tower model according to the candidate text and candidate visual information related to the candidate item, and the target multi-modal representation incorporates the feature information of the candidate text and the candidate visual information.

16. The retrieval method according to claim 15, wherein, The first tower model includes a feature expression layer and an encoding layer. Generating the target text representation of the current text includes: Use the feature expression layer of the first tower model to process the current text to obtain the initial text representation of the current text; Use the encoding layer of the first tower model to process the initial text representation of the current text to obtain the intermediate text representation of the current text, where the intermediate text representation incorporates the weight relationship between different sub-texts of the current text with respect to the initial text representation; Generate the target text representation of the current text according to the intermediate text representation of the current text.

17. The retrieval method according to claim 16, wherein, The first tower model further includes a normalization layer. Generating the target text representation of the current text according to the intermediate text representation of the current text includes: Use the normalization layer of the first tower model to perform normalization processing on the intermediate text features of the current text to obtain the target text representation of the current text.

18. The retrieval method according to any one of claims 15-17, wherein, The second tower model includes a feature expression layer, an encoding layer, and an output layer. The encoding layer includes a first encoding module, a second encoding module, and a cross-attention module, where: The feature representation layer of the second tower model is configured to generate an initial text representation of the candidate text and an initial visual representation of the candidate visual information respectively according to the candidate text and the candidate visual information, by using the feature representation layer of the second tower model; The first encoding module of the second tower model is configured to generate an intermediate text representation of the candidate text according to the initial text representation of the candidate text, wherein the intermediate text representation incorporates the weight relationship between different sub-texts of the candidate text with respect to the initial text representation; The second encoding module of the second tower model is configured to generate an intermediate visual representation of the candidate visual information according to the candidate visual information, wherein the intermediate visual representation incorporates the weight relationship between different pixel regions of the candidate visual information with respect to the initial visual representation; The cross-attention layer is configured to generate an intermediate multi-modal representation of the candidate item according to the intermediate text representation of the candidate text and the intermediate visual representation of the candidate visual information, wherein the intermediate multi-modal representation incorporates the weight relationship between different modalities; The output layer of the second tower model is configured to generate a target multi-modal representation of the candidate item according to the intermediate multi-modal representation of the candidate item.

19. The retrieval method according to claim 18, wherein The output layer of the second tower model is a normalization layer, and the normalization layer of the second tower model is configured to perform normalization processing on the intermediate multi-modal representation to obtain the target multi-modal representation of the candidate item.

20. The retrieval method according to any one of claims 15-17, wherein, Determining a target item from the candidate items includes: Determining the similarity between the target multi-modal representation of the candidate item and the target text representation of the current text; Determining, from the candidate items, the candidate items with a similarity greater than the similarity threshold as the target items.

21. A model training device, comprising: A first generation module, configured to generate a target text representation of a first training text for retrieval by using the first tower model of a two-tower model; A second generation module, configured to generate a target multi-modal representation of the training item according to a second training text related to the training item and training visual information, by using the second tower model of the two-tower model, wherein the target multi-modal representation integrates the feature information of the second training text and the training visual information, and the training item has an association relationship with the first training text; A training module, configured to train the two-tower model according to the target text representation of the first training text and the target multi-modal representation of the training item, wherein the trained two-tower model is configured to determine a target item from candidate items as the result of retrieval.

22. A retrieval device, comprising: An acquisition module, configured to acquire a current text for retrieval; A generation module, configured to generate a target text representation of the current text by using the first tower model of a two-tower model; A determination module, configured to determine a target item from the candidate items as a result of retrieval according to a target multi-modal representation of the candidate items and a target text representation of the current text, wherein the target multi-modal representation of the candidate items is generated by using a second tower model of the two-tower model according to candidate text and candidate visual information related to the candidate items, and the target multi-modal representation fuses feature information of the candidate text and the candidate visual information.

23. An electronic device, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute the model training method according to any one of claims 1 to 14 or the retrieval method according to any one of claims 15 to 20 based on instructions stored in the memory.

24. A computer-readable storage medium, on which computer program instructions are stored, and when the instructions are executed by a processor, the model training method according to any one of claims 1 to 14 or the retrieval method according to any one of claims 15 to 20 is implemented.

Citation Information

Patent Citations

  • Multi-modal pre-training model training method, application method and device thereof

    CN112990297A

  • Training method and training device of multi-modal pre-training model and electronic equipment

    CN113283551A