A Five-modal Commodity Pre-training Method and Retrieval System Based on Self-coordinated Contrastive Learning

Through the five-modal commodity pre-training method of self-coordinated comparison learning, the problem of low single-modal data retrieval effect in the existing technology is solved, efficient fusion and retrieval of multi-modal data is achieved, and the accuracy and generalization ability of commodity retrieval is improved.

CN114418032BActive Publication Date: 2025-06-20SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210164795.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-22
Publication Date
2025-06-20
Estimated Expiration
2042-02-22

AI Technical Summary

Technical Problem

The existing product search methods mainly rely on single-modal data and image-level search, resulting in low search results in large-scale real scenarios and lack of full utilization of multimodal data.

Method used

The five-modal commodity pre-training method based on self-coordinated contrast learning is adopted. By constructing a model feature encoding extractor and multimodal pre-training model, embedded representations of images, text, tables, videos and audio are extracted, and high-level semantic fusion and modal correlation correction are performed through self-supervised training and contrast learning methods.

Benefits of technology

A combined product search system with high generalization, high availability and high accuracy is realized, which can effectively utilize multimodal data information to improve the accuracy and effectiveness of product search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114418032B_ABST
    Figure CN114418032B_ABST
Patent Text Reader

Abstract

The present invention discloses a five-modal commodity pre-training method and a retrieval system based on self-coordinated contrast learning. The method is as follows: S1: Construct corresponding modal feature encoding extractors according to different modal data; S2: Combine the feature encodings, position encodings, and segment encodings of each modal data extracted by the modal feature encoding extractors to learn the embedding representations of different modal data; S3: Construct a multi-modal pre-training model for self-coordinated contrast learning; S4: Use the modal feature encoding extractors to learn the embedding representations of different modal data with occluded partial features, input them into the multi-modal pre-training model in step S3 for self-supervised training, perform high-level semantic fusion on each modal data, and continuously correct the inter-modal relevance using the self-coordinated contrast learning method, and recover the features at the corresponding positions during the learning process. The present invention realizes a combined commodity retrieval system with high generalization, high availability, and high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large-scale commodities, and more specifically, to a five-modal commodity pre-training method and retrieval system based on self-coordinated contrast learning. Background Art

[0002] The development of Internet technology has led to the rapid expansion of online shopping platforms, which are favored by more and more people due to their convenience. In online shopping platforms, the richness of commodity types and the shopping needs of users gradually increase over time. Given the diversity of online commodities, more commodities are presented in multiple modalities, that is, a commodity can be described by its display picture, description, parameters, and live broadcast. Therefore, how to use the data of these more modal information to serve large-scale commodity retrieval has become a major research issue. And in the case of large-scale data and in the real scenario lacking label annotation, how to perform large-scale commodity retrieval is a practical but unsolved problem.

[0003] Large-scale multi-modal commodity retrieval has high practical value and application prospects in the e-commerce field. First, it is conducive to improving the accuracy of commodity search and helping online users search for more accurate and specific commodities; second, it can be used to construct an e-commerce knowledge graph and mine commodity relationships; third, the matching commodities obtained through multi-modal fusion retrieval can be used for commodity recommendation to improve the recommendation effect of the shopping platform.

[0004] However, in the field of commodity retrieval, existing methods all train and extract features from single-modal data, such as a piece of text or a picture, and then perform matching search according to the features of the data stored in the retrieval library. However, in the e-commerce field, modal data such as pictures, texts, tables, videos, and audios widely exist in each commodity sample. Due to the lack of full utilization of multiple modal data, the current retrieval method greatly limits the effective improvement of the retrieval effect. More importantly, existing models focus on relatively simple situations, that is, picture-level retrieval. Picture-level retrieval cannot judge the attribute features of these commodities, while more modal data can provide commodity information other than image texture description, such as the color, origin, and material of the commodity. The single-modal data retrieval method lacks generalization in large-scale real-scenario datasets. Summary of the Invention

[0005] In order to solve the problem of low accuracy caused by the fact that existing commodity retrieval mainly relies on single-modal data and picture-level retrieval, the present invention provides a five-modal commodity pre-training method and retrieval system based on self-coordinated contrast learning, so as to realize a combined commodity retrieval system with high generalization, high availability, and high accuracy.

[0006] To achieve the above object of the present invention, the following technical solutions are adopted:

[0007] A five-modal commodity pre-training method based on self-coordinated contrast learning, the method comprising the following steps:

[0008] S1: Construct corresponding modal feature encoding extractors according to different modal data;

[0009] S2: Combine the feature encodings, position encodings, and segment encodings of each modal data extracted by the modal feature encoding extractor to learn the embedding representations of different modal data;

[0010] S3: Construct a multi-modal pre-training model for self-coordinated contrast learning;

[0011] S4: Use the modal feature encoding extractor to learn the embedding representations of different modal data with occluded part features, input them into the multi-modal pre-training model in step S3 for self-supervised training, perform high-level semantic fusion on each modal data, and use the self-coordinated contrast learning method to continuously correct the correlation between modalities, and recover the features at the corresponding positions during the learning process.

[0012] Preferably, the modal data includes five types of modal data: images, texts, tables, videos, and audios;

[0013] Use the bottom-up-attention network as the modal feature encoding extractor to obtain the bounding boxes of the images and the features of their coordinate positions;

[0014] Use word-piece as the modal feature encoding extractor to obtain the relationship features between different tokens of the text;

[0015] Use entity word-piece as the modal feature encoding extractor to obtain the encoding representation of the table modal data. Specifically, after splicing each row of data together, obtain the relationship features between different tokens;

[0016] Use the S3D network as the modal feature encoding extractor to obtain the video representation with spatio-temporal characteristics in the video;

[0017] Use MFCC as the modal feature encoding extractor to obtain the encoding representation of the audio modal data.

[0018] Further, in step S2, specifically learn the embedding representations of different modal data as follows:

[0019] For the bounding boxes and bounding box features output by the bottom-up-attention network, a 5D vector is used to calculate the position information of each bounding box, including the upper left coordinates, lower right coordinates of the bounding box, and the proportion of the bounding box in the entire image. This 5D vector is passed into a linear fully-connected layer to obtain a position encoding; 0 is used as the segment information and passed into a linear fully-connected layer to obtain a segment encoding; the bounding box features are passed into a linear fully-connected layer to obtain an encoding of the bounding box features; the position encoding, segment encoding, and feature encoding are added together to obtain an embedded representation of the image modality;

[0020] For the text sequence, an increasing natural number sequence is used to represent their position information, which is passed into a linear fully-connected layer to obtain a position encoding; 1 is used as the segment information and passed into a linear fully-connected layer to obtain a segment encoding; the text is passed into a linear fully-connected layer to obtain a feature encoding of the text; finally, the position encoding, segment encoding, and feature encoding are added together to obtain an embedded representation of the text;

[0021] For the table sequence, by stacking the table data in the same row and sharing the same encoder as the text sequence, an increasing natural number sequence is used to represent their position information, which is passed into a linear fully-connected layer to obtain a position encoding; 1 is used as the segment information and passed into a linear fully-connected layer to obtain a segment encoding; the table is passed into a linear fully-connected layer to obtain a feature encoding of the table; finally, the position encoding, segment encoding, and feature encoding are added together to obtain an embedded representation of the table;

[0022] For video data, first, the S3D network is used to extract video embedded features with spatio-temporal characteristics. According to the video embedded features, a natural number sequence is used to represent their position information, the sequential relationship of different frames is passed in, and this data is applied to a linear fully-connected layer to obtain a position encoding; 1 is used as the segment information and passed into a linear fully-connected layer to obtain a segment encoding; the video feature sequence is passed into a linear fully-connected layer to obtain a feature encoding of the text; finally, the position encoding, segment encoding, and feature encoding are added together to obtain an embedded representation of the video data;

[0023] For audio data, MFCC is used to extract the frequency domain features of the audio data. For each audio feature, a natural number sequence is used to represent their position information, the sequential relationship of different frames is passed in, and this data is applied to a linear fully-connected layer to obtain a position encoding; 1 is used as the segment information and passed into a linear fully-connected layer to obtain a segment encoding; the video feature sequence is passed into a linear fully-connected layer to obtain a feature encoding of the audio; finally, the position encoding, segment encoding, and feature encoding are added together to obtain an embedded representation of the audio data.

[0024] Furthermore, the multi-modal pre-training model includes

[0025] For each modality data, a Transformer comparison learning module is constructed between different modality data to learn the semantic alignment between different modality data;

[0026] A semantically aligned public multi-head self-attention network is obtained to extract retrieval features that are fully integrated among five modal data, wherein the input length of the public multi-head self-attention network is the feature length of the stacking of each modal data.

[0027] Furthermore, the public multi-head self-attention network concatenates text, image, table, video and audio features, uses Q and K to calculate the weight of each vector paying attention to all features, and then multiplies it by V to obtain a common feature representation of the five modal data, where Q, K, and V are obtained by concatenating the features of the five modal data.

[0028] Furthermore, the public multi-head self-attention network is repeatedly iterated and trained H times.

[0029] Furthermore, in step S4, the self-supervised training is specifically as follows:

[0030] By masking some features in each modal data, the modal data with masked features are input into the multimodal pre-training model. The multimodal pre-training model learns to restore the masked features during the training process, thereby extracting a feature representation with the modal data;

[0031] The multimodal pre-training model is trained using the contrastive learning loss function. For paired image and text pairs, their distance is shortened during the training process; for unpaired image and text pairs, their distance is increased during the training process, so as to learn discriminative image and text features.

[0032] Furthermore, one or several modal data of the image, text, table, video, and audio of the product data used for training are input into the multimodal pre-training model for training, and the retrieval features extracted from the training are stored in the retrieval library.

[0033] Furthermore, for the product sample data to be queried, it is first processed through steps S1 and S2, and then input into the multimodal pre-training model trained in step S4 to extract the separate retrieval features of each modal information and the modal features after all or part of the modal fusion, respectively, calculate the similarity between the queried features of the product and the single product features, and select the most similar single product as the result to be returned.

[0034] A retrieval system based on a five-modal product pre-training method based on self-coordinated contrastive learning, including

[0035] Modal feature encoding extractor, which is used to extract feature encoding, position encoding and segment encoding of each modal data, and learn the embedding representation of different modal data;

[0036] A multi-modal pre-training model module, which is used to implement self-supervised training, perform high-level semantic fusion on various modal data, and continuously correct the correlation between modalities using a self-consistent contrast learning method, and recover the features at the corresponding positions during the learning process.

[0037] The beneficial effects of the present invention are as follows:

[0038] 1. Compared with the image retrieval method based on annotation information, the present invention is trained in a self-consistent contrast learning manner, only using the semantic alignment relationship between different modal data, and using the self-learned semantic alignment information to further constrain multi-modal contrast learning and different masking tasks during the training process of the multi-modal pre-training model. Therefore, it has strong scalability and generalization, is easy to learn a more discriminative feature representation, and improves the effect of commodity retrieval.

[0039] 2. Compared with a single-modal information retrieval system, the present invention uses information of multiple modal data, can effectively utilize the complementary information between different modal data information, fuses the five modal data features, and can make up for the problem of incomplete semantic information of single-modal data by extracting simple fusion features of different modal data.

[0040] 3. Compared with most multi-modal pre-training models that only use text and image two modalities for training, the present invention uses five modal information for self-consistent contrast learning training, solves the problem of insufficient high-level semantics in the dual-modal training process, and at the same time uses the self-consistent contrast learning method to provide important modal guidance for different modal contrast learning and the masking tasks that cannot be performed in the multi-modal data contrast learning process using high-level semantic constraints, improves the feature representation effect of the multi-modal pre-training model, and is beneficial to improving the accuracy of large-scale commodity retrieval. Brief Description of the Drawings

[0041] Figure 1 It is a flowchart of the five-modal commodity pre-training method described in Embodiment 1.

[0042] Figure 2 It is a network schematic diagram of the multi-modal pre-training model described in Embodiment 1.

[0043] Figure 3 It is a structural block diagram of the retrieval system described in Embodiment 2. Detailed Embodiments

[0044] The following describes the present invention in detail with reference to the drawings and specific embodiments.

[0045] Embodiment 1

[0046] As Figure 1As shown in the figure, a five-modal commodity pre-training method based on self-coordinated contrastive learning, and the method includes the following steps:

[0047] S1: Construct corresponding modal feature encoding extractors according to different modal data;

[0048] In a specific embodiment, the modal data includes five types of modal data: images, texts, tables, videos, and audios;

[0049] Use the bottom-up-attention network as the modal feature encoding extractor to obtain the bounding boxes of the images and the features of their coordinate positions;

[0050] Use word-piece as the modal feature encoding extractor to obtain the relationship features between different tokens of the text;

[0051] Use entity word-piece as the modal feature encoding extractor to obtain the encoded representation of the table modal data. Specifically, after splicing each row of data together, obtain the relationship features between different tokens;

[0052] Use the S3D network as the modal feature encoding extractor to obtain the video representation with spatio-temporal characteristics in the video;

[0053] Use MFCC as the modal feature encoding extractor to obtain the encoded representation of the audio modal data.

[0054] In this embodiment, for each commodity (I, T, Tab, V, A) composed of an image I, the corresponding title text T, the commodity table Tab, the commodity video V, and the commodity audio A, for the corresponding modal data of each commodity, use the bottom-up-attention network, word-piece character encoding, word-piece entity encoding, S3D network, and MFCC frequency domain as the modal feature encoding extractors to extract their corresponding feature encodings, which can be expressed as

[0055] S2: Combine the feature encodings, position encodings, and segment encodings of each modal data extracted by the modal feature encoding extractor to learn the embedding representations of different modal data.

[0056] In a specific embodiment, step S2 specifically learns the embedding representations of different modal data as follows:

[0057] For the bounding boxes output by the bottom-up-attention network and the bounding box features F=(f0, f1, f2,..., f m ), by calculating the area ratio of each box to the entire image, construct a 5-dimensional vector Using a 5 - dimensional vector, calculate the position information of each bounding box, including the upper - left coordinates, lower - right coordinates of the bounding box, and the proportion of the bounding box occupying the entire image. Pass this 5 - dimensional vector into a linear fully - connected layer to obtain the position encoding E p , and its calculation formula is as follows: where w1 and b1 are the parameters of the fully - connected layer. Pass 0 as the segment information into the linear fully - connected layer to obtain the segment encoding E s , and its calculation formula is: where w1 and b1 are the parameters of the fully - connected layer. Pass the bounding - box features into the linear fully - connected layer to obtain the encoding E of the bounding - box features f , and its calculation formula is: where w1 and b1 are the parameters of the fully - connected layer; add the position encoding, segment encoding, and feature encoding to obtain the embedded representation E1 of the image modality = E p +E s +E f , which is also expressed as E Ii =(e0,e1,e2,…,e m ).

[0058] Similarly, for text sequences, use an increasing natural - number sequence to represent their position information, pass it into a linear fully - connected layer to obtain the position encoding; use 1 as the segment information and pass it into the linear fully - connected layer to obtain the segment encoding; pass the text into the linear fully - connected layer to obtain the feature encoding of the text; finally, add the position encoding, segment encoding, and feature encoding to obtain the embedded representation E Ti =(e0,e1,e2,…,e m );

[0059] Similarly, for table sequences, stack the table data in the same row and share the same encoder as the text data. Use an increasing natural - number sequence to represent their position information, pass it into a linear fully - connected layer to obtain the position encoding; use 1 as the segment information and pass it into the linear fully - connected layer to obtain the segment encoding; pass the table into the linear fully - connected layer to obtain the feature encoding of the table; finally, add the position encoding, segment encoding, and feature encoding to obtain the embedded representation E tabi =(e0,e1,e2,…,e m );

[0060] Similarly, for video data, first use the S3D network to extract video embedding features with spatio-temporal characteristics. Represent their position information using a natural number sequence according to the video embedding features, input the sequential relationship of different frames, and apply this data to a linear fully connected layer to obtain a position encoding; input 1 as the segmentation information into the linear fully connected layer to obtain a segmentation encoding; input the video feature sequence into the linear fully connected layer to obtain a time-frequency feature encoding; finally, add the position encoding, segmentation encoding, and feature encoding to obtain the embedding representation E of the video data vi =(e0,e1,e2,…,e m );

[0061] Similarly, for audio data, use MFCC to extract the frequency domain features of the audio data. For each audio feature, represent their position information using a natural number sequence, input the sequential relationship of different frames, and apply this data to a linear fully connected layer to obtain a position encoding; input 1 as the segmentation information into the linear fully connected layer to obtain a segmentation encoding; input the audio feature sequence into the linear fully connected layer to obtain an audio feature encoding; finally, add the position encoding, segmentation encoding, and feature encoding to obtain the embedding representation E of the audio data ai =(e0,e1,e2,…,e m ).

[0062] S3: As Figure 2 shown, construct a multi-modal pre-training model for self-consistent contrastive learning;

[0063] In a specific embodiment, the multi-modal pre-training model includes

[0064] Construct a Transformer contrastive learning module between different modal data for each type of modal data, which is used to learn the semantic alignment between different modal data;

[0065] Obtain a common multi-head self-attention network with semantic alignment, which is used to extract retrieval features with comprehensive fusion among five types of modal data, where the input length of the common multi-head self-attention network is the stacked feature length of each type of modal data.

[0066] The common multi-head self-attention network splices the text, image, table, video, and audio features, calculates the weights of each vector's attention to all features using Q and K, and then multiplies by V to obtain the common feature representation of the five types of modal data, where Q, K, and V are obtained from the features spliced by the five modal data. For each type of modal data, the common multi-head self-attention network uses the multi-head attention mechanism to calculate the attention weights to all features of the five types of modal data, so as to obtain the features of each modal data after comprehensive fusion. The common multi-head self-attention network is repeatedly trained H times.

[0067] S4: The modal data with occluded features are extracted using the modal feature coder to learn the embedded representation, which is then input into the multimodal pre-training model of step S3 for self-supervised training. The modal data are semantically fused at a high level, and the self-coordinated contrast learning method is used to continuously correct the correlation between the modalities, reduce the impact of modal noise, and restore the features of the corresponding positions during the learning process.

[0068] In a specific embodiment, in step S4, the self-supervised training is specifically as follows:

[0069] By masking the words in the title text, the text sequence with the masked words is input into the multimodal pre-trained model. During the training process, the multimodal pre-trained model learns to restore the masked words, thereby extracting a feature representation with text information.

[0070] By masking the bounding box features in the image, the masked image box feature sequence is input into the multimodal pre-trained model. During the training process, the multimodal pre-trained model learns to restore the masked bounding box features, thereby extracting a feature representation with visual information.

[0071] By masking the entity words in the table, the masked table text sequence is input into the multimodal pre-training model. During the training process, the multimodal pre-training model learns to restore the masked entity word features, thereby extracting a feature representation of structured table information.

[0072] By masking the temporal embedding features in the video, the masked temporal feature sequence is input into the multimodal pre-trained model. During the training process, the multimodal pre-trained model learns to restore the masked temporal sequence features, thereby extracting a feature representation with spatial visual information.

[0073] By masking the frequency domain features in the audio data, the masked audio frequency domain sequence is input into the multimodal pre-training model. During the training process, the multimodal pre-training model learns to restore the masked frequency domain sequence features, thereby extracting a feature representation with visual information.

[0074] The multimodal pre-training model is trained using the contrastive learning loss function. For paired image and text pairs, their distance is shortened during the training process; for unpaired image and text pairs, their distance is increased during the training process, so as to learn discriminative image and text features.

[0075] Specifically, taking the text mode as an example, by learning to predict an argmax=Softmax(E t ) makes the predicted dictionary token of the position consistent with the original token, so that the model has a certain feature determination ability.

[0076] In this embodiment, a modal feature encoding extractor is used to extract encoded features from five modalities: images, text, tables, videos, and audio. Then, a multi-modal pre-training model is utilized to fully integrate the feature encodings, position encodings, segment encodings, and encoded feature representations of each modality as the input to the multi-modal pre-training model. The multi-modal pre-training model uses two types of network layers to extract retrieval features for images, text, tables, videos, audio, and their mutual fusion.

[0077] In a specific embodiment, one or several modality data of images, text, tables, videos, and audio in the commodity data for training are input into the multi-modal pre-training model for training, and the retrieval features extracted during training are stored in the retrieval library.

[0078] In a specific embodiment, for the commodity sample data to be queried, after being processed through steps S1 and S2, it is then input into the multi-modal pre-training model trained in step S4 to separately extract the retrieval features of each modality information alone and the modality features after the fusion of all or part of the modalities, calculate the similarity between the queried features of the commodity and the single-item features, and select the closest single item as the result to return.

[0079] The similarity between the queried features of the commodity and the single-item features is calculated according to the Cosine distance. The smaller the Cosine distance, the greater the similarity; and the retrieved samples that match the query are sorted from largest to smallest similarity and returned.

[0080] Embodiment 2

[0081] As Figure 3 shown, a retrieval system for a five-modal commodity pre-training method based on self-coordinated contrast learning includes

[0082] A modal feature encoding extractor, which is used to extract the feature encoding, position encoding, and segment encoding of each modality data and learn the embedding representation of different modality data;

[0083] A multi-modal pre-training model module, which is used to implement self-supervised training, perform high-level semantic fusion on each modality data, and continuously correct the relevance between modalities using the self-coordinated contrast learning method, and recover the features at the corresponding positions during the learning process.

[0084] Among them, the multi-modal pre-training model module includes a Transformer contrast learning module and a common multi-head self-attention network module;

[0085] The Transformer contrast learning module is used to learn the semantic alignment between different modality data;

[0086] The described common multi-head self-attention network module is used to extract retrieval features with comprehensive fusion among five types of modal data, where the input length of the described common multi-head self-attention network is the stacked feature length of each type of modal data.

[0087] Embodiment 3

[0088] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the implemented method steps are as follows:

[0089] S1: Construct corresponding modal feature encoding extractors according to different modal data;

[0090] S2: Combine the feature encodings, position encodings, and segment encodings of each modal data extracted by the modal feature encoding extractors to learn the embedding representations of different modal data;

[0091] S3: Construct a multi-modal pre-training model for self-consistent contrast learning;

[0092] S4: Use the modal feature encoding extractors for different modal data with occluded partial features to learn the obtained embedding representations and input them into the multi-modal pre-training model in step S3 for self-supervised training, perform high-level semantic fusion on each modal data, and continuously correct the inter-modal correlation using the self-consistent contrast learning method, and recover the features at the corresponding positions during the learning process.

[0093] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A five-modal commodity pre-training method based on self-coordinated contrast learning, characterized in that: The method described is implemented based on a retrieval system, which includes: A modal feature encoding extractor, which is used to extract the feature encoding, position encoding, and segment encoding of each modal data, and learn the embedding representation of different modal data; A multi-modal pre-training model module, which is used to implement self-supervised training, perform high-level semantic fusion on each modal data, and continuously correct the correlation between modalities using the self-consistent contrast learning method, and recover the features at the corresponding positions during the learning process; The method described includes the following steps: S1: Construct a corresponding modal feature encoding extractor according to the different modal data of the commodity sample data; the modal data includes five types of modal data: images, texts, tables, videos, and audios; Use the bottom-up-attention network as the modal feature encoding extractor to obtain the bounding boxes of the images and the features of their coordinate positions; Use word-piece as the modal feature encoding extractor to obtain the relationship features between different tokens of the text; Use the entity word-piece as the modal feature encoding extractor to obtain the encoding representation of the table modal data. Specifically, after concatenating each row of data together, obtain the relationship features between different tokens; Use the S3D network as the modal feature encoding extractor to obtain the video representation with spatio-temporal characteristics in the video; Use MFCC as the modal feature encoding extractor to obtain the encoding representation of the audio modal data; S2: Combine the feature encoding, position encoding, and segment encoding of each modal data extracted by the modal feature encoding extractor to learn the embedding representation of different modal data; S3: Construct a multi-modal pre-training model for self-consistent contrast learning; the multi-modal pre-training model includes For each type of modal data, construct a Transformer contrast learning module between different modal data to learn the semantic alignment between different modal data; Obtain a common multi-head self-attention network with semantic alignment, which is used to extract the retrieval features of the comprehensive fusion between the five types of modal data, where the input length of the common multi-head self-attention network is the stacked feature length of each type of modal data; S4: Use the modal feature encoding extractor to learn the embedding representation of the different modal data with occluded part features, and input it into the multi-modal pre-training model in step S3 for self-supervised training, perform high-level semantic fusion on each modal data, and continuously correct the correlation between modalities using the self-consistent contrast learning method, and recover the features at the corresponding positions during the learning process; For the commodity sample data to be queried, after being processed by steps S1 and S2, input it into the multi-modal pre-training model trained in step S4, extract the retrieval features of each type of modal information separately, and the modal features after the fusion of all or part of the modalities, calculate the similarity between the features queried for the commodity and the features of the single item, and select the closest single item as the result to return.

2. The five-modal commodity pre-training method based on self-coordinated contrast learning according to claim 1, characterized in that: In step S2, specifically learn the embedding representation of different modal data as follows: For the bounding boxes and bounding box features output by the bottom-up-attention network, a 5D vector is used to calculate the position information of each bounding box, including the upper left corner coordinates, the lower right corner coordinates of the bounding box, and the size ratio of the bounding box to the entire image. This 5D vector is passed into a linear fully-connected layer to obtain a position encoding; 0 is used as the segment information and passed into a linear fully-connected layer to obtain a segment encoding; the bounding box features are passed into a linear fully-connected layer to obtain an encoding of the bounding box features; the position encoding, the segment encoding, and the feature encoding are added together to obtain an embedded representation of the image modality; For the text sequence, an increasing natural number sequence is used to represent their position information, which is passed into a linear fully-connected layer to obtain a position encoding; 1 is used as the segment information and passed into a linear fully-connected layer to obtain a segment encoding; the text is passed into a linear fully-connected layer to obtain a feature encoding of the text; finally, the position encoding, the segment encoding, and the feature encoding are added together to obtain an embedded representation of the text; For the table sequence, by stacking the table data in the same row and sharing the same encoder as the text sequence, an increasing natural number sequence is used to represent their position information, which is passed into a linear fully-connected layer to obtain a position encoding; 1 is used as the segment information and passed into a linear fully-connected layer to obtain a segment encoding; the table is passed into a linear fully-connected layer to obtain a feature encoding of the table; finally, the position encoding, the segment encoding, and the feature encoding are added together to obtain an embedded representation of the table; For the video data, first, the S3D network is used to extract video embedded features with spatio-temporal features. According to the video embedded features, a natural number sequence is used to represent their position information, the sequential relationship of different frames is passed in, and this data is applied to a linear fully-connected layer to obtain a position encoding; 1 is used as the segment information and passed into a linear fully-connected layer to obtain a segment encoding; the video feature sequence is passed into a linear fully-connected layer to obtain a time-frequency feature encoding; finally, the position encoding, the segment encoding, and the feature encoding are added together to obtain an embedded representation of the video data; For the audio data, MFCC is used to extract the frequency domain features of the audio data. For each audio feature, a natural number sequence is used to represent their position information, the sequential relationship of different frames is passed in, and this data is applied to a linear fully-connected layer to obtain a position encoding; 1 is used as the segment information and passed into a linear fully-connected layer to obtain a segment encoding; the video feature sequence is passed into a linear fully-connected layer to obtain a feature encoding of the audio; finally, the position encoding, the segment encoding, and the feature encoding are added together to obtain an embedded representation of the audio data.

3. The five-modal commodity pre-training method based on self-coordinated contrast learning according to claim 2, characterized in that: The described common multi-head self-attention network concatenates the text, image, table, video, and audio features, uses Q and K to calculate the weights for each vector to attend to all features, and then multiplies by V to obtain a common feature representation of the five-modal data, where Q, K, and V are obtained from the features after concatenating the five-modal data.

4. The five-modal commodity pre-training method based on self-coordinated contrast learning according to claim 3, characterized in that: The described common multi-head self-attention network is repeatedly iteratively trained H times.

5. The five-modal commodity pre-training method based on self-coordinated contrast learning according to claim 4, characterized in that: Step S4, the self-supervised training is specifically as follows: By masking some features in each modal data, the modal data with masked features are input into the multimodal pre-training model. The multimodal pre-training model learns to restore the masked features during the training process, thereby extracting a feature representation with the modal data; Use contrastive learning loss function to train multimodal pre-trained models, shortening the distance between paired images and text during training; For unpaired image-text pairs, the distance between them is increased during the training process to learn discriminative image-text features.

6. The five-modal commodity pre-training method based on self-coordinated contrast learning according to any one of claims 1 to 5, characterized in that: One or several modal data of the product data such as images, texts, tables, videos, and audios used for training are input into the multimodal pre-training model for training, and the retrieval features extracted from the training are stored in the retrieval library.

Citation Information

Patent Citations

  • Text-to-video cross-modal retrieval method based on multistage coding

    CN111309971A

  • Commodity mounting processing method and device, retrieval processing method and device, recommendation processing method and device and training processing method and device and electronic equipment

    CN113420166A