Training method of retrieval model, retrieval method, electronic equipment, medium and product

By acquiring and enhancing the video-text pair data and building a similarity loss function, the video text search model is trained, which solves the problem of poor retrieval performance of the existing search model and significantly improves the accuracy of the search results.

CN119938983AActive Publication Date: 2025-05-06CHINA MOBILE (XIONGAN) ICT CO LTD +3
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411894689.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-06
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

The search performance of existing video text retrieval models is poor and the search results are low, mainly due to the limited training sample information.

Method used

By obtaining the original video-text pair and the corresponding enhanced video-text pair, combining the video encoder, text encoder and similarity module, the search model is trained, and the loss function is constructed using the inter-modal similarity, the first modal similarity and the second modal similarity, and the model parameters are adjusted until the loss function converges.

Benefits of technology

Through the addition of enhanced data and the construction of loss functions, the enhanced data information is fully mined, and the search performance and result accuracy of the search model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938983A_ABST
    Figure CN119938983A_ABST
Patent Text Reader

Abstract

The invention discloses a retrieval model training method, a retrieval method, electronic equipment, a medium and a product, and the method comprises the steps: obtaining an original data set which comprises a plurality of original video-text pairs and corresponding classification marks, the classification mark is used for indicating whether the video content in the original video-text pair is consistent with the text content; performing enhancement processing on the basis of original video-text pairs in the original data set to obtain corresponding enhanced video-text pairs; taking the original video-text pair and the corresponding enhanced video-text pair as samples, taking the classification mark corresponding to the original video-text as a label, and training the retrieval model; constructing a loss function of the retrieval model based on the inter-modal similarity and the intra-modal similarity calculated based on the extracted features of the video-text pair; and training the retrieval model until the loss function converges to obtain a trained retrieval model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of retrieval technology, and in particular to a retrieval model training method, a retrieval method, an electronic device, a medium and a product. Background Art

[0002] With the rapid development of mobile Internet and the widespread popularity of smart phones, video, a vivid and intuitive form of information with strong information carrying capacity and good communication effect, is gradually becoming the mainstream in the mobile Internet era and playing an increasingly important role in social life. The video-text cross-modal retrieval model can effectively build a bridge between video data and text data in different modalities. It can retrieve related texts through video data, or retrieve related videos through text data, which is of great significance to information management and digital development.

[0003] The structure of the current mainstream video text retrieval model is a dual-stream structure that uses different encoders to extract features for video and text respectively. The retrieval model calculates the similarity between the two modal data, video data and text data, based on the features extracted by the two encoders to achieve retrieval. When training the retrieval model, the training samples contain limited information, resulting in poor retrieval performance and low accuracy of retrieval results in the trained retrieval model. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a retrieval model training method, a retrieval method, an electronic device, a medium and a product to solve the problem of poor retrieval performance of the retrieval model used for video text retrieval.

[0005] In order to solve the above technical problems, this specification is implemented as follows: In a first aspect, a retrieval model training method is provided, comprising: Acquire an original data set, wherein the original data set includes a plurality of original video-text pairs and corresponding classification marks, wherein the classification marks are used to indicate whether the video content and the text content in the original video-text pairs are consistent; Based on the original video-text pairs in the original data set, enhancement processing is performed respectively to obtain corresponding enhanced video-text pairs; Taking original video-text pairs and corresponding enhanced video-text pairs as samples and taking classification marks corresponding to the original video-text pairs as labels, a retrieval model is trained, wherein the retrieval model comprises a video encoder, a text encoder and a similarity module, wherein the video encoder is used to extract video features from each input original video-text pair and the corresponding enhanced video-text pair, the text encoder is used to extract text features from each input original video-text pair and the corresponding enhanced video-text pair, and the similarity module is used to calculate inter-modal similarities between text features and video features of each extracted original video-text pair and the corresponding enhanced video-text pair, first intra-modal similarities between text features of each extracted original video-text pair and text features of the corresponding enhanced video-text pair, and second intra-modal similarities between video features of each extracted original video-text pair and video features of the corresponding enhanced video-text pair; Constructing a loss function of the retrieval model based on the inter-modality similarity, the first intra-modality similarity, and the second intra-modality similarity, wherein the loss function represents the overall loss between the inter-modality similarity, the first intra-modality similarity, and the second intra-modality similarity and the corresponding label; The parameters of the video encoder and the text encoder are adjusted based on the overall loss, and the original video-text pairs and the enhanced video-text pairs after enhancement processing in the original data set continue to be used as samples, and the corresponding classification labels are used as labels to train the retrieval model until the loss function converges to obtain the trained retrieval model.

[0006] Optionally, constructing the loss function of the retrieval model based on the inter-modality similarity, the first intra-modality similarity, and the second intra-modality similarity includes: Calculate the cross entropy loss of inter-modal similarity, the cross entropy loss of similarity within the first modality, and the cross entropy loss of similarity within the second modality for each input original video-text pair and the corresponding enhanced video-text pair respectively; Based on each cross entropy loss, a loss function of the retrieval model is constructed.

[0007] Optionally, the video encoder extracts text features of a first number of multiple original video-text pairs and a first number of corresponding enhanced video-text pairs input in batches, respectively, to obtain the first number of original text features and the first number of enhanced text features; The text encoder extracts video features of a first number of multiple original video-text pairs and a corresponding first number of enhanced video-text pairs input in batches, respectively, to obtain the first number of original video features and the first number of enhanced video features.

[0008] Optionally, the similarity module, Based on the first number of original text features, the first number of enhanced text features, the first number of original video features and the first number of enhanced video features, a first inter-modal similarity matrix between each original video feature and each original text feature, a second inter-modal similarity matrix between each enhanced video feature and each enhanced text feature, a third inter-modal similarity matrix between each enhanced video feature and each original text feature, and a fourth inter-modal similarity matrix between each original video feature and each enhanced text feature are calculated respectively, and a first intra-modal similarity matrix between each original text feature and each enhanced text feature, and a second intra-modal similarity matrix between each original video feature and each enhanced video feature are calculated respectively.

[0009] Optionally, respectively calculating the cross entropy loss of inter-modality similarity, the cross entropy loss of similarity within the first modality, and the cross entropy loss of similarity within the second modality of the input video-text pair, comprises: The cross entropy loss of the i-th row and the i-th column in the target similarity matrix is ​​calculated respectively, wherein the target similarity matrix is ​​an N*N matrix, i is between 1 and N, and N represents the first number; wherein the target similarity matrix is ​​any one of the first inter-modality similarity matrix, the second inter-modality similarity matrix, the third inter-modality similarity matrix, the fourth inter-modality similarity matrix, the first intra-modality similarity matrix, and the second intra-modality similarity matrix, and the target similarity matrix is ​​the similarity of the first number of original video-text pairs and the first number of enhanced video-text pairs input in batches.

[0010] Optionally, the cross entropy loss of the target similarity matrix is ​​calculated by the following formula:

[0011]

[0012]

[0013]

[0014]

[0015] in, represents the cross entropy loss, represents the first value of the similarity calculated between the text features of the i-th row in the target similarity matrix and the video features of the N columns in the target similarity matrix after being normalized by softmax, represents the target similarity matrix, represents the second value of the similarity calculated between the video features of the jth column in the target similarity matrix and the text features of the N columns in the target similarity matrix after being normalized by softmax, Indicates based on Calculate the first cross entropy loss of the target similarity matrix, Indicates based on Calculate the second cross entropy loss of the target similarity matrix, represents the similarity of the i-th row and N-th column in the target similarity matrix, represents the similarity of the Nth row and jth column in the target similarity matrix, Represents an exponential function.

[0016] Optionally, the loss function of constructing the retrieval model based on the inter-modality similarity, the first intra-modality similarity, and the second intra-modality similarity further includes: Calculate the InfoNCE loss of the i-th row and the i-th column in the target similarity matrix respectively; The constructing a loss function of the retrieval model based on the cross entropy loss includes: Determine the overall loss of the target similarity matrix by linearly superimposing the cross entropy loss and the InfoNCE loss of the target similarity matrix; A loss function of the retrieval model is constructed based on the overall losses of the first inter-modality similarity matrix, the second inter-modality similarity matrix, the third inter-modality similarity matrix, the fourth inter-modality similarity matrix, the first intra-modality similarity matrix, and the second intra-modality similarity matrix.

[0017] Optionally, constructing the loss function of the retrieval model based on the overall losses of the first inter-modality similarity matrix, the second inter-modality similarity matrix, the third inter-modality similarity matrix, the fourth inter-modality similarity matrix, the first intra-modality similarity matrix, and the second intra-modality similarity matrix comprises: The loss function of the retrieval model is constructed based on the average of the overall losses of the first inter-modality similarity matrix, the second inter-modality similarity matrix, the third inter-modality similarity matrix, the fourth inter-modality similarity matrix, the first intra-modality similarity matrix and the second intra-modality similarity matrix.

[0018] Optionally, the performing enhancement processing on the target original video-text pair in the original data set to obtain a target enhanced video-text pair includes: Obtaining a target enhanced video-text pair after video enhancement of the target original video-text pair by performing operations including cropping, rotation, scaling, color conversion or Gaussian blurring on the video in the target original video-text pair in the original data set; By performing operations including synonym replacement, noise addition or machine translation on the text in the target original video-text pair in the original data set, a target enhanced video-text pair with text enhancement is obtained, and the target original video-text pair is any one or more original video-text pairs in the original data set.

[0019] In a second aspect, a video text retrieval method is provided, comprising: Inputting a target text and a plurality of videos into a retrieval model for retrieval to output the similarity between the target text and each video, and retrieving a video related to the target text based on the highest similarity; or Inputting a target video and a plurality of texts into a retrieval model for retrieval to output the similarity between the target video and each text, and retrieving the text related to the target video based on the highest similarity; Wherein, the retrieval model is trained according to the method described in the first aspect.

[0020] According to a third aspect, an electronic device is provided, comprising a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and wherein the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect or the second aspect.

[0021] In a fourth aspect, a readable storage medium is provided, characterized in that a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect or the second aspect are implemented.

[0022] In a fifth aspect, a computer program product is provided, the computer program product comprising a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to execute the steps of the method described in the first aspect or the second aspect.

[0023] In an embodiment of the present application, an original data set is obtained, wherein the original data set includes a plurality of original video-text pairs and corresponding classification labels, wherein the classification labels are used to indicate whether the video content and the text content in the original video-text pairs are consistent; enhancement processing is performed on the original video-text pairs in the original data set to obtain corresponding enhanced video-text pairs; the original video-text pairs and the corresponding enhanced video-text pairs are used as samples, and the classification labels corresponding to the original video-texts are used as labels to train a retrieval model, wherein the retrieval model includes a video encoder, a text encoder, and a similarity module, wherein the video encoder is used to extract video features from each input original video-text pair and the corresponding enhanced video-text pair, the text encoder is used to extract text features from each input original video-text pair and the corresponding enhanced video-text pair, and the similarity module is used to calculate the inter-modal similarity between the text features of each extracted original video-text pair and the corresponding enhanced video-text pair and the video features, and the text features of each extracted original video-text pair and the corresponding enhanced video-text pair. The method comprises the following steps: first, a first intra-modal similarity between text features of the video-text pairs, and a second intra-modal similarity between the extracted video features of each original video-text pair and the video features of the corresponding enhanced video-text pair; constructing a loss function of the retrieval model based on the inter-modal similarity, the first intra-modal similarity and the second intra-modal similarity, wherein the loss function represents the overall loss between the inter-modal similarity, the first intra-modal similarity and the second intra-modal similarity and the corresponding label; adjusting the parameters of the video encoder and the text encoder based on the overall loss, and continuing to use the original video-text pairs in the original data set and the enhanced video-text pairs after enhancement as samples, and using the corresponding classification labels as labels to train the retrieval model until the loss function converges to obtain the trained retrieval model, thereby combining the enhanced data with the original data, and constructing a loss function based on the inter-modal and intra-modal losses of the two modal data of video and text of the combined data, so as to fully mine the information contained in the enhanced data, so that the trained retrieval model obtains better retrieval performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 It is a flowchart of the training method of the video text retrieval model in the embodiment of the present application.

[0025] Figure 2 This is an example of the inter-modal similarity calculation principle diagram of the video text retrieval model of the embodiment of the present application.

[0026] Figure 3 It is a schematic diagram of the consistency distribution of inter-modal similarity of the video text retrieval model of the embodiment of the present application.

[0027] Figure 4 This is an example of a schematic diagram of the intra-modal similarity calculation principle of the video text retrieval model according to an embodiment of the present application.

[0028] Figure 5 It is a schematic diagram of the consistency distribution of intra-modal similarity of the video text retrieval model of the embodiment of the present application.

[0029] Figure 6 It is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0030] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application. The numbering of the drawings in this application is only used to distinguish the various steps in the scheme, and is not used to limit the execution order of the various steps. The specific execution order is subject to the description in the specification.

[0031] In order to solve the problems existing in the prior art, the present application embodiment provides a method for training a video text retrieval model, such as Figure 1 As shown, the following steps 102 to 110 are included.

[0032] Step 102: Acquire an original data set, wherein the original data set includes a plurality of original video-text pairs and corresponding classification marks, wherein the classification marks are used to indicate whether the video content and the text content in the original video-text pairs are consistent.

[0033] The original video-text pair can be referenced Figure 3 , where video 12 and video 14 are original videos, text 22 and text 24 are original texts, video 12 and text 22 constitute an original video-text pair, and video 14 and text 24 constitute an original video-text pair. The classification mark of each original video-text pair is used to indicate whether the video content and text content in the corresponding original video-text pair are consistent, for example Figure 3In the original video-text pair composed of video 12 and text 22, the content of text 22 is "dancers perform traditional dance on the stage", which is consistent with the content shown in video 12. If video 12 and text 24 constitute the original video-text pair, the content of text 24 is "a man walks in the rainy forest", which is inconsistent with the content shown in video 12. The classification label can be represented by 0, 1, for example.

[0034] Step 104 , performing enhancement processing on the original video-text pairs in the original data set to obtain corresponding enhanced video-text pairs.

[0035] In one embodiment, optionally, the enhancement processing is performed on the target original video-text pair in the original data set to obtain a target enhanced video-text pair, including: performing operations including cropping, rotation, scaling, color transformation or Gaussian blur on the video in the target original video-text pair in the original data set to obtain a target enhanced video-text pair with video enhancement of the target original video-text pair; performing operations including synonym replacement, noise addition or machine translation on the text in the target original video-text pair in the original data set to obtain a target enhanced video-text pair with text enhancement of the target original video-text pair, and the target original video-text pair is any one or more original video-text pairs in the original data set.

[0036] In the embodiment of the present application, data enhancement can be performed on all or part of the original video-text pairs in the original data set. For the original video, the enhancement process can be any one of cropping, rotation, scaling, color conversion and Gaussian blur; for the original text, the enhancement process can be any one of synonym replacement, noise addition or machine translation. By performing enhancement processing on the original video or the original text, an enhanced video or enhanced text similar to the original video or the original text can be obtained, and the scale of the original data set can be expanded.

[0037] Enhanced video-text pairs can be referenced Figure 3 , wherein video 12' and video 14' are enhanced videos, and text 22' and text 24' are enhanced texts. Video 12' and video 14' can be obtained by enhancing the color transformation of corresponding video 12 and video 14, and text 22' and text 24' can be obtained by enhancing the noise addition of corresponding text 22 and text 24.

[0038] Video 12' and text 22' constitute an enhanced video-text pair, and correspond to the original video-text pair including video 12 and text 22. Similarly, video 14' and text 24' constitute an enhanced video-text pair, and correspond to the original video-text pair including video 14 and text 24. The classification label of each enhanced video-text pair is the same as the classification label of the corresponding original video-text pair.

[0039] Step 106, using the original video-text pairs and the corresponding enhanced video-text pairs as samples and the classification labels corresponding to the original video-texts as labels, the retrieval model is trained, the retrieval model includes a video encoder, a text encoder and a similarity module, the video encoder is used to extract text features from each input original video-text pair and the corresponding enhanced video-text pair, the text encoder is used to extract video features from each input original video-text pair and the corresponding enhanced video-text pair, the similarity module is used to calculate the inter-modal similarity between the extracted text features and video features of each original video-text pair and the corresponding enhanced video-text pair, the first intra-modal similarity between the extracted text features of each original video-text pair and the text features of the corresponding enhanced video-text pair, and the second intra-modal similarity between the extracted video features of each original video-text pair and the video features of the corresponding enhanced video-text pair.

[0040] In an embodiment of the present application, the original video-text pairs and the corresponding enhanced video-text pairs in the original data set are used as samples, and the classification labels corresponding to each video-text pair are used as labels to train a retrieval model, and the retrieval model is used to retrieve the similarity between the video-text pairs and the input video-text pairs.

[0041] Specifically, the retrieval model includes a dual-stream structure of a video encoder and a text encoder, which are respectively used to extract video features and text features of an input video-text pair in a one-to-one correspondence, where the input video-text pair includes an original video-text pair and a corresponding enhanced video-text pair. The similarity module calculates the similarity between the two modal data, video data and text data, included in the corresponding video-text pair based on the video features and text features extracted by the video encoder and the text encoder to achieve retrieval.

[0042] In the embodiment of the present application, when the retrieval model is trained, an original video-text pair is input into the retrieval model and the corresponding enhanced video-text pair is input. The corresponding extracted features include original text features in the original video-text pair, enhanced text features in the enhanced video-text pair, original video features in the original video-text pair, and enhanced video features in the enhanced video-text pair.

[0043] The input video-text pairs are input in batches, including positive samples and negative samples. The video-text pair corresponding to one sample has text features and video features extracted accordingly, and multiple samples input in batches at the same time have corresponding numbers of text features and video features extracted accordingly.

[0044] Optionally, the video encoder extracts text features of a first number of multiple original video-text pairs input in batches and a corresponding first number of enhanced video-text pairs, respectively, to obtain the first number of original text features and the first number of enhanced text features; the text encoder extracts video features of a first number of multiple original video-text pairs input in batches and a corresponding first number of enhanced video-text pairs, respectively, to obtain the first number of original video features and the first number of enhanced video features.

[0045] like Figure 2 As shown, for example, 4 original video-text pairs and 4 corresponding enhanced video-text pairs are input in batches, and the original video features 10 extracted by the video encoder from the corresponding original videos include 4, namely Figure 2 The enhanced video features 10' extracted from the corresponding enhanced video include 4, namely, Figure 2 V1', V2', V3', V4'. The original text features 20 extracted by the text encoder from the corresponding original text include 4, namely Figure 2 The enhanced text features 20' extracted from the corresponding enhanced text include 4, namely, Figure 2 T1', T2', T3', T4'.

[0046] Correspondingly, the similarity module calculates, based on the first number of original text features, the first number of enhanced text features, the first number of original video features and the first number of enhanced video features, respectively, a first inter-modal similarity matrix between each original video feature and each original text feature, a second inter-modal similarity matrix between each enhanced video feature and each enhanced text feature, a third inter-modal similarity matrix between each enhanced video feature and each original text feature, and a fourth inter-modal similarity matrix between each original video feature and each enhanced text feature, and respectively calculates a first intra-modal similarity matrix between each original text feature and each enhanced text feature, and a second intra-modal similarity matrix between each original video feature and each enhanced video feature.

[0047] The similarity module is used to calculate the inter-modal similarity between the text features and video features extracted from the input video-text pair, including the original text features and the original video features, the original text features and the enhanced video features, the original video features and the enhanced text features, and the enhanced video features and the enhanced text features. The similarity module is also used to calculate the intra-modal similarity between the text features of the input original video-text pair and the text features of the enhanced video-text pair, and the intra-modal similarity between the video features of the input original video-text pair and the video features of the enhanced video-text pair.

[0048] by Figure 2 Taking the input of 4 original video-text pairs and 4 enhanced video-text pairs as an example, the similarity module calculates the inter-modal similarity between the extracted original text features (T1, T2, T3, T4) and the original video features (V1, V2, V3, V4). Then, by calculating the similarity between the original text feature T1 and the 4 original video features (V1, V2, V3, V4), 4 similarities are obtained. Similarly, the original text features T2, T3, T4 are calculated with the 4 original video features, and finally a 4*4 first inter-modal similarity matrix is ​​obtained.

[0049] Similarly, the similarity module can calculate a 4*4 second modality inter-similarity matrix, a 4*4 third modality inter-similarity matrix, a 4*4 fourth modality inter-similarity matrix, a 4*4 first modality intra-similarity matrix, and a 4*4 second modality intra-similarity matrix. For details, the calculation method of the first modality inter-similarity between the original text features and the original video features is the same, which will not be repeated here.

[0050] In the specific implementation, we can first perform a normalization operation of the softmax function on each row and each column of each similarity matrix, so as to convert each similarity value distribution into a probability distribution. Then, based on the probability distribution values ​​included in the similarity matrix, a cross entropy loss is performed between the probability distributions of the corresponding rows and columns.

[0051] Step 108: construct a loss function of the retrieval model based on the inter-modality similarity, the first intra-modality similarity and the second intra-modality similarity, wherein the loss function represents the overall loss between the inter-modality similarity, the first intra-modality similarity and the second intra-modality similarity and the corresponding labels.

[0052] Optionally, constructing the loss function of the retrieval model based on the inter-modality similarity, the similarity within the first modality and the similarity within the second modality includes: respectively calculating the cross entropy loss of the inter-modality similarity, the cross entropy loss of the similarity within the first modality and the cross entropy loss of the similarity within the second modality of each input original video-text pair and the corresponding enhanced video-text pair; and constructing the loss function of the retrieval model based on each cross entropy loss.

[0053] In the embodiment of the present application, based on the four inter-modality similarity matrices and two intra-modality similarity matrices of the input original video-text pair and the enhanced video-text pair, the cross entropy loss between the features of each similarity matrix is ​​calculated. The cross entropy loss can be based on Formula calculation, Indicates the calculation As a truth value, calculate and The cross entropy loss between two probability distributions. In the embodiment of the present application, represents the probability distribution corresponding to the similarity of the i-th row in each similarity matrix, Represents the probability distribution corresponding to the similarity of the i-th column in the similarity matrix.

[0054] Optionally, the respectively calculating the cross entropy loss of the inter-modal similarity, the cross entropy loss of the similarity within the first modality, and the cross entropy loss of the similarity within the second modality of the input video-text pairs includes: respectively calculating the cross entropy loss of the i-th row and the i-th column in the target similarity matrix, the target similarity matrix is ​​an N*N matrix, i is between 1 and N, and N represents the first number; wherein the target similarity matrix is ​​any one of the first inter-modal similarity matrix, the second inter-modal similarity matrix, the third inter-modal similarity matrix, the fourth inter-modal similarity matrix, the first intra-modal similarity matrix, and the second intra-modal similarity matrix, and the target similarity matrix is ​​the similarity of the first number of original video-text pairs and the first number of enhanced video-text pairs input in batches.

[0055] Taking the above batch input of 4 original video-text pairs and 4 enhanced video-text pairs as an example, it can be seen that 6 4*4 similarity matrices 100 are calculated accordingly. Figure 2 , the similarity between the original text features T1, T2, T3, T4 and the four original video features V1, V2, V3, V4 is calculated respectively, and the 4*4 first modality similarity matrix A can be obtained. Similarly, the 4*4 second modality similarity matrix B, the third modality similarity matrix C1, the fourth modality similarity matrix C2, and Figure 4 The first intra-modality similarity matrix F and the second intra-modality similarity matrix E are shown.

[0056] according to Figure 2 The method and formula shown in the similarity matrix 100 on the right , the cross entropy loss of the first row and column, the cross entropy loss of the second row and column, the cross entropy loss of the third row and column, and the cross entropy loss of the fourth row and column in the first inter-modal similarity matrix A can be calculated respectively. The cross entropy loss of the first inter-modal similarity matrix A can be obtained by averaging the sum of these cross entropy losses.

[0057] By using the formula When calculating the cross entropy loss of each similarity matrix, and When the positions are swapped, the calculated cross entropy loss results will also be different. Therefore, in one embodiment, the calculation and The cross entropy loss before and after the swap is to calculate the probability distribution of the rows and columns of the similarity matrix twice as the predicted probability and the true value respectively.

[0058] Optionally, the similarity module calculates the cross entropy loss of the target similarity matrix by the following formula:

[0059]

[0060]

[0061]

[0062]

[0063] in, represents the cross entropy loss, represents the first value of the similarity calculated between the text features of the i-th row in the target similarity matrix and the video features of the N columns in the target similarity matrix after being normalized by softmax, represents the target similarity matrix, represents the second value of the similarity calculated between the video features of the jth column in the target similarity matrix and the text features of the N columns in the target similarity matrix after being normalized by softmax, Indicates based on Calculate the first cross entropy loss of the target similarity matrix, Indicates based on Calculate the second cross entropy loss of the target similarity matrix, represents the similarity of the i-th row and N-th column in the target similarity matrix, represents the similarity of the Nth row and jth column in the target similarity matrix, Represents an exponential function.

[0064] From the above formula, we can see that the cross entropy loss of the similarity matrix is The i-th row and i-th column in the similarity matrix are respectively and , the average of the cross entropy losses before and after the swap.

[0065] The cross entropy loss of the six similarity matrices can be calculated in the above way. The cross entropy loss can effectively measure the difference between two probability distributions. When the similarity gap between a row and a column in the similarity matrix is ​​large, the cross entropy loss between the two probability distributions will also be large. Then, the gradient of the loss to the retrieval model parameters is calculated and the parameters of the model are optimized using the optimization algorithm. Finally, the retrieval model can continuously narrow the distance between the similarity of this row and this column, so that the similarity matrix has better symmetry along the main diagonal direction, which means that the feature space has better consistency for data of two different modalities. By calculating the symmetric consistency loss of the above cross entropy loss between the rows and columns of the similarity matrix, the discrimination between negative samples can be introduced, and this discrimination can be effectively constrained, which helps the retrieval model to fully utilize and learn to train a large amount of information containing the relationship between negative samples, thereby improving the retrieval performance of the trained retrieval model.

[0066] refer to Figure 3 ,The line segment connecting the video and the text represents the distance between the two ,the shorter the line segment, the closer the distance and the higher the similarity. Figure 3 It presents the symmetric consistency between the ideal enhanced data modalities. Figure 3 The solid line in the middle represents the original data, including the original video and original text; the dotted line represents the enhanced data, including the enhanced video and enhanced text. Figure 3 It can be seen that the distance between video 12 and text 24 of text 22 is basically the same as the distance between video 14 and text 22 of text 24. Figure 3 It can be seen that the distance between the video 12' of the text 22 and the text 24 is substantially the same as the distance between the video 14' of the text 24 and the text 22. Other details will not be described one by one.

[0067] refer to Figure 5 ,The line segments connecting videos and texts represent the distance between them. The shorter the line segment, the closer the distance and the higher similarity. Figure 5 It presents the ideal situation of enhancing the symmetric consistency within the data modality, from Figure 5It can be seen that the distance between video 12 and video 14' corresponding to video 12' is basically the same as the distance between video 14 and video 12' corresponding to video 14'; and the distance between text 22' and text 24' corresponding to text 22' is basically the same as the distance between text 24 and text 22' corresponding to text 24'.

[0068] Using cross entropy loss to construct the loss function of the retrieval model for training can effectively help the retrieval model learn the consistency of different modal data in the feature space.

[0069] In one embodiment, the loss function of the retrieval model constructed based on the inter-modality similarity, the first intra-modality similarity and the second intra-modality similarity also includes: respectively calculating the InfoNCE loss of the i-th row and the i-th column in the target similarity matrix; the loss function of the retrieval model constructed based on the cross entropy loss includes: determining the overall loss of the target similarity matrix by linearly superimposing the cross entropy loss and the InfoNCE loss of the target similarity matrix; constructing the loss function of the retrieval model based on the overall losses of the first inter-modality similarity matrix, the second inter-modality similarity matrix, the third inter-modality similarity matrix, the fourth inter-modality similarity matrix, the first intra-modality similarity matrix and the second intra-modality similarity matrix.

[0070] The above embodiment further combines the InfoNCE loss on the basis of the cross entropy loss. The InfoNCE loss is a loss function used for contrastive learning. It maximizes the similarity of similar sample pairs and minimizes the similarity of dissimilar sample pairs. In the embodiment of the present application, the similarity module calculates 6 similarity matrices in the above manner, and calculates the InfoNCE loss between the i-th row and the i-th column of each similarity matrix. For example, Figure 2 Taking the 4*4 similarity matrix 100 as an example, the InfoNCE loss between the corresponding rows and columns of the similarity matrix 100 and the preset matrix 200 is calculated respectively. The preset matrix 200 is a similarity matrix close to the ideal, for example, the distribution probability of the similarity at the diagonal position of the preset matrix is ​​1, and the distribution probability of the similarity at other positions is 0.

[0071] The InfoNCE loss of the first row of the similarity matrix 100 and the first column of the preset matrix 200, the InfoNCE loss of the second row of the similarity matrix 100 and the second column of the preset matrix 200, the InfoNCE loss of the third row of the similarity matrix 100 and the third column of the preset matrix 200, and the InfoNCE loss of the fourth row of the similarity matrix 100 and the fourth column of the preset matrix 200 are calculated respectively. The InfoNCE loss of the similarity matrix 100 is obtained by adding and averaging these InfoNCE losses.

[0072] After obtaining the InfoNCE loss, the cross entropy loss of the target similarity matrix is ​​combined to determine the overall loss of the target similarity matrix. For example, the target similarity matrix The overall loss is:

[0073] in, represents the cross entropy loss, represents the InfoNCE loss, and are preset coefficients respectively, so that the overall loss of each similarity matrix can be represented by linearly superimposing the cross entropy loss and the InfoNCE loss. The target similarity matrix here is any one of the similarity matrix between the first modality, the similarity matrix between the second modality, the similarity matrix between the third modality, the similarity matrix between the fourth modality, the similarity matrix within the first modality, and the similarity matrix within the second modality.

[0074] Based on the 6 overall losses corresponding to the 6 similarity matrices, the final loss function can be constructed.

[0075] Specifically, the loss function of the retrieval model is constructed based on the overall losses of the first inter-modal similarity matrix, the second inter-modal similarity matrix, the third inter-modal similarity matrix, the fourth inter-modal similarity matrix, the first intra-modal similarity matrix and the second intra-modal similarity matrix, including: constructing the loss function of the retrieval model based on the average of the overall losses of the first inter-modal similarity matrix, the second inter-modal similarity matrix, the third inter-modal similarity matrix, the fourth inter-modal similarity matrix, the first intra-modal similarity matrix and the second intra-modal similarity matrix.

[0076] For example, The first inter-modality similarity matrix, the second inter-modality similarity matrix, the third inter-modality similarity matrix, the fourth inter-modality similarity matrix, the first intra-modality similarity matrix, and the second intra-modality similarity matrix are represented one by one. Then the loss function of the retrieval model is It is expressed as follows:

[0077] By The overall losses corresponding to these six similarity matrices are added together and averaged to obtain the loss function .

[0078] Step 110, adjust the parameters of the video encoder and the text encoder based on the overall loss, continue to use the original video-text pairs in the original data set and the enhanced video-text pairs after enhancement processing as samples, use the corresponding classification labels as labels, and train the retrieval model until the loss function converges to obtain the trained retrieval model.

[0079] In one embodiment, the process may be stopped when the loss function converges to complete the retrieval model training, or may be stopped when the training iterations reach a preset number.

[0080] In an embodiment of the present application, an original data set is obtained, wherein the original data set includes a plurality of original video-text pairs and corresponding classification labels, wherein the classification labels are used to indicate whether the video content and the text content in the original video-text pairs are consistent; enhancement processing is performed on the original video-text pairs in the original data set to obtain corresponding enhanced video-text pairs; the original video-text pairs and the corresponding enhanced video-text pairs are used as samples, and the classification labels corresponding to the original video-texts are used as labels to train a retrieval model, wherein the retrieval model includes a video encoder, a text encoder, and a similarity module, wherein the video encoder is used to extract video features from each input original video-text pair and the corresponding enhanced video-text pair, the text encoder is used to extract text features from each input original video-text pair and the corresponding enhanced video-text pair, and the similarity module is used to calculate the inter-modal similarity between the text features of each extracted original video-text pair and the corresponding enhanced video-text pair and the video features, and the text features of each extracted original video-text pair and the corresponding enhanced video-text pair. The method comprises the following steps: first, a first intra-modal similarity between text features of the video-text pairs, and a second intra-modal similarity between the extracted video features of each original video-text pair and the video features of the corresponding enhanced video-text pair; constructing a loss function of the retrieval model based on the inter-modal similarity, the first intra-modal similarity and the second intra-modal similarity, wherein the loss function represents the overall loss between the inter-modal similarity, the first intra-modal similarity and the second intra-modal similarity and the corresponding label; adjusting the parameters of the video encoder and the text encoder based on the overall loss, and continuing to use the original video-text pairs in the original data set and the enhanced video-text pairs after enhancement as samples, and using the corresponding classification labels as labels to train the retrieval model until the loss function converges to obtain the trained retrieval model, thereby combining the enhanced data with the original data, and constructing a loss function based on the inter-modal and intra-modal losses of the two modal data of video and text of the combined data, so as to fully mine the information contained in the enhanced data, so that the trained retrieval model obtains better retrieval performance.

[0081] Optionally, an embodiment of the present application also provides a video text retrieval method, comprising: inputting a target text and multiple videos into a retrieval model for retrieval to output the similarity between the target text and each video, and retrieving videos related to the target text based on the highest similarity; or inputting a target video and multiple texts into a retrieval model for retrieval to output the similarity between the target video and each text, and retrieving texts related to the target video based on the highest similarity; wherein the retrieval model is trained according to any of the above-mentioned training methods for the video text retrieval model.

[0082] Alternatively, if Figure 6 As shown, an embodiment of the present application further provides an electronic device 2000, including a processor 2400 and a memory 2200, wherein the memory 2200 stores programs or instructions that can be executed on the processor 2400, and when the program or instructions are executed by the processor 2400, the various steps of the training method embodiment of the retrieval model of the video text or the retrieval method embodiment of the video text are implemented, and the same technical effect can be achieved. To avoid repetition, they are not described here.

[0083] The embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the training method embodiment of any of the above-mentioned video text retrieval model or the video text retrieval method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. Among them, the readable storage medium includes a computer-readable storage medium, such as a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.

[0084] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to enable a computer to implement the various processes of any of the above-mentioned video text retrieval model training method embodiments or video text retrieval method embodiments when executed, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0085] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0086] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0087] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A training method for a retrieval model, characterized in that: include: Acquire an original data set, wherein the original data set includes a plurality of original video-text pairs and corresponding classification marks, wherein the classification marks are used to indicate whether the video content and the text content in the original video-text pairs are consistent; Performing enhancement processing on the original video-text pairs in the original data set to obtain corresponding enhanced video-text pairs; Taking original video-text pairs and corresponding enhanced video-text pairs as samples and taking classification marks corresponding to the original video-text pairs as labels, a retrieval model is trained, wherein the retrieval model comprises a video encoder, a text encoder and a similarity module, wherein the video encoder is used to extract video features from each input original video-text pair and the corresponding enhanced video-text pair, the text encoder is used to extract text features from each input original video-text pair and the corresponding enhanced video-text pair, and the similarity module is used to calculate inter-modal similarities between text features and video features of each extracted original video-text pair and the corresponding enhanced video-text pair, first intra-modal similarities between text features of each extracted original video-text pair and text features of the corresponding enhanced video-text pair, and second intra-modal similarities between video features of each extracted original video-text pair and video features of the corresponding enhanced video-text pair; Constructing a loss function of the retrieval model based on the inter-modality similarity, the first intra-modality similarity, and the second intra-modality similarity, wherein the loss function represents the overall loss between the inter-modality similarity, the first intra-modality similarity, and the second intra-modality similarity and the corresponding label; The parameters of the video encoder and the text encoder are adjusted based on the overall loss, and the original video-text pairs and the enhanced video-text pairs after enhancement processing in the original data set continue to be used as samples, and the corresponding classification labels are used as labels to train the retrieval model until the loss function converges to obtain the trained retrieval model.

2. The method according to claim 1, characterized in that The loss function of the retrieval model constructed based on the inter-modality similarity, the first intra-modality similarity and the second intra-modality similarity includes: Calculate the cross entropy loss of inter-modal similarity, the cross entropy loss of similarity within the first modality, and the cross entropy loss of similarity within the second modality for each input original video-text pair and the corresponding enhanced video-text pair respectively; Based on each cross entropy loss, a loss function of the retrieval model is constructed.

3. The method according to claim 2, characterized in that The video encoder extracts text features of a first number of multiple original video-text pairs and a first number of corresponding enhanced video-text pairs input in batches, respectively, to obtain the first number of original text features and the first number of enhanced text features; The text encoder extracts video features of a first number of multiple original video-text pairs and a corresponding first number of enhanced video-text pairs input in batches, respectively, to obtain the first number of original video features and the first number of enhanced video features.

4. The method according to claim 3, characterized in that The similarity module, Based on the first number of original text features, the first number of enhanced text features, the first number of original video features and the first number of enhanced video features, a first inter-modal similarity matrix between each original video feature and each original text feature, a second inter-modal similarity matrix between each enhanced video feature and each enhanced text feature, a third inter-modal similarity matrix between each enhanced video feature and each original text feature, and a fourth inter-modal similarity matrix between each original video feature and each enhanced text feature are calculated respectively, and a first intra-modal similarity matrix between each original text feature and each enhanced text feature, and a second intra-modal similarity matrix between each original video feature and each enhanced video feature are calculated respectively.

5. The method according to claim 4, characterized in that The step of respectively calculating the cross entropy loss of the inter-modality similarity of the input video-text pair, the cross entropy loss of the similarity within the first modality, and the cross entropy loss of the similarity within the second modality includes: The cross entropy loss of the i-th row and the i-th column in the target similarity matrix is ​​calculated respectively, wherein the target similarity matrix is ​​an N*N matrix, i is between 1 and N, and N represents the first number; wherein the target similarity matrix is ​​any one of the first inter-modality similarity matrix, the second inter-modality similarity matrix, the third inter-modality similarity matrix, the fourth inter-modality similarity matrix, the first intra-modality similarity matrix, and the second intra-modality similarity matrix, and the target similarity matrix is ​​the similarity of the first number of original video-text pairs and the first number of enhanced video-text pairs input in batches.

6. The method according to claim 5, characterized in that The cross entropy loss of the target similarity matrix is ​​calculated by the following formula: in, represents the cross entropy loss, represents the first value of the similarity calculated between the text features of the i-th row in the target similarity matrix and the video features of the N columns in the target similarity matrix after being normalized by softmax, represents the target similarity matrix, represents the second value of the similarity calculated between the video features of the jth column in the target similarity matrix and the text features of the N columns in the target similarity matrix after being normalized by softmax, Indicates based on Calculate the first cross entropy loss of the target similarity matrix, Indicates based on Calculate the second cross entropy loss of the target similarity matrix, represents the similarity of the i-th row and N-th column in the target similarity matrix, represents the similarity of the Nth row and jth column in the target similarity matrix, Represents an exponential function.

7. The method according to claim 5 or 6, characterized in that: The loss function of the retrieval model constructed based on the inter-modality similarity, the first intra-modality similarity and the second intra-modality similarity also includes: Calculate the InfoNCE loss of the i-th row and the i-th column in the target similarity matrix respectively; The constructing a loss function of the retrieval model based on the cross entropy loss includes: Determine the overall loss of the target similarity matrix by linearly superimposing the cross entropy loss and the InfoNCE loss of the target similarity matrix; A loss function of the retrieval model is constructed based on the overall losses of the first inter-modality similarity matrix, the second inter-modality similarity matrix, the third inter-modality similarity matrix, the fourth inter-modality similarity matrix, the first intra-modality similarity matrix, and the second intra-modality similarity matrix.

8. The method according to claim 7, characterized in that The constructing the loss function of the retrieval model based on the overall losses of the first inter-modality similarity matrix, the second inter-modality similarity matrix, the third inter-modality similarity matrix, the fourth inter-modality similarity matrix, the first intra-modality similarity matrix, and the second intra-modality similarity matrix comprises: The loss function of the retrieval model is constructed based on the average of the overall losses of the first inter-modality similarity matrix, the second inter-modality similarity matrix, the third inter-modality similarity matrix, the fourth inter-modality similarity matrix, the first intra-modality similarity matrix and the second intra-modality similarity matrix.

9. The method according to claim 1, characterized in that: The enhancing process is performed based on the target original video-text pair in the original data set to obtain the target enhanced video-text pair, including: Obtaining a target enhanced video-text pair after video enhancement of the target original video-text pair by performing operations including cropping, rotation, scaling, color conversion or Gaussian blurring on the video in the target original video-text pair in the original data set; By performing operations including synonym replacement, noise addition or machine translation on the text in the target original video-text pair in the original data set, a target enhanced video-text pair with text enhancement is obtained, and the target original video-text pair is any one or more original video-text pairs in the original data set.

10. A video text retrieval method, characterized in that: include: Inputting a target text and a plurality of videos into a retrieval model for retrieval to output the similarity between the target text and each video, and retrieving a video related to the target text based on the highest similarity; or Inputting a target video and a plurality of texts into a retrieval model for retrieval to output the similarity between the target video and each text, and retrieving the text related to the target video based on the highest similarity; Wherein, the retrieval model is trained according to any one of the methods of claims 1-9.

11. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 9 or claim 10 are implemented.

12. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 9 or claim 10 are implemented.

13. A computer program product, characterized in that The computer program product comprises a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to perform the steps of the method according to any one of claims 1 to 9 or claim 10.

Citation Information

Patent Citations

  • Video text matching model training method and device and video text matching method and device

    CN115204301A

  • Cross-modal video text retrieval method, system and equipment and medium

    CN116910307A

  • Cross-modal video-text hash retrieval method based on supervised comparative learning

    CN117743639A

  • Cross-attention system and method for fast video-text retrieval task with image clip

    WO2022261570A1

  • Methods and systems for correlating video and text

    WO2023091507A1