Data processing method, device and equipment and readable storage medium

Through multimodal feature extraction and search instruction generation, the problem of low media tag search accuracy in the prior art is solved, and more efficient and accurate media tag search is achieved.

CN120353944APending Publication Date: 2025-07-22TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510510750.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing media tag prediction method has low retrieval accuracy because the tag only has text characteristics, and cannot effectively utilize the rich information of multimedia content.

Method used

By extracting multimodal feature of media data, generating modal search instructions, and performing feature search and tag fusion in multiple search data sets, generating target search results, and indirect search using multimodal features to avoid the information limitation of directly searching tags.

Benefits of technology

It improves the accuracy of media tag retrieval, reduces the risk of tag mismatch, reduces training costs and system complexity, and improves the efficiency and accuracy of retrieval tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353944A_ABST
    Figure CN120353944A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method, device and equipment and a readable storage medium, and the method comprises the steps: carrying out the feature extraction of media data, and obtaining M modal feature vectors; modal pair retrieval instructions corresponding to the N retrieval data sets are generated for each modal feature vector, and S modal pair retrieval instructions are obtained; performing feature retrieval in the N retrieval data sets based on modal feature vectors respectively contained in the S modal pair retrieval instructions to obtain retrieval results respectively corresponding to the S modal pair retrieval instructions; one retrieval result is obtained by performing feature retrieval on a modal feature vector contained in the retrieval instruction in a retrieval data set corresponding to a contained modal identifier based on one modal; and performing label fusion processing on the S retrieval results to obtain a target retrieval result. By adopting the method and the device, the accuracy of the retrieval result for the media tag can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular, to a data processing method, apparatus, device, and readable storage medium. Background Art

[0002] With the rapid growth of multimedia content, how to effectively predict media content tags has become an important research topic. Existing tag prediction often finds the set of tags most relevant to the input data among a large number of candidate tags. First, the features of the input data and the tags are extracted separately, and then the similarity between the features of the input data and the tags is calculated to select tags. Since tags are often words or short sentences and only have text features with limited information, the accuracy of tag retrieval is low. Summary of the Invention

[0003] Embodiments of the present application provide a data processing method, apparatus, device, and readable storage medium, which can improve the accuracy of retrieval results for media tags.

[0004] On the one hand, an embodiment of the present application provides a data processing method, including:

[0005] Obtain media data, perform feature extraction on the media data to obtain M modal feature vectors; M is a positive integer, and the modalities of the M modal feature vectors are different from each other;

[0006] Generate modality-pair retrieval instructions corresponding to N retrieval data sets for each modal feature vector, to obtain S modality-pair retrieval instructions; N is a positive integer, S is the product of M and N; the N retrieval data sets have different modality identifiers, the modalities of the retrieval data in different retrieval data sets are different from each other, and each modality-pair retrieval instruction includes a modal feature vector and a modality identifier;

[0007] Based on the modal feature vectors respectively included in the S modality-pair retrieval instructions, perform feature retrieval in the N retrieval data sets to obtain the retrieval results corresponding to the S modality-pair retrieval instructions respectively; a retrieval result is obtained by performing feature retrieval in the retrieval data set corresponding to the modality identifier included in a modality-pair retrieval instruction based on the modal feature vector included in the modality-pair retrieval instruction; a retrieval result includes the media tags carried by the retrieval data matched through feature retrieval;

[0008] Perform tag fusion processing on the S retrieval results to obtain a target retrieval result.

[0009] Among them, the M modal feature vectors include modal feature vector A i , where i is a positive integer less than or equal to M; generating modality-pair retrieval instructions corresponding to N retrieval data sets for each modal feature vector to obtain S modality-pair retrieval instructions includes:

[0010] Generate modal feature vector A i Retrieval instructions for N modalities of N retrieval datasets; the modal feature vectors included in the N-modal retrieval instructions are all modal feature vector A i , and the retrieval datasets associated with the N-modal retrieval instructions are different from each other;

[0011] Determine the N-modal retrieval instructions corresponding to the M modal feature vectors as S modal retrieval instructions.

[0012] Among them, the S modal retrieval instructions include modal retrieval instruction B j , where j is a positive integer less than or equal to S; based on the modal feature vectors included in the S modal retrieval instructions respectively, perform feature retrieval in the N retrieval datasets to obtain the retrieval results corresponding to the S modal retrieval instructions respectively, including:

[0013] Take the modal retrieval instruction B j The retrieval dataset associated with the modal identifier included therein is determined as the target retrieval dataset;

[0014] Obtain the retrieval feature vectors corresponding to the R retrieval data in the target retrieval dataset, and compare the modal feature vectors included in the modal retrieval instruction B j with the R retrieval feature vectors respectively to obtain R retrieval scores; R is a positive integer; one retrieval data includes one or more labeled media tags;

[0015] Sort the R retrieval scores, obtain K retrieval scores from the sorted R retrieval scores, and determine the media tags of the retrieval data corresponding to the K retrieval scores respectively as the retrieval result corresponding to the modal retrieval instruction B j , where K is a positive integer less than or equal to R.

[0016] Among them, perform label fusion processing on the S retrieval results to obtain the target retrieval result, including:

[0017] Count the number of the same media tags in the S retrieval results to obtain S statistical results;

[0018] Determine the media tags corresponding to the statistical results in the S statistical results where the number of media tags is greater than or equal to the quantity threshold as the target retrieval result.

[0019] Among them, the S retrieval results jointly include G mutually different media tags, where G is a positive integer; perform label fusion processing on the S retrieval results to obtain the target retrieval result, including:

[0020] Accumulate the retrieval scores of the same media tags in the S retrieval results to obtain the total retrieval scores corresponding to each of the G media tags;

[0021] Determine the media tags with total retrieval scores greater than or equal to the score threshold among the G media tags as the target retrieval results.

[0022] Among them, it also includes:

[0023] Obtain tag hint words, and input the media data, target retrieval results, and tag hint words into the sorting model; the tag hint words are used to instruct the sorting model to generate the matching verification results between the media data and the target retrieval results;

[0024] Extract features from the media data through the sorting model to obtain a mixed-modal feature vector, and generate similarity scores corresponding to each of the F sorting contents in the sorting model based on the mixed-modal feature vector; F is a positive integer, and one sorting content includes one or more labeled media tags;

[0025] Generate matching verification results corresponding to each media tag in the target retrieval results based on the similarity scores corresponding to the F sorting contents, the tag hint words, and the target retrieval results; the matching verification results are used to represent whether the media data is consistent or inconsistent with the media tags in the target retrieval results;

[0026] Determine the media tags corresponding to the matching verification results indicating consistency as the tags to be sorted, and generate relevance scores for the tags to be sorted based on the similarity scores corresponding to the sorting contents with the tags to be sorted among the F sorting contents;

[0027] Sort the tags to be sorted in the target retrieval results based on the relevance scores to obtain the sorting results.

[0028] Among them, extracting features from the media data to obtain M modal feature vectors includes:

[0029] Input the media data into the mixed-modal retrieval model; the mixed-modal retrieval model includes an encoding layer;

[0030] Extract features from the media data through the encoding layer to obtain P single-modal feature vectors; P is a positive integer less than or equal to M;

[0031] If the media data is video data, perform feature fusion processing on the P single-modal feature vectors to obtain a multi-modal feature vector, and determine the P single-modal feature vectors and the multi-modal feature vector as the M modal feature vectors;

[0032] If the media data is text data or image data, determine the P single-modal feature vectors as the M modal feature vectors.

[0033] Among them, the encoding layer includes P single-modal encoders with different modalities; the P single-modal data are respectively subjected to feature extraction through the encoding layer to obtain P single-modal feature vectors, including:

[0034] The media data are split according to the media type of the media data to obtain Q single-modal data with different media types; Q is a positive integer less than or equal to P;

[0035] The Q single-modal data are respectively input into the single-modal encoders matching the media types of the Q single-modal data. In the Q single-modal encoders, the Q single-modal data are respectively subjected to feature extraction to obtain Q single-modal feature vectors;

[0036] If Q is less than P, padding feature vectors corresponding to the single-modal encoders other than the Q single-modal encoders among the P single-modal encoders are generated, and the Q single-modal feature vectors and the padding feature vectors are determined as P single-modal feature vectors;

[0037] If Q is equal to P, the Q single-modal feature vectors are determined as P single-modal feature vectors.

[0038] Among them, the hybrid-modal retrieval model further includes an alignment processing layer; determining the Q single-modal feature vectors and the padding feature vectors as P single-modal feature vectors includes:

[0039] Inputting the Q single-modal feature vectors and the padding feature vectors into the alignment processing layer;

[0040] Based on the alignment projection matrix corresponding to the alignment processing layer, the Q single-modal feature vectors and the padding feature vectors are respectively linearly transformed to obtain P alignment feature vectors;

[0041] The P alignment feature vectors are determined as P single-modal feature vectors.

[0042] Among them, the Q single-modal data include image data, and the Q single-modal encodings include visual encoders corresponding to the image type; in the Q single-modal encoders, the Q single-modal data are respectively subjected to feature extraction to obtain Q single-modal feature vectors, including:

[0043] In the visual encoder, the image data is segmented to obtain T unit sampled images; T is a positive integer;

[0044] Feature extraction is performed on the T unit sampled images to obtain T unit visual feature vectors;

[0045] Based on the rearrangement factor, the vector elements representing the spatial channels in the T unit visual feature vectors are remapped to the feature channels to obtain T visual feature vectors to be matched; the rearrangement factor is determined based on the number T of unit visual feature vectors, and the vector length of the unit visual feature vector is greater than the vector length of the visual feature vector to be matched;

[0046] Perform attention processing on the T visual feature vectors to be matched to obtain a global visual feature vector, and determine the global visual feature vector as the single-modal feature vector corresponding to the image data.

[0047] Among them, performing attention processing on the T visual feature vectors to be matched to obtain a global visual feature vector includes:

[0048] Obtain the query parameter matrix, key parameter matrix, and value parameter matrix in the visual encoder; the query parameter matrix, key parameter matrix, and value parameter matrix are all matrices composed of learnable parameters;

[0049] Perform vector concatenation on the T visual feature vectors to be matched to obtain a visual feature sequence to be matched, perform a dot product operation on the visual feature sequence to be matched and the query parameter matrix to obtain a query vector, perform a dot product operation on the visual feature sequence to be matched and the key parameter matrix to obtain a key vector, and perform a dot product operation on the visual feature sequence to be matched and the value parameter matrix to obtain a value vector;

[0050] Generate an attention score vector based on the query vector and the key vector, perform dimensionality reduction processing on the attention score vector based on the vector length of the key vector, perform normalization processing on the dimensionality-reduced attention score vector to obtain an attention weight vector, and perform a dot product operation on the attention weight vector and the value vector to obtain a global visual feature vector.

[0051] On the one hand, an embodiment of the present application provides another data processing method, including:

[0052] Obtain sample media data and input the sample media data into the initial modal retrieval model;

[0053] Through the initial modal retrieval model, perform feature extraction on the sample media data to obtain M sample modal feature vectors; M is a positive integer, and the modalities of the M sample modal feature vectors are different from each other;

[0054] Generate sample modal pair retrieval instructions corresponding to N retrieval data sets for each sample modal feature vector to obtain S sample modal pair retrieval instructions; N is a positive integer, and S is the product of M and N; the N retrieval data sets have different modal identifiers, the retrieval data in different retrieval data sets have different modalities, and each sample modal pair retrieval instruction includes a sample modal feature vector and a modal identifier;

[0055] Based on the sample modal feature vectors included in the retrieval instructions for S sample modalities, perform feature retrieval in N retrieval data sets to obtain the sample retrieval results corresponding to the retrieval instructions for S sample modalities respectively; a sample retrieval result is obtained by performing feature retrieval on the sample modal feature vectors included in the retrieval instructions for a sample modality in the retrieval data set corresponding to the included modal identifier; a sample retrieval result includes the media tags carried by the retrieved data matched through feature retrieval.

[0056] Perform label fusion processing on the S sample retrieval results to obtain the target sample retrieval result, and based on the target sample retrieval result and the sample labels of the sample media data, adjust the initial modal retrieval model to obtain the hybrid modal retrieval model; the hybrid modal retrieval model is used to generate the target retrieval result of the media data.

[0057] One aspect of the embodiments of the present application provides a data processing device, including:

[0058] A feature extraction module, configured to obtain media data and perform feature extraction on the media data to obtain M modal feature vectors; M is a positive integer, and the modalities of the M modal feature vectors are different from each other;

[0059] An instruction generation module, configured to generate modal pair retrieval instructions corresponding to N retrieval data sets for each modal feature vector to obtain S modal pair retrieval instructions; N is a positive integer, S is the product of M and N; the N retrieval data sets have different modal identifiers, the modalities of the retrieval data in different retrieval data sets are different from each other, and each modal pair retrieval instruction includes a modal feature vector and a modal identifier;

[0060] An instruction retrieval module, configured to perform feature retrieval on the modal feature vectors included in the S modal pair retrieval instructions in N retrieval data sets to obtain the retrieval results corresponding to the S modal pair retrieval instructions respectively; a retrieval result is obtained by performing feature retrieval on the modal feature vector included in a modal pair retrieval instruction in the retrieval data set corresponding to the included modal identifier; a retrieval result includes the media tags carried by the retrieved data matched through feature retrieval.

[0061] A fusion processing module, configured to perform label fusion processing on the S retrieval results to obtain the target retrieval result.

[0062] In a possible implementation manner, the M modal feature vectors include modal feature vector A i , where i is a positive integer less than or equal to M; when the instruction generation module generates modal pair retrieval instructions corresponding to N retrieval data sets for each modal feature vector to obtain S modal pair retrieval instructions, it is specifically configured to perform the following operations:

[0063] Generate modal feature vector A i Retrieval instructions for N modalities of N retrieval data sets; the modal feature vectors included in the N-modal retrieval instructions are all modal feature vector A i , and the retrieval data sets associated with the N-modal retrieval instructions are different from each other;

[0064] Determine the N-modal retrieval instructions corresponding to the M modal feature vectors as S modal retrieval instructions.

[0065] In a possible implementation, the S modal retrieval instructions include modal retrieval instruction B j , where j is a positive integer less than or equal to S; when the instruction retrieval module is used to perform feature retrieval in the N retrieval data sets based on the modal feature vectors respectively included in the S modal retrieval instructions and obtain the retrieval results corresponding to the S modal retrieval instructions, it is specifically used to perform the following operations:

[0066] Determine the retrieval data set associated with the modal identifier included in modal retrieval instruction B j as the target retrieval data set;

[0067] Obtain the retrieval feature vectors corresponding to the R retrieval data in the target retrieval data set, and compare the modal feature vectors included in modal retrieval instruction B j with the R retrieval feature vectors respectively to obtain R retrieval scores; R is a positive integer; one retrieval data includes one or more labeled media tags;

[0068] Sort the R retrieval scores, obtain K retrieval scores from the sorted R retrieval scores, and determine the media tags of the retrieval data corresponding to the K retrieval scores as the retrieval result corresponding to modal retrieval instruction B j , where K is a positive integer less than or equal to R.

[0069] In a possible implementation, when the fusion processing module is used to perform label fusion processing on the S retrieval results to obtain the target retrieval result, it is specifically used to perform the following operations:

[0070] Count the number of the same media tags in the S retrieval results to obtain S statistical results;

[0071] Determine the media tags corresponding to the statistical results in the S statistical results where the number of media tags is greater than or equal to the quantity threshold as the target retrieval result.

[0072] In a possible implementation, the S retrieval results together include G mutually distinct media tags, where G is a positive integer; when the fusion processing module is used to perform tag fusion processing on the S retrieval results to obtain the target retrieval result, it is specifically used to perform the following operations:

[0073] Accumulate the retrieval scores of the same media tags in the S retrieval results to obtain the total retrieval scores corresponding to the G media tags respectively;

[0074] Determine the media tags in the G media tags whose total retrieval scores are greater than or equal to the score threshold as the target retrieval result.

[0075] In a possible implementation, the fusion processing module is further used to perform the following operations:

[0076] Obtain a tag prompt word, and input the media data, the target retrieval result, and the tag prompt word into a sorting model; the tag prompt word is used to instruct the sorting model to generate a matching verification result between the media data and the target retrieval result;

[0077] Extract features from the media data through the sorting model to obtain a mixed-modal feature vector, and based on the mixed-modal feature vector, generate similarity scores corresponding to the F sorting contents in the sorting model; F is a positive integer, and one sorting content includes one or more annotated media tags;

[0078] Based on the similarity scores corresponding to the F sorting contents respectively, the tag prompt word, and the target retrieval result, generate a matching verification result corresponding to each media tag in the target retrieval result; the matching verification result is used to represent whether the media data is consistent or inconsistent with the media tag in the target retrieval result;

[0079] Determine the media tags corresponding to the matching verification results indicating consistency as the tags to be sorted, and generate a relevance score for the tags to be sorted based on the similarity scores corresponding to the sorting contents with the tags to be sorted in the F sorting contents;

[0080] Based on the relevance scores, perform sorting processing on the tags to be sorted in the target retrieval result to obtain a sorting result.

[0081] In a possible implementation, when the feature extraction module is used to extract features from the media data to obtain M modal feature vectors, it is specifically used to perform the following operations:

[0082] Input the media data into a mixed-modal retrieval model; the mixed-modal retrieval model includes an encoding layer;

[0083] Extract features from the media data through the encoding layer to obtain P single-modal feature vectors; P is a positive integer less than or equal to M;

[0084] If the media data is video data, perform feature fusion processing on the P single-modal feature vectors to obtain multi-modal feature vectors, and determine the P single-modal feature vectors and the multi-modal feature vectors as M modal feature vectors;

[0085] If the media data is text data or image data, determine the P single-modal feature vectors as M modal feature vectors.

[0086] In a possible implementation, the encoding layer includes P single-modal encoders with different modalities; when the feature extraction module is used to perform feature extraction on P single-modal data through the encoding layer to obtain P single-modal feature vectors, it is specifically used to perform the following operations:

[0087] Perform data splitting on the media data according to the media type of the media data to obtain Q single-modal data with different media types; Q is a positive integer less than or equal to P;

[0088] Input the Q single-modal data into the single-modal encoders that match the media types of the Q single-modal data respectively, and perform feature extraction on the Q single-modal data in the Q single-modal encoders respectively to obtain Q single-modal feature vectors;

[0089] If Q is less than P, generate padding feature vectors corresponding to the single-modal encoders other than the Q single-modal encoders in the P single-modal encoders, and determine the Q single-modal feature vectors and the padding feature vectors as P single-modal feature vectors;

[0090] If Q is equal to P, determine the Q single-modal feature vectors as P single-modal feature vectors.

[0091] In a possible implementation, the hybrid-modal retrieval model further includes an alignment processing layer; when the feature extraction module is used to determine the Q single-modal feature vectors and the padding feature vectors as P single-modal feature vectors, it is specifically used to perform the following operations:

[0092] Input the Q single-modal feature vectors and the padding feature vectors into the alignment processing layer;

[0093] Based on the alignment projection matrix corresponding to the alignment processing layer, perform linear transformation on the Q single-modal feature vectors and the padding feature vectors respectively to obtain P alignment feature vectors;

[0094] Determine the P alignment feature vectors as P single-modal feature vectors.

[0095] In a possible implementation, the Q unimodal data includes image data, and the Q unimodal encodings include visual encoders corresponding to the image type. When the feature extraction module is used to extract features from the Q unimodal data in the Q unimodal encoders respectively to obtain Q unimodal feature vectors, it is specifically used to perform the following operations:

[0096] In the visual encoder, perform image segmentation on the image data to obtain T unit sampled images; T is a positive integer;

[0097] Extract features from the T unit sampled images to obtain T unit visual feature vectors;

[0098] Based on the rearrangement factor, remap the vector elements representing the spatial channels in the T unit visual feature vectors to the feature channels to obtain T visual feature vectors to be matched; the rearrangement factor is determined based on the number T of unit visual feature vectors, and the vector length of the unit visual feature vectors is greater than the vector length of the visual feature vectors to be matched;

[0099] Perform attention processing on the T visual feature vectors to be matched to obtain a global visual feature vector, and determine the global visual feature vector as the unimodal feature vector corresponding to the image data.

[0100] In a possible implementation, when the feature extraction module is used to perform attention processing on the T visual feature vectors to be matched to obtain a global visual feature vector, it is specifically used to perform the following operations:

[0101] Obtain the query parameter matrix, key parameter matrix, and value parameter matrix in the visual encoder; the query parameter matrix, key parameter matrix, and value parameter matrix are all matrices composed of learnable parameters;

[0102] Perform vector concatenation on the T visual feature vectors to be matched to obtain a visual feature sequence to be matched, perform dot product operation on the visual feature sequence to be matched and the query parameter matrix to obtain a query vector, perform dot product operation on the visual feature sequence to be matched and the key parameter matrix to obtain a key vector, and perform dot product operation on the visual feature sequence to be matched and the value parameter matrix to obtain a value vector;

[0103] Generate an attention score vector based on the query vector and the key vector, perform dimensionality reduction processing on the attention score vector based on the vector length of the key vector, perform normalization processing on the dimensionality-reduced attention score vector to obtain an attention weight vector, and perform dot product operation on the attention weight vector and the value vector to obtain a global visual feature vector.

[0104] On the one hand, an embodiment of the present application provides another data processing device, including:

[0105] A sample data acquisition module, configured to acquire sample media data and input the sample media data into an initial modality retrieval model;

[0106] A sample feature extraction module, configured to extract features from the sample media data through the initial modality retrieval model to obtain M sample modality feature vectors; M is a positive integer, and the modalities of the M sample modality feature vectors are different from each other;

[0107] A sample instruction generation module, configured to generate sample modality pair retrieval instructions corresponding to N retrieval data sets for each sample modality feature vector, to obtain S sample modality pair retrieval instructions; N is a positive integer, S is the product of M and N; the N retrieval data sets have different modality identifiers, the retrieval data in different retrieval data sets have different modalities, and each sample modality pair retrieval instruction includes a sample modality feature vector and a modality identifier;

[0108] A sample instruction retrieval module, configured to perform feature retrieval in the N retrieval data sets based on the sample modality feature vectors respectively included in the S sample modality pair retrieval instructions, to obtain sample retrieval results corresponding to the S sample modality pair retrieval instructions; a sample retrieval result is obtained by performing feature retrieval in the retrieval data set corresponding to the modality identifier included in a sample modality pair retrieval instruction based on the sample modality feature vector included in the sample modality pair retrieval instruction; a sample retrieval result includes media labels carried by the retrieval data matched through feature retrieval;

[0109] A sample training and processing module, configured to perform label fusion processing on the S sample retrieval results to obtain a target sample retrieval result, and adjust the initial modality retrieval model based on the target sample retrieval result and the sample label of the sample media data to obtain a hybrid modality retrieval model; the hybrid modality retrieval model is used to generate a target retrieval result of media data.

[0110] On the one hand, an embodiment of the present application provides a computer device, including: a processor, a memory, and a network interface;

[0111] The processor is connected to the memory and the network interface. Among them, the network interface is used to provide a data communication function, and the memory is used to store a computer program. When the computer program is executed by the processor, the computer device executes the method provided by the embodiment of the present application.

[0112] On the one hand, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor so that a computer device having the processor executes the method provided by the embodiment of the present application.

[0113] One aspect of the embodiments of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method provided by the embodiments of the present application.

[0114] In the embodiments of the present application, by extracting features from media data, M modality feature vectors with different modalities are obtained. The M modality feature vectors can include rich unimodal features and multimodal features, enabling the model to more comprehensively understand the input content. N retrieval data sets with different modality identifiers are obtained. One modality identifier can correspond to one modality feature, and the modalities of the retrieval data in different retrieval data sets are different. Generate modality pair retrieval instructions corresponding to the N retrieval data sets for each modality feature vector, obtaining S modality pair retrieval instructions. Each modality pair retrieval instruction includes a modality feature vector and a modality identifier. Based on the modality feature vectors included in the S modality pair retrieval instructions respectively, perform feature retrieval in the N retrieval data sets to obtain the retrieval results corresponding to the S modality pair retrieval instructions respectively. Each retrieval result includes the media tags carried by the retrieval data matched through feature retrieval. It can be seen that the present application can distinguish the retrieval tasks of different modality pairs through the modality pair retrieval instructions, and at the same time support various modality feature retrievals such as unimodal, cross-modal, and multimodal, flexibly adapting to the retrieval requirements between different modalities. In the retrieval task corresponding to a single modality pair retrieval instruction, the present application can learn and capture the fine-grained differences on a single modality feature by distinguishing the different features of the M modality feature vectors, and by indirectly retrieving the retrieval data similar to the media data in the retrieval data set, it can avoid the information limitations brought by directly retrieving the text semantics of the tags, reduce the risk of tag mismatch, and improve the accuracy of the retrieval results. It also realizes the integration of different modality retrieval tasks. Only one model can perform mutual retrieval between different modalities, and finally generate the retrieval results corresponding to the S modality pair retrieval instructions respectively, without training multiple modality pair retrieval models, reducing the training cost, reducing the complexity of the system, and the direct output of a single model also improves the efficiency of the retrieval task. At the same time, comprehensive feature retrieval of the M modality feature vectors of the media data is performed in the N retrieval data sets, making full use of the multiple modality features of the media data, obtaining the retrieval results corresponding to the S modality retrieval instructions respectively, and performing fusion processing on the S retrieval results to obtain the target retrieval result. The retrieval results corresponding to the S modality pairs can be comprehensively considered, and the media tags related to the media data can be accurately mined in terms of both depth and breadth, thereby improving the accuracy of the retrieval results for the media tags. BRIEF DESCRIPTION OF THE DRAWINGS

[0115] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0116] Figure 1 It is a schematic diagram of a network architecture provided by an embodiment of the present application;

[0117] Figure 2 It is a schematic diagram of a data processing scenario provided by an embodiment of the present application Figure 1 ;

[0118] Figure 3 It is a schematic flowchart of a data processing method provided by an embodiment of the present application Figure 1 ;

[0119] Figure 4 It is a schematic flowchart of a data processing method provided by an embodiment of the present application Figure 2 ;

[0120] Figure 5 It is a schematic diagram of the model structure of a hybrid modality retrieval model provided by an embodiment of the present application Figure 1 ;

[0121] Figure 6 It is a schematic diagram of the model structure of a hybrid modality retrieval model provided by an embodiment of the present application Figure 2 ;

[0122] Figure 7 It is a schematic diagram of a data processing scenario provided by an embodiment of the present application Figure 2 ;

[0123] Figure 8 It is a schematic diagram of a data processing scenario provided by an embodiment of the present application Figure 3 ;

[0124] Figure 9 It is a schematic diagram of a data processing scenario provided by an embodiment of the present application Figure 4 ;

[0125] Figure 10 It is a schematic flowchart of a data processing method provided by an embodiment of the present application Figure 3 ;

[0126] Figure 11 It is a schematic diagram of the structure of a data processing device provided by an embodiment of the present application Figure 1 ;

[0127] Figure 12It is a schematic structure of a data processing device provided by an embodiment of the present application Figure 2 ;

[0128] Figure 13 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0129] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0130] Please refer to Figure 1 , Figure 1 It is a schematic diagram of a network architecture provided by an embodiment of the present application. As Figure 1 shown, the network architecture may include a service server 100 and a cluster of terminal devices. The cluster of terminal devices may include terminal devices 10a, 10b,..., 10n. Among them, any terminal device in the cluster of terminal devices may have a communication connection with the service server 100. For example, there is a communication connection between terminal device 10a and service server 100, and there is a communication connection between terminal device 10b and service server 100. Among them, the above communication connection is not limited to the connection method, and can be directly or indirectly connected through a wired communication method, or can be directly or indirectly connected through a wireless communication method, or can be connected through other methods, which are not limited in the present application.

[0131] Among them, each terminal device in the cluster of terminal devices may include: intelligent terminals with data processing functions such as smart phones, tablet computers, laptop computers, desktop computers, intelligent voice interaction devices, smart home appliances (such as smart TVs), wearable devices, vehicle-mounted terminals, and aircraft. Among them, the vehicle-mounted terminal may be a terminal device in intelligent transportation scenarios and assisted driving scenarios. It should be understood that each terminal device in the cluster of terminal devices as Figure 1 shown may be installed with an application client with data processing functions. When the application client runs on each terminal device, it can perform data interaction with the above-mentioned Figure 1 shown service server 100 respectively.

[0132] The application client may specifically include: a vehicle client, a smart home client, an entertainment client (e.g., a game client), a multimedia client (e.g., a video client), a social client, and an information client (e.g., a news client), etc. The application client in the embodiment of the present application may be integrated in a client (e.g., a social client) or may be an independent client (e.g., a news client). The embodiment of the present application does not limit the type of the application client.

[0133] For ease of understanding, the terminal device 10a in the terminal device cluster is taken as an example for explanation. The object (user) can upload media data through the application client in the terminal device 10a. The media data can be data with one or more media types, and the media types can include video, audio, image, text and other types.

[0134] The business server 100 may be a server corresponding to the application client. The business server 100 may obtain media data and perform data retrieval on the media data in several retrieval data sets with different modality identifiers. For example, the retrieval data set may be used to retrieve retrieval data similar to the media data, and the media tag of the retrieval data may be determined as the retrieval result to obtain S retrieval results. The specific process of data retrieval may be referred to below. Figure 3 Specific description in the corresponding embodiment. Among them, the retrieval data set refers to a data set with a certain modality or media type, and each retrieval data in the retrieval data set can include one or more marked media tags. A modality identifier can correspond to a modality, and the modality can refer to the data type or the feature type of the data. For example, images, audio, and text can have different modalities, and natural scene images and text-embedded images can also have different modalities.

[0135] The service server 100 may perform fusion processing on the S search results. For example, the fusion processing may be to take the intersection or union of the media tags in the S search results to obtain the target search result. The service server 100 may send the target search result to the terminal device 10a, and the terminal device 10a may display the media tags of the media data in the target search result to the object through the application client.

[0136] The embodiment of the present application provides an indirect retrieval method supporting multiple modal features. By indirectly retrieving retrieval data similar to media data in multiple modalities, one or more media tags annotated for the retrieved retrieval data are determined as the retrieval result. The existing tags of the retrieval data can be directly reused without a large amount of manual annotation, which can significantly reduce the annotation time and labor consumption. At the same time, by retrieving similar content, the information limitation caused by directly retrieving the text semantics of tags can be avoided, the risk of tag mismatch can be reduced, and the accuracy of the retrieval result for media tags can be improved.

[0137] Please refer to Figure 2 , Figure 2 which is a schematic diagram of a data processing scenario provided by the embodiment of the present application. Figure 1 . As Figure 2 shown, for ease of understanding, taking the media data as video data as an example for illustration, the computer device can obtain video data (content production stage), input the video data into the hybrid modal retrieval model, and predict the video tags of the video data through the hybrid modal retrieval model (content processing stage). This process can also be called machine tagging. Among them, the hybrid modal retrieval model can be a model that supports joint retrieval in multiple modalities. For example, it can support retrieval in the same modality, cross-modal, multi-modal, etc. at the same time. The same modality retrieval means that the modality of the retrieval data and the target content is the same and both are single modalities. For example, it is the retrieval between images. The cross-modal retrieval means that the modality of the retrieval data and the target content are different from each other. For example, it is the retrieval between images and texts. The multi-modal retrieval means that the modality of the retrieval data and the target content are both multi-modal. For example, it is the retrieval between videos (including modalities such as texts, images, and audios) and videos. The computer device can be any one of the service server 100 or the terminal device cluster in the corresponding embodiment above. Figure 1 The computer device can extract the feature of the video data through the hybrid modal retrieval model to obtain M modal feature vectors. The computer device can generate a modal pair retrieval instruction for each retrieval data set for the M modal feature vectors. Among them, the retrieval data set refers to a data set with a certain modality or media type. Each retrieval data in the retrieval data set can include one or more annotated media tags. Each modal pair retrieval instruction includes a modal feature vector and a modal identifier.

[0138]

[0139] The computer device can perform feature retrieval on the retrieval instruction in each retrieval dataset through several modalities, and obtain S retrieval results corresponding to the retrieval instruction for different modality pairs. Taking a modality pair retrieval instruction as an example, the computer device can obtain the modality identifier corresponding to the modality pair retrieval instruction, and perform feature retrieval in the retrieval dataset with this modality identifier through the modality feature vector corresponding to the modality pair retrieval instruction, obtain the retrieval data similar to the modality feature vector, and determine the media label possessed by the retrieval data as the retrieval result corresponding to the modality pair retrieval instruction. The computer device can perform fusion processing on the S retrieval results. The fusion processing can be, for example, taking the intersection or union of the media labels in the S retrieval results to obtain the target retrieval result. Determine the media label included in the target retrieval result as the video label of the video data. The video label of the video data can be used to classify the video data and recommend the video data to the object (content distribution stage).

[0140] In the embodiments of the present application, the retrieval tasks of different modality pairs are distinguished by S modality pairs for the retrieval instruction, supporting feature retrieval of multiple modalities such as same modality, cross-modal, and multi-modal. Comprehensive feature retrieval can be performed on the M modality feature vectors of the media data in the N retrieval datasets, so as to make full use of the features of the M modality feature vectors in multiple modalities, and then generate S more accurate retrieval results, and perform fusion processing on the S retrieval results to obtain the target retrieval result, which can improve the accuracy of the retrieval result for the media label. At the same time, the embodiments of the present application can be applied in the field of label prediction for different media types, such as text labels, image labels, video labels, etc., and can also be applied in the retrieval field, such as content search, content understanding, and information extraction. Through the joint retrieval of multiple modality pairs, the accuracy of the retrieval is improved.

[0141] Please refer to Figure 3 , Figure 3 which is a flowchart of a data processing method provided by the embodiments of the present application Figure 1 This data processing method can be executed by a computer device, and the computer device can be any one of the service server 100 or the terminal device cluster shown in Figure 1 For example, it can be the terminal device 10a. The following will take the execution of this data processing method by the computer device as an example for description. Among them, this data processing method can at least include the following steps S101 - step S104:

[0142] Step S101, obtain media data, perform feature extraction on the media data, and obtain M modality feature vectors; M is a positive integer, and the modalities of the M modality feature vectors are different from each other;

[0143] Specifically, the computer device can obtain media data, which can be data with one or more media types. The media types can include video, audio, image, text, etc. The computer device can perform feature extraction on the media data to obtain P single-modal feature vectors. Taking the media data as video data as an example, the computer device can split the media data into text data (such as video titles, video introductions, etc.), image data (such as video frames, video covers, etc.), and audio data, and perform feature extraction on the text data, image data, and audio data respectively through encoders of corresponding modalities to obtain P single-modal feature vectors.

[0144] Among them, the modalities of the P single-modal feature vectors are different from each other. The modality can refer to the data type or the feature type that the data has. For example, images, audio, and text can have different modalities, and natural scene images and text-embedded images can also have different modalities. The encoder for extracting text data can be the text encoder in the CLIP encoder (Contrastive Language-Image Pretraining Vision Transformer, a pre-trained vision encoder for contrastive learning), the encoder for extracting image data can be the ViT encoder (Vision Transformer, a vision encoder), and the encoder for extracting audio data can be the CLAP encoder (Contrastive Language-Audio Pretraining, a pre-trained audio encoder for contrastive learning). The embodiments of the present application do not limit this here.

[0145] When the media data is multi-modal data such as video data or graphic content, the computer device can perform feature fusion processing on the P single-modal feature vectors to obtain multi-modal feature vectors. Optionally, the computer device can perform feature fusion processing on the P single-modal feature vectors to obtain multiple multi-modal feature vectors. For example, when the media data is video data, the text feature and the image feature can be fused into one multi-modal feature vector, and the text feature, image feature, and audio feature can be fused into another multi-modal feature vector. The feature fusion processing can be through attention processing or through cross-modal embedding mapping to map the features into the same semantic space. The embodiments of the present application do not limit this here.

[0146] The computer device can determine the P single-modal feature vectors and the multi-modal feature vectors as M modal feature vectors. It can be understood that when the media data has only one modality, the number of single-modal feature vectors and the number of modal feature vectors are both 1, that is, the single-modal feature vector can be determined as the modal feature vector.

[0147] Step S102: Generate modality-pair retrieval instructions corresponding to N retrieval data sets for each modality feature vector, obtaining S modality-pair retrieval instructions; N is a positive integer, and S is the product of M and N; the N retrieval data sets have mutually different modality identifiers, the modalities of the retrieval data in different retrieval data sets are mutually different, and each modality-pair retrieval instruction includes a modality feature vector and a modality identifier.

[0148] Specifically, the computer device can obtain N retrieval data sets. The N retrieval data sets have mutually different modality identifiers. One modality identifier can correspond to one modality. A retrieval data set refers to a data set with a certain modality or media type. Each retrieval data in the retrieval data set can include one or more labeled media tags. The retrieval data sets can be located in different databases or in the same database. The embodiments of the present application do not limit this here.

[0149] The computer device can generate modality-pair retrieval instructions corresponding to the N retrieval data sets for each modality feature vector, obtaining S modality-pair retrieval instructions. Among them, each modality-pair retrieval instruction includes a modality feature vector and a modality identifier. Taking the media data as video data as an example, the modality feature vector corresponding to the video data can include the single-modality feature vectors corresponding to the text, image, and audio modalities respectively, and the multi-modality feature vector obtained by fusing the text, image, and audio modalities. The retrieval data sets include four retrieval data sets: a text data set (with a text modality), an image data set (with an image modality), an audio data set (with an audio modality), and a video data set (with a multi-modality of text, image, and audio). The computer device can generate modality-pair retrieval instructions corresponding to the text data set, the image data set, the audio data set, and the video data set for the modality feature vector corresponding to the text modality (that is, 4 modality-pair retrieval instructions can be generated). Similarly, generating modality-pair retrieval instructions corresponding to the 4 retrieval data sets for the single-modality feature vector corresponding to the image, the single-modality feature vector corresponding to the audio, and the multi-modality feature vector, a total of 16 mutually different modality-pair retrieval instructions can be obtained. Among them, a modality pair refers to a combined pairing formed by two modalities. For example, "text + image" and "image + text" are mutually different modality pairs.

[0150] Among them, taking the text modality (Natural Language Processing, NLP) and the image modality (Computer Vision, CV) and their combination as examples, the modality pair retrieval instruction can refer to an instruction for indicating feature retrieval in a retrieval dataset with modality identifiers through certain modality feature vectors. The modality pair retrieval instruction can include the corresponding modality feature vector, modality identifier, and prompt words. When the modality pair is NLP-CV, the prompt words included in the modality pair retrieval instruction can specifically be "Please retrieve images that meet the requirements according to the text information". When the modality pair is NLP-CV+NLP, the prompt words included in the modality pair retrieval instruction can specifically be "Please retrieve images and text information that meet the requirements according to the text information". When the modality pair is CV-CV, the prompt words included in the modality pair retrieval instruction can specifically be "Please retrieve images that meet the requirements according to the image information". The specific form of the modality pair retrieval instruction in the embodiments of the present application is not limited herein.

[0151] Step S103: Based on the modality feature vectors respectively included in the S modality pair retrieval instructions, perform feature retrieval in the N retrieval datasets to obtain the retrieval results respectively corresponding to the S modality pair retrieval instructions; one retrieval result is obtained by performing feature retrieval in the retrieval dataset corresponding to the modality identifier included in the modality feature vector included in one modality pair retrieval instruction; one retrieval result includes the media tags carried by the retrieved data matched through feature retrieval.

[0152] Specifically, the computer device can perform feature retrieval in the N retrieval datasets based on the modality feature vectors respectively included in the S modality pair retrieval instructions to obtain the retrieval results respectively corresponding to the S modality pair retrieval instructions. Among them, one retrieval result is obtained by performing feature retrieval in the retrieval dataset corresponding to the modality identifier included in the modality feature vector included in one modality pair retrieval instruction.

[0153] Taking the modality pair retrieval instruction 1 in the S modality pair retrieval instructions as an example, the modality pair retrieval instruction 1 can include a modality pair of CV-NLP. The computer device can perform feature retrieval in the retrieval dataset with the text modality identifier based on the modality pair retrieval instruction 1 through the modality feature vector corresponding to the image modality to obtain the retrieval result corresponding to the modality pair retrieval instruction 1. For example, it can calculate the feature similarity between the modality feature vector and the retrieved data, determine the retrieved data with a feature similarity greater than a certain threshold as the retrieved data similar to the media data, and determine the media tag of the similar retrieved data as the retrieval result 1. Similarly, the computer device can perform feature retrieval in the retrieval datasets with the corresponding modality identifiers based on the S modality pair retrieval instructions through the included modality feature vectors to obtain the retrieval results respectively corresponding to the S modality pair retrieval instructions.

[0154] Step S104: Perform label fusion processing on the S retrieval results to obtain the target retrieval result.

[0155] Specifically, the computer device can perform label fusion processing on the S retrieval results. The fusion processing can, for example, take the intersection or union of the media labels in the S retrieval results to obtain the target retrieval result. Taking the intersection as an example, the S retrieval results may include Retrieval Result 1 (including Media Label 1, Media Label 2, and Media Label 3), Retrieval Result 2 (including Media Label 1 and Media Label 4), and Retrieval Result 3 (including Media Label 2, Media Label 4, and Media Label 5). The computer device can take the intersection of the media labels in Retrieval Result 1, Retrieval Result 2, and Retrieval Result 3 to obtain the target retrieval result including Media Label 1, Media Label 2, and Media Label 4.

[0156] In the embodiments of the present application, by extracting features from media data, M modality feature vectors with different modalities are obtained. The M modality feature vectors can include rich single-modal features and multi-modal features, enabling the model to more comprehensively understand the input content. N retrieval data sets with different modality identifiers are obtained. One modality identifier can correspond to one modality feature, and the modalities of the retrieval data in different retrieval data sets are different. Generate modality-pair retrieval instructions corresponding to the N retrieval data sets for each modality feature vector, obtaining S modality-pair retrieval instructions. Each modality-pair retrieval instruction includes a modality feature vector and a modality identifier. Based on the modality feature vectors included in the S modality-pair retrieval instructions, perform feature retrieval in the N retrieval data sets to obtain the retrieval results corresponding to the S modality-pair retrieval instructions respectively. Each retrieval result includes the media tags carried by the retrieved data matched through feature retrieval. It can be seen that the present application can distinguish the retrieval tasks of different modality pairs through the modality-pair retrieval instructions, and at the same time support feature retrieval of multiple modalities such as same modality, cross modality, and multi modality, flexibly adapting to the retrieval requirements between different modalities. For the retrieval task corresponding to a single modality-pair retrieval instruction, the present application can learn and capture the fine-grained differences on a single modality feature by distinguishing the different features of the M modality feature vectors, and by indirectly retrieving the retrieval data similar to the media data in the retrieval data set, it can avoid the information limitation brought by directly retrieving the text semantics of the tags, reduce the risk of tag mismatch, and improve the accuracy of the retrieval results. It also realizes the integration of different modality retrieval tasks. Only one model can perform mutual retrieval between different modalities, and finally generate the retrieval results corresponding to the S modality-pair retrieval instructions respectively, without training retrieval models for multiple modality pairs, reducing the training cost, reducing the complexity of the system, and the direct output of a single model also improves the efficiency of the retrieval task. At the same time, perform comprehensive feature retrieval on the M modality feature vectors of the media data in the N retrieval data sets, make full use of the multiple modality features of the media data, obtain the retrieval results corresponding to the S modality retrieval instructions respectively, and perform fusion processing on the S retrieval results to obtain the target retrieval result, which can comprehensively consider the retrieval results corresponding to the S modality pairs respectively, and accurately mine the media tags related to the media data both in depth and breadth, thereby improving the accuracy of the retrieval results for the media tags.

[0157] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a data processing method provided by an embodiment of the present application Figure 2 , and this data processing method can be executed by a computer device, and the computer device can be such as Figure 1The service server 100 shown or any terminal device in the terminal device cluster may be, for example, the terminal device 10a. The following will take the data processing method executed by a computer device as an example for explanation. The data processing method may at least include the following steps S201 to S206:

[0158] Step S201, obtaining media data;

[0159] For details, please refer to the above Figure 3 Detailed description of step S101 of the corresponding embodiment.

[0160] Step S202, inputting the media data into a hybrid modal retrieval model; the hybrid modal retrieval model includes a coding layer; extracting features from the media data through the coding layer to obtain P single-modal feature vectors; P is a positive integer less than or equal to M;

[0161] Specifically, the computer device can input the media data into a hybrid modal retrieval model. The hybrid modal retrieval model can be a model that supports the joint retrieval of multiple modalities, for example, it can simultaneously support homomodal, cross-modal, multi-modal and other retrieval. Homomodal retrieval means that the modality of the retrieval data and the target content is the same and both are single-modal, such as retrieval between images, cross-modal retrieval means that the modality of the retrieval data and the target content is different, such as retrieval between images and texts, and multi-modal retrieval means that the modality of the retrieval data and the target content are both multi-modal, such as retrieval between videos (including text, images, audio and other modalities) and videos.

[0162] Please also refer to Figure 5 , Figure 5 A model structure diagram of a hybrid modality retrieval model provided in an embodiment of the present application is shown in FIG. Figure 1 ,like Figure 5 As shown, the hybrid modality retrieval model may include a coding layer, an alignment processing layer, and a feature retrieval layer, and the coding layer includes P single-modality encoders with different modalities, for example, single-modality encoder 1, single-modality encoder 2, ..., single-modality encoder P. Among them, the encoder for extracting text data may be a text encoder in the CLIP encoder, the encoder for extracting image data may be a ViT encoder, and the encoder for extracting audio data may be a CLAP encoder, which is not limited in the embodiments of the present application.

[0163] The process by which a computer device extracts features from media data respectively through unimodal encoders with different modalities to obtain unimodal feature vectors can be as follows: The media data is split according to the media type of the media data to obtain Q unimodal data with different media types; Q is a positive integer less than or equal to P; The Q unimodal data are respectively input into unimodal encoders that match the media types of the Q unimodal data. Among the Q unimodal encoders, the Q unimodal data are respectively subjected to feature extraction to obtain Q unimodal feature vectors; If Q is less than P, padding feature vectors corresponding to the unimodal encoders other than the Q unimodal encoders among the P unimodal encoders are generated, and the Q unimodal feature vectors and the padding feature vectors are determined as P unimodal feature vectors; If Q is equal to P, the Q unimodal feature vectors are determined as P unimodal feature vectors.

[0164] Specifically, if the media data is unimodal image, text, or audio data, the computer device can input the media data into the unimodal encoder corresponding to this modality to obtain the unimodal feature vector corresponding to the media data. If the media data is multimodal data such as video data or graphic and text content, the computer device can split the media data according to the media type of the media data to obtain Q unimodal data with different media types. For example, the media data is split into text data (such as video titles, video introductions, etc.), image data (such as video frames, video covers, etc.), and audio data to obtain Q unimodal data with different media types. The computer device can respectively input the Q unimodal data into unimodal encoders that match the media types of the Q unimodal data. For example, the text data is input into the text encoder, and the image data is input into the image encoder, etc. Among the Q unimodal encoders, the Q unimodal data are respectively subjected to feature extraction to obtain Q unimodal feature vectors.

[0165] For ease of understanding, an example is given where there are Q single-modal data including image data (such as video frames that can be video data), and Q single-modal encodings include visual encoders corresponding to the image type. The process by which a computer device extracts features from image data through a visual encoder to obtain a single-modal feature vector can be as follows: In the visual encoder, the image data is segmented to obtain T unit sampling images; T is a positive integer; feature extraction is performed on the T unit sampling images to obtain T unit visual feature vectors; based on a rearrangement factor, the vector elements representing the spatial channels in the T unit visual feature vectors are remapped to the feature channels to obtain T visual feature vectors to be matched; the rearrangement factor is determined based on the number T of unit visual feature vectors, and the vector length of the unit visual feature vector is greater than the vector length of the visual feature vector to be matched; attention processing is performed on the T visual feature vectors to be matched to obtain a global visual feature vector, and the global visual feature vector is determined as the single-modal feature vector corresponding to the image data.

[0166] Specifically, the computer device can segment the image data to obtain T unit sampling images. Each unit sampling image can be flattened and mapped to a vector of a fixed dimension (such as 14×14 pixels), and at the same time, position encoding is added to retain the relative position information of the unit sampling images. In the visual encoder, the computer device can perform feature extraction on the T unit sampling images to obtain T unit visual feature vectors (visual tokens). The computer device can perform pixel rearrangement (Pixel Shuffle) on the T unit visual feature vectors to rearrange the elements of the unit visual feature vectors and redistribute the high-dimensional spatial channel information to the feature channel dimension. For example, the shape of the T unit sampling images can be a tensor of B×T×H×W×C. Among them, B (Batchsize, indicating the number of data processed by the model at one time), H (Height, the height of the image), and W (Width, the width of the image) are used to represent the spatial channels, and C (Channels, the number of feature channels) is used to represent the feature channels.

[0167] Taking the rearrangement factor r as an example, the rearrangement factor is determined based on the number T of unit visual feature vectors, and r is a positive integer less than T. Specifically, the computer device can rearrange the T unit visual feature vectors according to the rearrangement factor r, and remap the vector elements representing the spatial channels to the feature channels. Specifically, the computer device can rearrange the unit sampled images according to the rearrangement factor r, reduce the width and height of the images to 1 / r of the original respectively, and then rearrange these reduced image blocks into new feature channels, so that the number of feature channels becomes r^2. For example, the shape of the T visual feature vectors to be matched can be a tensor of B×T×(H / r)×(W / r)×(C×r^2), where the vector length of the unit visual feature vector is greater than the vector length of the visual feature vector to be matched. For example, if the rearrangement factor is selected as 2, for a 10-frame video A, each frame of the image will be scaled to 448×448, and then each frame of the image will be segmented into 1024 visual tokens. Each frame of the image can be processed by Pixel Shuffle to obtain 256 visual tokens. Then, the video A could originally include 10×1024 visual tokens and include 10×256 visual tokens after Pixel Shuffle. It can be understood that Pixel Shuffle can reduce the vector dimensions representing the image height and image width in the unit sampled image by rearranging the pixels of the image, thereby reducing the number of visual tokens.

[0168] The computer device can perform vector concatenation on the T visual feature vectors to be matched to obtain a visual feature sequence to be matched, and perform attention processing on the visual feature sequence to be matched to obtain a global visual feature vector. The computer device can obtain the query parameter matrix W in the visual encoder Q 、the key parameter matrix W K and the value parameter matrix W V , where the query parameter matrix W Q 、the key parameter matrix W K and the value parameter matrix W V are all matrices composed of learnable parameters.

[0169] The computer device performs a dot product operation on the visual feature sequence to be matched and the query parameter matrix W Q to obtain a query vector Q, performs a dot product operation on the visual feature sequence to be matched and the key parameter matrix W K to obtain a key vector K, and performs a dot product operation on the visual feature sequence to be matched and the value parameter matrix W VPerform a dot product operation to obtain the value vector V. The computer device can generate an attention score vector based on the query vector Q and the key vector K, perform dimensionality reduction processing on the attention score vector based on the dimension d of the key vector, perform normalization processing (Softmax) on the dimensionally reduced attention score vector to obtain a normalized vector, and perform a dot product operation on the normalized vector and the value vector V to obtain a global visual feature vector. The process can be as shown in formula (1):

[0170]

[0171] where K T is the transposed key vector, and d k is the dimension of the attention key vector K.

[0172] Optionally, the global visual feature vector can also be the feature vector corresponding to the CLS Token extracted by the ViT model.

[0173] Please also refer to Figure 6 , Figure 6 which is a schematic diagram of the model structure of a hybrid modality retrieval model provided by an embodiment of the present application Figure 2 , as Figure 6 shown. The P unimodal encoders with different modalities in the encoding layer can include an image encoder and a text encoder. If the media data only includes text data, the image encoder is a unimodal encoder without input data (i.e., if Q is less than P). The computer device can generate padding feature vectors (Padding Embedding) corresponding to the unimodal encoders other than the Q unimodal encoders among the P unimodal encoders. The padding feature vector can be a vector all of whose elements are zero. The text encoder extracts features from the text data to obtain a unimodal feature vector corresponding to the text features. The computer device can determine the Q unimodal feature vectors and the padding feature vectors as the P unimodal feature vectors. When all P unimodal encoders have input data (i.e., Q is equal to P), the media data can include unimodal data 1, unimodal data 2,..., unimodal data P. The computer device can perform encoding processing on each unimodal encoder corresponding to the unimodal data to obtain P unimodal feature vectors.

[0174] Please also refer to Figure 5 , as Figure 5As shown, the hybrid modality retrieval model may further include an alignment processing layer, which can be used to convert the outputs of P unimodal encoders into feature vectors of the same dimension. The process can be as follows: input Q unimodal feature vectors and padding feature vectors into the alignment processing layer; based on the alignment projection matrix corresponding to the alignment processing layer, perform linear transformation on the Q unimodal feature vectors and padding feature vectors respectively to obtain P aligned feature vectors; determine the P aligned feature vectors as P unimodal feature vectors.

[0175] Specifically, the computer device can input Q unimodal feature vectors and padding feature vectors into the alignment processing layer. The alignment processing layer can be used to convert the feature pairs into vectors of a fixed dimension. The alignment processing layer can be an MLP layer (Multi-Layer Perceptron). Taking the example of converting an A-dimensional unimodal feature vector into a B-dimensional unimodal feature vector, the alignment processing process can be as shown in formula (2):

[0176]

[0177] where i = 1, 2, …, B, j = 1, 2, …, A. α i is the i-th vector element of the B-dimensional unimodal feature vector, and β j is the j-th vector element of the A-dimensional unimodal feature vector. represents the element value of the j-th row and i-th column of the alignment projection matrix W AB , and the dimension of the alignment projection matrix W AB is A × B. The B-dimensional unimodal feature vector can be represented as {α1, α2, …, α B}.

[0178] Similarly, the computer device can obtain the alignment projection matrices corresponding to the vector dimensions of the Q unimodal feature vectors and padding feature vectors respectively, perform linear transformation on the Q unimodal feature vectors and padding feature vectors respectively to obtain P aligned feature vectors. The vector dimensions of the P aligned feature vectors are all vectors, and determine the P aligned feature vectors as P unimodal feature vectors. Through the alignment processing layer, the computer device can represent the Q unimodal feature vectors and padding feature vectors in the same feature space, improving the understanding ability of the model.

[0179] Step S203, if the media data is video data, perform feature fusion processing on the P unimodal feature vectors to obtain a multimodal feature vector, and determine the P unimodal feature vectors and the multimodal feature vector as M modal feature vectors; if the media data is text data or image data, determine the P unimodal feature vectors as M modal feature vectors.

[0180] Specifically, if the media data is multi-modal data such as video data or picture-and-text content, the computer device may perform feature fusion processing on P single-modal feature vectors to obtain multi-modal feature vectors. For example, when the media data is video data, the text feature and the image feature may be fused into one multi-modal feature vector, and the text feature, the image feature, and the audio feature may be fused into another multi-modal feature vector. The feature fusion processing may be through attention processing or through cross-modal embedding mapping to map the features into the same semantic space, which is not limited in the embodiments of the present application.

[0181] Optionally, the computer device may also perform causal transformation processing (CausalTransformer, a one-way masked attention processing) on the P single-modal feature vectors. The causal transformation processing restricts the information flow direction by introducing a causal mask (CausalMasking) to ensure that the model only uses the input at the historical or current moment to generate the output and avoid the leakage of future information. For example, in attention processing, the attention calculation at each position can only access the sequence information on its left (past) to ensure the temporal causality of the generation process. This makes it impossible for the previous tokens to obtain the information of the subsequent tokens. The computer device may obtain an attention result vector by performing causal transformation processing on the P single-modal feature vectors, and determine the last token in the attention result vector as the multi-modal feature.

[0182] If the media data is single-modal image, text, or audio data, the computer device may determine the P single-modal feature vectors as M modal feature vectors. It can be understood that when the media data has only one modality, the number of modal feature vectors is 1.

[0183] Step S204: Generate modality-pair retrieval instructions corresponding to N retrieval data sets for each modal feature vector, obtaining S modality-pair retrieval instructions; N is a positive integer, and S is the product of M and N; the N retrieval data sets have mutually different modality identifiers, the retrieval data in different retrieval data sets have different modalities, and each modality-pair retrieval instruction includes a modal feature vector and a modality identifier.

[0184] Specifically, the computer device may obtain N retrieval data sets, and the N retrieval data sets have mutually different modality identifiers. One modality identifier may correspond to one modality. A retrieval data set refers to a data set with a certain modality or media type, and each retrieval data in the retrieval data set may include one or more labeled media tags.

[0185] For ease of understanding, taking the M modal feature vectors including modal feature vector A iFor example, the process of generating modality pairs for retrieval instructions can be: generating modality feature vector A i N modality pairs of retrieval instructions for N retrieval data sets; the modality feature vectors included in the N modality pairs of retrieval instructions are all modality feature vector A i , and the retrieval data sets associated with the N modality pairs of retrieval instructions are different from each other; determine the N modality pairs of retrieval instructions corresponding to the M modality feature vectors as S modality pairs of retrieval instructions.

[0186] Specifically, the computer device can generate modality feature vector A i N modality pairs of retrieval instructions for N retrieval data sets, and the modality feature vectors included in the N modality pairs of retrieval instructions are all modality feature vector A i , and the retrieval data sets associated with the N modality pairs of retrieval instructions are different from each other. Please also refer to Figure 7 , Figure 7 which is a scenario schematic provided by an embodiment of the present application Figure 2 , such as Figure 7 shown, the retrieval data sets include 3 types of retrieval data sets: text data set (with text features), image data set (with image features), and video data set (with text + image features). The media data can include modality feature vectors corresponding to text features, modality feature vectors corresponding to image features, and modality feature vectors corresponding to text + image features. The computer device can be modality feature vector A i generate 3 modality pairs of retrieval instructions corresponding to the text data set, image data set, and video data set respectively.

[0187] Similarly, the computer device can generate N modality pairs of retrieval instructions for other modality feature vectors among the M modality feature vectors, and determine the N modality pairs of retrieval instructions corresponding to the M modality feature vectors as S modality pairs of retrieval instructions. Among them, the modality pair of retrieval instructions can refer to an instruction for indicating feature retrieval in a retrieval data set with a modality identifier through a certain modality feature vector. The modality pair of retrieval instructions can include the corresponding modality feature vector, modality identifier, and prompt word. For example, when the modality pair is NLP-CV, the prompt word included in the modality pair of retrieval instructions can specifically be "Please retrieve images that meet the requirements according to the text information". When the modality pair is NLP-CV + NLP, the prompt word included in the modality pair of retrieval instructions can specifically be "Please retrieve images and text information that meet the requirements according to the text information". When the modality pair is CV-CV, the prompt word included in the modality pair of retrieval instructions can specifically be "Please retrieve images that meet the requirements according to the image information". The specific form of the modality pair of retrieval instructions in the embodiments of the present application is not limited herein.

[0188] It can be understood that since the computer device can distinguish different retrieval tasks through modal instructions, the features of all modalities of the retrieval data in the retrieval dataset can be stored in the same database to improve the data management efficiency. The retrieval data in the retrieval dataset can also be stored in different databases, and the embodiments of the present application do not limit this here.

[0189] Step S205: Based on the modal feature vectors included in the retrieval instructions for S modalities, perform feature retrieval in N retrieval datasets to obtain the retrieval results respectively corresponding to the retrieval instructions for S modalities; one retrieval result is obtained by performing feature retrieval on the modal feature vectors included in the retrieval instructions for one modality in the retrieval dataset corresponding to the included modal identifier; one retrieval result includes the media tags carried by the retrieved data matched through feature retrieval.

[0190] Specifically, please also refer to Figure 5 , such as Figure 5 shown, the computer device can input the retrieval instructions for S modalities into the feature retrieval layer of the hybrid retrieval model. The feature retrieval layer can be a multimodal large language model (Large Language Model, LLM). The LLM can use large language models such as Llama (Large Language Model Meta Artificial Intelligence, a multimodal large language model) and qwen (Qwen Language Model, a multimodal large language model). Or, a pre-trained multimodal large language model can be directly used, such as the qwen-vl (Qwen Vision-Language Model, a pre-trained visual multimodal large language model) model, etc.

[0191] In the feature retrieval layer, the computer device can perform feature retrieval on the modal feature vectors included in the retrieval instructions for S modalities in N retrieval datasets to obtain the retrieval results respectively corresponding to the retrieval instructions for S modalities. Taking the retrieval instructions for S modalities including the retrieval instructions for modality pair B j as an example, the computer device can determine the retrieval dataset associated with the modal identifier included in the retrieval instructions for modality pair B j as the target retrieval dataset, obtain the retrieval feature vectors respectively corresponding to R retrieval data in the target retrieval dataset, and compare the modal feature vectors included in the retrieval instructions for modality pair B j with the R retrieval feature vectors respectively to obtain R retrieval scores. Among them, one retrieval data includes one or more annotated media tags.

[0192] The computer device can sort R retrieval scores. For example, it can be sorted in descending order according to the numerical value of the retrieval scores. Among the sorted R retrieval scores, the computer device can obtain the K retrieval scores with larger numerical values. K can be a preset fixed value. The computer device can determine the media tags of the retrieval data corresponding to the K retrieval scores as the modality pair retrieval instruction B j The corresponding retrieval result.

[0193] It can be understood that through the modality pair retrieval instruction, the computer device can flexibly specify the current retrieval task, distinguish the used modality and the modality to be retrieved. The hybrid modality retrieval model can execute the corresponding retrieval task according to the modality pair retrieval instruction and return the retrieval result of each modality pair retrieval instruction.

[0194] Step S206: Count the number of the same media tags in the S retrieval results to obtain S statistical results; determine the media tags corresponding to the statistical results in the S statistical results where the number of media tags is greater than or equal to the quantity threshold as the target retrieval results.

[0195] Specifically, the computer device can perform a fusion process on the media tags in the S retrieval results to obtain the target retrieval results. The fusion process can be to take the intersection or union of the media tags in the S retrieval results, or to perform a quantity statistics on the media tags in the S retrieval results to determine the target retrieval results.

[0196] For ease of understanding, taking the quantity statistics of the media tags in the S retrieval results as an example, please also refer to Figure 8 , Figure 8 which is a scenario schematic diagram provided by an embodiment of the present application Figure 3 ,such as Figure 8 shown, the computer device can count the number of the same media tags in the S retrieval results (including retrieval result 1, retrieval result 2,..., retrieval result S) to obtain S statistical results, and determine the media tags corresponding to the statistical results in the S statistical results where the number of media tags is greater than or equal to the quantity threshold as the target retrieval results.

[0197] For example, the S retrieval results may include retrieval result 1 (including media tag 1, media tag 2, media tag 3), retrieval result 2 (including media tag 1 and media tag 4), and retrieval result 3 (including media tag 1 and media tag 2). The computer device can count the media tags in retrieval result 1, retrieval result 2, and retrieval result 3 to obtain the statistical result corresponding to media tag 1 (quantity is 3), the statistical result corresponding to media tag 2 (quantity is 2), the statistical result corresponding to media tag 3 (quantity is 1), and the statistical result corresponding to media tag 4 (quantity is 1), as Figure 8As shown, the computer device can determine media tag 1 and media tag 2 with the number of media tags in the result being greater than or equal to a number threshold (e.g., 2) as the target retrieval result.

[0198] Optionally, if the S retrieval results together include G mutually distinct media tags, the computer device can obtain the retrieval scores respectively corresponding to each media tag in the S retrieval results. For example, the S retrieval results can include retrieval result 1, retrieval result 2, and retrieval result 3. Among them, retrieval result 1 contains media tag 1 (retrieval score: 88), media tag 2 (retrieval score: 72), and media tag 3 (retrieval score: 87); retrieval result 2 contains media tag 1 (retrieval score: 92) and media tag 4 (retrieval score: 88); retrieval result 3 contains media tag 1 (retrieval score: 78) and media tag 2 (retrieval score: 77). The computer device can accumulate the retrieval scores of the same media tags in the S retrieval results to obtain the total retrieval scores respectively corresponding to the G media tags. Among them, the total retrieval score corresponding to media tag 1 is 258, the total retrieval score corresponding to media tag 2 is 149, and the total retrieval score corresponding to media tag 3 is 78. The computer device can determine media tag 1 with the total retrieval score greater than or equal to a score threshold (e.g., 200) among the G media tags as the target retrieval result.

[0199] In the embodiments of the present application, by extracting features from media data, M modality feature vectors with different modalities are obtained. The M modality feature vectors can include rich single-modal features and multi-modal features, enabling the model to more comprehensively understand the input content. N retrieval data sets with different modality identifiers are obtained. One modality identifier can correspond to one modality feature, and the modalities of the retrieval data in different retrieval data sets are different from each other. Generate modality pair retrieval instructions corresponding to the N retrieval data sets for each modality feature vector, obtaining S modality pair retrieval instructions. Among them, each modality pair retrieval instruction includes a modality feature vector and a modality identifier. Based on the modality feature vectors included in the S modality pair retrieval instructions respectively, perform feature retrieval in the N retrieval data sets to obtain the retrieval results corresponding to the S modality pair retrieval instructions respectively. Each retrieval result includes the media tags carried by the retrieval data matched through feature retrieval. It can be seen that the present application can distinguish the retrieval tasks of different modality pairs through the modality pair retrieval instructions, and at the same time support various modality feature retrievals such as same modality, cross modality, and multi modality, flexibly adapting to the retrieval requirements between different modalities. In the retrieval task corresponding to a single modality pair retrieval instruction, the present application can learn and capture the fine-grained differences on a single modality feature by distinguishing the different features of the M modality feature vectors, and by indirectly retrieving the retrieval data similar to the media data in the retrieval data set, it can avoid the information limitation brought by directly retrieving the text semantics of the tags, reduce the risk of tag mismatch, and improve the accuracy of the retrieval results. It also realizes the integration of different modality retrieval tasks. Only one model is required to perform mutual retrieval between different modalities, and finally generate the retrieval results corresponding to the S modality pair retrieval instructions respectively, without training retrieval models for multiple modality pairs, reducing the training cost, reducing the complexity of the system, and the direct output of a single model also improves the efficiency of the retrieval task. At the same time, comprehensive feature retrieval of the M modality feature vectors of the media data is performed in the N retrieval data sets, making full use of the multiple modality features of the media data, obtaining the retrieval results corresponding to the S modality retrieval instructions respectively, and performing fusion processing on the S retrieval results to obtain the target retrieval result, which can comprehensively consider the retrieval results corresponding to the S modality pairs respectively, and accurately mine the media tags related to the media data both in depth and breadth, thereby improving the accuracy of the retrieval results for the media tags.

[0200] On the other hand, the embodiments of the present application utilize the powerful capabilities of the multi-modal large language model to extract more representative features when processing media data with multiple modality features, thereby improving the performance of the overall system. By integrating the retrieval tasks of multiple modality pairs, the complexity of the system is simplified and the operation efficiency is improved. At the same time, through the alignment processing layer, the M modality feature vectors can be represented in the same feature space, enabling the model to understand the detailed differences between different modality features and improving the model's feature understanding ability.

[0201] See also Figure 9 , Figure 9 This is a scenario diagram provided by the embodiment of the present application. Figure 4 ,like Figure 9 As shown, label prediction for media data may include a recall phase and a sorting phase. The recall phase may refer to the above Figure 4 The content of steps S201 to S206 in the corresponding embodiment, that is, feature retrieval of media data in several retrieval data sets is performed through a hybrid modal retrieval model to obtain target retrieval results (recall results). The sorting stage may refer to further screening and sorting of media tags in the target retrieval results through a sorting model, wherein the sorting stage may also be a multimodal large language model, which may have the same model architecture as the hybrid modal retrieval model, but its training objectives are different. The training of the sorting model adopts the autoregressive training method of the generative model, that is, predicting the next token. The content of the sorting model is not limited in the embodiments of the present application.

[0202] In the sorting stage, the computer device can obtain a label prompt word, which is used to ask whether the media label is related to the media data through a prompt. For example, it can specifically refer to the text "<video information> text: <text information> label: <label>, please ask whether the above label is a relevant label for the video, only output yes or no". Among them, the label prompt word is used to instruct the sorting model to generate a matching verification result between the media data and the target retrieval result.

[0203] The computer device can input media data, target search results and label prompt words into the sorting model, extract features of the media data through the sorting model, obtain a mixed modal feature vector, and generate similarity scores corresponding to F sorting contents in the sorting model based on the mixed modal feature vector. Among them, a sorting content includes one or more marked media tags, and the F sorting contents can refer to the vocabulary in the sorting model. Its dimension is [1,v_size], where v_size is the length of the vocabulary, that is, F.

[0204] The computer device can generate a matching verification result corresponding to each media label in the target retrieval result based on the similarity scores, label prompt words, and target retrieval results respectively corresponding to F sorted contents. The matching verification result is used to indicate whether the media data is consistent or inconsistent with the media label in the target retrieval result. For example, the matching verification result can be a token representing "yes" or "no". The computer device can determine the media label corresponding to the matching verification result indicating consistency as the label to be sorted, and perform normalization processing (Softmax) on the similarity scores corresponding to the sorted contents with the label to be sorted among the F sorted contents to generate the correlation score of the label to be sorted. The correlation score can represent the degree of the media label and the media data. The higher the score, the stronger the correlation between the label and the media content.

[0205] The computer device can sort the labels to be sorted in the target retrieval result based on the correlation scores. The computer device can determine the sorting result among the sorted labels to be sorted based on a correlation threshold of a certain correlation score. The correlation score of the media label in the sorting result can be greater than or equal to the correlation threshold, and the media label with the highest correlation score can be determined as the main label of the media data, and the media labels other than the main label in the sorting result can be determined as secondary labels.

[0206] It can be understood that after the target retrieval result is obtained through preliminary recall by the hybrid modality retrieval model, the sorting model further performs fine sorting on the recall result (target retrieval result). By scoring and sorting the correlation scores of each media label, the sorting model can screen out the most relevant media labels. This step further improves the metrics of the model, making the media labels in the finally output sorting result more accurate and reliable.

[0207] In the embodiments of the present application, the hybrid modality retrieval model extracts rich multi-modal features and single-modal features from the input media data. These features not only include text information but also visual information, enabling the model to more comprehensively understand the input content. In the retrieval stage, the hybrid modality retrieval model uses these multi-modal and single-modal features to retrieve content similar to the input content from the retrieval library. In the sorting stage, the sorting model can further sort the retrieved results, calculate the correlation score of each media label based on the media data and the label information in the sorting model vocabulary, so as to obtain a more accurate label prediction result.

[0208] Please refer to Figure 10 , Figure 10 which is a schematic flow of a data processing method provided by the embodiments of the present application Figure 3 This data processing method can be executed by a computer device, and the computer device can be such as Figure 1The service server 100 shown or any terminal device in the terminal device cluster may be, for example, the terminal device 10a. The following will take the data processing method executed by a computer device as an example for explanation. The data processing method may at least include the following steps S301 to S305:

[0209] Step S301, obtaining sample media data, and inputting the sample media data into an initial modality retrieval model;

[0210] For details, please refer to the above Figure 4 The specific description of step S201 of the corresponding embodiment will not be repeated here in the embodiment of the present application.

[0211] Step S302: extracting features from the sample media data through the initial modality retrieval model to obtain M sample modality feature vectors; M is a positive integer, and the modalities of the M sample modality feature vectors are different from each other;

[0212] For details, please refer to the above Figure 4 The specific description of step S202 and step S203 of the corresponding embodiment will not be repeated here in the embodiment of the present application.

[0213] Step S303, generating sample modality pair retrieval instructions corresponding to N retrieval data sets respectively for each sample modality feature vector, and obtaining S sample modality pair retrieval instructions; N is a positive integer, and S is the product of M and N; the N retrieval data sets have different modality identifiers, and the modalities of the retrieval data in different retrieval data sets are different, and each sample modality pair retrieval instruction includes a sample modality feature vector and a modality identifier;

[0214] For details, please refer to the above Figure 4 The specific description of step S204 of the corresponding embodiment will not be repeated here in the embodiment of the present application.

[0215] Step S304, based on the sample modality feature vectors respectively included in the retrieval instructions of the S sample modality pairs, feature retrieval is performed in the N retrieval data sets to obtain sample retrieval results respectively corresponding to the retrieval instructions of the S sample modality pairs; one sample retrieval result is obtained by performing feature retrieval in the retrieval data set corresponding to the modality identifier included based on the sample modality feature vector included in the retrieval instruction of one sample modality pair; one sample retrieval result includes the media tag carried by the retrieval data matched by the feature retrieval;

[0216] For details, please refer to the above Figure 4 The specific description of step S205 of the corresponding embodiment will not be repeated here in this embodiment of the present application.

[0217] Step S305: Perform label fusion processing on the S sample retrieval results to obtain the target sample retrieval result. Based on the target sample retrieval result and the sample labels of the sample media data, adjust the initial modal retrieval model to obtain a hybrid modal retrieval model; the hybrid modal retrieval model is used to generate the target retrieval result of the media data.

[0218] Specifically, the computer device can perform label fusion processing on the S sample retrieval results to obtain the target sample retrieval result. For the content of the target sample fusion result, refer to the specific description of step S206 in the corresponding embodiment above. This application embodiment will not be elaborated here. Figure 4 For the corresponding embodiment's step S206, it will not be repeated here.

[0219] The computer device can generate a model loss value based on the target sample retrieval result and the sample labels of the sample media data. The model loss value can be a noise contrastive estimation loss value (InfoNCE Loss). The calculation process of the model loss value L InfoNCE can be shown as formula (3):

[0220]

[0221] where B represents the batch size, τ represents the temperature parameter, which is used to control the sharpness of the distribution. The smaller the temperature parameter, the sharper the distribution, that is, the more obvious the contrast effect between sample pairs with high similarity. k(·,·) represents the similarity function, which is used to calculate the similarity between two inputs. LMM represents the initial modal retrieval model, which can be a multi-modal language model. q n represents the query sample, that is, the sample that needs to be compared with other samples in the current batch. c n represents the positive sample, the sample that is relevant or similar to the query sample q n In the noise contrastive learning loss function, the similarity between the positive sample and the query sample is expected to be maximized. c m represents the negative sample, the sample that is different from the query sample q nUnrelated or dissimilar samples. In the noise contrastive learning loss function, the similarity between negative samples and query samples is expected to be minimized. Negative samples typically include all other samples in the batch except for the positive samples. Positive samples are those that are consistent with the target class label, i.e., the target instances that need to be correctly identified. Negative samples are those that are inconsistent with the target class label, i.e., background or non-target instances. For example, in a binary classification task, if the label is "cat", then all samples with the "cat" label are positive samples, and samples without the "cat" label are negative samples.

[0222] The computer device can adjust the initial modal retrieval model based on the model loss value. When the initial modal retrieval model meets the training convergence condition, the initial modal retrieval model that meets the training convergence condition is determined as the hybrid modal retrieval model. Among them, the training convergence condition can refer to that during the training process, if the loss value of the loss function does not decrease significantly after multiple training batches (training rounds) or the training batch reaches a preset maximum value, etc., or a preset performance metric (such as recall rate, accuracy, etc.) is reached. The embodiments of the present application do not limit the training convergence condition here. The hybrid modal retrieval model is used to generate the target retrieval result of the media data.

[0223] Optionally, the embodiments of the present application can also combine sample data of different media types to obtain sample media data. For example, an image data and a text data can be determined as a sample media data, that is, a sample pair. Samples in the same training batch that belong to the same sample pair are positive samples, and samples that are in different sample pairs are negative samples for each other.

[0224] The embodiments of the present application obtain a hybrid modal retrieval model that can process different modal data through joint training of multimodal data pairs (such as related videos and texts). This model can perform retrieval between queries and candidate sets of different modalities. Using instruction fine-tuning, different retrieval tasks are represented by instruction information. By training the hybrid modal retrieval model through contrastive learning, the density estimation problem can be transformed into a binary classification task (distinguishing real samples from noise samples). In contrastive learning, only the similarity of positive and negative sample pairs needs to be calculated, avoiding traversing all data and bypassing the complex calculation of the normalization factor in traditional probability models, significantly reducing the training complexity. At the same time, by maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs, the model can learn more discriminative feature representations. By adjusting the similarity calculation method (such as cosine similarity, Euclidean distance), the trained hybrid modal retrieval model can adapt to feature contrast tasks of different modal data such as text, images, and videos, supporting multi-modal and cross-modal feature retrieval, and improving the feasibility and robustness of the embodiments of the present application.

[0225] Please refer to Figure 11 ,Figure 11 This is a schematic structure of a data processing device provided by an embodiment of the present application. Figure 1 As Figure 11 shown, the data processing device 1 includes a feature extraction module 910, an instruction generation module 920, an instruction retrieval module 930, and a fusion processing module 940.

[0226] The feature extraction module 910 is configured to obtain media data, extract features from the media data, and obtain M modal feature vectors; M is a positive integer, and the modalities of the M modal feature vectors are different from each other;

[0227] The instruction generation module 920 is configured to generate modality-pair retrieval instructions corresponding to N retrieval data sets for each modal feature vector, and obtain S modality-pair retrieval instructions; N is a positive integer, S is the product of M and N; the N retrieval data sets have different modality identifiers, the modalities of the retrieval data in different retrieval data sets are different from each other, and each modality-pair retrieval instruction includes a modal feature vector and a modality identifier;

[0228] The instruction retrieval module 930 is configured to perform feature retrieval in the N retrieval data sets based on the modal feature vectors included in the S modality-pair retrieval instructions respectively, and obtain retrieval results corresponding to the S modality-pair retrieval instructions; a retrieval result is obtained by performing feature retrieval in the retrieval data set corresponding to the modality identifier included in a modality-pair retrieval instruction based on the modal feature vector included in the modality-pair retrieval instruction; a retrieval result includes the media tags carried by the retrieval data matched through feature retrieval;

[0229] The fusion processing module 940 is configured to perform label fusion processing on the S retrieval results to obtain a target retrieval result.

[0230] In a possible implementation manner, the M modal feature vectors include a modal feature vector A i , where i is a positive integer less than or equal to M; when the instruction generation module 920 generates modality-pair retrieval instructions corresponding to N retrieval data sets for each modal feature vector and obtains S modality-pair retrieval instructions, it is specifically configured to perform the following operations:

[0231] Generate N modality-pair retrieval instructions for the N retrieval data sets for the modal feature vector A i ; the modal feature vectors included in the N modality-pair retrieval instructions are all the modal feature vector A i , and the retrieval data sets associated with the N modality-pair retrieval instructions are different from each other;

[0232] Determine the N modality-pair retrieval instructions corresponding to the M modal feature vectors respectively as the S modality-pair retrieval instructions.

[0233] In a possible implementation, the S modality pair retrieval instructions include the modality pair retrieval instruction B j , where j is a positive integer less than or equal to S; when the instruction retrieval module 930 is used to perform feature retrieval on the N retrieval data sets based on the modality feature vectors respectively included in the S modality pair retrieval instructions to obtain the retrieval results respectively corresponding to the S modality pair retrieval instructions, it is specifically used to perform the following operations:

[0234] Determine the retrieval data set associated with the modality identifier included in the modality pair retrieval instruction B j as the target retrieval data set;

[0235] Obtain the retrieval feature vectors respectively corresponding to the R retrieval data in the target retrieval data set, and compare the modality feature vectors included in the modality pair retrieval instruction B j with the R retrieval feature vectors respectively to obtain R retrieval scores; R is a positive integer; one retrieval data includes one or more labeled media tags;

[0236] Perform a sorting process on the R retrieval scores, obtain K retrieval scores from the sorted R retrieval scores, and determine the media tags of the retrieval data respectively corresponding to the K retrieval scores as the retrieval result corresponding to the modality pair retrieval instruction B j ; K is a positive integer less than or equal to R.

[0237] In a possible implementation, when the fusion processing module 940 is used to perform label fusion processing on the S retrieval results to obtain the target retrieval result, it is specifically used to perform the following operations:

[0238] Count the number of the same media tags in the S retrieval results to obtain S statistical results;

[0239] Determine the media tags corresponding to the statistical results in which the number of the media tags in the S statistical results is greater than or equal to the quantity threshold as the target retrieval result.

[0240] In a possible implementation, the S retrieval results together include G mutually different media tags, where G is a positive integer; when the fusion processing module 940 is used to perform label fusion processing on the S retrieval results to obtain the target retrieval result, it is specifically used to perform the following operations:

[0241] Accumulate the retrieval scores of the same media tags in the S retrieval results to obtain the total retrieval scores respectively corresponding to the G media tags;

[0242] Determine the media tags in the G media tags whose total retrieval scores are greater than or equal to the score threshold as the target retrieval result.

[0243] In a possible implementation, the fusion processing module 940 is further configured to perform the following operations:

[0244] Obtain label prompt words, and input the media data, the target retrieval result, and the label prompt words into a sorting model; the label prompt words are used to instruct the sorting model to generate a matching verification result between the media data and the target retrieval result;

[0245] Extract features from the media data through the sorting model to obtain a hybrid modality feature vector, and based on the hybrid modality feature vector, generate similarity scores corresponding to each of the F sorting contents in the sorting model; F is a positive integer, and one sorting content includes one or more labeled media labels;

[0246] Based on the similarity scores corresponding to the F sorting contents respectively, the label prompt words, and the target retrieval result, generate a matching verification result corresponding to each media label in the target retrieval result; the matching verification result is used to indicate whether the media data is consistent or inconsistent with the media label in the target retrieval result;

[0247] Determine the media labels corresponding to the matching verification results indicating consistency as the labels to be sorted, and based on the similarity scores corresponding to the sorting contents having the labels to be sorted among the F sorting contents, generate relevance scores for the labels to be sorted;

[0248] Based on the relevance scores, perform a sorting process on the labels to be sorted in the target retrieval result to obtain a sorting result.

[0249] In a possible implementation, when the feature extraction module 910 is configured to extract features from media data to obtain M modality feature vectors, it is specifically configured to perform the following operations:

[0250] Input the media data into a hybrid modality retrieval model; the hybrid modality retrieval model includes an encoding layer;

[0251] Extract features from the media data through the encoding layer to obtain P single-modality feature vectors; P is a positive integer less than or equal to M;

[0252] If the media data is video data, perform feature fusion processing on the P single-modality feature vectors to obtain a multi-modality feature vector, and determine the P single-modality feature vectors and the multi-modality feature vector as M modality feature vectors;

[0253] If the media data is text data or image data, determine the P single-modality feature vectors as M modality feature vectors.

[0254] In a possible implementation, the encoding layer includes P single-modal encoders with different modalities; when the feature extraction module 910 is used to perform feature extraction on P single-modal data through the encoding layer to obtain P single-modal feature vectors, it is specifically used to perform the following operations:

[0255] Split the media data according to the media type of the media data to obtain Q single-modal data with different media types; Q is a positive integer less than or equal to P;

[0256] Input the Q single-modal data into the single-modal encoders that match the media types of the Q single-modal data respectively. In the Q single-modal encoders, perform feature extraction on the Q single-modal data respectively to obtain Q single-modal feature vectors;

[0257] If Q is less than P, generate padding feature vectors corresponding to the single-modal encoders other than the Q single-modal encoders among the P single-modal encoders, and determine the Q single-modal feature vectors and the padding feature vectors as P single-modal feature vectors;

[0258] If Q is equal to P, determine the Q single-modal feature vectors as P single-modal feature vectors.

[0259] In a possible implementation, the hybrid-modal retrieval model further includes an alignment processing layer; when the feature extraction module 910 is used to determine the Q single-modal feature vectors and the padding feature vectors as P single-modal feature vectors, it is specifically used to perform the following operations:

[0260] Input the Q single-modal feature vectors and the padding feature vectors into the alignment processing layer;

[0261] Based on the alignment projection matrix corresponding to the alignment processing layer, perform linear transformation on the Q single-modal feature vectors and the padding feature vectors respectively to obtain P aligned feature vectors;

[0262] Determine the P aligned feature vectors as P single-modal feature vectors.

[0263] In a possible implementation, the Q single-modal data includes image data, and the Q single-modal encodings include visual encoders corresponding to the image type; when the feature extraction module 910 is used to perform feature extraction on the Q single-modal data respectively in the Q single-modal encoders to obtain Q single-modal feature vectors, it is specifically used to perform the following operations:

[0264] In the visual encoder, perform image segmentation on the image data to obtain T unit sampled images; T is a positive integer;

[0265] Perform feature extraction on the T unit sampled images to obtain T unit visual feature vectors;

[0266] Based on the rearrangement factor, the vector elements representing the spatial channels in the T unit visual feature vectors are remapped to the feature channels to obtain T visual feature vectors to be matched; the rearrangement factor is determined based on the number T of the unit visual feature vectors, and the vector length of the unit visual feature vectors is greater than the vector length of the visual feature vectors to be matched.

[0267] Perform attention processing on the T visual feature vectors to be matched to obtain a global visual feature vector, and determine the global visual feature vector as the single-modal feature vector corresponding to the image data.

[0268] In a possible implementation manner, when the feature extraction module 910 is used to perform attention processing on the T visual feature vectors to be matched to obtain a global visual feature vector, it is specifically used to perform the following operations:

[0269] Obtain the query parameter matrix, key parameter matrix, and value parameter matrix in the visual encoder; the query parameter matrix, key parameter matrix, and value parameter matrix are all matrices composed of learnable parameters.

[0270] Perform vector concatenation on the T visual feature vectors to be matched to obtain a visual feature sequence to be matched, perform dot product operation on the visual feature sequence to be matched and the query parameter matrix to obtain a query vector, perform dot product operation on the visual feature sequence to be matched and the key parameter matrix to obtain a key vector, and perform dot product operation on the visual feature sequence to be matched and the value parameter matrix to obtain a value vector.

[0271] Generate an attention score vector based on the query vector and the key vector, perform dimensionality reduction processing on the attention score vector based on the vector length of the key vector, perform normalization processing on the dimensionally reduced attention score vector to obtain an attention weight vector, and perform dot product operation on the attention weight vector and the value vector to obtain a global visual feature vector.

[0272] In the embodiments of the present application, by extracting features from media data, M modal feature vectors with different modalities are obtained. The M modal feature vectors can include rich unimodal features and multimodal features, enabling the model to more comprehensively understand the input content. N retrieval data sets with different modality identifiers are obtained. One modality identifier can correspond to one modal feature, and the modalities of the retrieval data in different retrieval data sets are different. Generate modality-pair retrieval instructions corresponding to the N retrieval data sets for each modal feature vector, obtaining S modality-pair retrieval instructions. Among them, each modality-pair retrieval instruction includes a modal feature vector and a modality identifier. Based on the modal feature vectors included in the S modality-pair retrieval instructions, perform feature retrieval in the N retrieval data sets to obtain the retrieval results corresponding to the S modality-pair retrieval instructions respectively. Each retrieval result includes the media tags carried by the retrieval data matched through feature retrieval. It can be seen that the present application can distinguish the retrieval tasks of different modality pairs through the modality-pair retrieval instructions, and at the same time support various modality feature retrievals such as unimodal, cross-modal, and multimodal, flexibly adapting to the retrieval requirements between different modalities. For the retrieval task corresponding to a single modality-pair retrieval instruction, the present application can learn and capture the fine-grained differences on a single modal feature by distinguishing the different features of the M modal feature vectors, and by indirectly retrieving the retrieval data similar to the media data in the retrieval data set, it can avoid the information limitations brought by directly retrieving the text semantics of the tags, reduce the risk of tag misalignment, and improve the accuracy of the retrieval results. It also realizes the integration of different modality retrieval tasks. Only one model can perform mutual retrieval between different modalities, and finally generate the retrieval results corresponding to the S modality-pair retrieval instructions respectively, without training retrieval models for multiple modality pairs, reducing the training cost, reducing the complexity of the system, and the direct output of a single model also improves the efficiency of the retrieval task. At the same time, comprehensively perform feature retrieval on the M modal feature vectors of the media data in the N retrieval data sets, make full use of the multiple modality features of the media data, obtain the retrieval results corresponding to the S modality retrieval instructions respectively, and perform fusion processing on the S retrieval results to obtain the target retrieval result, which can comprehensively consider the retrieval results corresponding to the S modality pairs respectively, and accurately mine the media tags related to the media data both in depth and breadth, thereby improving the accuracy of the retrieval results for the media tags.

[0273] On the other hand, the embodiments of the present application utilize the powerful capabilities of the multimodal large language model. When processing media data with multiple modal features, more representative features are extracted, thereby improving the performance of the overall system. By integrating the retrieval tasks of multiple modality pairs, the complexity of the system is simplified and the operation efficiency is improved. At the same time, through the alignment processing layer, the M modal feature vectors can be represented in the same feature space, enabling the model to understand the detailed differences of different modal features and improving the model's feature understanding ability.

[0274] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0275] Please refer to Figure 12 , Figure 12 is a structural schematic diagram of a data processing device provided by the embodiments of the present application Figure 2 . As Figure 12 shown, the data processing device 2 includes a sample data acquisition module 1010, a sample feature extraction module 1020, a sample instruction generation module 1030, a sample instruction retrieval module 1040, and a sample training processing module 1050.

[0276] The sample data acquisition module 1010 is used to acquire sample media data and input the sample media data into the initial modality retrieval model;

[0277] The sample feature extraction module 1020 is used to extract features from the sample media data through the initial modality retrieval model to obtain M sample modality feature vectors; M is a positive integer, and the modalities of the M sample modality feature vectors are different from each other;

[0278] The sample instruction generation module 1030 is used to generate sample modality pair retrieval instructions corresponding to N retrieval data sets for each sample modality feature vector, obtaining S sample modality pair retrieval instructions; N is a positive integer, S is the product of M and N; the N retrieval data sets have different modality identifiers, the retrieval data in different retrieval data sets have different modalities, and each sample modality pair retrieval instruction includes a sample modality feature vector and a modality identifier;

[0279] The sample instruction retrieval module 1040 is used to perform feature retrieval in the N retrieval data sets based on the sample modality feature vectors included in the S sample modality pair retrieval instructions respectively, obtaining sample retrieval results corresponding to the S sample modality pair retrieval instructions respectively; a sample retrieval result is obtained by performing feature retrieval in the retrieval data set corresponding to the modality identifier included in a sample modality pair retrieval instruction based on the sample modality feature vector included in the sample modality pair retrieval instruction; a sample retrieval result includes the media tags carried by the retrieval data matched through feature retrieval;

[0280] The sample training processing module 1050 is used to perform label fusion processing on the S sample retrieval results to obtain the target sample retrieval results, and based on the target sample retrieval results and the sample labels of the sample media data, adjust the initial modal retrieval model to obtain a hybrid modal retrieval model; the hybrid modal retrieval model is used to generate the target retrieval results of the media data.

[0281] In the embodiment of the present application, through the joint training of multimodal data pairs (such as related videos and texts), a hybrid modal retrieval model capable of processing different modal data is obtained. This model can perform retrieval between queries and candidate sets of different modalities. Using instruction fine-tuning, different retrieval tasks are represented by instruction information. By training the hybrid modal retrieval model through contrastive learning, the density estimation problem can be transformed into a binary classification task (distinguishing real samples from noise samples). In contrastive learning, only the similarity of positive and negative sample pairs needs to be calculated, avoiding the traversal of all data and bypassing the complex calculation of the normalization factor in traditional probability models, significantly reducing the training complexity. At the same time, by maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs, the model can learn more discriminative feature representations. By adjusting the similarity calculation method (such as cosine similarity, Euclidean distance), the trained hybrid modal retrieval model can adapt to feature comparison tasks of different modal data such as text, images, and videos, support multi-modal and cross-modal feature retrieval, and improve the feasibility and robustness of the embodiment of the present application.

[0282] In the embodiment of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the functions of the module or unit.

[0283] Please refer to Figure 13 , Figure 13 which is a schematic structural diagram of a computer device provided by the embodiment of the present application. As Figure 13As shown in the figure, the computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the computer device 1000 may further include: a user interface 1003 and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 may also be at least one storage device located far from the aforementioned processor 1001. As Figure 13 shown, the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0284] In the computer device 1000 as Figure 13 shown, the network interface 1004 may provide a network communication network element; while the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 may be used to call the device control application program stored in the memory 1005.

[0285] When the computer device 1000 is used to implement the data processing device 1, it performs:

[0286] Obtain media data, perform feature extraction on the media data to obtain M modal feature vectors; M is a positive integer, and the modalities of the M modal feature vectors are different from each other;

[0287] Generate modality-pair retrieval instructions corresponding to N retrieval data sets for each modal feature vector to obtain S modality-pair retrieval instructions; N is a positive integer, S is the product of M and N; the N retrieval data sets have different modality identifiers, the retrieval data in different retrieval data sets have different modalities, and each modality-pair retrieval instruction includes a modal feature vector and a modality identifier;

[0288] Based on the modal feature vectors respectively included in the S modality-pair retrieval instructions, perform feature retrieval in the N retrieval data sets to obtain retrieval results corresponding to the S modality-pair retrieval instructions respectively; a retrieval result is obtained by performing feature retrieval in the retrieval data set corresponding to the modality identifier included in a modality-pair retrieval instruction based on the modal feature vector included in the modality-pair retrieval instruction; a retrieval result includes the media tags carried by the retrieved data matched through feature retrieval;

[0289] Perform label fusion processing on S retrieval results to obtain target retrieval results.

[0290] When the computer device 1000 is used to implement the data processing device 2, it executes:

[0291] Obtain sample media data and input the sample media data into the initial modal retrieval model;

[0292] Extract features from the sample media data through the initial modal retrieval model to obtain M sample modal feature vectors; M is a positive integer, and the modalities of the M sample modal feature vectors are different from each other;

[0293] Generate sample modal pair retrieval instructions corresponding to N retrieval data sets for each sample modal feature vector to obtain S sample modal pair retrieval instructions; N is a positive integer, S is the product of M and N; the N retrieval data sets have different modality identifiers, the modalities of the retrieval data in different retrieval data sets are different from each other, and each sample modal pair retrieval instruction includes a sample modal feature vector and a modality identifier;

[0294] Based on the sample modal feature vectors included in the S sample modal pair retrieval instructions respectively, perform feature retrieval in the N retrieval data sets to obtain the sample retrieval results corresponding to the S sample modal pair retrieval instructions respectively; a sample retrieval result is obtained by performing feature retrieval in the retrieval data set corresponding to the modality identifier included in a sample modal pair retrieval instruction based on the sample modal feature vector included in the sample modal pair retrieval instruction; a sample retrieval result includes the media labels carried by the retrieved data that match through feature retrieval;

[0295] Perform label fusion processing on the S sample retrieval results to obtain a target sample retrieval result, and based on the target sample retrieval result and the sample label of the sample media data, adjust the initial modal retrieval model to obtain a hybrid modal retrieval model; the hybrid modal retrieval model is used to generate the target retrieval result of the media data.

[0296] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the data processing method in any of the foregoing Figure 3 , Figure 4 and Figure 10 corresponding embodiments, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.

[0297] In addition, it should be pointed out here that: the embodiments of the present application also provide a computer-readable storage medium, and a computer program is stored in the above computer-readable storage medium. When the above processor executes the above computer program, it can execute the foregoing Figure 3 , Figure 4 andFigure 10 The description of the above data processing method in any corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.

[0298] The computer-readable storage medium may be the data processing device provided in any of the aforementioned embodiments or the internal storage unit of the computer device, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (smart media card, SMC), a secure digital (secure digital, SD) card, a flash card (flash card), etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been displayed or is to be displayed.

[0299] In addition, it should be noted that: the embodiment of the present application also provides a computer program product, the computer program product includes a computer program, the computer program is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above Figure 3 , Figure 4 and Figure 10 Any method provided by the corresponding embodiment.

[0300] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, device, product, or equipment that includes a series of steps or units is not limited to the listed steps or modules, but optionally includes steps or modules that are not listed, or optionally includes other step units inherent to these processes, methods, devices, products, or equipment.

[0301] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to network elements in the above description. Whether these network elements are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described network elements for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0302] The methods and related devices provided in the embodiments of this application are described with reference to the method flowcharts and / or structural schematic diagrams provided in the embodiments of this application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic one block or multiple blocks.

[0303] The steps in the methods of the embodiments of this application can be adjusted, combined, and deleted according to actual needs.

[0304] The modules in the devices of the embodiments of this application can be combined, divided, and deleted according to actual needs.

[0305] The above-disclosed are only the preferred embodiments of this application. Of course, the scope of the rights of this application cannot be limited thereby. Therefore, equivalent changes made according to the claims of this application still fall within the scope covered by this application.

Claims

1. A data processing method, characterized in that, Including: Obtain media data, perform feature extraction on the media data to obtain M modal feature vectors; M is a positive integer, and the modalities of the M modal feature vectors are different from each other; Generate modality-pair retrieval instructions corresponding to N retrieval data sets for each modal feature vector, obtaining S modality-pair retrieval instructions; N is a positive integer, and S is the product of M and N; the N retrieval data sets have different modality identifiers, and the modalities of the retrieval data in different retrieval data sets are different from each other. Each modality-pair retrieval instruction includes a modal feature vector and a modality identifier; Based on the modal feature vectors included in the S modality-pair retrieval instructions respectively, perform feature retrieval in the N retrieval data sets to obtain the retrieval results corresponding to the S modality-pair retrieval instructions respectively; one retrieval result is obtained by performing feature retrieval in the retrieval data set corresponding to the modality identifier included in a modality-pair retrieval instruction based on the modal feature vector included in the modality-pair retrieval instruction; one retrieval result includes the media tags carried by the retrieved data matched through feature retrieval; Perform label fusion processing on the S retrieval results to obtain a target retrieval result.

2. The method according to claim 1, wherein The M modal feature vectors include the modal feature vector A i , where i is a positive integer less than or equal to M; generating, for each modal feature vector, modal pair retrieval instructions respectively corresponding to N retrieval data sets, to obtain S modal pair retrieval instructions, including: Generate the modal feature vector A i Retrieval instructions for N modalities of the N retrieval data sets; the modal feature vectors included in the retrieval instructions for the N modalities are all the modal feature vector A i , the retrieval data sets associated with the retrieval instructions for the N modalities are different from each other; Determine the N modality-pair retrieval instructions corresponding to the M modal feature vectors respectively as S modality-pair retrieval instructions.

3. The method according to claim 1, wherein The S modal pair retrieval instructions include the modal pair retrieval instruction B j , where j is a positive integer less than or equal to S; performing feature retrieval on the N retrieval data sets based on the modal feature vectors respectively included in the S modal pair retrieval instructions to obtain the retrieval results respectively corresponding to the S modal pair retrieval instructions, including: Determine the retrieval data set associated with the modality identifier included in the modality pair retrieval instruction B j as the target retrieval data set; Obtain the retrieval feature vectors corresponding to R retrieval data in the target retrieval dataset respectively, and pair the modality to the retrieval instruction B j Perform feature comparison between the modality feature vectors included and the R retrieval feature vectors respectively to obtain R retrieval scores; R is a positive integer; one retrieval data includes one or more labeled media tags; Sort the R retrieval scores, obtain K retrieval scores from the sorted R retrieval scores, and determine the media tags of the retrieval data corresponding to the K retrieval scores respectively as the retrieval result corresponding to the modality pair retrieval instruction B; K is a positive integer less than or equal to R. j The corresponding retrieval result; K is a positive integer less than or equal to R.

4. The method according to claim 1, characterized in that, The performing label fusion processing on the S retrieval results to obtain a target retrieval result includes: Count the quantities of the same media tags in the S retrieval results to obtain S statistical results; Determine the media tags corresponding to the statistical results in which the quantity of the media tags in the S statistical results is greater than or equal to a quantity threshold as the target retrieval result.

5. The method according to claim 3, characterized in that, The S retrieval results together include G mutually different media tags, G is a positive integer; the performing label fusion processing on the S retrieval results to obtain a target retrieval result includes: Accumulate the retrieval scores of the same media tags in the S retrieval results to obtain the total retrieval scores corresponding to the G media tags respectively; Determine the media tags in the G media tags whose total retrieval scores are greater than or equal to a score threshold as the target retrieval result.

6. The method according to claim 1, characterized in that, It also includes: Obtain label prompt words, and input the media data, the target retrieval result, and the label prompt words into a ranking model; The label prompt words are used to instruct the ranking model to generate a matching verification result between the media data and the target retrieval result; Perform feature extraction on the media data through the ranking model to obtain a mixed-modal feature vector, and based on the mixed-modal feature vector, generate similarity scores corresponding to F ranking contents in the ranking model; F is a positive integer, and one ranking content includes one or more annotated media tags; Based on the similarity scores corresponding to the F ranking contents respectively, the label prompt words, and the target retrieval result, generate matching verification results corresponding to each media tag in the target retrieval result; The matching verification result is used to represent whether the media data is consistent or inconsistent with the media tags in the target retrieval result; Determine the media tags corresponding to the consistent matching verification results as the tags to be sorted, and generate the relevance scores of the tags to be sorted based on the similarity scores of the sorting contents with the tags to be sorted among the F sorting contents; Based on the relevance scores, perform sorting processing on the tags to be sorted in the target retrieval results to obtain a sorting result.

7. The method according to claim 1, characterized in that The extracting M modal feature vectors from the media data includes: Input the media data into a hybrid modal retrieval model; the hybrid modal retrieval model includes an encoding layer; Extract features from the media data through the encoding layer to obtain P single-modal feature vectors; P is a positive integer less than or equal to M; If the media data is video data, perform feature fusion processing on the P single-modal feature vectors to obtain multi-modal feature vectors, and determine the P single-modal feature vectors and the multi-modal feature vectors as M modal feature vectors; If the media data is text data or image data, determine the P single-modal feature vectors as the M modal feature vectors.

8. The method according to claim 7, wherein The encoding layer includes P single-modal encoders with different modalities; the extracting P single-modal feature vectors by respectively performing feature extraction on P single-modal data through the encoding layer includes: Perform data splitting on the media data according to the media type of the media data to obtain Q single-modal data with different media types; Q is a positive integer less than or equal to P; Input the Q single-modal data into the single-modal encoders matching the media types of the Q single-modal data respectively, and respectively perform feature extraction on the Q single-modal data in the Q single-modal encoders to obtain Q single-modal feature vectors; If Q is less than P, generate the padding feature vectors corresponding to the single-modal encoders other than the Q single-modal encoders among the P single-modal encoders, and determine the Q single-modal feature vectors and the padding feature vectors as P single-modal feature vectors; If Q is equal to P, determine the Q single-modal feature vectors as P single-modal feature vectors.

9. The method according to claim 8, wherein The hybrid modal retrieval model further includes an alignment processing layer; the determining the Q single-modal feature vectors and the padding feature vectors as P single-modal feature vectors includes: Input the Q single-modal feature vectors and the padding feature vectors into the alignment processing layer; Based on the alignment projection matrix corresponding to the alignment processing layer, respectively perform linear transformation on the Q single-modal feature vectors and the padding feature vectors to obtain P aligned feature vectors; Determine the P aligned feature vectors as P single-modal feature vectors.

10. The method according to claim 8, wherein The Q single-modal data includes image data, and the Q single-modal encoders include a visual encoder corresponding to the image type; the extracting Q single-modal feature vectors by respectively performing feature extraction on the Q single-modal data in the Q single-modal encoders includes: Perform image segmentation on the image data in the visual encoder to obtain T unit sampled images; T is a positive integer; Extract features from the T unit sampled images to obtain T unit visual feature vectors; Based on the rearrangement factor, remap the vector elements representing the spatial channels in the T unit visual feature vectors to the feature channels to obtain T visual feature vectors to be matched; the rearrangement factor is determined based on the number T of the unit visual feature vectors, and the vector length of the unit visual feature vectors is greater than the vector length of the visual feature vectors to be matched; Perform attention processing on the T visual feature vectors to be matched to obtain a global visual feature vector, and determine the global visual feature vector as the single-modal feature vector corresponding to the image data.

11. The method according to claim 10, wherein The performing attention processing on the T visual feature vectors to be matched to obtain a global visual feature vector includes: Obtain the query parameter matrix, key parameter matrix, and value parameter matrix in the visual encoder; the query parameter matrix, the key parameter matrix, and the value parameter matrix are all matrices composed of learnable parameters; Perform vector concatenation on the T visual feature vectors to be matched to obtain a visual feature sequence to be matched, perform dot product operation on the visual feature sequence to be matched and the query parameter matrix to obtain a query vector, perform dot product operation on the visual feature sequence to be matched and the key parameter matrix to obtain a key vector, and perform dot product operation on the visual feature sequence to be matched and the value parameter matrix to obtain a value vector; Generate an attention score vector based on the query vector and the key vector, perform dimensionality reduction processing on the attention score vector based on the vector length of the key vector, perform normalization processing on the dimensionally reduced attention score vector to obtain an attention weight vector, and perform dot product operation on the attention weight vector and the value vector to obtain a global visual feature vector.

12. A data processing method, characterized in that, including: Obtain sample media data and input the sample media data into the initial modal retrieval model; Through the initial modal retrieval model, extract features from the sample media data to obtain M sample modal feature vectors; M is a positive integer, and the modalities of the M sample modal feature vectors are different from each other; Generate sample modal pair retrieval instructions corresponding to N retrieval data sets for each sample modal feature vector to obtain S sample modal pair retrieval instructions; N is a positive integer, and S is the product of M and N; the N retrieval data sets have different modal identifiers, the retrieval data in different retrieval data sets have different modalities, and each sample modal pair retrieval instruction includes a sample modal feature vector and a modal identifier; Based on the sample modal feature vectors included in the S sample modal pair retrieval instructions, perform feature retrieval in the N retrieval data sets to obtain the sample retrieval results corresponding to the S sample modal pair retrieval instructions respectively; a sample retrieval result is obtained by performing feature retrieval in the retrieval data set corresponding to the modal identifier included in a sample modal pair retrieval instruction based on the sample modal feature vector included in the sample modal pair retrieval instruction; A sample retrieval result includes the media labels carried by the retrieved data matched through feature retrieval; Perform label fusion processing on the retrieval results of S samples to obtain the retrieval results of the target samples. Based on the retrieval results of the target samples and the sample labels of the sample media data, adjust the initial modal retrieval model to obtain a hybrid modal retrieval model; the hybrid modal retrieval model is used to generate the target retrieval results of the media data.

13. A data processing device, characterized in that, Including: A feature extraction module, configured to obtain media data and perform feature extraction on the media data to obtain M modal feature vectors; M is a positive integer, and the modalities of the M modal feature vectors are different from each other; An instruction generation module, configured to generate modality-pair retrieval instructions corresponding to N retrieval data sets for each modal feature vector, obtaining S modality-pair retrieval instructions; N is a positive integer, and S is the product of M and N; the N retrieval data sets have different modality identifiers, the retrieval data in different retrieval data sets have different modalities, and each modality-pair retrieval instruction includes a modal feature vector and a modality identifier; An instruction retrieval module, configured to perform feature retrieval in the N retrieval data sets based on the modal feature vectors respectively included in the S modality-pair retrieval instructions, obtaining the retrieval results corresponding to the S modality-pair retrieval instructions respectively; a retrieval result is obtained by performing feature retrieval on the retrieval data corresponding to the modality identifier included in a modality-pair retrieval instruction based on the modal feature vector included in the modality-pair retrieval instruction; a retrieval result includes the media labels carried by the retrieval data matched through feature retrieval; A fusion processing module, configured to perform label fusion processing on the S retrieval results to obtain the target retrieval results.

14. A data processing device, characterized in that, Including: A sample data acquisition module, configured to obtain sample media data and input the sample media data into the initial modal retrieval model; A sample feature extraction module, configured to perform feature extraction on the sample media data through the initial modal retrieval model to obtain M sample modal feature vectors; M is a positive integer, and the modalities of the M sample modal feature vectors are different from each other; A sample instruction generation module, configured to generate sample modality-pair retrieval instructions corresponding to N retrieval data sets for each sample modal feature vector, obtaining S sample modality-pair retrieval instructions; N is a positive integer, and S is the product of M and N; the N retrieval data sets have different modality identifiers, the retrieval data in different retrieval data sets have different modalities, and each sample modality-pair retrieval instruction includes a sample modal feature vector and a modality identifier; A sample instruction retrieval module, configured to perform feature retrieval in the N retrieval data sets based on the sample modal feature vectors respectively included in the S sample modality-pair retrieval instructions, obtaining the sample retrieval results corresponding to the S sample modality-pair retrieval instructions respectively; a sample retrieval result is obtained by performing feature retrieval on the retrieval data corresponding to the modality identifier included in a sample modality-pair retrieval instruction based on the sample modal feature vector included in the sample modality-pair retrieval instruction; A sample retrieval result includes the media labels carried by the retrieval data matched through feature retrieval; A sample training processing module is used to perform label fusion processing on the retrieval results of S samples to obtain target sample retrieval results, and based on the target sample retrieval results and the sample labels of the sample media data, adjust the initial modal retrieval model to obtain a hybrid modal retrieval model; the hybrid modal retrieval model is used to generate target retrieval results of media data.

15. A computer device, characterized in that, It includes: a processor, a memory, and a network interface; The processor is connected to the memory and the network interface. Among them, the network interface is used to provide data communication functions, the memory is used to store computer programs, and the processor is used to call the computer programs so that the computer device executes the method according to any one of claims 1-12.

16. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is suitable for being loaded and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-12.

17. A computer program product, characterized in that, The computer program product includes a computer program, the computer program is stored in a computer-readable storage medium, and is suitable for being read and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-12.