Data processing method and apparatus, storage medium, and device
By performing feature encoding and fusion across multiple modalities and generating explicit fused feature representations using gating vectors, the problem of single-modal feature extraction is solved, thereby improving the accuracy of feature extraction and business processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2021-08-09
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies can only extract features from a single modality, resulting in generated objects with limited features and missing feature information across multiple dimensions, which affects the accuracy of business processing results.
By acquiring feature representations of the identified object in multiple modalities, feature encoding and fusion are performed. Gating vectors are used to highlight the explicit contrast relationship between modalities, generating explicit fused feature representations for identifying the similarity between the identified object and the object to be matched.
It improves the accuracy of feature extraction, ensures the accuracy of business processing results, and enables better identification of the correlation and similarity between objects.
Smart Images

Figure CN115705415B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a data processing method, apparatus, storage medium and device. Background Technology
[0002] Each object possesses multiple modalities of feature representation. These modalities are both independent and potentially interconnected. A modality is a way of representing information, such as images, text, and sound. Multimodal fusion aims to mimic the object perception and understanding process by leveraging the complementarity between modalities to eliminate redundant information or supplement missing information in a particular modality. Multimodality refers to various combinations of two or more modalities. For example, a scene can be perceived through visual, auditory, and even tactile modal signals, thereby obtaining rich scene perception information.
[0003] Currently, feature extraction can only be performed on the representation information of a single modality to generate object features. This results in relatively simple object features that are prone to missing feature information in many dimensions. When performing business processing based on these object features (such as obtaining the similarity between objects), the accuracy of the business processing results often cannot be guaranteed. Summary of the Invention
[0004] The technical problem to be solved by the embodiments of this application is to provide a data processing method, apparatus, storage medium, and device that can accurately obtain the data of the identified object in mode M. i The explicit fusion feature representation is used to ensure the accuracy of business processing results.
[0005] One embodiment of this application provides a data processing method, including:
[0006] Obtain the feature representations of the object to be identified in M modalities, and encode the feature representations in each of the M modalities to obtain the encoded features corresponding to the feature representations in the M modalities respectively; M is a positive integer;
[0007] Feature fusion is performed on M encoded features to obtain candidate fused feature representations corresponding to M modalities; the M modalities include modality M i , where i is a positive integer less than or equal to M;
[0008] Obtaining mode M i The corresponding gating vector; the gating vector is used to highlight different recognition objects in modality M. i The explicit contrast relationship between the similarity of objects;
[0009] According to mode M i The corresponding gating vector is used to encode the candidate fusion feature representation to obtain the object in modality M.i Explicit fusion feature representation under the following conditions; object recognition in modality M i The explicit fusion feature representation is used to identify the object to be identified and the object to be matched in mode M. i The similarity of objects.
[0010] Specifically, feature encoding is performed on the feature representations under each of the M modalities to obtain the encoded features corresponding to the feature representations under each of the M modalities, including:
[0011] Input the feature representations of the identified object in M modalities into the target autoencoder model;
[0012] Obtain the mode M in the target autoencoder model i The corresponding encoding weight matrix in the sub-encoder is used to apply the encoding weight matrix to mode M. i The feature representation below is used for feature encoding to obtain mode M. i The features below represent the corresponding candidate encoded features;
[0013] By using activation functions, the candidate coding features corresponding to the feature representations under each of the M modalities are activated respectively, thus obtaining the coding features corresponding to the feature representations under each of the M modalities.
[0014] Specifically, feature fusion is performed on the M encoded features to obtain candidate fused feature representations corresponding to the M modalities, including:
[0015] Input the M encoded features into the main encoder of the target autoencoder model, and perform attention feature encoding on the M encoded features to obtain the attention encoded features corresponding to the M encoded features respectively;
[0016] The M attention-encoded features are concatenated to obtain candidate fusion feature representations corresponding to the M modalities.
[0017] Specifically, attention feature encoding is performed on the M encoded features to obtain the attention encoded features corresponding to the M encoded features, including:
[0018] The main encoder generates a query vector, a key vector, and a value vector for each of the M encoded features;
[0019] Obtain the encoded feature T i Corresponding query vector and encoding feature T i The dot product of the eigenvectors of the corresponding key vectors is used to determine the encoded feature T. i The corresponding attention matrix; encoded feature T i It belongs to M encoded features, where i is a positive integer less than or equal to M;
[0020] Based on coding feature T i The corresponding attention matrix, for encoding feature T i Attention feature encoding is performed to obtain encoded feature T. i The corresponding attention encoding features.
[0021] Among them, mode M i The corresponding gating vectors include dominant sub-vectors and shielding sub-vectors;
[0022] According to mode M i The corresponding gating vector is used to encode the candidate fusion feature representation to obtain the feature recognition in mode M. i The explicit fusion feature representations include:
[0023] By using explicit sub-vectors, vector dot product is performed on the candidate fusion feature representations to obtain the explicit fusion features in the candidate fusion feature representations;
[0024] By using the masked sub-vector, vector dot product is performed on the candidate fusion feature representation to obtain the masked fusion feature in the candidate fusion feature representation;
[0025] The explicit fusion features and the masked fusion features are concatenated to obtain the object in modality M. i The explicit fusion feature representation.
[0026] The aforementioned data processing methods also include:
[0027] Obtain the object to be matched and identified in mode M i The explicit fusion feature representation;
[0028] The object to be matched and identified in mode M i The explicit fusion feature representation and recognition of the object in modality M i The explicit fusion feature representations are multiplied by vector dot product to obtain the vector dot product result;
[0029] The result of the vector dot product is used to determine the object similarity between the identified object and the object to be matched.
[0030] One embodiment of this application provides a data processing method, including:
[0031] Obtain the initial autoencoder model and S sample recognition objects; the S sample recognition objects include sample recognition object S j S is a positive integer, and j is a positive integer less than S.
[0032] The initial autoencoder model is used to identify the sample object S. j Feature fusion is performed on the encoded features of samples in M modalities to obtain the sample identification object S. jThe corresponding predicted candidate fusion feature representation, based on the sample identification object S j The corresponding predicted candidate fusion feature representation generates the first loss function; the M modes include mode M i M is a positive integer, and i is a positive integer less than or equal to M.
[0033] Obtaining mode M i The corresponding gate vector, based on mode M i The corresponding gating vectors are used to encode the predicted candidate fusion feature representations of the S sample recognition objects, respectively, to generate the S sample recognition objects in mode M. i The corresponding predicted explicit fusion feature representations are shown below, based on modality M. i The corresponding explicit contrastive labels and the S sample identification objects in modality M i The corresponding predicted explicit fusion feature representations are used to generate the second loss function;
[0034] Based on the first loss function and the second loss function, the model parameters of the initial autoencoder model are adjusted to obtain the target autoencoder model; the target autoencoder model is used to predict the explicit fusion feature representations corresponding to the feature representations of the identified object in M modalities.
[0035] Among them, the initial autoencoder model is used to identify the sample object S. j Feature fusion is performed on the encoded features of samples in M modalities to obtain the sample identification object S. j The corresponding predicted candidate fusion feature representation, based on the sample identification object S j The corresponding predicted candidate fusion feature representation generates the first loss function, including:
[0036] Obtain the sample recognition object S j The sample feature representation under M modalities is based on an initial autoencoder model, which is used to identify the sample object S. j Feature encoding is performed on the sample feature representations under M modalities to obtain the sample identification object S. j The sample encoding features under M modalities are used to identify the sample object S. j Feature fusion is performed on the encoded features of samples in M modalities to obtain the sample identification object S. j The corresponding predicted candidate fusion feature representation;
[0037] For sample identification object S j The corresponding predicted candidate fusion feature representation is used for feature decoding to obtain the sample identification object S. j Decoding feature representation in M modalities;
[0038] For sample identification object S jThe sample feature representations under M modalities are subjected to probability distribution transformation to obtain the sample identification object S. j The sample feature representations under M modalities correspond to the first probability distribution;
[0039] For sample identification object S j The probability distribution transformation is performed on the decoded feature representations under M modalities to obtain the sample identification object S. j The second probability distribution corresponding to the decoded feature representation in M modalities;
[0040] Based on the first probability distribution and the second probability distribution, generate the sample identification object S. j The corresponding object loss function;
[0041] The first loss function is generated based on the object loss function corresponding to each of the S samples.
[0042] Specifically, the initial autoencoder model is adjusted based on the first loss function and the second loss function to obtain the target autoencoder model, including:
[0043] Obtain the loss weights, and then apply the loss weights to the second loss function to obtain the weighted second loss function.
[0044] The first loss function and the weighted second loss function are summed to obtain the total loss function.
[0045] Based on the total loss function, the model parameters of the initial autoencoder model are adjusted to obtain the target autoencoder model.
[0046] Specifically, based on the total loss function, the initial autoencoder model is adjusted to obtain the target autoencoder model, including:
[0047] Based on the total loss function, determine the target model adjustment parameters used to adjust the parameters of the sub-encoder and master encoder in the initial autoencoder model;
[0048] The initial model parameters of the sub-encoder and master encoder in the initial autoencoder model are adjusted to the target model adjustment parameters to obtain the parameter-adjusted initial autoencoder model;
[0049] When the initial autoencoder model after parameter adjustment meets the training convergence condition, the initial autoencoder model after parameter adjustment is determined as the target autoencoder model.
[0050] Among them, the S sample identification objects also include sample identification object S j-1 and sample identification object S j+1 ;
[0051] According to mode M iThe corresponding explicit contrastive labels and the S sample identification objects in modality M i The corresponding predicted explicit fusion feature representations are used to generate a second loss function, including:
[0052] For sample identification object S j In mode M i The predictive explicit fusion feature representation, and the sample identification object S j-1 In mode M i The predicted explicit fusion feature representation is multiplied by a vector to obtain the sample identification object S. j With sample identification object S j-1 The first predicted similarity between them;
[0053] For sample identification object S j In mode M i The predictive explicit fusion feature representation, and the sample identification object S j+1 In mode M i The predicted explicit fusion feature representation is multiplied by a vector to obtain the sample identification object S. j With sample identification object S j+1 The second predicted similarity between them;
[0054] The predictive comparison relationship between the first and second predicted similarities is determined as the sample identification object S. j-1 Sample identification object S j and sample identification object S j+1 In mode M i The predictive explicit contrast relationship between object similarity;
[0055] According to mode M i The corresponding explicit contrastive relationship labels and predicted explicit contrastive relationships are used to generate a second loss function.
[0056] Among them, mode M i The corresponding explicit comparison labels include the comparison label between the first object similarity and the second object similarity, where the first object similarity is the sample identification object S. j-1 and sample identification object S j The similarity labels between the samples are used to identify the objects S. j and sample identification object S j+1 Similarity tags between them;
[0057] According to mode M i The corresponding explicit contrastive relation labels and predicted explicit contrastive relations are used to generate a second loss function, including:
[0058] If the explicit contrast relationship label is that the similarity of the first object is less than that of the second object, then the difference between the contrast learning threshold and the second predicted similarity is obtained;
[0059] The difference is summed with the first predicted similarity to obtain the similarity parameter;
[0060] A second loss function is generated based on the similarity parameter.
[0061] One embodiment of this application provides a data processing apparatus, including:
[0062] The first feature encoding module is used to obtain the feature representation of the object to be identified in M modalities, and to encode the feature representations in the M modalities respectively to obtain the encoded features corresponding to the feature representations in the M modalities; M is a positive integer;
[0063] The feature fusion module is used to fuse M encoded features to obtain candidate fused feature representations corresponding to M modalities; the M modalities include modality M i , where i is a positive integer less than or equal to M;
[0064] The first acquisition module is used to acquire mode M. i The corresponding gating vector; the gating vector is used to highlight different recognition objects in modality M. i The explicit contrast relationship between the similarity of objects;
[0065] The second feature encoding module is used to determine the modality M. i The corresponding gating vector is used to encode the candidate fusion feature representation to obtain the object in modality M. i Explicit fusion feature representation under the following conditions; object recognition in modality M i The explicit fusion feature representation is used to identify the object to be identified and the object to be matched in mode M. i The similarity of objects.
[0066] The first feature encoding module includes:
[0067] The input unit is used to input the feature representations of the object in M modalities into the target autoencoder model;
[0068] Feature encoding units are used to obtain the modes M in the target autoencoder model. i The corresponding encoding weight matrix in the sub-encoder is used to apply the encoding weight matrix to mode M. i The feature representation below is used for feature encoding to obtain mode M. i The features below represent the corresponding candidate encoded features;
[0069] The activation processing unit is used to activate the candidate coding features corresponding to the feature representations of the M modalities respectively through the activation function, so as to obtain the coding features corresponding to the feature representations of the M modalities respectively.
[0070] The feature fusion module includes:
[0071] The attention feature encoding unit is used to input M encoded features into the main encoder of the target autoencoder model, and perform attention feature encoding on the M encoded features to obtain the attention encoded features corresponding to the M encoded features respectively.
[0072] The first feature concatenation unit is used to concatenate the M attention-encoded features to obtain the candidate fusion feature representations corresponding to the M modalities.
[0073] Specifically, the attention feature encoding unit is used for:
[0074] The main encoder generates a query vector, a key vector, and a value vector for each of the M encoded features;
[0075] Obtain the encoded feature T i Corresponding query vector and encoding feature T i The dot product of the eigenvectors of the corresponding key vectors is used to determine the encoded feature T. i The corresponding attention matrix; encoded feature T i It belongs to M encoded features, where i is a positive integer less than or equal to M;
[0076] Based on coding feature T i The corresponding attention matrix, for encoding feature T i Attention feature encoding is performed to obtain encoded feature T. i The corresponding attention encoding features.
[0077] Among them, mode M i The corresponding gating vectors include dominant sub-vectors and shielding sub-vectors;
[0078] The second feature encoding module includes:
[0079] The first vector dot product unit is used to perform vector dot product on the candidate fusion feature representation using explicit sub-vectors to obtain the explicit fusion features in the candidate fusion feature representation;
[0080] The second vector dot product unit uses a masked sub-vector to perform a vector dot product on the candidate fusion feature representation to obtain the masked fusion feature in the candidate fusion feature representation.
[0081] The second feature splicing unit splices the explicit fusion features and the masked fusion features to obtain the object in modality M. i The explicit fusion feature representation.
[0082] The aforementioned data processing device also includes:
[0083] The second acquisition module is used to acquire the object to be matched and identified in mode M. i The explicit fusion feature representation;
[0084] The vector dot product module is used to identify the object to be matched in modality M. i The explicit fusion feature representation and recognition of the object in modality M i The explicit fusion feature representations are multiplied by vector dot product to obtain the vector dot product result;
[0085] The determination module is used to determine the similarity between the identified object and the object to be matched by the vector dot product result.
[0086] One embodiment of this application provides a data processing apparatus, including:
[0087] The third acquisition module is used to acquire the initial autoencoder model and S sample recognition objects; the S sample recognition objects include sample recognition object S j S is a positive integer, and j is a positive integer less than S.
[0088] The first generation module is used to identify the sample object S using an initial autoencoder model. j Feature fusion is performed on the encoded features of samples in M modalities to obtain the sample identification object S. j The corresponding predicted candidate fusion feature representation, based on the sample identification object S j The corresponding predicted candidate fusion feature representation generates the first loss function; the M modes include mode M i M is a positive integer, and i is a positive integer less than or equal to M.
[0089] The second generation module is used to obtain mode M. i The corresponding gate vector, based on mode M i The corresponding gating vectors are used to encode the predicted candidate fusion feature representations of the S sample recognition objects, respectively, to generate the S sample recognition objects in mode M. i The corresponding predicted explicit fusion feature representations are shown below, based on modality M. i The corresponding explicit contrastive labels and the S sample identification objects in modality M i The corresponding predicted explicit fusion feature representations are used to generate the second loss function;
[0090] The model parameter adjustment module is used to adjust the model parameters of the initial autoencoder model according to the first loss function and the second loss function to obtain the target autoencoder model. The target autoencoder model is used to predict the explicit fusion feature representations corresponding to the feature representations of the identified object in M modalities.
[0091] The first generation module includes:
[0092] The sample feature fusion module is used to obtain the sample recognition object S. j The sample feature representation under M modalities is based on an initial autoencoder model, which is used to identify the sample object S. j Feature encoding is performed on the sample feature representations under M modalities to obtain the sample identification object S. j The sample encoding features under M modalities are used to identify the sample object S. j Feature fusion is performed on the encoded features of samples in M modalities to obtain the sample identification object S. j The corresponding predicted candidate fusion feature representation;
[0093] Feature decoding unit, used to identify object S from samples j The corresponding predicted candidate fusion feature representation is used for feature decoding to obtain the sample identification object S. j Decoding feature representation in M modalities;
[0094] The first probability distribution transformation unit is used to identify the object S from the sample. j The sample feature representations under M modalities are subjected to probability distribution transformation to obtain the sample identification object S. j The sample feature representations under M modalities correspond to the first probability distribution;
[0095] The second probability distribution transformation unit is used to identify the object S from the sample. j The probability distribution transformation is performed on the decoded feature representations under M modalities to obtain the sample identification object S. j The second probability distribution corresponding to the decoded feature representation in M modalities;
[0096] The first generation unit is used to generate a sample recognition object S based on a first probability distribution and a second probability distribution. j The corresponding object loss function;
[0097] The second generation unit is used to generate the first loss function based on the object loss functions corresponding to the S samples.
[0098] The model parameter adjustment module includes:
[0099] The weighted processing unit is used to obtain the loss weights and apply the loss weights to the second loss function to obtain the weighted second loss function.
[0100] The summation processing unit is used to sum the first loss function and the weighted second loss function to obtain the total loss function;
[0101] The model parameter adjustment unit is used to adjust the model parameters of the initial autoencoder model according to the total loss function to obtain the target autoencoder model.
[0102] Specifically, the model parameter adjustment unit is used for:
[0103] Based on the total loss function, determine the target model adjustment parameters used to adjust the parameters of the sub-encoder and master encoder in the initial autoencoder model;
[0104] The initial model parameters of the sub-encoder and master encoder in the initial autoencoder model are adjusted to the target model adjustment parameters to obtain the parameter-adjusted initial autoencoder model;
[0105] When the initial autoencoder model after parameter adjustment meets the training convergence condition, the initial autoencoder model after parameter adjustment is determined as the target autoencoder model.
[0106] Among them, the S sample identification objects also include sample identification object S j-1 and sample identification object S j+1 ;
[0107] The second generation module includes:
[0108] The third vector dot product unit is used to identify the sample object S. j In mode M i The predictive explicit fusion feature representation, and the sample identification object S j-1 In mode M i The predicted explicit fusion feature representation is multiplied by a vector to obtain the sample identification object S. j With sample identification object S j-1 The first predicted similarity between them;
[0109] The fourth vector dot product unit is used to identify the sample object S. j In mode M i The predictive explicit fusion feature representation, and the sample identification object S j+1 In mode M i The predicted explicit fusion feature representation is multiplied by a vector to obtain the sample identification object S. j With sample identification object S j+1 The second predicted similarity between them;
[0110] The determining unit is used to determine the predictive comparison relationship between the first predicted similarity and the second predicted similarity as the sample identification object S. j-1 Sample identification object S j and sample identification object S j+1 In mode M i The predictive explicit contrast relationship between object similarity;
[0111] The third generation unit is used to generate based on mode M. i The corresponding explicit contrastive relationship labels and predicted explicit contrastive relationships are used to generate a second loss function.
[0112] Among them, mode M i The corresponding explicit comparison labels include the comparison label between the first object similarity and the second object similarity, where the first object similarity is the sample identification object S. j-1 and sample identification object S j The similarity labels between the samples are used to identify the objects S. j and sample identification object S j+1 Similarity tags between them;
[0113] The third generation unit includes:
[0114] Obtain a sub-unit, which is used to obtain the difference between the contrast learning threshold and the second predicted similarity if the explicit contrast relationship label is that the similarity of the first object is less than that of the second object.
[0115] The summation processing subunit is used to sum the difference with the first predicted similarity to obtain the similarity parameter;
[0116] A sub-unit is generated to produce a second loss function based on the similarity parameter.
[0117] One embodiment of this application provides a computer device, including: a processor and a memory;
[0118] The processor is connected to a memory, which stores a computer program. When the computer program is executed by the processor, it causes the computer device to perform the method provided in the embodiments of this application.
[0119] One aspect of this application provides a computer-readable storage medium storing a computer program adapted to be loaded and executed by a processor, so that a computer device having the processor performs the method provided in this application.
[0120] One embodiment of this application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in this application embodiment.
[0121] In this embodiment, feature representations of the object to be identified in M modalities are obtained. These representations are then encoded to obtain coded features corresponding to each of the M modalities. Finally, these coded features are fused to obtain candidate fused feature representations for the M modalities. By fusing the feature representations of the M modalities to obtain candidate fused feature representations for the M modalities, the complementarity between the M modalities can be utilized to remove redundant information from each modality or to supplement missing information, thereby improving the accuracy of feature extraction. (Modality M is then obtained.) i The corresponding gating vector is used to highlight different recognition objects in modality M. i The explicit contrastive relationship between object similarities under modality M. i The corresponding gating vector is used to encode the candidate fusion feature representation to obtain the object in modality M. i The explicit fusion feature representation under the following conditions allows for the identification of objects in modality M. i The explicit fusion feature representation is used to identify the object to be identified and the object to be matched in mode M. i Object similarity is assessed. By using gating vectors to encode candidate fusion feature representations, the resulting explicit fusion feature representation retains the correlation between the identified object and other identified objects. Even if the resulting explicit fusion feature representation contains the explicit characteristics of the original feature representation, it can further improve the accuracy of feature fusion, thereby accurately determining the similarity of the identified object in modality M. i The explicit fusion feature representation is used to ensure the accuracy of business processing results. Attached Figure Description
[0122] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0123] Figure 1 This is a schematic diagram of the architecture of a data processing system provided in an embodiment of this application;
[0124] Figure 2 This is a schematic diagram illustrating an application scenario of data processing provided in an embodiment of this application;
[0125] Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application;
[0126] Figure 4 This is a schematic diagram illustrating an embodiment of the present application for obtaining candidate fusion feature representations corresponding to M modalities;
[0127] Figure 5 This is a schematic diagram illustrating how to obtain an explicit fusion feature representation based on a gating vector, as provided in an embodiment of this application.
[0128] Figure 6 This is a schematic diagram illustrating an embodiment of obtaining an explicit fusion feature representation provided in this application;
[0129] Figure 7 This is a flowchart illustrating a data processing method provided in an embodiment of this application;
[0130] Figure 8 This is a schematic diagram illustrating how to determine object similarity based on explicit fusion feature representation, as provided in an embodiment of this application.
[0131] Figure 9 This is a flowchart illustrating a data processing method provided in an embodiment of this application;
[0132] Figure 10 This is a schematic diagram illustrating how to obtain the decoded feature representation of a sample identification object in M modalities, as provided in an embodiment of this application.
[0133] Figure 11 This is a schematic diagram illustrating a method for training an initial autoencoder model to obtain a target autoencoder model, as provided in an embodiment of this application.
[0134] Figure 12 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;
[0135] Figure 13 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;
[0136] Figure 14 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application;
[0137] Figure 15 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0138] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0139] See Figure 1 , Figure 1 This is a schematic diagram of the structure of a data processing system provided in an embodiment of this application. For example... Figure 1 As shown, the data processing system may include server 10 and a user terminal cluster. The user terminal cluster may include one or more user terminals; the number of user terminals is not limited here. Figure 1 As shown, it can specifically include user terminal 100a, user terminal 100b, user terminal 100c, ..., user terminal 100n. Figure 1 As shown, user terminals 100a, 100b, 100c, ..., 100n can each connect to the server 10 via a network, so that each user terminal can interact with the server 10 through the network connection.
[0140] Each user terminal in this user terminal cluster can include: smartphones, tablets, laptops, desktop computers, wearable devices, smart home devices, head-mounted devices, in-vehicle terminals, and other intelligent terminals capable of data processing. It should be understood that, for example... Figure 1 Each user terminal in the user terminal cluster shown can have the target application (i.e., the application client) installed. When the application client runs on each user terminal, it can interact with the aforementioned... Figure 1 Data interaction occurs between the servers 10 shown.
[0141] Among them, such as Figure 1 As shown, the server 10 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0142] For ease of understanding, the embodiments of this application may be described in detail below. Figure 1 From the plurality of user terminals shown, one user terminal is selected as the target user terminal. The target user terminal may include: a smartphone, tablet computer, laptop computer, desktop computer, smart TV, or other smart terminal with data processing capabilities. For example, for ease of understanding, embodiments of this application may... Figure 1 The user terminal 100a shown serves as the target user terminal. User terminal 100a can acquire feature representations of an identified object in M modalities. The identified object can refer to a user, an item, or other object. A modality is a way of representing information; each source or form of information can be called a modality. The feature representations in the M modalities can refer to images, text, sound, etc., where M is a positive integer, such as 1, 2, 3, etc. User terminal 100a can acquire the feature representations of the identified object in the M modalities through its included application client, or it can acquire the feature representations of the identified object in the M modalities through input from business personnel. After obtaining the feature representations of the identified object in the M modalities, user terminal 100a can send them to server 10. After receiving the feature representations of the identified object in the M modalities sent by user terminal 100a, server 10 can perform feature encoding on the feature representations in the M modalities to obtain the encoded features corresponding to each of the M feature representations. Server 10 can perform feature fusion on M coded features to obtain candidate fused feature representations corresponding to M modalities, and generate the recognition object in modality M based on the candidate fused feature representations. i The explicit fusion feature representation under the following conditions indicates that the identified object is in modality M i The explicit fusion feature representation can be used to determine the relationship between the identified object and the object to be matched in mode M. i The similarity of objects.
[0143] For better understanding, please refer to [link / reference]. Figure 2 , Figure 2 This is a schematic diagram illustrating an application scenario of data processing provided in an embodiment of this application. Wherein, as... Figure 2 The server 20c shown can be the server 10 mentioned above, such as... Figure 2 The target user terminal 20a shown can be the one described above. Figure 1 Any user terminal in the user terminal cluster shown, for example, target user terminal 20a can be the aforementioned user terminal 100a. Figure 2 As shown, the target user terminal 20a can receive the feature representation of the identified object in M modalities input by the business personnel 20b. The identified object can refer to a user or an item. The target user terminal 20a can send the acquired feature representation of the identified object in M modalities to the server 20c. After receiving the feature representation of the identified object in M modalities, the server 20c can input the feature representation of the identified object in M modalities into the target autoencoder model 20d, and the target autoencoder model 20d will output the feature representation of the identified object in modalities M. i The explicit fusion feature representation under the target autoencoder model 20d is used for feature fusion to obtain the explicit fusion feature representation under each modality, and the object to be identified is in modality M.i The explicit fusion feature representation can be used to determine the relationship between the identified object and the object to be matched in mode M. i Object similarity in the video modality. For example, by performing a vector dot product between the explicit fusion feature representation of the identified object in the video modality and the explicit fusion feature representation of the object to be matched in the video modality, the object similarity between the identified object and the object to be matched in the video modality can be obtained. Based on the object similarity in the video modality, video service recommendations can be made for the identified object and the object to be matched.
[0144] Among them, server 20c obtains the object to be identified in modality M i After representing the explicit fusion features, the object can be identified in modality M. i The explicit fusion feature representation is returned to the target user terminal 20a, and the target user terminal can obtain the object to be matched and identified in modality M using the method described above. i The explicit fusion feature representation 20e is obtained under the target user terminal 20a in modality M. i The explicit fusion feature representation under the condition, and the object to be matched and identified in modality M i After representing the explicit fusion features, the object can be identified in modality M. i The explicit fusion feature representation and the object to be matched and identified in modality M i The explicit fusion feature representation is multiplied by a vector to obtain the object similarity between the identified object and the object to be matched. This object similarity can provide a reference for the relevant business processing of the identified object and the object to be matched. For example, if the identified object is user 1 and the object to be matched is user 2, if the object similarity between user 1 and user 2 is high, then in a friend recommendation scenario, user 2 can be recommended to user 1, or user 1 can be recommended to user 2. For example, if the object similarity between user 1 and user 2 is high, then in a book recommendation scenario, books liked by user 1 can be recommended to user 2, or books liked by user 2 can be recommended to user 1.
[0145] Please see Figure 3 , Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application. This data processing method can be executed by a computer device, which can be a server (as described above). Figure 1 Server 10 in the middle), or user terminal (as described above) Figure 1 This application does not limit the scope to any user terminal in a user terminal cluster, or a system consisting of a server and user terminals. Figure 3 As shown, the data processing method may include steps S101-S104.
[0146] S101, obtain the feature representation of the object under M modalities, encode the feature representation under M modalities respectively, and obtain the encoded features corresponding to the feature representation under M modalities respectively.
[0147] Specifically, computer equipment can perform feature fusion on the feature representations of an object in M modalities. By utilizing the complementarity between modalities, it can eliminate redundant information in a modality or supplement missing information in a certain modality. This makes the resulting explicit fused feature representation more accurately represent the features of the object and can also represent the explicit correlation between multiple objects. M is a positive integer, such as M can take values of 1, 2, 3, etc. The computer equipment can acquire the feature representations of an object in M modalities. The object can refer to a user, such as a video user, an instant messaging application user, or other application users. The feature representations in M modalities can refer to the way the features of the object are represented. For example, humans have touch, hearing, vision, and smell. Feature information obtained through touch, hearing, vision, and smell can all become feature representations in a modality. For example, information media include voice, video, and text. Voice can refer to a feature representation in a modality, video can refer to a feature representation in a modality, and text can refer to a feature representation in a modality. Furthermore, features collected by radar, infrared, and acceleration can also refer to feature representations in different modalities. Multimodality refers to the feature representation of the same thing under different characteristics. After a computer device obtains the feature representation of an object in M modalities, it can encode the feature representations in each of the M modalities to obtain the encoded features corresponding to the feature representations in the M modalities respectively.
[0148] Optionally, the specific method by which the computer device encodes the feature representations under M modalities may include: inputting the feature representations of the object to be identified under M modalities into the target autoencoder model, and obtaining the modalities M in the target autoencoder model. i The corresponding encoding weight matrix in the sub-encoder is used to apply the encoding weight matrix to mode M. i The feature representation below is used for feature encoding to obtain mode M. i The candidate coding features corresponding to the feature representations under each of the M modalities are obtained by activating the candidate coding features corresponding to the feature representations under each of the M modalities using activation functions.
[0149] Specifically, a computer device can input the feature representations of the object to be recognized in M modalities into a target autoencoder model. The target autoencoder model can include an autoencoder network structure, also known as an autoencoder, which is an unsupervised neural network model. The autoencoder includes an encoder, which is responsible for compressing the input feature representations into a low-dimensional vector space. After inputting the feature representations of the object to be recognized in M modalities into the target autoencoder model, the computer device can obtain the M modalities in the target autoencoder model. i The corresponding encoding weight matrix in the sub-encoder. The target autoencoder model includes multiple sub-encoders; the feature representation of each input modality is encoded by a separate sub-encoder. The computer device can employ modality M... i The corresponding encoding weight matrix in the sub-encoder, for mode M i The feature representation below is used for feature encoding to obtain mode M. i The features below represent the corresponding candidate coding features. That is, through the coding weight matrix in the sub-encoder, the modality M is... i Feature encoding can be performed on the feature representation below to encode the mode M. i The feature representations are compressed into low-dimensional candidate encoded features. In this way, through the sub-encoders in the target autoencoder, the feature representations of the object under M modalities can be reduced in dimensionality to reduce subsequent storage pressure.
[0150] After obtaining candidate coding features corresponding to the feature representations under M modalities, the computer equipment can activate these candidate coding features using activation functions to obtain the coding features corresponding to the feature representations under the M modalities. In this way, activation functions are used to nonlinearly combine the candidate coding features corresponding to the feature representations under the M modalities, i.e., to add or remove nonlinear factors from the candidate coding features, thereby improving the feature extraction effect.
[0151] S102, perform feature fusion on the M encoded features to obtain candidate fused feature representations corresponding to the M modalities.
[0152] Specifically, after obtaining the encoded features of the object to be identified in M modalities, the computer device can perform feature fusion on the M encoded features to obtain candidate fused feature representations corresponding to the M modalities. Here, a feature representation for one modality corresponds to one encoded feature, and the M modalities include modality M... i Here, i is a positive integer less than or equal to M, such as i can take values of 1, 2, 3, ... For example, two modalities include modality M1 and modality M2. Feature fusion refers to combining vector representations of different types into a single vector representation, such as combining multiple 128-dimensional vector representations into a single 128-dimensional vector representation. For the feature representation under each modality... k represents the dimension of the feature representation; each modality's feature representation has a corresponding dimension. A computer device performs feature fusion on the encoded features from M modalities, obtaining candidate fused feature representations corresponding to the M modalities. Make That is, the fused candidate fused feature representation h can be made fusion The dimension of the fusion feature is less than the sum of the dimensions of the encoded features in each of the M modalities. This allows for dimensionality reduction of the encoded features of the identified object across the M modalities, reducing storage space for feature representations and subsequent computation. Furthermore, feature fusion of the encoded features across the M modalities removes noise signals. Additionally, the candidate fusion feature representation h corresponding to the M modalities... fusion The ability to implicitly contain feature representations of the object to be identified in M modalities can be achieved by fusing candidate fusion feature representations h corresponding to the M modalities. fusion In the decoder of the input autoencoder, the feature representations of the object to be recognized in M modalities are recovered through the decoder. Therefore, the computer device can retrieve the candidate fused feature representations h corresponding to the M modalities. fusion Instead of using the feature representations of the identified object across M modalities, downstream computation is performed. For example, the candidate fusion feature representations h corresponding to the M modalities can be used. fusion As object features for identification, these are input into other network structures to predict relevant features of the identified object. For example, if the identified object is a new user, then the candidate fusion features corresponding to the M modalities are represented as h. fusion As input features for user cold starts, to solve the cold start problem for new users.
[0153] Optionally, the specific method by which the computer device performs feature fusion on the M encoded features may include: inputting the M encoded features into the main encoder of the target autoencoder model, performing attention feature encoding on the M encoded features to obtain attention-encoded features corresponding to the M encoded features respectively, and concatenating the M attention-encoded features to obtain candidate fusion feature representations corresponding to the M modalities.
[0154] Specifically, after obtaining M encoded features, the computer device can input these M encoded features into the main encoder of the target autoencoder model. The main encoder performs attention feature encoding on the M encoded features to obtain attention-encoded features corresponding to each of the M encoded features. In the main encoder, the M attention-encoded features are concatenated to obtain candidate fusion feature representations corresponding to the M modalities.
[0155] Optionally, after obtaining M attention-encoded features, the computer device can add the M attention-encoded features to obtain candidate fusion feature representations corresponding to the M modalities. Alternatively, the computer device can directly concatenate or add the M encoded features to obtain candidate fusion feature representations corresponding to the M modalities.
[0156] Optionally, the specific method by which the computer device obtains the attention-encoded features corresponding to the M encoded features may include: generating a query vector, a key vector, and a value vector corresponding to each of the M encoded features through the main encoder. Obtaining the encoded features T i Corresponding query vector and encoding feature T i The dot product of the eigenvectors of the corresponding key vectors is used to determine the encoded feature T. i The corresponding attention matrix encodes feature T i It belongs to M coding features, where i is a positive integer less than or equal to M. Based on coding feature T... i The corresponding attention matrix, for encoding feature T i Attention feature encoding is performed to obtain encoded feature T. i The corresponding attention encoding features.
[0157] Specifically, the computer device can utilize the first, second, and third matrices in the attention layer of the main encoder. The first matrix can refer to the query weight matrix W. Q The second matrix can refer to the key weight matrix W. K The third matrix can be the value weight matrix W. V The parameters in the first, second, and third matrices can be defined by the user and adjusted later based on the model loss. This query weight matrix W is used. Q Key weight matrix W K Value weight matrix W V Each encoded feature is weighted (i.e., linearly mapped) to generate a query vector, a key vector, and a value vector for each of the M encoded features. Since there is a similarity between each query vector and key vector, this similarity can be used as a weight to weight the value vectors, resulting in the attention-based encoded features. The computer device can then acquire the encoded features T. i The corresponding query vector and encoding feature T i The dot product of the eigenvectors of the corresponding key vectors is used to determine the encoded feature T. i The corresponding attention matrix. The computer device obtains the encoded features T. iAfter obtaining the corresponding attention matrix, the encoded feature T can be used as a basis. i The corresponding attention matrix, for encoding feature T i Perform attention feature encoding, that is, encode feature T i After weighting, the encoded feature T is obtained. i The corresponding attention encoding features.
[0158] The network structures of the sub-encoder and main encoder in the target autoencoder model can be set according to specific circumstances, and the embodiments of this application do not impose any restrictions here.
[0159] like Figure 4 As shown, Figure 4 This is a schematic diagram illustrating an embodiment of the present application for obtaining candidate fusion feature representations corresponding to M modalities, as shown below. Figure 4 As shown, after the computer device acquires the feature representations of the object to be recognized in M modalities, it can input the feature representations in the M modalities into the target autoencoder model. The computer device inputs the feature representation 40a in modality M1, the feature representation 40b in modality M2, ..., the feature representation in modality M... m In the 40k input target autoencoder model, one of the M sub-encoders is used to encode the feature representation of a single modality. For example... Figure 4 As shown, sub-encoder 40d is used to perform feature encoding on feature representation 40a under mode M1, sub-encoder 40e is used to perform feature encoding on feature representation 40b under mode M2, and sub-encoder 40f is used to perform feature encoding on feature representation 40a under mode M1. m The feature representation 40c is used for feature encoding, etc. Through the M sub-encoders in the target autoencoder model, M encoded features corresponding to the feature representations of the M modalities are output, namely encoded feature 40g, encoded feature 40h… encoded feature 40i. After obtaining the M feature codes, these M feature codes can be output to the main encoder 40j in the target autoencoder model 40k, and feature fusion is performed on these M feature codes to obtain the candidate fused feature representation 40l corresponding to the M modalities.
[0160] S103, Obtain mode M i The corresponding gating vector.
[0161] Specifically, computer devices can acquire modal M i The corresponding gating vector is used to highlight different recognition objects in modality M. i The explicit contrastive relationship between object similarities under modality M. i And its characteristics, setting this mode M iEach mode has its own corresponding gating vector, and the gating vectors for any two modes are different. The elements in the gating vector for each mode can take values of 0 or 1. Element 0 is used to mask redundant feature information, and element 1 is used to highlight important feature information.
[0162] S104, according to mode M i The corresponding gating vector is used to encode the candidate fusion feature representation to obtain the object in modality M. i The explicit fusion feature representation.
[0163] Specifically, computer equipment can be based on modal M i The corresponding gating vector is used to encode the candidate fusion feature representation to obtain the object in modality M. i The explicit fusion feature representation under the modality M. The object being identified is in modality M. i The explicit fusion feature representation retains the feature information and display characteristics of the recognized object in M modalities, thus eliminating the need for the decoder in the autoencoder to re-encode the recognized object in modal M. i Decoding the explicit fusion features under the given conditions allows for direct identification of the object in modality M. i Extracting the object in modality M from the explicit fusion feature representation. i The display characteristics are used for downstream calculations. If it is not necessary to identify the object in modality M... i Decoding is performed using the explicit fusion feature representation under the given conditions, which can be directly based on the object being identified in mode M. i The explicit fusion feature representation under the condition determines the identification object and the object to be matched in mode M. i Object similarity under different modalities. i The corresponding gating vector has the same dimension as the candidate fusion feature representation, and the modality M i The corresponding gating vector is a vector in one dimension, used to encode features in one dimension of the candidate fused feature representation.
[0164] Optional, mode M i The corresponding gating vectors include dominant sub-vectors and masking sub-vectors. The computer device obtains the object to be identified in mode M. i The specific methods for representing explicit fusion features can include: using explicit sub-vectors to perform vector dot products on the candidate fusion feature representations to obtain explicit fusion features in the candidate fusion feature representations; using masked sub-vectors to perform vector dot products on the candidate fusion feature representations to obtain masked fusion features in the candidate fusion feature representations; and concatenating the explicit and masked fusion features to obtain the object in modality M. i The explicit fusion feature representation.
[0165] Specifically, the explicit sub-vector can refer to vector 1, and the masked sub-vector can refer to vector 0. The computer device can use the explicit sub-vector to perform a vector dot product on the candidate fusion feature representation to obtain the explicit fusion feature in the candidate fusion representation. The computer device can also use the masked sub-vector to perform a vector dot product on the candidate fusion feature representation to obtain the masked fusion feature in the candidate fusion feature representation. The explicit and masked fusion features are then concatenated, or added together, to obtain the object in modality M. i The explicit fusion feature representation.
[0166] like Figure 5 As shown, Figure 5 This is a schematic diagram illustrating how to obtain an explicit fusion feature representation based on a gating vector, as provided in an embodiment of this application. Figure 5 As shown, mode M i The corresponding gate vector is 50a, i.e., [0,1,1], and the candidate fusion feature representations for the M modalities are 50b, i.e., [0.1,0.3,0.5]. For example... Figure 5 As shown, mode M i The corresponding gate vector 50a includes two dominant sub-vectors and one shielding sub-vector, using mode M. i Each subvector in the corresponding gating vector 50a performs a vector dot product on the candidate fused feature representation, i.e., modality M i The corresponding gate vector g i A vector in the 50b candidate fusion feature representation is multiplied by a vector dot product with an element of the 50b candidate fusion feature representation. For example, modality M i The first sub-vector 0 (i.e., the masking sub-vector) in the corresponding gate vector 50a is multiplied by the element 0.1 in the candidate fusion feature representation 50b to obtain the masked fusion feature 0 in the candidate fusion representation 50b; Modality M i The second sub-vector 1 (i.e., the explicit sub-vector) in the corresponding gating vector 50a is multiplied by the element 0.3 in the candidate fusion feature representation 50b to obtain the explicit fusion feature 0.3 in the candidate fusion representation 50b; Modality M i The third sub-vector 1 (i.e., the dominant sub-vector) in the corresponding gating vector 50a is multiplied by the element 0.5 in the candidate fusion feature representation 50b to obtain the explicit fusion feature 0.5 in the candidate fusion representation 50b. The explicit fusion feature and the masked fusion feature are then concatenated to obtain the object in modality M. i The dominant fusion feature is represented as 50c, i.e., [0, 0.3, 0.5].
[0167] like Figure 6 As shown, Figure 6 This is a schematic diagram illustrating an embodiment of obtaining an explicit fusion feature representation, as provided in this application. Figure 6 As shown, the computer device can acquire the first-order feature representation 61a of the identified object in modality 1, the second-order feature representation 61b in modality 2, and the social application group feature representation 61c in modality 3. The first-order feature representation of the identified object in modality 1 can be used to characterize the social interests of the identified object, the second-order feature representation of the identified object in modality 2 can be used to characterize the social friend information of the identified object, and the social application group feature representation of the identified object in modality 3 can be used to characterize the social application group information possessed by the identified object. The first-order feature representation 61a of the identified object in modality 1, the second-order feature representation 61b of the identified object in modality 2, and the social application group feature representation 61c of the identified object in modality 3 can be input into the sub-encoder in the target autoencoder model 65. That is, the first-order feature representation 61a of the identified object in modality 1 is input into sub-encoder 62a, the second-order feature representation 61b of the identified object in modality 2 is input into sub-encoder 62b, and the social application group feature representation 61c of the identified object in modality 3 is input into sub-encoder 62c. The first-order feature representation, second-order feature representation, and social application group feature representation are feature encoded by each sub-encoder in the target autoencoder model 65, resulting in the feature codes corresponding to the first-order feature representation, second-order feature representation, and social application group feature representation, respectively. The specific process of feature encoding by the sub-encoders can be found in [link to documentation]. Figure 3 The specific details of step S101 will not be elaborated here.
[0168] In this process, after the computer device obtains the feature codes corresponding to the first-order feature representation, the second-order feature representation, and the social application group feature representation, it can input these feature codes into the main encoder 63 in the target autoencoder model 65. The main encoder 63 then performs feature fusion on the feature codes corresponding to the first-order feature representation, the second-order feature representation, and the social application group feature representation to obtain candidate fused feature representations for the three modalities. The specific process of feature fusion by the main encoder can be found in [link to documentation]. Figure 3 The specific details of step S102 will not be elaborated here. Figure 6 As shown, the computer device acquires the gating vector 64a corresponding to mode 1, and uses this gating vector 64a to encode the candidate fusion features, obtaining the explicit fusion feature representation 66a of the object to be identified in mode 1. The computer device can acquire the gating vector 64b corresponding to mode 2, and encode the candidate fusion feature representation, obtaining the explicit fusion feature representation 66b of the object to be identified in mode 2. The computer device can acquire the gating vector 64c corresponding to mode 3, and encode the candidate fusion feature representation, obtaining the explicit fusion feature representation 66c of the object to be identified in mode 3.
[0169] The computer device can employ the aforementioned method for obtaining explicit fusion feature representations of the identified object to acquire explicit fusion feature representations corresponding to the object to be matched in modality 1, modality 2, and modality 3, respectively. By performing a vector dot product between the explicit fusion feature representation of the identified object in modality 1 and the explicit fusion feature representation of the object to be matched in modality 1, a first-order similarity is obtained between the identified object and the object to be matched. This first-order similarity can be used to determine whether the identified object and the object to be matched are friends. Similarly, a second-order similarity can be obtained between the identified object and the object to be matched in modality 2, which is used to determine the probability that the identified object and the object to be matched are friends. Likewise, a social application group similarity can be obtained between the identified object and the object to be matched in modality 3, which can be used to determine the probability that the identified object and the object to be matched have joined the same social application group.
[0170] In this embodiment, feature representations of the object to be identified in M modalities are obtained. These representations are then encoded to obtain coded features corresponding to each of the M modalities. Finally, these coded features are fused to obtain candidate fused feature representations for the M modalities. By fusing the feature representations of the M modalities to obtain candidate fused feature representations for the M modalities, the complementarity between the M modalities can be utilized to remove redundant information from each modality or to supplement missing information, thereby improving the accuracy of feature extraction. (Modality M is then obtained.) i The corresponding gating vector is used to highlight different recognition objects in modality M. i The explicit contrastive relationship between object similarities under modality M. i The corresponding gating vector is used to encode the candidate fusion feature representation to obtain the object in modality M. i The explicit fusion feature representation under the following conditions allows for the identification of objects in modality M. i The explicit fusion feature representation is used to identify the object to be identified and the object to be matched in mode M. i Object similarity is assessed. By using gating vectors to encode candidate fusion feature representations, the resulting explicit fusion feature representation retains the correlation between the identified object and other identified objects. Even if the resulting explicit fusion feature representation contains the explicit characteristics of the original feature representation, it can further improve the accuracy of feature fusion, thereby accurately determining the similarity of the identified object in modality M. iThe explicit fusion feature representation is used to ensure the accuracy of business processing results. In addition, by performing feature fusion on the feature representations of the identified object in M modalities, the high-dimensional feature representations in M modalities can be merged into a lower-dimensional vector space. That is, the high-dimensional feature representations in M modalities are merged into a low-dimensional explicit fusion feature representation, which can reduce feature storage space and remove noise signals in the feature representations in M modalities.
[0171] Please see Figure 7 , Figure 7 This is a flowchart illustrating a data processing method provided in an embodiment of this application. This data processing method can be executed by a computer device, which can be a server (as described above). Figure 1 Server 10 in the middle), or user terminal (as described above) Figure 1 This application does not limit the scope to any user terminal in a user terminal cluster, or a system consisting of a server and user terminals. Figure 7 As shown, the data processing method may include steps S201-S207.
[0172] S201, obtain the feature representation of the object under M modalities, encode the feature representation under M modalities respectively, and obtain the encoded features corresponding to the feature representation under M modalities respectively.
[0173] S202, perform feature fusion on the M encoded features to obtain candidate fused feature representations corresponding to the M modalities.
[0174] S203, Obtain Mode M i The corresponding gating vector.
[0175] S204, according to mode M i The corresponding gating vector is used to encode the candidate fusion feature representation to obtain the object in modality M. i The explicit fusion feature representation.
[0176] Specifically, the details of steps S201-S204 in the embodiments of this application can be found in [reference needed]. Figure 3 The contents of steps S101-S104 in the embodiments will not be repeated here.
[0177] S205, Obtain the object to be matched and identified in modality M i The explicit fusion feature representation.
[0178] Specifically, computer devices can use the above-mentioned methods to acquire and identify objects in modal M. i The method of explicit fusion feature representation under the condition is used to obtain the target object in modality M. iThe explicit fusion feature representation is used to obtain the target object in modality M. i The representation of the dominant fusion feature can be found in [reference]. Figure 2 The content described herein will not be repeated here in the embodiments of this application.
[0179] S206, the object to be matched and identified is in modality M i The explicit fusion feature representation and recognition of the object in modality M i The explicit fusion feature representations are multiplied by vector dot product to obtain the vector dot product result.
[0180] Specifically, the computer device obtains the object to be identified in mode M i The explicit fusion feature representation under the condition, and the object to be matched and identified in modality M i After explicit fusion feature representation, the object to be matched and identified can be in modality M. i The explicit fusion feature representation and recognition of the object in modality M i The explicit fusion feature representation under modality M1 is then subjected to vector dot product to obtain the vector dot product result. For example, the explicit fusion feature representation of the identified object under modality M1 is: That is, [0.3, 0.2, 0.7], the explicit fusion features of the object to be matched and identified in modality M1 are represented as follows: That is, [0.1, 0.5, 0.3], then for and Perform vector dot product, i.e. The vector dot product is 3.4, which is 0.3*0.1+0.2*0.5+0.7*0.3=3.4.
[0181] S207, determine the vector dot product result as the object similarity between the identified object and the object to be matched.
[0182] Specifically, after the computer device obtains the vector dot product result of the identified object and the support of the object to be matched, it can determine the object similarity between the identified object and the object to be matched.
[0183] For example, after obtaining the explicit fusion feature representations of User 1 and User 2 in Modality 1 using the above method, the computer device can perform a vector dot product on these representations to obtain the first-order similarity between User 1 and User 2 in Modality 1. This first-order similarity is used to determine whether the two users are friends. If the two users are friends, their first-order similarity is greater than that between any two non-friends. For instance, it can be determined whether the first-order similarity between User 1 and User 2 in Modality 1 is greater than a target first-order similarity threshold. If it is, User 1 and User 2 are friends; otherwise, they are not friends.
[0184] For example, a computer device can obtain the explicit fusion feature representations of User 1 and User 2 in the social application group modality using the method described above. Then, it performs a vector dot product on these two representations to obtain the similarity between User 1 and User 2 in the social application group modality. This similarity can be used to determine the probability that two users have joined the same social application group. A higher similarity between User 1 and User 2 in the social application group modality indicates a higher probability that they have joined the same group; conversely, a lower similarity indicates a lower probability.
[0185] like Figure 8 As shown, Figure 8 This is a schematic diagram illustrating how to determine object similarity based on explicit fusion feature representation, as provided in an embodiment of this application. Figure 8 The feature representation shown in Modality 1 can be used to reflect the probability that objects are friends. That is, if two objects are friends, their object similarity in Modality 1 is relatively high; if two objects are not friends, their object similarity in Modality 1 is relatively low. Figure 8 The feature representation shown in Modality 2 can be used to reflect the probability of objects becoming friends. The more mutual friends two objects have, the greater their object similarity in Modality 2, and the higher the probability that they will become friends; conversely, the fewer mutual friends two objects have, the lower their object similarity in Modality 2, and the lower the probability that they will become friends. For example... Figure 8The feature representation in Modality 3 shown can be used to reflect the number of times objects share the same social application groups. The more social application groups two objects share, the greater the similarity of their social application groups in Modality 3, indicating a higher probability that the two objects have joined the same social application groups. Conversely, the fewer social application groups two objects share, the lower the similarity of their social application groups in Modality 3, indicating a lower probability that the two objects have joined the same social application groups.
[0186] After obtaining the feature representations 81a corresponding to the object to be identified in modal 1, modal 2, and modal 3, the computer device can input these feature representations 81a into the target autoencoder model 81b. The target autoencoder model 81b then inputs the explicit fusion feature representations 81c, 81d, and 81e of the object to be identified in modal 1. Similarly, after obtaining the feature representations 82a corresponding to the object to be matched in modal 1, modal 2, and modal 3, the computer device can input these feature representations 82a into the target autoencoder model 82b. The target autoencoder model then outputs the explicit fusion feature representations 82c, 82d, and 82e of the object to be matched in modal 1, modal 2, and modal 3.
[0187] In this process, after obtaining the explicit fusion feature representation 81c of the identified object in modality 1 and the explicit fusion feature representation 82c of the object to be matched in modality 1, the computer device can perform a vector dot product on the explicit fusion feature representation 81c of the identified object in modality 1 and the explicit fusion feature representation 82c of the object to be matched in modality 1 to obtain the first-order similarity 83a between the identified object and the object to be matched in modality 1. This first-order similarity 83a can be used to reflect the probability that the identified object and the object to be matched are friends. If the first-order similarity 83a is greater than the target first-order similarity, it can be determined that the identified object and the object to be matched are friends; if the first-order similarity 83a is less than or equal to the target first-order similarity, it can be determined that the identified object and the object to be matched are not friends.
[0188] In this process, after obtaining the explicit fusion feature representation 81d of the identified object in modality 2 and the explicit fusion feature representation 82d of the object to be matched in modality 2, the computer device can perform a vector dot product on the explicit fusion feature representation 81d of the identified object in modality 2 and the explicit fusion feature representation 82d of the object to be matched in modality 2 to obtain the second-order similarity 83b between the identified object and the object to be matched in modality 2. This second-order similarity 83b can be used to reflect the probability that the identified object and the object to be matched can become friends. If the second-order similarity 83b is greater than the target second-order similarity, it can be determined that the identified object and the object to be matched can become friends, and the object to be matched can be recommended to the identified object, or the identified object can be recommended to the object to be matched. If the second-order similarity 83b is less than or equal to the target second-order similarity, it can be determined that the identified object and the object to be matched cannot become friends.
[0189] In this process, after obtaining the explicit fusion feature representation 81e of the identified object in modality 3 and the explicit fusion feature representation 82e of the object to be matched in modality 3, the computer device can perform a vector dot product on the explicit fusion feature representation 81e of the identified object in modality 2 and the explicit fusion feature representation 82e of the object to be matched in modality 3 to obtain the social application group similarity 83c between the identified object and the object to be matched in modality 3. This social application group similarity 83c can be used to reflect the probability that the identified object and the object to be matched can join the same social application group. If the social application group similarity 83c between the identified object and the object to be matched is greater than the target social application group similarity, it can be determined that the identified object and the object to be matched can join the same social application group. If the social application group similarity 83c between the identified object and the object to be matched is less than or equal to the target social application group similarity, it can be determined that the identified object and the object to be matched cannot join the same social application group.
[0190] In this embodiment, feature representations of the object to be identified in M modalities are obtained. These representations are then encoded to obtain coded features corresponding to each of the M modalities. Finally, these coded features are fused to obtain candidate fused feature representations for the M modalities. By fusing the feature representations of the M modalities to obtain candidate fused feature representations for the M modalities, the complementarity between the M modalities can be utilized to remove redundant information from each modality or to supplement missing information, thereby improving the accuracy of feature extraction. (Modality M is then obtained.) i The corresponding gating vector is used to highlight different recognition objects in modality M. i The explicit contrastive relationship between object similarities under modality M.i The corresponding gating vector is used to encode the candidate fusion feature representation to obtain the object in modality M. i The explicit fusion feature representation under the following conditions allows for the identification of objects in modality M. i The explicit fusion feature representation is used to identify the object to be identified and the object to be matched in mode M. i The similarity between objects is calculated. The similarity between the objects to be matched and identified in modality M is obtained. i The explicit fusion feature representation under the condition is based on the object being identified in modality M. i The explicit fusion feature representation and the object to be matched and identified in modality M i The explicit fusion feature representation determines the object similarity between the identified object and the object to be matched. This object similarity can be used for business recommendation or object matching. Thus, by using gating vectors to encode the candidate fusion feature representations, the resulting explicit fusion feature representation can be used to determine the object similarity between the identified object and other identified objects in various modalities. This preserves the correlation between the identified object and other identified objects in the feature representation before fusion, and even though the resulting explicit fusion feature representation contains the explicit characteristics of the feature representation before fusion, it can further improve the accuracy of feature fusion, thereby accurately determining the similarity of the identified object in modality M. i The explicit fusion feature representation is used to ensure the accuracy of business processing results. In addition, by performing feature fusion on the feature representations of the identified object in M modalities, the high-dimensional feature representations in M modalities can be merged into a lower-dimensional vector space. That is, the high-dimensional feature representations in M modalities are merged into a low-dimensional explicit fusion feature representation, which can reduce feature storage space and remove noise signals in the feature representations in M modalities.
[0191] Please see Figure 9 , Figure 9 This is a flowchart illustrating a data processing method provided in an embodiment of this application. This data processing method can be executed by a computer device, which can be a server (as described above). Figure 1 Server 10 in the middle), or user terminal (as described above) Figure 1 This application does not limit the scope to any user terminal in a user terminal cluster, or a system consisting of a server and user terminals. Figure 9 As shown, the data processing method may include steps S301-S304.
[0192] S301, Obtain the initial autoencoder model and S sample recognition objects; the S sample recognition objects include sample recognition object S j .
[0193] Specifically, the computer equipment can train an initial autoencoder model to obtain a target autoencoder model. The target autoencoder model is then used to fuse the feature representations of the object in M modalities to obtain the object's representation in modality M. i The explicit fusion feature representation is given below, where i is a positive integer less than or equal to M, such as i can take values of 1, 2, 3, ... The computer device can acquire an initial autoencoder model and S sample recognition objects, which include sample recognition object S. j S is a positive integer, such as S can take the value 1, 2, 3..., and j is a positive integer less than S, such as j can take the value 1, 2, 3...
[0194] S302, using an initial autoencoder model to identify sample object S j Feature fusion is performed on the encoded features of samples in M modalities to obtain the sample identification object S. j The corresponding predicted candidate fusion feature representation, based on the sample identification object S j The corresponding predicted candidate fusion feature representation generates the first loss function.
[0195] Specifically, the computer device can use an initial autoencoder model to identify the sample object S. j Feature encoding is performed on the sample feature representations under M modalities to obtain the sample encoded features corresponding to the sample feature representations under each of the M modalities. That is, one sample feature representation under one modality corresponds to one sample encoded feature. The M modalities include modality M... i M is a positive integer, and i is a positive integer less than or equal to M, such as two modes including mode M1 and mode M2. Feature fusion is performed on the encoded features of the M samples to obtain the sample identification object S. j The corresponding predicted candidate fusion feature representations are used to obtain the predicted candidate fusion feature representations for each of the S sample recognition objects. For details on feature encoding and feature fusion performed by the computer equipment, please refer to [link to relevant documentation]. Figure 3 The content described in steps S101-S102 is not repeated here in this embodiment. The computer device obtains the sample identification object S. j After representing the corresponding predicted candidate fusion features, the object S can be identified from the sample. j The corresponding predicted candidate fusion feature representation is decoded to obtain the decoded predicted candidate fusion feature representation. Based on the decoded predicted candidate fusion feature representation and the sample identification object S, the feature representation is then used to identify the target object S. j Based on the feature representation of samples in M modalities, the sample identification object S is determined. j The corresponding object loss function. That is, identifying object S based on the fused samples. jThe difference between the sample feature representations in M modalities and the decoded predicted candidate fusion feature representations is used to determine the sample identification object S. j The corresponding object loss function. After obtaining the object loss function corresponding to each sample object, the computer device can sum the object loss functions corresponding to each sample object to generate a first loss function. In this way, by generating the first loss function based on the object loss function corresponding to each sample object, the training accuracy of the model can be improved when adjusting the model parameters of the initial autoencoder model based on the first loss function.
[0196] Optionally, the specific method by which the computer device generates the first loss function may include: acquiring the sample identification object S. j The sample feature representation under M modalities is based on an initial autoencoder model, which is used to identify the sample object S. j Feature encoding is performed on the sample feature representations under M modalities to obtain the sample identification object S. j The sample encoding features under M modalities are used to identify the sample object S. j Feature fusion is performed on the encoded features of samples in M modalities to obtain the sample identification object S. j The corresponding predicted candidate fusion feature representation. For the sample identification object S... j The corresponding predicted candidate fusion feature representation is used for feature decoding to obtain the sample identification object S. j Decoding feature representation in M modalities. For sample identification object S... j The sample feature representations under M modalities are subjected to probability distribution transformation to obtain the sample identification object S. j The sample feature representations under M modalities correspond to the first probability distribution. For the sample identification object S... j The probability distribution transformation is performed on the decoded feature representations under M modalities to obtain the sample identification object S. j The second probability distribution corresponding to the decoded feature representations under M modalities is used to generate the sample recognition object S based on the first probability distribution and the second probability distribution. j The corresponding object loss function is used to generate the first loss function based on the object loss functions corresponding to the S samples.
[0197] Specifically, computer equipment can identify sample objects S j The corresponding predicted candidate fusion feature representation is used for feature decoding to obtain the sample identification object S. j Decoding feature representation in M modalities, i.e., for sample identification object S j The corresponding predicted candidate fusion feature representation is used for feature decoding, that is, the sample identification object S jThe corresponding predicted candidate fusion feature representation reconstructs the feature representation of the original input, that is, restores the original signal. Computer devices can use the softmax function to identify the object S from the sample. j The sample feature representations under M modalities are subjected to probability distribution transformation to obtain the sample identification object S. j The sample feature representations under M modalities correspond to the first probability distribution. The softmax function, also known as the normalization exponential function, is used to divide the sample into three modalities and the corresponding probability distributions. j The feature representations of samples in M modalities are mapped to the interval (0,1) to obtain the first probability distribution. Computer devices can also use the softmax function to identify the object S from the samples. j The decoded feature representations under M modalities are mapped to the (0,1) interval to obtain the second probability distribution. The computer device can use the KL divergence loss function to measure the difference between the first and second probability distributions, generating a sample object S for identification. j The corresponding object loss function is the first loss function. After the computer device obtains the object loss function corresponding to each sample object, it can add the object loss functions corresponding to each sample object to generate the first loss function. In this way, by generating the first loss function based on the object loss function corresponding to each sample object, the training accuracy of the model can be improved when using the first loss function to input model parameters into the initial autoencoder model.
[0198] Optionally, the expression for the object loss function corresponding to the sample identification object can be the following formula (1):
[0199]
[0200] Among them, in formula (1) Refers to the sample identification object S j In M i Decoding feature representation under each modality Refers to the sample identification object S j In M i The feature representation under each modality is defined by softmax() for probability distribution transformation, KL divergence loss function to measure the difference between the first and second distribution probabilities, and I refers to the M modalities. The object loss function corresponding to the sample identification object can also be the MSE loss function or other loss functions; this embodiment does not impose such limitations.
[0201] like Figure 10 As shown, Figure 10 This is a schematic diagram illustrating how to obtain the decoded feature representation of a sample recognition object in M modalities, as provided in an embodiment of this application. Figure 10As shown, the computer device can represent the sample features of the object in mode M1 as 101a, the sample features in mode M2 as 101b, and so on, in mode M... m The sample feature representation 101c is input into sub-encoders 102a, 102b, and 102c in the initial autoencoder model to obtain sample coding features 103a, 103b, and 103c. The computer device can then input these sample coding features 103a, 103b, and 103c into the main encoder 104a in the initial autoencoder model. The main encoder 104a outputs the predicted candidate fusion feature representation 105a for sample recognition across M modalities. For details on the implementation, please refer to [link to documentation / documentation]. Figure 3 The content of step S102 will not be described again in this embodiment. The computer device can input the predicted candidate fusion feature representation 105a into sub-decoders 106a, 106b, and 106c respectively for feature decoding, obtaining the decoded feature representation 107a of the sample identification object in mode M1, the decoded feature representation 107b in mode M2, and so on in mode M... m The decoding feature is represented as 107c.
[0202] S303, Obtain Mode M i The corresponding gate vector, based on mode M i The corresponding gating vectors are used to encode the predicted candidate fusion feature representations of the S sample recognition objects, respectively, to generate the S sample recognition objects in mode M. i The corresponding predicted explicit fusion feature representations are shown below, based on modality M. i The corresponding explicit contrastive labels and the S sample identification objects in modality M i The corresponding predicted explicit fusion feature representations are used to generate the second loss function.
[0203] Specifically, the computer device acquires mode M out of M modes. i The corresponding gate vector, based on mode M i The corresponding gating vectors are used to encode the predicted explicit candidate fusion feature representations corresponding to the S sample recognition objects, generating the S sample recognition objects in modality M. i The corresponding predicted explicit fusion feature representations are shown below. That is, modality M is used. i The corresponding gating vectors are used to perform vector dot products on the predicted candidate fusion feature representations corresponding to the S sample recognition objects, with each sample recognition object corresponding to one predicted explicit fusion feature representation. This is based on the S sample recognition objects in modality M. i The corresponding predictive explicit fusion feature representations are used to determine the S sample objects in modality M.i The predicted explicit contrastive relationship between object similarities is determined based on this predicted explicit contrastive relationship and modality M. i The corresponding explicit contrast labels are used to generate a second loss function.
[0204] Optionally, the S sample identification objects also include sample identification object S. j-1 and sample identification object S j+1 The computer equipment determines the S sample recognition objects in mode M. i The specific methods for predicting explicit contrastive relationships between object similarities can include: identifying object S in the sample. j In mode M i The predictive explicit fusion feature representation, and the sample identification object S j-1 In mode M i The predicted explicit fusion feature representation is multiplied by a vector to obtain the sample identification object S. j With sample identification object S j-1 The first predicted similarity between the samples. For the sample object S... j In mode M i The predictive explicit fusion feature representation, and the sample identification object S j+1 In mode M i The predicted explicit fusion feature representation is multiplied by a vector to obtain the sample identification object S. j With sample identification object S j+1 The second predicted similarity between the first and second predicted similarities is used to determine the predictive comparison relationship between the first and second predicted similarities, which is then used to identify the sample object S. j-1 Sample identification object S j and sample identification object S j+1 In mode M i The predictive explicit contrast relationship between object similarity.
[0205] Specifically, computer equipment can identify sample objects S j In mode M i The predictive explicit fusion feature representation, and the sample identification object S j-1 In mode M i The predicted explicit fusion feature representation is multiplied by a vector to obtain the sample identification object S. j With sample identification object S j-1 The first predicted similarity between them. For example, the predicted explicit fusion features of sample object S2 in modality M1 are represented as follows: The predicted explicit fusion feature of sample object S1 in modality M1 is represented as follows: The first predicted similarity is Computer equipment identifies sample object S jIn mode M i The predictive explicit fusion feature representation, and the sample identification object S j+1 In mode M i The predicted explicit fusion feature representation is multiplied by a vector to obtain the sample identification object S. j With sample identification object S j+1 The second predicted similarity between them. For example, the predicted explicit fusion feature of sample identification object S2 in modality M1 is represented as: The predicted explicit fusion feature of sample object S3 in modality M1 is represented as follows: The second predicted similarity is After obtaining the first and second predicted similarities, the computer device can determine the predictive comparison relationship between the first and second predicted similarities as the sample identification object S. j-1 Sample identification object S j and sample identification object S j+1 In mode M i The predictive comparison relationship between object similarities is defined. For example, if the first predicted similarity is less than the second predicted similarity, then the sample identifies object S. j-1 Sample identification object S j and sample identification object S j+1 In mode M i The predictive explicit contrastive relationship between the similarity of objects can be:
[0206] Specifically, computer devices can employ contrastive learning algorithms to contrast modalities M. i The corresponding explicit contrast relationship labels, and the sample identification object S j-1 Sample identification object S j and sample identification object S j+1 In mode M i The difference between the predicted explicit contrastive relationships of object similarity is used to generate a second loss function corresponding to the initial autoencoder model. Contrastive learning enables objects of the same class to have similar feature representations, while objects of different classes have different feature representations.
[0207] Optional, mode M i The corresponding explicit comparison labels include the comparison label between the first object similarity and the second object similarity, where the first object similarity is the sample identification object S. j-1 and sample identification object S j The similarity labels between the samples are used to identify the objects S. j and sample identification object S j+1 Similarity tags between them. Computer devices based on modality M iThe specific methods for generating the second loss function based on the corresponding explicit contrast relationship labels and the predicted explicit contrast relationships can include: if the explicit contrast relationship label indicates that the similarity of the first object is less than that of the second object, then the difference between the contrast learning threshold and the second predicted similarity is obtained. This difference is then summed with the first predicted similarity to obtain a similarity parameter, and the second loss function is generated based on this similarity parameter.
[0208] Specifically, if the explicit comparison relationship label is that the similarity of the first object is less than that of the second object, that is, the sample identification object S... j-1 and sample identification object S j In mode M i The similarity of the first object is less than that of the sample object S. j and sample identification object S j+1 In mode M i The second object similarity is then used to obtain the difference between the contrast learning threshold and the second predicted similarity. The contrast learning threshold can be a threshold greater than 0. This allows us to obtain the sample identification object S. j and sample identification object S j+1 In mode M i The difference between the second predicted similarity and the contrastive learning threshold is calculated. This difference is then compared with the sample identification object S. j-1 and sample identification object S j In mode M i The similarity parameter is obtained by summing the first predicted similarity scores. Alternatively, the computer device can obtain the difference between the first and second predicted similarities, sum this difference with the contrastive learning threshold, and obtain the similarity parameter. This similarity parameter can be used to determine the magnitude of the first and second predicted similarities. After obtaining this similarity parameter, the computer device can compare its magnitude with zero to generate a second loss function.
[0209] Optionally, when the explicit comparison label is that the similarity of the first object is less than the similarity of the second object (i.e., The computer device generates the second loss function as shown in formula (2):
[0210]
[0211] In formula (2), M is the contrastive learning threshold, v i Refers to the sample identification object S j , Refers to the sample identification object S j+1 , Refers to the sample identification object S j-1 , Refers to the sample identification object S j With sample identification object Sj+1 The second predicted similarity between them Refers to the sample identification object S j With sample identification object S j-1 The first predicted similarity between them Refers to the sample identification object S j In mode M i The predicted explicit fusion feature representation is as follows: Refers to the sample identification object S j+1 In mode M i The predicted explicit fusion feature representation is as follows: Refers to the sample identification object S j-1 In mode M i The predicted explicit fusion feature representation is used. Wherein, when the explicit contrastive relation label changes, the functional form of the second loss function can be other forms of contrastive learning loss function, which is not limited in this embodiment.
[0212] S304. Based on the first loss function and the second loss function, the model parameters of the initial autoencoder model are adjusted to obtain the target autoencoder model.
[0213] Specifically, the computer equipment can adjust the parameters of the initial autoencoder model based on the first loss function and the second loss function to obtain the target autoencoder model.
[0214] Optionally, the specific method by which the computer device adjusts the parameters of the initial autoencoder model to obtain the target autoencoder model may include: obtaining loss weights; using these loss weights to weight the second loss function to obtain a weighted second loss function; summing the first loss function and the weighted second loss function to obtain the total loss function; and using the gradient descent algorithm to adjust the model parameters of the initial autoencoder model based on the total loss function to obtain the target autoencoder model.
[0215] The formula for obtaining the total loss function by the computer can be shown in formula (3) below:
[0216] L total =L ae +αL c (3)
[0217] Wherein, L in formula (3) ae This refers to the first loss function, i.e., L. c This refers to the second loss function, where α is the loss weight, and α > 0.
[0218] Optionally, the specific method by which the computer device adjusts the parameters of the initial autoencoder model to obtain the target autoencoder model may include: determining the target model adjustment parameters for adjusting the parameters of the sub-encoder and master encoder in the initial autoencoder model based on the total loss function; adjusting the initial model parameters of the sub-encoder and master encoder in the initial autoencoder model to the target model adjustment parameters to obtain the parameter-adjusted initial autoencoder model; and determining the parameter-adjusted initial autoencoder model as the target autoencoder model when the parameter-adjusted initial autoencoder model satisfies the training convergence condition. The training convergence condition may refer to the loss value of the initial autoencoder model calculated by the total loss function being less than or equal to a target loss threshold, or the training convergence condition may refer to the number of iterations of the initial autoencoder model reaching a target number.
[0219] like Figure 11 As shown, Figure 11 This is a schematic diagram illustrating a method for training an initial autoencoder model to obtain a target autoencoder model, as provided in an embodiment of this application. Figure 11 As shown, the computer device can input the feature representations 111a of sample recognition object 1 in M modalities, the feature representations 111b of sample recognition object 2 in M modalities, and the feature representations 111c of sample recognition object 3 in M modalities into the sub-encoder 112a and main encoder 113a of the initial autoencoder model to obtain the predicted candidate fusion feature representation corresponding to each sample recognition object. For details on the implementation, please refer to [link / reference needed]. Figure 3 The contents described in steps S102 and S102 will not be repeated here in this embodiment. After the computer device obtains the predicted candidate fusion feature representation corresponding to each sample identification object, it can perform feature decoding on the predicted candidate fusion feature representation corresponding to each sample identification object through the sub-decoder 114a to obtain the decoded feature representation 115a of sample identification object 1 in M modalities, the decoded feature representation 115b of sample identification object 2 in M modalities, and the decoded feature representation 115c of sample identification object 3 in M modalities. For specific implementation details, please refer to [link to relevant documentation]. Figure 9 The content described in step S303 will not be repeated here in this embodiment of the application. By comparing the differences between the feature representation 111a of sample identification object 1 in M modalities and the decoded feature representation 115a of sample identification object 1 in M modalities, the differences between the feature representation 111b of sample identification object 2 in M modalities and the decoded feature representation 115b of sample identification object 2 in M modalities, and the differences between the feature representation 111c of sample identification object 3 in M modalities and the decoded feature representation 115c of sample identification object 3 in M modalities, a first loss function 116a is generated.
[0220] like Figure 11As shown, after the computer device obtains the predicted candidate fusion feature representation corresponding to each sample object, it can use the gating vector 117a in contrastive learning to encode the predicted candidate fusion feature representation corresponding to each sample object, thereby obtaining the sample object 1 in mode M. i The explicit fusion feature representation 118a, sample identification object 2 in modality M i The explicit fusion feature representation 118b and the sample identification object 2 in modality M i The explicit fusion feature representation is 118c. Based on the sample, object 1 is identified in modality M. i The explicit fusion feature representation 118a, sample identification object 2 in modality M i The explicit fusion feature representation 118b and the sample identification object 2 in modality M i The explicit fusion feature representation 118c under the condition is used to determine the sample identification object 1, sample identification object 2, and sample identification object 3 in modality M. i The predictive explicit contrast relationship between object similarity is shown in 119a. For details on its implementation, please refer to [link / reference needed]. Figure 9 The content described in step S304 will not be repeated here in this embodiment. The computer device can acquire modal M. i The corresponding explicit contrast relationship label 1110a is used to generate a second loss function 1111a by comparing the difference between the predicted explicit contrast relationship 119a and the explicit contrast relationship label 1110a.
[0221] like Figure 11 As shown, after the computer device generates a first loss function 116a and a second loss function 1111a, it can adjust the model parameters in the sub-encoder and master encoder of the initial autoencoder model based on the first loss function 116a and the second loss function 1111a to obtain the target autoencoder model. Specifically, the computer device can weight the second loss function 1111a using loss weights to obtain a weighted second loss function. By summing the first loss function 116a and the weighted second loss function, a total loss function is generated. Using the total loss function, the model parameters in the initial autoencoder model are adjusted to generate the target autoencoder model.
[0222] In this embodiment of the application, an initial autoencoder model and S sample recognition objects are obtained, and the initial autoencoder model is used to identify the S sample recognition objects. j Feature encoding is performed on the sample feature representations under M modalities to obtain the sample encoded features corresponding to the sample feature representations under the M modalities. Feature fusion is then performed on the M sample encoded features to obtain the sample recognition object S. jThe corresponding predicted candidate fusion feature representation. Based on the sample identification object S... j Sample feature representation in M modalities and the sample identification object S j The corresponding predicted candidate fusion feature representation is used to generate the first loss function. Thus, by calculating the sample identification object S... j Sample feature representation in M modalities and the sample identification object S j The difference between the corresponding predicted candidate fusion feature representations is used to train the initial autoencoder model, enabling the trained target autoencoder model to accurately reconstruct the feature representations before fusion. Modality M is obtained from the M modalities. i The corresponding gating vector, based on the mode M i The corresponding gating vectors are used to encode the predicted candidate fusion feature representations corresponding to the S sample recognition objects, respectively, to generate the S sample recognition objects in the modality M. i The corresponding predicted explicit fusion feature representations are given below, and the objects identified based on the S samples are in the modality M. i The corresponding predictive explicit fusion feature representations are used to determine the S sample identification objects in the modality M. i The predictive explicit contrast relationship between object similarity under the given modality M. i The corresponding explicit contrastive relation labels and the predicted explicit contrastive relations are used to generate a second loss function. By incorporating contrastive learning and gating vectors, the differences between the explicit contrastive relation labels of the identified objects before fusion and the predicted explicit contrastive labels after fusion are compared. Based on this difference, the model parameters of the initial autoencoder model are adjusted, enabling the trained target autoencoder model to accurately reproduce the explicit characteristics (such as the correlation between the identified objects) of each identified object before fusion. Based on the first loss function and the second loss function, the model parameters of the initial autoencoder model are adjusted to obtain the target autoencoder model; the target autoencoder model is used to predict the explicit fused feature representations corresponding to the feature representations of the identified objects in M modalities.
[0223] Please see Figure 12 , Figure 12 This is a schematic diagram of the structure of a data processing apparatus 1 provided in an embodiment of this application. The data processing apparatus 1 can be a computer program (including program code) running on a computer device; for example, the data processing apparatus 1 is an application software. The data processing apparatus 1 can be used to execute corresponding steps in the data processing method provided in the embodiments of this application. Figure 12As shown, the data processing device 1 may include: a first feature encoding module 11, a feature fusion module 12, a first acquisition module 13, a second feature encoding module 14, a second acquisition module 15, a vector dot product module 16, and a first determination module 17.
[0224] The first feature encoding module 11 is used to obtain the feature representation of the object to be identified in M modalities, and to encode the feature representations in the M modalities respectively to obtain the encoded features corresponding to the feature representations in the M modalities respectively; M is a positive integer;
[0225] Feature fusion module 12 is used to fuse M encoded features to obtain candidate fused feature representations corresponding to M modalities; the M modalities include modality M i , where i is a positive integer less than or equal to M;
[0226] The first acquisition module 13 is used to acquire mode M. i The corresponding gating vector; the gating vector is used to highlight different recognition objects in modality M. i The explicit contrast relationship between the similarity of objects;
[0227] The second feature encoding module 14 is used to encode the modality M. i The corresponding gating vector is used to encode the candidate fusion feature representation to obtain the object in modality M. i Explicit fusion feature representation under the following conditions; object recognition in modality M i The explicit fusion feature representation is used to identify the object to be identified and the object to be matched in mode M. i The similarity of objects.
[0228] The first feature encoding module 11 includes:
[0229] Input unit 1101 is used to input the feature representation of the object in M modalities into the target autoencoder model;
[0230] Feature encoding unit 1102 is used to obtain mode M in the target autoencoder model. i The corresponding encoding weight matrix in the sub-encoder is used to apply the encoding weight matrix to mode M. i The feature representation below is used for feature encoding to obtain mode M. i The features below represent the corresponding candidate encoded features;
[0231] The activation processing unit 1103 is used to activate the candidate coding features corresponding to the feature representations of the M modalities respectively through the activation function, so as to obtain the coding features corresponding to the feature representations of the M modalities respectively.
[0232] The feature fusion module 12 includes:
[0233] Attention feature encoding unit 1201 is used to input M encoded features into the main encoder of the target autoencoder model, perform attention feature encoding on the M encoded features, and obtain attention encoded features corresponding to the M encoded features respectively.
[0234] The first feature concatenation unit 1202 is used to concatenate M attention-encoded features to obtain candidate fusion feature representations corresponding to M modalities.
[0235] Specifically, the attention feature encoding unit 1201 is used for:
[0236] The main encoder generates a query vector, a key vector, and a value vector for each of the M encoded features;
[0237] Obtain the encoded feature T i Corresponding query vector and encoding feature T i The dot product of the eigenvectors of the corresponding key vectors is used to determine the encoded feature T. i The corresponding attention matrix; encoded feature T i It belongs to M encoded features, where i is a positive integer less than or equal to M;
[0238] Based on coding feature T i The corresponding attention matrix, for encoding feature T i Attention feature encoding is performed to obtain encoded feature T. i The corresponding attention encoding features.
[0239] Among them, mode M i The corresponding gating vectors include dominant sub-vectors and shielding sub-vectors;
[0240] The second feature encoding module 14 includes:
[0241] The first vector dot product unit 1401 is used to perform vector dot product on the candidate fusion feature representation using explicit sub-vectors to obtain the explicit fusion features in the candidate fusion feature representation;
[0242] The second vector dot product unit 1402 uses a masked sub-vector to perform a vector dot product on the candidate fusion feature representation to obtain the masked fusion feature in the candidate fusion feature representation.
[0243] The second feature splicing unit 1403 splices the explicit fusion feature and the masked fusion feature to obtain the object in modality M. i The explicit fusion feature representation.
[0244] The data processing device 1 mentioned above also includes:
[0245] The second acquisition module 15 is used to acquire the object to be matched and identified in mode M. i The explicit fusion feature representation;
[0246] Vector dot product module 16 is used to identify the object to be matched in modality M. i The explicit fusion feature representation and recognition of the object in modality M i The explicit fusion feature representations are multiplied by vector dot product to obtain the vector dot product result;
[0247] The determination module 17 is used to determine the vector dot product result as the object similarity between the identified object and the object to be matched.
[0248] In this embodiment, feature representations of the object to be identified in M modalities are obtained. These representations are then encoded to obtain coded features corresponding to each of the M modalities. Finally, these coded features are fused to obtain candidate fused feature representations for the M modalities. By fusing the feature representations of the M modalities to obtain candidate fused feature representations for the M modalities, the complementarity between the M modalities can be utilized to remove redundant information from each modality or to supplement missing information, thereby improving the accuracy of feature extraction. (Modality M is then obtained.) i The corresponding gating vector is used to highlight different recognition objects in modality M. i The explicit contrastive relationship between object similarities under modality M. i The corresponding gating vector is used to encode the candidate fusion feature representation to obtain the object in modality M. i The explicit fusion feature representation under the following conditions allows for the identification of objects in modality M. i The explicit fusion feature representation is used to identify the object to be identified and the object to be matched in mode M. i The similarity between objects is calculated. The similarity between the objects to be matched and identified in modality M is obtained. i The explicit fusion feature representation under the condition is based on the object being identified in modality M. i The explicit fusion feature representation and the object to be matched and identified in modality M iThe explicit fusion feature representation determines the object similarity between the identified object and the object to be matched. This object similarity can be used for business recommendation or object matching. Thus, by using gating vectors to encode the candidate fusion feature representations, the resulting explicit fusion feature representation can be used to determine the object similarity between the identified object and other identified objects in various modalities. This preserves the correlation between the identified object and other identified objects in the feature representation before fusion, and even though the resulting explicit fusion feature representation contains the explicit characteristics of the feature representation before fusion, it can further improve the accuracy of feature fusion, thereby accurately determining the similarity of the identified object in modality M. i The explicit fusion feature representation is used to ensure the accuracy of business processing results. In addition, by performing feature fusion on the feature representations of the identified object in M modalities, the high-dimensional feature representations in M modalities can be merged into a lower-dimensional vector space. That is, the high-dimensional feature representations in M modalities are merged into a low-dimensional explicit fusion feature representation, which can reduce feature storage space and remove noise signals in the feature representations in M modalities.
[0249] Please see Figure 13 , Figure 13 This is a schematic diagram of the structure of a data processing device 2 provided in an embodiment of this application. The data processing device 2 can be a computer program (including program code) running on a computer device; for example, the data processing device 2 is an application software. The data processing device 2 can be used to execute corresponding steps in the data processing method provided in the embodiments of this application. Figure 13 As shown, the data processing device 2 may include: a third acquisition module 21, a first generation module 22, a second generation module 23, and a model parameter adjustment module 24.
[0250] The third acquisition module 21 is used to acquire the initial autoencoder model and S sample recognition objects; the S sample recognition objects include sample recognition object S j S is a positive integer, and j is a positive integer less than S.
[0251] The first generation module 22 is used to identify the sample object S using an initial autoencoder model. j Feature fusion is performed on the encoded features of samples in M modalities to obtain the sample identification object S. j The corresponding predicted candidate fusion feature representation, based on the sample identification object S j The corresponding predicted candidate fusion feature representation generates the first loss function; the M modes include mode M i M is a positive integer, and i is a positive integer less than or equal to M.
[0252] The second generation module 23 is used to obtain the mode M. iThe corresponding gate vector, based on mode M i The corresponding gating vectors are used to encode the predicted candidate fusion feature representations of the S sample recognition objects, respectively, to generate the S sample recognition objects in mode M. i The corresponding predicted explicit fusion feature representations are shown below, based on modality M. i The corresponding explicit contrastive labels and the S sample identification objects in modality M i The corresponding predicted explicit fusion feature representations are used to generate the second loss function;
[0253] The model parameter adjustment module 24 is used to adjust the model parameters of the initial autoencoder model according to the first loss function and the second loss function to obtain the target autoencoder model; the target autoencoder model is used to predict the explicit fusion feature representations corresponding to the feature representations of the identified object in M modalities.
[0254] The first generation module 22 includes:
[0255] The sample feature fusion module 2201 is used to obtain the sample recognition object S. j The sample feature representation under M modalities is based on an initial autoencoder model, which is used to identify the sample object S. j Feature encoding is performed on the sample feature representations under M modalities to obtain the sample identification object S. j The sample encoding features under M modalities are used to identify the sample object S. j Feature fusion is performed on the encoded features of samples in M modalities to obtain the sample identification object S. j The corresponding predicted candidate fusion feature representation;
[0256] Feature decoding unit 2202 is used to identify object S from the sample. j The corresponding predicted candidate fusion feature representation is used for feature decoding to obtain the sample identification object S. j Decoding feature representation in M modalities;
[0257] The first probability distribution transformation unit 2203 is used to identify the object S from the sample. j The sample feature representations under M modalities are subjected to probability distribution transformation to obtain the sample identification object S. j The sample feature representations under M modalities correspond to the first probability distribution;
[0258] The second probability distribution transformation unit 2204 is used to identify the object S from the sample. j The probability distribution transformation is performed on the decoded feature representations under M modalities to obtain the sample identification object S. j The second probability distribution corresponding to the decoded feature representation in M modalities;
[0259] The first generation unit 2205 is used to generate a sample recognition object S based on a first probability distribution and a second probability distribution. j The corresponding object loss function;
[0260] The second generation unit 2206 is used to generate a first loss function based on the object loss function corresponding to each of the S samples.
[0261] The model parameter adjustment module 24 includes:
[0262] The weighted processing unit 2401 is used to obtain the loss weights and apply the loss weights to the second loss function to obtain the weighted second loss function.
[0263] The summation processing unit 2402 is used to sum the first loss function and the weighted second loss function to obtain the total loss function;
[0264] The model parameter adjustment unit 2403 is used to adjust the model parameters of the initial autoencoder model according to the total loss function to obtain the target autoencoder model.
[0265] Specifically, the model parameter adjustment unit 2403 is used for:
[0266] Based on the total loss function, determine the target model adjustment parameters used to adjust the parameters of the sub-encoder and master encoder in the initial autoencoder model;
[0267] The initial model parameters of the sub-encoder and master encoder in the initial autoencoder model are adjusted to the target model adjustment parameters to obtain the parameter-adjusted initial autoencoder model;
[0268] When the initial autoencoder model after parameter adjustment meets the training convergence condition, the initial autoencoder model after parameter adjustment is determined as the target autoencoder model.
[0269] Among them, the S sample identification objects also include sample identification object S j-1 and sample identification object S j+1 ;
[0270] The second generation module 23 includes:
[0271] The third vector dot product unit 2301 is used to identify the sample object S. j In mode M i The predictive explicit fusion feature representation, and the sample identification object S j-1 In mode M i The predicted explicit fusion feature representation is multiplied by a vector to obtain the sample identification object S. j With sample identification object S j-1 The first predicted similarity between them;
[0272] The fourth vector dot product unit 2302 is used to identify the sample object S. j In mode M i The predictive explicit fusion feature representation, and the sample identification object S j+1 In mode M i The predicted explicit fusion feature representation is multiplied by a vector to obtain the sample identification object S. j With sample identification object S j+1 The second predicted similarity between them;
[0273] The determining unit 2303 is used to determine the prediction comparison relationship between the first predicted similarity and the second predicted similarity as the sample identification object S. j-1 Sample identification object S j and sample identification object S j+1 In mode M i The predictive explicit contrast relationship between object similarity;
[0274] The third generation unit is used to generate based on mode M. i The corresponding explicit contrastive relationship labels and predicted explicit contrastive relationships are used to generate a second loss function.
[0275] Among them, mode M i The corresponding explicit comparison labels include the comparison label between the first object similarity and the second object similarity, where the first object similarity is the sample identification object S. j-1 and sample identification object S j The similarity labels between the samples are used to identify the objects S. j and sample identification object S j+1 Similarity tags between them;
[0276] The third generation unit includes:
[0277] Obtain a sub-unit, which is used to obtain the difference between the contrast learning threshold and the second predicted similarity if the explicit contrast relationship label is that the similarity of the first object is less than that of the second object.
[0278] The summation processing subunit is used to sum the difference with the first predicted similarity to obtain the similarity parameter;
[0279] A sub-unit is generated to produce a second loss function based on the similarity parameter.
[0280] In this embodiment of the application, an initial autoencoder model and S sample recognition objects are obtained, and the initial autoencoder model is used to identify the S sample recognition objects. jFeature encoding is performed on the sample feature representations under M modalities to obtain the sample encoded features corresponding to the sample feature representations under the M modalities. Feature fusion is then performed on the M sample encoded features to obtain the sample recognition object S. j The corresponding predicted candidate fusion feature representation. Based on the sample identification object S... j Sample feature representation in M modalities and the sample identification object S j The corresponding predicted candidate fusion feature representation is used to generate the first loss function. Thus, by calculating the sample identification object S... j Sample feature representation in M modalities and the sample identification object S j The difference between the corresponding predicted candidate fusion feature representations is used to train the initial autoencoder model, enabling the trained target autoencoder model to accurately reconstruct the feature representations before fusion. Modality M is obtained from the M modalities. i The corresponding gating vector, based on the mode M i The corresponding gating vectors are used to encode the predicted candidate fusion feature representations corresponding to the S sample recognition objects, respectively, to generate the S sample recognition objects in the modality M. i The corresponding predicted explicit fusion feature representations are given below, and the objects identified based on the S samples are in the modality M. i The corresponding predictive explicit fusion feature representations are used to determine the S sample identification objects in the modality M. i The predictive explicit contrast relationship between object similarity under the given modality M. i The corresponding explicit contrastive relation labels and the predicted explicit contrastive relations are used to generate a second loss function. By incorporating contrastive learning and gating vectors, the differences between the explicit contrastive relation labels of the identified objects before fusion and the predicted explicit contrastive labels after fusion are compared. Based on this difference, the model parameters of the initial autoencoder model are adjusted, enabling the trained target autoencoder model to accurately reproduce the explicit characteristics (such as the correlation between the identified objects) of each identified object before fusion. Based on the first loss function and the second loss function, the model parameters of the initial autoencoder model are adjusted to obtain the target autoencoder model; the target autoencoder model is used to predict the explicit fused feature representations corresponding to the feature representations of the identified objects in M modalities.
[0281] Please see Figure 14 , Figure 14 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 14As shown, the computer device 1000 may include a processor 1001, a network interface 1004, and a memory 1005. Furthermore, the computer device 1000 may also include a target user interface 1003 and at least one communication bus 1002. The communication bus 1002 is used to enable communication between these components. The target user interface 1003 may include a display screen and a keyboard; optionally, the target user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the processor 1001. Figure 14 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a target user interface module, and a device control application.
[0282] exist Figure 14 In the computer device 1000 shown, the network interface 1004 provides network communication functionality; the target user interface 1003 is mainly used to provide an input interface for the target user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:
[0283] Obtain the feature representations of the object to be identified in M modalities, and encode the feature representations in each of the M modalities to obtain the encoded features corresponding to the feature representations in the M modalities respectively; M is a positive integer;
[0284] Feature fusion is performed on M encoded features to obtain candidate fused feature representations corresponding to M modalities; the M modalities include modality M i , where i is a positive integer less than or equal to M;
[0285] Obtaining mode M i The corresponding gating vector; the gating vector is used to highlight different recognition objects in modality M. i The explicit contrast relationship between the similarity of objects;
[0286] According to mode M i The corresponding gating vector is used to encode the candidate fusion feature representation to obtain the object in modality M. i Explicit fusion feature representation under the following conditions; object recognition in modality M i The explicit fusion feature representation is used to identify the object to be identified and the object to be matched in mode M. iThe similarity of objects.
[0287] It should be understood that the computer device 1000 described in the embodiments of this application can execute the foregoing text. Figure 3 The description of the data processing method in the corresponding embodiments can also be performed as described above. Figure 12 The description of the data processing device 1 in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated here.
[0288] Please see Figure 15 , Figure 15 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 15 As shown, the computer device 2000 may include a processor 2001, a network interface 2004, and a memory 2005. Furthermore, the computer device 2000 may also include a user interface 2003 and at least one communication bus 2002. The communication bus 2002 is used to enable communication between these components. The user interface 2003 may include a display screen and a keyboard; optionally, the user interface 2003 may also include a standard wired interface or a wireless interface. Optionally, the network interface 2004 may include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 2005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 2005 may also be at least one storage device located remotely from the processor 2001. Figure 15 As shown, the memory 2005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0289] In such Figure 15 In the computer device 2000 shown, the network interface 2004 provides network communication functionality; the user interface 2003 is mainly used to provide an input interface for the user; and the processor 2001 can be used to call the device control application program stored in the memory 2005 to achieve:
[0290] Obtain the initial autoencoder model and S sample recognition objects; the S sample recognition objects include sample recognition object S j S is a positive integer, and j is a positive integer less than S.
[0291] The initial autoencoder model is used to identify the sample object S. j Feature fusion is performed on the encoded features of samples in M modalities to obtain the sample identification object S. jThe corresponding predicted candidate fusion feature representation, based on the sample identification object S j The corresponding predicted candidate fusion feature representation generates the first loss function; the M modes include mode M i M is a positive integer, and i is a positive integer less than or equal to M.
[0292] Obtaining mode M i The corresponding gate vector, based on mode M i The corresponding gating vectors are used to encode the predicted candidate fusion feature representations of the S sample recognition objects, respectively, to generate the S sample recognition objects in mode M. i The corresponding predicted explicit fusion feature representations are shown below, based on modality M. i The corresponding explicit contrastive labels and the S sample identification objects in modality M i The corresponding predicted explicit fusion feature representations are used to generate the second loss function;
[0293] Based on the first loss function and the second loss function, the model parameters of the initial autoencoder model are adjusted to obtain the target autoencoder model; the target autoencoder model is used to predict the explicit fusion feature representations corresponding to the feature representations of the identified object in M modalities.
[0294] It should be understood that the computer device 2000 described in the embodiments of this application can execute the foregoing text. Figure 7 The description of the data processing method in the corresponding embodiments can also be performed as described above. Figure 13 The description of the data processing device 2 in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated.
[0295] Furthermore, it should be noted that this application embodiment also provides a computer-readable storage medium storing a computer program executed by the aforementioned data processing device. The computer program includes program instructions, which, when executed by a processor, enable the execution of the aforementioned data processing device. Figure 3 , Figure 7 or Figure 9 The description of the data processing method in the corresponding embodiments is already provided and will not be repeated here. Similarly, the beneficial effects of using the same method will not be repeated here. For technical details not disclosed in the computer-readable storage medium embodiments related to this application, please refer to the description of the method embodiments of this application. As an example, program instructions can be deployed and executed on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network. These multiple computing devices distributed across multiple locations and interconnected via a communication network can constitute a blockchain system.
[0296] Furthermore, it should be noted that this application also provides a computer program product or computer program, which may include computer instructions, which may be stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor may execute the computer instructions, causing the computer device to perform the aforementioned actions. Figure 3 , Figure 7 or Figure 9 The description of the data processing method in the corresponding embodiments is already provided and will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer program products or computer program embodiments related to this application, please refer to the description of the method embodiments of this application.
[0297] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0298] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.
[0299] The modules in the device of this application embodiment can be merged, divided, and deleted according to actual needs.
[0300] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0301] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A data processing method, characterized in that, include: The feature representations of the identified object in M modalities are obtained, and feature encoding is performed on the feature representations in the M modalities respectively to obtain the encoded features corresponding to the feature representations in the M modalities respectively; M is a positive integer; the M modalities include images, text and sound; the identified object is a user or an item; Feature fusion is performed on the M encoded features to obtain candidate fused feature representations corresponding to the M modalities; The M modes include mode M i , where i is a positive integer less than or equal to M; Obtain the mode M i The corresponding gating vector; the gating vector is used to highlight different recognition objects in the modality M. i The explicit contrastive relationship between object similarity; the modality M i The corresponding gating vectors include dominant sub-vectors and shielding sub-vectors; Using the explicit sub-vector, a vector dot product is performed on the candidate fusion feature representation to obtain the explicit fusion feature in the candidate fusion feature representation; Using the masked sub-vector, a vector dot product is performed on the candidate fusion feature representation to obtain the masked fusion feature in the candidate fusion feature representation; The explicit fusion feature and the masked fusion feature are concatenated to obtain the identified object in modality M. i The explicit fusion feature representation; The identified object is in mode M i The explicit fusion feature representation is used to identify the object to be identified and the object to be matched in the modality M. i The similarity of objects.
2. The method according to claim 1, characterized in that, The step of encoding the feature representations under the M modalities to obtain the encoded features corresponding to the feature representations under the M modalities respectively includes: The feature representations of the identified object in M modalities are input into the target autoencoder model; Obtain the mode M in the target autoencoder model i The corresponding encoding weight matrix in the sub-encoder is used to apply the encoding weight matrix to the mode M. i The feature representation below is used for feature encoding to obtain the mode M. i The features below represent the corresponding candidate encoded features; By using activation functions, the candidate coding features corresponding to the feature representations of the M modalities are activated respectively, thereby obtaining the coding features corresponding to the feature representations of the M modalities respectively.
3. The method according to claim 1, characterized in that, The feature fusion of the M encoded features to obtain candidate fused feature representations corresponding to the M modalities includes: M encoded features are input into the main encoder of the target autoencoder model, and attention feature encoding is performed on the M encoded features to obtain the attention encoded features corresponding to the M encoded features respectively; The M attention-encoded features are concatenated to obtain the candidate fusion feature representations corresponding to the M modalities.
4. The method according to claim 3, characterized in that, The step of performing attention feature encoding on the M encoded features to obtain attention encoded features corresponding to the M encoded features respectively includes: The main encoder generates a query vector, a key vector, and a value vector for each of the M encoded features; Obtain the encoded feature T i The corresponding query vector and the encoded feature T i The dot product of the eigenvectors between the corresponding key vectors is used to determine the encoded feature T. i The corresponding attention matrix; the encoded feature T i Belonging to the M encoded features, where i is a positive integer less than or equal to M; According to the coding feature T i The corresponding attention matrix, for the encoded feature T i Attention feature encoding is performed to obtain the encoded feature T. i The corresponding attention encoding features.
5. The method according to claim 1, characterized in that, The method further includes: Obtain the object to be matched and identified in the modality M i The explicit fusion feature representation; For the object to be matched and identified in the modality M i The explicit fusion feature representation of the identified object in modality M i The explicit fusion feature representations are multiplied by vector dot product to obtain the vector dot product result; The result of the vector dot product is determined as the object similarity between the identified object and the object to be matched.
6. A data processing method, characterized in that, include: Obtain an initial autoencoder model and S sample recognition objects; the S sample recognition objects include sample recognition object S. j S is a positive integer, and j is a positive integer less than S. The initial autoencoder model is used to identify the sample object S. j Feature fusion is performed on the encoded features of samples in M modalities to obtain the sample recognition object S. j The corresponding predicted candidate fusion feature representation, based on the sample identification object S j The corresponding predicted candidate fusion feature representation generates the first loss function; the M modes include mode M i M is a positive integer, and i is a positive integer less than or equal to M; the M modalities include images, text, and sound; the sample recognition object S j For users or items; Obtain the mode M i The corresponding gating vector, the mode M i The corresponding gating vectors include dominant sub-vectors and shielding sub-vectors; Using the explicit sub-vectors, vector dot products are performed on the predicted candidate fusion feature representations corresponding to the S sample identification objects to obtain the explicit fusion features corresponding to the S predicted candidate fusion feature representations respectively; Using the masked sub-vector, vector dot product is performed on the predicted candidate fusion feature representations corresponding to the S sample identification objects to obtain the masked fusion features corresponding to the S predicted candidate fusion feature representations respectively; The explicit fusion feature and the masked fusion feature in the same predicted candidate fusion feature representation are concatenated to obtain the S sample identification objects in modality M. i The corresponding predicted explicit fusion feature representations are shown below, based on the modality M. i The corresponding explicit contrastive relationship labels and the S sample identification objects in the modality M i The corresponding predicted explicit fusion feature representations are used to generate the second loss function; Based on the first loss function and the second loss function, the model parameters of the initial autoencoder model are adjusted to obtain the target autoencoder model; the target autoencoder model is used to predict the explicit fusion feature representations corresponding to the feature representations of the identified object in M modalities.
7. The method according to claim 6, characterized in that, The initial autoencoder model is used to identify the sample object S. j Feature fusion is performed on the encoded features of samples in M modalities to obtain the sample recognition object S. j The corresponding predicted candidate fusion feature representation, based on the sample identification object S j The corresponding predicted candidate fusion feature representation generates the first loss function, including: Obtain the sample recognition object S j The sample feature representation under M modalities is used, employing the initial autoencoder model, to identify the sample object S respectively. j Feature encoding is performed on the sample feature representations under M modalities to obtain the sample identification object S. j The sample encoding features under M modalities are used to identify the sample object S. j Feature fusion is performed on the encoded features of samples in M modalities to obtain the sample recognition object S. j The corresponding predicted candidate fusion feature representation; For the sample identification object S j The corresponding predicted candidate fusion feature representation is used for feature decoding to obtain the sample identification object S. j Decoding feature representation in M modalities; For the sample identification object S j The sample feature representations under M modalities are subjected to probability distribution transformation to obtain the sample identification object S. j The sample feature representations under M modalities correspond to the first probability distribution; For the sample identification object S j The probability distribution transformation is performed on the decoded feature representations under M modalities to obtain the sample recognition object S. j The second probability distribution corresponding to the decoded feature representation in M modalities; The sample identification object S is generated based on the first probability distribution and the second probability distribution. j The corresponding object loss function; The first loss function is generated based on the object loss function corresponding to each of the S sample objects.
8. The method according to claim 6, characterized in that, The step of adjusting the model parameters of the initial autoencoder model according to the first loss function and the second loss function to obtain the target autoencoder model includes: Obtain the loss weights, and use the loss weights to weight the second loss function to obtain the weighted second loss function; The first loss function and the weighted second loss function are summed to obtain the total loss function. Based on the total loss function, the model parameters of the initial autoencoder model are adjusted to obtain the target autoencoder model.
9. The method according to claim 8, characterized in that, The step of adjusting the model parameters of the initial autoencoder model according to the total loss function to obtain the target autoencoder model includes: Based on the total loss function, determine the target model adjustment parameters for adjusting the parameters of the sub-encoder and main encoder in the initial autoencoder model; The initial model parameters of the sub-encoder and the main encoder in the initial autoencoder model are adjusted to the target model adjustment parameters to obtain the parameter-adjusted initial autoencoder model; When the initial autoencoder model with adjusted parameters meets the training convergence condition, the initial autoencoder model with adjusted parameters is determined as the target autoencoder model.
10. The method according to claim 6, characterized in that, The S sample identification objects also include sample identification object S j-1 and sample identification object S j+1 ; According to the mode M i The corresponding explicit contrastive relationship labels and the S sample identification objects in the modality M i The corresponding predicted explicit fusion feature representations are used to generate a second loss function, including: For the sample identification object S j In the mode M i The predicted explicit fusion feature representation, and the sample identification object S j-1 In the mode M i The predicted explicit fusion feature representation is multiplied by a vector to obtain the sample identification object S. j With the sample identification object S j-1 The first predicted similarity between them; For the sample identification object S j In the mode M i The predicted explicit fusion feature representation, and the sample identification object S j+1 In the mode M i The predicted explicit fusion feature representation is multiplied by a vector to obtain the sample identification object S. j With the sample identification object S j+1 The second predicted similarity between them; The prediction comparison relationship between the first predicted similarity and the second predicted similarity is determined as the sample identification object S. j-1 The sample identification object S j and the sample identification object S j+1 In the mode M i The predictive explicit contrast relationship between object similarity; According to the mode M i The corresponding explicit contrastive relation labels and the predicted explicit contrastive relation are used to generate a second loss function.
11. The method according to claim 10, characterized in that, The mode M i The corresponding explicit comparison labels include the comparison label between the first object similarity and the second object similarity, where the first object similarity is the sample identification object S. j-1 and the sample identification object S j The similarity label between the two objects, wherein the second object similarity is the sample identification object S. j and the sample identification object S j+1 Similarity tags between them; According to the mode M i The corresponding explicit contrastive relation labels and the predicted explicit contrastive relation are used to generate a second loss function, including: If the explicit contrast relationship label is that the similarity of the first object is less than that of the second object, then the difference between the contrast learning threshold and the second predicted similarity is obtained; The difference is summed with the first predicted similarity to obtain the similarity parameter; A second loss function is generated based on the similarity parameters.
12. A data processing apparatus, characterized in that, include: The first feature encoding module is used to acquire the feature representation of the object to be identified in M modalities, and to encode the feature representations in the M modalities respectively to obtain the encoded features corresponding to the feature representations in the M modalities respectively; M is a positive integer; the M modalities include images, text and sound; the object to be identified is a user or an item; The feature fusion module is used to fuse M encoded features to obtain candidate fused feature representations corresponding to the M modalities; The M modes include mode M i , where i is a positive integer less than or equal to M; Acquisition module, used to acquire the mode M i The corresponding gating vector; the gating vector is used to highlight different recognition objects in the modality M. i The explicit contrastive relationship between object similarity; the modality M i The corresponding gating vectors include dominant sub-vectors and shielding sub-vectors; The second feature encoding module is used to perform a vector dot product on the candidate fusion feature representation using the explicit sub-vector to obtain the explicit fusion feature in the candidate fusion feature representation; to perform a vector dot product on the candidate fusion feature representation using the masking sub-vector to obtain the masked fusion feature in the candidate fusion feature representation; and to concatenate the explicit fusion feature and the masked fusion feature to obtain the recognition object in modality M. i The explicit fusion feature representation; The identified object is in mode M i The explicit fusion feature representation is used to identify the object to be identified and the object to be matched in the modality M. i The similarity of objects.
13. A computer device, characterized in that, include: Processor and memory; The processor is connected to a memory, wherein the memory is used to store a computer program, and the processor is used to invoke the computer program to cause the computer device to perform the method according to any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded and executed by a processor to cause a computer device having the processor to perform the method of any one of claims 1-11.