Method and apparatus for processing multimodal data
By feature screening the initial features of multimodal data, redundant features are removed, and cross-modal joint features are generated, the inefficiency problem caused by redundant features in multimodal data processing is solved, and more efficient data processing is achieved.
Patent Information
- Application Number
- CN202411925166.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-12-25
AI Technical Summary
In multimodal data processing, redundant features may exist in the extracted local features, resulting in low processing efficiency.
By obtaining the initial features of multimodal data, feature screening is performed to remove redundant features, key data features are obtained, and cross-modal joint features are generated using these features.
It effectively removes redundant features and improves the processing accuracy and efficiency of multimodal data.
Smart Images

Figure CN119357904B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computers, and more specifically, to a method and device for processing multimodal data. Background Art
[0002] In the related art, when processing multimodal data, redundant features may exist in the local features of the extracted multimodal data, which may affect the processing accuracy of the multimodal data, and further lead to the problem of low processing efficiency of the multimodal data. Therefore, there is a problem of low processing efficiency of multimodal data.
[0003] For the above problems, no effective solution has been proposed yet. Summary of the Invention
[0004] The embodiments of the present application provide a method and device for processing multimodal data to at least solve the problem of low processing efficiency of multimodal data in the related art.
[0005] According to an embodiment of the present application, a method for processing multimodal data is provided, including:
[0006] Obtaining input multimodal data, where the multimodal data includes at least two types of modal data; obtaining multiple initial data features corresponding to each type of modal data in the at least two types of modal data; performing feature screening on the multiple initial data features to obtain at least one key data feature, where the key data feature is a feature in the initial data features that can represent preset information of the multimodal data; and using the key data feature to output a cross-modal joint feature corresponding to the multimodal data.
[0007] According to another embodiment of the present application, a device for processing multimodal data is provided, including:
[0008] A first obtaining unit for obtaining input multimodal data, where the multimodal data includes at least two types of modal data; a second obtaining unit for obtaining multiple initial data features corresponding to each type of modal data in the at least two types of modal data; a screening unit for performing feature screening on the multiple initial data features to obtain at least one key data feature, where the key data feature is a feature in the initial data features that can represent preset information of the multimodal data; and an output unit for using the key data feature to output a cross-modal joint feature corresponding to the multimodal data.
[0009] As an alternative solution, the above-mentioned screening unit includes: a construction module for constructing corresponding multiple key weights for each of the multiple initial data features, where the key weights are used to represent the retention degree of the data features; an integration module for integrating each of the initial data features with the corresponding key weights to obtain an integration result corresponding to each of the initial data features; and a screening module for using the integration result to perform feature screening on the multiple initial data features to obtain the at least one key data feature.
[0010] As an alternative solution, the above-mentioned construction module includes: an input sub-module for inputting the multiple initial data features into a gated recurrent unit; a capture sub-module for capturing long-term dependencies in the feature sequence through the gating mechanism of the gated recurrent unit; a first acquisition sub-module for using the long-term dependencies to acquire the hidden states output by the gated recurrent unit at each time step, where the hidden states contain all historical information up to the current time step; and a mapping sub-module for mapping the hidden states into weight vectors, where each element in the weight vectors corresponds to each of the multiple key weights.
[0011] As an alternative solution, the above-mentioned construction module includes: a regulation sub-module for using the weight coefficients of L1 regularization to regulate the sparsity regularization intensity of the multiple key weights, where the weight coefficients are iteratively updated by applying a gradient descent algorithm.
[0012] As an alternative solution, the above-mentioned output unit includes: an extraction module for performing multi-level feature extraction on the key data features to obtain multi-level features; a decoupling module for decoupling each layer of the multi-level features using different decoupling modules to obtain independent semantic features at each level; and an output module for using the independent semantic features to output the cross-modal joint features.
[0013] As an alternative solution, the above-mentioned extraction module includes: a first extraction sub-module for performing multi-level encoding on the text key features in the key data features to extract the context information of each word; and an encoding sub-module for layer-by-layer encoding of the text key features in combination with the context information to obtain text features at different semantic levels.
[0014] As an alternative solution, the above-mentioned extraction module includes: a second extraction sub-module for extracting low-level features and high-level features obtained after multi-layer convolution from the image key features in the key data features, where the features extracted after the multi-layer convolution represent feature representations at different detail granularities.
[0015] As an alternative solution, the above output unit includes: a first fusion module for performing feature fusion on multiple first features belonging to the same modality type among the above key data features to obtain N first cross-modal features, where N is a positive integer; a second fusion module for performing feature fusion on multiple second features belonging to different modality types among the above key data features to obtain M second cross-modal features, where M is a positive integer; a first acquisition module for obtaining the above cross-modal joint features based on the above N first cross-modal features and the above M second cross-modal features.
[0016] As an alternative solution, the above first fusion module includes: a determination sub-module for determining a first current feature and a second current feature from the above multiple first features; a calculation sub-module for calculating the dot product of a query matrix and a key matrix to obtain a dot product result, where the above query matrix includes the above first current feature and the above second current feature, and the above key matrix is used to calculate similarity with the above query matrix; a conversion sub-module for converting the above dot product result into an attention weight, where the above attention weight is used to represent the feature similarity between the above first current feature and the above second current feature; a summation sub-module for performing weighted summation on the above attention weight and a value matrix to obtain a weighted summation result, where the above value matrix includes the specific information of the above first current feature and the above second current feature; a first fusion sub-module for using the above weighted summation result to fuse the above first current feature and the above second current feature to obtain the above first cross-modal feature.
[0017] As an alternative solution, the above second fusion module includes: a second acquisition sub-module for obtaining a similarity matrix between the above multiple second features, where the above similarity matrix is used to reflect the similarity between the above multiple second features; a second fusion sub-module for fusing the features with high similarity among the above multiple second features based on the above similarity matrix to obtain the above M second cross-modal features.
[0018] As an alternative solution, the above first acquisition unit includes: a second acquisition module for obtaining multi-modal sample data input to a cross-modal model, where the above cross-modal model is used to output corresponding descriptive text from an image input or generate a corresponding descriptive image from a text input; an iteration module for iteratively updating the model parameters of the above cross-modal model until a model convergence condition is satisfied to obtain the trained above cross-modal model.
[0019] According to another embodiment of the present application, there is also provided a computer-readable storage medium in which a computer program is stored, where the above computer program is configured to execute the steps in any one of the above method embodiments when running.
[0020] According to another embodiment of the present application, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0021] According to another embodiment of the present application, a computer program product is further provided, including a computer program. The computer program is configured to execute the steps in any one of the above method embodiments when being executed by a processor.
[0022] Through the present application, input multimodal data is obtained, where the multimodal data includes at least two types of modal data; multiple initial data features corresponding to each type of modal data in the at least two types of modal data are obtained; the multiple initial data features are screened to obtain at least one key data feature, where the key data feature is a feature in the initial data features that can represent preset information of the multimodal data; and a cross-modal joint feature corresponding to the multimodal data is output by using the key data feature.
[0023] Specifically, after the input multimodal data is obtained, the initial data features in the multimodal data are extracted, and then the initial features are screened to remove redundant features, and then key data features that can represent the multimodal data are obtained, and then cross-modal joint features are obtained according to the key data features. Since redundant features are effectively removed through feature screening, the processing accuracy of multimodal data is improved, thereby solving the problem of low processing efficiency of multimodal data, and further achieving the technical effect of improving the processing efficiency of multimodal data. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a schematic diagram of an application environment of a method for processing multimodal data according to an embodiment of the present application;
[0025] Figure 2 is a flowchart of a method for processing multimodal data according to an embodiment of the present application;
[0026] Figure 3 is a schematic diagram of a method for processing multimodal data according to an embodiment of the present application;
[0027] Figure 4 is a schematic diagram of a method for processing multimodal data according to an embodiment of the present application;
[0028] Figure 5 is a schematic diagram of a method for processing multimodal data according to an embodiment of the present application;
[0029] Figure 6 is a schematic diagram of a method for processing multimodal data according to an embodiment of the present application;
[0030] Figure 7 is a schematic diagram of a method for processing multimodal data according to an embodiment of the present application;
[0031] Figure 8 is a schematic diagram of a method for processing multimodal data according to an embodiment of the present application;
[0032] Figure 9 is a structural block diagram of a device for processing multimodal data according to an embodiment of the present application. Detailed implementation manners
[0033] In the following, embodiments of the present application will be described in detail with reference to the accompanying drawings and in conjunction with the embodiments.
[0034] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence.
[0035] The method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking the operation on a server device as an example, Figure 1 is a hardware structural block diagram of a server device for a method of processing multimodal data according to an embodiment of the present application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in Figure 1 processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in Figure 1 is only schematic and does not limit the structure of the above-mentioned server device. For example, the server device may further include more or fewer components than
[0036] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the multi-modal data processing method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0037] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0038] In this embodiment, a method for processing multi-modal data is provided. Figure 2 is a flowchart of the method for processing multi-modal data according to the embodiments of the present application, as Figure 2 shown, and the process includes the following steps:
[0039] Step S202, obtain the input multi-modal data, where the multi-modal data includes at least two modal type data;
[0040] In an alternative embodiment, the multi-modal data may but is not limited to refer to data containing at least two modal type data, and the modal type data may but is not limited to different modal type data such as text, image, audio, video, etc.
[0041] Step S204, obtain a plurality of initial data features corresponding to each modal type data in the at least two modal type data;
[0042] In an alternative embodiment, the initial data feature may but is not limited to be understood as a data feature directly extracted from the original modal type data.
[0043] Step S206: Perform feature screening on multiple initial data features to obtain at least one key data feature, where the key data feature is a feature in the initial data features that can represent the preset information of the multimodal data;
[0044] In an alternative embodiment, the key data feature can be, but is not limited to, a data feature obtained after performing feature screening on the initial data features, and can be, but is not limited to, a feature that can represent the preset information of the multimodal data.
[0045] It should be noted that when processing multimodal data, the initial data features of the obtained multimodal data often contain a lot of redundant information. If the initial data features are not screened, the redundant information will affect the processing process of the multimodal data, and further lead to low efficiency and accuracy in multimodal data processing.
[0046] Step S208: Use the key data features to output the cross-modal joint features corresponding to the multimodal data.
[0047] In an alternative embodiment, the cross-modal joint feature can be, but is not limited to, understood as fusing or combining the key data features of different modalities to generate a joint feature that can simultaneously represent different types of modal data.
[0048] It should be noted that based on the initial data features corresponding to each modal data in at least two modal type data, the system further performs feature screening and retains those features that can reflect the preset information during the screening process, so as to obtain the key data features. Through screening, the most relevant and representative features in at least two modal type data can be obtained, thereby removing redundant features that may have a negative impact on the processing process of multimodal data, reducing the influence of redundant features, improving the model's ability to capture key information of multimodal data, and further improving the accuracy and efficiency of multimodal data processing.
[0049] Through the above steps, obtain the input multimodal data, where the above multimodal data includes at least two modal type data; obtain multiple initial data features corresponding to each modal type data in the above at least two modal type data; perform feature screening on the above multiple initial data features to obtain at least one key data feature, where the above key data feature is a feature in the above initial data features that can represent the preset information of the above multimodal data; use the above key data features to output the cross-modal joint features corresponding to the above multimodal data.
[0050] Specifically, after obtaining the input multi-modal data, initial data features in the multi-modal data are extracted. Then, the initial features are screened to remove redundant features, and key data features that can represent the multi-modal data are obtained. Next, cross-modal joint features are obtained based on the key data features. Since redundant features are effectively removed through feature screening, the technical objective of improving the processing accuracy of multi-modal data is achieved, and further the technical effect of improving the processing efficiency of multi-modal data is achieved.
[0051] Among them, the execution subject of the above steps can be a server, a terminal, etc., but is not limited thereto.
[0052] Before screening multiple initial data features to obtain at least one key data feature, the method further includes:
[0053] S1-1, constructing corresponding multiple key weights for each initial data feature among the multiple initial data features, where the key weights are used to represent the retention degree of the data features;
[0054] S1-2, integrating each initial data feature with the corresponding key weight to obtain an integration result corresponding to each initial data feature;
[0055] Screening multiple initial data features to obtain at least one key data feature includes:
[0056] S1-3, using the integration result to screen the multiple initial data features to obtain at least one key data feature.
[0057] In an alternative embodiment, the key weight can be, but is not limited to, the weight value assigned to each initial data feature during the feature screening process, and can be, but is not limited to, used to represent the importance or retention degree of the data feature.
[0058] In an alternative embodiment, the integration process can be, but is not limited to, the processing method of combining the initial data feature with the corresponding key weight before the feature screening, and can also be, but is not limited to, adjusting the key weight corresponding to the initial data feature.
[0059] It should be noted that before feature screening, key weights need to be assigned to each initial data feature to represent the retention degree of the feature. Then, through the integration process, an integration result is obtained. Finally, feature screening is performed based on the integration result. By introducing key weights and the integration process, it is possible to more carefully control which features should be retained and which features are redundant features during the screening process, so as to more accurately generate key features that can represent multi-modal data, and further ensure the effectiveness and accuracy of feature screening.
[0060] Through the embodiments of the present application, a plurality of corresponding key weights are constructed for each of the plurality of initial data features, where the key weights are used to represent the retention degree of the data features; each initial data feature is integrated with the corresponding key weight to obtain an integration result corresponding to each initial data feature; and the integration result is used to perform feature screening on the plurality of initial data features to obtain at least one key data feature. Thus, the technical effect of more accurately generating key features that can represent multimodal data is achieved, and further, the effectiveness and accuracy of feature screening can be ensured.
[0061] As an alternative solution, constructing a plurality of corresponding key weights for each of the plurality of initial data features includes:
[0062] S2-1, inputting the plurality of initial data features into a gated recurrent unit;
[0063] S2-2, capturing long-term dependencies in the feature sequence through the gating mechanism of the gated recurrent unit;
[0064] S2-3, using the long-term dependencies to obtain the hidden states output by the gated recurrent unit at each time step, where the hidden state contains all historical information up to the current time step;
[0065] S2-4, mapping the hidden state to a weight vector, where each element in the weight vector corresponds to each of the plurality of key weights.
[0066] In an alternative embodiment, the gated recurrent unit (Gated Recurrent Unit, abbreviated as GRU) may but is not limited to include a gating mechanism, may but is not limited to be used to control the flow and update of information, and may but is not limited to capture long-term dependencies. The gating mechanism may but is not limited to include an update gate and a reset gate. The update gate may but is not limited to be used to determine how the unit state at the current moment is updated, and the reset gate may but is not limited to be used to determine how to forget or reset the old state information.
[0067] In an alternative embodiment, the long-term dependencies may but are not limited to refer to the relationship between the current state and earlier states as the time step progresses.
[0068] In an alternative embodiment, the hidden state may but is not limited to include the encoding of all historical information or historical information up to the current time step, and may but is not limited to be updated as the sequence progresses.
[0069] It should be noted that when the gated recurrent unit processes the feature sequence, its hidden state is updated at each time step based on the long-term dependencies and the current input. The hidden state encodes the historical information of the sequence at each time point in the sequence and is the key for the gated recurrent unit to capture the sequence dependencies.
[0070] In an alternative embodiment, the weight vector can be, but is not limited to, a vector generated by mapping the hidden state of the gated recurrent unit, and the elements of the weight vector can be, but are not limited to, the key weights corresponding to the initial data features, that is, each of the key weights among multiple key weights.
[0071] It should be noted that by inputting the feature sequence into the gated recurrent unit and using its gating mechanism to process the sequence features, the context relationships and long-term dependencies between the features can be effectively captured. These hidden states processed by the gated recurrent unit are then mapped to a weight vector, where each element represents the key weight corresponding to the initial feature. By using the gated recurrent unit and its gating mechanism, the feature screening process can be more finely controlled to ensure that the retained features are relatively important for the processing of multimodal data, thereby improving the effectiveness of cross-modal joint feature representation of multimodal data. At the same time, through the generation of the weight vector, quantitative control over the degree of feature retention can be achieved, further optimizing the feature screening and cross-modal fusion processes, and thus improving the processing efficiency of multimodal data.
[0072] Through the embodiments of the present application, multiple initial data features are input into the gated recurrent unit; the long-term dependencies in the feature sequence are captured through the gating mechanism of the gated recurrent unit; the hidden states output by the gated recurrent unit at each time step are obtained by using the long-term dependencies, where the hidden state contains all the historical information up to the current time step; the hidden state is mapped to a weight vector, where each element in the weight vector corresponds to each of the multiple key weights. Thus, the technical purpose of realizing quantitative control over the degree of feature retention is achieved, and furthermore, the technical effect of improving the processing efficiency of multimodal data is realized.
[0073] As an alternative solution, in the process of constructing corresponding multiple key weights for each of the multiple initial data features, the method further includes:
[0074] Using the weight coefficient of L1 regularization to adjust the sparse regularization strength of the multiple key weights, where the weight coefficient is iteratively updated by applying the gradient descent algorithm.
[0075] In an alternative embodiment, L1 regularization can, but is not limited to, be understood as a regularization technique that can, but is not limited to, ensure the sparsity of model weights by adding the sum of the absolute values of the weights (L1 norm) to the loss function, and can, but is not limited to, be used to reduce redundant features for feature selection.
[0076] In an alternative embodiment, the gradient descent algorithm can, but is not limited to, be an optimization algorithm that can, but is not limited to, be used to minimize the loss function of a machine learning model. It can, but is not limited to, calculate the gradient of the loss function with respect to the model weights and then update the weights at a certain learning rate, moving in the direction of reducing the loss. In this embodiment, the gradient descent algorithm can, but is not limited to, be used to iteratively update the weight coefficients used in L1 regularization to achieve the optimal sparse regularization strength.
[0077] In an alternative embodiment, the sparse regularization strength can, but is not limited to, refer to the degree of sparsification influence of L1 regularization on model weights.
[0078] It should be noted that L1 regularization and iterative updates of weight coefficients are used to optimize the feature screening process to achieve more accurate feature selection and model sparsification. In cross-modal learning tasks, the model usually needs to process a large amount of initial data features, which contain a lot of redundant and irrelevant information. Through L1 regularization, the weights can be selectively compressed, making the weights of unimportant features close to zero, thus achieving feature selection. The weight coefficients, on the other hand, play a role in adjusting the sparse regularization strength. By iteratively updating this coefficient through the gradient descent algorithm, a balance point can be found that can remove redundant features without over-sparsifying important features. This optimization process helps to improve the model's ability to capture key information, reduce waste of computing resources, and improve the efficiency and accuracy of cross-modal learning tasks when training or optimizing a multi-modal data processing model.
[0079] Through the embodiments of the present application, the sparse regularization strength of multiple key weights is adjusted using the weight coefficients of L1 regularization, where the weight coefficients are obtained by iteratively updating using the gradient descent algorithm. Thus, the technical objective of improving the ability to capture key information is achieved, and further the technical effect of improving the efficiency and accuracy of cross-modal learning tasks is realized.
[0080] As an alternative solution, using key data features, cross-modal joint features corresponding to multi-modal data are output, including:
[0081] S3-1, performing multi-level feature extraction on the key data features to obtain multi-level features;
[0082] S3-2, decoupling each layer of features in the multi-level features using different decoupling modules to obtain independent semantic features for each level;
[0083] S3-3. Output cross-modal joint features using independent semantic features.
[0084] In an optional embodiment, the multi-level feature extraction may, but is not limited to, refer to using different levels of a deep neural network to capture multi-level information of features during the feature extraction stage.
[0085] In an optional embodiment, in a deep learning model, the decoupling module may, but is not limited to, be used to separate different components of features of different modalities so that each component can be independently represented and optimized, thereby helping the model to more accurately understand the relationship between features of different modalities, avoiding cross-interference of information, and improving the performance of the model.
[0086] In an optional embodiment, the independent semantic features may, but are not limited to, be understood as the features processed by the decoupling module. The independent semantic features obtained at each layer may, but are not limited to, refer to the feature representations that can clearly reflect the specific semantic information of that layer.
[0087] In an optional embodiment, performing multi-level feature extraction on key data features to obtain multi-layer features may, but is not limited to, performing multi-level deep feature extraction on key data features to capture semantic information at different levels. By using different layers of a deep neural network, local features at low levels can be gradually ascended to global features at high levels layer by layer.
[0088] For further illustration, taking image features as an example, the shallow layers of a convolutional neural network can be used to capture edge and texture information, while the deep layers are used to extract semantic object and scene information. For text features, the underlying layers of a gated recurrent unit or a transformer model can be used to encode lexical information, while the high layers are used to capture the context relationships of sentences and paragraphs.
[0089] In an optional embodiment, decoupling each layer of features in the multi-layer features using different decoupling modules to obtain independent semantic features at each level may, but is not limited to, be understood as that after obtaining the multi-layer features, it is necessary to further use the decoupling module to decouple each layer of features to separate out independent semantic features. The purpose of decoupling is to ensure that the features at each level can be independently represented and optimized, avoiding cross-interference of information between features at different levels, which is very important for cross-modal understanding tasks.
[0090] In an optional embodiment, outputting cross-modal joint features using independent semantic features may, but is not limited to, be understood as fusing the decoupled independent semantic features to generate the final cross-modal joint features.
[0091] It should be noted that the representation of multimodal data is optimized by means of multi-level feature extraction and decoupling, as well as the fusion of independent semantic features, so as to improve the performance of cross-modal learning tasks. First, through different layers of a deep neural network, key data features are extracted layer by layer, capturing multi-level information of the features from low-level local details to high-level global semantics. Then, a decoupling module is used to decouple the features of each layer, separating out independent semantic features, avoiding information chaos between layers, and enabling each layer of features to independently represent specific semantic information. Finally, through a fusion strategy, the independent multi-layer features are synthesized into cross-modal joint features, which can comprehensively reflect the multi-layer semantic information of images and texts, providing a more accurate and comprehensive representation for cross-modal tasks. The technical solution of the present invention significantly improves the depth and accuracy of cross-modal understanding through hierarchical feature extraction and decoupling, laying a foundation for building a more intelligent multimodal learning system.
[0092] Through the embodiments of the present application, multi-level feature extraction is performed on key data features to obtain multi-layer features; for each layer of features in the multi-layer features, different decoupling modules are used for decoupling to obtain independent semantic features at each level; and cross-modal joint features are output using the independent semantic features. Thus, the technical purpose of obtaining cross-modal joint features is achieved, and furthermore, the technical effect of providing a more accurate and comprehensive representation for cross-modal tasks is realized.
[0093] As an alternative solution, performing multi-level feature extraction on key data features to obtain multi-layer features includes:
[0094] S4-1, performing multi-level encoding on the text key features in the key data features to extract the context information of each word;
[0095] S4-2, combining the context information and performing layer-by-layer encoding on the text key features to obtain text features at different semantic levels.
[0096] In an alternative embodiment, the file key features may but are not limited to being the key features screened out that contain text information.
[0097] In an alternative embodiment, multi-level encoding may but is not limited to being understood as, in the field of natural language processing, multi-level encoding refers to using the multi-layer capabilities of a deep neural network to perform layer-by-layer encoding on text data to capture semantic information at different levels.
[0098] In an alternative embodiment, the semantic level may but is not limited to referring to the semantic structure from the levels of words, phrases to sentences and paragraphs
[0099] It should be noted that context information is crucial for understanding the meaning of words. By encoding the key features of the text layer by layer, each layer of encoding takes into account the position of the word in the text, the semantics of the surrounding words, and their contribution to the overall meaning of the text, thereby extracting deeper context information.
[0100] Furthermore, by combining context information, the key features of the text are encoded layer by layer to obtain text features at different semantic levels. By combining context information and multi-level encoding techniques, the model can mine multi-level semantic structures from the text, deeply understand from words to phrases, sentences and even paragraphs, generate text features that contain both the surface meaning of the words and reflect their deep functions, and significantly improve the performance of cross-modal learning tasks.
[0101] Through the embodiments of this application, the key features of the text in the key data features are encoded at multiple levels to extract the context information of each word; by combining the context information, the key features of the text are encoded layer by layer to obtain text features at different semantic levels. Thus, the technical purpose of generating text features that contain both the surface meaning of the words and reflect their deep functions is achieved, and furthermore, the technical effect of improving the performance of cross-modal learning tasks is realized.
[0102] As an optional solution, multi-level feature extraction is performed on the key data features to obtain multi-level features, including:
[0103] From the key image features in the key data features, low-level features and high-level features are extracted after multiple layers of convolution of the image, where the features extracted after multiple layers of convolution respectively represent feature representations at different detail granularities.
[0104] In an optional embodiment, the key image features can be, but are not limited to, the features screened from the key data features for the image modality, and can include, but are not limited to, the main content and features of the image.
[0105] In an optional embodiment, in a convolutional neural network, multiple layers of convolution can be, but are not limited to, referring to that the network contains multiple convolutional layers, and each layer of convolution is responsible for capturing different features in the image. Shallow convolution can be, but is not limited to, used to capture low-level features, while deep convolution can be, but is not limited to, used to extract high-level features.
[0106] In an optional embodiment, the low-level features can be, but are not limited to, the features extracted from the shallow layer of a multi-layer convolutional neural network, and can be, but are not limited to, used to represent the local details of the image. The high-level features can be, but are not limited to, the features extracted from the deep layer of the key image features in a multi-layer convolutional neural network, and can be, but are not limited to, used to represent the global structure of the image.
[0107] It should be noted that low-level features and high-level features are extracted from the key features of an image through a multi-layer convolutional neural network. In the image feature extraction stage, the shallow convolutional layers of the multi-layer convolutional neural network are usually responsible for capturing the low-level features of the image, which are the basis for image understanding. The deep convolutional layers further process these low-level features to extract high-level features. The feature representations output by each convolutional layer reflect the image features at different levels of detail, from local details to global semantics, constructing a hierarchical feature space and providing a richer and more comprehensive image feature representation for cross-modal learning tasks.
[0108] Through the embodiments of the present application, low-level features and high-level features are extracted from the key features of an image in the key data features. Among them, the features extracted after multi-layer convolution respectively represent feature representations at different levels of detail. Thus, the technical purpose of constructing a hierarchical feature space is achieved, and furthermore, the technical effect of providing a richer and more comprehensive image feature representation for cross-modal learning tasks is achieved.
[0109] As an optional solution, using the key data features, cross-modal joint features corresponding to multi-modal data are output, including:
[0110] S5-1, performing feature fusion on multiple first features belonging to the same modal type in the key data features to obtain N first cross-modal features, where N is a positive integer;
[0111] S5-2, performing feature fusion on multiple second features belonging to different modal types in the key data features to obtain M second cross-modal features, where M is a positive integer;
[0112] S5-3, obtaining cross-modal joint features based on the N first cross-modal features and the M second cross-modal features.
[0113] In an optional embodiment, the multiple first features of the same modal type may but are not limited to referring to the key data features belonging to the same modality.
[0114] In an optional embodiment, the second features of different modal types may but are not limited to referring to the key data features belonging to different modalities, and may but are not limited to representing the main content of data of different modal types.
[0115] In an optional embodiment, feature fusion may but is not limited to referring to the process of integrating multiple feature vectors into a comprehensive feature representation through a certain method, and may but is not limited to performing feature fusion through methods such as concatenation and weighted summation.
[0116] It should be noted that for the key data features belonging to the same modality type, feature fusion processing is performed to obtain N first cross-modal features, which helps to integrate the information within the same modality and construct a more comprehensive and integrated intra-modal feature representation. Then, for the key data features from different modality types, cross-modal fusion is performed to obtain M second cross-modal features. These features integrate the information of different modalities and construct a cross-modal integrated representation space. Finally, by further integrating the N first cross-modal features and the M second cross-modal features, cross-modal joint features are generated. This feature representation not only contains the rich information of multi-modal data, but also ensures the diversity and depth of features through a hierarchical fusion strategy, providing a more accurate and comprehensive feature basis for cross-modal tasks, significantly improving the cross-modal understanding and matching ability of the model, and thus improving the processing efficiency of multi-modal data.
[0117] Through the embodiments of the present application, feature fusion is performed on multiple first features belonging to the same modality type in the key data features to obtain N first cross-modal features, where N is a positive integer; feature fusion is performed on multiple second features belonging to different modality types in the key data features to obtain M second cross-modal features, where M is a positive integer; based on the N first cross-modal features and the M second cross-modal features, cross-modal joint features are obtained. Thus, the technical purpose of generating cross-modal joint features is achieved, and the technical effect of improving the processing efficiency of multi-modal data is realized.
[0118] As an optional solution, performing feature fusion on multiple first features belonging to the same modality type in the key data features to obtain N first cross-modal features includes:
[0119] Execute the following steps until N first cross-modal features are obtained:
[0120] S6-1, determine a first current feature and a second current feature from the multiple first features;
[0121] S6-2, calculate the dot product of the query matrix and the key matrix to obtain a dot product result, where the query matrix includes the first current feature and the second current feature, and the key matrix is used to calculate similarity with the query matrix;
[0122] S6-3, convert the dot product result into an attention weight, where the attention weight is used to represent the feature similarity between the first current feature and the second current feature;
[0123] S6-4, perform weighted summation on the attention weight and the value matrix to obtain a weighted summation result, where the value matrix includes the specific information of the first current feature and the second current feature;
[0124] S6-5 Utilize the weighted summation result to fuse the first current feature and the second current feature to obtain the first cross-modal feature.
[0125] In an alternative embodiment, the query matrix can be, but is not limited to, a matrix used to represent what needs to be focused on currently. The key matrix can be, but is not limited to, a matrix used to calculate similarity with the query matrix to assist in determining the allocation of attention. The value matrix can be, but is not limited to, a matrix that contains specific information of features and can be, but is not limited to, used for weighted summation according to the attention weights.
[0126] In an alternative embodiment, the dot product can be, but is not limited to, understood as a process of multiplying the corresponding elements of the rows of the first matrix and the columns of the second matrix, and then adding the results to obtain a new matrix or a scalar value.
[0127] In an alternative embodiment, the attention weight can be, but is not limited to, understood as a weight representing the relative importance between different features transformed from the dot product result, and can be, but is not limited to, used to represent the feature similarity between the first current feature and the second current feature.
[0128] In an alternative embodiment, determining the first current feature and the second current feature from the multiple first features can be, but is not limited to, understood as that in the input stage of the self-attention mechanism, the model needs to select two features from the multiple first features of the same modality: the first current feature and the second current feature. These two features will be used to calculate the attention weights to determine their relative importance and similarity, which is the basis of feature fusion.
[0129] In an alternative embodiment, calculate the dot product of the query matrix and the key matrix to obtain the dot product result, where the query matrix contains the first current feature and the second current feature, and the key matrix is used to calculate similarity with the query matrix. It can be, but is not limited to, understood as determining the similarity between the first current feature and the second current feature through the dot product calculation of the query matrix and the key matrix.
[0130] In an alternative embodiment, convert the dot product result into an attention weight, where the attention weight is used to represent the feature similarity between the first current feature and the second current feature. It can be, but is not limited to, understood as converting the dot product result into an attention weight to represent the similarity between the first current feature and the second current feature.
[0131] In an alternative embodiment, a weighted sum of the attention weights and the value matrix is calculated to obtain a weighted sum result. The value matrix contains specific information about the first current feature and the second current feature. It can be understood, but not limited to, that the value matrix is weighted by the attention weights to generate a fused feature representation. The value matrix contains specific information about the first and second current features, and the attention weights guide the fusion of these features. The weighted sum result represents the first cross-modal feature after fusion, which synthesizes multiple information of the current features.
[0132] In an alternative embodiment, the weighted sum result is used to fuse the first current feature and the second current feature to obtain the first cross-modal feature. It can be understood, but not limited to, that the weighted sum result is used to fuse the first current feature and the second current feature to generate the first cross-modal feature. The weighted sum result contains comprehensive information of the two features. Through further fusion processing, a more comprehensive and integrated feature representation can be obtained, which helps the model to better understand the structure and content of data in the same modality.
[0133] It should be noted that through the calculation of the query matrix, key matrix, and value matrix of the self-attention mechanism, the model can intelligently determine the relative importance between different features, perform feature fusion based on these importances, reduce information redundancy, and improve the expressiveness of features and the generalization ability of the model. In multi-modal tasks such as image description generation and cross-modal retrieval, it can significantly improve the model's in-depth understanding ability of data in the same modality, lay a solid foundation for generating higher-quality cross-modal joint features, and further improve the accuracy and robustness of multi-modal data processing.
[0134] Through the embodiments of the present application, the first current feature and the second current feature are determined from multiple first features; the dot product of the query matrix and the key matrix is calculated to obtain a dot product result, where the query matrix contains the first current feature and the second current feature, and the key matrix is used to calculate similarity with the query matrix; the dot product result is converted into attention weights, where the attention weights are used to represent the feature similarity between the first current feature and the second current feature; a weighted sum of the attention weights and the value matrix is calculated to obtain a weighted sum result, where the value matrix contains specific information about the first current feature and the second current feature; the weighted sum result is used to fuse the first current feature and the second current feature to obtain the first cross-modal feature. Thus, the technical purpose of obtaining the first cross-modal feature is achieved, and further the technical effect of improving the accuracy and robustness of multi-modal data processing is realized.
[0135] As an alternative solution, feature fusion is performed on multiple second features belonging to different modality types in the key data features to obtain M second cross-modal features, including:
[0136] S7-1, Obtain the similarity matrix among multiple second features, where the similarity matrix is used to reflect the similarity degree among the multiple second features;
[0137] S7-2, Based on the similarity matrix, fuse the features with high similarity among the multiple second features to obtain M second cross-modal features.
[0138] In an optional embodiment, the similarity matrix can be, but is not limited to, a matrix for quantifying and describing the similarity degree between different features.
[0139] It should be noted that in cross-modal feature fusion, the construction of the similarity matrix is based on the dot product between features or other similarity measurement methods, and is used to quantify the similarity degree between the second features. By calculating the similarity scores between features, a matrix that quantifies the relationship between features can be obtained, and this matrix will be used to guide the subsequent feature fusion process to ensure that the model can identify and integrate features with high correlation.
[0140] Furthermore, fuse the second features with high similarity to generate second cross-modal features. The fusion process usually includes methods such as weighted summation, concatenation, or using a graph convolutional neural network to integrate the feature pairs with higher scores to generate M second cross-modal features. These features synthesize the key information between different modalities, improve the accuracy and robustness of cross-modal understanding and matching, and thus improve the processing accuracy of multi-modal data.
[0141] Through the embodiments of the present application, obtain the similarity matrix among multiple second features, where the similarity matrix is used to reflect the similarity degree among the multiple second features; based on the similarity matrix, fuse the features with high similarity among the multiple second features to obtain M second cross-modal features. Thus, the technical purpose of fusing features with high similarity based on the similarity matrix is achieved, and the technical effect of improving the processing accuracy of multi-modal data is further realized.
[0142] As an optional solution, the above multi-modal data processing method further includes:
[0143] S8-1, Obtain the input multi-modal data, including: obtain the multi-modal sample data input to the cross-modal model, where the cross-modal model is used to output the corresponding description text from the image input, or generate the corresponding description image from the text input;
[0144] In the process of using the key data features to output the cross-modal joint features corresponding to the multi-modal data, the method further includes:
[0145] S8-2, Iteratively update the model parameters of the cross-modal model until the model convergence condition is satisfied to obtain the trained cross-modal model.
[0146] In an alternative embodiment, the cross-modal model can be, but is not limited to, a neural network model capable of processing and understanding different-modal data.
[0147] In an alternative embodiment, the modal sample data can be, but is not limited to, a sample data set containing multiple-modal information, and can be, but is not limited to, used for training a cross-modal model to enable it to understand and generate associated information between different modalities.
[0148] In an alternative embodiment, the model parameters can be, but are not limited to, learnable variables such as weights and biases in a neural network, and can be, but is not limited to, used to determine the behavior and performance of the model.
[0149] In an alternative embodiment, the model convergence condition can be, but is not limited to, a criterion or threshold. When the performance metric of the model reaches this condition, it is considered that the model has learned the features of the sample data, and the iterative update can be stopped to avoid overfitting or resource waste.
[0150] It should be noted that by obtaining multi-modal sample data including images and texts, or texts and images, these data are the basis for training the cross-modal model. Then, the model uses the key data features in these data for learning to generate cross-modal joint features. This process involves multiple steps such as feature extraction, feature fusion, and feature optimization. During the model training process, it is necessary to continuously iterate and update the model parameters. Through optimization algorithms such as backpropagation and gradient descent, the performance of the model is gradually improved until the model convergence condition is met. When the performance metric of the model on the training set no longer changes significantly, it is considered that the model has learned the features of the sample data. At this time, the training can be stopped to obtain the trained cross-modal model. Through this strategy of iterative update and convergence check, it can be ensured that the model can be continuously optimized during the training process, improving its cross-modal understanding and generation ability, thereby improving the processing efficiency of multi-modal data.
[0151] Through the embodiments of the present application, the input multi-modal data is obtained, including: obtaining multi-modal sample data input into the cross-modal model, where the cross-modal model is used to output corresponding descriptive texts from image inputs, or generate corresponding descriptive images from text inputs; in the process of using the key data features to output the cross-modal joint features corresponding to the multi-modal data, the method further includes: iteratively updating the model parameters of the cross-modal model until the model convergence condition is met to obtain the trained cross-modal model. Thus, the technical purpose of training the cross-modal model through multi-modal data is achieved, and further the technical effect of improving the processing efficiency of multi-modal data is realized.
[0152] As an alternative solution, for ease of understanding, the multi-modal data processing method is applied in the scenario of cross-modal learning.
[0153] In an alternative embodiment, for the method of jointly understanding text and image based on cross-modal learning, aiming at the problem of information redundancy in the local features of images and texts extracted during the process of feature decoupling and reconstruction, through a feature selection module and sparse regularization: the feature selection module screens important features during the decoupling process, and sparse regularization further weakens redundant information in the feature distribution, making the feature representation more concise and effective; through multi-layer decoupling and multi-modal contrast attention: multi-layer decoupling can capture semantics at different levels, and combining with the cross-modal contrast attention mechanism helps to achieve more detailed feature alignment and fusion, comprehensively optimizing the graphic-text relationship structure; in the cross-modal retrieval task, the use of the feature selection module and sparse regularization can effectively reduce redundant features, thereby optimizing the effect of feature decoupling and reconstruction. The following are the specific implementation steps:
[0154] Step S9-1, Feature extraction
[0155] Use a bidirectional gated recurrent unit, recurrent neural network, and Transformer model to extract the forward and backward features of each word, take their mean as the word feature, and then combine them to obtain the local text feature. Use the object detection model Faster R-CNN to extract the target region features of the image. Map the target region features to the local image features through a fully connected network. Transformer can be, but is not limited to, an architecture based on the attention mechanism.
[0156] Step S9-2, Construct a feature selection module
[0157] Introduce a gating mechanism: use the gating mechanism in the gated recurrent unit to dynamically control the flow of features, screen out the features that contribute to cross-modal retrieval, construct a weight (or gate) for each image and text feature, and these weights can be generated by a separate network layer. Use the sigmoid activation function to generate weights between 0 and 1, indicating the degree of feature retention;
[0158] Apply the feature selection weights: multiply the local features of the image and text by the generated weights, screen out more important features, and filter out irrelevant or unimportant features. Suppose there are image features V and text features T. After using the feature selection weights, the obtained features are V' = V·G(V) and T' = T·G(T), where G represents the weights generated by the feature selection module, thereby further removing redundant image and text features and retaining more discriminative features;
[0159] For further illustration, steps S9-1 and S9-2 are shown in Figure 3. First, text feature extraction is performed. After text feature extraction, local text features are generated via a bidirectional gated recurrent unit or a Transformer model. While text feature extraction is being carried out, image feature extraction is also performed, and then local image features are obtained via the object extraction model Faster R-CNN. Then, the local text features and local image features are input into the feature selection module to generate feature selection weights, and then weighted features are obtained. Next, sparse regularization processing is carried out, and finally, the features are fused to generate joint features.
[0160] Step S9-3, sparse regularization processing
[0161] Apply L1 regularization: Add L1 regularization to the feature selection module to sparsify the weight distribution, prompting some weight values to approach zero, further reducing redundant information. By introducing the L1 regularization term into the loss function, the generated feature selection weight matrix becomes sparse, thereby controlling the number of selected features.
[0162] The goal is to minimize the loss function L = L task + λ∑|G|, where L task is the task loss (such as classification loss or retrieval loss), and λ is the weight coefficient of L1 regularization, which is used to control the intensity of sparse regularization.
[0163] Gradient descent update: In each training iteration, apply the gradient descent algorithm to update the model parameters and feature selection weights simultaneously. L1 regularization will cause the weights of irrelevant features to gradually tend to zero, thus achieving a sparsification effect.
[0164] For further illustration, step S9-3 is as Figure 4 shown. First, the features are input, then via the feature selection module, feature selection weights are generated. Next, feature weighting is performed, and the weighted features are added with L1 regularization to calculate the regularization loss. Then, the objective loss function is optimized, and through gradient descent update, sparse features are obtained. Finally, the refined features are output.
[0165] Step S9-4, reconstructed feature and joint feature generation
[0166] Combine the image and text local features after feature selection and sparse regularization processing to form a more refined feature vector for cross-modal feature fusion.
[0167] Use cross-modal fusion methods such as graph convolutional neural network (GCN) to model the relationship between image and text features and generate joint features with context and semantic information.
[0168] This combined feature contains the key relational structure of images and text, and removes redundant information, which helps to improve the accuracy of cross-modal retrieval;
[0169] For further illustration, step S9-3 is as Figure 5 shown. The refined image and text features are input into the graph convolutional neural network to generate combined features, and then combined features are generated through cross-modal relationship modeling, and finally cross-modal retrieval optimization is performed.
[0170] Step S9-5, Training and Optimization
[0171] Joint training is carried out with the objective loss function combined with sparse regularization, and the feature selection module and feature fusion model are optimized using backpropagation and gradient descent algorithms. During the optimization process, the L1 regularization weight coefficient is adjusted to ensure the balance between the effect of feature sparsification and the performance of the model.
[0172] In an alternative embodiment, the graphic-text relational structure is optimized through multi-layer decoupling and multi-modal contrastive attention. The specific steps include:
[0173] Step S10-1, Multi-layer Decoupling Module
[0174] Multi-layer decoupling is to capture semantic information in multi-modalities (such as images and text) at different levels, enabling the model to perform more refined decoupling at the feature level to enhance the understanding of multi-level semantics. The following are the specific implementation methods:
[0175] 1) Hierarchical feature extraction. In terms of text, bidirectional LSTM, GRU or Transformer encoders are used to perform multi-level encoding on the text to extract the context information of each word, and then layer-by-layer encoding is performed to generate features at different semantic levels. In terms of images, low-level and high-level features of the image can be extracted through a convolutional neural network. The features extracted after each layer of convolution represent feature representations with different detail granularities (such as edges, textures, semantic objects, etc.);
[0176] 2) Selection of decoupling layers. For the features extracted at each layer, different decoupling modules (such as self-attention or residual connections) are applied to decouple independent feature components, ensuring that the semantic features at each level can be separately represented and optimized. The feature decoupling loss is used to further achieve the decoupling of feature levels and avoid information redundancy between features at different levels;
[0177] For further illustration, step S10-1 is as Figure 6As shown, when decoupled output of text features is required, text feature extraction is first performed, and hierarchical text features are obtained through bidirectional LSTM / GTU / Transformer. Then, the text features are optimized through a decoupling module, and finally, the text features are decoupled and output. When decoupled output of image features is required, image feature extraction is first performed, and hierarchical image features are obtained through a convolutional neural network. Then, the image features are optimized through a decoupling module, and finally, the image features are decoupled and output.
[0178] Step S10-2, Cross-modal Contrastive Attention Mechanism
[0179] The cross-modal contrastive attention mechanism aims to achieve finer-grained feature alignment and information fusion between different modalities to comprehensively optimize the relationship structure between text and images. The implementation methods include:
[0180] 1) Intra-modal contrastive attention. In text features and image features, self-attention calculations are respectively performed on the features of each modality, and the similarity between features is calculated through a dot product method to capture the mutual relationship of intra-modal features, calculate the weight of each feature vector within the modality, and obtain an intra-modal enhanced feature representation. The self-attention formula is used:
[0181]
[0182] Q: Query matrix (Query), which is the input feature to be aligned;
[0183] K: Key matrix (Key), used to calculate similarity;
[0184] V: Value matrix (Value), which contains information of the input feature;
[0185] d k : The dimension of the key vector, used to scale to avoid too large or too small gradients;
[0186] Example: Suppose a text sequence is processed, and the feature representations of 4 words are input, and each feature is a 3D vector. Suppose the 4 words become the following matrix X after passing through the embedding layer:
[0187]
[0188] To generate the query, key, and value matrices, it is necessary to define the query weight matrix W Q 、key weight matrix W K and value weight matrix W V , assuming these weight matrices are:
[0189]
[0190] Calculate the query matrix Q:
[0191]
[0192] Calculate the key matrix K:
[0193]
[0194] Calculate the value matrix V:
[0195]
[0196] Calculate the attention scores. Calculate the dot product of Q and K, then scale it, and then use the softmax function to get the attention scores:
[0197]
[0198] Assume d k = 3, divide the dot product result by ≈1.732, where the square root of 3 is approximately equal to 1.732, and then apply softmax to each row:
[0199]
[0200] After the softmax transformation, an attention score matrix is obtained. Multiply the attention score matrix by the value matrix V to get the self-attention output:
[0201]
[0202] Through this process, the self-attention mechanism can fuse information according to the similarity weights, focus relevant features together, and achieve feature alignment and optimization.
[0203] 2) Cross-modal contrastive attention. Use an interactive attention layer to establish corresponding relationships between different modal features. Contrastive attention first calculates the similarity matrix of image features and text features, which reflects the similarity between image regions and text words. Based on the similarity matrix, features with high similarity are aligned and fused to generate cross-modal features. In the implementation, the method of "similarity assignment" is used to strengthen high-similarity regions and weaken low-similarity regions, and a cross-modal contrastive loss (such as contrastive learning loss) is introduced to ensure that the aligned image and text features are more consistent in similarity.
[0204] For further illustration, the above step S10-2 is as Figure 7As shown, self-attention is calculated for text features and image features simultaneously to obtain intra-text contrastive attention and intra-image contrastive attention respectively. Then, enhanced text features are obtained from the intra-text contrastive attention, and enhanced image features are obtained from the intra-image contrastive attention. Next, interaction attention is obtained by calculating a similarity matrix based on the enhanced text features, and interaction attention is also obtained based on the enhanced image features. Then, aligned image and text features are obtained according to the interaction attention. Finally, cross-modal feature output is performed.
[0205] Step S10-3, Fusion and Feature Alignment
[0206] Fuse the cross-modal aligned features at each level to further optimize the relationship structure between text and image. Aggregate the aligned features at each level (such as weighted summation, concatenation) to integrate semantic information at different levels, and use residual connections or skip connections for each layer of features to retain the original low-level and high-level semantic details.
[0207] Use a graph convolutional neural network (GCN) to structurally reconstruct the aligned features to better capture the text-image relationship. The reconstructed features are processed through non-linear activation to improve the expressive ability of the features.
[0208] For further illustration, the above step S10-3 is as Figure 8 shown. Perform weighted summation / concatenation based on the aligned text features and the aligned image features to obtain the fused features. Then, reconstruct them through a graph convolutional network to obtain the structured feature output. Then, perform feature activation processing. Finally, obtain the final joint features.
[0209] Step S10-4, Training and Optimization
[0210] Comprehensively use multiple loss functions such as cross-modal contrastive loss, feature alignment loss, and reconstruction loss, so that the model can optimize the decoupling and alignment effects simultaneously during training. During backpropagation, update the weights in the multi-layer decoupling and contrast attention modules layer by layer to ensure that the decoupling and contrast attention at each layer can be fully trained, thereby achieving high-quality text-image feature alignment at each level.
[0211] Through the above embodiments, multi-layer decoupling combined with multi-modal contrastive attention can capture fine-grained semantic information at different levels, optimize the relationship structure of the text-image modality, and make the accuracy of cross-modal retrieval results higher.
[0212] Through the embodiments of the present application, first, through the combination of the feature selection module and sparse regularization, the solution can effectively screen out the most representative features during the decoupling process. At the same time, through the sparse regularization mechanism, redundant information is further weakened, making the feature representation more concise and efficient. The feature selection module can automatically identify and retain important features related to the task, thereby reducing the interference of irrelevant or redundant features and ensuring that the model focuses on the features that contribute the most to the task. Sparse regularization introduces an L1 regularization term into the loss function, making the weights of unimportant features gradually tend to zero, further compressing the feature space, reducing the computational complexity, and improving the interpretability of the model. In this way, the feature representation becomes more concise and sparse, avoiding the accumulation of information redundancy, promoting more accurate decision-making and reasoning of the model in the high-dimensional feature space, enhancing the generalization ability of the model, reducing the risk of overfitting, and thus improving the effect and accuracy of cross-modal retrieval.
[0213] Furthermore, through the multi-layer decoupling and multi-modal contrast attention solution, different levels of semantic information of the image and text modalities can be effectively captured and fused, thereby improving the accuracy and robustness of cross-modal retrieval. Multi-layer decoupling captures the multi-level and fine-grained semantics in the features, enabling the model to identify the hierarchical relationships within the modality and avoiding information loss and redundancy caused by simple feature fusion. Combined with the cross-modal contrast attention mechanism, the contrast learning enhances the model's ability to identify the similarities and differences between modalities, enabling the model to more precisely align similar content in images and texts, while distinguishing irrelevant features in different modalities and effectively reducing noise. This solution makes full use of the relational structure information of each modality to form a more expressive joint feature representation in the semantic space, which helps to improve the fineness of feature alignment and the understanding of complex image-text relationships, thereby enhancing the accuracy and generalization of the cross-modal retrieval system.
[0214] From the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0215] In this embodiment, a processing device for multi-modal data is further provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0216] Figure 9 is a structural block diagram of a processing device for multi-modal data according to an embodiment of the present application. As Figure 9 shown, the device includes:
[0217] A first acquisition unit 902, configured to acquire input multi-modal data, where the multi-modal data includes at least two types of modal data;
[0218] A second acquisition unit 904, configured to acquire a plurality of initial data features corresponding to each type of modal data among at least two types of modal data;
[0219] A screening unit 906, configured to perform feature screening on the plurality of initial data features to obtain at least one key data feature, where the key data feature is a feature in the initial data features that can characterize the preset information of the multi-modal data;
[0220] An output unit 908, configured to use the key data feature to output a cross-modal joint feature corresponding to the multi-modal data.
[0221] As an optional solution, the screening unit 906 includes: a construction module, configured to construct a corresponding plurality of key weights for each of the plurality of initial data features, where the key weights are used to represent the retention degree of the data features; an integration module, configured to perform integration processing on each initial data feature and the corresponding key weight to obtain an integration result corresponding to each initial data feature; and a screening module, configured to perform feature screening on the plurality of initial data features by using the integration result to obtain at least one key data feature.
[0222] For specific embodiments, reference can be made to the examples shown in the above multi-modal data processing method, and details will not be repeated here.
[0223] As an alternative solution, a construction module includes: an input sub-module for inputting a plurality of initial data features into a gated recurrent unit; a capture sub-module for capturing long-term dependencies in a feature sequence through the gating mechanism of the gated recurrent unit; a first acquisition sub-module for using the long-term dependencies to obtain the hidden states output by the gated recurrent unit at each time step, where the hidden state contains all historical information up to the current time step; and a mapping sub-module for mapping the hidden state to a weight vector, where each element in the weight vector corresponds to each key weight among a plurality of key weights.
[0224] For specific embodiments, reference may be made to the examples shown in the above-mentioned multi-modal data processing method, and details are not described herein again in this example.
[0225] As an alternative solution, a construction module includes: an adjustment sub-module for adjusting the sparse regularization strength of a plurality of key weights by using the weight coefficient of L1 regularization, where the weight coefficient is iteratively updated by applying a gradient descent algorithm.
[0226] For specific embodiments, reference may be made to the examples shown in the above-mentioned multi-modal data processing method, and details are not described herein again in this example.
[0227] As an alternative solution, an output unit 908 includes: an extraction module for performing multi-level feature extraction on key data features to obtain multi-level features; a decoupling module for decoupling each layer of features in the multi-level features using different decoupling modules to obtain independent semantic features at each level; and an output module for outputting cross-modal joint features by using the independent semantic features.
[0228] For specific embodiments, reference may be made to the examples shown in the above-mentioned multi-modal data processing method, and details are not described herein again in this example.
[0229] As an alternative solution, the extraction module includes: a first extraction sub-module for performing multi-level encoding on text key features in the key data features to extract the context information of each word; and an encoding sub-module for performing layer-by-layer encoding on the text key features in combination with the context information to obtain text features at different semantic levels.
[0230] For specific embodiments, reference may be made to the examples shown in the above-mentioned multi-modal data processing method, and details are not described herein again in this example.
[0231] As an alternative solution, the extraction module includes: a second extraction sub-module for extracting low-level features and high-level features obtained after multi-layer convolution from image key features in the key data features, where the features extracted after multi-layer convolution represent feature representations with different levels of detail granularity.
[0232] For specific embodiments, reference may be made to the examples shown in the above-mentioned method for processing multimodal data, which will not be elaborated herein.
[0233] As an alternative solution, the output unit 908 includes: a first fusion module for performing feature fusion on multiple first features belonging to the same modal type among the key data features to obtain N first cross-modal features, where N is a positive integer; a second fusion module for performing feature fusion on multiple second features belonging to different modal types among the key data features to obtain M second cross-modal features, where M is a positive integer; and a first acquisition module for obtaining cross-modal joint features based on the N first cross-modal features and the M second cross-modal features.
[0234] For specific embodiments, reference may be made to the examples shown in the above-mentioned method for processing multimodal data, which will not be elaborated herein.
[0235] As an alternative solution, the first fusion module includes: a determination sub-module for determining a first current feature and a second current feature from the multiple first features; a calculation sub-module for calculating the dot product of a query matrix and a key matrix to obtain a dot product result, where the query matrix includes the first current feature and the second current feature, and the key matrix is used to calculate similarity with the query matrix; a conversion sub-module for converting the dot product result into an attention weight, where the attention weight is used to represent the feature similarity between the first current feature and the second current feature; a summation sub-module for performing weighted summation on the attention weight and a value matrix to obtain a weighted summation result, where the value matrix includes the specific information of the first current feature and the second current feature; and a first fusion sub-module for using the weighted summation result to fuse the first current feature and the second current feature to obtain a first cross-modal feature.
[0236] For specific embodiments, reference may be made to the examples shown in the above-mentioned method for processing multimodal data, which will not be elaborated herein.
[0237] As an alternative solution, the second fusion module includes: a second acquisition sub-module for obtaining a similarity matrix between the multiple second features, where the similarity matrix is used to reflect the similarity between the multiple second features; and a second fusion sub-module for fusing the features with high similarity among the multiple second features based on the similarity matrix to obtain M second cross-modal features.
[0238] For specific embodiments, reference may be made to the examples shown in the above-mentioned method for processing multimodal data, which will not be elaborated herein.
[0239] As an alternative solution, the first acquisition unit 902 includes: a second acquisition module, configured to acquire multimodal sample data input into a cross-modal model, where the cross-modal model is used to output corresponding descriptive text from an image input, or generate a corresponding descriptive image from a text input; an iteration module, configured to iteratively update the model parameters of the cross-modal model until a model convergence condition is satisfied, to obtain a trained cross-modal model.
[0240] For specific embodiments, reference may be made to the examples shown in the above-mentioned method for processing multimodal data, and details are not described herein again in this example.
[0241] It should be noted that the above-mentioned various virtual devices (modules, units, sub-modules, sub-units, components, etc.) can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited thereto: the above-mentioned virtual devices are all located in the same processor; or, the above-mentioned various virtual devices are respectively located in different processors in any combination form.
[0242] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0243] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs for short), random access memories (RAMs for short), mobile hard disks, magnetic disks, or optical discs, and other various media that can store computer programs.
[0244] An embodiment of the present application further provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0245] In an exemplary embodiment, the above-mentioned electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above-mentioned processor, and the input / output device is connected to the above-mentioned processor.
[0246] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary embodiments, and details are not described herein again in this embodiment.
[0247] An embodiment of the present application further provides a computer program product, including a computer program, where the computer program is configured to execute the steps in any one of the above method embodiments when executed by a processor.
[0248] For the specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary embodiments, and they will not be elaborated herein again.
[0249] Obviously, those skilled in the art should understand that the above-mentioned virtual devices or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present application is not limited to any specific combination of hardware and software.
[0250] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for processing multimodal data, characterized in that: include: Acquire input multimodal data, wherein the multimodal data includes at least two text modality type data and image modality type data; Acquire a plurality of initial data features corresponding to each modality type data in the multimodal data; Constructing a corresponding plurality of key weights for each of the plurality of initial data features, wherein the key weights are used to indicate a degree of retention of the data features; Using L1 regularized weight coefficients, adjusting the sparse regularization strength of the plurality of key weights, wherein the weight coefficients are obtained by iteratively updating using a gradient descent algorithm; Performing feature screening on the multiple initial data features to obtain at least one key data feature, wherein the key data feature is a feature of the initial data features that can characterize preset information of the multimodal data; Using the key data features, outputting cross-modal joint features corresponding to the multimodal data; The utilizing the key data features to output cross-modal joint features corresponding to the multimodal data includes: Performing multi-layer feature extraction on the key data features to obtain multi-layer features; Decoupling each layer of features in the multi-layer features using different decoupling modules to obtain independent semantic features at each level, wherein the decoupling modules are used to separate different components of features of different modalities; Decoupling the feature hierarchy using feature decoupling loss, wherein decoupling the feature hierarchy using feature decoupling loss is used to avoid information redundancy between features at different levels; Outputting the cross-modal joint feature by using the independent semantic feature; The multi-layered feature extraction of the key data features to obtain multi-layer features includes: For the key features of the text in the key data features, a bidirectional LSTM encoder is used to perform multi-level encoding on the text to extract the context information of each word, and then the features of different semantic levels are generated layer by layer to extract the context information of each word; Combining the context information, encoding the key features of the text layer by layer to obtain text features at different semantic levels; Extracting low-level features and high-level features of the image after multi-layer convolution from the key features of the image in the key data features, wherein the features extracted after the multi-layer convolution respectively represent feature representations of different detail granularities; The step of constructing a corresponding plurality of key weights for each of the plurality of initial data features comprises: Inputting the plurality of initial data features into a gated recurrent unit; Capturing long-term dependencies in feature sequences through the gating mechanism of the gated recurrent unit; Using the long-term dependency, obtaining the hidden state output by the gated recurrent unit at each time step, wherein the hidden state includes all historical information up to the current time step; The hidden state is mapped into a weight vector, wherein each element in the weight vector corresponds to a respective key weight of the plurality of key weights.
2. The method according to claim 1, characterized in that: Before screening the multiple initial data features to obtain at least one key data feature, the method further includes: Integrate the initial data features with the corresponding key weights to obtain integration results corresponding to the initial data features; The performing feature screening on the multiple initial data features to obtain at least one key data feature includes: utilizing the integration result to perform feature screening on the multiple initial data features to obtain the at least one key data feature.
3. The method according to claim 1, characterized in that The utilizing the key data features to output cross-modal joint features corresponding to the multimodal data includes: Performing feature fusion on multiple first features of the same modality type in the key data features to obtain N modality internal features, where N is a positive integer; Performing feature fusion on multiple second features belonging to different modal types in the key data features to obtain M modal comprehensive features, where M is a positive integer; Based on the N modality internal features and the M modality comprehensive features, the cross-modality joint features are obtained.
4. The method according to claim 3, characterized in that: The step of fusing multiple first features of the same modality type in the key data features to obtain N modality internal features includes: Perform the following steps until the N modal internal features are obtained: Determining a first current feature and a second current feature from the plurality of first features; Calculating a dot product of a query matrix and a key matrix to obtain a dot product result, wherein the query matrix includes the first current feature and the second current feature, and the key matrix is used to calculate similarity with the query matrix; Converting the dot product result into an attention weight, wherein the attention weight is used to represent the feature similarity between the first current feature and the second current feature; Performing weighted summation on the attention weight and the value matrix to obtain a weighted summation result, wherein the value matrix includes specific information of the first current feature and the second current feature; The weighted sum result is used to fuse the first current feature and the second current feature to obtain the modal internal feature.
5. The method according to claim 3, characterized in that: The step of fusing multiple second features of different modal types in the key data features to obtain M modal comprehensive features includes: Acquire a similarity matrix between the plurality of second features, wherein the similarity matrix is used to reflect the similarity between the plurality of second features; Based on the similarity matrix, the features with high similarity among the multiple second features are fused to obtain the M modal comprehensive features.
6. The method according to any one of claims 1 to 5, characterized in that The obtaining of input multimodal data includes: obtaining multimodal sample data for input into a cross-modal model, wherein the cross-modal model is used to output corresponding description text from image input, or to generate corresponding description images from text input; In the process of using the key data features to output the cross-modal joint features corresponding to the multimodal data, the method also includes: iteratively updating the model parameters of the cross-modal model until the model convergence conditions are met to obtain the trained cross-modal model.
7. A multimodal data processing device, characterized in that: include: A first acquisition unit, configured to acquire input multimodal data, wherein the multimodal data includes at least two text modality type data and image modality type data; A second acquisition unit, configured to acquire a plurality of initial data features corresponding to each modality type data in the multimodal data; A construction module, used to construct a corresponding plurality of key weights for each of the plurality of initial data features, wherein the key weights are used to indicate a degree of retention of the data features; The construction module includes: an adjustment submodule, which is used to adjust the sparse regularization strength of the multiple key weights using the weight coefficient of L1 regularization, wherein the weight coefficient is obtained by iteratively updating by applying a gradient descent algorithm; The device is further used to: perform feature screening on the multiple initial data features to obtain at least one key data feature, wherein the key data feature is a feature of the initial data features that can characterize the preset information of the multimodal data; An output unit, configured to output a cross-modal joint feature corresponding to the multimodal data using the key data feature; The output unit includes: an extraction module, which is used to perform multi-layer feature extraction on the key data features to obtain multi-layer features; a decoupling module, which is used to decouple the features of each layer in the multi-layer features using different decoupling modules to obtain independent semantic features of each layer, wherein the decoupling module is used to separate different components of features of different modalities; an output module, which is used to output the above-mentioned cross-modal joint features using the independent semantic features; The device is further used to: decouple the feature hierarchy using feature decoupling loss, wherein the decoupling of the feature hierarchy using feature decoupling loss is used to avoid information redundancy between features at different levels; The device is also used for: For the key features of the text in the key data features, a bidirectional LSTM encoder is used to perform multi-level encoding on the text to extract the context information of each word, and then the features of different semantic levels are generated layer by layer to extract the context information of each word; Combining the context information, encoding the key features of the text layer by layer to obtain text features at different semantic levels; Extracting low-level features and high-level features of the image after multi-layer convolution from the key features of the image in the key data features, wherein the features extracted after the multi-layer convolution respectively represent feature representations of different detail granularities; The step of constructing a corresponding plurality of key weights for each of the plurality of initial data features comprises: Inputting the plurality of initial data features into a gated recurrent unit; Capturing long-term dependencies in feature sequences through the gating mechanism of the gated recurrent unit; Using the long-term dependency, obtaining the hidden state output by the gated recurrent unit at each time step, wherein the hidden state includes all historical information up to the current time step; The hidden state is mapped into a weight vector, wherein each element in the weight vector corresponds to a respective key weight of the plurality of key weights.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 6 when executed by a processor.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Multi-modal feature processing method and device, storage medium and electronic equipment
CN116861363A
Sensitive information classification method and device based on multi-modal attention fusion
CN117235605A
Multi-modal information fusion method and device, equipment and storage medium
CN117851967A