Multimodal data interactive information extraction method, device, equipment, product and network training method

By training the interactive information extraction network and utilizing modal coding and random masking techniques, the problem of insufficient collaborative information capture in multimodal representation learning is solved, and more efficient multimodal data collaborative information extraction and task adaptability are achieved.

CN120524448BActive Publication Date: 2025-09-26CHENGDU EVERIMAGING SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511029791.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-09-26
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

Existing multimodal representation learning methods struggle to effectively capture collaborative information, resulting in poor performance in tasks where collaborative interaction is crucial.

Method used

By training the interactive information extraction network, the modal encoder is used to encode the modal data, the modal features are fused to extract redundant, collaborative and independent information, and the total training loss is minimized through random masking and feature enhancement methods, and the mutual information is calculated to optimize the network.

Benefits of technology

It improves the accuracy and robustness of multimodal data in collaborative information extraction and enhances the performance in different task scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120524448B_ABST
    Figure CN120524448B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, equipment, product and network training method for extracting interactive information from multimodal data, which relates to the field of multimodal representation learning and is used to effectively extract collaborative interactive information from multimodal data. The present invention utilizes multimodal data samples to train a network, and utilizes the trained network to extract collaborative interactive information from multimodal data. The network performs data enhancement and encoding on each modal data input, obtains redundant information by fusing each modal feature, obtains collaborative information by fusing each modal feature after random masking, and obtains independent information by performing feature enhancement on each modal feature separately. The network is trained with the goal of minimizing the training loss of the three types of information. The present invention can extract rich collaborative interactive information and reduce the computational complexity of network training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal representation learning, and in particular to a method, apparatus, and device for extracting interactive information from multimodal data, a computer program product, and a network training method for extracting interactive information. Background Art

[0002] Multimodal representation learning integrates data from different modalities (such as text, images, and audio) into a unified latent space and uses contrastive loss to align features from these modalities. In multimodal representation learning, the synergistic interaction between multimodal data not only provides complementary information but also creates unique effects through specific interaction patterns that cannot be achieved with single-modal data.

[0003] Multimodal representation learning involves learning three aspects of multimodal data: redundancy, independence, and collaboration. Redundancy refers to the ability of one modality to independently perform a task due to overlapping shared information. Uniqueness describes the fact that only one modality possesses all the information required to complete the task. Collaboration is the most important yet elusive of the three, as the modalities provide complementary information that must be integrated to achieve the goal. These types of interactions are not static; their dominance depends on the specific target task, which adds a certain level of complexity to multimodal representation learning. For example, one task may rely heavily on redundant information in a specific scenario, while another task may require the collaborative integration of multiple modalities to succeed. Therefore, task-independent multimodal representations must encompass the full spectrum of interactions beyond information redundancy.

[0004] Existing methods may struggle to effectively capture the full spectrum of collaborative information, resulting in poor performance in tasks where collaborative interactions are crucial. For example, CLIP and ALIGN have been applied to vision-language tasks. These models have demonstrated the effectiveness of capturing shared patterns between multimodal data sources by aligning multimodal features, thereby enabling diverse downstream applications. However, current methods often rely on the restrictive multi-view redundancy assumption, which assumes that data from one modality is sufficient to predict downstream tasks and contains the same task-relevant information. This assumption originates from multi-view learning and is limited in real-world multimodal settings because many multimodal tasks involve very little shared information. Summary of the Invention

[0005] The object of the present invention is to provide a method, device, equipment, product and network training method for extracting interactive information from multimodal data to effectively extract collaborative interactive information from multimodal data in order to address all or part of the above-mentioned problems.

[0006] The technical solution adopted in the present invention is as follows:

[0007] A method for extracting interactive information from multimodal data, wherein the interactive information includes redundant information, collaborative information, and independent information; the method comprises:

[0008] Using multimodal data samples to train an interactive information extraction network; using the trained interactive information extraction network to extract at least collaborative information from the multimodal data for a target task; wherein the interactive information extraction network is configured to:

[0009] Perform data enhancement on the input modal data;

[0010] Use the modal encoder to encode each modal data separately to obtain the corresponding modal features;

[0011] Redundant information is obtained by fusing each modal feature, collaborative information is obtained by fusing each modal feature after random masking, and independent information is obtained by enhancing each modal feature separately.

[0012] Furthermore, the method of training the interactive information extraction network using multimodal data samples includes:

[0013] Inputting the multimodal data sample into the interaction information extraction network to obtain corresponding interaction information;

[0014] The mutual information extraction network is trained with the goal of minimizing the total training loss, wherein the total training loss includes the contrast loss within the redundant information in the mutual information, the contrast loss between the redundant information and the independent information, and the contrast loss between the redundant information and the collaborative information.

[0015] Furthermore, the training of the interactive information extraction network with the goal of minimizing the total training loss includes:

[0016] The total training loss is minimized by minimizing the contrast loss between redundant information and collaborative information; the contrast loss between redundant information and system information is obtained by calculating the mutual information between collaborative information and redundant information.

[0017] Furthermore, the method for calculating the mutual information between the collaborative information and the redundant information includes:

[0018] Perform multiple rounds of random masking on each modal feature, and fuse each modal feature after random masking in each round to obtain the collaborative information of that round;

[0019] Calculate the mean mutual information between redundant information and collaborative information in each round.

[0020] Furthermore, the method for calculating the mutual information between the collaborative information and the redundant information includes:

[0021] The mutual information between collaborative information and redundant information is obtained by calculating the lower bound of the mutual information between collaborative information and redundant information.

[0022] Furthermore, data enhancement is performed on the input modal data, including:

[0023] At least two rounds of data enhancement are performed on the input modal data to obtain at least two sets of data-enhanced multimodal data.

[0024] The present invention also provides an interactive information extraction network training method to train an interactive information extraction network for extracting interactive information from multimodal data for a target task; the interactive information includes redundant information, collaborative information, and independent information;

[0025] The interaction information extraction network is configured as follows:

[0026] Perform data enhancement on the input modal data;

[0027] Use the modal encoder to encode each modal data separately to obtain the corresponding modal features;

[0028] Redundant information is obtained by fusing each modal feature, collaborative information is obtained by fusing each modal feature after random masking, and independent information is obtained by enhancing each modal feature separately.

[0029] Training methods include:

[0030] Inputting the multimodal data samples into the interaction information extraction network to obtain corresponding interaction information;

[0031] The mutual information extraction network is trained with the goal of minimizing the total training loss, wherein the total training loss includes the contrast loss within the redundant information in the mutual information, the contrast loss between the redundant information and the independent information, and the contrast loss between the redundant information and the collaborative information.

[0032] The present invention also provides a device for extracting interactive information from multimodal data, wherein the interactive information includes redundant information, collaborative information, and independent information; the device comprises:

[0033] A network training module is used to train the interaction information extraction network using multimodal data samples;

[0034] a network inference module for extracting at least collaborative information from multimodal data for a target task using the trained interactive information extraction network;

[0035] The interaction information extraction network is configured as follows:

[0036] Perform data enhancement on the input modal data;

[0037] Use the modal encoder to encode each modal data separately to obtain the corresponding modal features;

[0038] Redundant information is obtained by fusing each modal feature, collaborative information is obtained by fusing each modal feature after random masking, and independent information is obtained by enhancing each modal feature separately.

[0039] The present invention also provides a multimodal data interaction information extraction device, including a processor and a storage medium, wherein the storage medium stores computer instructions, and when the processor runs the computer instructions, it can execute the above-mentioned multimodal data interaction information extraction method.

[0040] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, can execute the above-mentioned multimodal data interaction information extraction method.

[0041] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0042] The present invention randomly masks a considerable portion of each modal feature during the process of fusing multimodal features. This masking retains only partial feature information, thereby creating a fused representation with different collaborative patterns. Subsequently, the unmasked fused representation is aligned with these masked representations by maximizing mutual information to encode comprehensive collaborative information. The present invention performs a large number of rounds of feature masking, and the infinite masking strategy enables the present invention to capture richer collaborative interaction information by exposing the network to a variety of partial modal information combinations during training. On this basis, the present invention approximates the network loss by calculating the lower bound of the mutual information, solving the problem of difficult estimation of mutual information. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The present invention will now be described by way of example with reference to the accompanying drawings, in which:

[0044] Figure 1 This is a flow chart of the method for extracting interactive information from multimodal data.

[0045] Figure 2 It is the network architecture diagram of the interactive information extraction network.

[0046] Figure 3 This is a graph showing the impact of the number of mask views on performance indicators.

[0047] Figure 4 It is a curve diagram showing the influence of masking rate on performance indicators. DETAILED DESCRIPTION

[0048] All features disclosed in this specification, or all steps in the disclosed methods or processes, except mutually exclusive features and / or steps, can be combined in any manner.

[0049] Unless otherwise stated, any feature disclosed in this specification (including any appended claims and abstract) may be replaced by other equivalent or similar features. In other words, unless otherwise stated, each feature is only an example of a series of equivalent or similar features.

[0050] In order to solve the problem of poor ability to extract collaborative interaction information of multimodal data in the existing technology, the embodiments of the present application propose a multimodal data interaction information extraction method, device, equipment, product and network training method, aiming to effectively extract the collaborative interaction information of multimodal data.

[0051] like Figure 1 As shown, the multimodal data interaction information extraction method provided in the embodiment of the present application includes the following process:

[0052] S1. Use multimodal data samples to train the interactive information extraction network.

[0053] The so-called multimodal data refers to data types of different modalities, such as two or more modal data including text data, image data (including video data), voice data / audio data, table data, etc. Different modal data may be able to complete the target task independently, or they may have to work together to advance the target task. For the latter, it is necessary to extract collaborative information between multimodal data. This collaborative information is obtained by fusing the modal features of each modal data. The collaborative information can achieve an effect that cannot be achieved by any single modal feature. Since there are different ways to fuse multimodal features, and the contributions of different methods to achieving the target task vary greatly, it is necessary to consider how to effectively extract collaborative information for the target task.

[0054] For example, in one optional embodiment of the present application, a multimodal data interaction information extraction method is used to extract human-computer interaction collaborative information in human-computer interaction scenarios such as intelligent voice assistants (such as Siri and Alexa) and virtual reality (VR) / augmented reality (AR) systems. Specifically, this embodiment of the present application provides a multimodal data interaction information extraction method for intelligent human-computer interaction. The multimodal data (including sample data and multimodal data for target tasks) includes voice data indicating user commands, text data, and gesture image data. The voice data is collected using a microphone, the text data is collected using a sensor, and the gesture image data is captured using a camera. For example, a user may speak a voice command to turn on the air conditioner, a temperature sensor may collect the indoor temperature, and the user may gesture toward a specific air conditioner to turn it on. The fused and extracted collaborative information is used to determine user needs. For another example, in another optional embodiment of the present application, the multimodal data interaction information extraction method is used to extract environmental interaction collaborative information in autonomous driving scenarios (such as cars and drones). Specifically, this embodiment of the present application also provides a multimodal data interaction information extraction method for autonomous driving. The multimodal data includes visual image data collected from the environment, LiDAR (Light Laser Detection and Ranging) data, and ultrasonic sensor data. The visual image data is captured using a camera, the ultrasonic sensor data is acquired by scanning the environment using an ultrasonic sensor, and the LiDAR data is acquired by scanning the environment using a LiDAR. The fused and extracted collaborative information is used for path planning and obstacle avoidance. For example, in another optional embodiment, the multimodal data interactive information extraction method is used to extract interactive collaborative information from medical records in health management scenarios such as disease-assisted diagnosis, telemedicine, and patient monitoring. Specifically, the present application also provides a multimodal data interactive information extraction method for health management. The multimodal data includes medical imaging data (such as X-rays, MRIs, and CT scans), text data such as medical reports, and sensor data from clinical examinations (such as heart rate and blood pressure). Medical imaging data is collected using medical testing equipment (such as X-rays and MRIs), text data is retrieved from a medical record management system, and sensor data is acquired using medical testing equipment (such as ultrasound and electrocardiograms). The fused and extracted collaborative information is used to assist in early diagnosis of a patient's health status. Similarly, in similar scenarios, it can also assist in diagnosing the mental health of patients (such as children with ADHD). The corresponding multimodal data includes: voice data from children responding to physicians' questions, text data from children's responses to scale assessments, and sensor data from physiological signals collected through audio and video stimulation of the children.Among them, voice data is collected by a microphone, text data is obtained by a scale evaluation system or a scanned paper quality chart, and sensor data is collected by medical detection equipment (heart rate monitor, brain cap, etc.). The fused and extracted collaborative information is used for early auxiliary judgment of children's mental health status. In addition, the multimodal data interaction information extraction method can also be applied to scenarios such as content recommendation to extract human-computer interaction collaborative information, that is, the embodiment of the present application also provides a multimodal data interaction information extraction method for content recommendation. Multimodal data includes: text data browsed by users, image data, and sensor data of clicking on the interactive interface. The fused and extracted collaborative information is used for user preference analysis.

[0055] Multimodal data samples are historically collected multimodal data. There are multiple groups of multimodal data samples, each corresponding to a different task. For example, in the medical field, multimodal data includes text-based data such as test reports, laboratory results, medical records, and diagnostic reports; image-based data such as radiographs, electrocardiograms, and ultrasound images; and audio-based data such as medical consultation records and voice interaction records. The same applies to multimodal data acquired for target scenarios.

[0056] For interactive information extraction networks, such as Figure 2 As shown, it is configured as:

[0057] 1) Perform data enhancement on the input modal data.

[0058] Data enhancement for each modality can be performed through existing data enhancement methods, such as synonym replacement for text data, flipping and cropping for image data, and noise addition / reduction for audio data.

[0059] In some possible implementations, such as Figure 2 As shown, data enhancement for each modal data includes performing at least two rounds of data enhancement on the input modal data to obtain at least two sets of data-enhanced multimodal data. Each set of multimodal data is subjected to subsequent processing operations.

[0060] Assume that the input multimodal data is X, , where n represents the number of types of modal data, represents the i-th type of modal data, By performing data enhancement on the multimodal data X, the first multimodal data and the second multimodal data .

[0061] It should be noted that the data enhancement performed on each modal data is highly related to the corresponding task. Assume that the multimodal transformation performed on the input multimodal data X is T. For any transformation ,have ,in Represents the mutual information between multimodal data X and data y, and Y represents the corresponding downstream task. Each modality data is included transformation.

[0062] like Figure 2 As shown, after data enhancement, the first multimodal data obtained Expressed as: , the second multimodal data Expressed as: .

[0063] The extraction of interactive data from multimodal data includes the extraction of three types of interactive information: (1) redundant information R, which is shared between each modal data; (2) independent information U, which is specific to each modal data, that is, it only exists in a single modal data; (3) collaborative information S, which is complementary information that only appears when combining the modal data. In this application, the collaborative information extracted by the interactive information extraction network is Indicates that, represents the network parameters, Indicates the network parameters The collaborative information extracted under Represents the potential representation extracted from the input multimodal data X. One of the purposes of this application is to retain as many features related to the target task as possible between the multimodal data when combining the multimodal data, so that , where Z is the collaborative information extracted by the interactive information extraction network mentioned above.

[0064] 2) Use the modal encoder to encode each modal data separately to obtain the corresponding modal features.

[0065] The modal encoder needs to be configured with a corresponding type of modal encoding module for different types of modal data, such as a modal encoding module for text data, a modal encoding module for image data, etc. The modal encoder can use an existing modal encoder, and there is no special limitation on this in the embodiments of the present application.

[0066] After being encoded by the modal encoder, each modal data obtains a corresponding feature vector, which is called modal feature.

[0067] 3) Redundant information is obtained by fusing each modal feature, collaborative information is obtained by fusing each modal feature after random masking, and independent information is obtained by enhancing each modal feature separately.

[0068] In the modal encoder, the modal features of each modal data are obtained respectively, and the interactive information extraction network obtains the interactive information by independently processing or combining the multimodal features.

[0069] Specifically, the interactive information extraction network consists of three branches. The first branch is responsible for enhancing each modal feature separately to obtain independent information from each modal data; the second branch is responsible for fusing each modal feature to obtain redundant information from multimodal data; and the third branch is responsible for fusing each modal feature after random masking to obtain collaborative information from multimodal data. Random masking refers to randomly blocking (for example, setting to zero or other elimination) some feature elements (a predetermined number and random positions) in the modal features.

[0070] For scenarios where collaborative information is needed to complete the target task, after training the aforementioned network, the third branch extracts collaborative information from the multimodal data in the target scenario. Correspondingly, if the target scenario requires independent or redundant information, the first or second branch of the network is used to extract the independent or redundant information from the multimodal data.

[0071] Since the interactive information extraction network has three branches, its training loss can be composed of the losses of the three branches.

[0072] In some optional implementations, the process of training the interaction information extraction network using multimodal data samples includes:

[0073] S11. Input the multimodal data samples into the interactive information extraction network to obtain corresponding interactive information. The interactive information includes redundant information, independent information, and collaborative information. At any training stage, the interactive information extraction network can obtain these three aspects of information from the input multimodal data.

[0074] S12. Train the interactive information extraction network with the goal of minimizing the total training loss. The total training loss includes the contrast loss within the redundant information in the interactive information, the contrast loss between the redundant information and the independent information, and the contrast loss between the redundant information and the collaborative information.

[0075] As an optional implementation, the network loss is characterized by mutual information. That is, the total training loss includes the mutual information within the redundant information in the interactive information, the mutual information between the redundant information and the independent information, and the mutual information between the redundant information and the collaborative information.

[0076] Expressed in formula, the total training loss is expressed as:

[0077] Formula (1): ;

[0078] Where, They represent the total training loss, the contrastive loss within redundant information, the contrastive loss between redundant information and independent information, and the contrastive loss between redundant information and collaborative information, respectively.

[0079] The total training loss is expressed in terms of mutual information:

[0080] Formula (2): ;

[0081] Where, They represent the mutual information within redundant information, the mutual information between redundant information and independent information, and the mutual information between redundant information and collaborative information, respectively.

[0082] The first multimodal data obtained in the data enhancement stage in the previous article and the second multimodal data For example, after the three branches of the interactive information extraction network are processed, the corresponding independent information extracted is and , the extracted redundant information is and , the extracted collaborative information is and , k represents the total number of rounds of random masking. Figure 2 middle, Represents the modal feature group of the j-th round of random masking.

[0083] For example, if Figure 2 As shown in FIG, as an optional embodiment, for the three branches in the interactive information extraction network, the multimodal features encoded by the modal encoder (the first multimodal data and the second multimodal data ), which is connected to the Transformer blocks in three branches. The Transformer blocks extract the modal features transmitted by the three branches in parallel. Among them, the first branch transmits each modal feature to the Transformer block, and the Transformer block generates single modal features respectively. and ; The second branch respectively converts the first multimodal data (Second multimodal data Similarly), the various modal features in the model are connected and passed to the Transformer block, which outputs the first fusion feature. and the second fusion feature ; The third branch respectively processes the first multimodal data (Second multimodal data Similarly, each modal feature of the random mask is connected and passed to the Transformer block, which outputs the first mask fusion feature. Fusion features with the second mask .

[0084] In some optional implementations, the losses in formula (2) are expressed as:

[0085] Formula (3): .

[0086] Formula (3) is equivalent to the total mutual information contained in redundant information, collaborative information and independent information.

[0087] Formula (4): ;

[0088] If the number of data augmentation groups is more than two, then the formula (3) is adjusted to the mean of the mutual information of the corresponding number of groups. Formula (4) is equivalent to the total mutual information contained in the redundant information and independent information.

[0089] Formula (5): .

[0090] Represents the integral operation performed in the process of calculating the mutual information between collaborative information and redundant information. When calculating the mutual information between collaborative information and redundant information using formula (5), each modal feature can be subjected to multiple rounds of random masking, and each modal feature after random masking in each round is fused to obtain the collaborative information of that round, and then the mean of the mutual information between the redundant information and the collaborative information of each round is calculated. That is:

[0091] .

[0092] The total training loss in the form of formula (2) is expressed as:

[0093] Formula (6): .

[0094] It can be found from formula (6) that the impact of the first two losses on the total training loss depends almost only on the input multimodal data, and is relatively less affected by the network parameters. Therefore, the optimization of the mutual information extraction network mainly depends on the third loss. Therefore, in some optional implementations, the total training loss is minimized by minimizing the contrast loss between redundant information and collaborative information. Among them, the contrast loss between redundant information and system information is obtained by calculating the mutual information between collaborative information and redundant information. According to formula (6), it is obtained by maximizing the mutual information between collaborative information and redundant information. to minimize the total training loss.

[0095] In practice, the more rounds of random masking performed on multimodal features, the more effective it is for extracting collaborative information. However, when masking is infinite (ideally, infinite rounds of random masking are performed), the computational cost of formula (5) is very high, which reduces the practical value of the network. To address this problem, in some optional embodiments of the present application, the mutual information between collaborative information and redundant information is approximated by calculating the lower bound of the mutual information between collaborative information and redundant information, that is, the comparison loss between redundant information and collaborative information is approximated.

[0096] The above formula (5) can be decomposed into two terms, namely , the calculation principle of the two items is exactly the same, so one of them (for example, the first ) is replaced with an equivalent formula.

[0097] By applying the principle of inequality to concave functions, we can Expressed as:

[0098] Formula (7):

[0099] ;

[0100] In the above formula (7), represents the collaborative information vector obtained by fusing the modal features of each random mask, Represents the negative sample vector in the same batch, Indicates Integral operation within the range, Representation characteristics Obey the first fusion feature The distribution probability of , is the temperature coefficient.

[0101] Since 1) mask vectors tend to cluster around a central value in vector space, as they all reflect the semantic characteristics of the query in some aspects; 2) the variance across feature dimensions can be interpreted as a representation of semantic differentiation in the surrounding space, which is consistent with the established principles in distributional semantics. Following the Gaussian distribution, it is expressed as ,in Respectively The mean vector and covariance matrix of . Based on this feature, it can be concluded that:

[0102] Formula (8): ;

[0103] According to formula (8), the lower bound of the mutual information between collaborative information and redundant information can be calculated without exhaustively calculating the mutual information under all possible random masks according to theoretical calculation methods, thereby significantly reducing the calculation cost and improving the feasibility of the application scheme.

[0104] S2. Utilize the trained interactive information extraction network to extract at least collaborative information from the multimodal data for the target task.

[0105] After training, the interactive information extraction network can extract one or more of redundant information, independent information, and collaborative information from the input multimodal data. Since the original intention of this application is to extract collaborative interactive information from multimodal data, in some optional implementations, after inputting the multimodal data into the interactive information extraction network, at least the collaborative information is extracted.

[0106] This application also verifies the effectiveness of the proposed method:

[0107] To evaluate the method's ability to extract redundant, independent, and collaborative information, the present invention's embodiments extracted corresponding interactive information based on the Trifeature dataset in a controlled environment. Furthermore, the method's generalization capabilities were validated in real-world scenarios across multiple widely used multimodal benchmark datasets, spanning diverse fields such as healthcare and robotics, enabling a comprehensive assessment of the proposed method's representational capabilities across diverse scenarios.

[0108] 1) Experimental verification on the Trifeature dataset.

[0109] Following the experimental design of the Trifeature dataset in the network model CoMM, the embodiment of the present application conducted a control experiment on a synthetic dataset derived from the Trifeature dataset. The embodiment of the present application evaluated the network's ability to learn uniqueness, redundancy, and synergy through two separate experiments. In terms of uniqueness and redundancy, the embodiment of the present application defined shape as a redundant feature and texture as a unique feature. This task involves two subtasks:

[0110] (1) Identify shared shapes (redundant parts) between two images, and (2) predict the (uniqueness) of textures with respect to the first image (or the second image). In both cases, the baseline of random guessing corresponds to 10%. As for synergy, the embodiment of the present application processes color by defining a mapping relationship M in the training set (for example, blue maps to triangles, stripes map to red) to artificially introduce strong correlations between textures, so that the baseline of random guessing is 50%. The model follows this mapping and is trained on image pairs. The task is to determine whether a given image pair satisfies the mapping Y = 1 (texture (X1), color (X2) ∈ M), thereby evaluating the model's ability to capture collaborative information across modal interactions. The comparative experimental modality configuration is shown in Table 1.

[0111] Table 1 Network parameter configuration table

[0112]

[0113] This approach freezes the pre-trained interaction feature / information extraction network (model) and trains a linear classifier (or regressor, depending on the task) on the learned features. The performance of the linear classifier on downstream tasks serves as an indicator of the quality of the learned interaction information. Table 2 shows the comparative experimental results.

[0114] Table 2 Comparative experimental performance table

[0115]

[0116] As can be seen from Table 2, the Cross model performs best in capturing redundant information, but performs poorly in terms of uniqueness and synergy. The FactorCL model and the "Cross+Self" model have improved in terms of uniqueness, but still perform poorly in terms of synergy. The CoMM model performs well in all three types of interactive information, but there is still considerable room for improvement in capturing synergy information. In contrast, the method proposed in this application performs well in capturing redundancy and synergy, which is 3.8% and 5.6% higher than CoMM, respectively. This is sufficient to demonstrate the effectiveness of the method proposed in this application in extracting information in three aspects (especially synergy information).

[0117] 2) Experimental verification on real-scene datasets.

[0118] The present embodiment further evaluates the performance of the proposed method on several real-world multimodal datasets provided by Multibench. These datasets span different modality combinations and task types, providing a comprehensive benchmark for evaluating the model's ability to learn effective multimodal representations.

[0119] 2.1) Dual-modal experiments on Multibench.

[0120] The embodiments of this application follow the data preprocessing procedures of previous work and conduct experiments using the same modality encoder, modality configuration, and training models based on encoded inputs of different modalities. "Cross", "Cross+Self", FactorCL, and CoMM are used as comparison baselines. As shown in Table 3, the method proposed in this application achieved the best performance on all benchmark datasets. In the binary classification task, this application outperformed the strongest baseline CoMM by 0.35%, 5.3%, 1.0%, and 2.4% on the MIMIC, MOSI, UR-FUNNY, and MUSTARD datasets, respectively. For regression tasks, this application also achieved superior results, leading in MSE compared to the suboptimal model. These experimental results demonstrate that the proposed method is very effective in capturing bimodal interaction information. Furthermore, its consistently strong performance across different datasets highlights the generalization and robustness of the proposed method in real-world bimodal scenarios.

[0121] Table 3 MSE performance test table

[0122]

[0123] 2.2) Trimodal experiments on Multibench.

[0124] This example evaluates the generalization ability of the proposed method in learning multimodal representations beyond bimodality. Specifically, experiments were conducted on two datasets: Vision & Touch (for contact prediction tasks, with visual, force, and proprioceptive modalities) and UR-FUNNY (with visual, textual, and audio modalities). CMC and CoMM were selected as baseline models in the trimodal setting.

[0125] The experimental results are shown in Table 4. In order to more intuitively evaluate the information gain brought by the introduction of the third modality, the embodiment of the present application also shows the results of using CoMM, "Cross" and "Cross+Self" in a dual-modal training scenario. Specifically, the training scenarios include: (1) training these baselines on the image and proprioception modalities of the Vision&Touch dataset, and (2) training these baselines on the image and text modalities of the UR-FUNNY dataset. The experiments show that adding the third modality significantly enhances the performance of CoMM and the present application's method. CoMM as the baseline shows a performance gain of 8.8% and 1.5% on Vision&Touch and UR-FUNNY, respectively. Although the performance gain obtained by the present application from adding the third modality is not much better than the baseline model CoMM, it is still comparable to CoMM on the Vision&Touch dataset. On the UR-FUNNY dataset, the present application's method achieved the best performance (an increase of 0.8% relative to the CoMMe model).

[0126] Table 4 Accuracy comparison table

[0127]

[0128] 2.3) Bimodal experiments on Multimodal IMDb.

[0129] Multimodal IMDb (MM-IMDb) is a real-world multimodal, multi-label dataset for movie genre classification. It presents two major challenges: significant class imbalance, with comedies and dramas dominating the label distribution, and significant semantic differences between the visual (poster) and textual (plot description) modalities. Genre predictions based on a single modality are often unreliable, while a combination of the two modalities can achieve better predictions. This highlights the need to effectively model multimodal interaction information. This application selects both unimodal and multimodal baselines. For unimodal models, SimCLR (image-only) and CLIP (pre-trained on image-text pairs) are selected as baseline models. For multimodal models, baseline models include CLIP, SLIP, and CoMM. The experimental results are shown in Table 5.

[0130] Table 5 Performance comparison table

[0131]

[0132] As shown in Table 5, our method consistently outperforms all baseline models compared to CoMM, further validating the importance of optimizing multimodal representation learning. Our method achieves the best overall performance, improving CoMM by 1.31% in weighted F1 score and 2.14% in macro F1 score. Notably, CLIP achieves a weighted F1 score of 58.9% using its original common weights, outperforming CLIP fine-tuned on MM-IMDb (54.59%). This suggests that learning redundant information is not always beneficial for complex tasks that require alignment of complementary modalities, such as genre prediction. These experimental results demonstrate our robustness and generalization capabilities in handling imbalanced, semantically heterogeneous, and multi-label, multimodal classification tasks.

[0133] 3) Ablation experiment.

[0134] The embodiment of the present application also conducted an ablation experiment to verify the necessity and effectiveness of the interactive information extraction method designed in the above embodiment.

[0135] In this embodiment of the application, we focus on the design of three key components: the total training loss, the optimal number of mask views, and the masking rate.

[0136] 3.1) Total training loss.

[0137] This embodiment of the application conducts an ablation study on the Trifeature dataset to evaluate the effectiveness of different loss combinations in capturing cross-modal interaction information. As shown in Table 6, the complete total training loss (λ1=λ2=λ3=1, λ1, λ2, λ3 are The weight of ) achieves the highest 77.0% in terms of synergy, while maintaining balanced performance in other indicators. The comparison loss between removing redundant information and collaborative information (λ1=0) significantly reduces the accuracy of the collaborative performance. Only using CoMM loss (λ1=0,λ2=1,λ3=1) yielded a 71.4% synergy, while using only (λ1=0,λ2=0,λ3=1) is further reduced to 58.7%, indicating that the CoMM loss without the view diversity brought by the mask is not enough. (λ3=0) reduces the cooperativity to 69.24% despite high redundancy and uniqueness scores, which shows that all three types of losses are important for effectively modeling multimodal interaction information.

[0138] Table 6 Ablation experiment results

[0139]

[0140] 3.2) Number of mask views.

[0141] The number of mask views is the number of rounds of random masking. Each round of random masking will obtain a mask fusion feature, which is a mask view. Figure 3 As shown in the figure, the number of masked views affects the ability of the proposed method to learn robust multimodal representations. On the MOSI dataset, performance gradually improves with increasing the number of views, peaking at 7 masked views before declining slightly. A similar trend is observed on the MIMIC dataset. These results indicate that a moderate number of masked views is beneficial for the model to learn diverse collaborative patterns while also avoiding overfitting to specific mask configurations.

[0142] 3.3) Masking rate.

[0143] Figure 4 The effects of different masking rates on learning are demonstrated. For both experimental datasets, optimal performance was achieved at masking rates of 0.5 or 0.6, demonstrating that controlled masking rates encourage the network to reason about cross-modal collaborative information, thereby enhancing its ability to learn more generalizable and robust representations. However, performance plummets when masking rates exceed this range, indicating that excessively high masking rates can lead to a breakdown in the task of extracting interactive information.

[0144] The above experiments show that the design of random mask features in the method proposed in this application is crucial for extracting interactive information from multimodal data.

[0145] According to the concept of this application, the embodiment of this application also proposes a method for training an interactive information extraction network to train an interactive information extraction network for extracting interactive information from multimodal data for a target task. As mentioned above, interactive information includes redundant information, collaborative information, and independent information.

[0146] The configuration of the interactive information extraction network refers to the previous embodiment, that is, it is configured as follows: performing data enhancement on the input modal data; using the modal encoder to encode each modal data separately to obtain the corresponding modal features; obtaining redundant information by fusing each modal feature, obtaining collaborative information by fusing each modal feature after random masking, and obtaining independent information by performing feature enhancement on each modal feature separately.

[0147] The training method in the embodiment of the present application includes:

[0148] Multimodal data samples are input into the mutual information extraction network to obtain corresponding mutual information. The mutual information extraction network is trained with the goal of minimizing the total training loss. The total training loss includes the contrastive loss within the redundant information in the mutual information, the contrastive loss between the redundant information and the independent information, and the contrastive loss between the redundant information and the collaborative information.

[0149] According to the concept of the present application, an embodiment of the present application further provides a multimodal data interaction information extraction device, which includes:

[0150] The network training module is used to train the interactive information extraction network using multimodal data samples. The configuration of the interactive information extraction network is as described in the previous embodiment.

[0151] The network reasoning module is used to extract at least collaborative information from the multimodal data for the target task using the trained interactive information extraction network.

[0152] In addition, an embodiment of the present application also provides a multimodal data interaction information extraction device, including a processor and a storage medium, wherein the storage medium stores computer instructions, and when the processor runs the computer instructions, it can execute the multimodal data interaction information extraction method of the above embodiment.

[0153] In addition, an embodiment of the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the multimodal data interaction information extraction method of the above embodiment can be executed.

[0154] The present invention is not limited to the aforementioned specific embodiments, but extends to any new features or any new combination disclosed in this specification, as well as any new method or process steps or any new combination disclosed.

Claims

1. A method for extracting interactive information from multimodal data, wherein the multimodal data includes two or more modal data selected from the group consisting of text data, image data, voice data / audio data, and tabular data; The interactive information includes redundant information, collaborative information and independent information; it is characterized in that, Methods include: Training an interactive information extraction network using multimodal data samples includes: inputting the multimodal data samples into the interactive information extraction network to obtain corresponding interactive information; training the interactive information extraction network with the goal of minimizing total training loss, wherein the total training loss includes contrast loss within redundant information in the interactive information, contrast loss between redundant information and independent information, and contrast loss between redundant information and collaborative information; extracting at least collaborative information from multimodal data for a target task using the trained interactive information extraction network; wherein the interactive information extraction network is configured to: Perform data enhancement on the input modal data; Use the modal encoder to encode each modal data separately to obtain the corresponding modal features; Redundant information is obtained by fusing each modal feature, collaborative information is obtained by fusing each modal feature after random masking, and independent information is obtained by enhancing each modal feature separately.

2. The method for extracting interactive information from multimodal data according to claim 1, wherein: The training of the interactive information extraction network with the goal of minimizing the total training loss includes: The total training loss is minimized by minimizing the contrast loss between redundant information and collaborative information; the contrast loss between redundant information and system information is obtained by calculating the mutual information between collaborative information and redundant information.

3. The method for extracting interactive information from multimodal data according to claim 2, wherein: Methods for calculating the mutual information between collaborative information and redundant information include: Perform multiple rounds of random masking on each modal feature, and fuse each modal feature after random masking in each round to obtain the collaborative information of that round; Calculate the mean mutual information between redundant information and collaborative information in each round.

4. The method for extracting interactive information from multimodal data according to claim 2, wherein: Methods for calculating the mutual information between collaborative information and redundant information include: The mutual information between collaborative information and redundant information is obtained by calculating the lower bound of the mutual information between collaborative information and redundant information.

5. The method for extracting interactive information from multimodal data according to claim 1, wherein: Perform data enhancement on the input modal data, including: At least two rounds of data enhancement are performed on the input modal data to obtain at least two sets of data-enhanced multimodal data.

6. A method for training an interactive information extraction network for extracting interactive information from multimodal data for a target task; the multimodal data includes two or more modal data selected from the group consisting of text data, image data, speech data / audio data, and tabular data; the interactive information includes redundant information, collaborative information, and independent information; characterized in that: The interaction information extraction network is configured as follows: Perform data enhancement on the input modal data; Use the modal encoder to encode each modal data separately to obtain the corresponding modal features; Redundant information is obtained by fusing each modal feature, collaborative information is obtained by fusing each modal feature after random masking, and independent information is obtained by enhancing each modal feature separately. Training methods include: Inputting the multimodal data samples into the interaction information extraction network to obtain corresponding interaction information; The mutual information extraction network is trained with the goal of minimizing the total training loss, wherein the total training loss includes the contrast loss within the redundant information in the mutual information, the contrast loss between the redundant information and the independent information, and the contrast loss between the redundant information and the collaborative information.

7. A device for extracting interactive information from multimodal data, wherein the multimodal data includes two or more modal data selected from the group consisting of text data, image data, voice data / audio data, and table data; the interactive information includes redundant information, collaborative information, and independent information; and wherein: The device includes: a network training module for training an interactive information extraction network using multimodal data samples, comprising: inputting the multimodal data samples into the interactive information extraction network to obtain corresponding interactive information; and training the interactive information extraction network with the goal of minimizing a total training loss, wherein the total training loss includes a contrast loss within redundant information in the interactive information, a contrast loss between redundant information and independent information, and a contrast loss between redundant information and collaborative information; a network inference module for extracting at least collaborative information from multimodal data for a target task using the trained interactive information extraction network; The interaction information extraction network is configured as follows: Perform data enhancement on the input modal data; Use the modal encoder to encode each modal data separately to obtain the corresponding modal features; Redundant information is obtained by fusing each modal feature, collaborative information is obtained by fusing each modal feature after random masking, and independent information is obtained by enhancing each modal feature separately.

8. A multimodal data interactive information extraction device, comprising a processor and a storage medium, characterized in that: The storage medium stores computer instructions, and when the processor runs the computer instructions, it executes the multimodal data interaction information extraction method according to any one of claims 1 to 5.

9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the multimodal data interactive information extraction method according to any one of claims 1 to 5 is executed.