Multimodal representation decoupling method, multimodal data reasoning method, and corresponding apparatus
Patent Information
- Application Number
- CN202610991846.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-07-06
AI Technical Summary
固定拼接后接线性探针的方法会对所有样本应用同一个全局重组规则,因此无法反映各单模态数据特定的信息组成
本申请所提供的多模态表示解耦方案,基于PID的多模态自监督框架,比较每个单模态数据单独能够学习到的信息与所有模态数据联合能够学习到的信息。对于每个单模态数据,学习一个独有因子,用于保留单模态数据特定证据,同时避免共享信号。对于冗余信息,提取所有单模态数据一致支持的信号,并进一步使不同冗余分支彼此保持一致。对于协同信息,从完整的多模态数据中学习,同时移除任何单模态数据已经能够解释的信息。通过这种方式,将“仅联合可得”的协同信息转化为一个显式且可复用的因子,而不是将其隐藏在融合表示(即多模态数据的整体信息)内部,以为下游推理任务的精细化执行提供细分的、显式的、非冗余的三种信息论成分,以供精准推理。此外,对于下游推理任务,本申请还引入了一种无参数的实例级策略,该策略根据相对条件信息增益的代理指标,自适应地重组解耦后的各类信息,使最终的表示向量能够反映特定实例的三种信息论成分的组成,能够持续提升多模态表示质量,尤其在标签依赖跨模态协同的推理场景中,能够取得显著增益。
Smart Images

Figure CN122529103B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer recognition technology, and in particular to a multimodal representation decoupling method and apparatus, as well as a multimodal data reasoning method and apparatus. Background Technology
[0002] In inference tasks (such as classification and prediction) involving multimodal data (each type of data being a single modality) comprised of text, images, audio, time series, and tabular data, it is necessary to extract multimodal representation information to accurately execute the predetermined task in downstream inference models (such as classifier models and prediction models). The accuracy of multimodal representation information extraction directly determines the accuracy and reliability of the downstream inference model in completing the inference task.
[0003] According to the Partial Information Decomposition (PID) theory, the task-related information carried by multimodal data can be decomposed into three information theory components: redundant information, unique information, and collaborative information. Unique information represents evidence that can only be obtained from individual modal data; for example, speech tone is unique audio evidence. Redundant information represents evidence shared by individual modal data and that can be recovered from more than one modality; for example, object categories appearing simultaneously in images and text are often redundant information. Collaborative information represents evidence that can only be obtained when some modal data (i.e., not all modal data) from more than one modality are jointly interpreted (i.e., simultaneously involved in task execution); for example, irony may only manifest when there is a mismatch between words (text) and facial expressions (images). For different reasoning tasks, accurately identifying (decoupling) and preserving the multimodal representations of multimodal data can effectively improve the execution accuracy of the reasoning task.
[0004] However, most known multimodal representation decoupling methods only consider the unique information of single-modal data and the overall information of multimodal data, without explicitly distinguishing between redundant and collaborative information. This results in the extracted multimodal representations being less accurate when facing later inference, affecting the accuracy of the inference task.
[0005] Furthermore, beyond accurate coupling of multimodal representations, the downstream method of recombining these representations also impacts the accuracy of downstream inference tasks. Empirical PID estimations on real multimodal benchmarks show that different datasets may be dominated by different components, such as uniqueness in MIMIC, synergy in MUSTARD, and strong "jointly available only" information in VQA / CLEVR. Methods that use fixed-joint wired probes apply the same global recombination rule to all samples, thus failing to reflect the specific information composition of each single-modal data. From an information theory perspective, the effective contribution of a component can be considered as incremental evidence it provides beyond other components. Therefore, decoupling alone is insufficient; downstream representations also require a method capable of adaptively recombining these separated information types to improve inference accuracy. Summary of the Invention
[0006] The purpose of this invention is to provide a multimodal representation decoupling method, a multimodal data reasoning method, and a corresponding apparatus to address all or part of the problems mentioned above, thereby indirectly or directly improving the reasoning accuracy of multimodal data reasoning tasks.
[0007] The technical solution adopted in this invention is as follows: In a first aspect, this application provides a multimodal representation decoupling method for collaborative perception tasks; the multimodal representation includes redundant information, unique information, and collaborative information of multimodal data; the unique information represents evidence that can only be obtained from each individual modal data; the redundant information represents evidence shared by each individual modal data and that can be recovered from more than one modal data; the collaborative information represents evidence that can only be obtained when partial modal data from more than one modal data are jointly interpreted; the multimodal data includes at least two of the following: text, image, audio, time series, or tabular data; Decoupling methods include: Acquire multimodal data samples, and construct a first enhanced view and a second enhanced view for the multimodal data samples, which preserve labels. The first augmented view and the second augmented view are encoded using an encoder to obtain the unimodal features of each unimodal data sample, the first multimodal features of the first augmented view, and the second multimodal features of the second augmented view; Based on the single-modal features, the first multimodal features, and the second multimodal features, the projection head infers the unique information, redundant information, and cooperative information, respectively. The parameters of the encoder and the projection head are updated based on the total inference loss; wherein the total loss includes the InfoNCE loss and decoupling loss of the unique information, redundant information and cooperative information; The target multimodal data is acquired, and the target-specific information, target redundancy information, and target cooperative information are obtained using the encoder and mapping head.
[0008] Optionally, the InfoNCE loss of the unique information, redundant information, and collaborative information includes the InfoNCE loss of the unique information, the InfoNCE loss of the redundant information, and the InfoNCE loss of the collaborative information; wherein, the InfoNCE loss of the redundant information includes the prediction loss of predicting the redundant information from the unique information, and the consistency loss of the redundant information.
[0009] Optionally, the unique information is lost via InfoNCE. for: ; The predicted loss for: ; The consistency loss for: ; The InfoNCE loss of the collaborative information for: ; in, This represents the InfoNCE loss function for A and B. , ; For the first enhanced view The single-modal features of the i-th single-modal data sample. This represents the i-th single-modal data sample in the first augmented view; For the second enhanced view The single-modal features of the i-th single-modal data sample. This represents the i-th single-modal data sample in the second enhanced view; This is to remove redundant information from the i-th and j-th single-modal data samples in the first enhanced view. This indicates redundant information in the i-th and j-th single-modal data samples in the second enhanced view; This represents the unique information of the i-th single-modal data sample in the first enhanced view; This represents unique information about the i-th single-modal data sample in the second enhanced view; This indicates collaborative information for the first enhanced view; This represents the collaborative information of the second enhanced view; M is the total number of modalities in the multimodal data samples; This represents the number of combinations of any two single-modal data points randomly selected from M types of single-modal data.
[0010] Optionally, the decoupling loss of the unique information, redundant information, and cooperative information includes a first separation loss between the unique information and the redundant information, and a second separation loss between the cooperative information and the multimodal representation.
[0011] Optionally, the first separation loss for: ; Second separation loss for: ; in, For multimodal data samples Total number of modes, This represents the i-th single-modal data sample in the multimodal data sample; The cosine hinge penalty is defined as follows: ,in, They represent The b-th vector in the array, , ; The set threshold; This indicates the inner product operation; This is the characteristic entropy estimation function; All are hyperparameters; For the first enhanced view The single-modal features of the i-th single-modal data sample. This represents the i-th single-modal data sample in the first augmented view; For the second enhanced view The single-modal features of the i-th single-modal data sample. This represents the i-th single-modal data sample in the second enhanced view; This represents the first multimodal feature of the first enhanced view; This represents the second multimodal feature of the second enhanced view; This represents the unique information of the i-th single-modal data sample in the first enhanced view; This represents unique information about the i-th single-modal data sample in the second enhanced view; This represents redundant information for the j-th single-modal data sample in the first enhanced view; This represents redundant information for the j-th single-modal data sample in the second enhanced view; This indicates collaborative information for the first enhanced view; This indicates collaborative information for the second enhanced view.
[0012] Optionally, the total loss includes the InfoNCE loss of the unique information, the InfoNCE loss of the redundant information, the InfoNCE loss of the collaborative information, the first separation loss, and the second separation loss, all weighted and fused.
[0013] In a second aspect, this application also provides a multimodal representation decoupling device for collaborative perception tasks; the device includes a first processor and a first storage medium, the first storage medium storing a first computer program, the first processor running the first computer program to execute the above-described multimodal representation decoupling method.
[0014] In a third aspect, this application also provides a multimodal data reasoning method for collaborative perception tasks; it includes: The above-mentioned multimodal representation decoupling method is used to obtain target-specific information, target redundancy information, and target cooperative information from the target multimodal data; The target-specific information, target redundancy information, and target cooperative information are weighted and concatenated to obtain a concatenated representation vector. The concatenated representation vector is input into the downstream inference model to obtain the inference result.
[0015] Optionally, the method for calculating the splicing weights in the weighted splicing includes: Aggregate target redundancy information from each single modality of data; The target-specific information, the aggregated target redundancy information, and the target cooperative information are used to construct a candidate set. , Represents the candidate set The k-th candidate factor, K is the total number of candidate factors in the candidate set; For the candidate set Each candidate factor in The candidate factor is calculated using the following method. splicing weight : ; in, This represents the temperature parameter of the softmax function; Represents the cosine similarity function. Both represent candidate factor indices. Representing the candidate set respectively The first in One candidate factor.
[0016] In a fourth aspect, this application also provides a multimodal data reasoning apparatus for collaborative perception tasks; the apparatus includes a second processor and a second storage medium, the second storage medium storing a second computer program, the second processor running the second computer program to execute the aforementioned multimodal data reasoning method.
[0017] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: The multimodal representation decoupling scheme provided in this application, based on a PID-based multimodal self-supervised framework, compares the information that can be learned individually from each single-modal data point with the information that can be learned jointly from all modal data points. For each single-modal data point, a unique factor is learned to preserve specific evidence from that single-modal data point while avoiding shared signals. For redundant information, signals consistently supported by all single-modal data points are extracted, and different redundant branches are further aligned with each other. For collaborative information, it is learned from the complete multimodal data while removing any information that can already be explained by the single-modal data points. In this way, the collaborative information, which is "only jointly obtainable," is transformed into an explicit and reusable factor, rather than being hidden within the fused representation (i.e., the overall information of the multimodal data). This provides three distinct, explicit, and non-redundant information-theoretic components for the refined execution of downstream inference tasks, enabling accurate inference. Furthermore, for downstream inference tasks, this application introduces a parameterless instance-level strategy. This strategy adaptively reorganizes various types of decoupled information based on the proxy index of relative conditional information gain, so that the final representation vector can reflect the composition of the three information theory components of a specific instance. This can continuously improve the quality of multimodal representation, especially in label-dependent cross-modal collaborative inference scenarios, where significant gains can be achieved. Attached Figure Description
[0018] The present invention will be described by way of example and with reference to the accompanying drawings, wherein: Figure 1 This is a flowchart illustrating the implementation of the multimodal representation decoupling method in one embodiment.
[0019] Figure 2 This is a data flow graph of a multimodal representation decoupling method in one embodiment.
[0020] Figure 3 This is a diagram showing the test results of three sets of decoupling on the MOSI dataset.
[0021] Figure 4 This is a construction diagram of a multimodal representation decoupling device in one embodiment.
[0022] Figure 5 This is a flowchart illustrating the implementation of a multimodal data reasoning method in one embodiment.
[0023] Figure 6 This is a structural diagram of a multimodal data inference device in one embodiment.
[0024] Figure 7 The results of the sensitivity analysis on decoupling regularization on UR-FUNNY are shown in the figure. The left subplot (a) is the sensitivity test plot for entropy_weight, and the right subplot (b) is the sensitivity test plot for lambda_repel. Detailed Implementation
[0025] All features disclosed in this specification, or steps in all methods or processes disclosed herein, may be combined in any way, except for mutually exclusive features and / or steps.
[0026] Any feature disclosed in this specification (including any appended claims and abstract) may be replaced by other equivalent or similar features, unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is merely one example of a series of equivalent or similar features.
[0027] Multimodal data, which consists of at least two (single) modal data such as text, images, audio, time series, and tabular data, when time-aligned, can generally improve the accuracy of reasoning tasks such as data classification and data prediction compared to data based on single modal data.
[0028] According to the PID theory, the task-related information carried by multimodal data can be decomposed into three information theory components: redundant information, unique information, and cooperative information. Accurately and explicitly decoupling these three information theory components can typically significantly improve the inference quality for downstream inference stages.
[0029] Current understanding of multimodal representation decoupling typically focuses only on the unique information of unimodal data and the overall / shared information of multimodal data, without explicitly decoupling the collaborative information of multimodal data. Explicitly modeling collaborative information is crucial for downstream refined inference tasks because many multimodal data labels are driven by cross-modal interactions. In humor, satire, and multimodal sentiment analysis, decisive cues often lie in the mismatches or complementarities between different modalities, rather than originating from any single modality. Hate expressions require joint interpretation of images and text; while contact-rich manipulation tasks often rely on evidence that cannot be resolved by either vision or touch alone.
[0030] Therefore, in response to the current situation of explicit decoupling of the three information theory components in multimodal representations, this application proposes a multimodal representation decoupling method and apparatus for collaborative perception tasks (i.e., reasoning task scenarios, such as emotion perception tasks—e.g., emotion recognition tasks, and contact-intensive perception tasks—e.g., robot operation tasks), and a multimodal data reasoning method and apparatus for collaborative perception tasks, aiming to indirectly or directly improve the accuracy of downstream reasoning tasks for multimodal data. By explicitly learning unique information, redundant information, and collaborative information from multimodal data, in addition to refining the identification of shared information, the collaborative information, which is "only jointly obtainable," is transformed into an explicit and reusable factor, rather than being hidden inside the fused representation, thereby supporting refined reasoning downstream. Furthermore, for downstream reasoning, a parameterless instance-level routing strategy is introduced, enabling each single-modal data to reflect its own unique information, redundant information, and collaborative information composition while retaining each information theory component. This allows the final concatenated representation vector to adapt to the three parts of information of each sample, completing the adaptive reasoning task and improving the reasoning accuracy gain.
[0031] like Figure 1 , Figure 2 As shown, in one optional implementation, the multimodal representation decoupling method proposed in this application includes the following stages: I. Model Training Phase.
[0032] The encoder and projector are optimized and updated during the model training phase.
[0033] The encoder takes multimodal data as input, internally encodes both single-modal and multimodal data, and outputs the single-modal features of each single-modal data point, as well as the overall multimodal features of the multimodal data. For example... Figure 2 As shown, in one alternative implementation, the encoder uses a CoMM-style fusion backbone network to convert the input multimodal features into token sequences through modality-specific encoders and latent transformers. These sequences are then concatenated and fed into a lightweight Transformer fusion module containing a learnable [CLS] token.
[0034] The projection head consists of three branches: a unique branch, a redundant branch, and a cooperative branch. These branches are connected to the backbone network and are responsible for obtaining the encoded features (single-modal features and multi-modal features) from the encoder. The unique information, redundant information, and cooperative information are then projected from the three branches respectively.
[0035] The encoder and projector are trained using multimodal data samples. After training, the target multimodal data is input into the encoder to obtain the corresponding target single-modal features (multiple types) and multimodal features. Then, the projector is used to obtain the corresponding target-specific information, target redundancy information, and target cooperative information.
[0036] (1) Knowledge preparation.
[0037] Obtain multimodal data samples, which are represented as follows: It contains a total of M types of single-modal data samples (each type of single-modal data has 1 sample). This represents the observation of the i-th unimodal data sample in a multimodal data sample, such as text data, image data, audio data, time series data, or tabular data; i is the index of the unimodal data sample. .
[0038] For example, in emotion-aware tasks such as intelligent customer service or companion robot systems, it's necessary to determine a user's true emotions based on their input text, facial expressions, and tone of voice. Multimodal data includes the user's text data (which can be manually entered or transcribed from speech), facial expression frames during user interaction (i.e., image data), and audio data generated during user interaction. Where audio data is present, text data can be omitted, or it can be retained. For example, when a user says "Great, this is exactly what I wanted," the generated text data (which can be omitted), image data, and audio data are all present. For unimodal data, text data can be encoded using a BERT-based text encoder, image data can be encoded using a ResNet-50 or ViT visual encoder, and audio data can be encoded using a wav2vec 2.0 audio encoder. Subsequently, unique, redundant, and collaborative information is extracted and dynamically recombined layer by layer. From the text alone, the above statements contain obvious positive vocabulary and are easily judged as positive emotions. However, if the user frowns and rolls their eyes in the visual modality, and the tone is flat, with abnormal emphasis and impatience in the audio modality, the true label may be sarcasm or negative emotions. In this sample, the positive vocabulary in the text, the negative facial expressions in the image, and the impatient tone in the audio each constitute unique information for each single modality. In ordinary non-sarcasm samples, when multiple modalities consistently express the same emotion, redundant information is formed. The sarcastic meaning generated by the cross-modal conflict between "positive text" and "negative facial expressions / tones" belongs to collaborative information that cannot be recovered from single modality data and can only be obtained through joint understanding.
[0039] Alternatively, for robot operation and perception tasks, such as grasping, pushing, or assembling tasks with a robotic arm, the robot needs to use multiple sensors to determine whether it will make stable contact with the target object in the next moment, or to predict the position of the end effector. A multimodal data sample can consist of the current RGB image (2D image data), depth map (3D image data), force sensor readings (numerical data), and proprioceptive state (tabular data), where the proprioceptive state includes joint angles, end effector position, and velocity. For RGB images or depth maps, ResNet-18 or ViT visual encoders can be used for encoding; for force sensor readings, MLP (Multilayer Perceptron) force encoders can be used; and for proprioceptive state, MLP proprioceptive encoders can be used. Subsequently, unique information, redundant information, and collaborative information are extracted and dynamically recombined layer by layer. For example, in the task of pushing a cup, the RGB image shows the robotic gripper approaching the cup, the depth map shows the distance is very close, the force signal shows a slight increase, and the proprioceptive state shows the end effector moving towards the cup. Relying solely on vision can only determine spatial proximity, but cannot confirm actual contact; relying solely on force perception may fail to distinguish the object being contacted; and relying solely on proprioception cannot understand the positional relationships of external objects. Only by combining visual geometry, depth distance, force changes, and end-effector motion can the model determine whether contact is imminent and whether it is stable. In this scenario, the object's appearance and geometry are unique to vision, changes in contact pressure are unique to force perception, and end-effector motion is unique to proprioception; multiple sensors collectively reflect "the robotic gripper is approaching the object," creating redundant information; and the stable contact state comes from the combined relationship between vision, force perception, and proprioception, which is collaborative information.
[0040] make This represents potential downstream semantic variables. From the perspective of the information bottleneck in representation learning, the goal of multimodal representation learning is to learn an encoder. ( (representing encoder parameters), Mapping to representation vector , making Able to retain Information relevant to the reasoning task. Ideally, this would include: , This represents the calculation of mutual information. However, in self-supervised learning, During the pre-training phase, it is unobservable. Therefore, the core challenge lies in constructing a system that can retain the... Self-supervised signals of relevant information, while avoiding the collapse of different types of information into a single entangled representation.
[0041] Multimodal representations are not homogeneous; different modal data may provide evidence of overlap (i.e., redundant information), modality-specific evidence (i.e., unique information), or information that can only be obtained after joint observation (i.e., collaborative information). For example, for bimodal data... According to PID theory, the task-related information in the joint observation task is decomposed into: ; in, This represents the semantic variables corresponding to downstream variables. right Decoupled multimodal representation Represents single-modal data The unique information, R represents Shared redundant information; S represents Collaborative information.
[0042] The corresponding consensus equation is: , , The results show that single-modal data can access redundant information and its own unique information, but cannot access collaborative information that is "jointly available only." For the general M-modal case, reachability notation is used. Represents single-modal data Accessible information, and use This represents the reachability information of complete multimodal data.
[0043] In multimodal contrastive learning methods, different modalities are typically treated as different views of the same semantic entity, maximizing modality consistency. Based on this assumption, given another modality, the actual task-related relevance of any given modality is very small, i.e. , The threshold is set, typically a minimum value. Under this assumption, most task-related information is shared across modal data, so cross-modal alignment primarily preserves redundant components R. However, real-world multimodal inference tasks often violate this assumption, as multimodal data labels (i.e., truth values) may depend on unique information specific to each individual modality. It may also rely on collaborative information S that is unavailable from any single-modal data. Therefore, merely aligning the representation learning objectives of multimodal data is insufficient to learn the full range of multimodal table interaction information.
[0044] because This is not available in supervised pre-training; therefore, in this application, another label is used to maintain the view instead. .make This represents a multimodal enhancement function. Let be a set of multimodal enhancement functions. Representing multimodal data First enhanced view, , This represents the i-th single-modal data sample in the first augmented view; Representing multimodal data The second enhanced view, , This represents the i-th unimodal data sample in the second augmented view; the two multimodal augmentation functions are different, and two reference representation views are constructed in this way. The multimodal augmentation functions retain task-related information, i.e. , This enhancement operation enables It can serve as a self-supervised alternative signal for Y. Unlike the single-modal enhancement hypothesis, It is not limited to transforming each single-modal data independently; that is, it can achieve label preservation by transforming the multimodal data as a whole. It can include transformations that can preserve multimodal semantics while changing modality-specific interference factors.
[0045] According to the encoder definition above, we have ,but Parameters omitted Then there is ,Right now Indicates the first enhanced view The first multimodal feature; Indicates the second enhanced view The second multimodal feature. Due to (or )Depend on (or The mapping yields the corresponding Markov chain as follows: This section uses... If we take this as an example, then the data processing boundary is: Therefore, maximizing (or ) will prompt and (or This captures stability information shared between two label-preserving views. Under the label-preserving strategy, this stability information corresponds to... Task-related information.
[0046] (2) Encoding steps.
[0047] Based on the label preservation strategy, multimodal data samples were generated. Construct the first enhanced view Second Enhanced View Then, the encoder is used to enhance the first view. Second Enhanced View Encode it.
[0048] For the first enhanced view The i-th single-mode data After being exposed to the encoder (e.g., by masking other modal data), the corresponding single-modal features are obtained. The encoder provides the first enhanced view. The first multimodal feature of the overall mapping is represented as Similarly, for the second enhanced view... The i-th single-mode data After being exposed to the encoder, the corresponding single-modal features are obtained. The encoder for the second enhanced view The second multimodal feature of the overall mapping is represented as .
[0049] The encoder then uses the aforementioned reference representation view to supervise a unique branch corresponding to each modal data, a set of redundant branches identified by the consistency between reachable paths of single modal data, and a cooperative branch learned from the full modal data while excluding any information already available from single modal data.
[0050] From the perspective of the execution process, the above learning process involves two decomposed steps: and The first step separates shared consensus from modality-specific content within each unimodal data set. The second step removes recoverable information from the unimodal data from the complete multimodal representation, leaving the remaining components to correspond to collaborative information.
[0051] (3) Projection steps.
[0052] Based on unimodal features, first multimodal features, and second multimodal features, the projection head infers unique information, redundant information, and cooperative information, respectively.
[0053] In the encoding step, the encoder for the first enhanced view Second Enhanced View The encoding yielded single-modal features. Single-modal features First multimodal features Second multimodal features Based on these characteristics, the projector head projects the unique information corresponding to each single-modal data, the redundant information corresponding to each single-modal data, and the cooperative information corresponding to the multimodal data from three branches.
[0054] definition Indicates the first enhanced view The unique information of the i-th single-modal data sample; Indicates the second enhanced view The unique information of the i-th single-modal data sample; Indicates the first enhanced view Redundant information of the i-th single-modal data sample. Indicates the first enhanced view Redundant information of the j-th single-modal data sample. Indicates the second enhanced view Redundant information of the i-th single-modal data sample. Indicates the second enhanced view Redundant information of the j-th single-modal data sample; ; Indicates the first enhanced view Collaborative information; Indicates the second enhanced view Collaborative information.
[0055] (4) Loss calculation and model update.
[0056] The encoder and projector head parameters are updated based on the total inference loss. The total training loss includes InfoNCE loss for unique information, redundant information, and cooperative information, as well as decoupling loss.
[0057] As an optional implementation, the InfoNCE loss of unique information, redundant information, and collaborative information includes the InfoNCE loss of unique information, the InfoNCE loss of redundant information, and the InfoNCE loss of collaborative information. Specifically, the InfoNCE loss of redundant information includes the prediction loss from predicting redundant information using unique information, and the consistency loss of redundant information.
[0058] In addition, the decoupling loss of unique information, redundant information and cooperative information includes the first separation loss of unique information and redundant information, and the second separation loss of cooperative information from multimodal representation.
[0059] To more clearly describe the generation process of each loss, based on the two steps decomposed above, the losses in the above optional embodiments are described according to the two stages of "coupling decomposition of unique information and redundant information" and "extraction of cooperative information".
[0060] (a) Decomposition of unique information and redundant information.
[0061] As mentioned earlier, each unimodal feature exposes both redundant information and its unique information, namely... Therefore, simply matching a branch with a single-modal feature cannot distinguish between unique and redundant information in modal data. In one optional implementation, this application performs a joint decomposition of redundant and unique information: redundant information is identified through cross-modal data consistency, while unique information is obtained by removing redundant information from the reachable information of single-modal data.
[0062] First, stable information accessible from the i-th single-modal data is preserved within the unique information, with the ideal goal being... This design utilizes the unique information lost through InfoNCE. Apply constraints: .
[0063] in, This represents the InfoNCE loss function. The above loss preserves... Unique information is then separated by eliminating its overlap with redundant information. To identify redundant information, each redundant branch must be predictable by each single-modal feature; therefore, in one alternative implementation, a prediction loss is constructed that predicts redundant information using unique information. : .
[0064] Since modality-specific information is not consistently supported across all single-modal data, the information reinforced in the pairing of all single-modal data with redundant branches is the shared component, i.e., the redundant information R. Therefore, a consistency loss mechanism for the redundant information is constructed. : ; in, This represents the number of combinations of any two single-modal data points randomly selected from M types of single-modal data.
[0065] Then the InfoNCE loss of redundant information . Used to extract shared knowledge across single-modal data, Then regularize the different redundant branches to make them consistent with each other.
[0066] Finally, in the decomposition of the coupling of unique and redundant information, it will also and Separate. Directly minimize. The reliability is poor because collapsed representations may also have low mutual information, and because... and The separation. Subject to Inspiration ( (representing the calculation of information entropy), this application maintains... While increasing the amount of information, the first separation loss is achieved by constructing a system that separates unique information from redundant information. This reduces the overlap between unique information and redundant information families (i.e., redundant information corresponding to each single-modal data).
[0067] ; in, The cosine hinge penalty is defined as follows: ,in, They represent The b-th vector in the array, ; This is the characteristic entropy estimation function; All are hyperparameters; As mentioned earlier, the threshold value is usually a minimum value.
[0068] The aforementioned first separation loss The interval-based penalty term follows the idea of using angular separation as a tractable decoupling agent: it does not directly estimate mutual information, but rather penalizes over-alignment between unique and redundant information. This penalty is only activated when two normalization factors are too similar; once their similarity is below a certain threshold... The penalty will then cease to apply. This is because the item is updated simultaneously. and Unique information will move away from redundant information, and redundant information will also remove modality-specific remnants. These factors together contribute to... Decomposed into redundant information and unique information .
[0069] (b) Extraction of collaborative information.
[0070] Only full-modal paths can access collaborative information; therefore, collaborative branches must be learned from the joint input of multimodal data, rather than from any single-modal data. Thus, this application first enables collaborative information to characterize the stable semantics present in the multimodal view, i.e., maximizing... , Refers to the collaborative information of multimodal data. This represents the multimodal features of multimodal data. By matching the projection heads of two augmented views, only the portions that remain stable between the label-based augmentations are retained, thereby constructing the InfoNCE loss for collaborative information. : .
[0071] The aforementioned loss term can capture stable joint information from multimodal data, including cooperative information S. However, using this term alone cannot separate S. Since the full-modal path also contains redundant and unique information, optimizing only... We will still get a normal fusion representation, that is .
[0072] To explicitly separate the true cooperative information, this application then removes all information that can be interpreted from the reachable paths of single-modal data, i.e. All single-modal features jointly cover Therefore, the remaining portion of the multimodal features is S. Based on this concept, the goal of information exclusion is... However, simply minimizing this mutual information is not sufficient, because Collapse may occur. Therefore, this application combines the exclusion of single-modal features with the anchoring of multimodal features, in order to... While removing unimodal features, it still maintains their association with complete multimodal features.
[0073] by Inspired by this, this application approximates this tradeoff by adding an entropy term to the alignment with the relative multimodal features. This results in a multimodal feature anchoring term: , This represents the inner product operation. This is a hyperparameter. Its purpose is to remove singleton reachability information while simultaneously enabling... It remains bound to multimodal features. For the loss term excluded by single-modal features, a second separation loss is used to separate the multimodal representation from collaborative information. : .
[0074] The aforementioned second separation loss In this context, the first term prevents the collapse of collaborative information. The second term grounds it in complete multimodal features. The third term applies an angular interval between collaborative information and all unimodal features, thereby removing information accessible to unimodal data. Because each... cover ,Will Removing features from all unimodal features will remove them. With the information still anchored by multimodal features, the remaining content is the collaborative information S. Therefore, we have →S.
[0075] Combining the various losses constructed in the aforementioned stages of decomposing unique and redundant information and extracting collaborative information, we have: This represents the InfoNCE loss function for A and B. , ; , , .
[0076] The aforementioned losses are used to train the encoder and projector to decouple unique and redundant information, as well as to explicitly separate cooperative information. The total inference loss consists of these loss terms.
[0077] In one alternative implementation, the total loss includes the InfoNCE loss of unique information from weighted fusion, the InfoNCE loss of redundant information, the InfoNCE loss of collaborative information, the first separation loss, and the second separation loss. Specifically, the total loss... Represented as: ; in, These represent the preset redundant information InfoNCE loss weight, first separation loss weight, collaborative information InfoNCE loss weight, and second separation loss weight, respectively. In the total loss... middle, Used to decompose single-modal data reachability information into unique information. and redundant information ; This is used to extract collaborative information from multimodal features by removing reachable information from unimodal data.
[0078] Using multimodal data samples, the encoder and projector are iteratively updated based on the total loss until convergence or the maximum number of iterations is reached. The encoder and projector parameters are saved for identifying the multimodal representation of the target multimodal data to be encoded. In one optional implementation, the encoder and projector updates are implemented using PyTorch and PyTorchLightning. All models use AdamW, with a fixed training epoch and an early stopping strategy based on the validation set.
[0079] The training process for the encoder and projector head can be referenced in the following pseudocode:
[0080]
[0081] In this embodiment of the application, different datasets (multimodal data samples) can be used for training. The hyperparameters used by the encoder and projector on different datasets are shown in Table 1.
[0082] Table 1. Hyperparameters used on different datasets
[0083] Epsilon ( ) represents the cosine repulsion interval, entropy_weight( ) Control entropy regularization, align_weight ( ) controls weak prototype alignment, lambda_repel( The control of cosine repulsion, R, RU_dis, S, and S_dis correspond to respectively .
[0084] II. Model Reasoning Stage.
[0085] (5) Multimodal representation reasoning.
[0086] After training the encoder and projector (cascaded relationship), the target multimodal data is acquired, which is the multimodal data to be decoupled for multimodal representation. The encoder and projector are used to obtain the target-specific information, target redundancy information and target cooperative information.
[0087] The above method can explicitly decouple the unique information, redundant information, and collaborative information of multimodal data, rather than simply classifying them into two categories: unique features and overall multimodal features. It refines the unique features of each single-modal data when exposed individually and when jointly interpreted, providing downstream inference tasks with highly interpretable, multidimensional, and independent multi-information components to improve inference gain.
[0088] Taking MOSI data as an example, which contains bimodal data of text and image, the qualitative visualization decoupling result of the learned representation structure obtained by applying the above method is as follows: Figure 3 As shown, unique, redundant, and collaborative information form explicitly distinguishable regions in the t-SNE space, indicating that the learned branches encode different types of information rather than collapsing into a single shared embedding. In particular, the collaborative sub-information is clearly separated from the information accessible from the unimodal data (redundant information + unique information), consistent with the design goal of removing information that can already be recovered from the unimodal data. This decoupling result is relevant to the corresponding task of MOSI, as sentiment prediction often relies on the interaction between textual content and visual sentiment cues.
[0089] Based on the ideas of this application, embodiments of this application also provide a multimodal representation decoupling device for collaborative perception tasks, such as... Figure 4 As shown, it includes a first processor and a first storage medium. The first storage medium stores a first computer program, and the first processor runs the first computer program to execute the multimodal representation decoupling method of any of the above optional embodiments, so as to achieve decoupling output of target-specific information, target redundancy information and target cooperative information of target multimodal data.
[0090] Furthermore, this application also provides a multimodal data reasoning method for collaborative perception tasks, such as... Figure 5 As shown, the method includes the following stages: (6) Using the multimodal representation decoupling method of the above optional embodiments, obtain the target-specific information, target redundancy information and target cooperative information of the target multimodal data.
[0091] For target multimodal data to be reasoned about (such as classification, prediction, etc.), its construction is the same as the multimodal data samples mentioned above, and it can be constructed from text, images, audio, time series, tabular data, etc.
[0092] (7) Weighted concatenation of target-specific information, target-redundant information and target-cooperative information is performed to obtain the concatenated representation vector.
[0093] When target-specific, target-redundant, and target-cooperative information is obtained, the method of simply splicing and connecting linear probes applies the same fixed fusion rule to each set of decoupled information. This approach has limitations because different multimodal data may contain different proportions of redundant, unique, and cooperative information. Therefore, for a given target multimodal data, an informative factor should be one that provides evidence not yet covered by other factors. To address this, this application employs a weighted splicing of target-specific, target-redundant, and target-cooperative information.
[0094] In one optional implementation, the method for calculating the splicing weights of the weighted splicing includes: Aggregate the target redundancy information from each single-modal data. That is, all... Clustering into a single candidate factor , here This represents the target redundancy information of the i-th single-modal data of the target multimodal data; M still represents the total number of modes of the target multimodal data.
[0095] A candidate set is constructed by combining the target-specific information, the aggregated target redundancy information, and the target cooperative information. , Represents the candidate set The k-th candidate factor, K represents the total number of candidate factors in the candidate set. Corresponding to actual information construction, then... ; here Indicates target collaboration information. This represents the target-specific information of the i-th single-modal data in the target multimodal data.
[0096] For the candidate set Each candidate factor in The candidate factor is calculated using the following method. splicing weight : ; in, This represents the temperature parameter of the softmax function; Represents the cosine similarity function. Both represent candidate factor indices. Representing the candidate set respectively The first in One candidate factor.
[0097] Based on the calculated connection weight The final constructed concatenated representation vector This preserves the identity of each decoupled piece of information while allowing their contributions to vary across different samples. In this way, this application replaces a single global fusion rule with an instance-specific recombination mechanism, thereby reflecting the composition of the three-component information inherent in each multimodal data set.
[0098] (8) Input the concatenated representation vector into the downstream inference model to obtain the inference result.
[0099] Depending on the reasoning task, the choice of downstream reasoning model also varies. For classification tasks, a pre-trained classifier model is used; for prediction / regression tasks, a pre-trained prediction model or a temporal regression model is used. That is, the specific choice of reasoning model is determined by the reasoning task, which is not an improvement of this application. This application mainly optimizes the explicit decoupling of unique information, redundant information and collaborative information, as well as the adaptive weighted concatenation of the three components of information. Step (8) is only a step implemented to maintain the integrity of the reasoning task execution and is not within the scope of optimization considered in this application. Considering only the example, for the emotion perception task, an MLP classifier model can be used to infer positive or negative emotions, or to infer ironic or non-ironic emotions. For the robot operation perception task, an MLP classifier or regressor can be used. The classifier is used to identify contact or non-contact, and the regressor is used to infer the position of the end effector.
[0100] Based on the ideas in this application, such as Figure 6 As shown, this application also provides a multimodal data reasoning apparatus for collaborative perception tasks, which includes a second processor and a second storage medium. The second storage medium stores a second computer program, and the second processor runs the second computer program to execute the multimodal data reasoning method of any of the optional embodiments described above.
[0101] To verify the effectiveness of the proposed method, the model was evaluated using a frozen linear probe on a synthetic benchmark and six real-world datasets. Accuracy, weighted / macro-average F1, or MSE are reported according to task type, and comparisons are made with representative baseline methods covering multimodal contrastive learning, decoupled representation learning, cross-modal interaction modeling, and strongly pre-trained vision-language models. Datasets and hyperparameter settings are shown in Table 1. The baseline methods include Cross, Cross+self, FactorCL, DisentangledSSL, CoMM, CMC, SimCLR, CLIP, and SLIP.
[0102] The Cross baseline uses only cross-modal pairing for contrastive learning, following the CLIP-style cross-modal alignment objective. It adopts the Cross naming and benchmarking protocol from previous multimodal work. Specifically, this method aligns the representations of the target stimuli from different modalities with each other, without explicitly modeling intramodal interactions.
[0103] Cross+Self, building upon the Cross baseline, introduces both cross-modal and intra-modal contrastive objectives, embodying the idea of multimodal contrastive training. It follows the same baseline naming and protocol as previous multimodal work. In addition to aligning representations from different modalities, the model also imposes consistency constraints between augmented views within the same modality. This dual objective improves the robustness of representations while maintaining the basic cross-modal alignment mechanism.
[0104] FactorCL introduces a factorized contrastive learning framework that separates modality-specific components from shared components during representation learning. By fusing the decoupling between modality-dependent and modality-invariant features, FactorCL aims to improve the quality of multimodal embeddings while preserving complementary information from each modality.
[0105] DisentangledSSL learns multimodal representations by explicitly decomposing features into shared and modality-specific components. It combines contrastive learning and decoupling constraints to reduce interference between public and private factors and improve robustness under modal heterogeneity. Compared to the ternary decomposition proposed in this application, this method serves as a strong binary / private decoupling baseline.
[0106] CoMM is a contrastive multimodal representation learning framework designed to explicitly model different types of consistent interactions. CoMM first aligns all shared information consistently, then distinguishes between co-information and collaborative information components through a carefully designed masking strategy and contrastive objectives. The model employs a Transformer-based fusion module to integrate modality-specific representations.
[0107] Contrastive Multiview Coding (CMC) learns representations by maximizing mutual information between different views of the same data instance. In a multimodal scenario, each modality is treated as a distinct view, and the goal is to encourage consistency in representations across modalities. This method focuses on learning shared representations between modalities while ignoring modality-specific variations.
[0108] SimCLR is a widely used self-supervised representation learning framework based on contrastive learning. It generates two augmented views for the same input and learns the representation by maximizing their consistency in the latent space. Although SimCLR was originally designed for unimodal data, it can be adapted to multimodal settings by treating modality-specific signals as different views.
[0109] CLIP learns a joint image-text representation by aligning visual and text embeddings through large-scale contrastive learning. The model consists of independent image and text encoders, trained by maximizing the similarity between matching pairs while minimizing the similarity between non-matching pairs. CLIP representation has demonstrated strong transfer capabilities across a wide range of multimodal tasks.
[0110] SLIP combines contrastive language-image pre-training with self-supervised visual representation learning. In addition to the cross-modal alignment objective used in CLIP, SLIP introduces a SimCLR-style objective in the visual modality. This hybrid training strategy improves the quality of visual representations while maintaining cross-modal alignment.
[0111] The final experimental results are as follows: The test results on the Trifeature dataset are shown in Table 2.
[0112] Table 2. Linear detection accuracy (%) on Trifeature
[0113] The test results on the MultiBench dataset are shown in Table 3.
[0114] Table 3. Linear Probe MSE (×10) for Regression Task on MultiBench 4 ) and a table of top-1 accuracy (%) for classification tasks
[0115] The test results on Vision&Touch (V&T) and UR-FUNNY are shown in Table 4.
[0116] Table 4. Top-1 accuracy (%) of linear detection on classification tasks on Vision&Touch and UR-FUNNY
[0117] The test results on MM-IMDb are shown in Table 5.
[0118] Table 5. Linear detection F1 score (weighted and macro average) (%) in MM-IMDb indoor environment.
[0119]
[0120] Furthermore, ablation analysis was conducted on three key design choices of the method model: three-part weighted splicing, decoupling regularization, and three-part decoupling. These ablation experiments collectively elucidate how adaptive routing, decoupling constraints, and explicit three-branching affect the final performance. Fusion Method. This application compares Conditional-Entropy Routing (the method described in this application) with the Raw Concat baseline on MOSI, UR-FUNNY, MUSTARD, and UR-FUNNY3. On these datasets, cross-modal interaction makes the fusion effect more significant. In Raw Concat, the frozen three-component features are directly input into a linear probe, resulting in a dataset-level fusion rule. Table 6 shows that the routing mechanism brings improvements on all four datasets, increasing the average result from 65.30 to 66.14, with the largest gain on MUSTARD.
[0121] Table 6 Ablation Experiment Table of Fusion Strategy (Results are the mean ± standard deviation of five runs)
[0122] Regularization sensitivity. This application investigates decoupling sensitivity on UR-FUNNY by varying `entropy_weight` and `lambda_repel`. Figure 7 As shown, both coefficients have an intermediate optimal value, which indicates that decoupling regularization is useful.
[0123] The components are represented. This application ablates the ternary decomposition on Trifeature by removing one component at a time. Table 7 shows that the complete model performs best on average. Removing cooperating branches causes the largest decrease, especially on cooperating targets, from 83.56 to 70.20; while removing a unique branch mainly affects its corresponding unique probe.
[0124] Table 7 shows the ablation experiments of the three components on Trifeature. One branch is removed from each row, and the linear detection accuracy (%) is reported.
[0125] As demonstrated by the comparative experiments with the baseline model and the ablation experiments of the proposed method, this application proposes a multimodal self-supervised framework capable of explicitly separating unique, redundant, and collaborative information. It learns modality-specific signals, shared signals, and "jointly available only" signals from unimodal and multimodal data, and recombines these three components through an instance-level parameterless conditional entropy routing strategy guided by conditional entropy reduction. Experiments on controlled and real-world benchmarks show that this application can improve the quality of multimodal representations, especially when downstream tasks involve modality-specific evidence or cross-modal interactions.
[0126] This invention is not limited to the specific embodiments described above. The invention extends to any new feature or combination disclosed in this specification, as well as any new method or process step or combination disclosed herein.
Claims
1. A multimodal representation decoupling method for collaborative perception tasks; characterized in that, The multimodal representation includes redundant information, unique information, and collaborative information from the multimodal data; the unique information represents evidence that can only be obtained from each individual modal data; the redundant information represents evidence shared by each individual modal data and that can be recovered from more than one modal data; the collaborative information represents evidence that can only be obtained when partial modal data from more than one modal data are jointly interpreted; the multimodal data includes at least two of the following: text, image, audio, time series, or tabular data. Decoupling methods include: Acquire multimodal data samples, and construct a first enhanced view and a second enhanced view for the multimodal data samples, which preserve labels. The first augmented view and the second augmented view are encoded using an encoder to obtain the unimodal features of each unimodal data sample, the first multimodal features of the first augmented view, and the second multimodal features of the second augmented view; Based on the single-modal features, the first multimodal features, and the second multimodal features, the unique information, redundant information, and cooperative information are inferred from the three branches of the projection head, respectively. The parameters of the encoder and the projection head are updated based on the total inference loss; wherein the total loss includes the InfoNCE loss of the unique information, the InfoNCE loss of the redundant information, the InfoNCE loss of the cooperative information, the first separation loss between the unique information and the redundant information, and the second separation loss of the cooperative information from the multimodal representation; the InfoNCE loss of the redundant information includes the prediction loss of the redundant information predicted by the unique information, and the consistency loss of the redundant information; The target multimodal data is acquired, and the target-specific information, target redundancy information, and target cooperative information are obtained using the encoder and projector.
2. The multimodal representation decoupling method as described in claim 1, characterized in that, The loss of the unique information in InfoNCE for: ; The predicted loss for: ; The consistency loss for: ; The InfoNCE loss of the collaborative information for: ; in, This represents the InfoNCE loss function for A and B. , ; For the first enhanced view The single-modal features of the i-th single-modal data sample. This represents the i-th single-modal data sample in the first augmented view; For the second enhanced view The single-modal features of the i-th single-modal data sample. This represents the i-th single-modal data sample in the second enhanced view; This is to remove redundant information from the i-th and j-th single-modal data samples in the first enhanced view. This indicates redundant information in the i-th and j-th single-modal data samples in the second enhanced view; This represents the unique information of the i-th single-modal data sample in the first enhanced view; This represents unique information about the i-th single-modal data sample in the second enhanced view; This indicates collaborative information for the first enhanced view; This represents the collaborative information of the second enhanced view; M is the total number of modalities in the multimodal data samples; This represents the number of combinations of any two single-modal data points randomly selected from M types of single-modal data.
3. The multimodal representation decoupling method as described in claim 1, characterized in that, First separation loss for: ; Second separation loss for: ; in, For multimodal data samples Total number of modes, This represents the i-th single-modal data sample in the multimodal data sample; The cosine hinge penalty is defined as follows: ,in, They represent The b-th vector in the array, , ; The set threshold; This indicates the inner product operation; This is the characteristic entropy estimation function; All are hyperparameters; For the first enhanced view The single-modal features of the i-th single-modal data sample. This represents the i-th single-modal data sample in the first augmented view; For the second enhanced view The single-modal features of the i-th single-modal data sample. This represents the i-th single-modal data sample in the second enhanced view; This represents the first multimodal feature of the first enhanced view; This represents the second multimodal feature of the second enhanced view; This represents the unique information of the i-th single-modal data sample in the first enhanced view; This represents unique information about the i-th single-modal data sample in the second enhanced view; This represents redundant information for the j-th single-modal data sample in the first enhanced view; This represents redundant information for the j-th single-modal data sample in the second enhanced view; This indicates collaborative information for the first enhanced view; This indicates collaborative information for the second enhanced view.
4. A multimodal representation decoupling device for collaborative sensing tasks; characterized in that, It includes a first processor and a first storage medium, the first storage medium storing a first computer program, the first processor running the first computer program to perform the multimodal representation decoupling method as described in any one of claims 1-3.
5. A multimodal data reasoning method for collaborative perception tasks; characterized in that, include: The multimodal representation decoupling method described in any one of claims 1-3 is used to obtain the target-specific information, target redundancy information, and target cooperative information of the target multimodal data; The target-specific information, target redundancy information, and target cooperative information are weighted and concatenated to obtain a concatenated representation vector. The concatenated representation vector is input into the downstream inference model to obtain the inference result.
6. The multimodal data inference method as described in claim 5, characterized in that, The method for calculating the splicing weight in the weighted splicing includes: Aggregate target redundancy information from each single modality of data; The target-specific information, the aggregated target redundancy information, and the target cooperative information are used to construct a candidate set. , Represents the candidate set The k-th candidate factor, K is the total number of candidate factors in the candidate set; For the candidate set Each candidate factor in The candidate factor is calculated using the following method. splicing weight : ; in, This represents the temperature parameter of the softmax function; Represents the cosine similarity function. Both represent candidate factor indices. Representing the candidate set respectively The first in One candidate factor.
7. A multimodal data inference device for collaborative perception tasks; characterized in that, It includes a second processor and a second storage medium, the second storage medium storing a second computer program, the second processor running the second computer program to execute the multimodal data inference method as described in claim 5 or 6.