Multi-source heterogeneous data representation consistency modeling method based on attention fusion

By employing collaborative noise separation, unified metric projection, and an end-to-end alignment-attention training framework, combined with a conflict-driven closed-loop calibration mechanism, the noise interference and consistency issues in multi-source heterogeneous data fusion are resolved, achieving adaptive alignment and robustness improvement of multimodal features.

CN121524904APending Publication Date: 2026-02-13MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511491528.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-19
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing technologies cannot effectively block the cooperative interference of cross-modal noise in the fusion of multi-source heterogeneous data, resulting in systematic distortion and robustness degradation. Furthermore, the lack of end-to-end consistency guarantees leads to a high cross-modal representation conflict rate.

Method used

We employ a collaborative noise separation mechanism, a latent space projection mechanism with a unified metric, and an end-to-end alignment-attention joint training framework. Combined with a conflict-driven closed-loop calibration mechanism, we achieve adaptive alignment of cross-modal features and decision consistency through a multi-head cross-modal attention mechanism and an entropy calibration function.

Benefits of technology

It effectively suppresses cross-modal noise interference, achieves adaptive alignment and robustness improvement of multimodal features, reduces data representation fragmentation rate, and improves the accuracy and system robustness of cross-media retrieval tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121524904A_ABST
    Figure CN121524904A_ABST
Patent Text Reader

Abstract

The invention provides a multi-source heterogeneous data representation consistency modeling method based on attention fusion, and relates to the field of artificial intelligence, and the method comprises the steps: carrying out the collaborative denoising of multi-source heterogeneous data, and obtaining a heterogeneous data stream; carrying out preprocessing and feature conversion on the heterogeneous data stream by adopting a multi-mode encoder to obtain a heterogeneous feature vector; performing cross-modal mapping on the heterogeneous feature vectors through a dynamic routing network, and establishing representation association among heterogeneous data in a unified hidden space by using a multi-head cross-modal attention mechanism to obtain joint representation of heterogeneous features; on the basis of a generalization control framework capable of implementing conflict resolution on the multi-modal joint embedding space based on a micrologic rule, data representation conflict detection is performed on joint representation of heterogeneous features, and a structured conflict report is obtained; and based on the structured conflict report, correcting the multi-modal decision confidence coefficient through an adaptive weight distribution mechanism, outputting a structured decision report of the confidence coefficient after multi-modal calibration, and attaching a global confidence coefficient index and a weight adjustment suggestion.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence, and particularly relates to a multi-source heterogeneous data representation consistency modeling method based on attention fusion. BACKGROUND

[0002] In recent years, with the development of information technology, the data involved in the data collection and connection process of various business systems contains data from multiple sources and multiple structures. The related technology for fusing multi-modal data in a multi-source heterogeneous scene continues to develop. Although the ImageBind launched by Meta in 2024 has constructed a six-modal unified embedding space, it cannot solve the problem that the single-modal denoising mechanism cannot block the cross-domain pollution of voice background noise on visual edge features. Although the PaLI-X system developed by Google in the same year introduces cross-modal attention gating technology, the lightweight routing network cannot reconcile the structural heterogeneity of voice waveform sequences, text symbol streams and image pixel matrices, resulting in the cross-modal mapping and attention fusion process being mutually fragmented, which causes systematic faults in the cross-modal representation transmission link.

[0003] The current mainstream systems such as the Sora video generation framework of OpenAI and the GPT-4V multi-modal reasoning engine lack full-link consistency guarantee, and the data representation conflict rate continues to be higher than 22%. Especially when the behavior-object combination violates the common sense rules, the traditional Euclidean space linear measurement cannot sensitively capture the contradiction signals. Moreover, the representation conflict recognition result and the decision confidence calibration have long been operated as independent modules, which causes the system robustness to deteriorate sharply under noise interference, and finally leads to the systematic amplification of the cross-modal gap from the bottom feature mapping to the high-level representation generation, resulting in representation conflicts in the joint embedding space. SUMMARY

[0004] Based on the defects described in the background, the present application aims to solve at least one of the following core problems: 1. Synergistic interference suppression of cross-modal noise Existing single-modal denoising technology cannot effectively block the cross-domain coupling effect between voice, image and text noise, such as the pollution of voice background noise on visual edge features, resulting in systematic distortion of heterogeneous data in the cross-modal mapping stage. The present application establishes a synergistic noise separation mechanism to eliminate the cross-domain noise interference source while preserving the correlation of multi-modal representations.

[0005] 2. Feature space compatibility defects When the lightweight routing network processes voice waveform sequences, text symbol streams and image pixel matrices at the same time, due to the essential differences in data structure and dimension, the feature space generated by the encoder has an intractable heterogeneity contradiction. The present application constructs a unified metric hidden space projection mechanism to realize the dimension decoupling and structure alignment of cross-modal features.

[0006] 3. Attention mechanism and cross-modal mapping coordination-induced data representation fault Although the current mainstream multi-head attention architecture can capture cross-modal interaction, it lacks joint optimization with the underlying cross-modal mapping process, resulting in a large deviation between high-level fusion features and original data. The invention designs an alignment-attention joint training framework to ensure the consistency of the whole link from the bottom feature mapping to the high-level data representation generation.

[0007] 4. Conflict-driven closed-loop calibration mechanism The existing system cannot dynamically adjust the modal weight in real time using conflict information due to the independent operation of the data representation conflict identification and confidence calibration modules, limiting the system's adaptability and robustness. The invention uses a conflict-driven closed-loop calibration mechanism to input the report generated by the data representation conflict detection into the confidence calibration module, and uses an entropy calibration function and a double-gated residual update mechanism to achieve progressive correction, complete conflict resolution and weight modulation in a unified framework, form a closed-loop control link, and ensure decision consistency and adaptability.

[0008] To this end, the invention provides the following technical solutions: a multi-source heterogeneous data representation consistency modeling method based on attention fusion, comprising the following steps: Step 1: Perform collaborative denoising on the input multi-source heterogeneous data to obtain denoised heterogeneous data streams; Step 2: Preprocess and feature convert the obtained denoised heterogeneous data streams using a multi-modal encoder to obtain heterogeneous feature vectors; Step 3: Perform cross-modal mapping on the obtained heterogeneous feature vectors through a dynamic routing network, establish representation association between heterogeneous data streams in a unified hidden space using a multi-head cross-modal attention mechanism, and fuse heterogeneous features based on the established representation association to obtain joint representation of heterogeneous features; Step 4: Based on the generalization control architecture that uses a differentiable logic rule to resolve conflicts in the multi-modal joint embedding space, perform data representation conflict detection on the obtained joint representation of heterogeneous features to obtain a structured conflict report; Step 5: Based on the obtained structured conflict report, correct the multi-modal decision confidence through an adaptive weight allocation mechanism, output a structured decision report of the multi-modal calibrated confidence, and attach global confidence indicators and weight adjustment suggestions.

[0009] Further, the multi-source heterogeneous data includes three types of heterogeneous data: speech, text, and image.

[0010] Further, in step 1, the denoising of speech data includes: Step 111, the input original speech signal is converted into a complex spectrum including an amplitude spectrum and a phase spectrum through a short-time Fourier transform; Step 112, the one-dimensional convolutional neural network is used to process the spliced features of the amplitude spectrum and the phase spectrum of the converted complex spectrum, and a complex mask matrix is output; Step 113, the output complex mask matrix is used to filter the converted complex spectrum to obtain a filtered spectrum; Step 114, the filtered spectrum is subjected to inverse Fourier transform to reconstruct a time domain signal, and a time coherence optimization is performed through a double-path recurrent neural network, and a speech signal with time dimension and channel characteristics is output.

[0011] Further, in step 1, the denoising of the text data includes: Step 121, performing a structured noise operation on the input text sequence; wherein the structured noise operation includes synonym replacement and entity mask masking; Step 122, generating a noise text based on the generator for the text sequence after the structured noise operation, and the discriminator reconstructs the data representation through the loss function; Step 123, constructing a noise-clean text pair, and performing similarity optimization in a data representation space to maximize the data representation embedding similarity of the denoised text and the original clean text and minimize the relevance to the noise sample through a contrastive learning mechanism, and outputting a text signal composed of discrete symbol streams of words or characters.

[0012] Further, in step 1, the denoising of the image data includes: Step 131, performing a two-dimensional fast Fourier transform on the input image signal to perform frequency domain decomposition, and obtaining low-frequency components and high-frequency components based on a cutoff frequency; Step 132, inputting the obtained low-frequency components into a U-Net network to generate a structure reconstruction result , and inputting the obtained high-frequency components into a gated convolutional network to generate a detail reconstruction result ; Step 133, weighting and summing the generated structure reconstruction result and the detail reconstruction result to obtain a reconstructed image; Step 134, generating a spatial attention map based on the obtained low-frequency components and high-frequency components, and the generated spatial attention map is used to perform pixel-level correction on the obtained reconstructed image.

[0013] Further, step 2 includes: Step 21, determining the modal type of the input denoised heterogeneous data stream through a dynamic routing decision network, which identifies the modal type to which the current input belongs based on data characteristics and generates a modal identification code; Step 22, based on the identified modality type to which the current input belongs, perform data representation labeling operation: For denoised speech data, extract phoneme-syllable topology to generate acoustic event labels; For denoised image data, generate visual concept labels through edge gradient distribution; Step 23, according to the generated modality identification code and label set, route the denoised heterogeneous data stream to the corresponding encoder branch, and perform feature conversion on the structure of the heterogeneous data to obtain a heterogeneous feature vector including speech spectrum features, text data representation features, and image visual features.

[0014] Further, in step 23, the denoised heterogeneous data stream is routed to the corresponding encoder branch, and the structure of the heterogeneous data is preprocessed and feature-converted, including: The denoised speech data is input into a spectrum encoder; The denoised text data is input into an encoder based on the Transformer architecture; The denoised image data is input into a convolutional neural network encoder.

[0015] Further, step 3 includes: Step 31, map the obtained heterogeneous feature vector to a unified hidden space of a unified dimension through a trainable projection matrix; Step 32, construct a multi-head attention mechanism, each attention head can independently learn the fine-grained dependency relationship between different modalities: Take text features as query source, speech features as key source, and image features as value source, and calculate the interaction weight matrix through scaled dot-product attention; Step 33, calculate the cosine similarity between speech features and text features, and the similarity between image features and text features, and based on the calculated cosine similarity between speech features and text features, and the similarity between image features and text features, generate dynamic fusion weight coefficients of speech features and text features, and dynamic fusion weight coefficients of image features and text features, respectively; Step 34, concatenate and fuse the weighted heterogeneous features to obtain a joint representation of the heterogeneous features.

[0016] Further, step 4 includes: Step 41, map the obtained joint representation of the heterogeneous features to a Riemannian manifold with constant negative curvature, and measure the logical compatibility between the heterogeneous features based on the hyperbolic distance function; Step 42, when the hyperbolic space projection determines a logical conflict event, use a causal reasoning engine to locate the root cause of the conflict, and output a structured conflict report containing conflict type, output confidence score, and associated feature position index.

[0017] Further, step 5 comprises: Step 51, constructing an entropy calibration function based on the obtained joint representation of heterogeneous features and structured conflict reports; Step 52, generating progressively modified intermediate confidence based on the constructed entropy calibration function, using a double-gated residual update mechanism including update gate control, reset gate control, and residual correction; Step 53, calculating a dynamic temperature parameter based on feature information entropy, performing temperature scaling operation on the intermediate confidence to obtain the modified confidence; Step 54, normalizing the obtained modified confidence through a softmax function; outputting a structured decision report containing calibrated confidence of voice, text, and image three modalities.

[0018] Compared with the prior art, the core advantage of the present application is: 1. Cross-modal noise collaborative suppression and feature fidelity enhancement Through the triple denoising mechanism of frequency domain-time domain-space domain, the correlation of multi-modal representation is preserved while the pollution of voice background noise to visual features is eliminated. The voice double-path recurrent network optimizes the timing coherence and blocks the propagation of cross-domain interference; the image frequency domain decomposition and dynamic weight fusion accurately separate high-frequency noise and edge details; the text adversarial training reconstructs the pure data representation space. Collaborative guarantee the signal-to-noise ratio improvement of multi-modal input data, and establish a high-fidelity foundation for data representation alignment.

[0019] 2. Heterogeneous feature adaptive alignment and data representation conflict resolution The dynamic routing network analyzes the structural differences of modal data, and realizes the dimension decoupling of voice waveform, text symbol and image pixel through unified hidden space projection. The multi-head cross-modal attention mechanism establishes a fine-grained correlation channel between text, vision and voice, and completes adaptive feature fusion in the unified embedding space. Based on the hyperbolic geometric manifold, the conflict perception sensitivity is enhanced, and the behavior-object conflict combination is accurately located by combining the causal reasoning engine, so that the data representation fault rate is reduced.

[0020] 3. Closed-loop decision robustness and scene generalization leap The conflict-driven confidence calibration mechanism connects the "detection-revision" closed-loop link, the entropy calibration function dynamically modulates the voice / text / image modal weight, and the double-gated residual update realizes the progressive revision under noise interference. This design improves the data representation uniformity understanding accuracy in cross-media retrieval tasks, and reduces the system robustness entropy value in complex noise environment. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 The overall flowchart of the multi-source heterogeneous data representation consistency modeling method based on attention fusion of the present application; Figure 2 Flow chart of multi-modal denoising structure of the application; Figure 3 Flow chart of multi-modal data acquisition and preprocessing operation of the application; Figure 4 Flow chart of multi-head attention cross-modal association operation of the application; Figure 5 Flow chart of data representation conflict detection structure of the application. DETAILED DESCRIPTION

[0022] The application first performs collaborative denoising on three types of heterogeneous data: speech, text, and image, while ensuring that the speech dual-path recurrent network optimizes the timing coherence, the text adversarial training reconstructs the complete data representation, and the image frequency domain decomposition preserves the edge details, eliminating cross-modal noise interference; then through a dynamic routing network, cross-modal mapping is realized, and a multi-head cross-modal attention mechanism is used to establish fine-grained data representation association in a unified hidden space; further, based on hyperbolic geometric manifold, data representation conflict perception is enhanced, and logical contradictions are located by combining a causal reasoning engine; finally, through a conflict-driven closed-loop calibration mechanism, the modal weight is dynamically modulated, and the entropy calibration function and double-gated residual update are used to realize end-to-end decision robustness.

[0023] The overall process of the multi-source heterogeneous data representation consistency modeling method based on attention fusion of the application is as follows Figure 1 .

[0024] Each step in the application will be described in detail below in conjunction with the relevant drawings. Figure 1

[0025] 1. Multi-modal data denoising 1.1 Speech data denoisingAfter receiving the original speech data from the interface, this step performs denoising based on frequency domain transformation and convolutional neural network processing.

[0026] The original speech signal is received and decomposed into a complex spectrum through short-time Fourier transform: wherein, is the frequency point index, used to identify the frequency domain coordinates, corresponding to the energy distribution position of the signal, is the time frame number, used to locate the timing evolution characteristics, reflecting the window position of the framing processing, is the amplitude spectrum, is the phase spectrum. Then the one-dimensional convolutional neural network processes the spliced features of the amplitude spectrum and the phase spectrum and , outputting a complex mask matrix . Next, the filtered spectrum ​The time domain signal is reconstructed by inverse Fourier transform, and the time series coherence optimization is performed by a double-path recurrent neural network, and the output is: wherein the summation sign is a convolution operation, represents a time domain kernel, represents the processing result of the signal, and each convolution item represents a pair of filters, ; this defines the reconstructed time domain speech signal As a linear superposition of multiple convolution operations, the essence of convolution operation is to capture the time series dependence of speech signal through local weighting and effectively model long and short term coherence, while suppressing high frequency noise components and preserving key features such as fundamental frequency and formant of speech.

[0027] 1.2 Text data denoising After receiving the text raw data from the interface, this step performs denoising based on data representation reconstruction.

[0028] The input text sequence is subjected to structured noise operation, including synonym replacement and entity mask masking. Next, based on the generator noise text is generated, and the discriminator performs data representation reconstruction through the loss function L.

[0029] wherein, is the expected function of the input text sequence .

[0030] The loss function forces the generator to reconstruct the text expression consistent with the original data representation distribution through the adversarial scoring of the discriminator D on the noise text and the clean text, thereby driving the recovery of data representation integrity. Next, construct the noise-clean text pair , wherein represents the noise sample, represents the clean sample. Define the loss , and perform similarity optimization in the data representation space: wherein, is the vector representation in the encoder network, is a stability coefficient, , T is the transpose, is the clean signal after denoising for the noisy input , which is used for subsequent cross-modal mapping and attention fusion.

[0031] The purpose of performing similarity optimization is to maximize the similarity of data representation embedding between the denoised text and the original clean text through the contrastive learning mechanism, and to minimize the relevance to the noise samples, so as to force the denoising model to capture the essential data representation features and filter the data representation distortion components. Finally, the text dependency tree structure is parsed, and the syntax conflict nodes are filtered.

[0032] 1.3 Image data denoising After receiving the image raw data from the interface, the image representation reconstruction-based denoising work is performed.

[0033] In order to eliminate high-frequency noise interference and preserve image edge detail features, it is necessary to decompose the original input into frequency domain components that can be processed independently, so first perform frequency domain decomposition on the input image with height and width of H and W respectively Perform two-dimensional fast Fourier transform: Where u and v are the coordinates of the frequency components in the image, , .

[0034] For the cut-off frequency , the low-frequency component Input U-Net network to generate structure reconstruction results , and the high-frequency component is generated by the gating convolutional network to generate detail reconstruction results . Next, output the image , where is the dynamic weight , where is the frequency domain discriminator. Finally, generate the spatial attention map A: Where, is the scaled dot product attention function, which is used to perform pixel-level correction, is the gradient operator.

[0035] The specific operation process of this step is shown in Figure 2 .

[0036] 2. Multimodal data acquisition and preprocessing This step receives a heterogeneous data stream from the output of the previous denoising module, which contains three types of structured information. One is the speech waveform signal, which has time dimension and channel characteristics; two is the text symbol sequence, which is a discrete symbol stream composed of words or characters; three is the image pixel matrix, which contains a three-dimensional data structure of height, width and color channel.

[0037] To achieve the structured feature conversion of heterogeneous data and eliminate the dimensional conflict between modalities, it is necessary to accurately identify the structural differences of speech, text and image. Therefore, first, the input data is subjected to modal type determination through a dynamic routing decision network, which identifies the current input belonging to which modal type based on data features and generates a modal identification code. Then, data representation labeling operation is performed: for speech data, phoneme-syllable topology structure is extracted to generate acoustic event labels such as fundamental frequency and intensity; for text data, syntactic dependency relations are parsed to extract entity and action labels such as person name, location and object action; for image data, visual concept labels such as shape, material and spatial relationship are generated through edge gradient distribution; finally, according to the modal identification code and the label set, the data is routed to the corresponding encoder branch, i.e. the speech data is input into the spectral encoder, the text data is input into the encoder based on the Transformer architecture, and the image data is input into the convolutional neural network encoder, thereby realizing the structured preprocessing and feature conversion of heterogeneous data.

[0038] The specific operation process of this step is shown in Figure 3

[0039] 3. Multi-head cross-modal association This step receives the heterogeneous feature vectors output from the multi-modal encoder in the previous step, including speech spectrum features, text data representation features and image visual features.

[0040] Multi-head cross-modal attention mechanism (MCA) is a heterogeneous data fusion architecture that decouples the associated dimensions of heterogeneous data through parallel weight matrix groups. It decouples the associated dimensions of heterogeneous modal data through parallel attention head groups, each of which independently learns a specific modal interaction mode, including spatial positioning, voiceprint fluctuation and data representation analysis, and realizes fine-grained feature coupling in a unified hidden space. This mechanism can significantly improve the completeness and generalization of cross-modal data representation alignment.

[0041] In this step, first, a trainable projection matrix is used to map each modal feature to a unified hidden space of the same dimension, eliminating alignment bias caused by differences in original feature dimensions. The projection of speech features is represented as: The projection of text features is represented as: The projection of image features is represented as: wherein, are the projection matrices of speech, text and image respectively, which map the heterogeneous features to the same dimension ​the latent space of, are the feature matrices of the three modalities respectively, representing the unstructured feature expression of the encoder output, , are the bias vectors of the three modalities respectively, used to enhance the expression ability of the model and compensate for the projection loss. This step ensures that all features are in the same mathematical space for subsequent calculations. Then a multi-head attention mechanism is constructed, allowing each attention head to independently learn the fine-grained dependency between different modalities: taking the text features as the query source, the speech features as the key source, and the image features as the value source, the interaction weight matrix is calculated through the scaled dot-product attention: wherein, is the dimension of the feature space, is a scaling factor that constrains the variance to ensure the stability of backpropagation, denotes the attention function, is a function that maps an arbitrary real number vector to a weight / probability distribution with non-negative components and a sum of 1, used to convert the score into a weight. This process captures the local data representation correlation across modalities through multiple parallel attention heads, enhancing the expression ability of feature interaction.

[0042] Further, based on the feature similarity, the fusion of cross-modal data representation is realized. In the multi-head cross-modal correlation process, the feature similarity refers to the data representation correlation strength of different modal feature vectors in the unified latent space, which is quantitatively calculated by cosine similarity, and its mathematical essence is the alignment degree of two feature vectors in direction. The cosine similarity between speech and text features is calculated as and the similarity between image and text features is , according to which the dynamic fusion weight coefficient is generated: wherein, denotes the sigmoid function, denotes the dynamic fusion weight coefficient between speech and text features, denotes the dynamic fusion weight coefficient between image and text features, , is a learnable fusion weight matrix. Finally, the weighted cross-modal features are concatenated and fused to output the joint representation F: wherein, denotes the concatenation of different modal vectors in the feature dimension. This feature preserves the single-modal specificity and cross-modal data representation consistency, providing a structured input for subsequent data representation consistency modeling.

[0043] The specific operation process of this step is shown in Figure 4 .

[0044] 4. Data representation conflict detection This step performs multi-level logical consistency verification on cross-modal fusion features.

[0045] The data representation consistency constraint (SCC) is a generalization control architecture that resolves conflicts in a multi-modal joint embedding space through differentiable logical rules. Its core is to implement closed-loop conflict resolution in the multi-modal joint embedding space through a differentiable logical rule engine. This architecture contains three progressive control mechanisms.

[0046] The conflict perception layer monitors the logical compatibility of cross-modal features in the unified embedding space in real time, uses hyperbolic geometric manifold to perform nonlinear divergence processing on contradictory features, and is used to enhance the sensitivity of contradictory signals, instead of the linear measurement method in traditional Euclidean space. The dynamic correction layer constructs a rule-driven residual modulation channel. When a data representation conflict is detected, such as a behavior-object combination violating common sense rules, a feature correction gradient is automatically generated to dynamically adjust the encoder parameters through backpropagation. The consensus optimization layer establishes a cross-modal consensus confidence evaluation system, converts the conflict resolution process into a collaborative optimization problem of multi-modal decision weights, and ensures that the output data representation meets the logical consistency standards that can be understood by the human brain. Its structure is shown in Figure 5 .

[0047] First, perform hyperbolic space projection to map the input feature vector to a Riemannian manifold with constant negative curvature. In this geometric space, the feature pairs of data representation conflicts show exponential distance growth. The projection process realizes feature transformation through a trainable parameter matrix and measures the logical compatibility between features based on the hyperbolic distance function: wherein represent different modal sub-features, such as text description features and corresponding image region features. When the distance value D exceeds the preset threshold (default value 0.75), it is determined as a logical conflict event.

[0048] When the hyperbolic space projection determines a logical conflict event, the root cause of the conflict needs to be further analyzed. The hyperbolic distance can only determine the contradiction strength and position between feature pairs, but it does not reveal which specific data representation rules are violated; therefore, a causal reasoning engine is introduced to map abstract contradictions to explainable data representation conflict types based on a common sense knowledge base. Next, the causal reasoning engine is used to locate the conflict source. The engine loads a pre-constructed common sense knowledge graph, which contains object-behavior-scene three-level association rules, thereby generating a causal constraint mask matrix. The element-wise product of the feature attention weight and the causal mask is calculated: wherein, represents the conflict mask matrix generated by the causal reasoning engine, homomorphic to the attention scores, for element-wise suppression on cross-modal pairs violating common sense / rule to resolve data representation conflicts and improve consistency.

[0049] Identify contradictory combinations that violate common sense rules, such as detecting the simultaneous presence of "action" and "empty action object", outputting the "missing action object" conflict label. Finally, output a structured conflict report containing conflict types, including object conflicts, behavior conflicts, and scene conflicts. Also output confidence scores and associated feature position indexes.

[0050] 5. Dynamic confidence calibration This step receives the conflict report and original fusion features output in the previous step, and corrects the multi-modal decision confidence through an adaptive weight distribution mechanism. To avoid weight distribution imbalance caused by excessive confidence or excessive uncertainty in multi-modal conflict resolution, the invention introduces an entropy calibration function in loss construction. This function constrains the entropy value of the confidence distribution, prompting each modality to maintain a moderate level of uncertainty, thereby preventing a single modality from dominating or completely failing in fusion decision-making. In other words, entropy calibration here plays the role of a confidence regulator, making the decision weight distribution of different modalities after conflict fusion more balanced and stable. First, construct the entropy calibration function : wherein, represents the original confidence distribution of the modality (voice / text / image), is the target distribution after conflict report correction, and the KL divergence quantifies the distribution deviation, is the modality trustworthiness coefficient (initial value 1.0), is the system global trust benchmark, is the regularization intensity factor (default 0.1), represents the decision strategy distribution in the multi-modal calibration process. This function dynamically reduces the decision weight of high-conflict modalities through entropy constraint. The entropy calibration function directly quantifies the deviation between the original confidence distribution and the target distribution through KL divergence, where is the central confidence, dynamically adjusting the confidence core value to ensure that the decision weight of high-conflict modalities is suppressed, thereby matching the correction target and maintaining distribution consistency.

[0051] Further, a double-gated residual update mechanism is adopted to generate the calibrated confidence. It mainly includes three modules: update gate control, reset gate control and residual correction.

[0052] The update gate control is to generate an update gate signal between 0 and 1 by a sigmoid function, which is generated by the linear transformation of the spliced result of the fusion feature vector and the conflict report, to control the writing strength of the new confidence. The reset gate control is to generate a reset gate signal by another sigmoid function to determine the retention proportion of the original confidence. The residual correction is to perform element-wise multiplication of the nonlinear activation transformation of the original confidence vector and the update gate signal, and then add the original confidence regulated by the reset gate to output the gradually corrected confidence intermediate value. On this basis, it is further necessary to depict the uncertainty degree of the calibrated distribution of each modality, so the calculation of the feature information entropy is introduced, which quantifies the residual uncertainty of the modality distribution by measuring the entropy value, and serves as an important basis for subsequent confidence reset and decision optimization. The feature information entropy is an index for calculating the uncertainty of the modality distribution, defined as the entropy value , which is used to quantify the degree of information disorder; in the temperature scaling operation, the value of the feature information entropy directly affects the temperature parameter, thereby scaling the intermediate confidence value, so that the confidence of the modality with high entropy value (high uncertainty) is more uniformly distributed, reducing the decision bias.

[0053] Based on the calculation of the dynamic temperature parameter, the greater the information entropy, the higher the temperature, and the temperature scaling operation is performed on the intermediate confidence: first, the intermediate value is multiplied by the temperature scaling factor, which is derived from the modality trustworthiness coefficient, and then divided by the temperature parameter; finally, the softmax function is normalized to ensure that the sum of all modality confidences is strictly equal to 1, while maintaining the unit characteristics of the probability distribution. The output is a structured decision report containing the calibrated confidence of the three modalities of speech, text and image, as well as the global confidence index and weight adjustment suggestions.

[0054] Note that any combination of the technical features of the above embodiments can be made. In order to make the description simple, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the description. The above embodiments only express several implementation ways of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, some modifications and improvements can be made, which are all within the scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for modeling consistency of multi-source heterogeneous data representation based on attention fusion, characterized in that, The method includes the following steps: Step 1: Perform collaborative denoising on the input multi-source heterogeneous data to obtain a denoised heterogeneous data stream; Step 2: The obtained denoised heterogeneous data stream is preprocessed and feature transformed using a multimodal encoder to obtain heterogeneous feature vectors; Step 3: The obtained heterogeneous feature vectors are mapped across modalities through a dynamic routing network. In a unified latent space, a multi-head cross-modal attention mechanism is used to establish representational associations between heterogeneous data. Based on the established representational associations, heterogeneous features are fused to obtain a joint representation of heterogeneous features. Step 4: Based on the differentiable logic rules, a generalized control architecture for conflict resolution is implemented in the multimodal joint embedding space. Data representation conflict detection is performed on the joint representation of the obtained heterogeneous features to obtain a structured conflict report. Step 5: Based on the obtained structured conflict report, the multimodal decision confidence is corrected through an adaptive weight allocation mechanism, and a structured decision report with multimodal calibration confidence is output, along with a global confidence index and weight adjustment suggestions.

2. The method as described in claim 1, characterized in that, Multi-source heterogeneous data includes three types of heterogeneous data: speech, text, and images.

3. The method as described in claim 2, characterized in that, In step 1, the noise reduction of the speech includes: Step 111: The input raw speech signal is transformed into a complex spectrum including amplitude spectrum and phase spectrum by short-time Fourier transform; Step 112: Process the amplitude spectrum and phase spectrum splicing features of the transformed complex spectrum using a one-dimensional convolutional neural network, and output a complex mask matrix. Step 113: The complex spectrum obtained by transformation is filtered using the output complex mask matrix to obtain the filtered spectrum; Step 114: The filtered spectrum is reconstructed into a time-domain signal through inverse Fourier transform, and temporal coherence optimization is performed through a dual-path recurrent neural network to output a speech signal with time dimension and vocal tract characteristics.

4. The method as described in claim 2, characterized in that, In step 1, the text denoising includes: Step 121: Perform structured noise operation on the input text sequence; wherein, the structured noise operation includes synonym replacement and entity masking. Step 122: Generate noisy text based on the generator for the text sequence after performing structured noise operation, and the discriminator reconstructs the data representation through the loss function; Step 123: Construct noise-clean text pairs, perform similarity optimization in the data representation space to maximize the data representation embedding similarity between the denoised text and the original clean text through a contrastive learning mechanism, while minimizing the correlation with noise samples, and output a text signal of a discrete symbol stream consisting of words or characters.

5. The method as described in claim 2, characterized in that, In step 1, image denoising includes: Step 131: The input image signal is decomposed in the frequency domain by two-dimensional fast Fourier transform, and the low-frequency component and high-frequency component are obtained based on the cutoff frequency. Step 132: Input the obtained low-frequency components into the U-Net network to generate the structural reconstruction results. The obtained high-frequency components are then processed through a gated convolutional network to generate detailed reconstruction results. ; Step 133: The generated structural reconstruction results and detail reconstruction results are weighted and summed to obtain the reconstructed image; Step 134: Based on the obtained low-frequency and high-frequency components, a spatial attention map is generated. The generated spatial attention map is used to perform pixel-level correction on the obtained reconstructed image.

6. The method according to any one of claims 2-5, characterized in that, Step 2 includes: Step 21: The input data is modally determined by a dynamic routing decision network. This network identifies the modality type to which the current input belongs based on data features and generates a modality identifier code. Step 22: Based on the identified modality type to which the current input belongs, perform data representation labeling operation: For extracting phoneme-syllable topology from speech data to generate acoustic event labels; For image data, visual concept labels are generated through edge gradient distribution; Step 23: Based on the generated modal identifier and tag set, the data is routed to the corresponding encoder branch, and the structured features of the heterogeneous data are transformed to obtain heterogeneous feature vectors including speech spectrum features, text data representation features and image visual features.

7. The method as described in claim 6, characterized in that, In step 23, the data is routed to the corresponding encoder branch, and the heterogeneous data undergoes preprocessing and feature transformation to achieve structural transformation, including: Input the voice data into the spectrum encoder; Input text data into an encoder based on the Transformer architecture; The image data is then input into the convolutional neural network encoder.

8. The method as described in claim 2, characterized in that, Step 3 includes: Step 31: Map the obtained heterogeneous feature vectors to a unified latent space of a unified dimension using a trainable projection matrix; Step 32: Construct a multi-head attention mechanism, where each attention head can independently learn fine-grained dependencies between different modalities: Using text features as the query source, speech features as the key source, and image features as the value source, the interaction weight matrix is ​​calculated through scaling dot product attention. Step 33: Calculate the cosine similarity between speech features and text features, and the similarity between image features and text features. Based on the calculated cosine similarity between speech features and text features, and the similarity between image features and text features, generate dynamic fusion weight coefficients for speech features and text features, and dynamic fusion weight coefficients for image features and text features, respectively. Step 34: The weighted heterogeneous features are spliced ​​and fused to obtain a joint representation of the heterogeneous features.

9. The method as described in claim 2, characterized in that, Step 4 includes: Step 41: Map the joint representation of the obtained heterogeneous features to a Riemannian manifold with constant negative curvature, and measure the logical compatibility between heterogeneous features based on the hyperbolic distance function. Step 42: After the hyperbolic space projection determines the logical conflict event, the causal reasoning engine is used to locate the root cause of the conflict and output a structured conflict report containing the conflict type, output confidence score and related feature location index.

10. The method as described in claim 2, characterized in that, Step 5 includes: Step 51: Construct an entropy calibration function based on the joint characterization of the obtained heterogeneous features and the structured conflict report; Step 52: A dual-gated residual update mechanism, including update gate control, reset gate control, and residual correction, is used to generate the progressively corrected intermediate confidence level. Step 53: Calculate dynamic temperature parameters based on feature information entropy, perform temperature scaling operation on intermediate confidence scores to obtain corrected confidence scores; Step 54: Normalize the obtained corrected confidence scores using the softmax function; output a structured decision report containing the calibrated confidence scores for the three modalities of speech, text, and image.