An intelligent fusion processing method and system based on cross-modal semantic alignment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-11
AI Technical Summary
这种单纯基于静态规则或全局统计的方法忽略了异构模态内复杂的空间拓扑约束与局部语义关联,且在跨模态对齐时缺乏对亲和度分布自适应的稀疏化能力,导致在构建跨模态关联时引入大量伪连接与冗余交互,从而限制了跨模态特征融合的准确性与鲁棒性
本发明通过引入基于改进EdgeConv模型的特征同质化免疫拮抗机制与跨模态动态双随机软对齐策略,显著提高了多源异构数据跨模态融合过程中特征表征的异质判别力与跨模态语义对齐的精准度。通过创新性地构建基于交互增量分布特征计算动态免疫阈值的拮抗掩码,有效抑制了图聚合过程中的同质化冗余成分,并与原节点特征残差拼接强化了异质性判别力,实现了模态内特征增强与抗同质化衰退的统一约束。同时,采用基于亲和度分布动态稀疏化与行列交替归一化迭代的跨模态软对齐机制,将零元素固定排除在非线性变换之外,精准收敛于非零子空间双随机矩阵,彻底消除了跨模态伪关联与噪声干扰,并在跨模态融合更新中基于分布特征动态计算模态平衡系数,确保了模态内固有特征与跨模态交互特征的自适应权重融合,从而显著提升了跨模态融合决策的准确度、鲁棒性与深层语义交互的可靠性
Smart Images

Figure CN122548644A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing and semantic understanding, and in particular to an intelligent fusion processing method and system based on cross-modal semantic alignment. Background Technology
[0002] In the fields of multimodal data fusion and graph neural networks, with the explosive growth of heterogeneous perceptual data, traditional methods face severe representational challenges in cross-modal semantic alignment and deep feature interactions. Existing cross-modal fusion algorithms, such as attention-based graph fusion models, while utilizing feature mapping to achieve unified modeling of heterogeneous data, mainly rely on global feature interactions or static threshold truncation for cross-modal alignment and graph construction. This method, which is based solely on static rules or global statistics, ignores the complex spatial topological constraints and local semantic associations within heterogeneous modalities. Furthermore, it lacks the ability to adapt to the sparsity of affinity distribution during cross-modal alignment, leading to the introduction of numerous pseudo-connections and redundant interactions when constructing cross-modal associations, thus limiting the accuracy and robustness of cross-modal feature fusion. In addition, traditional graph aggregation methods often struggle to effectively suppress feature homogenization caused by neighborhood interactions during feature propagation and lack a dynamic balance mechanism between intrinsic features within a modality and cross-modal interaction features. This results in the fused features losing their modal heterogeneity discriminative power, increasing the risk of misjudgment in downstream decisions based on multi-source heterogeneous data.
[0003] Therefore, how to provide an intelligent fusion processing method and system based on cross-modal semantic alignment is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0004] This invention proposes an intelligent fusion processing method and system based on cross-modal semantic alignment. Through a cross-modal graph fusion optimization process based on an improved EdgeConv model and a feature homogenization immune antagonism mechanism, entity deconstruction and feature mapping are performed on multi-source heterogeneous modal data, and neighborhood interaction and context aggregation based on feature differences are performed on the intra-modal semantic graph. A feature homogenization immune antagonism mechanism is introduced, calculating a dynamic immune threshold based on the distribution characteristics of interaction increments to generate an antagonistic mask, suppressing homogenization redundancy components, and concatenating it with the feature residuals of the original nodes to enhance heterogeneity discrimination. Simultaneously, dynamic sparsification and row-column alternating normalization iterations are performed based on the inter-modal affinity distribution characteristics to construct a cross-modal soft alignment matrix and a cross-modal fusion graph with non-zero subspace double random constraints, performing feature weighted aggregation and dynamic balance fusion of intrinsic intra-modal features on heterogeneous nodes. This mechanism, by establishing a closed-loop interactive pathway from "dynamic immune antagonistic shielding" to "cross-modal double-random soft alignment," effectively eliminates feature homogenization decay and cross-modal pseudo-association interference during graph neural network aggregation. It ensures that the generated deep-fused cross-modal entity features dynamically maintain the heterogeneous discriminative power of each modality and the consistency of cross-modal semantics, achieving the technical effect of improving the accuracy of cross-modal fusion decision-making while suppressing intra-modal redundancy. This invention overcomes the limitations of traditional methods, such as rigid cross-modal alignment, homogenized feature fusion, and neglect of dynamic distribution constraints, providing an efficient solution for intelligent decision-making with multi-source heterogeneous data.
[0005] An intelligent fusion processing method based on cross-modal semantic alignment according to an embodiment of the present invention includes the following steps: S1. Perform entity deconstruction on multi-source heterogeneous modal data to extract heterogeneous entities and associated structural locations, map heterogeneous entities to a unified feature space to output a heterogeneous feature matrix, and use the associated structural locations as the heterogeneous location set. S2. Based on the heterogeneous location set and heterogeneous feature matrix, determine the topological association and semantic association within the modality to construct the intramodal adjacency matrix, and combine the heterogeneous feature matrix to construct the intramodal semantic graph; S3. Based on the improved EdgeConv model, neighborhood interaction and context aggregation based on feature differences are performed on the intramodal semantic graph. A feature homogenization immune antagonism mechanism is introduced. The dynamic immune threshold is calculated based on the distribution features of the interaction increment to generate an antagonistic mask, suppress homogenization redundant components, and concatenate with the feature residuals of the original node to enhance the heterogeneity discrimination power, and output the intramodal entity augmentation feature tensor. S4. Determine the cross-modal affinity between the intramodal entity enhancement feature matrices of different modalities to construct the cross-modal affinity matrix, and perform normalization iteration based on row and column distribution constraints to output the cross-modal soft alignment matrix; S5. Dynamic sparsification is performed based on the statistical distribution characteristics of the cross-modal soft alignment matrix to construct a cross-modal weighted adjacency matrix, and the intra-modal adjacency matrix and the intra-modal entity enhancement feature matrix are combined to construct a cross-modal fusion graph. S6. Based on the cross-modal fusion graph and the cross-modal weighted adjacency matrix, perform feature weighted aggregation on the associated heteromodal nodes to generate cross-modal neighborhood interaction features, and combine them with the intramodal entity enhancement features corresponding to the current node for fusion and update, outputting a deep fusion cross-modal entity feature matrix; S7. Perform global feature aggregation on the deep fusion cross-modal entity feature matrix to generate a global cross-modal fusion vector, and perform classification mapping to output the decision result.
[0006] Optionally, S1 specifically includes: S11. Perform modal analysis and boundary detection on multi-source heterogeneous modal data to extract discrete entity units inside each modality and topological connection edges between entities; S12. Extract the spatial coordinates of the endpoints of the topological connection edges into the associated structure positions and summarize them into a heterogeneous position set. Map the topological connection edges into the adjacency relationships between heterogeneous entities and output them as a topological adjacency matrix. S13. Define discrete entity units as heterogeneous entities, extract the original modal attribute parameters of each heterogeneous entity corresponding to the heterogeneous location set, perform vectorization transformation on the original modal attribute parameters to generate initial attribute vectors, statistically analyze the dimensional distribution characteristics of the initial attribute vectors in each modality to dynamically generate cross-modal target alignment dimensions, perform linear projection transformation on the initial attribute vectors based on the target alignment dimensions to generate same-dimensional feature vectors, and perform dynamic distribution normalization based on the statistical mean and variance of the same-dimensional feature vectors in each modality to output heterogeneous feature matrices.
[0007] Optionally, S2 specifically includes: S21. Calculate the spatial distance between entity pairs within a modality based on the heterogeneous location set to generate spatial proximity, statistically analyze the distribution characteristics of spatial proximity to dynamically calculate the topological association determination threshold, perform binarization filtering on the spatial proximity based on the topological association determination threshold, and generate the spatial topology matrix within the modality. S22. Calculate the feature interaction covariance between entity pairs within a modality based on the heterogeneous feature matrix to generate feature correlation degree. Statistically calculate the global distribution features of feature correlation degree to dynamically calculate the semantic correlation determination threshold. Based on the semantic correlation determination threshold, perform mask filtering to retain continuous values on the feature correlation degree to generate a semantic correlation matrix within the modality. S23. Perform a logical AND operation on the modal spatial topology matrix and the modal semantic association matrix to generate a modal topology semantic adjacency matrix to determine the existence of node edges. Directly extract the element values of the corresponding edge existence in the modal semantic association matrix to determine the edge weights. Using heterogeneous entities as graph nodes, determine the row vectors of the heterogeneous feature matrix as node features, and combine them to construct the output modal semantic graph.
[0008] Optionally, the improved EdgeConv model includes a neighborhood query layer, a feature difference interaction layer, a nonlinear mapping layer, a context aggregation layer, and a homogeneity suppression calibration layer: The neighborhood query layer is used to locate the current node and its corresponding set of neighboring nodes based on the intramodal semantic graph and the intramodal adjacency matrix, and to extract the feature tensor of the current node and the feature tensor of the neighboring nodes. The feature difference interaction layer is used to calculate the feature difference between the feature tensor of the neighboring nodes and the feature tensor of the current node to obtain the relative feature tensor of the neighboring nodes, and then concatenate the relative feature tensor of the neighboring nodes and the feature tensor of the current node along the feature dimension to output the interactively concatenated feature tensor. The nonlinear mapping layer is used to input the interactive splicing feature tensor into the multilayer perceptron to perform nonlinear feature transformation, and then process the transformed tensor through the activation function to output the neighborhood interactive edge feature tensor. The context aggregation layer is used to perform aggregation operations on all neighborhood interaction edge feature tensors of the current node, concatenate the aggregation result with the feature tensor of the current node using residuals, and output the in-modality entity enhancement feature tensor of the current node through normalization processing. The homogeneous inhibition calibration layer is used to introduce a characteristic homogeneous immune antagonistic mechanism in cellular immunology, and the specific execution process includes: Calculate the feature difference between the entity enhancement feature tensor within the current node modality and the current node feature tensor extracted by the neighborhood query layer, and use it as the interaction increment tensor; The distribution characteristics of the change magnitude of the statistical interaction increment tensor across each feature dimension are analyzed. The dynamic immune threshold is calculated based on the mean and standard deviation of the distribution characteristics. Feature dimensions with change magnitudes higher than the dynamic immune threshold are identified as homogeneous redundant components caused by excessive interaction. A redundant feature mask is generated based on the determination results. The redundant feature mask is then multiplied element-wise with the entity enhancement feature tensor within the current node modality to filter out homogeneous redundant components and obtain the antagonistic purification feature tensor. The heterogeneity discrimination features retained in the antagonistic purification feature tensor are extracted, and the heterogeneity discrimination features are concatenated with the current node feature tensor extracted by the neighborhood query layer to restore and enhance the unique discriminative power of the node itself. Finally, the intramodal entity enhancement feature tensor calibrated by homogeneity suppression is output.
[0009] Optionally, S4 specifically includes: S41. Reshape the intramodal entity enhancement feature tensor into an intramodal entity enhancement matrix. Calculate the feature cosine distance between heterogeneous entity pairs based on the intramodal entity enhancement matrices of different modalities to generate an affinity matrix. Statistically calculate the distribution characteristics of non-zero elements in the affinity matrix to dynamically calculate the affinity filtering threshold. Perform mask filtering on the affinity matrix based on the affinity filtering threshold to retain continuous values, setting elements below the threshold to zero and retaining the original values of elements above the threshold, and then updating the affinity matrix. S42. Calculate the row smoothing coefficient dynamically based on the row distribution characteristics of each non-zero row of the affinity matrix, and calculate the column smoothing coefficient dynamically based on the column distribution characteristics of each non-zero column. Perform nonlinear activation and scaling on the affinity matrix based on the row smoothing coefficient and column smoothing coefficient, and fix the zero elements to exclude them from the transformation to generate the initial normalized affinity matrix. S43. Perform row-column alternating normalization iteration on the initial normalized affinity matrix until it converges to a non-zero subspace double random matrix, and output the cross-modal soft alignment matrix.
[0010] Optionally, S5 specifically includes: S51. Calculate the cross-modal sparsity threshold dynamically based on the global element distribution characteristics of the cross-modal soft alignment matrix, and perform mask filtering to retain continuous values on the cross-modal soft alignment matrix based on the cross-modal sparsity threshold to generate a cross-modal weighted adjacency matrix. S52. Extract the intramodal spatial topology matrix and intramodal semantic association matrix from the intramodal semantic graphs of different modalities. Perform non-zero element binarization on the cross-modal weighted adjacency matrix to extract the cross-modal topological support matrix. Perform block-diagonal combination of the intramodal spatial topology matrix and the cross-modal topological support matrix to generate the global topology matrix. Perform block-diagonal combination of the intramodal semantic association matrix and the cross-modal weighted adjacency matrix to generate the global association attribute matrix. S53. Extract the row vectors of the intra-modal entity enhancement feature matrix of different modalities and concatenate them to generate a global node feature matrix. Perform element-wise multiplication of the global topology matrix and the global association attribute matrix to generate a global weighted adjacency matrix. Determine the existence of node edges based on the position of non-zero elements in the global weighted adjacency matrix, and directly extract the values of non-zero elements to determine the edge weights. Combine the global node feature matrix to construct and output a cross-modal fusion graph.
[0011] Optionally, S6 specifically includes: S61. Based on the global weighted adjacency matrix in the cross-modal fusion graph, extract the source node features and target node features corresponding to the connecting edges, and use the non-zero element values in the global weighted adjacency matrix as scaling factors to perform weighted aggregation on the source node features to generate cross-modal neighborhood interaction features of the target node. S62. Read the entity enhancement feature matrix within the modality, extract the feature vectors corresponding to each node in its original modality as intrinsic features within the modality, and dynamically calculate the modal balance coefficient by statistically analyzing the distribution characteristics of cross-modal neighborhood interaction features and intrinsic features within the modality. S63. Based on the modal balance coefficient, perform element-wise weighted summation on the cross-modal neighborhood interaction features and the intrinsic features within the modality to generate node-level cross-modal fusion features. Then, stitch together the node-level cross-modal fusion features of all nodes in the cross-modal fusion graph to output a deep fusion cross-modal entity feature matrix.
[0012] Optionally, S7 specifically includes: S71. Based on the deep fusion cross-modal entity feature matrix, the energy distribution characteristics of each node feature vector are statistically analyzed to dynamically calculate the feature saliency weight. Based on the feature saliency weight, the deep fusion cross-modal entity feature matrix is weighted and aggregated to output the global cross-modal fusion vector. S72. Calculate the normalized scaling factor dynamically based on the numerical distribution characteristics of the global cross-modal fusion vector, and perform affine transformation and nonlinear activation on the global cross-modal fusion vector based on the normalized scaling factor to generate global decision features. S73. Extract global decision features, perform linear mapping to output classification confidence, and determine the decision result based on the classification confidence.
[0013] An intelligent fusion processing system based on cross-modal semantic alignment according to an embodiment of the present invention includes the following modules: The multi-source heterogeneous data entity deconstruction and mapping module is used to deconstruct multi-source heterogeneous modal data, extract heterogeneous entities and associated structure locations, map heterogeneous entities to a unified feature space, output heterogeneous feature matrix, and use associated structure locations as heterogeneous location set; The intramodal topological and semantic association graph building module is used to determine the intramodal topological associations and semantic associations based on heterogeneous location sets and heterogeneous feature matrices to construct an intramodal adjacency matrix, and to construct an intramodal semantic graph by combining the heterogeneous feature matrix. The homogenization immune antagonism feature enhancement module is used to perform neighborhood interaction and context aggregation based on feature differences on the semantic graph within the modality, based on the improved EdgeConv model. It introduces a feature homogenization immune antagonism mechanism, calculates a dynamic immune threshold based on the distribution features of the interaction increment to generate an antagonistic mask, suppresses homogenization redundancy components, and concatenates them with the feature residuals of the original nodes to enhance the heterogeneity discrimination power, and outputs an intramodal entity enhancement feature tensor. The cross-modal affinity calculation and soft alignment module is used to determine the cross-modal affinity between intra-modal entity enhancement feature matrices of different modalities to construct a cross-modal affinity matrix, and to perform normalization iteration based on row and column distribution constraints to output a cross-modal soft alignment matrix; The cross-modal dynamic sparsification and fusion graph construction module is used to perform dynamic sparsification based on the statistical distribution characteristics of the cross-modal soft alignment matrix to construct a cross-modal weighted adjacency matrix, and to construct a cross-modal fusion graph by combining the intra-modal adjacency matrix and the intra-modal entity enhancement feature matrix. The cross-modal neighborhood interaction and feature fusion update module is used to perform feature weighted aggregation on associated heteromodal nodes based on the cross-modal fusion graph and the cross-modal weighted adjacency matrix to generate cross-modal neighborhood interaction features, and to perform fusion update by combining the intramodal entity enhancement features corresponding to the current node, and output a deep fusion cross-modal entity feature matrix. The global feature aggregation and classification decision module is used to perform global feature aggregation on the deep fusion cross-modal entity feature matrix to generate a global cross-modal fusion vector, and perform classification mapping to output the decision result.
[0014] The beneficial effects of this invention are: This invention significantly improves the heterogeneity discrimination power of feature representation and the accuracy of cross-modal semantic alignment during the cross-modal fusion of multi-source heterogeneous data by introducing a feature homogenization immune antagonism mechanism based on an improved EdgeConv model and a cross-modal dynamic double-random soft alignment strategy. By innovatively constructing an antagonistic mask based on interactive incremental distribution features to calculate the dynamic immune threshold, it effectively suppresses homogenization redundancy components in the graph aggregation process and strengthens heterogeneity discrimination power by concatenating it with the feature residuals of the original nodes, thus achieving a unified constraint on intra-modal feature enhancement and resistance to homogenization decay. Meanwhile, a cross-modal soft alignment mechanism based on affinity distribution dynamic sparsification and alternating row and column normalization iteration is adopted. Zero elements are fixed and excluded from nonlinear transformations, ensuring precise convergence to a non-zero subspace double random matrix. This completely eliminates cross-modal pseudo-correlation and noise interference. Furthermore, in the cross-modal fusion update, modal balance coefficients are dynamically calculated based on distribution characteristics, ensuring adaptive weight fusion of intrinsic features within a modality and cross-modal interaction features. This significantly improves the accuracy, robustness, and reliability of deep semantic interactions in cross-modal fusion decisions. Attached Figure Description
[0015] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is an overall flowchart of an intelligent fusion processing method based on cross-modal semantic alignment proposed in this invention; Figure 2 This is a flowchart illustrating the working principle of the improved EdgeConv model, which is based on a cross-modal semantic alignment intelligent fusion processing method proposed in this invention. Figure 3 This is a schematic diagram of the structure of an intelligent fusion processing system based on cross-modal semantic alignment proposed in this invention. Detailed Implementation
[0016] The invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0017] refer to Figure 1 and Figure 2 A smart fusion processing method based on cross-modal semantic alignment includes the following steps: S1. Perform entity deconstruction on multi-source heterogeneous modal data to extract heterogeneous entities and associated structural locations, map heterogeneous entities to a unified feature space to output a heterogeneous feature matrix, and use the associated structural locations as the heterogeneous location set. S2. Based on the heterogeneous location set and heterogeneous feature matrix, determine the topological association and semantic association within the modality to construct the intramodal adjacency matrix, and combine the heterogeneous feature matrix to construct the intramodal semantic graph; S3. Based on the improved EdgeConv model, neighborhood interaction and context aggregation based on feature differences are performed on the intramodal semantic graph. A feature homogenization immune antagonism mechanism is introduced. The dynamic immune threshold is calculated based on the distribution features of the interaction increment to generate an antagonistic mask, suppress homogenization redundant components, and concatenate with the feature residuals of the original node to enhance the heterogeneity discrimination power, and output the intramodal entity augmentation feature tensor. S4. Determine the cross-modal affinity between the intramodal entity enhancement feature matrices of different modalities to construct the cross-modal affinity matrix, and perform normalization iteration based on row and column distribution constraints to output the cross-modal soft alignment matrix; S5. Dynamic sparsification is performed based on the statistical distribution characteristics of the cross-modal soft alignment matrix to construct a cross-modal weighted adjacency matrix, and the intra-modal adjacency matrix and the intra-modal entity enhancement feature matrix are combined to construct a cross-modal fusion graph. S6. Based on the cross-modal fusion graph and the cross-modal weighted adjacency matrix, perform feature weighted aggregation on the associated heteromodal nodes to generate cross-modal neighborhood interaction features, and combine them with the intramodal entity enhancement features corresponding to the current node for fusion and update, outputting a deep fusion cross-modal entity feature matrix; S7. Perform global feature aggregation on the deep fusion cross-modal entity feature matrix to generate a global cross-modal fusion vector, and perform classification mapping to output the decision result.
[0018] In this embodiment, S1 specifically includes: S11. Read the data stream of multi-source heterogeneous modal data, set the sliding window length to 256 sampling units and the step size to 64 sampling units to perform boundary detection, and extract the discrete entity units and their spatial coordinates inside each modality; for a single modality, calculate the Euclidean distance between the spatial coordinates of adjacent discrete entity units, set the distance threshold to 50.0 distance units, and determine that there is a topological connection edge when the Euclidean distance is less than 50.0 distance units, and output the discrete entity unit, spatial coordinates and the topological connection edge within the modality.
[0019] S12. Summarize the spatial coordinates of discrete entity units in each modality and output them as a heterogeneous location set. Based on the index of the discrete entity units, construct a zero matrix with a row and column size equal to the number of entities as a candidate topological adjacency matrix. Map the topological connection edges in the modality to the elements of the corresponding row and column positions and their transpose positions in the candidate topological adjacency matrix, set the value of the element to 1, and output the topological adjacency matrix.
[0020] S13. Define discrete entity units as heterogeneous entities, extract the original modal attribute parameters of each heterogeneous entity corresponding to the heterogeneous location set, perform vectorization transformation on the original modal attribute parameters to generate initial attribute vectors, count the dimension values of the initial attribute vectors in each modality, calculate the minimum value of all modal dimension values and round down as the cross-modal target alignment dimension, perform linear projection dimension reduction transformation on the initial attribute vectors of each modality based on the cross-modal target alignment dimension to generate same-dimensional feature vectors, calculate the statistical mean and standard deviation of the same-dimensional feature vectors in each modality, subtract the statistical mean from the same-dimensional feature vectors and divide by the standard deviation to perform dynamic distribution normalization, and output the heterogeneous feature matrix.
[0021] In this embodiment, S2 specifically includes: S21. Read the heterogeneous location set output from the previous step, calculate the Euclidean distance between spatial coordinates of entity pairs within the modality, set the reciprocal of the Euclidean distance as the spatial proximity, calculate the arithmetic mean and standard deviation of the spatial proximity of all entity pairs, add 1.5 times the standard deviation to the arithmetic mean and set it as the upper limit threshold of spatial redundancy, subtract 1.0 times the standard deviation from the arithmetic mean and set it as the lower limit threshold of spatial isolation, set the element with spatial proximity greater than the upper limit threshold of spatial redundancy to the value 0 to remove redundancy that is too close in space, set the element with spatial proximity less than the lower limit threshold of spatial isolation to the value 0 to remove isolated points that are too far in space, set the element with spatial proximity greater than or equal to the lower limit threshold of spatial isolation and less than or equal to the upper limit threshold of spatial redundancy to the value 1 to retain a moderate spatial long-distance dependency, and generate the spatial topology matrix within the modality.
[0022] S22. Read the heterogeneous feature matrix output from the previous step, calculate the cosine similarity of the feature row vectors between entity pairs within the modality as the feature correlation degree, calculate the arithmetic mean and standard deviation of the feature correlation degree of all entity pairs, subtract 1.0 times the standard deviation from the arithmetic mean and set it as the lower bound threshold of semantic correlation, set the element with feature correlation degree less than the lower bound threshold of semantic correlation to the value 0 to filter out weak semantic correlation, and retain the original continuous value of feature correlation degree for the element with feature correlation degree greater than or equal to the lower bound threshold of semantic correlation to retain strong semantic intensity, and generate the semantic correlation matrix within the modality.
[0023] S23. Perform an element-wise logical AND operation on the modal spatial topology matrix and the modal semantic association matrix. Output the value 1 when the element of the modal spatial topology matrix is 1 and the element of the modal semantic association matrix is non-zero; otherwise, output the value 0. Generate a modal topology semantic adjacency matrix to determine the existence of node edges. Extract the continuous values of the original feature correlation degree corresponding to the position of the modal topology semantic adjacency matrix with the value of 1 in the modal semantic association matrix and determine them as edge weights. Using heterogeneous entities as graph nodes, determine the row vectors of the heterogeneous feature matrix as node features and combine them to construct the output modal semantic graph.
[0024] In this embodiment, the improved EdgeConv model includes a neighborhood query layer, a feature difference interaction layer, a nonlinear mapping layer, a context aggregation layer, and a homogeneity suppression calibration layer: The neighborhood query layer is used to locate the current node and its corresponding set of neighboring nodes based on the intramodal semantic graph and the intramodal adjacency matrix, and to extract the feature tensor of the current node and the feature tensor of the neighboring nodes. The feature difference interaction layer is used to calculate the feature difference between the feature tensor of the neighboring nodes and the feature tensor of the current node to obtain the relative feature tensor of the neighborhood. The relative feature tensor of the neighborhood and the feature tensor of the current node are then concatenated along the feature dimension to output the interactively concatenated feature tensor. The nonlinear mapping layer is used to input the interactive concatenation feature tensor into the multilayer perceptron to perform nonlinear feature transformation, and then process the transformed tensor through the activation function to output the neighborhood interactive edge feature tensor. The context aggregation layer is used to perform aggregation operations on all neighborhood interaction edge feature tensors of the current node, concatenate the aggregation result with the feature tensor of the current node using residuals, and output the in-modality entity augmentation feature tensor of the current node through normalization processing. The homogeneity inhibition calibration layer is used to introduce characteristic homogeneous immune antagonistic mechanisms of cellular immunology. The specific execution process includes: Calculate the feature difference between the entity enhancement feature tensor within the current node modality and the current node feature tensor extracted by the neighborhood query layer, and use it as the interaction increment tensor; The distribution characteristics of the change magnitude of the statistical interaction increment tensor across each feature dimension are analyzed. The dynamic immune threshold is calculated based on the mean and standard deviation of the distribution characteristics. Feature dimensions with change magnitudes higher than the dynamic immune threshold are identified as homogeneous redundant components caused by excessive interaction. A redundant feature mask is generated based on the determination results. The redundant feature mask is then multiplied element-wise with the entity enhancement feature tensor within the current node modality to filter out homogeneous redundant components and obtain the antagonistic purification feature tensor. The heterogeneity discrimination features retained in the antagonistic purification feature tensor are extracted, and the heterogeneity discrimination features are concatenated with the current node feature tensor extracted by the neighborhood query layer to restore and enhance the unique discriminative power of the node itself. Finally, the intramodal entity enhancement feature tensor calibrated by homogeneity suppression is output.
[0025] The improved EdgeConv model proposed in this step is similar to the traditional EdgeConv model in that it is based on the spatial domain message passing and feature transformation theory of graph neural networks. That is, by locating the central node and its neighborhood set, the feature differences between the neighborhood and the central node are calculated for interactive representation, nonlinear feature mapping is performed using multilayer perceptron and activation function, and the enhanced feature representation of the node is output through residual connection and normalization operation.
[0026] The difference lies in that this invention breaks through the limitations of the traditional EdgeConv model, which is prone to over-smoothing and feature homogenization due to pure feature aggregation. It adds a homogenization suppression calibration layer to introduce a feature homogenization immune antagonism mechanism. It calculates the interaction increment tensor between the aggregated features and the original features, statistically calculates the distribution features, dynamically calculates the immune threshold, accurately identifies dimensions above the threshold as homogenized redundant components, generates an antagonistic mask for element-by-element filtering and purification, and finally concatenates the retained heterogeneity discrimination features with the residuals of the original features, rather than directly outputting the context aggregation residual results.
[0027] The beneficial effects of the improvements are that, by using dynamic immune antagonistic masks and residual splicing calibration, the feature heterogeneity preservation constraint is hard-embedded into the forward propagation of graph convolution. This breaks the limitation of traditional methods where node features converge and lose unique discriminative power after multi-layer interactions, and achieves a precise conversion from indiscriminate aggregation to redundant immune purification. This design significantly enhances the defense against homogenization decay caused by topological information overload, can accurately remove homogenization noise in high-dimensional feature space, and, combined with heterogeneous feature enhancement splicing, effectively improves the model's sensitivity to local structural differences and the robustness of cross-modal entity feature discrimination.
[0028] In this embodiment, S4 specifically includes: S41. Reshape the intramodal entity enhancement feature tensor into an intramodal entity enhancement feature matrix. Calculate the feature cosine distance between heterogeneous entity pairs based on the intramodal entity enhancement feature matrices of different modalities to generate an affinity matrix. Statistically calculate the affinity filtering threshold based on the distribution characteristics of non-zero elements in the affinity matrix. Perform mask filtering on the affinity matrix based on the affinity filtering threshold, setting elements below the threshold to zero and retaining the original value of elements above the threshold, and update the affinity matrix. Decompose the graph topology dimension of the previously output intramodal entity enhancement feature tensor and flatten the graph nodes into independent entity rows. Reconstruct the intramodal entity enhancement feature matrices of modality A and modality B according to the feature dimension. Calculate the cosine similarity between entity feature vectors of modality A and modality B pairwise to generate an affinity matrix. Statistically calculate the mean and standard deviation of the non-zero elements in the affinity matrix. Use the mean plus 1.5 times the standard deviation as the affinity filtering threshold, setting elements below the threshold to zero and retaining the original value of elements not below the threshold. Perform mask filtering and output the updated affinity matrix.
[0029] S42. Calculate the row smoothing coefficient dynamically based on the intra-row distribution characteristics of each non-zero row in the affinity matrix, and calculate the column smoothing coefficient dynamically based on the intra-column distribution characteristics of each non-zero column. Perform nonlinear activation and scaling on the affinity matrix based on the row and column smoothing coefficients, fixing zero elements and excluding them from the transformation, to generate an initial normalized affinity matrix. Read the updated affinity matrix from the previous output, calculate the mean of each non-zero row and column, divide by the standard deviation, and add a zero constant of 0.001, using these as the row and column smoothing coefficients respectively. For each non-zero element, calculate the product of its row and column smoothing coefficients as a scaling factor. Multiply the original value by the scaling factor and input it into a Logistic function with a gain of 2.0 and a midpoint of 0.5. Multiply the output result by the original value to complete the scaling calibration, keeping zero elements fixed at 0 and excluding them from the transformation. Integrate to generate the initial normalized affinity matrix.
[0030] S43. Perform alternating row and column normalization iterations on the initial normalized affinity matrix until it converges to a non-zero subspace double random matrix, and output the cross-modal soft alignment matrix; read the initial normalized affinity matrix output from the previous step, extract the positions of zero elements to generate a zero element mask, and perform alternating row and column non-zero element normalization. After each iteration, use the zero element mask to force the original zero element positions to be set to zero to eliminate errors, so that all non-zero elements in the matrix are converted into non-negative probability values and satisfy the probability transition constraint that the row sum is 1 and the column sum is 1. Calculate the sum of the absolute values of the differences between the current and previous round matrices. Stop the iteration when the sum is less than the convergence threshold of 0.001, and output the converged non-zero subspace double random matrix that satisfies the probability transition characteristics as the cross-modal soft alignment matrix.
[0031] The cross-modal affinity calculation and soft alignment process proposed in this step is similar to the traditional cross-modal feature alignment process in that it is based on the cross-modal feature similarity measurement and probability mapping theory. That is, by calculating the distance or similarity between features of different modal entities, an affinity matrix is constructed, and a normalization constraint is applied to the affinity matrix to transform it into an alignment matrix that satisfies the probability distribution characteristics, thereby establishing a soft matching mapping relationship between cross-modal entities.
[0032] The difference lies in that this invention breaks through the limitation of traditional global Softmax normalization, which forcibly assigns non-zero probabilities, leading to spurious correlation noise. It adds dynamic masking and zero-element immune calibration steps, statistically calculates affinity filtering thresholds based on the distribution of non-zero elements, performs continuous value masking to absolutely zero low-confidence elements, and performs nonlinear activation and scaling under the guidance of smoothing coefficients dynamically calculated from the distribution characteristics within rows and columns, fixing and excluding zero elements from the transformation. Finally, it performs alternating row and column normalization iterations on the initial normalized affinity matrix to converge to a non-zero subspace double random matrix, rather than a single global Softmax or linear normalization prediction.
[0033] The beneficial effects of the improvements are that, through dynamic mask filtering and zero-element immune calibration, the double random constraints of the non-zero subspace are forcibly embedded into the iterative solution of the alignment matrix. This breaks the limitations of traditional methods that cause cross-modal pseudo-associations and semantic drift due to the forced activation of irrelevant entities, and realizes the transformation from global rigid probability allocation to local elastic precise alignment. This design significantly enhances the defense capability against cross-modal heterogeneous noise and weak correlation perturbations, and can accurately converge the true matching probability in the non-zero subspace. Combined with row and column bidirectional constraints, it effectively improves the purity of cross-modal soft alignment and the absolute reliability of semantic mapping.
[0034] In this embodiment, S5 specifically includes: S51. Calculate the cross-modal sparsity threshold dynamically based on the global element distribution characteristics of the cross-modal soft alignment matrix. Perform mask filtering to retain continuous values on the cross-modal soft alignment matrix based on the cross-modal sparsity threshold to generate a cross-modal weighted adjacency matrix. Read the cross-modal soft alignment matrix output from the previous step, calculate the mean and standard deviation of all non-zero elements, and use the mean minus 0.5 times the standard deviation as the cross-modal sparsity threshold. Traverse the cross-modal soft alignment matrix, set elements below the threshold to zero, and retain the original values of elements not below the threshold. Perform mask filtering to output the cross-modal weighted adjacency matrix.
[0035] S52. Extract the intramodal spatial topology matrix and intramodal semantic association matrix from the intramodal semantic graphs of different modalities. Perform non-zero element binarization on the cross-modal weighted adjacency matrix to extract the cross-modal topological support matrix. Perform block-diagonal combination of the intramodal spatial topology matrix and the cross-modal topological support matrix to generate the global topology matrix. Perform block-diagonal combination of the intramodal semantic association matrix and the cross-modal weighted adjacency matrix to generate the global association attribute matrix. Read the cross-modal weighted adjacency matrix output from the previous step and replace its non-zero elements with the value 1. Zero elements are kept at 0. The cross-modal topological support matrix is extracted. The intramodal spatial topological matrix and intramodal semantic association matrix in the previously constructed intramodal semantic graph are read. The intramodal spatial topological matrices of modality A and modality B are placed in the diagonal position, and the cross-modal topological support matrix is placed in the off-diagonal position. The global topological structure matrix is generated by combining the blocks diagonally. The intramodal semantic association matrix of modality A and modality B is placed in the diagonal position, and the cross-modal weighted adjacency matrix is placed in the off-diagonal position. The global association attribute matrix is generated by combining the blocks diagonally.
[0036] S53. Extract the row vectors of the intra-modal entity enhancement feature matrices of different modalities and concatenate them to generate a global node feature matrix. Perform element-wise multiplication of the global topology matrix and the global association attribute matrix to filter association attributes using topological connectivity gating to generate a global weighted adjacency matrix. Determine the existence of node edges based on the positions of non-zero elements in the global weighted adjacency matrix and directly extract the values of non-zero elements to determine the edge weights. Combine the global node feature matrix to construct the output cross-modal fusion graph. Read the intra-modal entity enhancement feature matrices of modal A and modal B from the previous output and concatenate them in row order to generate a global node feature matrix. Read the global topology matrix and the global association attribute matrix from the previous output and perform element-wise multiplication. Use the connectivity of the global topology matrix to shield unreachable paths and retain only the association weights with topological support to generate a global weighted adjacency matrix. Use the positions of non-zero elements in this matrix as edges and the values as edge weights. Combine the node attributes of the global node feature matrix to construct the output cross-modal fusion graph.
[0037] In this embodiment, S6 specifically includes: S61. Based on the global weighted adjacency matrix in the cross-modal fusion graph, extract the source node features and target node features corresponding to the edges. Use the non-zero element values in the global weighted adjacency matrix as scaling factors to perform weighted aggregation on the source node features to generate the cross-modal neighborhood interaction features of the target nodes. Read the global weighted adjacency matrix and global node feature matrix in the cross-modal fusion graph output by the previous step. Locate the source node and target node of the edge according to the non-zero elements in the global weighted adjacency matrix. Extract the source node feature vector from the global node feature matrix. Use the non-zero element values as scaling factors to perform weighted summation on the source node feature vector. Assign the aggregation result to the target node to generate the cross-modal neighborhood interaction features of each target node.
[0038] S62. Read the entity enhancement feature matrix within the modality and extract the feature vectors corresponding to each node in its original modality as intrinsic features within the modality. Statistically calculate the distribution characteristics of cross-modal neighborhood interaction features and intrinsic features within the modality to dynamically calculate the modal balance coefficient. Read the entity enhancement feature matrices within the modality of modality A and modality B before the pre-splicing. Extract the feature vectors corresponding to each node in its original modality according to the node splicing order as intrinsic features within the modality. Calculate the L2 norm of the cross-modal neighborhood interaction features and the L2 norm of the intrinsic features output by the pre-splicing. Divide the L2 norm of the interaction features by the sum of the L2 norm of the interaction features and the L2 norm of the intrinsic features, add 0.001 (a zero constant), and use the result as the modal balance coefficient.
[0039] S63. Based on the modal balance coefficient, perform element-wise weighted summation on the cross-modal neighborhood interaction features and the intrinsic features within the modality to generate node-level cross-modal fusion features. Concatenate the node-level cross-modal fusion features of all nodes in the cross-modal fusion graph to output a deep fusion cross-modal entity feature matrix. Read the modal balance coefficient, cross-modal neighborhood interaction features, and intrinsic features output from the previous step. Multiply the modal balance coefficient by the difference between the cross-modal neighborhood interaction features plus 1 and the modal balance coefficient, multiply by the intrinsic features within the modality, and perform element-wise weighted summation to generate node-level cross-modal fusion features. Concatenate the node-level cross-modal fusion features of all nodes in the global graph in node order to output a deep fusion cross-modal entity feature matrix.
[0040] The cross-modal neighborhood interaction and feature fusion update process proposed in this step is similar to the traditional cross-modal graph feature fusion process in that it is based on graph message passing and feature aggregation theory. That is, by extracting the features of source nodes and target nodes connected by edges in the cross-modal graph, performing weighted aggregation on the features of source nodes using adjacency weights to capture cross-modal neighborhood interaction information, and fusing the interaction features with the original features of the target nodes to output a cross-modal entity representation.
[0041] The difference lies in that this invention breaks away from the limitations of traditional methods that simply concatenate cross-modal aggregated features with original modal features or statically weight them, which leads to the submergence of modal specificity. Instead, it adds a dynamic calibration step for modal balance, extracts the feature vector of the original modality to which the node belongs as the anchor point of the intrinsic feature within the modality, and dynamically calculates the modal balance coefficient by statistically analyzing the distribution characteristics of cross-modal neighborhood interaction features and intrinsic features within the modality. Based on this coefficient, it performs element-wise adaptive weighted summation on both, rather than using a fixed weight allocation or indiscriminate feature superposition.
[0042] The beneficial effects of the improvements are that, by anchoring intrinsic features within a modality and weighting with dynamic balance coefficients, this invention forcibly embeds modality-specific protective constraints into the cross-modal feature fusion process. This breaks through the limitations of traditional methods that are prone to modality swallowing and loss of self-characteristics when heterogeneous information is overloaded, and achieves a precise conversion from indiscriminate mixing to modality-aware adaptive equilibrium. This design significantly enhances the defense capability against cross-modal information redundancy, accurately maintains the original discriminative power of nodes in the feature fusion space, and, combined with the dynamic balance mechanism, effectively improves the robustness of deep fusion of cross-modal entity features in expressing heterogeneous semantic differences and the reliability of cross-modal decision-making.
[0043] In this embodiment, S7 specifically includes: S71. Based on the deep fusion cross-modal entity feature matrix, the energy distribution characteristics of each node's feature vector are statistically analyzed to dynamically calculate the feature saliency weights. Based on the feature saliency weights, the deep fusion cross-modal entity feature matrix is weighted and aggregated to output a global cross-modal fusion vector. The deep fusion cross-modal entity feature matrix output from the previous step is read, and the L2 norm of each node's feature vector is calculated as the energy value. The energy value is normalized using the Softmax function to generate the feature saliency weights of each node. The saliency weights are multiplied element-wise with the corresponding node feature vectors and then summed across all nodes. Weighted aggregation is performed to output a global cross-modal fusion vector.
[0044] S72. Calculate the normalization scaling factor dynamically based on the numerical distribution characteristics of the global cross-modal fusion vector. Perform affine transformation and nonlinear activation on the global cross-modal fusion vector based on the normalization scaling factor to generate global decision features. Read the pre-output global cross-modal fusion vector, calculate the mean and standard deviation of the vector elements, divide the mean by the standard deviation and add 0.001 (excluding zero constant) as the normalization scaling factor. Multiply the global cross-modal fusion vector by the normalization scaling factor and add the mean to complete the affine transformation. Input the affine transformation result into the LeakyReLU activation function and set the slope of the negative interval to 0.01 to output the global decision features.
[0045] S73. Extract global decision features, perform linear mapping to output classification confidence, and determine the decision result based on the classification confidence. Read the global decision features output from the previous step, multiply the global decision features by a weight matrix with the same dimension as the number of categories, add a bias vector, perform linear mapping, output the confidence values of each category, and select the category label with the largest confidence value as the decision result.
[0046] refer to Figure 3 An intelligent fusion processing system based on cross-modal semantic alignment specifically includes the following modules: The multi-source heterogeneous data entity deconstruction and mapping module is used to deconstruct multi-source heterogeneous modal data, extract heterogeneous entities and associated structure locations, map heterogeneous entities to a unified feature space, output heterogeneous feature matrix, and use associated structure locations as heterogeneous location set; The intramodal topological and semantic association graph building module is used to determine the intramodal topological associations and semantic associations based on heterogeneous location sets and heterogeneous feature matrices to construct an intramodal adjacency matrix, and to construct an intramodal semantic graph by combining the heterogeneous feature matrix. The homogenization immune antagonism feature enhancement module is used to perform neighborhood interaction and context aggregation based on feature differences on the semantic graph within the modality, based on the improved EdgeConv model. It introduces a feature homogenization immune antagonism mechanism, calculates a dynamic immune threshold based on the distribution features of the interaction increment to generate an antagonistic mask, suppresses homogenization redundancy components, and concatenates them with the feature residuals of the original nodes to enhance the heterogeneity discrimination power, and outputs an intramodal entity enhancement feature tensor. The cross-modal affinity calculation and soft alignment module is used to determine the cross-modal affinity between intra-modal entity enhancement feature matrices of different modalities to construct a cross-modal affinity matrix, and to perform normalization iteration based on row and column distribution constraints to output a cross-modal soft alignment matrix; The cross-modal dynamic sparsification and fusion graph construction module is used to perform dynamic sparsification based on the statistical distribution characteristics of the cross-modal soft alignment matrix to construct a cross-modal weighted adjacency matrix, and to construct a cross-modal fusion graph by combining the intra-modal adjacency matrix and the intra-modal entity enhancement feature matrix. The cross-modal neighborhood interaction and feature fusion update module is used to perform feature weighted aggregation on associated heteromodal nodes based on the cross-modal fusion graph and the cross-modal weighted adjacency matrix to generate cross-modal neighborhood interaction features, and to perform fusion update by combining the intramodal entity enhancement features corresponding to the current node, and output a deep fusion cross-modal entity feature matrix. The global feature aggregation and classification decision module is used to perform global feature aggregation on the deep fusion cross-modal entity feature matrix to generate a global cross-modal fusion vector, and perform classification mapping to output the decision result.
[0047] Example 1: To verify the feasibility of this invention in intelligent decision-making with multi-source heterogeneous data, the method of this invention was applied to a cross-modal auxiliary diagnostic platform of a smart medical system (hereinafter referred to as "Platform M"). In traditional medical auxiliary diagnostic systems, fully connected networks based on single-modal feature extraction or simple feature splicing and fusion models are usually used for lesion classification and diagnostic decisions. These methods not only ignore the complex spatial topology and semantic relationships within medical data, but also easily lead to feature homogenization decay during graph aggregation, generating a large number of pseudo-connections during cross-modal alignment, which easily results in limited diagnostic accuracy and a high misjudgment rate. To solve the above problems, Platform M decided to adopt the cross-modal fusion decision-making method proposed in this invention, based on an improved EdgeConv and a feature homogenization immune antagonism mechanism.
[0048] During implementation, platform M first utilizes its multimodal data acquisition system to acquire patients' heterogeneous sensory data in real time, including electronic medical record text, CT image sequences, and biochemical test indicators. Through entity deconstruction, extraction of associated structural locations, and unified feature space mapping, a high-quality heterogeneous feature matrix and heterogeneous location set are formed. Simultaneously, platform M accurately constructs intramodal semantic graphs by integrating the topological and semantic relationships within each modality.
[0049] Platform M, through an improved EdgeConv model, performs neighborhood interactions and context aggregation based on feature differences on the intramodal semantic graph. It innovatively introduces a feature homogenization immune antagonism mechanism, calculating a dynamic immune threshold based on the distribution features of interaction increments to generate an antagonistic mask. This effectively suppresses homogenization redundancy and strengthens heterogeneity discrimination by concatenating it with the feature residuals of the original nodes, outputting an intramodal entity enhancement feature matrix. On the other hand, it dynamically calculates affinity filtering thresholds and smoothing coefficients by statistically analyzing the distribution features of the cross-modal affinity matrix, excluding zero elements from the transformation. It then performs alternating row and column normalization iterations until convergence to a non-zero subspace double random matrix, accurately outputting a cross-modal soft alignment matrix and constructing a cross-modal fusion graph. Subsequently, it performs feature weighted aggregation on associated heteromodal nodes and dynamically balances and updates the fusion by combining intrinsic intramodal features. Finally, it performs global aggregation and classification mapping on the deeply fused cross-modal entity feature matrix, effectively improving the accuracy and robustness of cross-modal diagnostic decisions.
[0050] During implementation, the technical team of Platform M discovered that, compared with traditional single-modal and simple splicing fusion methods, the method of this invention significantly improves the accuracy of cross-modal assisted diagnosis and the discriminative power of feature representation. Traditional methods cannot effectively shield the feature homogenization phenomenon in graph aggregation and cross-modal alignment is subject to noise interference. However, the method of this invention effectively achieves deep semantic interaction and heterogeneity preservation of medical multimodal data through an immune antagonism mechanism and a dynamic double-random soft alignment strategy.
[0051] To further verify the actual performance of the method of the present invention, platform M conducted a detailed comparative test between the method of the present invention and the traditional method. The specific performance data is shown in Table 1: Table 1 Performance Comparison of Cross-Modal Auxiliary Diagnostic Methods on Platform M
[0052] As shown in Table 1, the performance of the cross-modal assisted diagnostic platform was comprehensively improved after applying the method of this invention. The diagnostic accuracy increased from 78.5% of the traditional method to 92.3%, and the recall increased from 72.1% to 89.6%, significantly improving the reliability of clinical diagnosis and directly reducing the risk of missed diagnoses. The cross-modal alignment pseudo-connection rate decreased from 18.5% to 4.2%, and the feature homogenization decay decreased from 0.45 to 0.12, demonstrating that the invention effectively eliminated pseudo-association interference and suppressed homogenization redundancy. The time consumed for a single diagnostic inference decreased from 350 milliseconds to 110 milliseconds, and the effective retention rate of multimodal feature fusion dimensions increased from 55.0% to 88.5%, significantly improving the efficiency and quality of feature extraction and fusion. Furthermore, the rare lesion detection rate jumped from 45.2% to 71.8%, and the satisfaction rate of the assisted diagnostic system increased from 68.0% to 94.5%, effectively improving the clinical work experience of doctors and the treatment outcomes for patients.
[0053] Through the method of this invention, platform M successfully realizes accurate cross-modal fusion and intelligent decision-making of multi-source heterogeneous medical data, effectively improving the accuracy of assisted diagnosis and the ability to distinguish heterogeneous features, ensuring the efficiency and safety of medical decision-making, significantly reducing noise interference and information loss in cross-modal feature fusion, enhancing the system's adaptive perception capability for complex heterogeneous data, and providing strong technical support for intelligent operation of smart healthcare.
[0054] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for intelligent fusion processing based on cross-modal semantic alignment, characterized in that, Includes the following steps: S1. Perform entity deconstruction on multi-source heterogeneous modal data to extract heterogeneous entities and associated structural locations, map heterogeneous entities to a unified feature space to output a heterogeneous feature matrix, and use the associated structural locations as the heterogeneous location set. S2. Based on the heterogeneous location set and heterogeneous feature matrix, determine the topological association and semantic association within the modality to construct the intramodal adjacency matrix, and combine the heterogeneous feature matrix to construct the intramodal semantic graph; S3. Based on the improved EdgeConv model, neighborhood interaction and context aggregation based on feature differences are performed on the intramodal semantic graph. A feature homogenization immune antagonism mechanism is introduced. The dynamic immune threshold is calculated based on the distribution features of the interaction increment to generate an antagonistic mask, suppress homogenization redundant components, and concatenate with the feature residuals of the original node to enhance the heterogeneity discrimination power, and output the intramodal entity augmentation feature tensor. S4. Reshape the intramodal entity enhancement feature tensor into an intramodal entity enhancement matrix, calculate cross-modal affinity to construct an affinity matrix, and output the cross-modal soft alignment matrix through row and column double normalization iteration; S5. Dynamic sparsification is performed based on the statistical distribution characteristics of the cross-modal soft alignment matrix to construct a cross-modal weighted adjacency matrix, and the intra-modal adjacency matrix and the intra-modal entity enhancement feature matrix are combined to construct a cross-modal fusion graph. S6. Based on the cross-modal fusion graph and the cross-modal weighted adjacency matrix, perform feature weighted aggregation on the associated heteromodal nodes to generate cross-modal neighborhood interaction features, and combine them with the intramodal entity enhancement features corresponding to the current node for fusion and update, outputting a deep fusion cross-modal entity feature matrix; S7. Perform global feature aggregation on the deep fusion cross-modal entity feature matrix to generate a global cross-modal fusion vector, and perform classification mapping to output the decision result. 2.The intelligent fusion processing method based on cross-modal semantic alignment according to claim 1, characterized in that, S1 includes: S11. Perform modal analysis and boundary detection on multi-source heterogeneous modal data to extract discrete entity units inside each modality and topological connection edges between entities; S12. Extract the spatial coordinates of the endpoints of the topological connection edges into the associated structure positions and summarize them into a heterogeneous position set. Map the topological connection edges into the adjacency relationships between heterogeneous entities and output them as a topological adjacency matrix. S13. Define discrete entity units as heterogeneous entities, extract the original modal attribute parameters of each heterogeneous entity corresponding to the heterogeneous location set, perform vectorization transformation on the original modal attribute parameters to generate initial attribute vectors, statistically analyze the dimensional distribution characteristics of the initial attribute vectors in each modality to dynamically generate cross-modal target alignment dimensions, perform linear projection transformation on the initial attribute vectors based on the target alignment dimensions to generate same-dimensional feature vectors, and perform dynamic distribution normalization based on the statistical mean and variance of the same-dimensional feature vectors in each modality to output heterogeneous feature matrices. 3.The intelligent fusion processing method based on cross-modal semantic alignment according to claim 1, characterized in that, S2 specifically includes: S21. Calculate the spatial distance between entity pairs within a modality based on the heterogeneous location set to generate spatial proximity, statistically analyze the distribution characteristics of spatial proximity to dynamically calculate the topological association determination threshold, perform binarization filtering on the spatial proximity based on the topological association determination threshold, and generate the spatial topology matrix within the modality. S22. Calculate the feature interaction covariance between entity pairs within a modality based on the heterogeneous feature matrix to generate feature correlation degree. Statistically calculate the global distribution features of feature correlation degree to dynamically calculate the semantic correlation determination threshold. Based on the semantic correlation determination threshold, perform mask filtering to retain continuous values on the feature correlation degree to generate a semantic correlation matrix within the modality. S23. Perform a logical AND operation on the modal spatial topology matrix and the modal semantic association matrix to generate a modal topology semantic adjacency matrix to determine the existence of node edges. Directly extract the element values of the corresponding edge existence in the modal semantic association matrix to determine the edge weights. Using heterogeneous entities as graph nodes, determine the row vectors of the heterogeneous feature matrix as node features, and combine them to construct the output modal semantic graph.
4. The intelligent fusion processing method based on cross-modal semantic alignment according to claim 1, characterized in that, The improved EdgeConv model includes a neighborhood query layer, a feature difference interaction layer, a nonlinear mapping layer, a context aggregation layer, and a homogeneity suppression calibration layer. The neighborhood query layer is used to locate the current node and its corresponding set of neighboring nodes based on the intramodal semantic graph and the intramodal adjacency matrix, and to extract the feature tensor of the current node and the feature tensor of the neighboring nodes. The feature difference interaction layer is used to calculate the feature difference between the feature tensor of the neighboring nodes and the feature tensor of the current node to obtain the relative feature tensor of the neighboring nodes, and then concatenate the relative feature tensor of the neighboring nodes and the feature tensor of the current node along the feature dimension to output the interactively concatenated feature tensor. The nonlinear mapping layer is used to input the interactive splicing feature tensor into the multilayer perceptron to perform nonlinear feature transformation, and then process the transformed tensor through the activation function to output the neighborhood interactive edge feature tensor. The context aggregation layer is used to perform aggregation operations on all neighborhood interaction edge feature tensors of the current node, concatenate the aggregation result with the feature tensor of the current node using residuals, and output the in-modality entity enhancement feature tensor of the current node through normalization processing. The homogeneous inhibition calibration layer is used to introduce a characteristic homogeneous immune antagonistic mechanism in cellular immunology, and the specific execution process includes: Calculate the feature difference between the entity enhancement feature tensor within the current node modality and the current node feature tensor extracted by the neighborhood query layer, and use it as the interaction increment tensor; The distribution characteristics of the change magnitude of the statistical interaction increment tensor across each feature dimension are analyzed. The dynamic immune threshold is calculated based on the mean and standard deviation of the distribution characteristics. Feature dimensions with change magnitudes higher than the dynamic immune threshold are identified as homogeneous redundant components caused by excessive interaction. A redundant feature mask is generated based on the determination results. The redundant feature mask is then multiplied element-wise with the entity enhancement feature tensor within the current node modality to filter out homogeneous redundant components and obtain the antagonistic purification feature tensor. The heterogeneity discrimination features retained in the antagonistic purification feature tensor are extracted, and the heterogeneity discrimination features are concatenated with the current node feature tensor extracted by the neighborhood query layer to restore and enhance the unique discriminative power of the node itself. Finally, the intramodal entity enhancement feature tensor calibrated by homogeneity suppression is output.
5. The intelligent fusion processing method based on cross-modal semantic alignment according to claim 1, characterized in that, S4 specifically includes: S41. Reshape the intramodal entity enhancement feature tensor into an intramodal entity enhancement matrix. Calculate the feature cosine distance between heterogeneous entity pairs based on the intramodal entity enhancement matrices of different modalities to generate an affinity matrix. Statistically calculate the distribution characteristics of non-zero elements in the affinity matrix to dynamically calculate the affinity filtering threshold. Perform mask filtering on the affinity matrix based on the affinity filtering threshold to retain continuous values, setting elements below the threshold to zero and retaining the original values of elements above the threshold, and then updating the affinity matrix. S42. Calculate the row smoothing coefficient dynamically based on the row distribution characteristics of each non-zero row of the affinity matrix, and calculate the column smoothing coefficient dynamically based on the column distribution characteristics of each non-zero column. Perform nonlinear activation and scaling on the affinity matrix based on the row smoothing coefficient and column smoothing coefficient, and fix the zero elements to exclude them from the transformation to generate the initial normalized affinity matrix. S43. Perform row-column alternating normalization iteration on the initial normalized affinity matrix until it converges to a non-zero subspace double random matrix, and output the cross-modal soft alignment matrix.
6. The intelligent fusion processing method based on cross-modal semantic alignment according to claim 1, characterized in that, S5 specifically includes: S51. Calculate the cross-modal sparsity threshold dynamically based on the global element distribution characteristics of the cross-modal soft alignment matrix, and perform mask filtering to retain continuous values on the cross-modal soft alignment matrix based on the cross-modal sparsity threshold to generate a cross-modal weighted adjacency matrix. S52. Extract the intramodal spatial topology matrix and intramodal semantic association matrix from the intramodal semantic graphs of different modalities. Perform non-zero element binarization on the cross-modal weighted adjacency matrix to extract the cross-modal topological support matrix. Perform block-diagonal combination of the intramodal spatial topology matrix and the cross-modal topological support matrix to generate the global topology matrix. Perform block-diagonal combination of the intramodal semantic association matrix and the cross-modal weighted adjacency matrix to generate the global association attribute matrix. S53. Extract the row vectors of the intra-modal entity enhancement feature matrix of different modalities and concatenate them to generate a global node feature matrix. Perform element-wise multiplication of the global topology matrix and the global association attribute matrix to generate a global weighted adjacency matrix. Determine the existence of node edges based on the position of non-zero elements in the global weighted adjacency matrix, and directly extract the values of non-zero elements to determine the edge weights. Combine the global node feature matrix to construct and output a cross-modal fusion graph.
7. The intelligent fusion processing method based on cross-modal semantic alignment according to claim 1, characterized in that, S6 specifically includes: S61. Based on the global weighted adjacency matrix in the cross-modal fusion graph, extract the source node features and target node features corresponding to the connecting edges, and use the non-zero element values in the global weighted adjacency matrix as scaling factors to perform weighted aggregation on the source node features to generate cross-modal neighborhood interaction features of the target node. S62. Read the entity enhancement feature matrix within the modality, extract the feature vectors corresponding to each node in its original modality as intrinsic features within the modality, and dynamically calculate the modal balance coefficient by statistically analyzing the distribution characteristics of cross-modal neighborhood interaction features and intrinsic features within the modality. S63. Based on the modal balance coefficient, perform element-wise weighted summation on the cross-modal neighborhood interaction features and the intrinsic features within the modality to generate node-level cross-modal fusion features. Then, stitch together the node-level cross-modal fusion features of all nodes in the cross-modal fusion graph to output a deep fusion cross-modal entity feature matrix.
8. The intelligent fusion processing method based on cross-modal semantic alignment according to claim 1, characterized in that, Specifically, S7 includes: S71. Based on the deep fusion cross-modal entity feature matrix, the energy distribution characteristics of each node feature vector are statistically analyzed to dynamically calculate the feature saliency weight. Based on the feature saliency weight, the deep fusion cross-modal entity feature matrix is weighted and aggregated to output the global cross-modal fusion vector. S72. Calculate the normalized scaling factor dynamically based on the numerical distribution characteristics of the global cross-modal fusion vector, and perform affine transformation and nonlinear activation on the global cross-modal fusion vector based on the normalized scaling factor to generate global decision features. S73. Extract global decision features, perform linear mapping to output classification confidence, and determine the decision result based on the classification confidence.
9. An intelligent fusion processing system based on cross-modal semantic alignment, comprising executing the intelligent fusion processing method based on cross-modal semantic alignment as described in any one of claims 1 to 8, characterized in that, Includes the following modules: The multi-source heterogeneous data entity deconstruction and mapping module is used to deconstruct multi-source heterogeneous modal data, extract heterogeneous entities and associated structure locations, map heterogeneous entities to a unified feature space, output heterogeneous feature matrix, and use associated structure locations as heterogeneous location set; The intramodal topological and semantic association graph building module is used to determine the intramodal topological associations and semantic associations based on heterogeneous location sets and heterogeneous feature matrices to construct an intramodal adjacency matrix, and to construct an intramodal semantic graph by combining the heterogeneous feature matrix. The homogenization immune antagonism feature enhancement module is used to perform neighborhood interaction and context aggregation based on feature differences on the semantic graph within the modality, based on the improved EdgeConv model. It introduces a feature homogenization immune antagonism mechanism, calculates a dynamic immune threshold based on the distribution features of the interaction increment to generate an antagonistic mask, suppresses homogenization redundancy components, and concatenates them with the feature residuals of the original nodes to enhance the heterogeneity discrimination power, and outputs an intramodal entity enhancement feature tensor. The cross-modal affinity calculation and soft alignment module is used to determine the cross-modal affinity between intra-modal entity enhancement feature matrices of different modalities to construct a cross-modal affinity matrix, and to perform normalization iteration based on row and column distribution constraints to output a cross-modal soft alignment matrix; The cross-modal dynamic sparsification and fusion graph construction module is used to perform dynamic sparsification based on the statistical distribution characteristics of the cross-modal soft alignment matrix to construct a cross-modal weighted adjacency matrix, and to construct a cross-modal fusion graph by combining the intra-modal adjacency matrix and the intra-modal entity enhancement feature matrix. The cross-modal neighborhood interaction and feature fusion update module is used to perform feature weighted aggregation on associated heteromodal nodes based on the cross-modal fusion graph and the cross-modal weighted adjacency matrix to generate cross-modal neighborhood interaction features, and to perform fusion update by combining the intramodal entity enhancement features corresponding to the current node, and output a deep fusion cross-modal entity feature matrix. The global feature aggregation and classification decision module is used to perform global feature aggregation on the deep fusion cross-modal entity feature matrix to generate a global cross-modal fusion vector, and perform classification mapping to output the decision result.