Protein network-based multi-modal data fusion analysis method and system

Through the multimodal data fusion analysis method based on protein network, the problems of heterogeneity gap, redundancy, noise and interpretability of multimodal data fusion in the prior art are solved, efficient integration and semantic understanding are achieved, and the robustness and interpretability of data fusion are improved.

CN120126540APending Publication Date: 2025-06-10CHONGQING RUANJIANG TURING ARTIFICIAL INTELLIGENCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510181618.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing multimodal data fusion methods face problems such as heterogeneity divide, redundancy and noise, and lack of interpretability, and fail to make full use of the deep correlation between modes.

Method used

The multimodal data fusion analysis method based on protein network is adopted, and the multimodal data is obtained for preprocessing, the feature vector is mapped as protein network nodes, the protein network topology structure is constructed, and the network topology characteristics are optimized through the dynamic network evolution model, combining modular fusion technology and cross-modal alignment loss function for feature fusion.

Benefits of technology

It realizes efficient integration and semantic understanding of cross-modal data, improves the robustness and interpretability of data fusion, and is suitable for fields such as intelligent medical diagnosis, autonomous driving perception enhancement, and cross-media content understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126540A_ABST
    Figure CN120126540A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal data fusion analysis method and system based on a protein network, and the method comprises the steps: obtaining multi-modal data, and carrying out the preprocessing, including noise filtering, normalization and feature extraction, and obtaining the feature vector of each modal; the feature vectors are mapped into protein network nodes, the connection probability is calculated based on the similarity between the nodes, a protein network topological structure is constructed, and cross-modal interaction can be achieved through the protein network topological structure; optimizing the protein network topological structure through a dynamic network evolution model, and updating network topological characteristics to obtain a protein network; and fusing the features in the protein network in combination with a modular fusion technology and a cross-modal alignment loss function to generate a multi-modal data fusion result. According to the method, efficient integration and semantic understanding of multi-modal data can be realized, and the robustness and interpretability of data fusion are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of biological networks, and in particular, to a multimodal data fusion analysis method and system based on a protein network. Background Art

[0002] Multimodal data fusion refers to integrating data from different modalities (such as images, texts, audios, videos, etc.) to extract richer and more comprehensive information. In the prior art, the following problems exist in multimodal data fusion:

[0003] First, the heterogeneity gap: The data structures and semantic spaces of different modalities, such as images, audios, and texts, are significantly different and difficult to directly align;

[0004] Second, redundancy and noise: Low-quality data leads to a decrease in the credibility of the fusion result;

[0005] Third, lack of interpretability: It is difficult to trace the decision basis for black-box models.

[0006] In addition, traditional multimodal data fusion methods, such as early fusion (feature splicing) and late fusion (decision weighting), do not fully utilize the deep associations between modalities. Recent research attempts to introduce graph neural networks, but they are still limited by fixed topological structures and insufficient biological rationality.

[0007] Therefore, there is an urgent need for a multimodal data fusion analysis method that can achieve efficient integration and semantic understanding of cross-modal data, and improve the robustness and interpretability of data fusion. Summary of the Invention

[0008] Based on this, it is necessary to provide a multimodal data fusion analysis method and system based on a protein network for the above technical problems.

[0009] A multimodal data fusion analysis method based on a protein network includes the following steps: obtaining multimodal data and performing preprocessing, where the preprocessing includes noise filtering, normalization, and feature extraction to obtain feature vectors of each modality; mapping the feature vectors to protein network nodes, calculating connection probabilities based on the similarity between nodes, and constructing a protein network topological structure that can achieve cross-modal interaction; optimizing the protein network topological structure through a dynamic network evolution model, updating network topological characteristics, to obtain a protein network; and fusing the features in the protein network by combining modular fusion technology and a cross-modal alignment loss function to generate a multimodal data fusion result.

[0010] In one embodiment, the feature extraction uses a multi-scale encoder, which is:

[0011]

[0012] Wherein, f i represents the extracted feature, represents feature splicing, and CNN, BiLSTM, and GAT are respectively used to extract local features, temporal features, and graph structure features.

[0013] In one embodiment, the connection probability is calculated based on the similarity between nodes, and the formula is:

[0014]

[0015] Wherein, P ij is the connection probability, is the cosine similarity between nodes i and j, f i and f j respectively represent the features of nodes i and j, and θ(t) is the connection threshold dynamically adjusted with time t, and the adjustment formula is:

[0016] θ(t) = θ 0 + α · (1 - e -βt );

[0017] Wherein, θ 0 is the initial threshold, α and β are attenuation coefficients, τ is the temperature parameter used to control the smoothness of the probability distribution, and σ(·) is the Sigmoid function.

[0018] In one embodiment, the optimization objective of the dynamic network evolution model is:

[0019]

[0020] Wherein, is the global efficiency, N is the total number of nodes, d ij is the shortest path length between nodes i and j, is the modularity, λ 1 and λ 2 are weight coefficients, C is the constraint on the total number of network connections, L represents the total number of edges in the network, l c represents the number of edges within community c, and k c represents the sum of the degrees of all nodes within community c.

[0021] In one embodiment, the cross-modal alignment loss function is:

[0022]

[0023] Wherein, and are respectively the feature embeddings of the visual and auditory modalities, is an indicator function used to force cross-modal feature alignment of similar samples, y i and y j are the outputs of nodes i and j respectively.

[0024] In one embodiment, the modular fusion technology and the cross-modal alignment loss function are combined to fuse the features in the protein network to generate a multi-modal data fusion result, including: dividing the protein network into multiple modules, and determining a set of nodes with similar functions or semantics corresponding to the multiple modules according to the connection probability; using the cross-modal alignment loss function to align and fuse the features within each module to obtain a modular feature representation; based on the modular fusion technology, fusing the modular feature representations to output a multi-modal data fusion result.

[0025] In one embodiment, after generating the multi-modal data fusion result, it further includes: scoring the multi-modal data fusion result to obtain a scoring result, and the formula is:

[0026]

[0027] In the formula, φ k (·) is the k-th evaluation index, and w k is the weight based on information entropy, and the formula is:

[0028]

[0029] In the formula, φ a represents the a-th evaluation index, and K represents the probability occupied by the evaluation index; the multi-modal data fusion result is evaluated according to the scoring result to obtain an evaluation result.

[0030] In one embodiment, it further includes: determining edge weights according to the connection probability, mapping node colors according to centrality indicators, and generating a node-edge heat map, and the formula is:

[0031]

[0032] In the formula, C i is the centrality indicator.

[0033] A multi-modal data fusion analysis system based on a protein network for implementing a multi-modal data fusion analysis method as described above, comprising: a multi-modal data preprocessing module for acquiring multi-modal data and performing preprocessing, the preprocessing including noise filtering, normalization, and feature extraction to obtain feature vectors of each modality; a topological structure construction module for mapping the feature vectors to protein network nodes, calculating connection probabilities based on the similarity between nodes, and constructing a protein network topological structure that can achieve cross-modal interaction; a protein network optimization module for optimizing the protein network topological structure through a dynamic network evolution model, updating network topological characteristics to obtain a protein network; and a fusion result generation module for fusing the features in the protein network by combining modular fusion technology and a cross-modal alignment loss function to generate a multi-modal data fusion result.

[0034] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: By acquiring multi-modal data and performing preprocessing operations such as noise filtering, normalization, and feature extraction, feature vectors of each modality are obtained. The feature vectors are mapped to protein network nodes, connection probabilities are calculated based on the similarity between nodes, and a protein network topological structure is constructed, which can achieve cross-modal interaction and introduce biological network characteristics, helping to improve the robustness and interpretability of data fusion. By optimizing the protein network topological structure through a dynamic network evolution model, updating network topological characteristics to obtain a protein network, the dynamic relationship between data can be captured by simulating the dynamic evolution process of the protein network. By combining modular fusion technology and a cross-modal alignment loss function, the features in the protein network are fused to generate a multi-modal data fusion result, thereby achieving efficient integration and semantic understanding of multi-modal data, and being applicable to fields such as intelligent medical diagnosis, enhanced perception for autonomous driving, and cross-media content understanding. Brief Description of the Drawings

[0035] Figure 1 It is a schematic flowchart of a multi-modal data fusion analysis method based on a protein network in an embodiment;

[0036] Figure 2 It is a schematic flowchart of the processing of multi-modal data in an embodiment;

[0037] Figure 3 It is a schematic flowchart of the construction of a protein network topological structure in an embodiment;

[0038] Figure 4 It is a schematic flowchart of the implementation of an alignment loss and a scoring model in an embodiment;

[0039] Figure 5Schematic diagram of the structure of a multi-modal data fusion analysis system based on a protein network in an embodiment. Detailed implementation manners

[0040] Before describing the detailed implementation manners of the present invention, the overall concept of the present invention will be described as follows:

[0041] The present invention is mainly developed based on the multi-modal data fusion process. At present, there are disadvantages such as heterogeneous gap, redundancy and noise, and lack of interpretability in multi-modal data fusion.

[0042] Therefore, the present invention proposes a multi-modal data fusion analysis method based on a protein network. By obtaining multi-modal data and performing preprocessing operations such as noise filtering, normalization, and feature extraction, feature vectors of each modality are obtained. The feature vectors are mapped to protein network nodes, and the connection probability is calculated based on the similarity between nodes to construct the protein network topology structure. Moreover, it can achieve cross-modal interaction, introduce biological network characteristics, and help improve the robustness and interpretability of data fusion. By optimizing the protein network topology structure through a dynamic network evolution model and updating the network topology characteristics, a protein network is obtained. By simulating the dynamic evolution process of the protein network, the dynamic relationship between data can be captured. Combining modular fusion technology and cross-modal alignment loss function, the features in the protein network are fused to generate the multi-modal data fusion result, thereby realizing the efficient integration and semantic understanding of multi-modal data, and being applicable to fields such as intelligent medical diagnosis, enhanced perception of autonomous driving, and cross-media content understanding.

[0043] After introducing the overall concept of the present invention, in order to make the purpose, technical solution and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below through specific implementation manners in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0044] In one embodiment, as Figure 1 shown, a multi-modal data fusion analysis method based on a protein network is provided, including the following steps:

[0045] Step S110, obtain multi-modal data and perform preprocessing, where the preprocessing includes noise filtering, normalization, and feature extraction to obtain feature vectors of each modality.

[0046] Specifically, obtain multi-modal data, such as visual data, auditory data, and language data, etc., and perform preprocessing operations such as noise filtering, normalization, and feature extraction on the multi-modal data to obtain feature vectors of each modality.

[0047] For data of different modalities, representative features need to be extracted. As Figure 2As shown, for example, for visual data, a convolutional neural network (CNN) can be used to extract the depth features of an image, or ResNet-50 can be used to extract the global features of the image, and SuperPoint can be used to extract local key point descriptors; for auditory data, a recurrent neural network (RNN) can be used to extract the time series features of the audio, or a Mel spectrogram can be input into a Conv1D network to extract the time series spectral features; for language data, natural language processing techniques (such as BERT) can be used to extract the semantic features of the text, word vectors can be obtained through the BERT model, and the syntactic dependencies can be encoded through a graph attention network (GAT).

[0048] Among them, when performing feature extraction, a multi-scale encoder is adopted, and the formula is:

[0049]

[0050] In the formula, f i represents the extracted features, represents feature concatenation, and CNN, BiLSTM, and GAT are respectively used to extract local features, time series features, and graph structure features.

[0051] Specifically, through multi-scale encoding, feature extraction in multiple aspects such as local, time series, and graph structure of multi-modal data can be realized, so as to obtain more effective feature information, which is convenient for subsequent data fusion.

[0052] Step S120: Map the feature vector to a protein network node, calculate the connection probability based on the similarity between nodes, and construct a protein network topology structure, which can realize cross-modal interaction.

[0053] Specifically, after the feature vectors of multi-modal data are extracted, they are mapped to protein network nodes, and a protein network is constructed by using the method of probabilistic connection. The features of each modal data are regarded as a node in the protein network, and the connection relationship between nodes is determined according to the similarity and relevance between the data. The connection probability is calculated for the similarity between any two nodes (i.e., the features of different modal data), and the similarity can be calculated in various ways, such as Euclidean distance, cosine similarity, etc., and the protein network topology structure is constructed according to the obtained connection probability. Cross-modal data interaction can be realized through this protein network topology structure.

[0054] Among them, the formula for calculating the connection probability based on the similarity between nodes is:

[0055]

[0056] In the formula, P ij is the connection probability, is the cosine similarity between nodes i and j, and f i and f j represent the features of nodes i and j respectively, θ(t) is the connection threshold dynamically adjusted with time t, and the adjustment formula is:

[0057] θ(t) = θ 0 + α·(1 - e -βt );

[0058] In the formula, θ 0 is the initial threshold, α and β are decay coefficients, τ is the temperature parameter used to control the smoothness of the probability distribution, and σ(·) is the Sigmoid function.

[0059] Specifically, θ 0 can be set to 0.7, the decay parameter is set as α = 0.2, β = 0.1, to ensure that the network gradually sparsifies as the number of iterations t increases. The connection probability is calculated through the cosine similarity between nodes. Of course, the Euclidean distance or other similarity calculation methods can also be used to obtain the connection probability to realize the construction of the protein network topology structure. Through the above method, a protein network with probabilistic connections can be constructed, where the connection strength between nodes reflects the similarity and correlation between data.

[0060] Step S130, optimize the protein network topology structure through the dynamic network evolution model, update the network topology characteristics, and obtain the protein network.

[0061] Specifically, adopt the dynamic network evolution model, simulate the dynamic evolution process of the protein network, analyze the changes in the network topology characteristics to capture the dynamic relationships between data, so as to realize the optimization of the protein network topology structure and obtain new network topology characteristics, including network global efficiency, modularity, and centrality indicators, etc., to form the protein network.

[0062] By defining the dynamic network evolution model, update the network topology structure according to the current state of the network and external inputs (such as new modality data). For example, when new modality data enters the protein network, the connection relationship can be dynamically adjusted according to its similarity with the existing nodes, and the network topology characteristics are updated to obtain the protein network, so as to realize the modeling and analysis of the dynamic relationships of multi-modal data.

[0063] Among them, the optimization objective of the dynamic network evolution model is:

[0064]

[0065] Among them, is the global efficiency, N is the total number of nodes, d ij is the shortest path length between nodes i and j, is the modularity, λ1 and λ 2 are weight coefficients, C is the constraint on the total number of connections in the network, L represents the total number of edges in the network, and l c represents the number of edges within community c, and k c represents the sum of the degrees of all nodes within community c.

[0066] Specifically, set λ 1 = 0.6, λ 2 = 0.4. Under the constraint that the total number of connections C = 10^4, the gradient ascent method is used to optimize the global efficiency and modularity, as Figure 3 shown. Since proteins have dynamic characteristics in organisms and their topological structures change with time and environment, based on this characteristic, by simulating the dynamic evolution process of protein networks, analyzing the changes in the topological characteristics of protein networks, and thus capturing the dynamic relationships between data. By setting optimization goals, the topological structure of the protein network is optimized until it meets the set standards, and a protein network is obtained.

[0067] Step S140: Combine the modular fusion technology and the cross-modal alignment loss function to fuse the features in the protein network to generate a multi-modal data fusion result.

[0068] Specifically, according to the obtained protein network, combine the modular fusion technology and the cross-modal alignment loss function to fuse the features in the protein network to generate a multi-modal data fusion result, achieving the efficient integration and semantic understanding of multi-modal data. By introducing the characteristics of biological networks, the robustness and interpretability of data fusion are significantly improved.

[0069] Among them, step S140 includes: dividing the protein network into multiple modules, determining a group of nodes with similar functions or semantics corresponding to the multiple modules according to the connection probability; using the cross-modal alignment loss function to align and fuse the features within each module to obtain a modular feature representation; based on the modular fusion technology, fusing the modular feature representations to output a multi-modal data fusion result.

[0070] Specifically, when performing multi-modal data fusion, the protein network is divided into multiple modules, and a set of nodes with similar functions or semantics corresponding to the multiple modules is determined based on the connection probability, that is, the nodes with the highest connection probability corresponding to the module are selected as the nodes with similar functions or semantics. The cross-modal alignment loss function is used to align and fuse the features within each module to obtain a modular feature representation. Finally, based on the modular fusion technology, the modular feature representations are fused to output the multi-modal data fusion result. The feature representations can be fused using weighted summation, and the weights can be dynamically adjusted according to the importance of the modules, so as to achieve the efficient integration and semantic understanding of multi-modal data.

[0071] Among them, the cross-modal alignment loss function is:

[0072]

[0073] In the formula, and are the feature embeddings of the visual and auditory modalities respectively, is an indicator function used to enforce cross-modal feature alignment of similar samples, and y i and y j are the outputs of nodes i and j respectively.

[0074] Specifically, the cross-modal alignment loss function is used to map the features of different modalities to a shared semantic space and ensure the alignment of relevant samples. The cross-modal alignment loss function includes contrastive loss, triplet loss, cosine similarity loss, cross-entropy loss, etc., so as to achieve the alignment and fusion of the features within each module.

[0075] As Figure 4 shown, in medical diagnosis, the CT images (visual) and pathology reports (text) of the same patient are forced to align in the embedding space.

[0076] Among them, after step S140, it further includes: scoring the multi-modal data fusion result to obtain a scoring result, and the formula is:

[0077]

[0078] In the formula, φ k (·) is the k-th evaluation index, and w k is the weight based on information entropy, and the formula is:

[0079]

[0080] In the formula, φ a represents the a-th evaluation index, and K represents the probability occupied by the evaluation index; the multi-modal data fusion result is evaluated according to the scoring result to obtain the optimal fusion result.

[0081] Specifically, after obtaining the multi-modal data fusion result, the fusion result can be scored according to the set evaluation index. The evaluation index can be classification confidence, reconstruction error, etc. The multi-modal data fusion result is evaluated according to the scoring result to obtain the evaluation result, so as to ensure the interpretability of the fusion result.

[0082] In this embodiment, by obtaining multi-modal data and performing preprocessing operations such as noise filtering, normalization, and feature extraction, the feature vectors of each modality are obtained. The feature vectors are mapped to protein network nodes, the connection probability is calculated based on the similarity between nodes, and the protein network topology structure is constructed. Moreover, it can achieve cross-modal interaction, introduce biological network characteristics, and help improve the robustness and interpretability of data fusion; the protein network topology structure is optimized through a dynamic network evolution model, the network topology characteristics are updated to obtain a protein network, and by simulating the dynamic evolution process of the protein network, the dynamic relationship between data can be captured; combining modular fusion technology and cross-modal alignment loss function, the features in the protein network are fused to generate a multi-modal data fusion result, thereby realizing the efficient integration and semantic understanding of multi-modal data, and being applicable to fields such as intelligent medical diagnosis, autonomous driving perception enhancement, and cross-media content understanding.

[0083] In one embodiment, it further includes: determining edge weights according to the connection probability, mapping node colors according to centrality indicators, and generating a node-edge heat map. The formula is:

[0084]

[0085] where C i is the centrality indicator.

[0086] Specifically, the edge weights are determined according to the connection probability, the node colors are mapped according to the centrality indicators, and a node-edge heat map is generated to explain the basis for the fusion decision. The centrality indicators can be indicators of network topology such as degree centrality and betweenness centrality.

[0087] As Figure 5 shown, a multi-modal data fusion analysis system 50 based on a protein network is provided, which is used to implement a multi-modal data fusion analysis method based on a protein network as described above, including: a multi-modal data preprocessing module 51, a topology structure construction module 52, a protein network optimization module 53, and a fusion result generation module 54, where:

[0088] The multi-modal data preprocessing module 51 is used to obtain multi-modal data and perform preprocessing, and the preprocessing includes noise filtering, normalization, and feature extraction to obtain the feature vectors of each modality;

[0089] A topology construction module 52 is used to map feature vectors into protein network nodes, calculate connection probabilities based on the similarity between nodes, and construct a protein network topology, which can achieve cross-modal interaction;

[0090] A protein network optimization module 53 is used to optimize the protein network topology through a dynamic network evolution model, update the network topology characteristics, and obtain a protein network;

[0091] A fusion result generation module 54 is used to combine modular fusion technology and a cross-modal alignment loss function to fuse the features in the protein network and generate a multi-modal data fusion result.

[0092] In one embodiment, the fusion result generation module 54 is specifically used for: dividing the protein network into multiple modules, determining a set of nodes with similar functions or semantics corresponding to the multiple modules according to the connection probabilities; using a cross-modal alignment loss function to align and fuse the features within each module to obtain a modular feature representation; based on modular fusion technology, fusing the modular feature representations and outputting a multi-modal data fusion result.

[0093] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0094] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a computer storage medium (ROM / RAM, magnetic disk, optical disc) and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. Therefore, the present invention is not limited to any specific combination of hardware and software.

[0095] The above content is a further detailed description of the present invention in combination with specific implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A multimodal data fusion analysis method based on protein network, characterized in that: The following steps are involved: Acquire multimodal data and perform preprocessing, wherein the preprocessing includes noise filtering, normalization and feature extraction to obtain a feature vector of each modality; Mapping the feature vectors into protein network nodes, calculating connection probabilities based on similarities between nodes and constructing a protein network topology structure, wherein the protein network topology structure can achieve cross-modal interaction; Optimizing the protein network topology structure through a dynamic network evolution model, updating the network topology characteristics, and obtaining a protein network; Combining modular fusion technology and cross-modal alignment loss function, the features in the protein network are fused to generate multimodal data fusion results.

2. The multimodal data fusion analysis method based on protein network according to claim 1 is characterized in that: The feature extraction adopts a multi-scale encoder, which is: In the formula, f i represents the extracted features, It represents feature concatenation. CNN, BiLSTM and GAT are used to extract local features, temporal features and graph structure features respectively.

3. The multimodal data fusion analysis method based on protein network according to claim 2 is characterized in that: The connection probability is calculated based on the similarity between nodes, and the formula is: Where P ij is the connection probability, is the cosine similarity between nodes i and j, f i and f j Represent the characteristics of nodes i and j respectively, θ(t) is the connection threshold that is dynamically adjusted over time t, and the adjustment formula is: θ(t)=θ0+α·(1-e -βt ); Where θ0 is the initial threshold, α and β are the attenuation coefficients, τ is the temperature parameter used to control the smoothness of the probability distribution, and σ(·) is the Sigmoid function.

4. The multimodal data fusion analysis method based on protein network according to claim 3 is characterized in that: The optimization goal of the dynamic network evolution model is: in, is the global efficiency, N is the total number of nodes, d ij is the shortest path length between nodes i and j, is the modularity, λ1 and λ2 are weight coefficients, C is the total number of network connections, L represents the total number of edges in the network, l c represents the number of edges within community c, k c represents the sum of the degrees of all nodes in community c.

5. The multimodal data fusion analysis method based on protein network according to claim 1, characterized in that: The cross-modal alignment loss function is: In the formula, and are the feature embeddings of visual and auditory modalities, is an indicator function used to force similar samples to align cross-modal features, y i and j are the outputs of nodes i and j respectively.

6. The multimodal data fusion analysis method based on protein network according to claim 3 is characterized in that: The modular fusion technology and the cross-modal alignment loss function are combined to fuse the features in the protein network to generate a multimodal data fusion result, including: Dividing the protein network into multiple modules, and determining a group of nodes with similar functions or semantics corresponding to the multiple modules according to the connection probabilities; A cross-modal alignment loss function is used to align and fuse the features in each module to obtain a modular feature representation. Based on modular fusion technology, the modular feature representations are fused and output to obtain a multimodal data fusion result.

7. The multimodal data fusion analysis method based on protein network according to claim 1, characterized in that: After generating the multimodal data fusion result, the method further includes: The multimodal data fusion result is scored to obtain a scoring result, and the formula is: In the formula, φ k (·) is the kth evaluation index, w k is the weight based on information entropy, and the formula is: In the formula, φ a represents the ath evaluation indicator, and K represents the probability of the evaluation indicator; The multimodal data fusion result is evaluated according to the scoring result to obtain an evaluation result.

8. The multimodal data fusion analysis method based on protein network according to claim 4, characterized in that: Also includes: The edge weight is determined according to the connection probability, the node color is mapped according to the centrality index, and the node-edge heat map is generated. The formula is: In the formula, C i is a centrality indicator.

9. A multimodal data fusion analysis system based on protein network, characterized in that: A method for implementing a protein network-based multimodal data fusion analysis method as claimed in any one of claims 1 to 8, comprising: A multimodal data preprocessing module is used to obtain multimodal data and perform preprocessing, wherein the preprocessing includes noise filtering, normalization and feature extraction to obtain a feature vector of each modality; A topology structure construction module, used for mapping the feature vectors into protein network nodes, calculating the connection probability based on the similarity between the nodes and constructing a protein network topology structure, wherein the protein network topology structure can realize cross-modal interaction; A protein network optimization module, used to optimize the protein network topology structure through a dynamic network evolution model, update the network topology characteristics, and obtain a protein network; The fusion result generation module is used to combine modular fusion technology and cross-modal alignment loss function to fuse the features in the protein network and generate multimodal data fusion results.