Lightweight design method for multi-modal fusion network architecture
By optimizing the multimodal fusion network structure through neural architecture search, locally interpretable models, and expert prompting mechanisms, the problems of cross-modal alignment and fusion imbalance are solved, generating a lightweight model suitable for edge devices such as autonomous driving, mobile robots, and drones.
Patent Information
- Application Number
- CN202511608670.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2025-12-09
AI Technical Summary
Existing multimodal fusion methods suffer from problems such as cross-modal alignment and fusion effectiveness, uneven utilization of modal information, and high model computational complexity, making it difficult to meet the real-time and resource efficiency requirements of complex tasks.
The neural architecture search (NAS) technique is used to determine the backbone structure of the cellular network for each modality, construct residual connections, evaluate the contribution of residual structures by combining the locally interpretable model (LIME), introduce an expert prompting mechanism (MoPE) to dynamically adjust the modality fusion ratio, and generate a lightweight student model through knowledge distillation.
It enables lightweight deployment of multimodal fusion networks, improves the model's generalization ability and decision reliability in complex scenarios, meets the real-time and low-power requirements of edge devices, and broadens application scenarios.
Smart Images

Figure CN121094064A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of artificial intelligence and deep learning, and in particular to a lightweight design method for a multimodal fusion network architecture. Background Technology
[0002] In the fields of artificial intelligence and deep learning, single-modal technologies have significant limitations due to their reliance on a single data source or sensor for information processing and decision-making. Problems such as incomplete information acquisition, poor system robustness, high data sparsity, and significant uncertainty are particularly pronounced in complex tasks requiring high accuracy, generalization ability, and environmental adaptability, making them unsuitable for practical applications.
[0003] To address the shortcomings of single-modal technologies, multimodal fusion technology has gradually become a research hotspot. This technology integrates information from various heterogeneous data sources such as images, text, audio, and sensors to construct richer and more comprehensive feature representations, effectively improving model prediction performance and decision reliability. Furthermore, when data from one modality is missing or interfered with, it can be supplemented by other modalities, significantly enhancing the system's fault tolerance and operational stability. This has significant implications for complex and dynamic scenarios such as autonomous driving, mobile robots, and drones.
[0004] However, current multimodal fusion methods face several key technical challenges. First, different modalities differ significantly in data format, semantic representation, and feature space, and the effectiveness of cross-modal alignment and fusion has not been fully resolved. Second, in practical applications, the frequency or quantity of data collection for each modality often varies significantly, which can easily lead to the model over-relying on some modalities while ignoring the information value of other modalities, affecting the overall modeling effect. Third, as application scenarios increase their requirements for real-time performance and resource efficiency, traditional multimodal fusion network architectures that rely on manual design are no longer able to meet the growing demands for flexibility and deployment due to their long development cycles, high costs, and difficulty in ensuring structural optimality.
[0005] In recent years, Neural Architecture Search (NAS) technology has provided new possibilities for the automated design of deep neural networks, but its application in multimodal tasks is still in the early stages of exploration. On the one hand, there is a lack of effective mechanisms to quantify the contribution of different modalities and their internal structures (such as residual connections) to the overall performance; on the other hand, existing models have high computational complexity and a large number of parameters, which limits their deployment capabilities on edge devices or real-time systems. Summary of the Invention
[0006] The purpose of this invention is to provide a lightweight design method for multimodal fusion network architecture, which aims to systematically solve the bottleneck problems of existing multimodal fusion models in terms of structural design, modal balance, interpretability and lightweight deployment.
[0007] To address the aforementioned technical issues, a lightweight design method for a multimodal fusion network architecture is provided, comprising:
[0008] S1. Perform neural architecture search on multiple different modalities to determine the backbone structure of the cellular network for each modality;
[0009] S2. Based on the backbone structure of the cell network with multiple different modalities, a residual structure is constructed using a general convolution operation to realize residual connections on the backbone of the cross-modal cell network.
[0010] S3. Complete the automatic design of multiple different modal cell libraries, and establish a cell library containing multiple different modalities from the resulting multiple cell network backbone structures;
[0011] S4. Evaluate the backbone structure of the cell network in the Cell library using topological complexity and accuracy;
[0012] S5. Select and mutate cell structures in the Cell library;
[0013] S6. If the current iteration number is an integer multiple of the time window, the contribution of the residual structure in the cell network backbone structure is evaluated using a locally interpretable model, the local interpretation weight of the residual structure to the unit performance is generated, and the mutation probability of the residual structure is dynamically adjusted according to the local interpretation weight. The steps from S4 to S6 are repeated to traverse the cell network backbone structure in the Cell library for each modality.
[0014] S7. Use the LIME method to perform local interpretability modeling on the residual structures in each modality cell library, obtain their local contribution weights to the changes in model output, and select modalities that have significant contributions in the current task based on the local contribution weights. Then, perform a fusion operation on the selected high-contribution modalities to construct a new overall structure search space cell library.
[0015] S8. By perturbing the residual structure in a local region and using LIME to model its impact on the model output, the local interpretation weights of each residual structure are obtained.
[0016] S9. Construct the population set for the overall architecture;
[0017] S10. Repeat steps S4 to S6 continuously to further optimize the cell structure and the overall architecture;
[0018] S11. Perform knowledge distillation on the optimal overall architecture obtained in step S10 to obtain a lightweight neural network structure.
[0019] Further, step S7 includes: during the search process of the multimodal fusion network structure, introducing a locally interpretable model to model the local interpretability of the cell structure corresponding to each modality, simulating its impact on the overall model performance by small-range perturbation of the residual structure input, and using a linear regression model to fit the local response function to obtain the local interpretive weight of each residual structure on the change of the model output, calculating the overall contribution of each modality based on the local interpretive weight, setting a LIME weight threshold, and screening out the feature modalities whose LIME local contribution weight in the multimodal cell structure is higher than the threshold, re-fusioning the screened modalities, constructing a new overall search space for the multimodal fusion network, and searching for the optimal multimodal fusion network structure using a neural network structure search method.
[0020] Furthermore, the method for calculating the LIME weight threshold includes: for the first unit within a certain cell... For each residual structure, the local interpretation weights corresponding to each residual structure are obtained using the LIME method. ( i =1,2,..., n ),in This represents the local contribution of the residual structure to the change in model output in the current task; the average absolute weight of all residual structures is calculated:
[0021] ;
[0022] The LIME weight threshold is set as follows: ,in Offset coefficient ,
[0023] It is all The standard deviation.
[0024] Furthermore, residual structure The local contribution weight in the Cell structure can be expressed as:
[0025] ;
[0026] in, This represents a multimodal fusion neural network model to be explained; Representing residual structure Input features or activation state; Representing residual structure Local interpretation weights for the model output.
[0027] Furthermore, the construction of the new multimodal fusion network overall search space includes: the search space is dynamically adjusted during the iteration process according to the local interpretation weights of the residual branches in each modal cell structure, specifically including:
[0028] S701. Whenever t%Tk==0, a local perturbation experiment is performed on the residual structure within the cell using a locally interpretable model to generate the local interpretation weights of each residual structure on the unit performance.
[0029] S702. Calculate the LIME weight threshold based on the local interpretation weight, and filter out high-contribution residual structures with a contribution higher than the LIME weight threshold;
[0030] S703. Based on the screening results, the mutation probability is dynamically adjusted, and the cells are re-fused based on the high contribution modalities to construct a new overall structural search space cell library.
[0031] Further, step S11 includes:
[0032] S111. Construct a simplified student model, wherein the student model reduces at least one of the following in the backbone chain structure: number of operation layers, number of residual branches, or channel width.
[0033] S112. Use the teacher model to infer the training dataset and generate soft labels;
[0034] S113. Compare the output of the student model with the output of the teacher model, and optimize the joint loss function by combining the real labels;
[0035] S114. Iteratively train the student model until its performance is close to that of the teacher model and its inference speed meets the actual deployment requirements.
[0036] Furthermore, the student model is configured to have fewer layers, channels, or parameters. During the training of the student model, both the original task labels and the probability distribution output by the teacher model are used as supervision signals to improve the generalization ability of the student model by minimizing the difference between the two.
[0037] Furthermore, step S7 also includes: using an expert-guided technique to dynamically determine the weight ratios of locally interpretable models for different modalities.
[0038] Furthermore, the expert suggestion technology includes multiple suggestion expert modules oriented towards specific tasks or application scenarios. Each suggestion expert module predicts the relative importance of each modality to the overall performance improvement in the current scenario based on the current task type, environmental characteristics, or data distribution characteristics. Based on the modality importance prediction results output by the suggestion experts, and combined with the LIME local interpretation weights of the residual branches in the cell structure of each modality, the contribution ratio of different modalities in the fusion process is dynamically adjusted. Modality screening and fusion are performed according to the dynamically adjusted LIME weight ratios to generate a multimodal fusion overall structure search space cell library.
[0039] Implementing the embodiments of the present invention will have the following beneficial effects:
[0040] The lightweight design method for the multimodal fusion network architecture in this embodiment integrates key technologies such as Neural Architecture Search (NAS), Locally Interpretable Models (LIME), Expert Hints (MoPE), and knowledge distillation. This achieves fully automated optimization of the multimodal fusion network from structural design to lightweight deployment. Firstly, NAS automatically searches for the backbone structure of each modality and the overall fusion architecture, eliminating reliance on expert experience, significantly shortening the development cycle and reducing design costs. Simultaneously, by combining LIME's quantitative evaluation of the residual structure's contribution, the mutation strategy is dynamically optimized to guide the network towards a high-performance structure, ensuring that the final architecture is closer to the global optimum in terms of task adaptability. Secondly, LIME is introduced to conduct local perturbation experiments on the residual structure, accurately selecting high-value modalities that significantly contribute to the task, avoiding the problem of the model over-relying on a certain type of modality while ignoring other information. Combined with the Expert Hints (MoPE), the fusion ratio of each modality is dynamically adjusted according to the task scenario and environmental characteristics, achieving efficient alignment and collaboration of cross-modal information, and improving the model's generalization ability and decision reliability in complex scenarios. Thirdly, by generating local interpretation weights for the residual structure using LIME, the influence of each component on the model output is quantified, solving the "black box" problem of traditional multimodal networks and making the architecture optimization process interpretable and traceable. Fourthly, by using knowledge distillation technology, the optimal architecture is compressed into a lightweight student model. While maintaining prediction accuracy close to that of the teacher model, the number of model layers, parameters, and computational overhead are reduced, significantly improving inference speed. This enables it to meet the deployment requirements of edge devices (such as autonomous driving sensors, mobile robots, and drones) for real-time performance and low power consumption, thus broadening the application scenarios of multimodal fusion technology. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a flowchart of the lightweight design method for the multimodal fusion network architecture described in the embodiments of the present invention. Figure 1 ;
[0043] Figure 2 This is a flowchart of step S7 of the present invention;
[0044] Figure 3 This is a flowchart of the method for step S11 as described in an embodiment of the present invention;
[0045] Figure 4 This is a flowchart of the lightweight design method for the multimodal fusion network architecture described in the embodiments of the present invention. Figure 2 ;
[0046] Figure 5 This is a flowchart illustrating the design process of a single-modality Cell module according to an embodiment of the present invention.
[0047] Figure 6 This is a schematic diagram of the encoding form of the multimodal fusion network Cell structure according to an embodiment of the present invention;
[0048] Figure 7 This is a schematic diagram illustrating the automatic design principle of the multimodal fusion network Cell structure according to an embodiment of the present invention.
[0049] Figure 8 This is an automatic design diagram of the overall structure of the multimodal fusion neural network described in the embodiments of the present invention;
[0050] Figure 9 This is a diagram illustrating the process of searching for the multimodal fusion neural network structure according to an embodiment of the present invention;
[0051] Figure 10 This is a process diagram of model knowledge distillation as described in an embodiment of the present invention;
[0052] Figure 11 This is a flowchart illustrating the process of MoPE dynamically determining the LIME weight ratios for different modes, as described in an embodiment of the present invention. Detailed Implementation
[0053] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Preferred embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the invention.
[0054] It should be noted that when a component is said to be "fixed to" another component, it can be directly attached to the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0056] Please refer to Figures 1-11 This invention provides a lightweight design method for a multimodal fusion network architecture, comprising:
[0057] S1. Perform neural architecture search on multiple different modalities to determine the backbone structure of the cellular network for each modality. For example, for the various heterogeneous modalities involved in the task (such as images, text, audio, sensor data, etc.), determine the original input format for each modality (e.g., pixel matrix of images, waveform signal of audio, word vector of text, time-series data of sensors, etc.) to provide a targeted input foundation for subsequent searches. Pre-set a set of basic operation units suitable for the data characteristics of each modality, forming an "operation pool." These operation units include the basic layers or components of the neural network. Run the neural architecture search algorithm independently for each modality, automatically selecting operation units from the pre-set operation pool and combining them into different network topologies, generating a large number of candidate "cellular network backbone structures." During the search process, the algorithm explores different combinations of operation units (e.g., the order and parameter configuration of "convolutional layer + pooling layer + activation layer") through evolutionary strategies, reinforcement learning, or random search. For the candidate cellular network backbone structures generated for each modality, performance is evaluated using training data. Core metrics include accuracy, topological complexity, and feature representation ability. Accuracy refers to the basic recognition precision for tasks within a modality (such as image classification and text semantic extraction). Topological complexity refers to the number of parameters and computational cost (FLOPs) of the structure, avoiding excessive complexity that could lead to difficulties in subsequent fusion. Feature representation capability refers to whether the extracted features can effectively support subsequent cross-modal fusion. Based on the evaluation results, the optimal cellular network backbone structure with suitable complexity is selected for each modality as the basic framework for that modality. After completing the search and selection for all modalities, the "cellular network backbone structure" corresponding to each modality is obtained.
[0058] S2. Based on the backbone structures of various modal cell networks, a general convolutional operation is used to construct residual structures to achieve residual connections on the backbones of cross-modal cell networks. Using the backbone structures of each modality determined in step S1 as a foundation (e.g., convolutional backbone structure for image modality, one-dimensional convolutional backbone structure for audio modality, Transformer backbone structure for text modality, etc.), the cross-modal combinations requiring connections (e.g., "image-audio", "text-image", etc.) are identified. The goal is to allow features from different modalities to interact through additional connection channels, enhancing the correlation between cross-modal features. A general convolutional operation suitable for multimodal feature processing is used as the core component for constructing the residual structures. For cross-modal backbone structures, a "backbone-residual branch" topology is constructed. Specifically, this involves: designing residual branches for the two modalities that need to be connected, with common convolutional operations (e.g., "1x1 convolution → activation function → 3x3 convolution") chained within the branches to achieve feature transformation and enhancement; setting connection points at corresponding levels of the two modalities (e.g., mid- or high-level feature extraction layers) to ensure that the features transmitted by the residual branches are at similar semantic levels (avoiding ineffective fusion due to excessive level differences); and element-wise adding or concatenating the output of the modality A residual branch with the features of the modality B backbone to achieve direct injection of cross-modal information (e.g., features from the image backbone + features from the audio residual branch). After construction, the effectiveness of the cross-modal residual structure is verified through preliminary training.
[0059] S3. Complete the automatic design of multiple modal cell libraries, and establish a cell library containing multiple modalities from the resulting multiple cell network backbone structures; collect the modality-specific cell network backbone structures determined by neural architecture search (NAS) in step S1, and integrate the cross-modal residual structures constructed in step S2 with the corresponding modal backbone structures to form a "complete cell structure (Cell) with residual connections". A structured storage strategy for the Cell library was developed to ensure efficient retrieval and retrieval of Cells across multiple modalities. This included: classifying Cells by modality (image, audio, text, etc.) and labeling the core function of each Cell (e.g., "image feature extraction Cell," "audio temporal feature Cell"); recording key metrics for each Cell, such as parameter count, computational complexity (FLOPs), and accuracy in single-modal tasks, to provide a basis for subsequent evaluation and selection; and specifying the input / output feature dimensions and compatible cross-modal residual connection types for each Cell (e.g., "image Cell output dimension 256, can be connected to audio Cell's 128-dimensional input via 1x1 convolution") to ensure structural compatibility during fusion. All single-modal Cells with integrated residual structures were then aggregated according to the above rules to form a Cell library containing multiple modalities, and its completeness was verified.
[0060] S4. The backbone structures of cell networks in the Cell library are evaluated using topological complexity and accuracy. Topological complexity assesses the "simplicity" and "computational efficiency" of the backbone structure, with specific metrics including the number of parameters, computational cost, and network depth and width. The number of parameters refers to the total number of learnable parameters in the backbone structure (e.g., weights of convolutional layers, weights and biases of fully connected layers); a lower number of parameters indicates a more concise structure. Computational cost (FLOPs) refers to the number of floating-point operations required for the backbone structure to complete one forward inference (e.g., FLOPs of a convolutional layer = number of input channels × number of output channels × kernel size × output feature map size); a lower computational cost indicates higher inference efficiency. Network depth and width refer to the number of operational layers in the backbone structure (e.g., the total number of convolutional and pooling layers) and the channel width of each layer, avoiding structural redundancy due to excessive depth or width. The topological complexity of each backbone structure is quantified by comprehensively scoring these metrics (lower scores indicate more efficient structures). Each backbone structure in the Cell library is treated as an independent model and fine-tuned on the corresponding modality's training data (with the structure fixed, only the parameters are optimized) to ensure that the structure can learn the core features of the modality. The trained backbone structure is then used for inference on the validation set to calculate its accuracy in basic tasks (such as classification accuracy for image modality, semantic matching accuracy for text modality, and event recognition accuracy for audio modality). Through multiple training and testing sessions, the fluctuation of accuracy is observed, and backbone structures with stable performance (small fluctuations) and high accuracy are selected.
[0061] S5. Select and mutate cell structures in the Cell library; sort cell structures in the Cell library in descending order based on a weighted comprehensive score of topological complexity (low is better) and accuracy (high is better); select the top K cell structures and retain them directly (e.g., retain the top 30% of structures) to ensure that high-quality structures enter the next round of optimization; prioritize retaining structures that perform well in specific tasks (e.g., a cell structure in an image has extremely high accuracy in a classification task, so it is still retained even if its complexity is slightly higher), to avoid missing high-quality structures due to a single sort. Mutate the retained cell structures to generate new candidate structures by adjusting structural details to explore better topologies. Specific mutation methods include: replacing some basic operations in the backbone chain structure (e.g., replacing a 3x3 convolution with a 3x3 depthwise separable convolution, or adding / deleting pooling layers); increasing / decreasing the number of operation layers in the backbone structure (e.g., adding 1 convolution layer on top of the existing 4 layers), or adjusting the channel width of each layer (e.g., adjusting the number of channels in a layer from 64 to 96). Add / remove cross-modal residual branches (e.g., add a residual connection to an image-audio cell structure), or change the type of convolution operation in the residual branch (e.g., replace 1x1 convolution with 3x3 convolution); adjust the input / output connection points of the residual branch (e.g., change from a connection at layer 2 of the backbone to a connection at layer 3), or change the residual fusion method (from element-wise addition to feature concatenation). Set differentiated mutation probabilities for different structural components (e.g., low mutation probability for core high-contribution operations, high mutation probability for minor operations) to avoid excessive mutation destroying the core features of high-quality structures. Quickly evaluate the topological complexity of the new structure (e.g., number of parameters, computational cost), and eliminate overly complex and invalid mutated structures; perform simple performance tests on the mutated structures (e.g., accuracy after training on a small amount of data), and retain candidate structures whose performance has not significantly decreased; merge the selected mutated structures with the retained original high-quality structures, update the Cell library, and provide richer structural samples for the next round of iteration optimization (step S6 and repeated process).
[0062] S6. If the current iteration count is an integer multiple of the time window, a locally interpretable model is used to evaluate the contribution of the residual structure within the cell network backbone, generating local interpretation weights for the residual structure's contribution to unit performance. The mutation probability of the residual structure is dynamically adjusted based on these local interpretation weights, and steps S4 to S6 are repeated to traverse the cell network backbone structures in each modality's Cell library. For example, the current iteration count is an integer multiple of the time window, i.e., t%Tk==0. Iterative optimization of the network architecture (selection and mutation operations in steps S4-S5) needs to be continuously performed to improve performance. However, if the contribution of the residual structure is evaluated in each iteration (LIME analysis in step S6), the computational overhead will significantly increase due to frequent calls to the interpretable model, slowing down the overall optimization process. By setting "t%Tk==0", LIME evaluation is triggered only after Tk iterations (e.g., every 10 or 20 iterations), ensuring effective tracking of the residual structure's contribution while avoiding efficiency losses due to excessively frequent evaluations, achieving a balance between "optimization accuracy" and "computational cost". By setting a periodic evaluation node through a "time window Tk", the local explanatory weights of residual structures can be calculated using LIME based on the phased optimization results (such as cell structure performance after Tk iterations), thereby dynamically adjusting the mutation probability (high-contribution residual structures have fewer mutations, and low-contribution structures have more mutations). The setting of "t%Tk==0" controls the evaluation frequency through a fixed time window (Tk), enabling LIME evaluation to capture the stable contribution characteristics of residual structures after multiple iterations, ensuring a more accurate judgment of the importance of cross-modal residual connections, and providing a reliable basis for subsequent modality screening and fusion (step S7).
[0063] S7. Using the LIME method, perform local interpretability modeling on the residual structures in each modality's Cell library to obtain their local contribution weights to changes in model output. Based on these local contribution weights, select modalities that make significant contributions to the current task. Merge the selected high-contribution modalities to construct a new overall structure search space cell library. Perform small-scale perturbations on the input features or activation states of each residual structure (e.g., slightly modifying image pixels or adjusting audio waveform segments) to simulate their impact on the model output. Use simple models such as linear regression to fit the relationship between the perturbated input and the changes in model output to generate local interpretation weights for each residual structure. The higher the weight value, the greater the impact of the residual structure on the model output (the higher the contribution). Summarize the local interpretation weights of all residual structures within the same modality to calculate the overall contribution of the modality (e.g., through average weighting or weighted summation, reflecting the comprehensive value of the modality to the current task).
[0064] S8. By perturbing the input of the residual structure in a local region and modeling its impact on the model output using LIME, the local interpretation weights of each residual structure are obtained. For the residual structures (including cross-modal residual connections and intra-modal residual branches) in each modal cell structure (such as cells of image, audio, and text modalities) in the Cell library, the specific residual structure to be evaluated and its corresponding input feature region are identified. The input features of the residual structure are systematically perturbed in the defined local region to simulate its potential impact on the model output. Model inference is performed on the perturbed input features. The output results before and after the perturbation are compared to capture the impact of the residual structure on the model decision. Based on the perturbation experimental data, the relationship between the residual structure and the model output change is fitted using the LIME (Locally Interpretable Model - Agnostic Interpretation) method to quantify its contribution.
[0065] S9. Construct the population set for the overall architecture;
[0066] S10. Continue repeating steps S4 to S6 to further optimize cell structure and overall architecture;
[0067] S11. Perform knowledge distillation on the optimal overall architecture obtained in step S10 to obtain a lightweight neural network structure.
[0068] Please refer to Figure 1 and Figure 4 This invention proposes a lightweight design method for a multimodal fusion network architecture. Specifically, the invention first employs a NAS algorithm to independently search the data for each modality to determine the optimal backbone structure of the cell network corresponding to each modality, and then constructs a cell library containing multimodal feature expressions. Subsequently, a Locally Interpretable Model (LIME) is introduced to conduct local perturbation experiments on the residual connections in the cell structures of each modality, obtaining their local contribution weights to the overall task performance. Further, a Mixture of Prompt Experts (MoPE) mechanism is combined to dynamically adjust the fusion ratio between modalities according to different task scenarios, thereby generating a task-adaptive multimodal fusion cell library. During the fusion process, a preset LIME weight threshold is used to filter out high-value modalities that significantly contribute to the current task, and a new overall structure search space is reconstructed based on these modalities. Finally, the NAS algorithm is again used to efficiently search within this fusion search space to determine the globally optimal multimodal fusion neural network architecture.
[0069] Furthermore, after completing the optimal structure search, this invention introduces knowledge distillation technology to compress the obtained high-performance teacher model into a lightweight student model, enabling it to maintain high prediction accuracy while having stronger deployment capabilities to meet the operational needs of edge devices or real-time tasks.
[0070] In summary, by introducing multiple key technologies such as LIME, MoPE, and knowledge distillation, this invention achieves fully automated design and optimization from single-modal structure search and cross-modal fusion to model lightweighting, significantly improving the modeling efficiency, generalization ability, and adaptability of multimodal neural networks.
[0071] A primitive operation pool is defined for searching multimodal fusion neural network architectures. This pool contains a set of selectable basic operation units. These basic operation units consist of neural network layers suitable for multimodal data processing, and can be combined in different ways and with different connection structures to construct a variety of multimodal fusion neural network architectures with different forms and functions.
[0072] When automatically designing the structure of a multimodal fusion network cell, the first step is to construct the backbone chain structure, and then introduce residual connections. These residual structures can be flexibly embedded into the backbone structure with various combinations of input and output positions, significantly improving the diversity and expressive power of the network structure. This invention allows each cell to have up to three residual branches, which have the ability to adjust the dimensionality of the input data, further enhancing the model's nonlinear modeling capabilities and structural design flexibility, while also introducing higher design complexity.
[0073] Please refer to Figure 6 To enhance the flexibility and diversity of network structure design, this invention endows each multimodal unit (Cell) with the ability to make autonomous decisions, namely, whether to introduce additional branches and how to organize the connections between these branches. These optional branches can be regarded as additional functional modules within the multimodal unit, capable of adjusting the dimension of input features, enhancing the model's expressive power, or increasing the network depth, thereby effectively improving the overall architecture's nonlinear modeling capability and task adaptability. In specific implementation, the backbone structure of each Cell is constructed using a chain-like connection method and consists of a set of basic operation units. The basic operations include: 1x1 convolution, 3x3 convolution, 5x5 convolution, 3x3 depthwise separable convolution (DW), and 5x5 depthwise separable convolution (DW), etc. These operations are the most basic and widely used components in current deep neural networks, possessing good versatility and computational efficiency.
[0074] Please refer to Figure 6To enhance the flexibility and diversity of multimodal neural network architecture design, this invention proposes a design strategy for automatically generating multimodal cells with various connection methods. Before generating these cells, the system first explicitly defines the backbone structure of each cell to ensure good versatility and scalability. This backbone structure is built upon the most basic and widely used operations in neural networks, including 1x1 convolution, 3x3 convolution, 5x5 convolution, 3x3 depthwise separable convolution (DW), and 5x5 depthwise separable convolution (DW), forming a chain structure containing four layers of basic operations as the basic framework of the multimodal cells. Based on this, the system empowers the machine to freely select and combine these basic operations according to task requirements, thereby supporting the construction of more complex and efficient network topologies. A highly flexible design mechanism enables the system to automatically adjust and optimize the network architecture based on input data characteristics and specific task objectives, ultimately generating efficient models more suitable for specific application scenarios. Furthermore, by introducing a skip operation, direct information transfer between different layers is achieved, further enhancing the connectivity and expressive power of the network structure.
[0075] Furthermore, the residual structure can be embedded in the backbone structure in various forms, specifically manifested as different combinations of input and output positions. This invention allows up to three residual branches to be added to each cell. These branches not only change the dimensionality of the input data but also enhance the model's nonlinear modeling capability, thereby improving the overall network's expressiveness and adaptability.
[0076] Furthermore, this method employs a Locally Interpretable Model (LIME) to conduct local perturbation experiments on the residual structures, thereby evaluating the contribution of each residual branch to the performance improvement of the cell structure. In this method, the system perturbs the input of multiple residual structures within a cell, simulating their impact on the overall model performance. A local response function is then fitted using linear regression or other simple models to obtain the local explanatory weight of each residual structure for the change in model output. This local explanatory weight reflects the degree of influence of the residual structure on the model's prediction results under the current task and can be used to quantify its relative importance to overall performance. Specifically, within a local region, the existence state or parameter configuration of a residual branch is changed, and the changes in the model output are recorded. This is used to train a locally explanatory model to approximate the behavior of the original network in that region. Based on the obtained local explanatory weights, the contribution ratio of each residual structure can be further calculated, and its mutation probability during the neural architecture search process can be dynamically adjusted accordingly.
[0077] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the technical scope disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention. The multimodal fusion neural network proposed in this invention is an intelligent perception system capable of efficiently integrating data from multiple sensors (such as visual images, LiDAR point clouds, radar signals, sonar information, etc.), aiming to provide a more comprehensive, accurate, and robust environmental understanding capability. This technology, through automated modeling and structural optimization, achieves deep perception of complex dynamic scenes, demonstrating broad application prospects in multiple cutting-edge artificial intelligence application fields such as autonomous driving, mobile robots, and drones. In the field of autonomous driving, the multimodal fusion network constructed in this invention can effectively integrate data from multiple sensing devices such as cameras, LiDAR, and millimeter-wave radar, forming an all-round, multi-layered environmental perception capability. By using a Locally Interpretable Model (LIME) to analyze the contribution of the residual structure, the system can identify key feature modalities and optimize their fusion ratio, thereby improving the recognition and tracking accuracy of static targets (such as traffic signs and road markings) and dynamic targets (such as pedestrians and vehicles). Meanwhile, even under extreme weather conditions (such as rain, snow, and fog), the system can still maintain stable operation through data from other modalities, significantly enhancing the overall system's fault tolerance and reliability, even if the performance of some sensors degrades. In the field of mobile robots, this invention combines a Mixture of Prompt Experts (MoPE) mechanism to automatically adjust the weight ratio of each modality according to different task scenarios (such as indoor navigation, outdoor inspection, and human-computer interaction), enabling the robot to more accurately fuse distance information from LiDAR and visual perception data to achieve high-quality map construction and path planning. Furthermore, this mechanism also improves the robot's autonomous decision-making and obstacle avoidance capabilities in complex environments such as densely populated areas and dynamic changes. In UAV applications, this invention employs multi-source sensor fusion technology to uniformly model information from cameras, radar, GPS, etc., enabling high-precision positioning and stable flight in complex environments such as urban buildings and mountainous terrain. By introducing knowledge distillation technology, the high-performance teacher model is further compressed to generate a lightweight student model, allowing the system to maintain high-precision perception while possessing stronger edge computing capabilities, meeting the requirements for real-time and low-power deployment.
[0078] In summary, this invention not only overcomes the limitations of traditional single-modal perception systems, but also constructs a multimodal fusion neural network architecture with adaptability, efficiency, and generalization capabilities by introducing LIME interpretability analysis, MoPE task adaptation mechanism, and knowledge distillation lightweight strategy, providing a more intelligent, safe, and reliable solution for various intelligent perception systems.
[0079] The lightweight design method for the multimodal fusion network architecture in this embodiment integrates key technologies such as Neural Architecture Search (NAS), Locally Interpretable Models (LIME), Expert Hints (MoPE), and knowledge distillation. This achieves fully automated optimization of the multimodal fusion network from structural design to lightweight deployment. Firstly, NAS automatically searches for the backbone structure of each modality and the overall fusion architecture, eliminating reliance on expert experience, significantly shortening the development cycle and reducing design costs. Simultaneously, by combining LIME's quantitative evaluation of the residual structure's contribution, the mutation strategy is dynamically optimized to guide the network towards a high-performance structure, ensuring that the final architecture is closer to the global optimum in terms of task adaptability. Secondly, LIME is introduced to conduct local perturbation experiments on the residual structure, accurately selecting high-value modalities that significantly contribute to the task, avoiding the problem of the model over-relying on a certain type of modality while ignoring other information. Combined with the Expert Hints (MoPE), the fusion ratio of each modality is dynamically adjusted according to the task scenario and environmental characteristics, achieving efficient alignment and collaboration of cross-modal information, and improving the model's generalization ability and decision reliability in complex scenarios. Thirdly, by generating local interpretation weights for the residual structure using LIME, the influence of each component on the model output is quantified, solving the "black box" problem of traditional multimodal networks and making the architecture optimization process interpretable and traceable. Fourthly, by using knowledge distillation technology, the optimal architecture is compressed into a lightweight student model. While maintaining prediction accuracy close to that of the teacher model, the number of model layers, parameters, and computational overhead are reduced, significantly improving inference speed. This enables it to meet the deployment requirements of edge devices (such as autonomous driving sensors, mobile robots, and drones) for real-time performance and low power consumption, thus broadening the application scenarios of multimodal fusion technology.
[0080] In one possible implementation, step S7 includes: during the search for the multimodal fusion network structure, introducing a locally interpretable model to model the local interpretability of the cell structure corresponding to each modality; simulating the impact on the overall model performance by small-range perturbation of the residual structure input; and fitting the local response function using a linear regression model to obtain the local interpretive weight of each residual structure on the model output change; calculating the overall contribution of each modality based on the local interpretive weight; setting a LIME weight threshold; and selecting feature modalities whose LIME local contribution weight in the multimodal cell structure is higher than the threshold; re-fusioning the selected modalities to construct a new overall search space for the multimodal fusion network; and searching for the optimal multimodal fusion network structure using a neural network structure search method. For example, by conducting small-range perturbation experiments on the residual structure input and combining the local response function with a linear regression model, the specific impact weight of each residual structure on the model output can be quantified. Further calculating the overall contribution of each modality based on this weight can clearly distinguish the actual value of different modalities in the task (e.g., the contribution of the image modality to visual features, and the contribution of the audio modality to sound features). This quantification mechanism effectively avoids the problem of "over-reliance on some modalities while ignoring others" in traditional multimodal models, ensuring that high-value modalities are fully utilized and low-contribution modalities do not occupy redundant resources, thus achieving balanced collaboration among modalities. Furthermore, by setting a LIME weight threshold and filtering out feature modalities with weights higher than the threshold, modalities or residual structures with weak contributions to the task can be eliminated, concentrating search resources on the fusion optimization of high-value modalities. Since the locally explained weights are obtained based on perturbation experiments under specific task scenarios, and the selected high-contribution modalities directly match the current task requirements, the reconstructed network architecture can better adapt to task characteristics.
[0081] In one possible implementation, the LIME weight threshold is calculated by: for the first LIME weight within a certain unit... For each residual structure, the local interpretation weights corresponding to each residual structure are obtained using the LIME method. ( i =1,2,..., n ),in This represents the local contribution of the residual structure to the change in model output in the current task; the average absolute weight of all residual structures is calculated:
[0082] ;
[0083] Set the LIME weight threshold as follows: ,in Offset coefficient ,
[0084] It is all The standard deviation.
[0085] In one possible implementation, the residual structure The local contribution weight in the Cell structure can be expressed as:
[0086] ;
[0087] in, This represents a multimodal fusion neural network model to be explained; Representing residual structure Input features or activation state; Representing residual structure Local interpretation weights for the model output.
[0088] Please refer to Figure 2 In one possible implementation, constructing the new overall search space for the multimodal fusion network includes: dynamically adjusting the search space during iteration based on the local interpretation weights of the residual branches in each modal cell structure, specifically including:
[0089] S701. Whenever t%Tk==0, a local perturbation experiment is performed on the residual structure within the cell using a locally interpretable model to generate the local interpretation weights of each residual structure on the unit performance.
[0090] S702. Calculate the LIME weight threshold based on the local interpretation weight, and select high-contribution residual structures with a contribution higher than the LIME weight threshold.
[0091] S703. Based on the screening results, the mutation probability is dynamically adjusted, and the cells are re-fused based on the high contribution modalities to construct a new overall structural search space cell library.
[0092] Please refer to Figure 3 and Figure 10 In one possible implementation, step S11 includes:
[0093] S111. Construct a simplified student model, wherein the student model reduces at least one of the following in the backbone chain structure: number of operation layers, number of residual branches, or channel width.
[0094] S112. Use the teacher model to infer the training dataset and generate soft labels;
[0095] S113. Compare the output of the student model with the output of the teacher model, and optimize the joint loss function by combining the real labels;
[0096] S114. Iteratively train the student model until its performance is close to that of the teacher model and its inference speed meets the actual deployment requirements.
[0097] Please refer to Figure 10 In one possible implementation, the student model is configured to have fewer layers, channels, or parameters. During the training of the student model, both the original task labels and the probability distributions output by the teacher model are used as supervision signals to improve the generalization ability of the student model by minimizing the difference between the two.
[0098] Please refer to Figure 10 The teacher model is an optimal multimodal fusion network obtained through multiple rounds of optimization. Its output soft labels not only contain information on "whether it is correctly classified," but also include the probability distribution between different categories (e.g., "90% probability of being a cat, 8% probability of being a dog, and 2% probability of being another animal"). This probability distribution implies the teacher model's subtle judgments on data features and the correlation between categories—"tacit knowledge"—far richer than the true labels that only represent the final result (e.g., "cat"). By learning the soft labels, the student model can inherit the deep cognitive ability of the teacher model and perform closer to it in terms of performance. The probability distribution of the soft labels is more robust to noisy data or fuzzy samples. For example, for a fuzzy image with "mixed cat and dog features," the soft label output by the teacher model may contain trade-off information on the fuzzy features, while the true label can only give a single category. By minimizing the difference with the soft labels (e.g., KL divergence), the student model can learn tolerance to data variation and generalization ability, performing more stably when faced with unseen new data. The goal of knowledge distillation is to compress complex teacher models into lightweight student models (e.g., reducing the number of layers and parameters). Soft labels, as the "refined output" of the teacher model, can provide the student model with more nuanced optimization directions than the true labels. By simultaneously fitting soft labels and true labels (joint loss function optimization), the student model can retain key feature extraction capabilities to the greatest extent while simplifying the structure, avoiding a significant performance drop due to compression.
[0099] Please refer to Figure 11 In one possible implementation, step S7 further includes: using an expert-guided technique to dynamically determine the weight ratios of locally interpretable models for different modalities.
[0100] Please refer to Figure 11In one possible implementation, the expert prompting technology includes multiple prompting expert modules tailored to specific tasks or application scenarios. Each prompting expert module predicts the relative importance of each modality to the overall performance improvement in that scenario based on the current task type, environmental characteristics, or data distribution characteristics. Based on the modal importance prediction results output by the prompting experts, and combined with the LIME local interpretation weights of the residual branches in each modal cell structure, the contribution ratio of different modalities in the fusion process is dynamically adjusted. Modal selection and fusion are performed according to the dynamically adjusted LIME weight ratios to generate a multimodal fusion overall structure search space cell library. For example, each prompting expert module can specifically predict the relative importance of each modality based on the current task type (e.g., autonomous driving, robot navigation), environmental characteristics (e.g., rainy / snowy weather, indoor / outdoor scenarios), or data distribution characteristics (e.g., sparse data for a certain modality). The expert prompting technology combines the "scenario experience prediction of the expert module" with the "LIME residual branch data contribution weights." When the two are combined, the dynamically adjusted fusion ratio avoids the local optima that pure data-driven approaches may fall into (such as over-reliance due to data redundancy in a certain modality), and also makes up for the limitations of pure experience-based guidance lacking data support, making modality screening and fusion more scientific and accurate.
[0101] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A lightweight design method for a multimodal fusion network architecture, characterized in that, include: S1. Perform neural architecture search on multiple different modalities to determine the backbone structure of the cellular network for each modality; S2. Based on the backbone structure of the cell network with multiple different modalities, a residual structure is constructed using a general convolution operation to realize residual connections on the backbone of the cross-modal cell network. S3. Complete the automatic design of multiple different modal cell libraries, and establish a cell library containing multiple different modalities from the resulting multiple cell network backbone structures; S4. Evaluate the backbone structure of the cell network in the Cell library using topological complexity and accuracy; S5. Select and mutate cell structures in the Cell library; S6. If the current iteration number is an integer multiple of the time window, the contribution of the residual structure in the cell network backbone structure is evaluated using a locally interpretable model, the local interpretation weight of the residual structure to the unit performance is generated, and the mutation probability of the residual structure is dynamically adjusted according to the local interpretation weight. The steps from S4 to S6 are repeated to traverse the cell network backbone structure in the Cell library for each modality. S7. Use the LIME method to perform local interpretability modeling on the residual structures in each modality cell library, obtain their local contribution weights to the changes in model output, and select modalities that have significant contributions in the current task based on the local contribution weights. Then, perform a fusion operation on the selected high-contribution modalities to construct a new overall structure search space cell library. S8. By perturbing the residual structure in a local region and using LIME to model its impact on the model output, the local interpretation weights of each residual structure are obtained. S9. Construct the population set for the overall architecture; S10. Repeat steps S4 to S6 continuously to further optimize the cell structure and the overall architecture; S11. Perform knowledge distillation on the optimal overall architecture obtained in step S10 to obtain a lightweight neural network structure.
2. The lightweight design method for the multimodal fusion network architecture according to claim 1, characterized in that, Step S7 includes: during the search process of the multimodal fusion network structure, a locally interpretable model is introduced to model the local interpretability of the cell structure corresponding to each modality. By perturbing the input of the residual structure within a small range, the impact on the overall model performance is simulated. A linear regression model is used to fit the local response function to obtain the local interpretive weight of each residual structure on the change of the model output. Based on the local interpretive weight, the overall contribution of each modality is calculated, and a LIME weight threshold is set. Feature modalities with LIME local contribution weights higher than the threshold in the multimodal cell structure are selected. The selected modalities are re-fused to construct a new overall search space for the multimodal fusion network. The optimal multimodal fusion network structure is searched by a neural network structure search method.
3. The lightweight design method for the multimodal fusion network architecture according to claim 2, characterized in that, The method for calculating the LIME weight threshold includes: for the first unit within a certain cell... For each residual structure, the local interpretation weights corresponding to each residual structure are obtained using the LIME method. ( i =1,2,..., n ),in This represents the local contribution of the residual structure to the change in model output in the current task; the average absolute weight of all residual structures is calculated: ; The LIME weight threshold is set as follows: ,in Offset coefficient , It is all The standard deviation.
4. The lightweight design method for the multimodal fusion network architecture according to claim 3, characterized in that, Residual Structure The local contribution weight in the Cell structure can be expressed as: ; in, This represents a multimodal fusion neural network model to be explained; Representing residual structure Input features or activation state; Representing residual structure Local interpretation weights for the model output.
5. The lightweight design method for the multimodal fusion network architecture according to claim 2, characterized in that, The construction of the new multimodal fusion network overall search space includes: the search space is dynamically adjusted during the iteration process according to the local interpretation weights of the residual branches in each modal cell structure, specifically including: S701. Whenever t%Tk==0, a local perturbation experiment is performed on the residual structure within the cell using a locally interpretable model to generate the local interpretation weights of each residual structure on the unit performance. S702. Calculate the LIME weight threshold based on the local interpretation weight, and filter out high-contribution residual structures with a contribution higher than the LIME weight threshold; S703. Based on the screening results, the mutation probability is dynamically adjusted, and the cells are re-fused based on the high contribution modalities to construct a new overall structural search space cell library.
6. The lightweight design method for the multimodal fusion network architecture according to claim 1, characterized in that, Step S11 includes: S111. Construct a simplified student model, wherein the student model reduces at least one of the following in the backbone chain structure: number of operation layers, number of residual branches, or channel width. S112. Use the teacher model to infer the training dataset and generate soft labels; S113. Compare the output of the student model with the output of the teacher model, and optimize the joint loss function by combining the real labels; S114. Iteratively train the student model until its performance is close to that of the teacher model and its inference speed meets the actual deployment requirements.
7. The lightweight design method for the multimodal fusion network architecture according to claim 6, characterized in that, The student model is configured to have fewer layers, channels, or parameters. During the training of the student model, both the original task labels and the probability distribution output by the teacher model are used as supervision signals. The generalization ability of the student model is improved by minimizing the difference between the two.
8. The lightweight design method for the multimodal fusion network architecture according to claim 7, characterized in that, Step S7 also includes: using an expert-guided technique to dynamically determine the weight ratios of locally interpretable models for different modalities.
9. The lightweight design method for the multimodal fusion network architecture according to claim 8, characterized in that, The expert suggestion technology includes multiple suggestion expert modules tailored to specific tasks or application scenarios. Each suggestion expert module predicts the relative importance of each modality to the overall performance improvement in that scenario based on the current task type, environmental characteristics, or data distribution characteristics. Based on the modal importance prediction results output by the suggestion experts, and combined with the LIME local interpretation weights of the residual branches in each modal cell structure, the contribution ratio of different modalities in the fusion process is dynamically adjusted. Modal selection and fusion are performed according to the dynamically adjusted LIME weight ratios to generate a multimodal fusion overall structure search space cell library.