Multi-modal data reasoning method and device, equipment and storage medium

By deploying a lightweight multimodal data inference model on edge devices, combining data cleaning, fusion and incremental learning, the problem of high cloud computing latency is solved, and efficient and real-time data inference and user experience improvement is achieved.

CN120278280APending Publication Date: 2025-07-08SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510512519.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing smart assistants mainly rely on cloud computing mode when making data inference, resulting in high latency and inability to ensure the timeliness of problem solving, reducing user experience.

Method used

By deploying a lightweight multimodal data inference model on edge devices, using frequency recursive convolutional recurrent network and dynamic time adjustment algorithm for data cleaning and fusion, and configuring incremental learning algorithms for model optimization and parameter updates, reducing computing resource requirements.

Benefits of technology

It realizes efficient data inference locally, reduces latency, improves the real-time data processing and targetedness of inference results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278280A_ABST
    Figure CN120278280A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data reasoning method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining initial multi-modal data, and carrying out the data cleaning of the initial multi-modal data, so as to obtain target multi-modal data; performing lightweight processing on the initial data reasoning large model to obtain a target data reasoning large model, and analyzing the target multi-modal data by using the target data reasoning large model to determine a data reasoning demand; wherein the target data reasoning large model is a large model deployed locally; and fusing the target multi-modal data based on a preset dynamic time adjustment algorithm to obtain corresponding fused data, and performing data reasoning on the fused data by using the target data reasoning large model and the data reasoning demand to obtain a corresponding data reasoning result. Data reasoning is performed by using a local data reasoning large model of the edge device, so that the delay of data processing is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a multi-modal data reasoning method, device, equipment and storage medium. Background Art

[0002] With the development of artificial intelligence technology, using intelligent assistants for data analysis and reasoning operations has become a current trend.

[0003] However, currently, intelligent assistants mainly rely on the cloud computing mode when performing data reasoning. There are problems with relatively high latency in performing data reasoning operations through the cloud computing mode, and the timeliness of problem-solving cannot be guaranteed, thus reducing the user experience. Therefore, how to improve the real-time performance of data reasoning and reduce the latency in the output processing process has become a technical problem to be solved currently. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a multi-modal data reasoning method, device, equipment and storage medium, which can perform data reasoning by using a data reasoning large model on the edge device locally, reducing the latency of data processing. The specific scheme is as follows:

[0005] In the first aspect, the present application provides a multi-modal data reasoning method, which is applied to an edge device and includes:

[0006] Obtain the initial multi-modal data input by the target user, and perform data cleaning on the initial multi-modal data to obtain the corresponding target multi-modal data;

[0007] Perform lightweight processing on the preset initial data reasoning large model to obtain the corresponding target data reasoning large model, use the local target intelligent assistant to call the target data reasoning large model, and use the target data reasoning large model to analyze the target multi-modal data to determine the data reasoning requirements corresponding to the target user; wherein, the target data reasoning large model is a large model deployed locally;

[0008] Based on the preset dynamic time adjustment algorithm, fuse the target multi-modal data to obtain the corresponding fused data, and use the target data reasoning large model and the data reasoning requirements to perform data reasoning on the fused data to obtain the data reasoning result corresponding to the fused data; wherein, an incremental learning algorithm is pre-configured in the target intelligent assistant.

[0009] Optionally, the performing data cleaning on the initial multi-modal data includes:

[0010] Using a preset frequency recursive convolutional network to perform noise reduction processing on the voice data in the initial multimodal data, and performing adaptive gamma correction on the video data in the initial multimodal data;

[0011] Performing semantic analysis on the text data in the initial multimodal data, and modifying typos in the text data according to corresponding semantic analysis results;

[0012] Determine sensor data in the initial multimodal data, judge whether there is abnormal data in the sensor data, and if there is abnormal data in the sensor data, remove the abnormal data.

[0013] Optionally, the lightweight processing of the preset initial data inference large model to obtain the corresponding target data inference large model includes:

[0014] The precision of the output layer and the attention mechanism layer of the preset initial data reasoning large model is set to half precision, and the precision of the convolution layer and the fully connected layer of the preset initial data reasoning large model is quantized to eight bits to obtain the corresponding model to be pruned;

[0015] The model to be pruned is pruned to obtain a corresponding pruned model, and the pruned model is trained based on knowledge distillation technology to obtain a corresponding large model for inference of the target data.

[0016] Optionally, the fusing the target multimodal data based on a preset dynamic time adjustment algorithm includes:

[0017] The time series corresponding to each modality data in the target multimodal data are obtained, the alignment points between the time series are obtained using the preset dynamic time adjustment algorithm, and the target multimodal data are fused based on the alignment points.

[0018] Optionally, after performing data reasoning on the fused data using the target data reasoning large model and the data reasoning requirements, the method further includes:

[0019] Evaluate the data inference result to obtain a corresponding evaluation result;

[0020] The target parameters of the target intelligent assistant are regularly updated based on the evaluation results corresponding to the data inference results and the incremental learning algorithm.

[0021] Optionally, the multimodal data reasoning method further includes:

[0022] Monitor the local data changes in real time, so as to determine the newly added data locally based on the data changes in the local target storage space, and use the breakpoint resumption technology to upload the newly added data to the target cloud platform.

[0023] Optionally, the multimodal data inference method further includes:

[0024] Retrieve the target data in the local target storage space, quantify the sensitivity levels corresponding to the target data, and determine the target data with a sensitivity level less than the preset sensitivity threshold as non-sensitive data;

[0025] Compress the non-sensitive data and upload the compressed non-sensitive data to the target cloud platform.

[0026] In a second aspect, the present application provides a multimodal data inference device, which is applied to an edge device and includes:

[0027] A data cleaning module, configured to obtain the initial multimodal data input by a target user and perform data cleaning on the initial multimodal data to obtain corresponding target multimodal data;

[0028] A data analysis module, configured to perform lightweight processing on a preset initial data inference large model to obtain a corresponding target data inference large model, use the local target intelligent assistant to call the target data inference large model, and use the target data inference large model to analyze the target multimodal data to determine the data inference requirements corresponding to the target user; wherein, the target data inference large model is a large model deployed locally;

[0029] A data inference module, configured to fuse the target multimodal data based on a preset dynamic time adjustment algorithm to obtain corresponding fused data, and perform data inference on the fused data using the target data inference large model and the data inference requirements to obtain a data inference result corresponding to the fused data; wherein, an incremental learning algorithm is pre-configured in the target intelligent assistant.

[0030] In a third aspect, the present application provides an electronic device, including:

[0031] A memory, configured to store a computer program;

[0032] A processor, configured to execute the computer program to implement the foregoing multimodal data inference method.

[0033] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program, and when the computer program is executed by a processor, the foregoing multimodal data inference method is implemented.

[0034] This application first obtains the initial multimodal data input by the target user, and cleans the initial multimodal data to obtain the corresponding target multimodal data. Then, it performs lightweight processing on the preset initial data inference large model to obtain the corresponding target data inference large model. The local target intelligent assistant is used to call the target data inference large model, and the target data inference large model is used to analyze the target multimodal data to determine the data inference requirements corresponding to the target user. Among them, the target data inference large model is a large model deployed locally. Finally, based on the preset dynamic time adjustment algorithm, the target multimodal data is fused to obtain the corresponding fused data, and the target data inference large model and the data inference requirements are used to perform data inference on the fused data to obtain the data inference result corresponding to the fused data. Among them, an incremental learning algorithm is pre-configured in the target intelligent assistant. It can be seen that through the lightweight operation of the data inference large model, this application reduces the computing resources required by the large model, so that the large model can be deployed on edge devices. When the intelligent assistant calls the large model for data inference, there is no need to use the computing resources of the cloud, thereby reducing the latency in the data inference process. By configuring the incremental learning algorithm for the intelligent assistant, the intelligent assistant can continuously learn, thereby understanding the user's usage habits and improving the pertinence of the inference result. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0036] Figure 1 It is a flowchart of a multimodal data inference method disclosed in this application;

[0037] Figure 2 It is a flowchart of an intelligent assistant parameter update method disclosed in this application;

[0038] Figure 3 It is a schematic structural diagram of a multimodal data inference device disclosed in this application;

[0039] Figure 4 It is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] Currently, when intelligent assistants perform data reasoning, they mainly rely on the cloud computing mode, which has the problem of high latency in data reasoning operations. For this reason, this application provides a multi-modal data reasoning method, which reduces the latency of data processing by using the local data reasoning large model of edge devices.

[0042] See Figure 1 As shown, the embodiments of the present invention disclose a multi-modal data reasoning method, which is applied to edge devices and includes:

[0043] Step S11, obtain the initial multi-modal data input by the target user, and perform data cleaning on the initial multi-modal data to obtain the corresponding target multi-modal data.

[0044] The intelligent assistant system in this embodiment realizes efficient and intelligent user interaction and task management through technologies such as multi-modal data fusion, edge large model reasoning, personalized learning, and task automation. This embodiment first obtains environmental information through multi-modal data collection (including voice, video, text, sensor data, etc.), and improves the data quality through data preprocessing (such as noise reduction, feature extraction, etc.). Subsequently, the multi-modal data fusion module combines the cross-modal attention mechanism to deeply understand the user's intention, and with the support of edge large model reasoning, efficiently completes speech recognition, visual analysis, and natural language understanding locally. The system also has the ability of personalized learning, dynamically optimizing the model through incremental learning, enabling the intelligent assistant to adapt to user habits and provide accurate services. Finally, it combines the task automation and intelligent recommendation modules.

[0045] Based on environmental perception and user behavior prediction, the intelligent assistant can actively execute tasks, optimize the interaction experience, and interact with users in various ways, providing intelligent solutions for application scenarios such as smart homes, remote meetings, and medical assistance. For example, in the smart home scenario, the intelligent assistant can receive the user's instructions to turn on or off electrical appliances, or control the music player to play corresponding music according to the multi-modal data input by the user; in the field of medical assistance, the intelligent assistant can receive the multi-modal data input by the user, such as data such as medical images to be segmented and image segmentation requirements, and fuse the above multi-modal data based on a preset dynamic time adjustment algorithm, and assist in the segmentation process of medical images according to the fused multi-modal data and image segmentation requirements.

[0046] In this embodiment, first, it is necessary to obtain the initial multimodal data input by the user, such as audio, text, video and other data, or call the microphone array and camera to collect multimodal data such as voice and visual data, and use the FRCRN (Frequency Recurrent CRN) noise reduction algorithm to filter out environmental noise and extract the MFCC (Mel Frequency Cepstrum Coefficient) features of the multimodal data. It can be understood that the Mel Frequency Cepstrum Coefficient obtained after processing the initial multimodal data is one of the target multimodal data; correspondingly, the process of data cleaning for the initial multimodal data may specifically include: using a preset frequency recurrent convolutional recurrent network to perform noise reduction processing on the voice data in the initial multimodal data, and performing adaptive gamma correction on the video data in the initial multimodal data; performing semantic analysis on the text data in the initial multimodal data, and modifying the typos in the text data according to the corresponding semantic analysis results; determining the sensor data in the initial multimodal data, judging whether there is abnormal data in the sensor data, and if there is abnormal data in the sensor data, removing the abnormal data. By performing data cleaning operations on the obtained initial multimodal data, the reliability of the cleaned multimodal data is ensured, and thus the accuracy of the data inference result can be improved.

[0047] Step S12: Perform lightweight processing on the preset initial data inference large model to obtain the corresponding target data inference large model, use the local target intelligent assistant to call the target data inference large model, and use the target data inference large model to analyze the target multimodal data to determine the data inference requirements corresponding to the target user; wherein, the target data inference large model is a large model deployed locally.

[0048] In this embodiment, it is necessary to optimize the initial data inference large model to obtain the corresponding target data inference large model. It can be understood that the above target data inference large model is a large model in the intelligent assistant system. The steps for optimizing the initial data inference large model include:

[0049] Lightweight model deployment: Select a multimodal large model and compress the number of parameters through knowledge distillation.

[0050] Adopt mixed-precision quantization, such as FP16+INT8 (a data precision), to perform 8-bit quantization on non-sensitive layers (such as fully connected layers) to reduce the model size.

[0051] Inference acceleration technology: Use TensorRT (an inference model) or ONNX Runtime (an inference model) to deploy the model as the target data inference engine, and enable the INT8 acceleration core to improve the computing efficiency.

[0052] By pre - allocating the memory pool and fusing operators, such as merging Conv (convolution layer) + ReLU (activation layer), the inference latency is reduced.

[0053] Correspondingly, the process of lightweight processing of the pre - set initial data inference large model to obtain the corresponding target data inference large model can specifically include: setting the precision of the output layer and the attention mechanism layer of the pre - set initial data inference large model to half - precision, and performing 8 - bit quantization on the precision of the convolution layer and the fully - connected layer of the pre - set initial data inference large model to obtain the corresponding model to be pruned; performing model pruning on the model to be pruned to obtain the corresponding pruned model, and training the pruned model based on the knowledge distillation technology to obtain the corresponding target data inference large model; specifically, the process of performing lightweight operations on the initial data inference large model includes: based on the lightweight large - model optimization strategy, running an efficient deep - learning model on edge devices, and combining technologies such as knowledge distillation and pruning to optimize the computing efficiency, reduce the dependence on the cloud, improve the response speed and data security. The model is lightweight processed by using the method of mixed - precision quantization. First, perform hierarchical precision allocation on the large model. Precision - sensitive modules such as the model output layer and the attention mechanism layer retain half - precision (FP16) to ensure the integrity of key features. Compute - intensive modules such as the convolution layer and the fully - connected layer are quantized to 8 - bit integer (INT8) to reduce the amount of computation and memory occupancy. By statistically analyzing the activation value distribution, use asymmetric quantization to optimize the quantization range and reduce the precision loss. Use knowledge distillation to assist training. Use the cloud large model (FP32) as the teacher model to guide the mixed - precision student model (FP16 + INT8) at the edge. Transfer soft - label knowledge through the KL - divergence loss to improve the generalization ability. At the same time, add a pruning operation. First, perform structured pruning on the model (pruning rate 20% - 30%) to remove redundant parameters, and then perform grouped quantization on non - zero weights to further compress the model volume. Jointly optimize the task loss, the distillation loss, and the quantization loss:

[0054] ;

[0055] Among them, represents the task loss, which ensures the performance of the model on the target task; represents the knowledge distillation loss, which uses the knowledge of the teacher model to make up for the quantization loss; represents the quantization loss, which controls the quantization error and avoids model collapse. The weight coefficients are adjusted according to the task type and empirical values. By adjusting the weight coefficients, the best balance can be achieved among precision, compression rate, and generalization ability. For example, in a scenario with high - precision requirements, increase , while in a scenario with high - compression requirements, increase . In the later stage of training, , enhance the autonomy of the model.

[0056] By lightweight processing of the large data inference model, the large model's reliance on computing resources is reduced, enabling the large model to be deployed locally on edge devices. As a result, when using the intelligent assistant to perform data inference, it is no longer necessary to rely on the cloud, reducing data processing latency.

[0057] Step S13: Based on a preset dynamic time adjustment algorithm, fuse the target multimodal data to obtain corresponding fused data, and use the target data inference large model and the data inference requirements to perform data inference on the fused data to obtain the data inference result corresponding to the fused data; wherein, an incremental learning algorithm is pre-configured in the target intelligent assistant.

[0058] In this embodiment, the process of fusing the target multimodal data based on the preset dynamic time adjustment algorithm may specifically include: obtaining the time series corresponding to each modal data in the target multimodal data, using the preset dynamic time adjustment algorithm to obtain the alignment points between the time series, and fusing the target multimodal data based on the alignment points; specifically, constructing a two-stream Transformer architecture (a deep learning architecture), extracting multimodal information such as audio, video, text, and sensor data respectively, and through cross-modal attention mechanisms and multi-layer feature extraction networks, improving the data fusion effect, achieving more accurate environmental perception and user intention understanding, designing a dynamic timestamp calibration module to achieve spatio-temporal multimodal feature alignment. The DTW algorithm (Dynamic Time Warping) is an algorithm based on dynamic programming for measuring the similarity between two time series. An improved DTW algorithm (i.e., the preset dynamic time adjustment algorithm) is used to achieve cross-modal temporal synchronization, minimizing the cumulative distance of the feature sequences between modalities:

[0059] ;

[0060] Wherein, and respectively represent two input time series, represents finding the optimal path with the minimum cumulative distance among all possible warping paths. w represents aligning the two paths, , k is the path length, represents the alignment of the corresponding positions of the two sequences, is the square of the Euclidean distance between two alignment points.

[0061] This algorithm compensates for the difference in sensor sampling rates through dynamic path planning. For example, in a video conference, due to device latency, the audio and lip movements are out of sync, and it is necessary to align the speech and the picture in real time. Alignable features are extracted from the audio and video signals. The audio is collected at a sampling rate of 16 kHz, and the MFCC features are extracted every 10 ms. The video is collected at 30 fps, and the lip key points of each frame are extracted. Calculate the Euclidean distance between the feature sequences of the audio frames and the video frames, and find the path with the minimum cumulative distance through dynamic programming. According to the path mapping relationship, interpolate or dynamically change the speed of the audio to align it with the video frame rate. If the audio lags, temporarily store the video frames in a buffer and play the corresponding frames according to the path delay to achieve the time synchronization of the audio and video.

[0062] It should be noted that when performing multimodal data fusion in this embodiment, the weights of different modal data are also considered, and the data is fused based on the weights of each modal data. Specifically, feature interaction is achieved through the cross-attention mechanism, and the weight distribution formula is calculated:

[0063] ;

[0064] Among them, Q, K, and V are features from different modalities, that is, data of different modalities. Q (Query) is the query vector, representing the element or position that needs to be focused on currently. In the self-attention mechanism, K (Key) is the key vector, representing the element or position being queried. V (Value) is the value vector, representing the information related to the query. represents the transpose of K. is the scaling factor, usually taking the square root of the dimension of the key vector, which is used to prevent numerical stability problems caused by too large input to the softmax function. The softmax function is used to convert the input scores into a probability distribution, making the sum of all elements equal to 1 for subsequent processing, and splice the multimodal data (such as speech MFCC + image HOG features; HOG, that is, Histogram of Oriented Gradients, the histogram of oriented gradients), and input the target data to infer the large model for processing.

[0065] In this embodiment, after using the target data to infer the large model and the data inference requirements to perform data inference on the fused data, it further includes: evaluating the data inference results to obtain corresponding evaluation results; regularly updating the target parameters of the intelligent assistant based on the evaluation results corresponding to the data inference results and the incremental learning algorithm; that is, collecting the data and results generated during use, regularly performing data evaluation, and using the incremental learning strategy to regularly perform online updates on the model.

[0066] In addition, in this embodiment, it is also possible to monitor the local data change situation in real time, so as to determine the newly added local data based on the data change situation in the local target storage space, and use the breakpoint resumption technology to upload the newly added data to the target cloud platform; that is, through incremental upload (only transmitting differential data) and the breakpoint resumption mechanism to synchronize the data to the cloud, thereby reducing bandwidth consumption.

[0067] Moreover, in this embodiment, the target data in the local target storage space will also be retrieved simultaneously, the sensitivity corresponding to each target data will be quantified, and the target data with a sensitivity less than the preset sensitivity threshold will be determined as non-sensitive data; the non-sensitive data will be compressed and the compressed non-sensitive data will be uploaded to the target cloud platform; that is, sensitive data (such as face information) is only stored locally, and non-sensitive data (such as device logs) is compressed and uploaded to the cloud.

[0068] In addition, the edge device in this embodiment uses a local database (SQLite) encrypted by AES-256 (a kind of encryption technology) to store frequently accessed data (such as user preference configurations), and uses an edge caching strategy to retain hot data, reducing the frequency of cloud access.

[0069] Thus, through the lightweight operation of the data inference large model in this application, the computing resources required by the large model are reduced, so that the large model can be deployed on the edge device. When the intelligent assistant invokes the large model for data inference, it does not need to use the computing resources of the cloud, thereby reducing the latency in the data inference process; by configuring the incremental learning algorithm for the intelligent assistant, the intelligent assistant can continuously learn, thereby understanding the user's usage habits and improving the pertinence of the inference results.

[0070] Based on the previous embodiment, this application describes the overall process of fusing multi-modal data and using the fused multi-modal data for data inference. To make the technical solution in this application more complete, next, this application will elaborate on the process of updating the parameters of the intelligent assistant. See Figure 2 As shown, the embodiment of the present invention discloses a process for updating the parameters of an intelligent assistant, which is applied to an edge device and includes:

[0071] Step S21, use the target intelligent assistant to invoke the target data inference large model, use the target data inference large model to perform data inference on the fused multi-modal data to obtain corresponding data inference results, and evaluate the data inference results to obtain corresponding evaluation results.

[0072] In this embodiment, after obtaining the data inference result by using the target data to infer the large model, it is also necessary to evaluate the obtained data inference result to obtain the corresponding evaluation result, so as to adjust the target parameters of the intelligent assistant by using the above evaluation result and the incremental learning algorithm subsequently.

[0073] It should be noted that the incremental learning algorithm is pre-configured in the intelligent assistant in this embodiment. That is to say, the intelligent assistant in this embodiment has the ability of personalized learning for users and can automatically adjust its own behavior over time. Specifically, this embodiment proposes an intelligent assistant system (i.e., an incremental learning system) based on a dynamic evolvable network architecture and elastic parameter optimization. The above incremental learning algorithm is implemented through this system, and this system realizes the adaptive evolution of the model in a continuous data stream environment through multi-dimensional collaborative optimization. The core of the system adopts a modular neural network design and constructs a composite architecture composed of a fixed-parameter backbone feature extractor, an extensible dynamic branch network, and an adaptive routing controller.

[0074] The backbone feature extractor solidifies the core feature extraction ability through pre-training to ensure the stable retention of basic knowledge.

[0075] The dynamic branch network is composed of multiple lightweight convolution modules stacked together, and the topology of each module is dynamically optimized through neural architecture search. The specific implementation of the dynamic branch network is based on an improved ENAS (Efficient Neural Architecture Search) algorithm, and an adaptive topology optimization system is constructed by adopting a parameter sharing mechanism and a hardware-aware acceleration strategy.

[0076] The network architecture consists of three parts: a basic unit library, a search controller, and a dynamic compilation module. The basic unit library pre-sets 16 configurations of lightweight convolution modules, and each module contains a three-dimensional combination option of the number of channels {64, 128, 256}, the convolution kernel size {3×3, 5×5}, and the skip connection position (residual connection / dense connection / no connection). The search controller uses a double-layer LSTM (Long Short-Term Memory) network to generate architecture codes, and searches for the optimal configuration within a preset exploration space of 400 steps through MCTS (Monte Carlo Tree Search). The objective function simultaneously considers the multi-objective optimization of the validation set accuracy (weight 0.6), inference latency (weight 0.3), and energy consumption efficiency (weight 0.1). The dynamic compilation module integrates the just-in-time model conversion technology to convert the topological description of the selected architecture into an executable computation graph.

[0077] In terms of search acceleration, the system adopts a parameter inheritance strategy to reuse the weight parameters of the historical model, maps the old architecture parameters to the new topology through a weight transformation matrix, and adjusts the convolution kernel tensor by bilinear interpolation when the number of channels changes. The verification accuracy of the new architecture can reach more than 95% of the peak performance after 10 epochs of fine-tuning. The activation mechanism of the dynamic branch integrates a gated routing network. By analyzing the feature activation spectrum of the input data (using a sliding window to statistically analyze the L2 norm distribution of the first 128-dimensional features), a branch selection coefficient in the range of [0, 1] is generated in real time. When the coefficient exceeds the threshold it triggers the loading of the new branch and the redirection of the data stream, and at the same time the old branch enters the low-power freeze state until the next architecture update cycle.

[0078] By constructing the above incremental learning system, the intelligent assistant in this embodiment has the incremental learning function and can automatically adjust its own behavior over time, thereby improving the user experience.

[0079] Step S22: Regularly update the target parameters of the target intelligent assistant based on the evaluation result corresponding to the data inference result and the incremental learning algorithm.

[0080] The incremental learning system in this embodiment has a data real-time monitoring function. The realization of this function is based on a multi-modal data fusion architecture and a dynamic anomaly detection mechanism, and its core consists of a data acquisition layer, a real-time processing engine, and an adaptive feedback system. The data acquisition layer includes a variety of data sensors, which perform timestamp synchronization and normalization processing through edge computing nodes to monitor anomalies in real time. The dynamic threshold adjustment mechanism uses the sliding window statistical method to update the parameter reference value every 30 seconds. The adaptive adjustment rule for the window size is , the drift determination threshold is controlled by a dynamic P value, and FDR < 0.05 (False Discovery Rate, that is, the false discovery rate). When the KL divergence (Kullback-Leibler divergence, a statistical measure) between the data distribution of the newly input multi-modal data and the historical data (such as historical multi-modal data and the corresponding historical evaluation results of historical multi-modal data) exceeds the preset threshold (usually 0.3 - 0.5), it triggers the branch expansion mechanism to automatically generate a dedicated processing path adapted to the new features; the adaptive routing controller constructs a decision network using GRU (Gate Recurrent Unit), and dynamically allocates processing paths by analyzing the feature activation patterns of the input data, realizing on-demand scheduling of computing resources during the inference stage.

[0081] At the parameter optimization level, the system introduces an improved elastic weight consolidation algorithm and establishes a hierarchical constraint mechanism for parameter importance. Based on the second-order optimization theory, a parameter importance evaluation matrix is constructed. The weight sensitivity index is calculated by diagonal approximation of the Hessian matrix of the loss function. The specific calculation formula is:

[0082] ;

[0083] where represents the th parameter, L is the loss function, N represents the total number of samples, represents the partial derivative. A strong L2 regularization constraint is imposed on high-importance parameters (the top 20% quantile), and the constraint strength decays exponentially with the task interval time t. The decay coefficient is designed as , which can effectively inhibit catastrophic forgetting and avoid overly restricting the model adaptability.

[0084] In addition, the data management subsystem of the incremental learning system in this embodiment integrates a dual intelligent screening strategy to construct a collaborative working mechanism for a dynamic memory bank and real-time sample selection. The active learning filter uses the Monte Carlo Dropout method to estimate sample uncertainty and selects samples with a prediction entropy value higher than 2.0 for priority learning; the memory bank adopts a hierarchical storage structure, dividing the buffer area into a high-frequency area (storing the last 500 high-variance samples) and a schema area (storing representative samples after t-SNE clustering), and achieving the balance between old and new knowledge through an importance sampling weighted replay strategy. Memory replay introduces temperature-regulated Softmax weighted sampling, and the temperature coefficient is dynamically adjusted according to the model stability: . The storage capacity is dynamically adjusted according to the total number of processed samples, following the expansion rule, and setting a hardware safety upper limit of 2000 samples.

[0085] The system operation architecture adopts a three-layer optimization control system to achieve the dynamic balance of resources and performance: the micro layer implements online asynchronous stochastic gradient descent to support the instant parameter update of each batch of data; the middle layer realizes the sharded update of the model through a distributed parameter server, and uses the Ring-AllReduce communication protocol to reduce the gradient synchronization delay to the millisecond level; the macro layer deploys a hyperparameter optimizer driven by reinforcement learning, and automatically adjusts 10-dimensional hyperparameter spaces such as the learning rate and regularization strength every 24 hours based on the PPO algorithm (Proximal Policy Optimization).

[0086] It can be seen that in this application, by performing lightweight operations on the data inference large model, the computing resources required by the large model are reduced, so that the large model can be deployed on edge devices. When the intelligent assistant invokes the large model for data inference, it does not need to utilize the computing resources of the cloud, thereby reducing the latency in the data inference process; by configuring the intelligent assistant with an incremental learning algorithm, the intelligent assistant can continuously learn, thereby understanding the user's usage habits and improving the pertinence of the inference results.

[0087] See Figure 3 As shown, an embodiment of the present invention discloses a multimodal data inference device applied to an edge device, including:

[0088] A data cleaning module 11, configured to obtain initial multimodal data input by a target user, and perform data cleaning on the initial multimodal data to obtain corresponding target multimodal data;

[0089] A data analysis module 12, configured to perform lightweight processing on a preset initial data inference large model to obtain a corresponding target data inference large model, use a local target intelligent assistant to invoke the target data inference large model, and use the target data inference large model to analyze the target multimodal data to determine the data inference requirements corresponding to the target user; wherein, the target data inference large model is a large model deployed locally;

[0090] A data inference module 13, configured to fuse the target multimodal data based on a preset dynamic time adjustment algorithm to obtain corresponding fused data, and perform data inference on the fused data using the target data inference large model and the data inference requirements to obtain a data inference result corresponding to the fused data; wherein, an incremental learning algorithm is pre-configured in the target intelligent assistant.

[0091] It can be seen that in this application, by performing lightweight operations on the data inference large model, the computing resources required by the large model are reduced, so that the large model can be deployed on edge devices. When the intelligent assistant invokes the large model for data inference, it does not need to utilize the computing resources of the cloud, thereby reducing the latency in the data inference process; by configuring the intelligent assistant with an incremental learning algorithm, the intelligent assistant can continuously learn, thereby understanding the user's usage habits and improving the pertinence of the inference results.

[0092] In some specific embodiments, the data cleaning module 11 may specifically include:

[0093] A voice data cleaning unit, configured to perform noise reduction processing on the voice data in the initial multimodal data using a preset frequency recursive convolutional recurrent network, and perform adaptive gamma correction on the video data in the initial multimodal data;

[0094] A text data cleaning unit for performing semantic analysis on the text data in the initial multi-modal data and modifying typos in the text data according to the corresponding semantic analysis results;

[0095] A sensor data cleaning unit for determining the sensor data in the initial multi-modal data, judging whether there is abnormal data in the sensor data, and if there is abnormal data in the sensor data, removing the abnormal data.

[0096] In some specific embodiments, the data analysis module 12 may specifically include:

[0097] A model precision setting unit for setting the precision of the output layer and the attention mechanism layer of the preset initial data inference large model to half precision, and performing eight-bit quantization on the precision of the convolutional layer and the fully connected layer of the preset initial data inference large model to obtain a corresponding model to be pruned;

[0098] A model training unit for pruning the model to be pruned to obtain a corresponding pruned model, and training the pruned model based on the knowledge distillation technology to obtain the corresponding target data inference large model.

[0099] In some specific embodiments, the data inference module 13 may specifically include:

[0100] A data fusion unit for obtaining the time series corresponding to each modal data in the target multi-modal data, using the preset dynamic time adjustment algorithm to obtain the alignment points between the time series, and fusing the target multi-modal data based on the alignment points.

[0101] In some specific embodiments, the data inference module 13 further includes:

[0102] A result evaluation unit for evaluating the data inference result to obtain a corresponding evaluation result;

[0103] A parameter update unit for periodically updating the target parameters of the target intelligent assistant based on the evaluation result corresponding to the data inference result and the incremental learning algorithm.

[0104] In some specific embodiments, the multi-modal data inference device further includes:

[0105] A first data upload module for monitoring the local data change situation in real time, so as to determine the newly added data locally based on the data change situation in the local target storage space, and uploading the newly added data to the target cloud platform using the breakpoint resume technology.

[0106] In some specific embodiments, the multimodal data inference device further includes:

[0107] A data retrieval module, configured to retrieve target data in the local target storage space, quantify the sensitivity levels corresponding to the respective target data, and determine the target data with a sensitivity level less than a preset sensitivity threshold as non-sensitive data;

[0108] A second data upload module, configured to compress the non-sensitive data and upload the compressed non-sensitive data to the target cloud platform.

[0109] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 4 which is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be considered as any limitation to the scope of use of the present application.

[0110] Figure 4 This is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the multimodal data inference method disclosed in any of the foregoing embodiments. Additionally, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0111] In this embodiment, the power supply 23 is used to provide operating voltages for the various hardware devices on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitation is made here.

[0112] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.

[0113] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the multi-modal data inference method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.

[0114] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the multi-modal data inference method disclosed above is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0115] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0116] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0117] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0118] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0119] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A multimodal data inference method, characterized in that Applied to edge devices, including: Acquire initial multimodal data input by a target user, and perform data cleaning on the initial multimodal data to obtain corresponding target multimodal data; Lightweight processing is performed on the preset initial data inference big model to obtain the corresponding target data inference big model, the target data inference big model is called by the local target intelligent assistant, and the target data inference big model is used to analyze the target multimodal data to determine the data inference requirements corresponding to the target user; wherein the target data inference big model is a big model deployed locally; The target multimodal data is fused based on a preset dynamic time adjustment algorithm to obtain corresponding fused data, and data inference is performed on the fused data using the target data inference big model and the data inference requirements to obtain data inference results corresponding to the fused data; wherein an incremental learning algorithm is pre-configured in the target intelligent assistant.

2. The multimodal data inference method according to claim 1, wherein The performing data cleaning on the initial multimodal data includes: Using a preset frequency recursive convolutional network to perform noise reduction processing on the voice data in the initial multimodal data, and performing adaptive gamma correction on the video data in the initial multimodal data; Performing semantic analysis on the text data in the initial multimodal data, and modifying typos in the text data according to corresponding semantic analysis results; Determine sensor data in the initial multimodal data, judge whether there is abnormal data in the sensor data, and if there is abnormal data in the sensor data, remove the abnormal data.

3. The multimodal data inference method according to claim 1, wherein The lightweight processing of the preset initial data reasoning large model to obtain the corresponding target data reasoning large model includes: The precision of the output layer and the attention mechanism layer of the preset initial data reasoning large model is set to half precision, and the precision of the convolution layer and the fully connected layer of the preset initial data reasoning large model is quantized to eight bits to obtain the corresponding model to be pruned; The model to be pruned is pruned to obtain a corresponding pruned model, and the pruned model is trained based on knowledge distillation technology to obtain a corresponding large model for inference of the target data.

4. The multimodal data inference method according to claim 1, wherein The fusing the target multimodal data based on a preset dynamic time adjustment algorithm includes: The time series corresponding to each modality data in the target multimodal data are obtained, the alignment points between the time series are obtained using the preset dynamic time adjustment algorithm, and the target multimodal data are fused based on the alignment points.

5. The multimodal data reasoning method according to claim 1, wherein After performing data reasoning on the fused data using the target data reasoning large model and the data reasoning requirements, the method further includes: Evaluate the data inference result to obtain a corresponding evaluation result; The target parameters of the target intelligent assistant are regularly updated based on the evaluation results corresponding to the data inference results and the incremental learning algorithm.

6. The multimodal data reasoning method according to claim 1, characterized in that Also includes: Monitor the local data changes in real time to determine the newly added local data based on the data changes in the local target storage space, and use the breakpoint resumption technology to upload the newly added data to the target cloud platform.

7. The multimodal data reasoning method according to claim 6, wherein It further includes: Retrieve the target data in the local target storage space, quantify the sensitivity corresponding to each target data, and determine the target data with a sensitivity less than the preset sensitivity threshold as non-sensitive data; Compress the non-sensitive data and upload the compressed non-sensitive data to the target cloud platform.

8. A multimodal data inference device, characterized in that, Applied to edge devices, it includes: A data cleaning module for obtaining the initial multimodal data input by the target user and cleaning the initial multimodal data to obtain the corresponding target multimodal data; A data analysis module for lightweight processing of the preset initial data inference large model to obtain the corresponding target data inference large model, using the local target intelligent assistant to call the target data inference large model, and using the target data inference large model to analyze the target multimodal data to determine the data inference requirements corresponding to the target user; wherein, the target data inference large model is a large model deployed locally; A data inference module for fusing the target multimodal data based on the preset dynamic time adjustment algorithm to obtain the corresponding fused data, and using the target data inference large model and the data inference requirements to perform data inference on the fused data to obtain the data inference result corresponding to the fused data; wherein, an incremental learning algorithm is pre-configured in the target intelligent assistant.

9. An electronic device, characterized in that, It includes: A memory for storing computer programs; A processor for executing the computer programs to implement the multimodal data inference method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, For storing computer programs, the computer programs, when executed by the processor, implement the multimodal data inference method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Power system scheduling method and device based on multi-modal data fusion

    CN120879789A