Large model optimized integrated sensor multi-modal data edge computing system and method

Through the integrated sensor multimodal data edge computing system optimized by a large model, the problems of simple fusion strategy, lack of flexibility in modal weight adjustment and difficulty in introducing new modes in multimodal fusion are solved, and more efficient and flexible multimodal data processing and stronger adaptability are achieved.

CN120337131APending Publication Date: 2025-07-18WUHAN UNIV OF TECH

Patent Information

Application Number
CN202510391494.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In actual application of existing multimodal fusion technology, there are problems such as simple fusion strategy, lack of flexibility in modal weight adjustment and difficulty in introducing new modes, resulting in poor fusion effect and reduced model generalization ability.

Method used

A large-modal data edge computing system for integrated sensor optimization is designed. Through data preprocessing, modal specific coding, fusion strategy selection and execution, cross-modal coding and task processing modules, we deeply consider the relationship between modals, dynamically adjust the modal weights, optimize weight allocation, and reduce calculation costs and overfitting risks.

Benefits of technology

It improves the performance and adaptability of multimodal data processing, improves the fusion effect and flexibility between modes, and enhances the adaptability and generalization capabilities introduced by new modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337131A_ABST
    Figure CN120337131A_ABST
Patent Text Reader

Abstract

The invention discloses an integrated sensor multi-modal data edge computing system optimized by a large model. The integrated sensor multi-modal data edge computing system comprises a data preprocessing module which obtains pre-training parameters of each modal data through the large model; the modal specific coding module optimizes attention weight through an encoder branch to determine key information of data of each modal, and advanced feature representation of each modal is obtained; the fusion strategy selection and execution module is used for calculating the correlation between the rest modes and the final core mode, and inputting the advanced feature representation of the mode with the strongest correlation in the rest modes and the advanced feature representation of the final core mode into a cross-mode encoder to generate the fusion feature representation of the fused modes; continuing to perform modal fusion operation until all modalities are fused, and generating final fusion feature representation; and the cross-modal coding and task processing module finally generates a multi-modal fusion model prediction result through cross-modal feature representation. According to the invention, the performance and adaptability of multi-modal data processing and the data processing efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and specifically refers to an edge computing system and method for integrating sensor multimodal data optimized by a large model. Background Art

[0002] In today's digital age, multimodal data plays an increasingly important role in many key fields. With the continuous progress of sensor technology and the significant improvement of data acquisition capabilities, the quantity and variety of multimodal data are increasing rapidly, including image data, text data, audio data, video data, and sensor data. These data not only contain rich information but also can complement each other, providing a more comprehensive perspective for solving complex problems. However, there are many significant deficiencies in the existing multimodal fusion technologies in practical applications, which seriously restrict the in-depth development and wide application of multimodal data in integrated multi-sensor scenarios.

[0003] First, the fusion strategy is simple. Traditional multimodal fusion methods mostly adopt ways such as direct splicing or fixed-weight weighted averaging, without deeply considering the complex internal relationships between different modalities and unable to effectively mine potential features, resulting in poor fusion effects. For example, in the field of intelligent security, it is difficult to accurately capture the connection between the human behavior in the image and the sound event in the audio, and in the autonomous driving task, it is impossible to effectively integrate various modal information. For example, Chinese Patent Application No. CN202410073311.7, with the publication date of April 19, 2024, and the invention creation name of "A Video Content Analysis Method and System Based on Multimodal Fusion", although this solution is dedicated to multimodal fusion for video content analysis, the fusion strategy is relatively basic, without fully mining the deep-level correlation information between different modalities, and it is difficult to achieve high-precision video content analysis.

[0004] Second, the adjustment of modal weights lacks flexibility. Most of the existing technologies rely on fixed weights or weights set based on experience and cannot dynamically adjust according to the actual contributions of modalities in specific tasks. In actual multimodal task scenarios, the importance of different modalities varies greatly and continuously in different task links and environmental conditions. For example, in autonomous driving, the importance of camera images and radar data is different under different weather conditions; in medical diagnosis, the dependence of different disease diagnoses on medical images and medical record texts is different. Chinese Patent Application No. CN202311637372.2, with the publication date of March 22, 2024, and the invention creation name of "A Health Status Assessment Method and System for Multimodal Data Fusion", adopts pre-set weights to fuse multimodal data when conducting health status assessment and cannot dynamically adjust the weights of each modality according to the specific conditions of different patients and the health data at different stages.

[0005] Finally, it is difficult to introduce new modalities. When introducing new modalities in the application scenarios of integrating multi-sensors, existing methods often require large-scale adjustment and retraining of the fusion model, consuming a large amount of computing resources and time, and are also prone to overfitting problems, resulting in a decline in the generalization ability of the model. For example, when introducing a new environmental sensor modality in a smart home system, or a new behavior recognition sensor modality in a smart security system, existing methods cannot effectively fuse the new modality data, and may even reduce the system performance. Chinese Patent Application No. CN202311744302.6, with the publication date of April 12, 2024, and the invention creation name of "An Industrial Equipment Fault Diagnosis Method and System Based on Multi-Modal Data Fusion", in the industrial equipment fault diagnosis of this solution, when introducing a new sensor data modality, it is necessary to re-adjust and train the entire fusion model complexly, which is not only time-consuming and laborious, but also prone to overfitting, affecting the accuracy and generalization ability of fault diagnosis.

[0006] In summary, there are many significant deficiencies in the existing technologies in aspects such as fusion strategies, weight adjustment, and new modality processing. These problems are particularly prominent in the application scenarios of multi-modal data processing by integrating multi-sensors, severely restricting the in-depth development and wide application of multi-modal fusion technologies. Summary of the Invention

[0007] The purpose of the present invention is to provide an edge computing system and method for integrating sensor multi-modal data optimized by a large model, which improves the performance and adaptability of multi-modal data processing and data processing efficiency.

[0008] To achieve this purpose, the edge computing system for integrating sensor multi-modal data optimized by a large model designed by the present invention includes:

[0009] The data preprocessing module is used to collect multi-modal data through the data acquisition system, and the large model performs normalization and enhancement operations on different multi-modal data respectively to obtain the pre-training parameters of each multi-modal data;

[0010] The modality-specific encoding module is used to construct a corresponding encoder branch for each modality. The encoder branch corresponding to each modality adjusts the pre-training parameters of the corresponding modality data according to the characteristics of the corresponding modality data and the specific task requirements, and optimizes the attention weights of the corresponding modality data, and determines the key information of each modality data according to the attention weights, so as to obtain the high-level feature representation of each modality;

[0011] The fusion strategy selection and execution module is used to determine the final core modality by calculating the variance and information entropy of each modality data and combining expert knowledge, calculate the correlation between the remaining modalities except the final core modality and the final core modality through correlation analysis, and input the high-level feature representation of the modality with the strongest correlation among the remaining modalities and the high-level feature representation of the final core modality into the cross-modal encoder for modality fusion operation to generate the fusion feature representation of the fused modality. Calculate the correlation between the remaining modalities except the fused modality and the fused modality through correlation analysis. After each modality fusion operation, it is necessary to adjust the fusion weights of the fused modality and the modality with the strongest correlation among the remaining modalities using the multi-modal dataset. Input the high-level feature representation of the modality with the strongest correlation among the remaining modalities and the high-level feature representation of the fused modality into the cross-modal encoder for modality fusion operation to generate the fusion feature representation of the new fused modality until all modalities are fused and the final fusion feature representation is generated;

[0012] The cross-modal encoding and task processing module is used to input the final fusion feature representation into the cross-modal encoder to generate the cross-modal feature representation, and input the cross-modal feature representation into the task-specific decoder designed according to the specific task requirements to finally generate the prediction result of the multi-modal fusion model.

[0013] Advantages of the present invention:

[0014] The present invention can deeply consider the internal relationship between different modalities, effectively extract the key information of different modality data, and improve the fusion effect between modalities; the present invention can be dynamically adjusted according to the actual contribution of each modality data in the specific task requirements, improve the accuracy and flexibility of modality fusion, and enhance the performance and adaptability of multi-modal data processing; the present invention can deeply and comprehensively analyze the multi-dimensional correlation between the newly introduced modality and the fused modality and modality combination, optimize the weight allocation, and provide flexible fusion method suggestions according to the characteristics of the newly introduced modality and the specific task requirements, reduce the calculation cost and overfitting risk, and enhance the generalization ability of the multi-modal fusion model. Description of the drawings

[0015] Figure 1 is the structural schematic diagram of the present invention;

[0016] Figure 2 is the flow chart of the present invention. Detailed implementation manners

[0017] The following further elaborates the present invention in detail with reference to the drawings and specific embodiments:

[0018] Embodiment 1

[0019] An edge computing system for integrated sensor multi-modal data optimized by a large model, as Figure 1 shown, it includes:

[0020] The data preprocessing module is used to collect various modal data (including image data, text data, audio data, video data, and sensor data) through a data acquisition system. The large model performs normalization and enhancement operations on different modal data respectively to obtain the pre-training parameters of each modal data. This design normalizes and optimizes each modal data through normalization and enhancement operations, which can increase the diversity of each modal data, reduce the redundant information of each modal data, reduce the dimension of each modal data, and lighten the computational burden of subsequent processing;

[0021] In the above data preprocessing module, the large model automatically adapts and executes precise preprocessing strategies for various modal data according to a preset modal feature recognition model and a task requirement adaptation model; the modal feature recognition model can recognize features including but not limited to image data, text data, audio data, video data, and sensor data, and the task requirement adaptation model refines the preprocessing parameters according to the specific application scenario to maximize the utility of the modal data;

[0022] The above data acquisition system can widely receive multi-modal data, including image data, text data, audio data, video data, and sensor data; in the intelligent transportation scenario, the image data is obtained by on-vehicle cameras and is used to capture road conditions and information about surrounding vehicles and pedestrians; the text data can come from traffic rule documents and road condition descriptions; the audio data can be vehicle horn sounds and warning sounds; the video data records the entire process of vehicle driving; the sensor data includes the data collected by speed sensors and distance sensors, and these collected data together constitute the original source of multi-modal data;

[0023] The above large model plays a key role in the data preprocessing module. According to the characteristics of different modal data and task requirements, it provides precise strategies for the preprocessing of each modal data; in the image modality, it automatically determines the cropping area and scaling ratio to highlight key information; in the text modality, it conducts in-depth semantic analysis and cleaning with the help of a large-scale text corpus to remove redundant information; in the audio modality, it optimizes the frame extraction strategy according to the audio rhythm and frequency to accurately extract key information; it is also applied in the subsequent modality-specific encoding module, fusion strategy selection and execution module, and model training and optimization module, such as providing pre-training parameters for modality-specific encoding, assisting in determining the fusion strategy, and helping to adjust the modality weights;

[0024] The modality-specific encoding module is used to construct a corresponding encoder branch for each modality (the corresponding encoder branch is the corresponding Transformer encoder branch). The encoder branch corresponding to each modality adjusts the pre-trained parameters of the corresponding modality data according to the characteristics of the corresponding modality data and the specific task requirements, and optimizes the attention weights of the corresponding modality data. The key information of each modality data is determined according to the attention weights, and the high-level feature representations of each modality are obtained based on the key information of each modality data. This design can perform specialized processing and optimization for the characteristics of each modality data, so as to better extract and represent the key information of each modality data;

[0025] In the above modality-specific encoding module, the parameter adaptation process follows the modality characteristic analysis rule and the task-oriented optimization criterion, and performs modality-differentiated adjustment on the pre-trained parameters of each modality data obtained by the large model; the modality characteristic analysis rule is based on the intrinsic statistical characteristics of the data, and the task-oriented optimization criterion combines the specific task to specifically adjust the architecture parameters of the Transformer encoder to achieve efficient feature encoding of various modality data;

[0026] The above encoder branch constructs a corresponding Transformer encoder branch for each modality, that is, one modality corresponds to an exclusive Transformer encoder branch; taking the text, image, and audio modalities as examples, the text modality has a Transformer encoder branch specially constructed for it, and the image and audio modalities also have their own corresponding independent Transformer encoder branches; the encoder branch is adopted to process data of different modalities more specifically. Data of different modalities have different characteristics, and a single encoder is difficult to meet the complex needs of all modality data at the same time. The Transformer encoder branch corresponding to each modality can flexibly adjust parameters and architecture according to the unique properties of the corresponding modality data to achieve efficient feature encoding of each modality data and improve the performance and flexibility of the entire multi-modal data processing system;

[0027] The fusion strategy selection and execution module is used to determine the final core modality by calculating the variance and information entropy of each modality data and combining expert knowledge, calculate the correlation between the remaining modalities except the final core modality and the final core modality through correlation analysis, and input the high-level feature representation of the modality with the strongest correlation among the remaining modalities and the high-level feature representation of the final core modality into the cross-modal encoder for modality fusion operation to generate the fusion feature representation of the fused modality. Calculate the correlation between the remaining modalities except the fused modality and the fused modality through correlation analysis. After each modality fusion operation, it is necessary to use a multi-modal data set (here, the "multi-modal data set" specifically refers to the "multi-modal small data set", that is, a relatively small-scale and representative multi-modal data subset, which plays a unique role in the training and adjustment of the multi-modal fusion model) to adjust the fusion weights of the fused modality and the modality with the strongest correlation among the remaining modalities. Input the high-level feature representation of the modality with the strongest correlation among the remaining modalities and the high-level feature representation of the fused modality into the cross-modal encoder for modality fusion operation to generate the fusion feature representation of the new fused modality. Continue to calculate the correlation between the remaining modalities except the new fused modality and the new fused modality through correlation analysis. Input the high-level feature representation of the modality with the strongest correlation among the remaining modalities except the new fused modality and the high-level feature representation of the new fused modality into the cross-modal encoder for modality fusion operation to generate the fusion feature representation of the new round of fused modality until all modalities are fused and the final fusion feature representation is generated. Through scientific calculation methods and expert knowledge, this design can effectively determine the correlation between modalities, thereby achieving efficient and optimized fusion. In addition, gradually fusing the modalities with the strongest correlation and adjusting the fusion weights at the same time can gradually optimize the fusion effect, avoid information redundancy and noise interference caused by one-time fusion, reduce unnecessary calculations, and improve the fusion efficiency;

[0028] In the above fusion strategy selection and execution module, the correlation analysis automatically matches and executes the algorithms in the preset correlation algorithm library according to the type characteristics of the modality data; the correlation algorithm library includes a variety of correlation analysis methods, and dynamically adjusts the algorithm parameters and selection logic according to the attributes of the modality data and the modality association evaluation provided by the large model, and is applicable to various modality data to achieve in-depth correlation mining and fusion decision optimization between modalities;

[0029] The cross-modal encoding and task processing module is used to input the final fused feature representation into the cross-modal encoder to generate the cross-modal feature representation, and then input the cross-modal feature representation into the task-specific decoder designed according to specific task requirements. Finally, the prediction result of the multi-modal fusion model (the multi-modal fusion model is the model in the multi-modal data fusion processing system optimized by the large model) is generated. Through the design of cross-modal encoding and task-specific decoder, this design realizes the effective conversion of multi-modal data into the target results required by specific tasks, meets the requirements of different tasks, improves the flexibility and adaptability of modal data processing, and improves the accuracy and reliability of task processing;

[0030] In the above cross-modal encoding and task processing module, the design of the task-specific decoder follows the principle of task-driven architecture adaptation and the orientation of performance optimization; the architecture adaptation principle selects the corresponding decoder basic architecture according to the task type and adjusts it in combination with successful architecture cases of the large model in similar tasks; the performance optimization orientation focuses on improving the output quality and response speed of the decoder, and uses the pre-trained feature mapping and optimization strategies provided by the large model to optimize the decoder to ensure that various modal features can be accurately and efficiently mapped to the task target space;

[0031] In addition to the prediction results of the multi-modal fusion model classification task and the prediction results of the multi-modal fusion model generation task, in some complex tasks, the multi-modal fusion model needs to output structured information. For example, in the medical image analysis task, the multi-modal fusion model not only needs to judge the disease category, but also needs to output the location and severity of the disease; in the intelligent transportation scenario, the multi-modal fusion model needs to output a comprehensive assessment of the traffic conditions and coping strategies.

[0032] In the above technical solution, the model training and optimization module is used to calculate the loss function value of the multi-modal fusion model based on the true label of the specific task and the prediction result of the multi-modal fusion model, calculate the precision rate of each modality on the corresponding modality data validation set to adjust the weights of each modality, use the backpropagation algorithm to update the parameters of the multi-modal fusion model, and dynamically adjust the learning rate and use the regularization method to optimize the multi-modal fusion model; through calculating the loss function value, the above design can clarify the accuracy degree of the current prediction, so as to provide a direction for subsequent parameter update; using the regularization method can prevent the multi-modal fusion model from overfitting and improve the generalization ability of the multi-modal fusion model. Adjusting the weights according to the precision rate of each modality on the corresponding modality data validation set can make the multi-modal fusion model pay more attention to the modality data that can meet the specific task requirements, so as to optimize the fusion effect of multi-modal data and further improve the overall performance of the multi-modal fusion model;

[0033] In the above model training and optimization module, the adjustment of modal weights follows a dynamic feedback adjustment mechanism and a large model-assisted decision-making process; the dynamic feedback adjustment mechanism calculates the weight adjustment vector in real time based on the change in precision rate of the modality on the validation set; the large model-assisted decision-making process uses its overall understanding of the multi-modal fusion model and its understanding of the task objectives to analyze the factors affecting the change in precision rate of each modality, and provides policy suggestions for weight adjustment and parameter optimization solutions, which are applicable to various types of modal data to achieve precise adjustment of modal weights and continuous optimization of model performance;

[0034] The above-mentioned task true label is the real data label used to measure the prediction accuracy of the multi-modal fusion model during the training process of the multi-modal fusion model; in different specific tasks, the manifestation form of the task true label is different; in the image classification task, if it is for disease diagnosis and classification of medical images, the task true label is the disease category actually corresponding to the medical image, such as "normal", "pneumonia", "tumor"; in the text sentiment analysis task, the task true label can be the real sentiment tendency expressed by the text, such as "positive", "negative", "neutral"; the task true label is an objectively existing real data label, which is used to clarify the gap between the prediction result of the multi-modal fusion model and the real situation during training, so as to continuously adjust the parameters to improve the prediction accuracy.

[0035] In the above technical solution, the above-mentioned various modal data include image data, text data, audio data, video data, and sensor data; the above design can provide rich semantic information through multiple modal data, and the fusion of these information can more accurately understand the semantics of the scene or object. By fusing multiple modal data, the advantages of different modalities can be utilized to make up for the deficiencies of a single modality, and different modal data can meet the needs of different tasks, thereby improving the prediction accuracy of specific task requirements.

[0036] In the above technical solution, the specific process of obtaining the pre-training parameters of the above-mentioned various modal data is as follows:

[0037] For the image modality, based on the learning results of a large amount of image data, the large model determines the cropping area and scaling ratio according to the image content and specific task requirements. The large model uses a pre-trained image recognition model to extract key features, and the formula is expressed as:

[0038] I′=f crop-scale (I)

[0039] where I represents the original image, I′ represents the cropped and scaled image obtained after being processed by the large model, and f crop-scale represents the cropping and scaling function determined by the large model, and the cropped and scaled image I′ obtained after being processed by the large model is the pre-training parameter of the image modality;

[0040] For the text modality, the large model learns from a large-scale text corpus, performs semantic analysis and cleaning on the text. When tokenizing, it provides an accurate tokenization scheme based on the context, and uses a pre-trained word vector model to generate word vectors, which is expressed by the formula:

[0041] T′ = f clean (T)

[0042] T tokenized = f tokenize (T′)

[0043] where T represents the original text, T′ represents the text cleaned by the large model, T tokenized represents the result after tokenization, and f clean represents the cleaning function, and f tokenize represents the tokenization function. The text T′ cleaned by the large model passes through the tokenization function f tokenize to obtain the tokenized result T tokenized which is the pre-training parameter for the text modality;

[0044] For the audio modality, the large model optimizes the frame extraction strategy according to the audio rhythm and frequency, dynamically adjusts the frame extraction interval, and obtains a frame sequence reflecting the essential features of the audio, and assists in extracting key audio event features. The formula is expressed as:

[0045] A frames = f frame-optimize (A)

[0046] where A represents the original audio, and A frames represents the audio frame sequence after optimized frame extraction, and f frame-optimize is the frame extraction function optimized by the large model. The audio frame sequence A frames after optimized frame extraction is the pre-training parameter for the audio modality;

[0047] For the video modality, the large model extracts key frames and performs content analysis on the video, removing redundant information. The formula is expressed as:

[0048] V keyframes = f keyframe-extract (V)

[0049] where V represents the video, and V keyframes represents the set of extracted key frames, and f keyframe-extract is the key frame extraction function. The set of extracted key frames V keyframes is the pre-training parameter for the video modality;

[0050] For the sensor modality, the large model performs data calibration and feature engineering based on the sensor type and task requirements. The formula is expressed as:

[0051] S′ = f sensor-process (S)

[0052] Among them, S represents sensor data, S′ represents the processed sensor data, and f sensor-process is a processing function for the sensor data, where the processed sensor data S′ is the pre-training parameter of the sensor modality;

[0053] After completing the preprocessing of each modality data, data quality detection is performed on the preprocessed each modality data. If the data quality does not meet the standard, it will return to the data preprocessing module for reprocessing until the data quality meets the standard (the large model evaluates and gives feedback on the data quality based on a large amount of data and knowledge it has learned itself), and the pre-training parameters of each modality data are obtained. The above design constructs a comprehensive and efficient data input and preprocessing system through the refined preprocessing of different modality data and a strict data quality detection process. This system utilizes the powerful learning and analysis capabilities of the large model to fully explore the potential value of each modality data, obtains data from multiple data sources, performs targeted processing according to different modality characteristics, greatly improves the usability and accuracy of the data, not only effectively removes the noise and redundancy in the original data, but also enhances the key information, laying a solid and reliable foundation for subsequent multi-modal data fusion, analysis, and application assisted by the large model, and strongly promoting the practical application and in-depth development of multi-modal data processing technology in various fields.

[0054] In the above technical solution, the specific method for extracting the key information of each modality data to obtain the high-level feature representation of each modality is as follows:

[0055] Construct a corresponding Transformer encoder branch for each modality. Each Transformer encoder branch corresponding to each modality can adaptively learn and capture the key feature information in each modality data according to the characteristics of each modality data, and adjust and optimize the attention weights of the pre-training parameters of each modality data obtained through the large model in combination with the specific task requirements. The specific process is as follows:

[0056] For the image modality, according to the resolution and color channel number of the image, adjust the convolution kernel size and stride, and optimize the attention weights of different regions under the multi-head self-attention mechanism; the large model will analyze the resolution (H, W) and color channel number C of the image, and adjust the original convolution kernel size k and stride s in combination with the pre-training knowledge and the current task requirements; at the same time, divide the image into multiple regions [r1, r2,..., r m , analyze the semantic information of each region, and optimize the attention weight matrix A before adjustment. The formula is expressed as:

[0057] k′ = f image-k (H, W, C, k)

[0058] s' = f image-s (H, W, C, s)

[0059] A' = f image-attn-adjust (A, [r1, r2,..., r m )

[0060] Among them, (H, W) represents the resolution of the image; C represents the number of color channels; k represents the size of the original convolution kernel; s represents the original stride; k' represents the size of the adjusted convolution kernel; s' represents the adjusted stride; f image-k represents a function for adjusting the size of the convolution kernel based on the image modality features; f image-s represents a function for adjusting the stride based on the image modality features; A represents the attention weight matrix before adjustment; [r1, r2,..., r m represents multiple regions into which the image is divided; A' represents the adjusted attention weight matrix; f image-attn-adjust represents a function for adjusting the attention weight according to the image semantic information;

[0061] For the text modality, according to the length and vocabulary of the text, adjust the word embedding dimension and the number of attention heads, and optimize the attention weight between words; the large model evaluates the length L and vocabulary V of the text, and adjusts the original word embedding dimension d embedding and the number of attention heads n heads ; at the same time, analyze the semantic relevance of the word sequence [w1, w2,..., w n in the text, and optimize the attention weight matrix A before adjustment, which is expressed by the formula:

[0062] d' embedding = f text-d (L, V, d embedding )

[0063] n' heads = f text-n (L, V, n heads )

[0064] A' = f text-attn-adjust (A, [w1, w2,..., w n )

[0065] Among them, L represents the text length; V represents the vocabulary; d embedding represents the original word embedding dimension; n heads represents the original number of attention heads; d' embedding represents the adjusted word embedding dimension; n' heads represents the adjusted number of attention heads; f text-dA function for adjusting the word embedding dimension based on text modal features; f text-n A function for adjusting the number of attention heads based on text modal features; A represents the attention weight matrix before adjustment; [w1, w2,..., w n represents the word sequence in the text; A′ represents the attention weight matrix after adjustment; f text-attn-adjust A function for adjusting the attention weights according to the semantic relevance of the text;

[0066] For the audio modality, according to the frequency distribution and duration of the audio, adjust the window size and stride of audio feature extraction, and optimize the attention weights between different audio segments; the large model analyzes the frequency distribution F and duration T of the audio, and adjusts the original window size w and stride p; at the same time, the audio is divided into multiple segments [a1, a2,..., a q , analyze the semantic information of each segment, and optimize the attention weight matrix A before adjustment, which is expressed by the formula:

[0067] w′ = f audio-w (F, T, w)

[0068] p′ = f audio-p (F, T, p)

[0069] A′ = f audio-attn-adjust (A, [a1, a2,..., a q )

[0070] where, F represents the frequency distribution of the audio; T represents the audio duration; w represents the original window size; p represents the original stride; w′ represents the adjusted window size; p′ represents the adjusted stride; f audto-w represents a function for adjusting the window size based on audio modal features; f audio-p represents a function for adjusting the stride based on audio modal features; A represents the attention weight matrix before adjustment; [a1, a2,..., a q represents the multiple segments into which the audio is divided; A′ represents the attention weight matrix after adjustment; f audio-attn-adjust represents a function for adjusting the attention weights according to the audio semantic information;

[0071] By constructing corresponding Transformer encoder branches for each modality during the modality-specific encoding stage, adapting and optimizing the pre-trained parameters of each modality's data according to the characteristics of each modality's data and the specific task requirements, and accurately extracting the key information of the corresponding modality's data through the attention weights of each modality's data after adjustment to obtain the high-level feature representation of the corresponding modality; the above design gives full play to the pre-training advantages of the large model, realizes the accurate extraction of key information from multi-modal data, and starting from the data characteristics, adjusts the key parameters and optimizes the attention weights for the image modality, text modality, and audio modality respectively, can adaptively capture the key information in each modality's data, significantly improves the flexibility and pertinence of multi-modal data processing, and lays a solid foundation for the subsequent fusion and analysis of multi-modal data.

[0072] In the above technical solution, the specific process of generating the final fusion feature representation is as follows:

[0073] By calculating the variance (the variance reflects the degree of dispersion of the data) and information entropy (the information entropy measures the uncertainty and information content of the data) of each modality's data, select the modality with the largest amount of information as the initially selected core modality (in a multi-modal dataset containing images, text, and audio, if the variance of the image data is large and the information entropy shows that it contains more unique information, select the image modality as the initially selected core modality), and combine expert knowledge to determine the final core modality (experts will evaluate and correct the selection of the final core modality according to their professional experience in this field and their understanding of the specific task requirements to ensure that the selected final core modality meets the actual needs and data characteristics), which is expressed by the formula:

[0074] M core-initial

[0075] =f initial-select (Var1,Entropy1,Var2,Entropy2,...,Var n ,Entropy n )

[0076] M core =f expert-adjust (M core-initial )

[0077] Among them, Var i represents the variance of the i-th modality's data; Entropy i represents the information entropy of the i-th modality's data; M core-initial represents the initially selected core modality; M core represents the final core modality after expert adjustment; f initial-select represents the function of the initially selected core modality based on data statistical indicators; f expert-adjustA function for an expert to adjust the final core modality according to professional knowledge;

[0078] Calculate the correlation between the remaining modalities except the final core modality and the final core modality through correlation analysis and determine the modality fusion order. This process can be expressed by the following formula:

[0079] Corr i =f corr-select (M core ,M i )

[0080] Order=sort(Corr1,Corr2,...,Corr n )

[0081] Where Corr i represents the correlation between the i-th modality and the final core modality; M i represents the i-th modality; f corr-select represents a function that selects the corresponding correlation analysis method according to the type and characteristics of the modality data and calculates the correlation; Order represents the modality fusion order after sorting according to the correlation; sort(Corr1,Corr2,...,Corr n ) represents the modality fusion order after sorting according to the correlation, and only selects the modality with the strongest correlation for fusion;

[0082] The above correlation analysis method is a pre-set set of common correlation analysis methods, such as Pearson correlation coefficient, mutual information, and cosine similarity; according to the type and characteristics of the modality data, automatically select the appropriate method to calculate the correlation between modalities; for example, for numerical modality data, select the Pearson correlation coefficient; for discrete modality data, select mutual information; the large model plays an important role in this process, and it can provide suggestions for the selection of the correlation analysis method according to its own learning and understanding, and assist in the comprehensive evaluation of the correlation;

[0083] Starting from the determined final core modality, each time select the remaining modality with the strongest correlation with the fused modality among the remaining modalities except the already fused modalities for fusion; input the high-level feature representation of the modality with the strongest correlation among the remaining modalities and the high-level feature representation of the already fused modalities into the cross-modal Transformer encoder for modality fusion operation. The cross-modal Transformer encoder uses its multi-head self-attention mechanism and feed-forward neural network layer to interact and integrate the high-level feature representations of different modalities to generate a new fused feature representation;

[0084] Input the high-level feature representation H of the final core modality coreThe high-level feature representation H1 of the remaining modalities with the strongest correlation with the final core modality in all modalities except the final core modality is input into the cross-modal Transformer encoder to obtain the fused feature representation H of the fused modalities. fusion-1 , which is expressed by the formula:

[0085] H fusion-1 = f cross-modal-encoder (H core , H1)

[0086] where f cross-modal-encoder represents the fusion function of the cross-modal Transformer encoder;

[0087] After each modality fusion operation, it is necessary to use the multi-modal dataset to adjust the fusion weights of the fused modalities and the modality with the strongest correlation in the remaining modalities except the fused modalities. (Adjusting the fusion weights can enable the multi-modal fusion model to better adapt to the new fused features, and at the same time ensure that the fusion process conforms to the actual semantic logic. The large model can provide initialization parameters and adjustment strategies for the fusion weight adjustment process to help the multi-modal fusion model converge and optimize faster.) The adjusted fused feature representation can be expressed as:

[0088] H fusion-1-fine-tuned = f fine-tune (H fusion-1 , Small-dataset)

[0089] where f fine-tune represents the function of adjusting the fused features using the dataset, and Small-dataset represents the dataset used for adjustment;

[0090] Repeat the above steps of modality fusion and adjustment until all modalities are fused to generate the final fused feature representation H final-fusion (the final fused feature representation with high semantic abstraction and multi-modal information cohesion);

[0091] The entire fusion process can be summarized as an iterative process:

[0092] H fusion-j = f cross-modal-encoder (H fusion-(j-1) , H j )

[0093] H fusion-j-fine-tuned = f fine-tune (H fusion-j , Small-dataset)

[0094] where j represents the number of fusion rounds; H j represents the high-level feature representation of the j-th modality to be fused; Hfusion-j denotes the fused feature obtained after the j-th round of modality fusion; H fusion-(j-1) denotes interacting and integrating with H j to make the features of different modalities complement and fuse with each other, generating a new fused feature H fusion-j ; H fusion-j-fine-tuned denotes the result after the j-th round of fused feature is adjusted by the dataset; through continuous iteration, the information of all modalities is gradually fused and the final fused feature representation is generated; the above design provides a solid data foundation for the subsequent cross-modal encoding and task processing stages. By selecting an appropriate correlation analysis method according to the type and characteristics of the modality data and conducting comprehensive evaluation with the assistance of a large model, the correlation degree between modalities can be accurately grasped, so as to determine a reasonable fusion order, enabling the data of different modalities to be fused in an optimal way; the use of the cross-modal Transformer encoder and the adjustment of the fusion weights by the dataset before each fusion operation effectively enhances the interaction and integration between the features of different modalities, improving the quality and semantic expression ability of the fused features; the finally generated final fused feature representation fully integrates the information of multi-modal data, has high semantic abstraction and multi-modal information cohesion, provides a richer and more accurate data foundation for the subsequent cross-modal encoding and task processing stages, and helps to improve the performance and effect of the entire multi-modal data fusion processing system in various practical tasks.

[0095] In the above technical solution, the specific method of inputting the final fused feature representation into a cross-modal Transformer encoder to generate a cross-modal feature representation and then inputting the cross-modal feature representation into a task-specific decoder designed according to specific task requirements (which can perform targeted processing on the cross-modal feature representation to achieve efficient utilization and analysis of multi-modal data) to finally generate the prediction result of the multi-modal fusion model is as follows:

[0096] Input the final fused feature representation H final-fusion into the cross-modal Transformer encoder for cross-modal encoding. The cross-modal Transformer encoder uses the multi-head self-attention mechanism to capture the complex interaction relationships between the features of different modalities, strengthening and refining the information fusion between different modalities; when processing multi-modal data including images, text, and audio, the large model helps the cross-modal Transformer encoder identify the associations between the objects in the image, the relevant descriptions in the text, and the corresponding sounds in the audio, and strengthen the fusion of these associated information during encoding; the formula representation of this encoding process is:

[0097] H cross-modal = f cross-modal-transformer (H final-fusion )

[0098] Among them, H cross-modal is the cross-modal feature representation after cross-modal encoding; f cross-modal-transformer is the encoding function of the cross-modal Transformer encoder;

[0099] After encoding, based on the knowledge of the large model, additional semantic labels and feature annotations are added to the cross-modal feature representation H cross-modal (These semantic labels and feature annotations are helpful for subsequent task processing and can better understand the information contained in the cross-modal feature representation; for example, for the cross-modal feature representation containing a person image and a text description, the large model can add semantic labels such as "person identity", "action posture", "scene atmosphere", etc.).

[0100] According to the specific task requirements, if it is a classification task, the cross-modal feature representation H cross-modal is input into the fully connected layer for class prediction; the fully connected layer maps the cross-modal feature representation H cross-modal to different classes through a series of linear transformations and activation functions to obtain the required target result (in this process, referring to the output mapping method of the large model in similar tasks, the structure and parameters of the fully connected layer are optimized. The large model can provide prior knowledge related to the classification task such as the appropriate number of neurons and the selection of activation functions, improving the effectiveness of class prediction of the fully connected layer). The specific formula is expressed as:

[0101] y pred = f classification (H cross-modal )

[0102] where y pred is the prediction result of the multi-modal fusion model classification task; f classification is the classification task processing function;

[0103] If it is a generation task, the cross-modal feature representation H cross-modal is input into the generative decoder for target generation (such as generating text, images); (the large model learns the patterns and rules of different types of generation tasks, can provide more reasonable parameter settings for the decoder initialization, and give optimization suggestions for learning rate adjustment and loss function design during training, accelerating the training process and improving the generation quality). The generative decoder gradually generates the required target result based on the input cross-modal feature representation H cross-modal . The specific formula is expressed as:

[0104] G ou t put = f generation (H cross-modal )

[0105] where G outputGenerate prediction results for the multi-modal fusion model generation task; f generation is the generation function of the generative decoder; the above design realizes the deep fusion and effective utilization of multi-modal data. With the assistance of the large model, the cross-modal Transformer encoder accurately captures the correlations of different modal features, generates high-quality cross-modal feature representations, and the added semantic labels and feature annotations enhance the feature comprehensibility; the specific decoders or processing layers designed for different tasks, combined with the prior knowledge and optimization strategies provided by the large model, efficiently complete classification and generation tasks, improving the accuracy and quality of task processing; for example, in image and text classification tasks, it can classify more accurately, and in image generation tasks, it can generate high-quality images that better conform to semantic descriptions, improving the performance and effect of the entire multi-modal data fusion processing system in practical applications, and providing strong support for the application of multi-modal data in various fields.

[0106] In the above technical solution, the specific method for optimizing the model is as follows:

[0107] Calculate the loss function value of the multi-modal fusion model based on the true labels of the specific task and the prediction results of the multi-modal fusion model, where different types of tasks correspond to different loss functions;

[0108] For classification tasks, the cross-entropy loss function is used to measure the difference between the class probability distribution predicted by the multi-modal fusion model and the true class distribution. Suppose there are M modalities, the weight of the m-th modality is w m , the total number of samples is N, the total number of classes is C, and the class probability distribution predicted by the multi-modal fusion model for the i-th sample in the m-th modality is The true label of the task is y i , then the total loss L cls is:

[0109]

[0110] where i represents the i-th sample; j represents the j-th class; y ij represents the true label of the i-th sample belonging to the j-th class. If the i-th sample belongs to the j-th class, y ij =1; if the i-th sample does not belong to the j-th class, y ij =0; represents the predicted probability that the multi-modal fusion model predicts the i-th sample belonging to the j-th class in the m-th dimension, and the value range is between [0,1];

[0111] For generation tasks, the discriminator loss is used to optimize the model. Suppose the output of the discriminator for real samples is D(x), and the output for generated samples is D(G(z)), where x is the real sample, z is the random noise, G is the generator, and the loss L genCombined with the prior knowledge factor α provided by the large model, it can be expressed as:

[0112] L gen = -(1 - α)log(D(x)) - αlog(1 - D(G(z)))

[0113] Calculate the precision of each modality on the corresponding modality data validation set to adjust the weights of each modality. Let the initial modality weight be The adjusted weight is w m , and the adjustment coefficient β given by the large model according to the inter-modal correlation and specific task requirements m , can be expressed as:

[0114]

[0115] where P m is the precision of the m-th modality on the corresponding modality data validation set; is the average value of the precisions of all modalities; The large model will deeply analyze the reasons for the changes in the precisions of each modality, and combine the understanding of the relationships between modalities and specific task requirements to formulate a reasonable weight adjustment strategy; When the precision of a certain modality on the corresponding modality data validation set is low, the large model will analyze the interaction between this modality and other modalities, judge whether it is due to insufficient feature extraction or problems in modality fusion, and then optimize the multi-modal fusion model by reducing the weight of this modality or adjusting the feature extraction and fusion methods of this modality;

[0116] Before each modality fusion operation, use the dataset to adjust the fusion weights of the already fused modalities and the remaining modality with the strongest correlation with the already fused modalities among the remaining modalities except the already fused modalities in the multi-modal fusion model. Let the parameters of the multi-modal fusion model before the t-th adjustment be θ t , the loss function on the dataset is L small , the learning rate is η t , and the adjustment correction factor γ provided by the large model t , then the adjusted parameters θ t+1 of the multi-modal fusion model are:

[0117]

[0118] The large model provides initialization parameters and adjustment strategies for the adjustment process (ensuring that the multi-modal fusion model can converge quickly), and through a small number of training iterations on the dataset, makes the multi-modal fusion model adapt to new fusion features;

[0119] Use the backpropagation algorithm to calculate the gradients of the loss function with respect to the parameters of the multimodal fusion model (including each modal encoder, cross-modal encoder, and decoder), and use the Adam optimizer to update the parameters based on the gradients (the Adam optimizer combines the advantages of the AdaGrad optimizer and the RMSProp optimizer, and can adaptively adjust the learning rate of each parameter to improve the stability and efficiency of training). Let the first moment estimate of the gradient at the t-th iteration be m t and the second moment estimate of the gradient at the t-th iteration be v t The learning rate adjustment factor given by the large model according to the training status is δ t The original learning rate is α. Then the parameter update formula for the multimodal fusion model is:

[0120] m t+1 = β1m t + (1 - β1)g t

[0121]

[0122] where g t is the gradient at the t-th iteration; β1 is the first setting parameter of the Adam optimizer (β1 is used to control the weights of historical gradient information and current gradient information in the first moment estimate of the gradient, usually set to a value close to 1, and can be taken as 0.9); is the value of the first setting parameter β1 of the Adam optimizer at the (t + 1)-th iteration; β2 is the second setting parameter of the Adam optimizer (β2 is used to control the weights of historical gradient square information and current gradient square information in the second moment estimate of the gradient, usually set to a value close to 1, and can be taken as 0.999); is the value of the second setting parameter β2 of the Adam optimizer at the (t + 1)-th iteration; m t+1 represents the first moment estimate of the gradient at the (t + 1)-th iteration; v t+1 represents the second moment estimate of the gradient at the (t + 1)-th iteration; represents the first moment estimate of the gradient at the (t + 1)-th iteration after bias correction; represents the second moment estimate of the gradient at the (t + 1)-th iteration after bias correction; θ t represents the parameters of the multimodal fusion model before the t-th adjustment; θ t+1Denote the parameters of the updated multimodal fusion model; ∈ is a small constant to prevent the denominator from being zero; the large model provides a reference for optimizer selection based on its understanding of the performance of different optimizers in multimodal data processing and dynamically adjusts the learning rate; at the initial stage of training, a larger learning rate is adopted to enable the multimodal fusion model to quickly converge to a better solution; at the later stage of training, the learning rate is gradually decreased to avoid the multimodal fusion model oscillating near the local optimal solution; during the training process, the large model continuously monitors the training status of the multimodal fusion model, provides dynamic optimization suggestions based on real-time data and multimodal fusion model performance metrics, and when the performance of the multimodal fusion model on the validation sets of each modality data no longer improves, stops training in a timely manner to prevent overfitting;

[0123] Adopt Dropout and L2 regularization to prevent overfitting. Let the original loss be L, and the parameters of the m-th modality be θ m , and the L2 regularization coefficient set by the large model for the m-th modality is λ m , and the Dropout probability is p m , then the regularized loss L reg is:

[0124]

[0125] where is the L2 norm of the parameter θ m , and is the information entropy term related to the Dropout probability p m . The large model sets appropriate regularization parameters for each modality separately according to the characteristics of multimodal data and the model structure, dynamically adjusts the Dropout probability of different layers, and balances the complexity and generalization ability of the multimodal fusion model;

[0126] The large model regularly evaluates the modality contributions and correlations, introduces complex analysis methods and metrics, considers the synergistic effects and complementarities between modalities, and constructs a multimodal collaboration matrix S, where S mn represents the collaboration degree between the m-th modality and the n-th modality. When adjusting the modality weights, the influence of the collaboration matrix is considered. Let the current modality weight vector be w = (w1, w2, …, w M ) T , where w M represents the weight vector of the M-th modality, the adjusted weight vector is w′, and the collaboration matrix adjustment factor is ζ, then:

[0127] w′ = w + ζSw

[0128] ​When dynamically adjusting weights, the large model adopts a flexible adjustment strategy based on its real-time understanding of specific task requirements and modal relationships. In addition to the smooth adjustment strategy, it also makes adaptive adjustments according to the changing trends of modal contributions. When introducing adversarial training to optimize the fusion effect and weights, the large model optimizes the design and training of the discriminator network, providing accurate decision boundary guidance for the discriminator. At the same time, it optimizes the loss function design of the generator, guiding the generator to generate features that are closer to the real data distribution and have more natural inter-modal fusion, thus improving the overall performance of the multi-modal fusion model. The above design can fully learn the effective information in multi-modal data, continuously adjust its own parameters and modal weights, and significantly improve the performance and generalization ability in specific tasks. The assistance of the large model makes the training and optimization process more intelligent and efficient, and can effectively handle the complexity and diversity of multi-modal data.

[0129] Embodiment 2

[0130] An edge computing method for integrated sensor multi-modal data optimized by a large model, as Figure 2 shown, obtains the pre-training parameters of each modal data through the large model; optimizes the attention weights through the encoder branch to determine the key information of each modal data, and obtains the high-level feature representations of each modality; by calculating the correlation between the remaining modalities and the final core modality, inputs the high-level feature representations of the modality with the strongest correlation among the remaining modalities and the high-level feature representation of the final core modality into the cross-modal encoder to generate the fused feature representation of the fused modality, and continues the modal fusion operation until all modalities are fused, and generates the final fused feature representation; inputs the final fused feature representation into the cross-modal encoder to generate the cross-modal feature representation, and inputs the cross-modal feature representation into the task-specific decoder designed according to specific task requirements, and finally generates the prediction result of the multi-modal fusion model. The present invention improves the performance, adaptability and data processing efficiency of multi-modal data processing.

[0131] The specific method for edge computing of multi-modal data includes the following steps:

[0132] Collect each modal data through the data acquisition system, and the large model performs normalization and enhancement operations on different modal data respectively to obtain the pre-training parameters of each modal data;

[0133] Construct a corresponding encoder branch for each modality. Each encoder branch corresponding to a modality adjusts the pre-training parameters of the corresponding modal data according to the characteristics of the corresponding modal data and specific task requirements, and optimizes the attention weights of the corresponding modal data, and determines the key information of each modal data according to the attention weights, so as to obtain the high-level feature representations of each modality;

[0134] The final core modality is determined by calculating the variance and information entropy of each modality data and combining expert knowledge. The correlation between the remaining modalities except the final core modality and the final core modality is calculated through correlation analysis, and the high-level feature representations of the modality with the strongest correlation among the remaining modalities and the high-level feature representation of the final core modality are input into a cross-modal encoder for modality fusion operation to generate the fused feature representation of the fused modality. The correlation between the remaining modalities except the fused modality and the fused modality is calculated through correlation analysis. After each modality fusion operation, it is necessary to use a multi-modal dataset to adjust the fusion weights of the fused modality and the modality with the strongest correlation among the remaining modalities. The high-level feature representation of the modality with the strongest correlation among the remaining modalities and the high-level feature representation of the fused modality are input into a cross-modal encoder for modality fusion operation to generate the fused feature representation of the new fused modality until all modalities are fused and the final fused feature representation is generated;

[0135] The final fused feature representation is input into a cross-modal encoder to generate a cross-modal feature representation, and the cross-modal feature representation is input into a task-specific decoder designed according to specific task requirements to finally generate the prediction result of the multi-modal fusion model.

[0136] Embodiment 3

[0137] A computer program product includes a computer program, characterized in that the steps of the method described in Embodiment 2 are implemented when the computer program is executed by a processor.

[0138] The content not detailed in this specification belongs to the prior art well-known to those skilled in the art. Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0139] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1Apparatus for the functions specified in one or more boxes.

[0140] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one Figure 1 one or more processes and / or boxes Figure 1 or more boxes.

[0141] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one or more processes and / or boxes Figure 1 or more boxes.

[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the scope of its protection. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: after reading the present invention, those skilled in the art can still make various changes, modifications or equivalent replacements to the specific implementation manners of the invention, but these changes, modifications or equivalent replacements are all within the scope of the claims of the invention pending approval.

[0143] The content not described in detail in this specification belongs to the prior art well-known to those of ordinary skill in the art.

Claims

1. An edge computing system for integrated sensor multimodal data optimized by large models, characterized in that, Including: The data preprocessing module is used to collect various modal data through a data acquisition system. The large model performs normalization and enhancement operations on different modal data respectively to obtain the pre-training parameters of each modal data. The modality-specific encoding module is used to construct a corresponding encoder branch for each modality. Each encoder branch corresponding to a modality adjusts the pre-training parameters of the corresponding modality data according to the characteristics of the corresponding modality data and the specific task requirements, and optimizes the attention weights of the corresponding modality data. The key information of each modality data is determined according to the attention weights, so as to obtain the high-level feature representation of each modality. The fusion strategy selection and execution module is used to determine the final core modality by calculating the variance and information entropy of each modality data and combining expert knowledge, calculate the correlation between the remaining modalities except the final core modality and the final core modality through correlation analysis, and input the high-level feature representation of the modality with the strongest correlation among the remaining modalities and the high-level feature representation of the final core modality into the cross-modal encoder for modality fusion operation to generate the fusion feature representation of the fused modality. Calculate the correlation between the remaining modalities except the fused modality and the fused modality through correlation analysis. After each modality fusion operation, it is necessary to use the multi-modal dataset to adjust the fusion weights of the fused modality and the modality with the strongest correlation among the remaining modalities, and input the high-level feature representation of the modality with the strongest correlation among the remaining modalities and the high-level feature representation of the fused modality into the cross-modal encoder for modality fusion operation to generate the fusion feature representation of the new fused modality until all modalities are fused and the final fusion feature representation is generated. The cross-modal encoding and task processing module is used to input the final fusion feature representation into the cross-modal encoder to generate the cross-modal feature representation, and input the cross-modal feature representation into the task-specific decoder designed according to the specific task requirements to generate the multi-modal fusion model prediction result.

2. The edge computing system for integrated sensor multimodal data optimized by the large model according to claim 1, characterized in that, It also includes: The model training and optimization module is used to calculate the loss function value of the multi-modal fusion model according to the real label of the specific task and the prediction result of the multi-modal fusion model, calculate the precision rate of each modality on the corresponding modality data validation set to adjust the weights of each modality, update the parameters of the multi-modal fusion model using the backpropagation algorithm, dynamically adjust the learning rate, and optimize the multi-modal fusion model using the regularization method.

3. The edge computing system for integrated sensor multimodal data optimized by the large model according to claim 1, characterized in that: The various modal data include image data, text data, audio data, video data, and sensor data.

4. The edge computing system for integrated sensor multimodal data optimized by the large model according to claim 3, characterized in that: The specific process of obtaining the pre-training parameters of the various modal data is as follows: For the image modality, based on the learning results of a large amount of image data, the large model determines the cropping area and scaling ratio according to the image content and specific task requirements. The large model uses the pre-trained image recognition model to extract key features, which is expressed by the formula: I′ = f crop-scale (I) Among them, I represents the original image, I' represents the cropped and scaled image obtained after being processed by the large model, and f crop-scale represents the cropping and scaling function determined by the large model, where the cropped and scaled image I' obtained after being processed by the large model is the pre-training parameter of the image modality; For the text modality, the large model learns from a large-scale text corpus, performs semantic analysis and cleaning on the text. When performing word segmentation and tokenization, it provides an accurate word segmentation scheme according to the context, and uses the pre-trained word vector model to generate word vectors, which is expressed by the formula: T′ = f clean (T) T tokenized = f tokenize (T′) Among them, T represents the original text, and T' represents the text after being cleaned by the large model. T tokenized represents the result after word segmentation tokenization, and f clean represents the cleaning function, and f tokenize represents the word segmentation tokenization function. Among them, the text T' after being cleaned by the large model passes through the word segmentation tokenization function f tokenize to obtain the word segmentation tokenization result T tokenized is the pre-training parameter for the text modality; For the audio modality, the large model optimizes the frame extraction strategy according to the audio rhythm and frequency, dynamically adjusts the frame extraction interval, obtains a frame sequence reflecting the essential features of the audio, and assists in extracting the key audio event features. The formula is as follows: A frames = f frame-optimize (A) Among them, A represents the original audio, A frames represents the audio frame sequence after optimized frame extraction, f frame-optimize is the frame extraction function optimized by the large model, where the audio frame sequence A after optimized frame extraction frames is the pre-training parameter of the audio modality; For the video modality, the large model performs key frame extraction and content analysis on the video to remove redundant information. The formula is as follows: V keyframes = f keyframe-extract (V) Among them, V represents a video, and V keyframes represents the set of extracted key frames, and f keyframe-extract is the key frame extraction function, where the set of extracted key frames V keyframes is the pre-training parameter of the video modality; For the sensor modality, the large model performs data calibration and feature engineering based on the sensor type and task requirements. The formula is as follows: S′ = f sensor-process (S) Among them, S represents sensor data, S' represents the processed sensor data, and f sensor-process is a processing function for the sensor data, where the processed sensor data S' is the pre-training parameter of the sensor modality; After completing the preprocessing of each modality's data, perform data quality detection on the preprocessed data of each modality. If the data quality does not meet the standard, return it to the data preprocessing module for reprocessing until the data quality meets the standard, and obtain the pre-training parameters of each modality's data.

5. The edge computing system for integrated sensor multimodal data optimized by the large model according to claim 3, characterized in that: The specific method for extracting the key information of each modality's data to obtain the high-level feature representation of each modality is as follows: Construct a corresponding Transformer encoder branch for each modality. Each modality's corresponding Transformer encoder branch can adaptively learn and capture the key feature information in each modality's data according to the characteristics of each modality's data, and adjust and optimize the attention weights for the pre-training parameters of each modality's data obtained through the large model according to the specific task requirements. The specific process is as follows: For the image modality, according to the resolution and number of color channels of the image, the convolution kernel size and stride are adjusted to optimize the attention weights of different regions under the multi-head self-attention mechanism; the large model will analyze the resolution (H, W) and number of color channels C of the image, and combine the pre-trained knowledge and the current task requirements to adjust the original convolution kernel size k and stride s; at the same time, the image is divided into multiple regions [r1, r2,..., r m , analyze the semantic information of each region, and optimize the attention weight matrix A before adjustment, which is expressed by the formula: k′ = f image-k (H, W, C, k) s′ = f image-s (H, W, C, s) A′ = f image-attn-adjust (A, [r1, r2,..., r m ) Among them, (H, W) represents the resolution of the image; C represents the number of color channels; k represents the size of the original convolution kernel; s represents the original stride; k' represents the size of the adjusted convolution kernel; s' represents the adjusted stride; f image-k represents a function for adjusting the convolution kernel size based on the image modality features; f image-s represents a function for adjusting the stride based on the image modality features; A represents the attention weight matrix before adjustment; [r1, r2,..., r m represents multiple regions into which the image is divided; A' represents the adjusted attention weight matrix; f image-attn-adjust represents a function for adjusting the attention weights according to the image semantic information; For the text modality, according to the length and vocabulary of the text, adjust the word embedding dimension and the number of attention heads, and optimize the attention weights between words; the large model evaluates the length L and vocabulary V of the text, and adjusts the original word embedding dimension d embedding and the number of attention heads n heads accordingly; at the same time, analyze the semantic relevance of the word sequence [w1, w2,..., w n in the text, and optimize the attention weight matrix A before adjustment, which is expressed by the formula: d′ embedding = f text-d (L, V, d embedding ) n′ heads = f text-n (L, V, n heads ) A′ = f text-attn-adjust (A, [w1, w2,..., w n ) Among them, L represents the text length; V represents the vocabulary; d embedding represents the original word embedding dimension; n heads represents the original number of attention heads; d′ embedding represents the adjusted word embedding dimension; n′ heads represents the adjusted number of attention heads; f text-d represents a function for adjusting the word embedding dimension based on text modal features; f text-n represents a function for adjusting the number of attention heads based on text modal features; A represents the attention weight matrix before adjustment; [w1, w2,..., w n represents the word sequence in the text; A′ represents the adjusted attention weight matrix; f text-attn-adjust represents a function for adjusting the attention weights according to text semantic relevance; For the audio modality, according to the frequency distribution and duration of the audio, adjust the window size and stride of audio feature extraction, and optimize the attention weights between different audio segments; the large model analyzes the frequency distribution F and duration T of the audio, and adjusts the original window size w and stride p; at the same time, divide the audio into multiple segments [a1, a2,..., a q , analyze the semantic information of each segment, and optimize the attention weight matrix A before adjustment, which is expressed by the formula: w′ = f audio-w (F, T, w) p′ = f audio-p (F, T, p) A′ = f audio-attn-adjust (A, [a1, a2,..., a q ) Among them, F represents the frequency distribution of the audio; T represents the audio duration; w represents the original window size; p represents the original step size; w′ represents the adjusted window size; p′ represents the adjusted step size; f audio-w represents a function for adjusting the window size based on audio modal features; f audio-p represents a function for adjusting the step size based on audio modal features; A represents the attention weight matrix before adjustment; [a1, a2,..., a q represents multiple segments into which the audio is divided; A′ represents the adjusted attention weight matrix; f audio-attn-adjust represents a function for adjusting the attention weight according to the audio semantic information; By constructing a corresponding Transformer encoder branch for each modality in the modality-specific encoding stage, adapt and optimize the pre-training parameters of each modality's data according to the characteristics of each modality's data and the specific task requirements, and accurately extract the key information of the corresponding modality's data through the attention weights of the adjusted data of each modality to obtain the high-level feature representation of the corresponding modality.

6. The edge computing system for integrated sensor multimodal data optimized by a large model according to claim 3, characterized in that: The specific process for generating the final fusion feature representation is as follows: Select the modality with the largest amount of information as the initially selected core modality by calculating the variance and information entropy of each modality's data, and combine expert knowledge to determine the final core modality. The formula is as follows: M core-initial = f initial-select (Var1, Entropy1, Var2, Entropy2,..., Var n , Entropy n ) M core = f expert-adjust (M core-initial ) Among them, Var i represents the variance of the i-th modal data; Entropy i represents the information entropy of the i-th modal data; M core-initial represents the initially selected core mode; M core represents the final core mode after adjustment by experts; f initial-select represents the function of the core mode initially selected based on data statistical indicators; f expert-adjust represents the function for experts to adjust the final core mode according to professional knowledge; Calculate the correlation between the remaining modalities except the final core modality and the final core modality through correlation analysis and determine the modality fusion order. This process can be represented by the following formula: Corr i = f corr-select (M core , M i ) Order=sort(Corr1,Corr2,...,Corr n ) Among them, Corr i represents the correlation between the i-th modality and the final core modality; M i represents the i-th modality; f corr-select represents a function that selects the corresponding correlation analysis method according to the modality data type and characteristics and calculates the correlation; Order represents the modality fusion order after sorting according to the correlation; sort(Corr1, Corr2,..., Corr n ) represents the modality fusion order after sorting according to the correlation, and only the modality with the strongest correlation ranking is selected for fusion; Starting from the determined final core modality, each time select the remaining modality with the strongest correlation with the fused modality from the remaining modalities except the already fused modalities for fusion; input the high-level feature representation of the modality with the strongest correlation in the remaining modalities and the high-level feature representation of the already fused modality into the cross-modal Transformer encoder for modality fusion operation. The cross-modal Transformer encoder uses its multi-head self-attention mechanism and feed-forward neural network layer to interact and integrate the high-level feature representations of different modalities to generate a new fusion feature representation. The high-level feature representation H of the final core modality core and the high-level feature representation H1 of the remaining modality with the strongest correlation with the final core modality among the remaining modalities except the final core modality are input into the cross-modal Transformer encoder to obtain the fused feature representation H of the fused modality fusion-1 , which is expressed by the formula as: H fusion-1 = f cross-modal-encoder (H core , H1) Among them, f cross-modal-encoder represents the fusion function of the cross-modal Transformer encoder; After each modal fusion operation, it is necessary to use the multi-modal dataset to adjust the fusion weights of the fused modality and the modality with the strongest correlation among the remaining modalities except the fused modality. The adjusted fused feature representation H fusioN-1-fine-tuned can be expressed as: H fusion-1-fine-tuned = f fine-tune (H fusion-1 , Small-dataset) Among them, f fine-tune represents a function for adjusting the fused features using a dataset, and Small-dataset represents the dataset used for adjustment; Repeat the above steps of modal fusion and adjustment until all modalities are fused to generate the final fused feature representation H final-fusion ; The entire fusion process can be summarized as an iterative process: H fusion-j = f cross-modal-encoder (H fusion-(j-1) , H j ) H fusion-j-fine-tuned = f fine-tune (H fusion-j , Small-dataset) Among them, j represents the number of rounds of fusion; H j represents the high-level feature representation of the j-th modality to be fused; H fusion-j represents the fused feature obtained after the j-th round of modality fusion; H fusion-(j-1) represents interacting and integrating with H j to make the features of different modalities complement and fuse with each other, generating a new fused feature H fusion-j ; H fusion-j-fine-tuned represents the result after the j-th round of fused feature is adjusted by the dataset; through continuous iteration, the information of all modalities is gradually fused and the final fused feature representation is generated.

7. The edge computing system for integrated sensor multimodal data optimized by a large model according to claim 3, characterized in that: Input the final fusion feature representation into the cross-modal Transformer encoder to generate a cross-modal feature representation, and input the cross-modal feature representation into the task-specific decoder designed according to the specific task requirements. The specific method for finally generating the prediction result of the multi-modal fusion model is as follows: The final fused feature representation H final-fusion is input into a cross-modal Transformer encoder for cross-modal encoding. The cross-modal Transformer encoder utilizes the multi-head self-attention mechanism to capture the complex interaction relationships between features of different modalities, strengthening and refining the information fusion between different modalities. When processing multi-modal data containing images, text, and audio, the large model helps the cross-modal Transformer encoder identify the associations between objects in the image, text-related descriptions, and corresponding sounds in the audio, and enhances the fusion of these associated information, namely, objects in the image, text-related descriptions, and corresponding sounds in the audio, during encoding. The formula representation of this encoding process is as follows: H cross-modal = f cross-modal-transformer (H final-fusion ) Among which H cross-modal is the cross-modal feature representation after cross-modal encoding; f cross-modal-transformer is the encoding function of the cross-modal Transformer encoder; After encoding, based on the knowledge of the large model, for the cross-modal feature representation H cross-modal Add additional semantic tags and feature annotations; Process according to specific task requirements. For classification tasks, the cross-modal feature representation H cross-modal is input into the fully connected layer for class prediction; through a series of linear transformations and activation functions, the fully connected layer maps the cross-modal feature representation H cross-modal to different classes to obtain the desired target result, and the specific formula is expressed as: y pred = f classification (H cross-modal ) where y peed is the prediction result of the multi-modal fusion model classification task; f classification is the classification task processing function; For a generation task, the cross-modal feature representation H cross-modal is input into the generative decoder for target generation; the generative decoder generates the required target result step by step based on the input cross-modal feature representation H cross-modal The specific formula is expressed as: G output = f generation (H cross-modal ) Among which G outPut is the prediction result of the multi-modal fusion model generation task; f generation is the generation function of the generative decoder.

8. The edge computing system for integrated sensor multimodal data optimized by a large model according to claim 2, wherein: The specific method for optimizing the model is as follows: Calculate the loss function value of the multi-modal fusion model based on the true labels of specific tasks and the prediction results of the multi-modal fusion model, where different types of tasks correspond to different loss functions; For the classification task, the cross-entropy loss function is used to measure the difference between the class probability distribution predicted by the multi-modal fusion model and the true class distribution. Suppose there are M modalities in total, and the weight of the m-th modality is w m , the total number of samples is N, the total number of classes is C, and the class probability distribution predicted by the multi-modal fusion model for the i-th sample in the m-th modality is The true label of the task is y i , then the total loss L cls is: where, i represents the i-th sample; j represents the j-th class; y ij represents the ground truth label of the task that the i-th sample belongs to the j-th class. If the i-th sample belongs to the j-th class, y ij = 1; if the i-th sample does not belong to the j-th class, y ij = 0; represents the predicted probability that the multi-modal fusion model predicts the i-th sample belongs to the j-th class under the m-th dimension, and the value range is between [0, 1]; For the generation task, the discriminator loss is used to optimize the model. Let the output of the discriminator for real samples be D(x), and the output for generated samples be D(G(z)), where x is the real sample, z is the random noise, and G is the generator. The loss L of the generation task gen incorporates the prior knowledge factor α provided by the large model and can be expressed as: L gen = -(1 - α) log(D(x)) - α log(1 - D(G(z))) Calculate the precision rate of each modality on the corresponding modality data validation set to adjust the weights of each modality. Let the initial modality weight be The adjusted weight is w m , and the adjustment coefficient given by the large model according to the inter-modal correlation and specific task requirements is β m , which can be expressed as: where p m is the precision rate of the m-th modality on the corresponding modality data validation set; is the average value of the precision rates of all modalities; the large model will deeply analyze the reasons for the changes in the precision rates of each modality, and formulate a reasonable weight adjustment strategy in combination with the understanding of the relationships between modalities and specific task requirements; when the precision rate of a certain modality on the corresponding modality data validation set is low, the large model will analyze the interaction between this modality and other modalities, judge whether it is due to insufficient feature extraction or problems in modality fusion, and then optimize the multi-modal fusion model by reducing the weight of this modality or adjusting the feature extraction and fusion methods of this modality; After each modal fusion operation, the fusion weight of the remaining modality with the strongest correlation with the fused modality among the fused modalities and the remaining modalities other than the fused modalities in the multi-modal fusion model is adjusted using the multi-modal dataset. Let the parameters of the multi-modal fusion model before the t-th adjustment be θ t , the loss function on the dataset be L small , the learning rate be η t , and the adjustment correction factor provided by the large model be γ t , then the adjusted parameters θ t+1 of the multi-modal fusion model are as follows: The large model provides initialization parameters and adjustment strategies for the adjustment process. By performing a small number of training iterations on the dataset, the multi-modal fusion model is adapted to new fusion features; Use the backpropagation algorithm to calculate the gradient of the loss function with respect to the parameters of the multimodal fusion model, and use the Adam optimizer to update the parameters according to the gradient. Let the first-order moment estimate of the gradient at the t-th iteration be m t , and the second-order moment estimate of the gradient at the t-th iteration be v t , the learning rate adjustment factor given by the large model according to the training status is δ t , and the original learning rate is α. Then the parameter update formula for the multimodal fusion model is as follows: m t+1 = β1m t + (1 - β1)g t where g t is the gradient of the t-th iteration; β1 is the first set parameter of the Adam optimizer; is the value of the first set parameter β1 of the Adam optimizer at the (t + 1)-th iteration; β2 is the second set parameter of the Adam optimizer; is the value of the second set parameter β2 of the Adam optimizer at the (t + 1)-th iteration; m t+1 represents the first moment estimate of the gradient of the (t + 1)-th iteration; v t+1 represents the second moment estimate of the gradient of the (t + 1)-th iteration; represents the first moment estimate of the gradient of the (t + 1)-th iteration after bias correction; represents the second moment estimate of the gradient of the (t + 1)-th iteration after bias correction; θ t represents the parameters of the multi-modal fusion model before the t-th adjustment; θ t+1 represents the updated parameters of the multi-modal fusion model; ∈ is a small constant to prevent the denominator from being zero; The large model provides a reference for selecting the optimizer based on its understanding of the performance of different optimizers in multi-modal data processing and dynamically adjusts the learning rate; At the beginning of training, a larger learning rate is adopted to make the multi-modal fusion model quickly converge to a better solution; In the later stage of training, the learning rate is gradually decreased to avoid the multi-modal fusion model oscillating near the local optimal solution; During the training process, the large model continuously monitors the training status of the multi-modal fusion model, provides dynamic optimization suggestions based on real-time data and performance indicators of the multi-modal fusion model, and stops training in time when the performance of the multi-modal fusion model on the validation sets of each modal data no longer improves to prevent overfitting; Use Dropout and L2 regularization to prevent overfitting. Let the original loss be L, and the parameters of the m-th modality be θ m , the L2 regularization coefficient set by the large model for the m-th modality is λ m , and the Dropout probability is p m , then the regularized loss L reg is: where is the L2 norm of the parameter θ m , is the information entropy term related to the Dropout probability p m . According to the characteristics of multimodal data and the model structure, the large model sets appropriate regularization parameters for each modality separately, dynamically adjusts the Dropout probability of different layers, and balances the complexity and generalization ability of the multimodal fusion model; The large model regularly evaluates the modal contributions and correlations, introduces complex analysis methods and metrics, considers the synergistic effects and complementarities among modalities, and constructs a multimodal synergy matrix S, where S mn represents the degree of synergy between the m-th modality and the n-th modality. When adjusting the modal weights, the influence of the synergy matrix is considered. Let the current modal weight vector be w = (w1, w2, …, w M ) T , where w M represents the weight vector of the M-th modality, the adjusted weight vector is w′, and the synergy matrix adjustment factor is ζ. Then: w′ = w + ζSw When dynamically adjusting the weights, the large model adopts a flexible adjustment strategy based on the real-time understanding of the specific task requirements and modal relationships. In addition to the smooth adjustment strategy, it also makes adaptive adjustments according to the changing trend of modal contributions. When introducing adversarial training to optimize the fusion effect and weights, the large model optimizes the design and training of the discriminator network, provides accurate decision boundary guidance for the discriminator, and at the same time optimizes the loss function design of the generator to guide the generator to generate features closer to the real data distribution and with more natural inter-modal fusion, improving the overall performance of the multi-modal fusion model.

9. An edge computing method for integrated sensor multimodal data optimized by a large model, characterized in that, It includes the following steps: Collect multi-modal data through a data acquisition system. The large model performs normalization and enhancement operations on different types of modal data respectively to obtain the pre-training parameters of each type of modal data; Construct corresponding encoder branches for each modality. Each encoder branch corresponding to a modality adjusts the pre-training parameters of the corresponding modality data according to the characteristics of the corresponding modality data and the specific task requirements, and optimizes the attention weights of the corresponding modality data. Determine the key information of each modality data according to the attention weights, so as to obtain the high-level feature representations of each modality; Determine the final core modality by calculating the variance and information entropy of each modality data and combining expert knowledge. Calculate the correlation between the remaining modalities except the final core modality and the final core modality through correlation analysis, and input the high-level feature representations of the modality with the strongest correlation among the remaining modalities and the high-level feature representation of the final core modality into the cross-modal encoder for modal fusion operation to generate the fusion feature representation of the fused modality. Calculate the correlation between the remaining modalities except the fused modality and the fused modality through correlation analysis. After each modal fusion operation, it is necessary to use the multi-modal dataset to adjust the fusion weights of the fused modality and the modality with the strongest correlation among the remaining modalities. Input the high-level feature representation of the modality with the strongest correlation among the remaining modalities and the high-level feature representation of the fused modality into the cross-modal encoder for modal fusion operation to generate the fusion feature representation of the new fused modality until all modalities are fused and the final fusion feature representation is generated; Input the final fusion feature representation into the cross-modal encoder to generate the cross-modal feature representation, and input the cross-modal feature representation into the task-specific decoder designed according to the specific task requirements to generate the prediction results of the multi-modal fusion model.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in claim 9.

Citation Information

Patent Citations

  • Structured query statement analysis method and device, equipment and storage medium

    CN117453720A

  • Cyclic prefix type processing method and device and related equipment

    CN120342560A

Cited By

  • Kit raw material quality detection method based on multi-modal data fusion

    CN120510481A

  • A method for testing the quality of raw materials of test kits based on multimodal data fusion

    CN120510481B

  • Multi-modal general-purpose model collaborative reasoning method based on dynamic routing mechanism

    CN120597213A

  • Multi-modal general-special model collaborative reasoning method based on dynamic routing mechanism

    CN120597213B

  • Data processing method, target container determination method, device, equipment, medium and product

    CN120744860A