A data decision-making method and system based on multi-modal large model analysis
By introducing situational modeling modules and modal specialization optimization technology into the multimodal data decision-making method, dynamically adjusting modal weights and eliminating redundant information, the problems of insufficient context perception and low modal fusion efficiency in the existing technology are solved, and more accurate and efficient multimodal data decisions are achieved.
Patent Information
- Application Number
- CN202510072673.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-01-17
AI Technical Summary
Existing multimodal data decision-making methods fail to effectively filter or correct redundant and conflicting information during data fusion, resulting in limited generalization capabilities of decision-making models and weak understanding of context, making it difficult to dynamically adjust modal weights to adapt to specific scenarios or tasks.
A data decision-making method based on multimodal large model analysis is proposed. The context information of multimodal data is dynamically extracted through the context modeling module to guide modal optimization and decision-making. The method includes modal feature extraction and fusion, specialization optimization, modal interaction enhancement and global feature alignment, dynamic adjustment of modal weights and elimination of redundant information.
It has achieved higher context perception capabilities, more efficient modal fusion methods and more accurate task decision-making capabilities, breaking through the bottlenecks of existing technologies, and is suitable for complex scenarios such as medical care, transportation, and agriculture.
Smart Images

Figure CN119494079B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-modal large models, and particularly relates to a data decision-making method and system based on multi-modal large model analysis. Background Art
[0002] With the rapid development of artificial intelligence and big data technologies, multi-modal data analysis has gradually become an important direction for promoting the development of intelligent decision-making systems. Multi-modal data comes from different data modalities, such as text, image, speech, time series, etc., and its characteristic is rich information and diversity. By fusing multi-modal data, complex problems can be understood more comprehensively, thereby improving the scientificity and accuracy of decision-making. Especially in the fields of machine learning ensemble models and classification decision-making, using multi-modal data can significantly improve the prediction performance and classification accuracy of models. However, there are still many challenges in current multi-modal data decision-making methods:
[0003] There are often some redundant information in current multi-modal data (such as the repetition of text descriptions and the content presented in images), and even data conflicts may occur between different modalities (such as the inconsistency between sensor data and image information). Existing methods usually adopt simple splicing or alignment methods during data fusion, and fail to effectively filter or correct these redundant and conflicting information, resulting in limited generalization ability of the decision-making model.
[0004] At the same time, existing multi-modal decision-making models have weak understanding ability of context, and it is difficult to dynamically adjust modal weights according to specific scenarios or tasks. For example, in medical diagnosis, different diseases may require different modal features to be concerned (such as heart diseases are more concerned about time series modality, while lung diseases rely more on imaging modality). In existing technologies, large models often treat all modalities equally and fail to make full use of context information to improve the pertinence of decision-making. Especially in machine learning ensemble models and classification decision-making tasks, there is a lack of dynamic adjustment ability for specific scenarios.
[0005] Many multi-modal large models are designed for general scenarios and are difficult to adapt to the task requirements of specific fields. This general design leads to insufficient decision-making efficiency and accuracy of the model in specific tasks. For example, an ensemble classifier may not be able to effectively fuse information from different modalities when facing complex tasks, resulting in a decline in classification performance. In addition, semantic alignment between different modalities often relies on a unified feature space, ignoring the impact of semantic characteristics of specific tasks on decision-making.
[0006] Current multi-modal large models usually need to process a large number of data features. However, the current technology lacks a dynamic pruning mechanism for unimportant information during the modal fusion process, resulting in high system computational complexity and inability to meet the real-time decision-making requirements. For example, in the scenario of autonomous driving, a large amount of data from lidar, cameras, and sensors needs to be analyzed in real time, and the efficiency of existing technologies is insufficient to handle such requirements. Especially in integrated learning models, data from different modalities may bring excessive redundant information, affecting the efficiency of real-time decision-making.
[0007] Based on the above problems, the current multi-modal decision-making technology has obvious performance bottlenecks when facing complex tasks. To overcome these defects, an innovative system and method are needed that can achieve efficient analysis and optimized decision-making of multi-modal data, thereby fully leveraging the potential of large models in machine learning integration and classification decision-making. Summary of the Invention
[0008] The object of the present invention is to propose a data decision-making method and system based on multi-modal large model analysis, which can achieve higher context awareness, more efficient modal fusion methods, and more accurate task decision-making capabilities in multi-modal data decision-making, thereby breaking through the bottleneck of the existing technology and being widely applicable to complex scenarios such as medical, transportation, and agriculture.
[0009] To achieve the above object, in the first aspect of the present invention, a data decision-making method based on multi-modal large model analysis is provided, and the method includes:
[0010] S1. Collect multi-modal data and perform preprocessing, extract and fuse modal features from the preprocessed data to obtain multi-modal features fused through concatenation operations; wherein, the multi-modal features include text features, image features, and time series features;
[0011] S2. Perform feature concatenation and weighted combination on the multi-modal features to generate initial context features, and perform feature optimization on the initial context features to obtain final context features, representing the dynamic context information in the current task and scenario;
[0012] S3. Use the final context features to perform specialization optimization on each modal feature respectively, remove redundant information and strengthen task-related characteristics to generate final modal specialization optimization features, and the modal specialization optimization features include optimized specialized text features, optimized specialized image features, and optimized specialized time series features; the specialization optimization includes:
[0013] Generate dynamically weighted modal specialization according to the final context features;
[0014] Use the dynamically weighted modal specialization to perform specialization optimization on the multi-modal features respectively to generate optimized modal features, including optimized text features, image features, and time series features;
[0015] Introduce a modal interaction enhancement module to guide the high-order information interaction between modalities through the final context features, and obtain interaction features;
[0016] Fuse the interaction features with the optimized text features, image features, and time series features respectively to generate the final modality-specific optimized features, including optimized specific text features, optimized specific image features, and optimized specific time series features;
[0017] S4. Align the final modality-specific optimized features to generate the fused global features;
[0018] S5. When a specific task of multimodal analysis enters, optimize the global features according to the specific task, generate decision results and provide explanations, including:
[0019] Dynamically weight the global features according to the requirements of the current task, strengthen the information related to the task, and suppress redundant features at the same time;
[0020] If there are multi-objective scenarios in the task, design a collaborative optimization module to model the shared relationship between subtasks through the task collaboration matrix of the collaborative optimization module, and combine the specialized features of each subtask into the final task-specific features;
[0021] Utilize the task-specific features optimized by the task to generate the final output decision results through a task-specific decision network.
[0022] Furthermore, the multimodal data is represented as , where each data point contains three modalities: text modality , image modality and time series modality ;
[0023] The preprocessing includes:
[0024] For the text modality, convert the text modality into a word vector representation matrix , where is the number of words in the sentence, is the embedding dimension of each word;
[0025] Design a dynamic semantic aggregation module based on context attention to weight the word vectors and generate a text feature vector :
[0026] ;
[0027] Among them, It is a task-related semantic query vector used to guide the attention mechanism, and the attention weights are calculated by the following formula:
[0028] ;
[0029] where represents the semantic query, represents the embedding of the th word, is the attention weight of the th word;
[0030] The dynamic semantic aggregation module captures the most task-relevant information in the context by dynamically adjusting the task-related semantic query vector , is the final text feature, is the embedding dimension of the final text feature;
[0031] For the image modality, a hierarchical convolutional network is used to extract image features . A frequency-domain perturbation suppression term is added to the last layer of the convolutional network to suppress high-frequency noise by analyzing the frequency-domain information of the image:
[0032] ;
[0033] where is the frequency-domain representation of the image , is the regularization coefficient used to balance the influence of frequency-domain suppression on the features, enhancing the stability and effectiveness of the image features by reducing high-frequency noise, is the final image feature, is the embedding dimension of the final image feature;
[0034] For the time series modality, a multi-scale trend capture module is designed. The multi-scale trend capture module first decomposes the time series into a long-term trend and a short-term fluctuation:
[0035] ;
[0036] where represents the long-term trend part, represents the short-term fluctuation part;
[0037] Smoothing convolution and multi-scale window convolution are respectively used to extract 's global features and 's local features. The two parts of the features are fused through weighting to generate the final time series feature vector :
[0038] ;
[0039] Among them, is the weight of the trend feature, dynamically learned by minimizing the squared error of the local residual, , is the embedding dimension of the final time series feature.
[0040] Furthermore, the joint feature vector is constructed as follows:
[0041] Normalize the three types of modal features extracted , and so that they have a consistent distribution in the same feature space:
[0042] ;
[0043] Among them, and are the mean and standard deviation of modality respectively;
[0044] The normalized features are fused into a joint feature vector :
[0045] ;
[0046] Among them, represents vector concatenation, , .
[0047] Furthermore, the S2 specifically includes:
[0048] Design a task-aware self-attention mechanism to generate an attention weight vector and dynamically adjust the modal weights, expressed as:
[0049] ;
[0050] Among them are trainable parameters, is the attention weight vector, representing the dynamic weights of the text, image, and time series modalities respectively;
[0051] Weighted combination of each modal feature according to the generated attention weights to obtain the initial context feature, expressed as:
[0052] ;
[0053] Among them, is the initial context feature, Dynamic weights for text, image, and time series modalities respectively;
[0054] Introduce a modality interaction enhancement module according to the correlation between modalities, model the high-order relationship between modalities through tensorization, and obtain enhanced context features;
[0055] For the enhanced context features, Perform feature sparsity optimization through norm constraint, remove low-importance dimensions, and generate the final context features.
[0056] Furthermore, the modality interaction enhancement module is calculated as follows:
[0057] First, calculate the modality interaction tensor , capturing the bidirectional interaction between different modalities:
[0058] ;
[0059] Among them, Represents the outer product operation, Is the high-dimensional tensor of modality interaction information;
[0060] Use the high-dimensional tensor of modality interaction information To enhance the initial context features , and add a regularization term To suppress excessive interaction complexity:
[0061] ;
[0062] Among them, Is the linear mapping matrix, Is the regularization coefficient, Used to introduce non-linearity;
[0063] The final context features are expressed as:
[0064] ;
[0065] Among them, Is the sparsity constraint strength, Is the final context feature.
[0066] Furthermore, the modality-specific dynamic weight , is calculated as follows:
[0067] ;
[0068] Among them, Is the modality 's linear mapping matrix, used to convert context features into modality-specific weights, is a non - linear activation function used to suppress the influence of negative values, is a sparse smoothing term to ensure smooth adjustment of the weight matrix in cases of low task relevance;
[0069] The optimized modal feature is expressed as:
[0070] ;
[0071] where, is the modal feature;
[0072] The generation formula of the interaction feature is:
[0073] ;
[0074] where, is the interaction enhancement feature of modality generated by the weighted sum of the context feature and the optimized modal feature, capturing the implicit association between modalities;
[0075] The final modality - specific optimization feature is expressed as:
[0076] ;
[0077] where, is the balance coefficient of interaction enhancement, used to control the fusion degree of the optimization feature and the interaction feature.
[0078] Furthermore, the S4 specifically includes:
[0079] Project the final modality - specific optimization feature into a unified semantic space to obtain the aligned feature, expressed as:
[0080] ;
[0081] where, is the projection matrix of modality , is the bias term, represents the aligned feature;
[0082] Design a modality - complementary regularization term to encourage complementarity between modalities, expressed as:
[0083] ;
[0084] This regularization term reduces redundancy by constraining the dot - product of features between modalities, ensuring that the fused global feature retains both the unique information of modalities and can effectively combine the modality synergy; denotes the element-wise product, is the Frobenius norm.
[0085] Furthermore, the S4 further includes introducing a context-weighted fusion strategy to dynamically adapt to the current task and context requirements, expressed as:
[0086] ;
[0087] where, is the global feature, is the modality 's weight, controlled by the final context feature .
[0088] Furthermore, the dynamic weighting of the global feature, with the weight guided and generated by the final context feature , and the calculation of the weight is directly associated with the global feature , expressed as:
[0089] ;
[0090] where, is the final context feature, ensures the normalization of the weight ;
[0091] The specialized task-optimized feature is calculated by element-wise weighting:
[0092] ;
[0093] where represents element-wise multiplication;
[0094] The task-specialized feature is calculated as follows:
[0095] ;
[0096] where, is the specialized feature of the -th sub-task; the collaboration matrix is learnable, used to capture the implicit correlation between sub-tasks, and at the same time ensure the collaboration of specialized features between tasks;
[0097] During the specialization optimization process, to avoid the concentration of features on a few dimensions, a distribution balance regularization term is designed, expressed as:
[0098] ;
[0099] where, is a vector of all 1s, is the dimension of;
[0100] Finally, using the task-specialized features optimized by the task, , through the task-specific decision network , generate the final output decision :
[0101] ;
[0102] wherein, the structure of changes according to the task type.
[0103] In another aspect of the present invention, a data decision-making system based on multi-modal large model analysis is provided, and the system includes:
[0104] A multi-modal data acquisition unit, configured to acquire multi-modal data and perform preprocessing, extract and fuse modal features from the preprocessed data to obtain multi-modal features fused by splicing operation; wherein, the multi-modal features include text features, image features and time series features;
[0105] A multi-modal feature analysis unit, configured to perform feature splicing and weighted combination on the multi-modal features to generate initial context features, and perform feature optimization on the initial context features to obtain final context features, representing dynamic context information in the current task and scenario;
[0106] A multi-modal feature optimization unit, configured to use the final context features to respectively perform specialization optimization on each modal feature, remove redundant information and strengthen task-related characteristics, and generate final modal specialization optimization features, the modal specialization optimization features include optimized specialized text features, optimized specialized image features and optimized specialized time series features; the specialization optimization includes:
[0107] Generate modal-specialized dynamic weights according to the final context features;
[0108] Use the modal-specialized dynamic weights to respectively perform specialization optimization on the multi-modal features to generate optimized modal features, including optimized text features, image features and time series features;
[0109] Introduce a modal interaction enhancement module, and through the final context features, guide the high-order information interaction between modalities to obtain interaction features;
[0110] Respectively fuse the interaction features with the optimized text features, image features and time series features to generate final modal specialization optimization features, including optimized specialized text features, optimized specialized image features and optimized specialized time series features;
[0111] A multi-modal feature alignment unit for aligning the final modality-specific optimized features to generate fused global features;
[0112] A decision generation unit, when a specific task of multi-modal analysis enters, optimizes the global features according to the specific task, generates a decision result and provides an explanation, including:
[0113] Dynamically weight the global features according to the requirements of the current task, strengthen task-related information, and at the same time suppress redundant features;
[0114] If the task has a multi-objective scenario, design a collaborative optimization module to model the shared relationship between subtasks through the task collaboration matrix of the collaborative optimization module, and combine the specialized features of each subtask into the final task-specialized features;
[0115] Use the task-specialized features optimized by the task to generate the final output decision result through a task-specific decision network.
[0116] The beneficial technical effects of the present invention are at least as follows:
[0117] (1) The present invention dynamically extracts the context information of multi-modal data through the context modeling module to guide subsequent modality optimization and decision-making. This mechanism can adjust the modality weights according to the specific scenario, solving the problem of insufficient context awareness in the prior art. For example, in the traffic management scenario, the system can identify the context of "peak hours" and give priority to real-time sensor data rather than historical data; in medical diagnosis, the context model can automatically adjust the weights of image data and time series data according to the characteristics of different diseases to improve the accuracy of classification decisions. This mechanism is particularly applicable to machine learning integration models and classification decision tasks, enhancing the flexibility and pertinence of the model through dynamic perception of different contexts.
[0118] (2) In view of the characteristics of different modalities (such as the spatial characteristics of images and the semantic information of texts), the present invention uses a specialized optimization module to crop and enhance the modality features, and at the same time uses a modality conflict detection mechanism to eliminate the inconsistencies between modalities, thereby achieving more accurate feature fusion. This design overcomes the problem of decision failure caused by information redundancy and conflict in the prior art. Especially in the integration model, through efficient feature fusion and conflict handling, it avoids over-reliance on irrelevant information, thereby improving the stability and accuracy of the model.
[0119] (3) The present invention constructs a three - layer decision - making framework of "modal - independent reasoning - context - aware fusion - task - driven optimization". In the modal - independent reasoning stage, preliminary extraction of modal features is realized; in the context - aware fusion stage, dynamic feature weighting is completed; and in the task - driven optimization stage, semantic alignment related to specific tasks is further strengthened. This framework not only improves the task - targeting of the model but also significantly reduces the computational complexity by dynamically pruning redundant information, solving the problem of insufficient computational efficiency in the prior art. Especially in ensemble learning and classification decision - making tasks, through this framework, multi - modal data can be flexibly processed, task relevance can be enhanced, the computational process of the model can be optimized, and the efficiency of real - time decision - making can be improved.
[0120] (4) The present invention can achieve higher context - awareness capabilities, more efficient modal fusion methods, and more accurate task - decision - making capabilities in multi - modal data decision - making, thus breaking through the bottleneck of the prior art and being widely applicable to complex scenarios such as medical treatment, transportation, and agriculture.
[0121] In these scenarios, the innovative decision - making mechanism of the present invention can, through machine - learning ensemble models and classification decision - making methods, combine the characteristics of different modal data to achieve accurate and efficient intelligent decision - making, providing strong technical support for practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0122] The present invention is further described with reference to the accompanying drawings. However, the embodiments in the drawings do not constitute any limitation to the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the following drawings.
[0123] Figure 1 It is a flowchart of a data decision - making method based on multi - modal large - model analysis according to an embodiment of the present invention.
[0124] Figure 2 It is a framework diagram of a data decision - making system based on multi - modal large - model analysis according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0125] The embodiments of the present invention are described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.
[0126] As Figure 1 shown, a data decision - making method based on multi - modal large - model analysis provided by an embodiment of the present invention includes:
[0127] S1. Collect multimodal data and perform preprocessing. Extract and fuse modal features from the preprocessed data to obtain multimodal features fused through concatenation operations. Among them, the multimodal features include text features, image features, and time series features.
[0128] Specifically, input a multimodal data set , where each data point contains three modalities: text modality , image modality , and time series modality . There are obvious distribution differences and heterogeneous characteristics among multimodal data. For example, the text modality has semantic dependencies, the image modality has spatial characteristics, and the time series modality has time dependencies. Therefore, feature extraction and standardization must be performed separately for each modality.
[0129] Among them, text modality feature extraction:
[0130] Convert the input text into a word vector representation matrix , where is the number of words in the sentence, is the embedding dimension of each word.
[0131] Design a dynamic semantic aggregation module based on context attention to weight the word vectors and generate a text feature vector :
[0132] ;
[0133] Among them is a task-related semantic query vector used to guide the attention mechanism. The attention weights are calculated by the following formula:
[0134] ;
[0135] Among them represents the semantic query, represents the embedding of the th word, is the attention weight of the th word.
[0136] Capture the information most relevant to the task in the context by dynamically adjusting . is the final text feature.
[0137] Among them, image modality feature extraction:
[0138] Use a hierarchical convolutional network Extract image modality features 。
[0139] Add a frequency domain perturbation suppression term to the last layer of the convolutional network, and suppress high-frequency noise by analyzing the frequency domain information of the image:
[0140] ;
[0141] where is the frequency domain representation of the image , is the regularization coefficient, which is used to balance the influence of frequency domain suppression on features. This design enhances the stability and effectiveness of image features by reducing high-frequency noise. 。
[0142] Among them, time series modality feature extraction:
[0143] Design a multi-scale trend capture module for the input time series 。This module first decomposes the time series into a long-term trend and a short-term fluctuation: 。
[0144] ;
[0145] where represents the long-term trend part, represents the short-term fluctuation part.
[0146] Use smooth convolution and multi-scale window convolution to extract the global features of and the local features of respectively. The two parts of features are fused by weighting to generate the final time series feature vector :
[0147] ;
[0148] where is the weight of the trend feature, which is dynamically learned by minimizing the squared error of the local residual. 。
[0149] Furthermore, normalize the three extracted modality features , and to make them have a consistent distribution in the same feature space:
[0150] ;
[0151] where and are the mean and standard deviation of the modality respectively.
[0152] The normalized features are fused into a joint feature vector through a concatenation operation :
[0153] ;
[0154] where represents vector concatenation, , .
[0155] It can be understood that the output joint feature , which contains the standardized feature representations of multi-modal data, serves as the input to the subsequent context modeling module. Through specific processing of each modality, the generated features not only have high-quality expression capabilities but also can effectively reduce noise interference and unify the feature scales of different modalities.
[0156] S2. Perform feature concatenation and weighted combination on the multi-modal features to generate initial context features, and perform feature optimization on the initial context features to obtain the final context features, representing the dynamic context information in the current task and scenario.
[0157] Specifically, the information distribution in the multi-modal features may vary significantly depending on the task or scenario. For example, in traffic scheduling, the visual modality during the day may be more important than the sensor time series modality . Therefore, a task-aware self-attention mechanism is designed to dynamically adjust the modality weights.
[0158] First, input the concatenated multi-modal features into the attention weight generation module:
[0159] ;
[0160] where are trainable parameters, is the attention weight vector, representing the dynamic weights of the text, image, and time series modalities respectively.
[0161] Furthermore, according to the generated attention weights perform weighted combination on the features of each modality:
[0162] ;
[0163] where is the initially generated context feature, which has combined the dynamic importance of the multi-modal features.
[0164] Furthermore, considering that there may be strong correlations among multimodal features (e.g., the descriptions in the text modality may be closely related to the image modality information), a modality interaction enhancement module is introduced to model the high-order relationships among modalities in a tensored manner:
[0165] First, calculate the modality interaction tensor , capturing the bidirectional interaction between different modalities:
[0166] ;
[0167] where represents the outer product operation, and is the high-dimensional tensor of modality interaction information.
[0168] Furthermore, use to enhance the preliminary context features , and add a regularization term (Frobenius norm) to suppress excessive interaction complexity:
[0169] ;
[0170] where is the linear mapping matrix, is the regularization coefficient, and is used to introduce non-linearity.
[0171] Furthermore, in the generated context features , there may still be some noisy features or information redundancy. Therefore, a feature sparsification mechanism is designed to remove low-importance dimensions. Feature sparsification optimization is performed through norm constraint, and the final context feature is represented as:
[0172] ;
[0173] where is the sparsification constraint strength, and is the optimized context feature.
[0174] It can be understood that the output context feature represents the dynamic context information in the current task and scenario, integrating the dynamic importance of multimodal features, the interaction information among modalities, and removing redundant features. This feature will be used as the input for the next modality specialization optimization to guide the subsequent steps to more precisely adjust the modality features.
[0175] S3. Use the final context features to specialize and optimize each modal feature respectively, remove redundant information and strengthen task-related characteristics, and generate the final modal specialized and optimized features, where the modal specialized and optimized features include optimized and specialized text features, optimized and specialized image features, and optimized and specialized time series features.
[0176] Specifically, according to the context features , generate the dynamic weights for modal specialization ( ). The weights are calculated by the following formula:
[0177] ;
[0178] where is the linear mapping matrix of modality , which is used to convert the context features into modal specialized weights, is the non-linear activation function, which is used to suppress the influence of negative values, is the sparse smoothing term, which ensures that the weight matrix is smoothly adjusted when the task relevance is low. has the same size as the dimension of the modal feature .
[0179] Use the dynamic weights to specialize and optimize the modal feature , and generate the optimized modal feature :
[0180] ;
[0181] where represents the element-wise weighting operation. The optimized modal feature retains the task-related characteristics and suppresses the redundant information irrelevant to the current task.
[0182] Furthermore, to enhance the synergy between modalities, a modal interaction enhancement module is introduced to guide the high-order information interaction between modalities through the context features . The generation formula for the interaction features is:
[0183] ;
[0184] where is the interaction enhancement feature of modality , which is generated by the weighted sum of the context features and the optimized modal features, capturing the implicit associations between modalities.
[0185] Fuse the interaction features with the optimized features to generate the final modal specialized and optimized features :
[0186] ;
[0187] Among them is the balance coefficient of interaction enhancement, which is used to control the fusion degree of the optimized feature and the interaction feature.
[0188] Furthermore, output the finally optimized modal features , and , and splice them into a joint feature . The joint feature contains the modal information after removing redundancy and combines the interaction enhancement between modalities, providing high-quality input for the subsequent cross-modal alignment and fusion steps.
[0189] Through the modality specialization optimization and the interaction enhancement between modalities guided by the context feature , not only the expression ability of the single-modal feature is improved, but also the collaborative information between modalities is captured, ensuring that the final feature can better serve the multi-modal task.
[0190] S4. Align the finally modality-specialized optimized features to generate the fused global features.
[0191] Specifically, the feature distributions of different modalities may be located in heterogeneous semantic spaces, so they need to be aligned to a shared feature space. For this purpose, design a context-guided feature alignment module to project the modality features
[0192] into a unified semantic space:
[0193] ;
[0194] Among them, is the projection matrix of modality , is the bias term, represents the aligned feature. This step realizes the semantic alignment of the modality features, enabling them to have a consistent expression ability in the subsequent fusion.
[0195] Furthermore, the fused global features need to dynamically adapt to the current task and context requirements. For this purpose, introduce a context-weighted fusion strategy:
[0196] Among them is the weight of modality , which is controlled by the context feature . For example, can be calculated through , among which is modality The weight mapping vector. The importance of the weight dynamic adjustment mode, for example, increasing the contribution of the time series mode in time-sensitive scenarios.
[0197] Furthermore, in feature fusion, there may be redundant information between modes or a risk of over-reliance on a single mode. Therefore, a modal complementary regularization term is designed to encourage complementarity between modes:
[0198] ;
[0199] This regularization term reduces redundancy by constraining the dot product of features between modes, ensuring that the fused global feature retains both the unique information of the modes and can effectively combine the modal synergy. Here represents the element-wise product, and
[0200] is the Frobenius norm. It can be understood that the output global feature fuses the shared information of text, image, and time series modes, and achieves balance between modes through dynamic weights and regularization control. The global feature
[0201] S5. When specific modal analysis enters, optimize the global feature according to the specific task, generate decision results and provide explanations.
[0202] Specifically, according to the requirements of the current task, dynamically weight the global feature to strengthen task-related information and suppress redundant features at the same time. Design a task relevance optimization module, the core weight of which is guided by the context feature to generate, and the calculation of the weight can be directly associated with
[0203] ;
[0204] where is the context feature, ensuring the normalization of the weight . Here, the dependence on complex matrix mapping is reduced, and instead, the dynamic association between the task and the global feature is directly captured through the context feature.
[0205] Furthermore, the specialized task-optimized feature can be calculated by element-wise weighting:
[0206] ;
[0207] where Denotes element-wise multiplication. The generated features retain the task-related characteristics and exclude unnecessary noise.
[0208] Furthermore, if the task is a multi-objective scenario (e.g., coexistence of classification and prediction tasks), it is necessary to further coordinate the specialized features of different tasks. Design a collaborative optimization module that models the shared relationships between subtasks through a task collaboration matrix to combine the specialized features of each subtask into the final task-specialized features :
[0209] ;
[0210] where is the specialized feature of the th subtask. The collaboration matrix is learnable and used to capture the implicit correlations between subtasks while ensuring the collaboration of specialized features among tasks.
[0211] Furthermore, during the specialization optimization process, to avoid the concentration of features on a few dimensions, design a distribution balance regularization term :
[0212] ;
[0213] where is a vector of all 1s, is the dimension. This regularization term ensures the uniform distribution of task weights and avoids over-reliance on a single dimension, enhancing the robustness and generalization ability of the features.
[0214] Furthermore, using the specialized features optimized for the task , generate the final output through a task-specific decision network :
[0215] ;
[0216] where the structure of can vary according to the task type. For example, a fully connected layer is used in classification tasks, and a regression network is used in prediction tasks.
[0217] As Figure 2 shown, in another embodiment of the present invention, a data decision system based on multi-modal large model analysis is provided. The system includes:
[0218] The multimodal data acquisition unit 101 is used to acquire multimodal data, perform preprocessing, extract and fuse modal features from the preprocessed data to obtain multimodal features fused through splicing operations; wherein, the multimodal features include text features, image features, and time series features;
[0219] The multimodal feature analysis unit 102 is used to perform feature splicing and weighted combination on the multimodal features to generate initial context features, and perform feature optimization on the initial context features to obtain final context features, representing the dynamic context information in the current task and scenario;
[0220] The multimodal feature optimization unit 103 is used to specifically optimize each modal feature using the final context features, remove redundant information, and strengthen task-related characteristics to generate final specifically optimized modal features, where the specifically optimized modal features include specifically optimized text features, specifically optimized image features, and specifically optimized time series features; the specific optimization includes:
[0221] Generate dynamically weighted modal specialization according to the final context features;
[0222] Use the dynamically weighted modal specialization to specifically optimize the multimodal features respectively to generate optimized modal features, including optimized text features, image features, and time series features;
[0223] Introduce a modal interaction enhancement module to guide high-order information interaction between modalities through the final context features to obtain interaction features;
[0224] Fuse the interaction features with the optimized text features, image features, and time series features respectively to generate final specifically optimized modal features, including specifically optimized text features, specifically optimized image features, and specifically optimized time series features;
[0225] The multimodal feature alignment unit 104 is used to align the final specifically optimized modal features to generate fused global features;
[0226] The decision generation unit 105 is used to, when a specific task of multimodal analysis enters, optimize the global features according to the specific task, generate decision results and provide explanations, including:
[0227] Dynamically weight the global features according to the requirements of the current task, strengthen task-related information, and at the same time suppress redundant features;
[0228] If the task has a multi-objective scenario, design a collaborative optimization module, model the shared relationship between subtasks through the task collaboration matrix of the collaborative optimization module, and combine the specialized features of each subtask into the final task-specialized features;
[0229] Using the task specialization features optimized by the task, the final output decision result is generated through a task-specific decision network.
[0230] It should be noted that the above-described work process is only illustrative and does not limit the protection scope of the present invention. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and no limitation is imposed here.
[0231] In addition, for the technical details not described in detail in this embodiment, reference can be made to the parameter operation method provided in any embodiment of the present invention, and details will not be repeated here.
[0232] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including that element.
[0233] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0234] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as a read-only memory / random access memory, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0235] The above is only the preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, is equally included in the patent protection scope of the present invention.
Claims
1. A data decision method based on multimodal large model analysis, characterized in that: The method comprises: S1, collecting multimodal data and preprocessing it, extracting and fusing modal features of the preprocessed data to obtain multimodal features fused by splicing operations; wherein the multimodal features include text features, image features and time series features; S2, performing feature concatenation and weighted combination on the multimodal features to generate initial context features, and performing feature optimization on the initial context features to obtain final context features, representing dynamic context information under the current task and scenario; S3. Using the final context features, each modality feature is optimized separately, redundant information is removed and task-related characteristics are strengthened, and the final modality-specific optimization features are generated. The modality-specific optimization features include optimized specialized text features, optimized specialized image features, and optimized specialized time series features. The specialized optimization includes: Generate dynamic weights of modal specialization based on the final situational characteristics; Use the modality-specific dynamic weights to perform specialized optimization on the multimodal features respectively, and generate optimized modality features, including optimized text features, image features, and time series features; The modal interaction enhancement module is introduced to guide the high-order information interaction between modalities through the final context features to obtain the interaction features; The interactive features are respectively fused with the optimized text features, image features and time series features to generate the final modality-specific optimized features, including optimized specialized text features, optimized specialized image features and optimized specialized time series features; S4, aligning the final modality-specific optimization features to generate fused global features; S5. When a specific task of multimodal analysis comes in, optimize the global features according to the specific task, generate decision results and provide explanations, including: According to the needs of the current task, the global features are dynamically weighted to strengthen the task-related information while suppressing redundant features; If the task has a multi-objective scenario, a collaborative optimization module is designed to model the sharing relationship between subtasks through the task collaboration matrix of the collaborative optimization module, and combine the specialized features of each subtask into the final task specialized features; The task-specific features after task optimization are used to generate the final output decision result through the task-specific decision network.
2. The data decision method based on multimodal large model analysis according to claim 1 is characterized in that: Multimodal data is represented as , where each data point Contains three modes: text mode , Image Modality and time series modality ; The pre-processing comprises: For text mode, the text mode of the i-th data point Converted into word vector representation matrix ,in is the number of words in the sentence, is the embedding dimension of each word, and R is the real number space; Designing a dynamic semantic aggregation module based on contextual attention , weight the word vectors and generate text feature vectors : ; in, is the word vector representation matrix, is a task-related semantic query vector used to guide the attention mechanism. The attention weight is calculated by the following formula: ; in represents a semantic query, Indicates word embedding, T is the text modality, It is The attention weight of each word, is the normalization function, is the transpose of the semantic query vector; The dynamic semantic aggregation module By dynamically adjusting the task-related semantic query vector Capture the most relevant task-related information in context, is the final text feature, is the embedding dimension of the final text feature; For the image modality, a layered convolutional network is used Extracting image features , add the frequency domain disturbance suppression term to the last layer of the convolutional network to suppress high-frequency noise by analyzing the frequency domain information of the image: ; in, is an image The frequency domain representation of is the regularization coefficient, which is used to balance the influence of frequency domain suppression on features and enhance the stability and effectiveness of image features by reducing high-frequency noise. is the final image feature, is the embedding dimension of the final image feature; For time series modality, design a multi-scale trend capture module , the multi-scale trend capture module First, decompose the time series into long-term trends and short-term fluctuations: ; in, is the ith time series data point, Represents the long-term trend part, Represents the short-period fluctuation part; Use smooth convolution to extract Global features and multi-scale window convolution extraction The local features of the two parts are weighted fused to generate the final time series feature vector : ; in, is the weight of the trend feature, which is dynamically learned by minimizing the squared error of the local residual, , is the embedding dimension of the final time series feature, is the global eigenvector of the long-term trend part, It is the local eigenvector of the short-period fluctuation part.
3. The data decision method based on multimodal large model analysis according to claim 2 is characterized in that: The multimodal features are constructed as follows: Extracted text features , Image features and time series features Normalize them so that they have consistent distribution in the same feature space: ; in, is the normalized eigenvector of mode m, is the eigenvector of the original mode m, and are the mean and standard deviation of mode m respectively; The normalized features are fused into multimodal features through splicing operations : ; in, represents vector concatenation, , , is the dimension of the final multimodal feature, is the embedding dimension of the time series feature vector.
4. The data decision method based on multimodal large model analysis according to claim 3 is characterized in that: The S2 specifically includes: Design a task-aware self-attention mechanism, generate an attention weight vector, and dynamically adjust the modal weight, expressed as: ; in, is a trainable parameter, is the attention weight vector, which represents the dynamic weights of text, image and time series modalities respectively, is the normalization function; The features of each modality are weighted and combined according to the generated attention weights to obtain the initial context features, which are expressed as: ; in, is the initial situation feature, are the dynamic weights for text, image, and time series modalities, respectively; According to the correlation between modalities, a modal interaction enhancement module is introduced to model the high-order relationship between modalities through tensorization to obtain enhanced situational features. The enhanced situational features are The norm constraint is used to optimize feature sparsity, remove dimensions of low importance, and generate the final contextual features.
5. The data decision method based on multimodal large model analysis according to claim 4 is characterized in that: The modal interaction enhancement module is calculated as follows: First, calculate the modal interaction tensor , capturing the two-way interaction between different modalities: ; in, represents the outer product operation, is a high-dimensional tensor of modal interaction information; High-dimensional tensors using modality interaction information Initial situation characteristics Enhance and add regularization terms To suppress excessive interaction complexity: ; in, is the enhanced situational feature, is the linear mapping matrix, is the regularization coefficient; Represents a nonlinear activation function, which is used to introduce nonlinearity; The final situation feature is expressed as: ; in, is the sparsification constraint strength, is the final situational feature, is a coefficient that adjusts the strength of the interaction between modalities and is used to control the difference between the initial context features and the interaction tensor.
6. The data decision method based on multimodal large model analysis according to claim 5 is characterized in that: The linear mapping matrix , calculated as follows: ; in, is the linear mapping matrix of modality m, which is used to convert situational features into modality-specific weights. It is a nonlinear activation function used to suppress the influence of negative values. It is a sparse smoothing term, which ensures smooth adjustment of the weight matrix when the task correlation is low; The optimized modal characteristics , expressed as: ; in, is the modal feature; The generation formula of the interaction feature is: ; in, is the interactive enhancement feature of modality m, which is generated by the weighted sum of contextual features and optimized modality features, capturing the implicit association between modalities. is the optimized text modality feature, is the optimized image modality feature, is the optimized time series modal feature; The final modal-specific optimization feature , expressed as: ; in, It is the balance coefficient of interactive enhancement, which is used to control the degree of fusion between optimized features and interactive features.
7. The data decision method based on multimodal large model analysis according to claim 6 is characterized in that: The S4 specifically includes: The final modality-specific optimized features are projected into the unified semantic space to obtain the aligned features, which are expressed as: ; in, is the projection matrix of mode m, is the bias term, represents the aligned features; Designing the modal complementarity regularization term , encourages complementarity between modes, expressed as: ; Among them, m is mode m, n is mode n, and the regularization term By constraining the feature point product between modalities, redundancy is reduced to ensure the global features after fusion It not only retains the unique information of the modality, but also effectively combines the modal synergy; represents element-wise product, is the Frobenius norm.
8. The data decision method based on multimodal large model analysis according to claim 7 is characterized in that: The S4 also includes introducing a context-weighted fusion strategy to dynamically adapt to current tasks and context requirements, which is expressed as: ; in, is a global feature, is the weight of mode m, which is determined by the final context feature control.
9. The data decision method based on multimodal large model analysis according to claim 7 is characterized in that: The global features are dynamically weighted, and the weights are determined by the final situational features. Guided generation, weight calculation is directly related to global features Related, expressed as: ; in, is the final situational feature, is the weight of dynamic weighting, Guaranteed weight The normalization of Specialized task optimization features Calculated by element-wise weighting: ; in represents element-wise multiplication; The task-specific features The calculation is as follows: ; in, It is The specialized features of each subtask, For the The specialized features of each subtask, For the The specialized features of each subtask, For the Specialized features of each subtask; synergy matrix It is learnable and is used to capture the implicit correlation between subtasks while ensuring the synergy of specialized features between tasks; In the process of specialized optimization, in order to avoid the concentration of features in a few dimensions, a distribution balance regularization term is designed , expressed as: ; in, is a vector of all 1s, yes Dimensions; Finally, the task-specific features after task optimization are used , through a task-specific decision network , generating the final output decision : ; in, The structure varies according to the task type.
10. A data decision system based on multimodal large model analysis, characterized in that: The system comprises: A multimodal data acquisition unit, used to acquire multimodal data and perform preprocessing, extract modal features from the preprocessed data and fuse them to obtain multimodal features fused by splicing operations; wherein the multimodal features include text features, image features and time series features; A multimodal feature analysis unit, used to perform feature concatenation and weighted combination on the multimodal features to generate initial context features, and perform feature optimization on the initial context features to obtain final context features, representing dynamic context information under the current task and scenario; The multimodal feature optimization unit is used to perform specialized optimization on each modal feature using the final contextual feature, remove redundant information and strengthen task-related characteristics, and generate the final modal specialized optimization features, wherein the modal specialized optimization features include optimized specialized text features, optimized specialized image features, and optimized specialized time series features; the specialized optimization includes: Generate dynamic weights of modal specialization based on the final situational characteristics; Use the modality-specific dynamic weights to perform specialized optimization on the multimodal features respectively, and generate optimized modality features, including optimized text features, image features, and time series features; The modal interaction enhancement module is introduced to guide the high-order information interaction between modalities through the final context features to obtain the interaction features; The interactive features are respectively fused with the optimized text features, image features and time series features to generate the final modality-specific optimized features, including optimized specialized text features, optimized specialized image features and optimized specialized time series features; Multimodal feature alignment unit, used to align the final modality-specific optimization features and generate fused global features; The decision generation unit is used to optimize the global features according to the specific task when a specific task of multimodal analysis comes in, generate decision results and provide explanations, including: According to the needs of the current task, the global features are dynamically weighted to strengthen the task-related information while suppressing redundant features; If the task has a multi-objective scenario, a collaborative optimization module is designed to model the sharing relationship between subtasks through the task collaboration matrix of the collaborative optimization module, and combine the specialized features of each subtask into the final task specialized features; The task-specific features after task optimization are used to generate the final output decision result through the task-specific decision network.
Citation Information
Patent Citations
E-commerce marketing intelligent decision-making method and system based on multi-modal learning
CN118096205A
Transformation method and device of novel protective transformer
CN118801394A