A multimodal information fusion control system and method
By evaluating the differences in execution parameters within the control cycle and dynamically adjusting the modal weights, the problem of inflexibility and inaccuracy in the response of traditional control systems in complex environments is solved, achieving higher operational precision and response speed.
Patent Information
- Application Number
- CN202511277063.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Traditional control systems suffer from insufficient intermodal synchronization and fusion when dealing with complex and dynamically changing environments, resulting in inflexible or inaccurate responses and an inability to reflect environmental changes in a timely manner, thus affecting the safety and efficiency of the system.
The target difference assessment module calculates the path coordinate difference, category variation frequency and offset change magnitude, the modality replacement judgment module evaluates modality output error and matching score, the modality lag identification module identifies abnormal responses, and the modality weight reconstruction module dynamically adjusts modality weights to ensure the system's optimization and stability in complex environments.
It enables dynamic adjustment of the control system, improves the system's adaptability to environmental changes and decision-making accuracy, reduces erroneous decisions caused by modal mismatch, enhances the system's stability and reliability, and improves response speed and processing efficiency.
Smart Images

Figure CN120821188B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information fusion technology, and in particular to a multimodal information fusion control system and method. Background Technology
[0002] The field of information fusion encompasses the technology of jointly processing and analyzing information from multiple sensory sources. It primarily studies how to collaboratively process and fuse multimodal information (such as images, speech, text, motion signals, and physiological parameters) from different sources and with heterogeneous types. This technological field focuses on intermodal feature representation methods, temporal and spatial alignment mechanisms, intermodal weight allocation strategies, and fusion algorithms, including but not limited to specific methods such as feature-level fusion, decision-level fusion, attention mechanisms, cross-modal alignment networks, and graph neural networks based on deep learning. Its core technology lies in improving information complementarity, enhancing system robustness, and reducing modal redundancy, thereby achieving more efficient system response and accurate decision-making in scenarios such as intelligent control, context recognition, and human-computer interaction.
[0003] A multimodal information fusion control system is a system designed for control tasks with multiple input sources. Its main purpose is to take multiple modal data as input, perform collaborative analysis through fusion algorithms, and then output control signals to drive or schedule execution units. This system can be applied in fields such as intelligent manufacturing, autonomous driving, and intelligent robotics to improve the environmental perception capability, state discrimination accuracy, and action decision-making efficiency of the control system.
[0004] Traditional control systems face challenges in handling complex and dynamically changing environments due to insufficient intermodal synchronization and fusion. Fixed fusion strategies struggle to cope with rapidly changing environmental conditions, resulting in inflexible or inaccurate responses when processing real-time tasks. For example, in autonomous driving or smart manufacturing scenarios, single or fixed-weight modal fusion strategies may fail to reflect drastic environmental changes, such as sudden events or operational errors, in a timely manner. This deficiency can lead to decision-making delays or errors, impacting the overall system's safety and efficiency. Furthermore, traditional systems lack effective real-time evaluation and adjustment mechanisms for dynamic modal weight adjustment. This limits the system's optimal performance in volatile environments, preventing the full utilization of the complementary advantages of different modal data and affecting overall system performance and reliability. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing a multimodal information fusion control system and method.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a multimodal information fusion control system, the system comprising:
[0007] The target difference assessment module obtains the control execution parameters of multimodal tasks within adjacent control cycles, calculates the path coordinate difference, category variation frequency and offset change magnitude, calls the three difference thresholds of path, category and feedback for judgment, and generates the task difference change value.
[0008] The modality replacement judgment module compares the output error of the master control mode with the task difference change value to see if it is greater than the error response threshold and whether the modality matching score is lower than the modality preference threshold. If both conditions are met, the master control mode is marked as a state to be switched and a modality replacement judgment result is generated.
[0009] Based on the modal replacement determination result, the modal hysteresis identification module calculates the linkage ratio between modal response time and state change amplitude, determines whether the linkage ratio exceeds the modal disturbance judgment threshold, and obtains abnormal response modal identification information.
[0010] The modality weight reconstruction module calls the feature participation ratio of the corresponding modality in the fusion feature generation according to the abnormal response modality identification information, calculates the product between the time interval deviation value and the feature participation ratio, sets a weight adjustment factor for the modality whose product result is greater than the time contribution threshold, and performs new weight allocation according to the adjustment factor to obtain the modality participation weight reconstruction value.
[0011] The present invention is improved in that the task difference change value is specifically a path offset judgment value, a category mutation state value, and a feedback change amplitude value; the mode replacement judgment result includes a mode switching trigger identifier, a master mode degradation mark, and a candidate mode preference mark; the abnormal response mode identification information is specifically a state disturbance index, a response delay level, and a signal validity label; and the mode participation weight reconstruction value includes a participation weight adjustment coefficient, a feature contribution attenuation factor, and a time penalty allocation value.
[0012] The present invention is improved in that the target difference assessment module includes:
[0013] The path offset calculation submodule obtains the control execution parameters of the multimodal task within adjacent control cycles, extracts the coordinate vector values of the starting point coordinates and the ending point coordinates of the path, calculates the Euclidean distance difference between the coordinate vectors of the starting point and the ending point of the path in two cycles, and compares the absolute value difference between the difference result and the set path offset threshold to generate a path offset judgment value.
[0014] The category variation detection submodule calls the target category number in the two control cycles before and after the path offset judgment value, extracts whether the category number status has changed, counts the number of category numbers that have changed and performs periodic normalization processing, judges the relationship between the result and the category status mutation frequency threshold, and generates a category mutation status value.
[0015] The offset magnitude calculation submodule, based on the aforementioned category mutation state value, calls the two-period feedback offset values to obtain the maximum amplitude difference of the offset over the time series, and compares it with the feedback change amplitude threshold using the following formula: ;
[0016] The offset normalization amplitude value is calculated and then combined with the first three judgment values to obtain the task difference change value.
[0017] Where D represents the offset normalization amplitude value, This represents the coordinate vector value of the endpoint of the execution path in the current control cycle. This represents the coordinate vector value of the starting point of the execution path in the previous control cycle, and C represents the category mutation state value. This represents the feedback offset value at the i-th time point. Indicates the first The feedback offset values at each time point, where n represents the total number of data points in the feedback offset sequence.
[0018] The present invention is improved in that the modal replacement determination module includes:
[0019] The error extraction submodule obtains the output response sequence of the main control mode signal and the alternative mode signal in each task segment of the current control cycle based on the task difference change value, extracts the average error value of the corresponding feedback signal in each sequence, records the signal time window position corresponding to the error value, and judges the absolute value difference between the error value and the corresponding error response threshold to obtain the mode error deviation value.
[0020] The matching score judgment submodule, based on the modal error deviation value, calls the matching score value corresponding to the control target for each modal signal within the current control cycle, extracts the task relevance index, response stability coefficient, and structural coupling consistency index of the score value, and constructs a matching composite evaluation expression, using the formula: ;
[0021] The composite scoring error value is obtained through calculation. The magnitude of the composite scoring error value is compared with the modality optimization critical value to determine the modality scoring deviation state and generate the modality scoring difference value.
[0022] Where R represents the composite scoring error value, This represents the task relevance index corresponding to the modal signal. This represents the response stability coefficient of the j-th modal signal within the current control cycle. This represents an index indicating the consistency of the structural coupling between the modal signal and the current task. This represents the score value of the current modal signal. This represents the target expected mean of the modal scores within the current control cycle. This represents the adjustment constant for the matching score evaluation, and K represents the total number of terms in the stability coefficient sequence.
[0023] The switching state marking submodule determines whether the master control mode signal simultaneously satisfies the condition that the error deviation value is greater than the error response threshold and the score difference value is less than the preferred critical value, based on the modal score difference value. If both conditions are met, the master control mode signal is marked as a state to be switched, and a candidate list of modes is established to obtain the mode replacement determination result.
[0024] The present invention is improved in that the modal hysteresis identification module includes:
[0025] The timestamp extraction submodule obtains the response timestamp sequence of image modality, voice modality and action modality in the current control cycle based on the modality replacement determination result, performs signal statistical processing on each type of modality signal within a fixed time window, extracts the first response time and the last response time, calculates the average response interval within a unit time window, and generates a modality response interval value.
[0026] The state variation calculation submodule, based on the modal response interval value, calls the state change sequences of the three modes within the response time window, and extracts the state change frequency, the number of overlapping change directions, and the change amplitude values, respectively, using the formula: ;
[0027] The modal linkage offset value is obtained through calculation;
[0028] Where L represents the modal linkage offset value, Indicates the total duration of the statistical period. This represents the difference in the k-th state variation. This represents the variation value of the modal response direction vector. Indicators representing the consistency of modal change direction This represents the mean of the magnitudes of the modal state values. This represents the modal response interval value, and N represents the total number of state variation events;
[0029] The linkage ratio judgment submodule compares the modal linkage offset value with the set modal disturbance judgment threshold. If the value is greater than the judgment threshold, the corresponding modality is recorded as an unstable signal modality, and the response label is set as an abnormal state to obtain abnormal response modality identification information.
[0030] The present invention is improved in that the modal weight reconstruction module includes:
[0031] The interval extraction submodule obtains the abnormal response modality identification information, extracts the first response time and feedback completion time of each modality based on the control feedback timestamp of each modality signal in the previous control cycle, calculates the duration of action execution in the current cycle, and performs difference processing with the duration of the previous cycle to obtain the modality interval difference value.
[0032] The contribution product calculation submodule, based on the modality interval difference value, calls the feature participation weight data of each modality in the fusion feature generation process, using the formula: ;
[0033] The product evaluation value of modal interval deviation and feature participation ratio is obtained by calculation, and the modal product evaluation value is obtained.
[0034] in, This represents the modal product evaluation value. This indicates the modal feedback end time of the current control cycle. This indicates the start time of the modal feedback in the previous control cycle. This represents the contribution value of a single feature corresponding to the z-th mode. This indicates the total number of features participating in the modal types. This indicates the proportion of modality involved in the feature generation process. This indicates the stability deviation of the current cycle control feedback process;
[0035] The weight resetting generation submodule, based on the modal product evaluation value, determines the relationship between the evaluation value of each modality and the time contribution threshold, filters out modes that are higher than the threshold and sets their weight adjustment factors, and performs replacement and update on the original modal weights according to the adjustment factor ratio to obtain the modality participation weight reconstruction value.
[0036] The present invention has an improvement, wherein the system further includes:
[0037] The control command output module calls the modal fusion output vector in the current period according to the modal participation weight reconstruction value, weights and combines the fusion components of each modal dimension according to the updated weight, inputs the weighted fusion vector into the command mapping matrix, generates the fusion command code set of the task scenario in the current period, and obtains the control command code information of the current period.
[0038] The current cycle control instruction encoding information specifically refers to the fusion vector mapping number, task control instruction group code, and signal channel activation index.
[0039] The present invention is improved in that the control command output module includes:
[0040] The weight mapping submodule obtains the modal participation weight reconstruction value, calls the modal fusion output vector in the current control cycle, performs a positioning operation on the fusion component of each modal dimension, assigns the corresponding modal participation weight to the fusion component, and forms a weighted result matrix to obtain the modal weighted combination value.
[0041] The fusion encoding construction submodule constructs an instruction output vector within the task cycle based on the modal weighted combination value. It combines all modal weighted component values with the scene label matching index, structural information coupling factor, and instruction mapping template index using the following formula: ;
[0042] The operation yields the fused output vector value;
[0043] in, This represents the fused output vector value. This represents the value of the x-th fusion component. This represents the participation weight of the x-th dimension fusion component. This represents the structural coupling factor of the x-th mapping template. This represents the matching index of the x-th modality label. This represents the target value of the x-th scene task label. Indicates the number of dimensions of the fused output vector;
[0044] The task instruction generation submodule inputs the fused output vector value into the control instruction mapping matrix column by column, performs matrix channel matching and index replacement operations according to the instruction path configuration rules, obtains the set of task mapping results for each instruction channel, and establishes the current cycle control instruction encoding information.
[0045] A multimodal information fusion control method, which is executed based on the aforementioned multimodal information fusion control system, includes the following steps:
[0046] S1: Obtain the control execution parameters of the multimodal task within adjacent control cycles, calculate the path coordinate difference, category variation frequency and offset change magnitude, call the three difference thresholds of path, category and feedback for judgment, and generate the task difference change value;
[0047] S2: Based on the task difference change value, compare whether the output error of the master control mode is greater than the error response threshold and whether the mode matching score is lower than the mode preference threshold. If both conditions are met, mark the master control mode as a state to be switched and generate a mode replacement judgment result.
[0048] S3: Based on the modal replacement determination result, calculate the linkage ratio between modal response time and state change amplitude, determine whether the linkage ratio exceeds the modal disturbance judgment threshold, and obtain abnormal response modal identification information;
[0049] S4: Based on the abnormal response mode identification information, calculate the product between the time interval deviation value and the feature participation ratio, set a weight adjustment factor for modes whose product result is greater than the time contribution threshold, and perform new weight allocation according to the adjustment factor to obtain the mode participation weight reconstruction value.
[0050] S5: Based on the modal participation weight reconstruction value, the fusion components of each modal dimension are weighted and combined according to the updated weight. The weighted fusion vector is input into the instruction mapping matrix to generate the fusion instruction code set of the task scenario in the current period, and the control instruction code information of the current period is obtained.
[0051] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0052] In this invention, dynamic adjustment and optimization of the control system are achieved by comprehensively evaluating the differences in execution parameters within the control cycle. By calculating path coordinate differences, category variation frequency, and offset change magnitude, subtle changes in task execution can be reflected in detail, thereby accurately adjusting the control strategy and improving the system's adaptability to environmental changes and decision-making accuracy. By evaluating the output error and matching score between modes, the system can switch to a control mode that is more suitable for the current task environment in a timely manner, reducing erroneous decisions caused by mode mismatch and enhancing the stability and reliability of the system. By identifying abnormal response modes and reallocating mode weights, resource allocation is optimized, reaction speed and processing efficiency are improved, and the performance of the multimodal information fusion control system in complex environments is enhanced, ensuring higher operational accuracy and response speed. Attached Figure Description
[0053] Figure 1 This is a system flowchart of the present invention;
[0054] Figure 2 This is a flowchart of the target difference assessment module of the present invention;
[0055] Figure 3 This is a flowchart of the modal replacement judgment module of the present invention;
[0056] Figure 4 This is a flowchart of the modal hysteresis recognition module of the present invention;
[0057] Figure 5 This is a flowchart of the modal weight reconstruction module of the present invention;
[0058] Figure 6 This is a flowchart of the control command output module of the present invention. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0060] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0061] Please see Figure 1 This invention provides a technical solution: a multimodal information fusion control system, the system comprising:
[0062] The target difference assessment module obtains the control execution parameters of the multimodal task within adjacent control cycles, including execution path parameters, target classification number and feedback offset value, calculates path coordinate difference, category variation frequency and offset change magnitude, calls the three difference thresholds of path, category and feedback for single-item judgment, accumulates the judgment results and generates task difference change value;
[0063] The modal replacement judgment module calls up the historical output error value and modal matching score value corresponding to the master control modal signal and the candidate modal signal in the current control cycle based on the task difference change value. It compares whether the output error of the master control mode is greater than the error response threshold and whether the modal matching score is lower than the modal selection threshold. If both conditions are met, the master control mode is marked as a state to be switched and a modal replacement judgment result is generated.
[0064] Based on the modal replacement judgment result, the modal lag recognition module calls the response timestamp sequence of the image modality, voice modality and action modality received in the current period, extracts the frequency of change of state value, the proportion of overlap of change direction and the average value of change amplitude of each type of modal signal within a fixed time window, calculates the linkage ratio between modal response time and state change amplitude, and determines whether the linkage ratio exceeds the modal disturbance judgment threshold to obtain abnormal response modal recognition information;
[0065] The modal weight reconstruction module, based on the abnormal response modal identification information, calls the control feedback timestamp corresponding to each modal signal in the previous control cycle, extracts the first reaction time and feedback completion time, calculates the difference between the action execution duration interval and the time interval of the previous cycle, calls the feature participation ratio of the corresponding modality in the fusion feature generation, calculates the product between the time interval deviation value and the feature participation ratio, sets a weight adjustment factor for modalities whose product result is greater than the time contribution threshold, and performs new weight allocation according to the adjustment factor to obtain the modal participation weight reconstruction value;
[0066] The control command output module calls the modal fusion output vector in the current period according to the modal participation weight reconstruction value, and weights and combines the fusion components of each modal dimension according to the updated weight. The weighted fusion vector is then input into the command mapping matrix to generate the fusion command code set of the task scenario in the current period, and obtains the control command code information of the current period.
[0067] The task difference change value specifically includes path offset judgment value, category mutation state value, and feedback change amplitude value. The mode replacement judgment result includes mode switching trigger identifier, master mode degradation mark, and alternative mode selection mark. The abnormal response mode identification information specifically includes state disturbance index, response delay level, and signal validity label. The mode participation weight reconstruction value includes participation proportion adjustment coefficient, feature contribution attenuation factor, and time penalty allocation value. The current cycle control command encoding information specifically refers to fusion vector mapping number, task control command group code, and signal channel activation index.
[0068] Please see Figure 2 The target difference assessment module includes:
[0069] The path offset calculation submodule obtains the control execution parameters of the multimodal task within adjacent control cycles, extracts the coordinate vector values of the starting point coordinates and the ending point coordinates of the path, calculates the Euclidean distance difference between the coordinate vectors of the starting point and the ending point of the path in two cycles, and compares the absolute value difference between the difference result and the set path offset threshold to generate a path offset judgment value.
[0070] The path offset calculation submodule obtains the control execution parameters of the multimodal task within adjacent control cycles. It sequentially reads the starting and ending point coordinate vectors of the task path in the previous and current control cycles. In the example, the starting point coordinates of the previous cycle are (2,3,5), and the ending point coordinates are (8,6,10). The starting point coordinates of the current cycle are (2.5,3.2,5.1), and the ending point coordinates are (8.5,6.1,10.2). By comparing the starting and ending coordinates of the path in the previous and current cycles, the Euclidean distance between the vectors is calculated. The path length of the previous cycle is... The current cycle path length is The difference between the two cycle paths is The difference is compared with the path offset threshold of 0.05. The offset threshold is set based on the maximum allowable change of the path in history of 5%. The maximum length of the historical path is set to 10 units, and the threshold is set to 5%×10=0.5. Since the unit normalization is set to 0.05, it is determined whether the difference of 0.026 exceeds the threshold. The result is that it does not exceed the threshold, so the path offset judgment value is 0.
[0071] The category variation detection submodule calls the target category number in the two control cycles before and after the path offset judgment value, extracts whether the category number status has changed, counts the number of category numbers that have changed and performs cycle normalization processing, judges the relationship between the result and the category status mutation frequency threshold, and generates the category mutation status value.
[0072] The category variation detection submodule compares the task category numbers of the previous and current control cycles based on the path offset judgment value. The previous cycle's category number is C1, and the current cycle's is C2. It determines whether there has been a change. If the category numbers are inconsistent, it is determined to be a switch, and the number of switches is recorded. It is set that a total of 10 tasks are recorded in one sampling cycle, of which 3 tasks have undergone category switching, and the switching ratio is 3 / 10=0.3. The category status mutation frequency threshold is 0.2. The mutation frequency threshold is set based on the maximum mutation ratio allowed by the system. The upper limit of the mutation ratio in the collected historical running data is set to 20%, i.e., 0.2. The comparison result 0.3>0.2, and it is determined that a significant category mutation has occurred in the current cycle, generating a category mutation status value of 1.
[0073] The offset magnitude calculation submodule, based on the category mutation state value, calls the two-period feedback offset values to obtain the maximum amplitude difference of the offset over the time series, and compares it with the feedback change amplitude threshold using the following formula: ;
[0074] The offset normalization amplitude value is calculated and then combined with the first three judgment values to obtain the task difference change value.
[0075] Where D represents the offset normalization amplitude value, This represents the coordinate vector value of the endpoint of the execution path in the current control cycle. This represents the coordinate vector value of the starting point of the execution path in the previous control cycle, and C represents the category mutation state value. This represents the feedback offset value at the i-th time point. Indicates the first The feedback offset values at each time point, where n represents the total number of data points in the feedback offset sequence;
[0076] The offset amplitude calculation submodule extracts the two-cycle feedback offset data sequence based on the category mutation state value. An example feedback sequence is: the previous cycle feedback sequence {1.0, 1.2, 1.1}, the current cycle feedback sequence {1.3, 1.5, 1.6}, and the merged sequence {1.0, 1.2, 1.1, 1.3, 1.5, 1.6}. The maximum value is 1.6, the minimum value is 1.0, and the maximum amplitude difference is 1.6 - 1.0 = 0.6. The feedback change amplitude threshold is 0.5, which is set based on the maximum allowable fluctuation amplitude of the feedback control system. The maximum change in feedback amplitude in historical data is 0.5, so it is set as the feedback change amplitude threshold. Since 0.6 > 0.5, the offset amplitude calculation begins, and the values are substituted into the formula: ;
[0077] Let the coordinate vector of the current cycle endpoint be (8.5, 6.1, 10.2), and the sum of squares be... The starting coordinates of the previous cycle are (2,3,5), and the sum of squares is... , , The sequence differences are 0.2, 0.1, 0.2, 0.2, and 0.1, respectively, and the sum is 0.8.
[0078] ;
[0079] ;
[0080] ;
[0081] The offset normalization magnitude value D≈97.5. When D is added to the path offset judgment value 0 and the category mutation state value 1, the task difference change degree is 97.5+0+1=98.5. This result indicates that the current task state has changed significantly compared with the previous cycle, and dynamic adjustment needs to be performed.
[0082] Please see Figure 3 The modal replacement judgment module includes:
[0083] The error extraction submodule obtains the output response sequence of the master control mode signal and the alternative mode signal in each task segment of the current control cycle based on the task difference change value, extracts the average error value of the corresponding feedback signal in each sequence, records the signal time window position corresponding to the error value, and judges the absolute value difference between the error value and the corresponding error response threshold to obtain the mode error deviation value.
[0084] The error extraction submodule, based on the task difference variation value, sequentially obtains the output response sequences of the master control mode signal and the alternative mode signal within each task segment of the current control cycle. For each task segment, it extracts the master control mode signal feedback sequence and the actual response sequence, calculates the difference between the two sequences at corresponding time points within that task segment, and averages them. In the example, the feedback sequence for task segment 1 is {1.2, 1.3, 1.4}, the actual response sequence is {1.1, 1.2, 1.3}, the calculated difference sequence is {0.1, 0.1, 0.1}, and the average error value is 0.1. This error value is recorded relative to the task segment. The corresponding relationship of time window 1 is repeated to obtain the error value of task segment 2, which is 0.15. This error value is compared with the set error response threshold, which is set to 0.12. This threshold is set according to the task type and response tolerance range. The standard is 80% of the maximum allowable error amplitude of the task. For example, the maximum allowable error of the task is 0.15, so the threshold is 0.12. The absolute value difference judgment is performed. If 0.1 < 0.12, there is no deviation. If 0.15 > 0.12, there is a deviation. The deviation task segment and the deviation value of 0.15 are recorded. Finally, the modal error deviation value sequence {0.0, 0.03} is obtained.
[0085] The matching score judgment submodule, based on the modal error deviation value, calls the matching score value corresponding to the control target for each modal signal within the current control cycle. It then extracts the task relevance index, response stability coefficient, and structural coupling consistency index from the score value to construct a matching composite evaluation expression, using the formula: ;
[0086] The composite scoring error value is obtained through calculation. The magnitude of the composite scoring error value is compared with the modality optimization critical value to determine the modality scoring deviation state and generate the modality scoring difference value.
[0087] Where R represents the composite scoring error value, This represents the task relevance index corresponding to the modal signal. This represents the response stability coefficient of the j-th modal signal within the current control cycle. This represents an index indicating the consistency of the structural coupling between the modal signal and the current task. This represents the score value of the current modal signal. This represents the target expected mean of the modal scores within the current control cycle. This represents the adjustment constant for the matching score evaluation, and K represents the total number of terms in the stability coefficient sequence.
[0088] The matching score judgment submodule reads the matching score values of the main control and candidate modal signals in the current control cycle based on the modal error deviation value. The score value includes the task relevance index. Response stability coefficient sequence Coupling consistency index with structure In the example, the task relevance index of the master modality is 0.85, the response stability coefficient sequence is {0.9, 0.8, 0.85}, the coupling consistency index is 0.75, and the current modality score is invoked. Target expected mean Matching score adjustment constant This adjustment constant is set based on the system control sensitivity parameter, obtained by multiplying the upper limit of the system control error range by an adjustment factor of 1.1. The set range is 1.0 to 1.2, with 1.1 used as an example. Substituting into the formula: ;
[0089] ;
[0090] ;
[0091] ;
[0092] ;
[0093] ;
[0094] ;
[0095] Composite scoring error value This value is compared with the modal optimization threshold of 5.000. The optimization threshold is set as the 95th percentile of the historical task score, for example, 5.000. Since R < 5.000, it is determined that there is a scoring bias, and a modal score difference value of 1 is generated.
[0096] The switching state marking submodule determines whether the main control mode signal simultaneously satisfies the conditions that the error deviation value is greater than the error response threshold and the score difference value is less than the preferred critical value based on the modal score difference value. If both conditions are met, the main control mode signal is marked as a state to be switched, and a candidate list of modes is established to obtain the mode replacement determination result.
[0097] The switching status marking submodule, based on the modal score difference value, first determines whether the master control modal signal meets two conditions. Condition one is that the error deviation value is greater than the error response threshold. In the example, the maximum deviation value is 0.15, and the threshold is 0.12, thus satisfying condition one. Condition two is that the score difference value is less than the preferred threshold value. In the example, 1 < 5.000, which also satisfies condition two. When both conditions are met, the master control modality is marked as a state to be switched. Then, the current candidate modal signal list is obtained, and candidate modalities with a score value greater than 5.000 are filtered. In the example, candidate modal B1 has a score of 5.5, and B2 has a score of 4.8. Only B1 meets the condition, so a candidate list {B1} is established. Finally, the modal replacement judgment result is generated, and the current master control modality needs to enter the switching process. The candidate modality is B1.
[0098] Please see Figure 4 The modal hysteresis recognition module includes:
[0099] The timestamp extraction submodule obtains the response timestamp sequence of image modality, speech modality and action modality within the current control cycle based on the modality replacement judgment result. It performs signal statistical processing on each type of modality signal within a fixed time window, extracts the first response time and the last response time, calculates the average response interval within a unit time window, and generates the modality response interval value.
[0100] The timestamp extraction submodule operates on the modality replacement determination results. Within the current control cycle, it acquires the response timestamp sequences of the image modality, speech modality, and action modality respectively. For each modality signal, it performs signal statistical processing within a fixed time window. For example, the response timestamps of the image modality within a fixed time window are {100ms, 150ms, 200ms}. The first response time of 100ms and the last response time of 200ms are extracted from this sequence. The average response interval within a unit time window is calculated as (200ms-100ms) / 2=50ms. This process is repeated for the speech and action modalities. The first and last response times of the speech modality are 120ms to 220ms with an interval of 50ms, and the first and last response times of the action modality are 130ms to 230ms with an interval of 50ms. Finally, the response interval values for each modality are 50ms, 50ms, and 50ms respectively. These values provide a quantitative indicator for evaluating the time sensitivity and response efficiency of each modality.
[0101] The state variation calculation submodule, based on the modal response interval, calls the state change sequences of the three modes within the response time window, and extracts the state change frequency, the number of overlapping change directions, and the change amplitude values, using the following formula: ;
[0102] The modal linkage offset value is obtained through calculation;
[0103] Where L represents the modal linkage offset value, Indicates the total duration of the statistical period. This represents the difference in the k-th state variation. This represents the variation value of the modal response direction vector. Indicators representing the consistency of modal change direction This represents the mean of the magnitudes of the modal state values. This represents the modal response interval value, and N represents the total number of state variation events;
[0104] The state variation calculation submodule calls the state change sequences of the three modes within the response time window based on the modal response interval value. Taking the image modality as an example, the number of overlapping change directions is 5, the change amplitude value sequence is {2,3,2,4,5}, and the total statistical period is 100ms. Example calculation formula: ;
[0105] ;
[0106] set up ;
[0107] ;
[0108] The calculation yielded a linkage offset value of 0.226 for the image modality. This calculation process was repeated for the speech and action modalities.
[0109] The linkage ratio judgment submodule compares the modal linkage offset value with the set modal disturbance judgment threshold. If it is greater than the judgment threshold, the corresponding mode is recorded as an unstable signal mode, and the response label is set as an abnormal state to obtain abnormal response mode identification information.
[0110] The linkage ratio judgment submodule compares the modal linkage offset value with the set modal disturbance judgment threshold. For example, if the threshold is set to 0.2, and the modal linkage offset value 0.226 is greater than 0.2, the image modality is recorded as an unstable signal modality, and the response label is set as an abnormal state, thereby obtaining abnormal response modal identification information. This judgment sets the threshold based on the actual dynamic characteristics of the modal response and the expected stability requirements to ensure the accuracy of the system response and stable operation. The comparison result of the modal linkage offset value directly affects the modal application decision and the adjustment of subsequent control strategies.
[0111] Please see Figure 5 The modal weight reconstruction module includes:
[0112] The interval extraction submodule obtains abnormal response modal identification information. Based on the control feedback timestamp of each type of modal signal in the previous control cycle, it extracts the first response time and feedback completion time of each type of modality, calculates the duration of action execution in the current cycle, and performs difference processing with the duration of the previous cycle to obtain the modal interval difference value.
[0113] After the interval extraction submodule obtains the abnormal response modality recognition information, it sequentially obtains the control feedback timestamps of the image modality, speech modality, and action modality in the previous control cycle. It extracts the first response time and feedback completion time of each modality in the previous cycle. For example, if the first response time of the image modality in the previous cycle was 120ms, the feedback completion time was 320ms, and the action execution duration was 320ms-120ms=200ms, and the response time of the image modality in the current control cycle was 130ms, the feedback completion time was 350ms, and the duration was 220ms, the difference between the current cycle and the previous cycle was calculated as 220ms-200ms=20ms. This difference is used as the interval difference value of the image modality. Similarly, the difference values of speech modality and action modality are calculated as 15ms and 25ms, respectively. This operation quantifies the fluctuation of modality execution time and provides input for subsequent evaluation.
[0114] The contribution product calculation submodule, based on the modality interval difference value, calls the feature participation weight data of each modality in the fusion feature generation process, using the following formula: ;
[0115] The product evaluation value of modal interval deviation and feature participation ratio is obtained by calculation, and the modal product evaluation value is obtained.
[0116] in, This represents the modal product evaluation value. This indicates the modal feedback end time of the current control cycle. This indicates the start time of the modal feedback in the previous control cycle. This represents the contribution value of a single feature corresponding to the z-th mode. This indicates the total number of features participating in the modal types. This indicates the proportion of modality involved in the feature generation process. This indicates the stability deviation of the current cycle control feedback process;
[0117] The contribution product calculation submodule calls the feature participation weight data in the fusion feature generation process based on the modal interval difference value, and sets the current feedback end time of the image modality. Feedback start time of the previous cycle The sequence of individual feature contribution values is {0.3, 0.2}, and the total number of modal types is... The proportion of features involved Stability offset Substitute into the formula:
[0118] ;
[0119] ;
[0120] ;
[0121] The product evaluation value of the image modality is 132251.25. The speech modality and action modality are calculated according to the same logic, and the example results are 122000.0 and 138000.5 respectively. The modality product evaluation value reflects the combined effect of the time interval deviation and feature contribution of each modality.
[0122] The weight resetting generation submodule is based on the modal product evaluation value. It judges the relationship between the evaluation value of each type of modality and the time contribution threshold, filters out the modalities that are higher than the threshold and sets their weight adjustment factor. It then replaces and updates the original modal weights according to the adjustment factor ratio to obtain the modal participation weight reconstruction value.
[0123] The weight resetting generation submodule compares the product evaluation value of each modality with the time contribution threshold. The time contribution threshold is set to 130000.0. This threshold is set based on the minimum contribution value required for the fusion feature generation process multiplied by the maximum stable offset tolerance required for periodic performance. Assuming the minimum contribution value is 100000.0 and the stable offset tolerance factor is 1.3, the threshold is 130000.0. The image modality evaluation value (132251.25) > the threshold, satisfying the adjustment condition. The speech modality evaluation value (122000) is also acceptable. 0 < threshold, not satisfied; action modality 138000.5 > threshold, satisfied. Set weight adjustment factor of 1.1 for image modality and factor of 1.2 for action modality. The original weights are 0.3 for image and 0.4 for action modality. After adjustment, the image weight is 0.3 × 1.1 = 0.33 and the action modality weight is 0.4 × 1.2 = 0.48. Perform replacement update operation to obtain new modality participation weight reconstruction values: image 0.33, speech remains unchanged at 0.3, and action is 0.48.
[0124] Please see Figure 6 The control command output module includes:
[0125] The weight mapping submodule obtains the modal participation weight reconstruction value, calls the modal fusion output vector within the current control cycle, performs positioning operation on the fusion component of each modal dimension, assigns the corresponding modal participation weight to the fusion component, and forms a weighted result matrix to obtain the modal weighted combination value.
[0126] After obtaining the modality participation weight reconstruction values, the weight mapping submodule first calls the fusion output vector jointly generated by the image modality, speech modality, and action modality within the current control cycle, and performs component localization according to the modality dimension. For example, if the fusion output vector is a three-dimensional vector... The corresponding image modality is the first dimension, the speech modality is the second dimension, and the action modality is the third dimension. Combining the reconstructed modalities with the weights, the image modality has a weight of 0.33, the speech modality has a weight of 0.30, and the action modality has a weight of 0.48. The mapping operation between the components and their corresponding weights is performed sequentially. Each dimension's fused component is multiplied by its weight coefficient. The weighted image modality component is... The speech modality is Action mode is The weighted components are combined into a matrix. The results are recorded in the modal weighting result matrix and, combined with the control task identifier, the modal weighting combination value for the corresponding control cycle is generated.
[0127] The fusion encoding construction submodule constructs the instruction output vector within the task cycle based on the modal weighted combination value. It combines all modal weighted component values with the scene label matching index, structural information coupling factor, and instruction mapping template index using the following formula: ;
[0128] The operation yields the fused output vector value;
[0129] in, This represents the fused output vector value. This represents the value of the x-th fusion component. This represents the participation weight of the x-th dimension fusion component. This represents the structural coupling factor of the x-th mapping template. This represents the matching index of the x-th modality label. This represents the target value of the x-th scene task label. Indicates the number of dimensions of the fused output vector;
[0130] The fusion encoding construction submodule performs fusion processing on the modal weighted combination values according to dimensions within the task cycle. Let the first-dimensional weighted component of the image modality be... Weight Structural coupling factor Modal matching index Target value of task tag Substitute into the formula:
[0131] ;
[0132] ;
[0133] ;
[0134] ;
[0135] ;
[0136] ;
[0137] ;
[0138] ;
[0139] The fused output vector value is 0.34, which is used for subsequent instruction generation.
[0140] The task instruction generation submodule inputs the fused output vector value into the control instruction mapping matrix column by column, performs matrix channel matching and index replacement operations according to the instruction path configuration rules, obtains the set of task mapping results for each instruction channel, and establishes the current cycle control instruction encoding information.
[0141] The task instruction generation submodule inputs the fused output vector value 0.34 as the instruction parameter for the current cycle into the control instruction mapping matrix. Multiple preset instruction path mapping channels are arranged in columns within the matrix. The matrix channel matching process is executed. If the mapping matrix index range is 0.30 to 0.40, then 0.34 falls within this range, matching index number 5. An instruction path replacement operation is performed, loading the task control instruction template under index channel 5. The output encoded content is the task identifier A001, corresponding to the instruction "MOVE-35," and is matched with the task cycle identifier, completing the generation of control instruction encoding information within the control cycle.
[0142] Table 1. Examples of Modal Dimension Mapping and Fusion Operations
[0143]
[0144] Table 1 lists the participating terms and calculation results of each modal component in the fusion operation, combined with the fusion coding construction steps. The results show that the mapping index corresponding to the fusion output vector value of 0.34 matches channel 5, thus deriving the task instruction code for the current control cycle as "MOVE-35".
[0145] A multimodal information fusion control method includes the following steps:
[0146] S1: Obtain the control execution parameters of the multimodal task within adjacent control cycles, calculate the path coordinate difference, category variation frequency and offset change magnitude, call the three difference thresholds of path, category and feedback for judgment, and generate the task difference change value;
[0147] S2: Based on the task difference change value, compare whether the output error of the master mode is greater than the error response threshold and whether the mode matching score is lower than the mode selection threshold. If both conditions are met, mark the master mode as a state to be switched and generate a mode replacement judgment result.
[0148] S3: Based on the modal replacement judgment result, calculate the linkage ratio between modal response time and state change amplitude, determine whether the linkage ratio exceeds the modal disturbance judgment threshold, and obtain abnormal response modal identification information;
[0149] S4: Based on the abnormal response modality identification information, call the feature participation weight of the corresponding modality in the fusion feature generation, calculate the product between the time interval deviation value and the feature participation weight, set a weight adjustment factor for modalities whose product result is greater than the time contribution threshold, and allocate new weights according to the adjustment factor to obtain the modality participation weight reconstruction value.
[0150] S5: Based on the modal participation weight reconstruction value, the fusion components of each modal dimension are weighted and combined according to the updated weight. The weighted fusion vector is input into the instruction mapping matrix to generate the fusion instruction code set of the task scenario in the current cycle, and the control instruction code information of the current cycle is obtained.
[0151] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A multimodal information fusion control system, characterized in that, The system includes: The target difference assessment module obtains the control execution parameters of multimodal tasks within adjacent control cycles, calculates the path coordinate difference, category variation frequency and offset change magnitude, calls the three difference thresholds of path, category and feedback for judgment, and generates the task difference change value. The modality replacement judgment module compares the output error of the master control mode with the task difference change value to see if it is greater than the error response threshold and whether the modality matching score is lower than the modality preference threshold. If both conditions are met, the master control mode is marked as a state to be switched and a modality replacement judgment result is generated. Based on the modal replacement determination result, the modal hysteresis identification module calculates the linkage ratio between modal response time and state change amplitude, determines whether the linkage ratio exceeds the modal disturbance judgment threshold, and obtains abnormal response modal identification information. The modality weight reconstruction module calls the feature participation ratio of the corresponding modality in the fusion feature generation according to the abnormal response modality identification information, calculates the product between the time interval deviation value and the feature participation ratio, sets a weight adjustment factor for the modality whose product result is greater than the time contribution threshold, and performs new weight allocation according to the adjustment factor to obtain the modality participation weight reconstruction value. The modal weight reconstruction module includes: The interval extraction submodule obtains the abnormal response modality identification information, extracts the first response time and feedback completion time of each modality based on the control feedback timestamp of each modality signal in the previous control cycle, calculates the duration of action execution in the current cycle, and performs difference processing with the duration of the previous cycle to obtain the modality interval difference value. The contribution product calculation submodule, based on the modality interval difference value, calls the feature participation weight data of each modality in the fusion feature generation process, using the formula: ; The product evaluation value of modal interval deviation and feature participation ratio is obtained by calculation, and the modal product evaluation value is obtained. in, This represents the modal product evaluation value. This indicates the modal feedback end time of the current control cycle. This indicates the start time of the modal feedback in the previous control cycle. This represents the contribution value of a single feature corresponding to the z-th mode. This indicates the total number of features participating in the modal types. This indicates the proportion of modality involved in the feature generation process. This indicates the stability deviation of the current cycle control feedback process; The weight resetting generation submodule, based on the modal product evaluation value, determines the relationship between the evaluation value of each modality and the time contribution threshold, filters out modes that are higher than the threshold and sets their weight adjustment factors, and performs replacement and update on the original modal weights according to the adjustment factor ratio to obtain the modality participation weight reconstruction value. The task difference change value specifically includes path offset judgment value, category mutation state value, and feedback change amplitude value. The mode replacement judgment result includes mode switching trigger identifier, master mode degradation mark, and alternative mode selection mark. The abnormal response mode identification information specifically includes state disturbance index, response delay level, and signal validity label. The mode participation weight reconstruction value includes participation weight adjustment coefficient, feature contribution attenuation factor, and time penalty allocation value.
2. The multimodal information fusion control system according to claim 1, characterized in that, The target difference assessment module includes: The path offset calculation submodule obtains the control execution parameters of the multimodal task within adjacent control cycles, extracts the coordinate vector values of the starting point coordinates and the ending point coordinates of the path, calculates the Euclidean distance difference between the coordinate vectors of the starting point and the ending point of the path in two cycles, and compares the absolute value difference between the difference result and the set path offset threshold to generate a path offset judgment value. The category variation detection submodule calls the target category number in the two control cycles before and after the path offset judgment value, extracts whether the category number status has changed, counts the number of category numbers that have changed and performs periodic normalization processing, judges the relationship between the result and the category status mutation frequency threshold, and generates a category mutation status value. The offset magnitude calculation submodule, based on the aforementioned category mutation state value, calls the two-period feedback offset values to obtain the maximum amplitude difference of the offset over the time series, and compares it with the feedback change amplitude threshold using the following formula: ; The offset normalization amplitude value is calculated and then combined with the first three judgment values to obtain the task difference change value. Where D represents the offset normalization amplitude value, This represents the coordinate vector value of the endpoint of the execution path in the current control cycle. This represents the coordinate vector value of the starting point of the execution path in the previous control cycle, and C represents the category mutation state value. This represents the feedback offset value at the i-th time point. Indicates the first The feedback offset values at each time point, where n represents the total number of data points in the feedback offset sequence.
3. The multimodal information fusion control system according to claim 1, characterized in that, The modality replacement determination module includes: The error extraction submodule obtains the output response sequence of the main control mode signal and the alternative mode signal in each task segment of the current control cycle based on the task difference change value, extracts the average error value of the corresponding feedback signal in each sequence, records the signal time window position corresponding to the error value, and judges the absolute value difference between the error value and the corresponding error response threshold to obtain the mode error deviation value. The matching score judgment submodule, based on the modal error deviation value, calls the matching score value corresponding to the control target for each modal signal within the current control cycle, extracts the task relevance index, response stability coefficient, and structural coupling consistency index of the score value, and constructs a matching composite evaluation expression, using the formula: ; The composite scoring error value is obtained through calculation. The magnitude of the composite scoring error value is compared with the modality optimization critical value to determine the modality scoring deviation state and generate the modality scoring difference value. Where R represents the composite scoring error value, This represents the task relevance index corresponding to the modal signal. This represents the response stability coefficient of the j-th modal signal within the current control cycle. This represents an index indicating the consistency of the structural coupling between the modal signal and the current task. This represents the score value of the current modal signal. This represents the target expected mean of the modal scores within the current control cycle. This represents the adjustment constant for the matching score evaluation, and K represents the total number of terms in the stability coefficient sequence. The switching state marking submodule determines whether the master control mode signal simultaneously satisfies the condition that the error deviation value is greater than the error response threshold and the score difference value is less than the preferred critical value, based on the modal score difference value. If both conditions are met, the master control mode signal is marked as a state to be switched, and a candidate list of modes is established to obtain the mode replacement determination result.
4. The multimodal information fusion control system according to claim 1, characterized in that, The modal hysteresis recognition module includes: The timestamp extraction submodule obtains the response timestamp sequence of image modality, voice modality and action modality within the current control cycle based on the modality replacement determination result, performs signal statistical processing on each type of modality signal within a fixed time window, extracts the first response time and the last response time, calculates the average response interval within a unit time window, and generates a modality response interval value. The state variation calculation submodule, based on the modal response interval value, calls the state change sequences of the three modes within the response time window, and extracts the state change frequency, the number of overlapping change directions, and the change amplitude values, respectively, using the formula: ; The modal linkage offset value is obtained through calculation; Where L represents the modal linkage offset value, Indicates the total duration of the statistical period. This represents the difference in the k-th state variation. This represents the variation value of the modal response direction vector. Indicators representing the consistency of modal change direction This represents the mean of the magnitudes of the modal state values. This represents the modal response interval value, and N represents the total number of state variation events; The linkage ratio judgment submodule compares the modal linkage offset value with the set modal disturbance judgment threshold. If the value is greater than the judgment threshold, the corresponding modality is recorded as an unstable signal modality, and the response label is set as an abnormal state to obtain abnormal response modality identification information.
5. The multimodal information fusion control system according to claim 1, characterized in that, The system also includes: The control command output module calls the modal fusion output vector in the current period according to the modal participation weight reconstruction value, weights and combines the fusion components of each modal dimension according to the updated weight, inputs the weighted fusion vector into the command mapping matrix, generates the fusion command code set of the task scenario in the current period, and obtains the control command code information of the current period. The current cycle control instruction encoding information specifically refers to the fusion vector mapping number, task control instruction group code, and signal channel activation index.
6. The multimodal information fusion control system according to claim 5, characterized in that, The control command output module includes: The weight mapping submodule obtains the modal participation weight reconstruction value, calls the modal fusion output vector in the current control cycle, performs a positioning operation on the fusion component of each modal dimension, assigns the corresponding modal participation weight to the fusion component, and forms a weighted result matrix to obtain the modal weighted combination value. The fusion encoding construction submodule constructs an instruction output vector within the task cycle based on the modal weighted combination value. It combines all modal weighted component values with the scene label matching index, structural information coupling factor, and instruction mapping template index using the following formula: ; The operation yields the fused output vector value; in, This represents the fused output vector value. This represents the value of the x-th fused component. This represents the participation weight of the x-th dimension fusion component. This represents the structural coupling factor of the x-th mapping template. This represents the matching index of the x-th modality label. This represents the target value of the x-th scene task label. Indicates the number of dimensions of the fused output vector; The task instruction generation submodule inputs the fused output vector value into the control instruction mapping matrix column by column, performs matrix channel matching and index replacement operations according to the instruction path configuration rules, obtains the set of task mapping results for each instruction channel, and establishes the current cycle control instruction encoding information.
7. A multimodal information fusion control method, characterized in that, The multimodal information fusion control system according to any one of claims 1-6 is executed by including the following steps: S1: Obtain the control execution parameters of the multimodal task within adjacent control cycles, calculate the path coordinate difference, category variation frequency and offset change magnitude, call the three difference thresholds of path, category and feedback for judgment, and generate the task difference change value; S2: Based on the task difference change value, compare whether the output error of the master control mode is greater than the error response threshold and whether the mode matching score is lower than the mode preference threshold. If both conditions are met, mark the master control mode as a state to be switched and generate a mode replacement judgment result. S3: Based on the modal replacement determination result, calculate the linkage ratio between the modal response time and the state change amplitude, determine whether the linkage ratio exceeds the modal disturbance judgment threshold, and obtain abnormal response modal identification information; S4: Based on the abnormal response modality identification information, calculate the product between the time interval deviation value and the feature participation ratio, set a weight adjustment factor for the modality whose product result is greater than the time contribution threshold, and perform new weight allocation according to the adjustment factor to obtain the modality participation weight reconstruction value. S5: Based on the modal participation weight reconstruction value, the fusion components of each modal dimension are weighted and combined according to the updated weight. The weighted fusion vector is input into the instruction mapping matrix to generate the fusion instruction code set of the task scenario in the current period, and the control instruction code information of the current period is obtained.
Citation Information
Patent Citations
Multi-modal data fusion method and device, computer equipment and storage medium
CN119691687A
AI-based building design optimization system and method
CN119962060A