Multimodal perception fusion human-computer interaction intelligent recognition system
By introducing the fusion hysteresis drift failure model of the modal confidence imbalance coefficient and the modal missing detection delay coefficient, the problems of modal degradation and missing detection lag in the multimodal perception fusion system are solved, and the stability and accuracy of human-computer interaction recognition are improved.
Patent Information
- Application Number
- CN202510864576.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-06-26
AI Technical Summary
When modal degradation is not recognized in time and modal loss detection has a delayed response in the existing human-computer interaction system based on multimodal perception fusion, it leads to abnormal modal misdominance and strategy drift, affecting the accuracy and stability of interactive recognition.
The modal confidence imbalance coefficient and modal missing detection delay coefficient are introduced to construct a fusion hysteresis drift failure model. The modal contribution trust distribution is evaluated through the modal confidence imbalance coefficient. The modal missing is detected in real time and the modal missing detection delay is calculated. The fusion hysteresis drift failure coupling index is output for dynamic early warning.
It improves the recognition stability and accuracy of the multimodal fusion strategy in complex environments, avoids the abnormal mode from misleading, promptly identifies the missing mode, prevents the long-term drift of the fusion strategy, and enhances the system's adaptability to complex environments.
Smart Images

Figure CN120354100B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer interaction intelligent recognition, and more specifically, to a human-computer interaction intelligent recognition system based on multimodal perception fusion. Background Art
[0002] With the development of multimodal perception technology, human-computer interaction systems have gradually transitioned from single-channel recognition to intelligent recognition that integrates multimodal signals such as speech, vision, gesture, touch, and EEG. Multimodal perception fusion can improve interaction robustness and response accuracy in complex environments, and has demonstrated broad application potential in scenarios such as intelligent customer service, virtual assistants, medical assistance, and smart driving. Existing multimodal fusion methods typically rely on dynamic weighting of confidence estimates of each modal signal, supplemented by a modal state judgment mechanism to achieve fault tolerance for missing modalities.
[0003] However, as applications expand and usage scenarios become more complex, multimodal systems are experiencing internal defects such as failure to promptly identify modal degradation and delayed response to modal loss detection. This leads to hidden risks of "abnormal modal misleading" or "strategy drift reinforcement" in the fusion decision-making process. In particular, in actual operation, when modal data degrades or is completely missing due to factors such as occlusion, channel instability, or device failure, if the system fails to quickly detect and eliminate the modality and instead continues to rely on it for decision-making, it will seriously affect the accuracy of interactive recognition. More critically, if the system confidence assessment mechanism fails to promptly lower its fusion weight while modal loss detection is delayed, the two will form a coupled hysteresis chain, which will not only induce short-term recognition misjudgments, but may also continuously interfere with model learning through false feedback, causing long-term drift in the fusion strategy and forming a systemic risk of fusion hysteresis drift failure. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a multi-modal perception fusion human-computer interaction intelligent recognition system to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] The multimodal perception fusion human-computer interaction intelligent recognition system includes a multimodal data acquisition module, a modal state perception module, a modal loss detection and delay recognition module, a modal fusion hysteresis drift failure module, and a fusion strategy dynamic warning module;
[0007] Multimodal data acquisition module, used to obtain raw modal inputs from multiple channels and synchronize them through time stamps;
[0008] The modal state perception module is used to obtain the modal confidence imbalance information of multimodal data, establish the modal confidence imbalance function, calculate the modal confidence imbalance coefficient, and evaluate the degree of modal degradation imbalance;
[0009] A modal loss detection delay identification module is used to perform loss detection on the modal data link, obtain the modal loss detection delay information of the modal data link, calculate the modal loss detection delay coefficient, and evaluate the degree of modal loss detection delay;
[0010] The modal fusion hysteresis drift failure module is used to construct a modal fusion hysteresis drift failure model based on the modal confidence imbalance coefficient and the modal missing detection delay coefficient, output the fusion hysteresis drift failure coupling index, and determine the potential drift failure risk of the current fusion strategy;
[0011] The fusion strategy dynamic warning module is used to provide dynamic warnings for current human-computer interaction operations based on the potential drift failure risk of the fusion strategy.
[0012] In a preferred embodiment, the logic for obtaining the modal confidence imbalance coefficient is as follows:
[0013] At each moment, the confidence of all modalities involved in the fusion is obtained to form a confidence vector: ,in is the confidence of the nth mode participating in the fusion at time t, is the total number of modalities currently involved in fusion, , and the normalization conditions should be met before fusion: ;
[0014] In the time window Inside, construct a three-dimensional tensor ,in is a three-dimensional tensor consisting of elements composition, Represents the confidence value of the f-th feature of the n-th modality participating in the fusion at time t, is the time window length, The confidence level of each modality forms a feature dimension;
[0015] At each time instant t, a cooperative offset matrix is constructed based on the tensor slices. , the elements in the cooperative offset matrix are the cooperative degrees of the confidence difference between mode n and mode m. The calculation expression of the cooperative degree of the confidence difference between mode n and mode m is as follows: ,in is the confidence difference coordination degree, Represents the confidence value of the fth feature of the mth modality participating in the fusion at time t;
[0016] Calculate the tensor-form structural norm of all confidence difference synergies within the time window as the deviation coupling metric : ,in , ;
[0017] In order to enhance the sensitivity of the model to the difference in modal information, the confidence dispersion of the modal assurance criterion is defined as: ,in is the confidence dispersion of the nth mode participating in the fusion at time t, is the confidence ratio probability, ; and calculate the perturbation difference of the confidence dispersion within the time window : ,in is the confidence dispersion of the nth mode participating in fusion at time t+1; calculate the modal dispersion disturbance factor : ,in is the disturbance difference of the mth mode participating in the fusion, ;
[0018] Calculate the modal assurance criterion imbalance coefficient: ,in is the modal assurance imbalance coefficient of mode n.
[0019] In a preferred embodiment, the logic for obtaining the modality loss detection delay coefficient is as follows:
[0020] Constructing a modal propagation diagram ,in is a node set, each mode corresponds to a node, The information propagation relationship between each modality in the edge set is through the edge Indicates that the weight of each edge is expressed as the modal information offset strength. The calculation expression is as follows: ,in represents the gradient of the missing probability of mode n at time t, ,in is the missing state probability output of the system for mode n at time t, represents the gradient of the missing probability of mode m at time t;
[0021] The actual missing moment of the mode is determined based on the entropy mutation and the gradient of the missing probability of mode n at time t. The judgment logic is as follows: Indicates finding the earliest moment on the timeline , which makes the entropy value of mode n suddenly change Exceeding the preset mutation threshold And the gradient of the probability of mode n missing at time t Exceeding the preset gradient threshold ;in is the probability distribution entropy of mode n;
[0022] Modal missing response detection time calculate: Represents the probability output of the missing state of mode n at time t Exceeds the preset missing judgment threshold Or there is a mode m adjacent to mode n, whose modal information offset strength with mode n is Greater than the preset offset strength threshold And the missing state probability output of mode m Exceeds the preset missing judgment threshold ,in is the neighborhood mode set of mode n;
[0023] Calculate the modal absence detection delay coefficient: ,in is the mode loss detection delay coefficient of mode n, is the maximum delay that the system can tolerate.
[0024] In a preferred embodiment, a modal fusion hysteresis drift failure model is constructed according to the modal confidence imbalance coefficient and the modal missing detection delay coefficient, and a fusion hysteresis drift failure coupling index is output. The modal fusion hysteresis drift failure model is based on the following formula: , where is the fusion hysteresis drift failure coupling index, is the modal assurance imbalance coefficient of mode n, is the mode loss detection delay coefficient of mode n, is the total number of modalities currently involved in fusion, represent the preset proportional coefficients of the modal confidence imbalance coefficient and the modal missing detection delay coefficient, respectively, and Both are greater than 0.
[0025] In a preferred embodiment, the fusion hysteresis drift failure coupling index is compared with a preset fusion hysteresis drift failure coupling index threshold to determine the potential drift failure risk of the current fusion strategy, as follows:
[0026] If the fusion hysteresis drift failure coupling index is greater than the fusion hysteresis drift failure coupling index threshold, the current fusion strategy has a potential drift failure risk;
[0027] If the fusion hysteresis drift failure coupling index is less than or equal to the fusion hysteresis drift failure coupling index threshold, the current fusion strategy does not have an obvious potential drift failure risk and the current strategy can be maintained.
[0028] In a preferred embodiment, if the current fusion strategy has a potential drift failure risk, the fusion hysteresis drift failure coupling index is obtained and combined with the current human-computer interaction duration to obtain a warning coefficient: ,in is the warning coefficient, The duration of human-computer interaction.
[0029] In a preferred embodiment, the warning coefficient is compared with a preset warning coefficient threshold to perform a dynamic warning on the current human-computer interaction operation, as follows:
[0030] If the warning coefficient is greater than the warning coefficient threshold, a dynamic warning signal is generated;
[0031] If the warning coefficient is less than or equal to the warning coefficient threshold, there is no need to generate a dynamic warning signal.
[0032] The technical effects and advantages of the present invention are as follows:
[0033] 1. The present invention introduces the modal confidence imbalance coefficient and the modal missing detection delay coefficient, and based on this, constructs a fusion hysteresis drift failure model to achieve real-time quantitative assessment of the potential risks of multimodal fusion strategies, thereby effectively improving the stability and accuracy of interactive recognition in complex environments. First, the modal confidence imbalance coefficient is used to dynamically reflect the distribution of the contribution trust between each modality, avoiding the risk of wrong decision-making caused by "abnormal modal misleading"; secondly, by performing real-time missing detection on the modal data link and calculating the modal missing detection delay coefficient, the time interval required from the actual modal failure to the system capturing the failure is quantified. The fusion hysteresis drift failure coupling index output after coupling the two is used as a key indicator to measure the potential drift failure risk of the current fusion strategy. When the index exceeds the preset threshold, the fusion strategy dynamic warning module is triggered in time to issue a risk warning for human-computer interaction operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings;
[0035] Figure 1 Flowchart of the system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0037] Embodiment: The present invention provides Figure 1 The multimodal perception fusion human-computer interaction intelligent recognition system shown includes a multimodal data acquisition module, a modal state perception module, a modal loss detection delay recognition module, a modal fusion hysteresis drift failure module, and a fusion strategy dynamic warning module;
[0038] Multimodal data acquisition module, used to obtain raw modal inputs from multiple channels and synchronize them through time stamps;
[0039] The modal state perception module is used to obtain the modal confidence imbalance information of multimodal data, establish the modal confidence imbalance function, calculate the modal confidence imbalance coefficient, and evaluate the degree of modal degradation imbalance;
[0040] A modal loss detection delay identification module is used to perform loss detection on the modal data link, obtain the modal loss detection delay information of the modal data link, calculate the modal loss detection delay coefficient, and evaluate the degree of modal loss detection delay;
[0041] The modal fusion hysteresis drift failure module is used to construct a modal fusion hysteresis drift failure model based on the modal confidence imbalance coefficient and the modal missing detection delay coefficient, output the fusion hysteresis drift failure coupling index, and determine the potential drift failure risk of the current fusion strategy;
[0042] The fusion strategy dynamic warning module is used to provide dynamic warnings for current human-computer interaction operations based on the potential drift failure risk of the fusion strategy;
[0043] In the multimodal data acquisition module, a multi-channel input interface is established (supporting heterogeneous modal protocols such as USB video streaming, audio I2S, and EEG BLE) to obtain raw modal input from multiple channels. A modal identifier is defined for each modality, and modal sampling parameters are initialized, including sampling frequency, communication protocol specifications, data packet header format, and data frame structure. A unified system time base (such as UNIX_TIMESTAMP() or a system high-precision clock) is used to assign an acquisition timestamp to each modal data frame. Indicates the timestamp of the i-th mode collected at the k-th time step, establishes a buffer channel for the asynchronous mode, and records the time deviation : ,in As the reference modal time, a synchronous calibration mechanism is constructed, and a sliding window is used to maintain the synchronization of the modal sampling window: ,in Indicates the timestamp corresponding to the modal channel selected as the synchronization reference in the kth time step (used to find a modal channel i such that the acquisition timestamp of the modality is closest to the average of the acquisition timestamps of other modalities, thus serving as the synchronization reference for the current time step k). is the average value of all modal acquisition times at the kth time step: ,in is the number of modal types;
[0044] The modal state perception module is used to obtain the modal confidence imbalance information of multimodal data, establish the modal confidence imbalance function, calculate the modal confidence imbalance coefficient, and evaluate the degree of modal degradation imbalance;
[0045] The modal confidence imbalance coefficient in the present invention is a key metric used to measure whether there is a significant deviation in the confidence distribution between modalities in a multimodal perception fusion human-computer interaction intelligent recognition system. Its core role is to quantify the differences and bias in the actual influence of the information carried by different modalities in decision-making during the multimodal data fusion process. When a modality obtains an excessively high or low confidence value in system recognition due to reasons such as data quality, environmental interference, or algorithm mechanism, the system fusion layer will assign it an asymmetric fusion weight, thereby affecting the directionality of the overall recognition judgment. The modal confidence imbalance coefficient is a quantitative expression of this bias. In actual operation, a large modal confidence imbalance coefficient usually means that certain modalities have an excessively strong dominance over the system recognition results, which may mask the effective contribution of other modalities to the target. In particular, when the quality of the dominant modality is reduced due to factors such as occlusion, signal interference, resolution degradation, or equipment failure, it is still assigned a high weight by the system, which can easily lead to the problem of "abnormal modality misleading". Such risks are highly insidious and persistent, potentially leading not only to short-term misidentification but also to the reinforcement of erroneous patterns during long-term feedback training, causing structural shifts in the fusion strategy and irreversible policy drift. In contrast, a smaller modal confidence imbalance coefficient indicates that the confidence distributions across modalities are becoming more consistent, and the system shows no significant bias toward a particular modality during the perceptual fusion process. This indicates that the multimodal perception structure is relatively stable in the current environment, and that the fusion judgment mechanism is able to fully absorb the complementarity of multi-source information. Recognition results in this state are more holistic and robust, enhancing adaptability to complex environments while avoiding systematic recognition errors caused by fluctuations in a single modality. Assessing the potential drift failure risk of the current fusion strategy based on the modal confidence imbalance coefficient can provide the system with a mechanism for quantitatively analyzing the distribution of modal information contributions, enabling the system to perceive the actual impact of each modal signal state on the fusion strategy in real time during operation. Secondly, as an important indicator of the robustness of the fusion strategy, this coefficient can be used together with the modal loss detection delay coefficient to construct a fusion hysteresis drift failure model and output a fusion hysteresis drift failure coupling index, providing a reliable basis for high-order strategy optimization and interactive decision reconstruction. In summary, the modal confidence imbalance coefficient is not only a static analysis tool, but also a core variable for dynamic perception and risk prevention of fusion strategies. It enables the system to evolve from "passive fusion" to "adaptive optimization fusion" and is a key support for ensuring the long-term stable operation of multimodal human-computer interaction systems in complex and dynamic environments.
[0046] The logic for obtaining the modal assurance imbalance coefficient is as follows:
[0047] At each moment, the confidence of all modalities involved in the fusion is obtained to form a confidence vector: ,in is the confidence of the nth mode participating in the fusion at time t, is the total number of modalities currently involved in fusion, , and the normalization conditions should be met before fusion: ;
[0048] It should be noted that the confidence level of each modality can be calculated by different platforms based on their own actual conditions (e.g., using a deep learning model);
[0049] In the time window Inside, construct a three-dimensional tensor ( Represents a set of real numbers, that is, a three-dimensional tensor are all real numbers), where Is a three-dimensional tensor that represents the confidence set under all modes, all time frames, and all feature dimensions, consisting of elements composition, Represents the confidence value of the f-th feature of the n-th modality participating in the fusion at time t, is the time window length, The confidence level of each modality forms a feature dimension;
[0050] At each time instant t, a cooperative offset matrix is constructed based on the tensor slices. , the elements in the cooperative offset matrix are the cooperative degrees of the confidence difference between mode n and mode m. The calculation expression of the cooperative degree of the confidence difference between mode n and mode m is as follows: ,in is the confidence difference coordination degree, Represents the confidence value of the fth feature of the mth modality participating in the fusion at time t;
[0051] Calculate the tensor-form structural norm of all confidence difference synergies within the time window as the deviation coupling metric : ,in , ; The larger the deviation coupling metric value, the more persistent and strong the non-equilibrium coordination deviation between the multimodal confidences;
[0052] In order to enhance the sensitivity of the model to the difference in modal information, the confidence dispersion of the modal assurance criterion is defined as: ,in is the confidence dispersion of the nth mode participating in the fusion at time t, is the confidence ratio probability, ; and calculate the perturbation difference of the confidence dispersion within the time window : ,in is the confidence dispersion of the nth mode participating in fusion at time t+1; calculate the modal dispersion disturbance factor : ,in is the disturbance difference of the mth mode participating in the fusion, ;
[0053] Calculate the modal assurance criterion imbalance coefficient: ,in is the modal assurance imbalance coefficient of mode n;
[0054] It should be noted that the above formulas are all dimensionless and numerical calculations. Common dimensionless methods include Min-Max normalization and Z-Score normalization, which will not be described here.
[0055] A modal loss detection delay identification module is used to perform loss detection on the modal data link, obtain the modal loss detection delay information of the modal data link, calculate the modal loss detection delay coefficient, and evaluate the degree of modal loss detection delay;
[0056] The modality missing detection delay coefficient in the present invention is a quantitative indicator used to measure the degree of time delay required for the detection and identification of the missing state of modal data in a multimodal perception system. The core purpose is to accurately reflect the source of systemic risk of "abnormal modalities not identified in time" in the multimodal fusion decision-making mechanism. Specifically, the larger the modality missing detection delay coefficient, the slower the system responds to the missing of a certain modality, and there is a possibility that the invalid modality will still be involved in the fusion calculation for a long period of time, thereby increasing the risk of the system being dominated by the wrong modality in the fusion strategy; and the smaller the modality missing detection delay coefficient, the faster the system has the ability to detect and respond to abnormal modal changes, and can complete fusion control operations such as modal elimination and confidence weight adjustment in a timely manner, which helps to ensure the stability and accuracy of the fusion strategy. In the fusion scenario, modality missing may be caused by various factors such as sensor occlusion, channel interruption, hardware failure, and accidental shielding. If the system fails to promptly detect interruptions or distortions (i.e., missing) in the modal data stream and continues to treat it as participating in the fusion process, this will result in a "false contribution" from that modality in the decision output. This can easily lead to misleading identification, especially when its initial confidence is high or its weight is significant. This misleading behavior not only leads to short-term output errors but also triggers a chain of error feedback in interactive human-computer systems, causing cumulative offsets in long-term policy learning. The introduction of a modality loss detection delay coefficient provides a structural mitigation mechanism for this problem. By defining a modality loss detection delay coefficient, this invention enables the construction of a quantitative model of the delay response mechanism within the multimodal fusion system. This coefficient calculates the time difference between the actual occurrence of a modality loss and the system's confirmation of the missing state, and then corrects it based on the temporal continuity of the modal data and false triggering, ultimately outputting a stable and usable evaluation value. Furthermore, the modality loss detection delay coefficient can be coupled with the modal confidence imbalance coefficient to be input into the fusion hysteresis drift failure model to calculate the fusion hysteresis drift failure coupling index, thereby identifying potential systemic risk trends in the current fusion strategy at a global level. The beneficial effects of this method are that it can not only capture and respond to the functional attenuation process of local modes in a timely manner, but also identify the drift direction of the fusion mechanism in the evolution path in advance through coefficient evolution trend analysis, realize the prediction and early warning of fusion strategy deviation, enhance the system's temporal sensitivity to changes in modal input reliability, reduce false fusion and recognition deviation caused by detection lag, improve the system's ability to intervene in the abnormal perception chain, and provide a solid foundation for subsequent fusion strategy optimization, active modal screening and modal switching control.
[0057] The logic for obtaining the modality loss detection delay coefficient is as follows:
[0058] Constructing a modal propagation diagram ,in is a node set, each mode corresponds to a node, The information propagation relationship between each modality in the edge set is through the edge Indicates that the weight of each edge is expressed as the modal information offset strength. The calculation expression is as follows: ,in represents the gradient of the missing probability of mode n at time t, ,in is the missing state probability output of the system for mode n at time t, represents the gradient of the missing probability of mode m at time t;
[0059] The actual missing moment of the mode is determined based on the entropy mutation and the gradient of the missing probability of mode n at time t. The judgment logic is as follows: Indicates finding the earliest moment on the timeline , which makes the entropy value of mode n suddenly change Exceeding the preset mutation threshold And the gradient of the probability of mode n missing at time t Exceeding the preset gradient threshold ;in is the probability distribution entropy of mode n;
[0060] Modal missing response detection time calculate: Represents the probability output of the missing state of mode n at time t Exceeds the preset missing judgment threshold Or there is a mode m adjacent to mode n, whose modal information offset strength with mode n is Greater than the preset offset strength threshold And the missing state probability output of mode m Exceeds the preset missing judgment threshold ,in is the neighborhood mode set of mode n;
[0061] Calculate the modal absence detection delay coefficient: ,in is the mode loss detection delay coefficient of mode n, is the maximum delay that the system can tolerate;
[0062] It should be noted that the above formulas are all dimensionless and numerical calculations. Common dimensionless methods include Min-Max normalization and Z-Score normalization, which will not be described here.
[0063] The modal fusion hysteresis drift failure module is used to construct a modal fusion hysteresis drift failure model based on the modal confidence imbalance coefficient and the modal missing detection delay coefficient, output the fusion hysteresis drift failure coupling index, and determine the potential drift failure risk of the current fusion strategy;
[0064] According to the modal confidence imbalance coefficient and the modal missing detection delay coefficient, a modal fusion hysteresis drift failure model is constructed, and the fusion hysteresis drift failure coupling index is output. The modal fusion hysteresis drift failure model is based on the following formula: , where is the fusion hysteresis drift failure coupling index, is the modal assurance imbalance coefficient of mode n, is the mode loss detection delay coefficient of mode n, is the total number of modalities currently involved in fusion, represent the preset proportional coefficients of the modal confidence imbalance coefficient and the modal missing detection delay coefficient, respectively, and All greater than 0;
[0065] It should be noted that the above formulas are all dimensionless and numerical calculations. Common dimensionless methods include Min-Max normalization and Z-Score normalization, which will not be described here. Set it according to the actual situation. For example, adopt the expert empowerment method, that is, invite experts in related fields to determine the preset proportion coefficients of various indicators through professional opinion surveys and comprehensive evaluations, for example, It can be 0.5, 05;
[0066] From the above calculation expression, it can be seen that the larger the modal confidence imbalance coefficient and the larger the modal loss detection delay coefficient, the larger the fusion hysteresis drift failure coupling index. This indicates that the combined impact of modal signal quality deterioration and modal loss perception hysteresis in the current multimodal perception fusion human-computer interaction intelligent recognition system is more serious, and the risk of the fusion strategy being misled, drifting, or even failing is higher. Conversely, the smaller the modal confidence imbalance coefficient and the smaller the modal loss detection delay coefficient, the smaller the fusion hysteresis drift failure coupling index. This indicates that the contributions of each modality in the current multimodal perception fusion human-computer interaction intelligent recognition system are relatively balanced, the system responds to modal anomalies or loss more promptly and effectively, the overall fusion strategy remains stable and reliable, and the risk of recognition drift and fusion failure is low.
[0067] The fusion hysteresis drift failure coupling index is compared with the preset fusion hysteresis drift failure coupling index threshold to determine the potential drift failure risk of the current fusion strategy, as follows:
[0068] If the fusion hysteresis drift failure coupling index is greater than the fusion hysteresis drift failure coupling index threshold, it means that during the fusion process of the current system, the differences between the modal confidence levels are large and the fluctuations are unstable. At the same time, there is a significant delay in modal loss or anomaly detection. This indicates that the system fusion strategy is facing the coupling problem of "modal dominance imbalance" and "hysteresis response lag", which can easily lead to a series of potential risks such as fusion offset and reduced output stability. The current fusion strategy has the potential risk of drift failure.
[0069] If the fusion hysteresis drift failure coupling index is less than or equal to the fusion hysteresis drift failure coupling index threshold, it indicates that the current modal confidence distribution is generally balanced, the system has a high perception sensitivity and processing speed for modal anomalies and omissions, and the fusion process is not significantly disturbed. In this state, the system fusion strategy is stable and reliable, and the current fusion strategy does not have obvious potential drift failure risks, so the current strategy can be maintained.
[0070] The fusion strategy dynamic warning module is used to provide dynamic warnings for current human-computer interaction operations based on the potential drift failure risk of the fusion strategy. The details are as follows:
[0071] If the current fusion strategy has a potential drift failure risk, obtain the fusion hysteresis drift failure coupling index and combine it with the current human-computer interaction duration to obtain the warning coefficient: ,in is the warning coefficient, The duration of human-computer interaction;
[0072] It should be noted that through The influence of human-computer interaction time is introduced. The longer the interaction time, the more obvious the increase in warning value, reflecting the "accumulation effect of hysteresis drift" in the interaction scenario. This function avoids linear amplification of risk values during interaction time, ensuring that the system response is not overly sensitive.
[0073] Compare the warning coefficient with the preset warning coefficient threshold and issue a dynamic warning for the current human-computer interaction operation, as follows:
[0074] If the warning coefficient is greater than the warning coefficient threshold, it means that the current fusion strategy has obvious hysteresis drift failure risk accumulation, and has reached a level that poses a substantial threat to system stability and interactive reliability, and a dynamic warning signal is generated;
[0075] If the warning coefficient is less than or equal to the warning coefficient threshold, it means that the current fusion strategy does not have obvious hysteresis drift failure risk accumulation, and there is no need to generate a dynamic warning signal;
[0076] The present invention introduces the modal confidence imbalance coefficient and the modal missing detection delay coefficient, and based on this, constructs a fusion hysteresis drift failure model to achieve real-time quantitative assessment of the potential risks of multimodal fusion strategies, thereby effectively improving the stability and accuracy of interactive recognition in complex environments. First, the modal confidence imbalance coefficient is used to dynamically reflect the distribution of the contribution trust between each modality, avoiding the risk of wrong decision-making caused by "abnormal modal misleading". Secondly, by performing real-time missing detection on the modal data link and calculating the modal missing detection delay coefficient, the time interval from the actual modal failure to the system capturing the failure is quantified. The fusion hysteresis drift failure coupling index output after coupling the two is used as a key indicator to measure the potential drift failure risk of the current fusion strategy. When the index exceeds the preset threshold, the fusion strategy dynamic warning module is triggered in time to issue a risk warning for human-computer interaction operations.
[0077] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0078] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0079] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0080] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0081] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A multimodal perception-fusion human-computer interaction intelligent recognition system, characterized by: It includes multimodal data acquisition module, modal state perception module, modal loss detection delay recognition module, modal fusion hysteresis drift failure module, and fusion strategy dynamic warning module; Multimodal data acquisition module, used to obtain raw modal inputs from multiple channels and synchronize them through time stamps; The modal state perception module is used to obtain the modal confidence imbalance information of multimodal data, establish the modal confidence imbalance function, calculate the modal confidence imbalance coefficient, and evaluate the degree of modal degradation imbalance; A modal loss detection delay identification module is used to perform loss detection on the modal data link, obtain the modal loss detection delay information of the modal data link, calculate the modal loss detection delay coefficient, and evaluate the degree of modal loss detection delay; The modal fusion hysteresis drift failure module is used to construct a modal fusion hysteresis drift failure model based on the modal confidence imbalance coefficient and the modal missing detection delay coefficient, output the fusion hysteresis drift failure coupling index, and determine the potential drift failure risk of the current fusion strategy; The fusion strategy dynamic warning module is used to provide dynamic warnings for current human-computer interaction operations based on the potential drift failure risk of the fusion strategy; The logic for obtaining the modal assurance imbalance coefficient is as follows: At each moment, the confidence of all modalities involved in the fusion is obtained to form a confidence vector: ,in is the confidence of the nth mode participating in the fusion at time t, is the total number of modalities currently involved in fusion, , and the normalization conditions should be met before fusion: ; In the time window Inside, construct a three-dimensional tensor ,in is a three-dimensional tensor consisting of elements composition, Represents the confidence value of the f-th feature of the n-th modality participating in the fusion at time t, is the time window length, The confidence level of each modality forms a feature dimension; At each time instant t, a cooperative offset matrix is constructed based on the tensor slices. , the elements in the cooperative offset matrix are the cooperative degrees of the confidence difference between mode n and mode m. The calculation expression of the cooperative degree of the confidence difference between mode n and mode m is as follows: ,in is the confidence difference coordination degree, Represents the confidence value of the fth feature of the mth modality participating in the fusion at time t; Calculate the tensor-form structural norm of all confidence difference synergies within the time window as the deviation coupling metric : ,in , ; In order to enhance the sensitivity of the model to the difference in modal information, the confidence dispersion of the modal assurance criterion is defined as: ,in is the confidence dispersion of the nth mode participating in the fusion at time t, is the confidence ratio probability, ; And calculate the perturbation difference of the confidence dispersion within the time window : ,in is the confidence dispersion of the nth mode participating in the fusion at time t+1; calculate the modal dispersion disturbance factor : ,in is the disturbance difference of the mth mode participating in the fusion, ; Calculate the modal assurance criterion imbalance coefficient: ,in is the modal assurance imbalance coefficient of mode n; The logic for obtaining the modality loss detection delay coefficient is as follows: Constructing a modal propagation diagram ,in is a node set, each mode corresponds to a node, is an edge set, and the information propagation relationship between each modality is through the edge Indicates that the weight of each edge is expressed as the modal information offset strength. The calculation expression is as follows: ,in represents the gradient of the missing probability of mode n at time t, ,in is the missing state probability output of the system for mode n at time t, represents the gradient of the missing probability of mode m at time t; The actual missing moment of the mode is determined based on the entropy mutation and the gradient of the missing probability of mode n at time t. The judgment logic is as follows: Indicates finding the earliest moment on the timeline , which makes the entropy of mode n suddenly change Exceeding the preset mutation threshold And the gradient of the probability of mode n missing at time t Exceeding the preset gradient threshold ;in is the probability distribution entropy of mode n; Modal missing response detection time calculate: Represents the probability output of the missing state of mode n at time t Exceeds the preset missing judgment threshold Or there is a mode m adjacent to mode n, whose modal information offset strength with mode n is Greater than the preset offset strength threshold And the missing state probability output of mode m Exceeds the preset missing judgment threshold ,in is the neighborhood mode set of mode n; Calculate the modal absence detection delay coefficient: ,in is the mode loss detection delay coefficient of mode n, is the maximum delay that the system can tolerate; According to the modal confidence imbalance coefficient and the modal missing detection delay coefficient, a modal fusion hysteresis drift failure model is constructed, and the fusion hysteresis drift failure coupling index is output. The modal fusion hysteresis drift failure model is based on the following formula: , where is the fusion hysteresis drift failure coupling index, is the modal assurance imbalance coefficient of mode n, is the mode loss detection delay coefficient of mode n, is the total number of modalities currently involved in fusion, represent the preset proportional coefficients of the modal confidence imbalance coefficient and the modal missing detection delay coefficient, respectively, and Both are greater than 0.
2. The multimodal perception fusion human-computer interaction intelligent recognition system according to claim 1, characterized in that: The fusion hysteresis drift failure coupling index is compared with the preset fusion hysteresis drift failure coupling index threshold to determine the potential drift failure risk of the current fusion strategy, as follows: If the fusion hysteresis drift failure coupling index is greater than the fusion hysteresis drift failure coupling index threshold, the current fusion strategy has a potential drift failure risk; If the fusion hysteresis drift failure coupling index is less than or equal to the fusion hysteresis drift failure coupling index threshold, the current fusion strategy does not have an obvious potential drift failure risk and the current strategy can be maintained.
3. The multimodal perception fusion human-computer interaction intelligent recognition system according to claim 2, characterized in that: If the current fusion strategy has a potential drift failure risk, obtain the fusion hysteresis drift failure coupling index and combine it with the current human-computer interaction duration to obtain the warning coefficient: ,in is the warning coefficient, The duration of human-computer interaction.
4. The multimodal perception fusion human-computer interaction intelligent recognition system according to claim 3 is characterized by: Compare the warning coefficient with the preset warning coefficient threshold and issue a dynamic warning for the current human-computer interaction operation, as follows: If the warning coefficient is greater than the warning coefficient threshold, a dynamic warning signal is generated; If the warning coefficient is less than or equal to the warning coefficient threshold, there is no need to generate a dynamic warning signal.
Citation Information
Patent Citations
Dynamic multi-modal data fusion and real-time analysis method
CN120046119A