Multi-mode perception fused man-machine interaction intelligent identification system
By constructing a fusion hysteresis drift failure model, using the modal confidence imbalance coefficient and modal missing detection delay coefficient, the problem of modal degradation and missing detection lag in the multimodal fusion system is solved, real-time risk assessment and dynamic early warning of the multimodal fusion strategy are realized, and the stability and accuracy of human-computer interaction recognition are improved.
Patent Information
- Application Number
- CN202510864576.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-26
AI Technical Summary
The existing multimodal fusion method has problems in complex environments where modal degradation is not recognized in time and modal deletion detection lagging responses, resulting in abnormal modal misdog and policy drift strengthening, affecting the accuracy and stability of interaction recognition.
Modal confidence imbalance coefficient and modal missing detection delay coefficient are introduced, and a fusion hysteresis drift failure model is built. The modal confidence imbalance coefficient dynamically reflects the contribution trust of each mode through modal confidence imbalance coefficient, detects modal missing in real time and calculates the modal missing detection delay, and outputs the fusion hysteresis drift failure coupling index for dynamic early warning.
The identification stability and accuracy of multimodal fusion strategy in complex environments is improved, abnormal modal misleading is avoided, early warning is triggered in a timely manner, long-term drift failure is prevented, and the system's reliability and adaptability to modal inputs are enhanced.
Smart Images

Figure CN120354100A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer interaction intelligent recognition, and more specifically, to a human-computer interaction intelligent recognition system for multi-modal perception fusion. Background Art
[0002] With the development of multi-modal perception technology, human-computer interaction systems have gradually transformed from single-channel recognition to intelligent recognition that fuses multi-modal signals such as speech, vision, gesture, touch, and electroencephalogram. Multi-modal perception fusion can improve the interaction robustness and response accuracy in complex environments, and has shown extensive application potential especially in scenarios such as intelligent customer service, virtual assistants, medical assistance, and intelligent driving. Existing multi-modal fusion methods usually rely on the confidence estimation of each modal signal for dynamic weighting, supplemented by a modal state judgment mechanism to achieve fault tolerance processing for missing modalities.
[0003] However, with the expansion of the application scale and the complexity of the usage scenarios, internal defects such as unrecognized modal degradation and lagged response in modal missing detection have emerged during the operation of multi-modal systems, which in turn lead to potential hidden risks such as "abnormal modal misdomination" or "strategy drift reinforcement" in the fusion decision-making process. Especially in actual operation, when the quality of modal data deteriorates or is completely missing due to factors such as occlusion, unstable channels, or device failures, if the system fails to quickly detect and eliminate the modality and instead continues to rely on it for decision-making, it will seriously affect the interaction recognition accuracy. More critically, while the modal missing detection is delayed, if the system's confidence evaluation mechanism fails to timely reduce its fusion weight, the two will form a coupled lag chain, which will not only induce short-term recognition misjudgments, but may also continuously interfere with model learning through misfeedback, resulting in long-term drift of the fusion strategy and forming a systematic risk of fusion lag drift failure. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a human-computer interaction intelligent recognition system for multi-modal perception fusion to solve the problems raised in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions: A human-computer interaction intelligent recognition system for multi-modal perception fusion, including a multi-modal data acquisition module, a modal state perception module, a modal missing detection delay recognition module, a modal fusion lag drift failure module, and a fusion strategy dynamic warning module; The multi-modal data acquisition module is used to obtain the original modal inputs from multiple channels and synchronize and align them through timestamps; The modal state perception module is used to obtain the modal confidence imbalance information of the multi-modal data, establish a modal confidence imbalance function, calculate the modal confidence imbalance coefficient, and evaluate the degradation imbalance degree of the modality; The modal missing detection delay identification module is used to perform missing detection on the modal data link, obtain the modal missing detection delay information of the modal data link, calculate the modal missing detection delay coefficient, and evaluate the delay degree of the modal missing detection; The modal fusion hysteresis drift failure module is used to construct a modal fusion hysteresis drift failure model according to the modal confidence imbalance coefficient and the modal missing detection delay coefficient, output the fusion hysteresis drift failure coupling index, and determine the potential drift failure risk of the current fusion strategy; The fusion strategy dynamic warning module is used to dynamically warn the current human-computer interaction operation according to the potential drift failure risk of the fusion strategy.
[0006] In a preferred embodiment, the acquisition logic of the modal confidence imbalance coefficient is as follows: For each moment, obtain the confidence levels of all the modalities participating in the fusion to form a confidence vector: , where is the confidence level of the nth modality participating in the fusion at time t, is the total number of modalities currently participating in the fusion, , and the normalization condition should be satisfied before fusion: ; Within the time window , construct a three-dimensional tensor , where is a three-dimensional tensor composed of elements , represents the confidence level value of the fth feature of the nth modality participating in the fusion at time t, is the length of the time window, is the feature dimension composed of the confidence levels of each modality; At each moment t, construct a collaborative offset matrix based on tensor slicing. The element in the collaborative offset matrix is the confidence difference collaboration degree between modality n and modality m. The calculation expression of the confidence difference collaboration degree between modality n and modality m is as follows: , where is the confidence difference collaboration degree, represents the confidence level value of the fth feature of the mth modality participating in the fusion at time t; Calculate the tensor form structure norm of all the confidence difference collaboration degrees within the time window as the deviation coupling metric value : , where , ; To enhance the sensitivity of the model to the difference in modal information volume, define the confidence dispersion degree of the modal confidence level: , where is the confidence dispersion of the nth modality participating in the fusion at time t, is the confidence ratio probability, ; and calculate the perturbation difference of the confidence dispersion within the time window : , where is the confidence dispersion of the nth modality participating in the fusion at time t + 1; calculate the modality dispersion perturbation factor : , where is the perturbation difference of the mth modality participating in the fusion, ; Calculate the modality confidence imbalance coefficient: , where is the modality confidence imbalance coefficient of modality n.
[0007] In a preferred embodiment, the acquisition logic of the modality missing detection delay coefficient is as follows: Construct a modality propagation graph , where is the set of nodes, each modality corresponds to a node, is the set of edges, the information propagation relationship between each modality is represented by an edge , and the weight of each edge is represented by the modality information offset intensity. The calculation expression of the modality information offset intensity is as follows: , where represents the gradient of the missing probability of modality n at time t, , where is the output of the missing state probability of modality n by the system at time t, represents the gradient of the missing probability of modality m at time t; Based on the entropy value mutation and the gradient of the missing probability of modality n at time t, judge the actual missing time of the modality. The judgment logic is as follows: represents finding the earliest time on the time axis such that the entropy value mutation of modality n exceeds the preset mutation threshold and the gradient of the missing probability of modality n at time t exceeds the preset gradient threshold ; where is the probability distribution entropy of modality n; Modality missing response detection time Calculate: represents that when the missing state probability output of modality n at time t exceeds the preset missing determination threshold or there exists a modality m adjacent to modality n, and its modality information offset intensity with modality n Greater than the preset offset intensity threshold and the output of the missing state probability of mode m exceeds the preset missing determination threshold , where is the neighborhood mode set of mode n; Calculate the mode missing detection delay coefficient: , where is the mode missing detection delay coefficient of mode n, is the maximum tolerable delay of the system.
[0008] In a preferred embodiment, a mode fusion hysteresis drift failure model is constructed based on the mode confidence imbalance coefficient and the mode missing detection delay coefficient, and the fusion hysteresis drift failure coupling index is output. The formula on which the mode fusion hysteresis drift failure model is based is as follows , where in the formula is the fusion hysteresis drift failure coupling index, is the mode confidence imbalance coefficient of mode n, is the mode missing detection delay coefficient of mode n, is the total number of modes currently participating in the fusion, respectively represent the preset proportionality coefficients of the mode confidence imbalance coefficient and the mode missing detection delay coefficient, and are both greater than 0.
[0009] In a preferred embodiment, the fusion hysteresis drift failure coupling index is compared with the preset fusion hysteresis drift failure coupling index threshold to determine the potential drift failure risk of the current fusion strategy, specifically as follows: If the fusion hysteresis drift failure coupling index is greater than the fusion hysteresis drift failure coupling index threshold, there is a potential drift failure risk in the current fusion strategy; If the fusion hysteresis drift failure coupling index is less than or equal to the fusion hysteresis drift failure coupling index threshold, there is no obvious potential drift failure risk in the current fusion strategy, and the current strategy can be continued.
[0010] In a preferred embodiment, when there is a potential drift failure risk in the current fusion strategy, the fusion hysteresis drift failure coupling index is obtained and the warning coefficient is obtained in combination with the current human-computer interaction duration: , where is the warning coefficient, is the human-computer interaction duration.
[0011] In a preferred embodiment, the warning coefficient is compared with the preset warning coefficient threshold to perform dynamic warning on the current human-computer interaction operation, specifically as follows: If the warning coefficient is greater than the warning coefficient threshold, a dynamic warning signal is generated; If the warning coefficient is less than or equal to the warning coefficient threshold, there is no need to generate a dynamic warning signal.
[0012] Technical effects and advantages of the present invention: 1. By introducing the modal confidence imbalance coefficient and the modal missing detection delay coefficient, and constructing a fusion hysteresis drift failure model based on this, the present invention realizes the real-time quantitative evaluation of the potential risks of the multi-modal fusion strategy, thereby effectively improving the stability and accuracy of interactive recognition in complex environments. First, the modal confidence imbalance coefficient is used to dynamically reflect the distribution of the contribution trust degrees between various modalities, avoiding the risk of incorrect decisions caused by "abnormal modality misdomination"; second, by performing real-time missing detection on the modal data link and calculating the modal missing detection delay coefficient, the time interval required from the real failure of the modality to the system capturing this failure is quantified, and the fusion hysteresis drift failure coupling index output after coupling the two is used as the key index to measure the potential drift failure risk of the current fusion strategy. When the index exceeds the preset threshold, the dynamic warning module of the fusion strategy is triggered in a timely manner to issue a risk prompt for the human-computer interaction operation. Description of the drawings
[0013] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the drawings; Figure 1 It is the flowchart of the system in the embodiment of the present invention. Detailed implementation manners
[0014] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0015] Embodiment: The present invention provides a Figure 1 human-computer interaction intelligent recognition system for multi-modal perception fusion as shown in which includes a multi-modal data acquisition module, a modal state perception module, a modal missing detection delay recognition module, a modal fusion hysteresis drift failure module, and a fusion strategy dynamic warning module; The multi-modal data acquisition module is used to obtain the original modal inputs from multiple channels and synchronize and align them through timestamps; The modal missing detection delay identification module is used to perform missing detection on the modal data link to obtain the modal missing detection delay information of the modal data link, calculate the modal missing detection delay coefficient, and evaluate the delay degree of the modal missing detection; The modal fusion hysteresis drift failure module is used to construct a modal fusion hysteresis drift failure model based on the modal confidence imbalance coefficient and the modal missing detection delay coefficient, output the fusion hysteresis drift failure coupling index, and determine the potential drift failure risk of the current fusion strategy; The fusion strategy dynamic warning module is used to perform dynamic warning on the current human-computer interaction operation according to the potential drift failure risk of the fusion strategy; In the multi-modal data acquisition module, raw modal inputs from multiple channels are obtained by establishing multi-channel input interfaces (supporting heterogeneous modal protocols such as USB video stream, audio I2S, EEG BLE, etc.), modal identifiers are defined for each modality, and modal sampling parameters are initialized, including sampling frequency, communication protocol specifications, data packet header formats, and data frame structures; A unified system time reference (such as UNIX_TIMESTAMP() or the system high-precision clock) is used to assign acquisition timestamps to each modal data frame Denotes the timestamp of the i-th modality collected at the k-th time step, a buffer channel is established for asynchronous modalities, and the time deviation is recorded : , where is the reference modal time; A synchronous calibration mechanism is constructed, and a sliding window is used to maintain the synchronization of the modal sampling window: , where Denotes the timestamp corresponding to the modal channel selected as the synchronization reference in the k-th time step (used to find a modal channel i such that the acquisition timestamp of this modality is closest to the average value of the acquisition timestamps of other modalities, so as to be used as the synchronization reference for the current time step k), is the average value of the acquisition times of all modalities in the k-th time step: , where is the number of modal types; The modal state perception module is used to obtain the modal confidence imbalance information of multi-modal data, establish a modal confidence imbalance function, calculate the modal confidence imbalance coefficient, and evaluate the degradation imbalance degree of the modality; In the present invention, the modal confidence imbalance coefficient is a key metric for measuring whether there is a significant deviation in the confidence distribution among modalities in a multi-modal perception fusion-based human-computer interaction intelligent recognition system. Its core function is to quantify the differences and bias degrees in the actual influence of information carried by different modalities in the decision-making process during multi-modal data fusion. When a certain modality obtains an overly high or low confidence value in the system recognition due to reasons such as data quality, environmental interference, or algorithm mechanisms, the system fusion layer will assign asymmetric fusion weights to it, thereby affecting the directionality of the overall recognition judgment. The modal confidence imbalance coefficient is a quantitative expression of this bias degree. In actual operation, a relatively large modal confidence imbalance coefficient usually means that certain modalities have too strong a dominant position in the system recognition result, which may obscure the effective contributions of other modalities to the target. Especially when the quality of the dominant modality deteriorates due to factors such as occlusion, signal interference, resolution decline, or equipment failure, it is still assigned a high weight by the system, which is extremely likely to cause the problem of "abnormal modality misdomination". Such risks are highly concealed and persistent, not only may lead to short-term recognition misjudgments, but also may strengthen the wrong patterns in long-term feedback training, causing a structural shift in the fusion strategy and resulting in irreversible strategy drift failure. In contrast, a relatively small modal confidence imbalance coefficient indicates that the confidence distributions among modalities tend to be consistent, and the system does not significantly favor a certain modality during the perception fusion process, indicating that the multi-modal perception structure is relatively stable in the current environment, and the fusion judgment mechanism can fully absorb the complementarity of multi-source information. The recognition result in this state is more holistic and robust, which can not only enhance the adaptability to complex environments, but also avoid systematic recognition errors caused by the fluctuations of a single modality. Evaluating the potential drift failure risk of the current fusion strategy based on the modal confidence imbalance coefficient can provide the system with a mechanism for quantitatively analyzing the contribution distribution of modal information, enabling the system to perceive the actual impact of the signal states of each modality on the fusion strategy in real time during operation. Secondly, as an important indicator of the robustness of the fusion strategy, this coefficient can jointly construct a fusion lag drift failure model with the modal missing detection delay coefficient, output a fusion lag drift failure coupling index, and provide a reliable basis for high-order strategy optimization and interactive decision reconstruction. In summary, the modal confidence imbalance coefficient is not only a static analysis tool, but also a core variable for dynamic perception and risk prevention and control of the fusion strategy. It enables the system to evolve from "passive fusion" to "adaptive optimization fusion", and is the key support for ensuring the long-term stable operation of the multi-modal human-computer interaction system in complex and dynamic environments.
[0016] The acquisition logic of the modal confidence imbalance coefficient is as follows: At each moment, obtain the confidence levels of all modalities participating in the fusion to form a confidence vector: , where is the confidence level of the nth modality participating in the fusion at time t, is the total number of modalities currently participating in the fusion, , and should satisfy the normalization condition before fusion: ; It should be noted that the confidence of each modality can be calculated by different platforms according to their actual situations (such as using deep learning models); Within the time window , a three-dimensional tensor ( represents the set of real numbers, that is, all elements of the three-dimensional tensor are real numbers), where is a three-dimensional tensor representing the set of confidences under all modalities, all time frames, and all feature dimensions, composed of elements , represents the confidence value of the nth modality participating in the fusion at the fth feature at time t, is the length of the time window, is the feature dimension composed of the confidences of each modality; At each time t, a collaborative offset matrix is constructed based on tensor slicing. The elements in its collaborative offset matrix are the collaborative degrees of the confidence differences between modality n and modality m. The calculation expression of the collaborative degree of the confidence difference between modality n and modality m is as follows: , where is the collaborative degree of the confidence difference, represents the confidence value of the mth modality participating in the fusion at the fth feature at time t; Calculate the tensor form structure norm of all confidence differences within the time window as the deviation coupling metric value : , where , ; The larger the deviation coupling metric value, the greater the persistent and high-intensity non-equilibrium collaborative deviation between multimodal confidences; To enhance the sensitivity of the model to the differences in modal information content, the confidence dispersion of the modal confidence is defined: , where is the confidence dispersion of the nth modality participating in the fusion at time t, is the confidence ratio probability, ; And calculate the perturbation difference of the confidence dispersion within the time window : , where is the confidence dispersion of the nth modality participating in the fusion at time t + 1; Calculate the modal dispersion perturbation factor : , where is the perturbation difference of the mth modality participating in the fusion, ; Calculate the modal confidence imbalance coefficient: , where is the modal confidence imbalance coefficient of mode n; It should be noted that the above formulas are all dimensionless and take their numerical values for calculation. Common methods for removing dimensions include Min-Max normalization, Z-Score standardization, etc., which will not be elaborated here; The modal missing detection delay identification module is used to perform missing detection on the modal data link to obtain the modal missing detection delay information of the modal data link and calculate the modal missing detection delay coefficient to evaluate the delay degree of the modal missing detection; In the present invention, the modal loss detection delay coefficient is a quantitative index used to measure the time delay degree required for detecting and identifying the missing state of modal data in a multi-modal perception system. The core purpose is to accurately reflect, in the multi-modal fusion decision-making mechanism, the systematic risk source of "abnormal modality not being identified in a timely manner". Specifically, the larger the modal loss detection delay coefficient, the slower the system's response to the loss of a certain modality, and there is a possibility that the failed modality will still participate in the fusion calculation for a long period of time, thereby increasing the risk that the system is dominated by incorrect modalities in the fusion strategy; while the smaller the modal loss detection delay coefficient, the faster the system's detection and response ability to abnormal modal changes, and it can complete fusion control operations such as modal elimination and confidence weight adjustment in a timely manner, which helps to ensure the stability and accuracy of the fusion strategy. In the fusion scenario, modal loss may be caused by various factors such as sensor occlusion, channel interruption, hardware failure, accidental shielding, etc. If the system fails to detect the interruption or distortion (i.e., loss) of the modal data stream in a timely manner and still continues to treat it as normal and participate in the fusion, this will result in a "false contribution" of this modality in the decision output. Especially when its initial confidence is relatively high or the proportion is relatively large, it is very likely to induce misidentification. This kind of misidentification not only brings short-term output errors, but also triggers an error feedback chain in the interactive human-machine system, causing cumulative deviation in long-term strategy learning. The introduction of the modal loss detection delay coefficient is precisely to provide a structural mitigation mechanism for this problem. By defining the modal loss detection delay coefficient, the present invention can construct a quantitative model of the delay response mechanism within the multi-modal fusion system. This coefficient calculates the time difference between the actual occurrence time of modal loss and the time when the system completes the confirmation of the loss state, and is corrected in combination with the temporal continuity and false triggering of modal data, and finally outputs a stable and usable evaluation value. Further, the modal loss detection delay coefficient can be coupled with the modal confidence imbalance coefficient for modeling, and jointly input into the fusion hysteresis drift failure model to calculate the fusion hysteresis drift failure coupling index, so as to identify the potential systematic risk trend in the current fusion strategy at the global level. The beneficial effect of this method is that it can not only timely capture and respond to the functional attenuation process of local modalities, but also, through the analysis of the coefficient evolution trend, predict and warn in advance the drift direction of the fusion mechanism in the evolution path, realize the prediction and early warning of the deviation of the fusion strategy, strengthen the temporal sensitivity of the system to the change of the reliability of modal input, reduce the misfusion and identification deviation caused by detection lag, improve the system's intervention ability for abnormal perception chains, and provide a solid foundation for subsequent fusion strategy optimization, active modal screening and modal switching control.
[0017] The acquisition logic of the modal loss detection delay coefficient is as follows: Construct a modal propagation graph , where is a set of nodes, and each modality corresponds to a node, The information propagation relationship between each modality in the edge set is represented by an edge Each edge's weight is represented by the modality information offset intensity, and the modality information offset intensity is calculated as follows: , where represents the gradient of the missing probability of modality n at time t, , where is the output of the missing state probability of modality n by the system at time t, represents the gradient of the missing probability of modality m at time t; Based on the entropy value mutation and the gradient of the missing probability of modality n at time t, judge the actual missing time of the modality. The judgment logic is as follows: represents finding the earliest time on the time axis such that the entropy value mutation of modality n exceeds the preset mutation threshold and the gradient of the missing probability of modality n at time t exceeds the preset gradient threshold ; where is the probability distribution entropy of modality n; The missing response detection time of the modality is calculated as: represents that when the missing state probability output of modality n at time t exceeds the preset missing judgment threshold or there exists a modality m adjacent to modality n, and its modality information offset intensity with modality n is greater than the preset offset intensity threshold and the missing state probability output of modality m exceeds the preset missing judgment threshold , where is the neighborhood modality set of modality n; Calculate the modality missing detection delay coefficient: , where is the modality missing detection delay coefficient of modality n, is the maximum tolerable delay of the system; It should be noted that the above formulas are all calculated by taking the numerical values after dimensionless. Common methods for removing dimensions include Min-Max normalization, Z-Score standardization, etc., which will not be elaborated here; The modality fusion hysteresis drift failure module is used to construct a modality fusion hysteresis drift failure model based on the modality confidence imbalance coefficient and the modality missing detection delay coefficient, output the fusion hysteresis drift failure coupling index, and determine the potential drift failure risk of the current fusion strategy; Construct a modal fusion hysteresis drift failure model based on the modal confidence imbalance coefficient and the modal missing detection delay coefficient, and output the fusion hysteresis drift failure coupling index. The formula on which the modal fusion hysteresis drift failure model is based is as follows , where is the fusion hysteresis drift failure coupling index, is the modal confidence imbalance coefficient of mode n, is the modal missing detection delay coefficient of mode n, is the total number of modes currently participating in the fusion, respectively represent the preset proportionality coefficients of the modal confidence imbalance coefficient and the modal missing detection delay coefficient, and are all greater than 0; It should be noted that the above formulas are all dimensionless and take their numerical values for calculation. Common methods for removing dimensions include Min-Max normalization, Z-Score standardization, etc., which will not be elaborated here; is set according to the actual situation. For example, the expert weighting method is adopted, that is, experts in related fields are invited to determine the preset proportionality coefficients of various indicators through professional opinion surveys and comprehensive evaluations. For example, can be 0.5, 05; It can be seen from the above calculation expressions that the larger the modal confidence imbalance coefficient and the larger the modal missing detection delay coefficient, the larger the fusion hysteresis drift failure coupling index, indicating that in the current multi-modal perception fusion human-computer interaction intelligent recognition system, the combined effects caused by the deterioration of modal signal quality and the hysteresis of modal missing perception are more serious, and the risk of the fusion strategy being misled, drifting, and even failing is higher. On the contrary, the smaller the modal confidence imbalance coefficient and the smaller the modal missing detection delay coefficient, the smaller the fusion hysteresis drift failure coupling index, indicating that the contributions between various modes in the current multi-modal perception fusion human-computer interaction intelligent recognition system are relatively balanced, the system responds more timely and effectively to modal abnormalities or missing, the overall fusion strategy remains stable and reliable, and the risk of recognition drift and fusion failure is lower; Compare the fusion hysteresis drift failure coupling index with the preset fusion hysteresis drift failure coupling index threshold to determine the potential drift failure risk of the current fusion strategy, as follows: If the fusion hysteresis drift failure coupling index is greater than the fusion hysteresis drift failure coupling index threshold, it means that during the fusion process of the current system, the difference between modal confidences is large and fluctuates unstably, and at the same time, there is an obvious delay in modal missing or abnormality detection, indicating that the system fusion strategy has faced the coupling problem of "modal dominance imbalance" and "hysteresis response lag", which is likely to cause a series of potential risks such as fusion deviation and output stability decline. The current fusion strategy has a potential drift failure risk; If the fusion hysteresis drift failure coupling index is less than or equal to the fusion hysteresis drift failure coupling index threshold, it indicates that the overall confidence distribution of each current mode is balanced, the system has high perception sensitivity and processing speed for mode anomalies and deficiencies, and the fusion process is not significantly interfered. In this state, the system fusion strategy runs stably and reliably, there is no obvious potential drift failure risk in the current fusion strategy, and the current strategy can be continued; The fusion strategy dynamic warning module is used to dynamically warn the current human-computer interaction operation according to the potential drift failure risk of the fusion strategy, specifically as follows: When there is a potential drift failure risk in the current fusion strategy, obtain the fusion hysteresis drift failure coupling index and combine it with the current human-computer interaction duration to obtain the warning coefficient: , where is the warning coefficient, is the human-computer interaction duration; It should be noted that by introducing the influence of the human-computer interaction duration, the longer the interaction time, the more obvious the growth of the warning value, reflecting the "accumulation effect of hysteresis drift" in the interaction scenario. The function is used to avoid the risk of linearly amplifying the risk value by the interaction time and ensure that the system response is not overly sensitive; Compare the warning coefficient with the preset warning coefficient threshold to dynamically warn the current human-computer interaction operation, specifically as follows: If the warning coefficient is greater than the warning coefficient threshold, it indicates that there is an obvious accumulation of hysteresis drift failure risk in the current fusion strategy and it has reached the level of posing a substantial threat to the system stability and interaction reliability, and a dynamic warning signal is generated; If the warning coefficient is less than or equal to the warning coefficient threshold, it indicates that there is no obvious accumulation of hysteresis drift failure risk in the current fusion strategy and there is no need to generate a dynamic warning signal; In the present invention, by introducing the mode confidence imbalance coefficient and the mode missing detection delay coefficient, and based on this constructing a fusion hysteresis drift failure model, the real-time quantitative evaluation of the potential risks of the multi-modal fusion strategy is realized, so as to effectively improve the stability and accuracy of interaction recognition in a complex environment. First, the mode confidence imbalance coefficient is used to dynamically reflect the distribution of the contribution trust degrees among each mode, avoiding the risk of wrong decisions brought by "abnormal mode misdomination"; secondly, by performing real-time missing detection on the mode data link and calculating the mode missing detection delay coefficient, quantifying the time interval required from the real failure of the mode to the system capturing this failure, and using the fusion hysteresis drift failure coupling index output after coupling the two as the key index to measure the potential drift failure risk of the current fusion strategy, when the index exceeds the preset threshold, the fusion strategy dynamic warning module is triggered in time to issue a risk prompt for the human-computer interaction operation.
[0018] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0019] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0020] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0021] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc., which can store program codes.
[0022] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims described.
Claims
1. A human-computer interaction intelligent recognition system for multimodal perception fusion, characterized in that: It includes a multi-modal data acquisition module, a modal state perception module, a modal missing detection delay identification module, a modal fusion hysteresis drift failure module, and a fusion strategy dynamic warning module; The multi-modal data acquisition module is used to obtain the original modal inputs from multiple channels and synchronize and align them through timestamps; The modal state perception module is used to obtain the modal confidence imbalance information of the multi-modal data, establish a modal confidence imbalance function, calculate the modal confidence imbalance coefficient, and evaluate the degradation and imbalance degree of the modality; The modal missing detection delay identification module is used to perform missing detection on the modal data link, obtain the modal missing detection delay information of the modal data link, calculate the modal missing detection delay coefficient, and evaluate the delay degree of the modal missing detection; The modal fusion hysteresis drift failure module is used to construct a modal fusion hysteresis drift failure model based on the modal confidence imbalance coefficient and the modal missing detection delay coefficient, output the fusion hysteresis drift failure coupling index, and determine the potential drift failure risk of the current fusion strategy; The fusion strategy dynamic warning module is used to perform dynamic warning on the current human-computer interaction operation according to the potential drift failure risk of the fusion strategy.
2. The multi-modal perception fusion human-computer interaction intelligent recognition system according to claim 1, characterized in that: The acquisition logic of the modal confidence imbalance coefficient is as follows: At each moment, obtain the confidence levels of all modalities participating in the fusion to form a confidence vector: , where is the confidence level of the nth modality participating in the fusion at time t, is the total number of modalities currently participating in the fusion, , and the normalization condition should be satisfied before fusion: ; Within the time window construct a three-dimensional tensor where is a three-dimensional tensor composed of elements where represents the confidence value of the f-th feature of the n-th modality participating in the fusion at time t, is the length of the time window, is the feature dimension composed of the confidence of each modality; At each moment t, a collaborative offset matrix is constructed based on tensor slices , and the elements in the collaborative offset matrix are the collaborative degrees of the confidence differences between mode n and mode m. The calculation expression of the collaborative degree of the confidence difference between mode n and mode m is as follows: , where is the collaborative degree of the confidence difference, represents the confidence value of the f-th feature of the m-th mode participating in the fusion at moment t; Calculate the tensor form structural norm of the synergy of all confidence differences within the calculation time window as the deviation coupling metric value : , where , ; To enhance the model's sensitivity to the differences in modal information volume, the confidence dispersion of modal confidence is defined as: , where is the confidence dispersion of the nth modality participating in the fusion at time t, is the probability of confidence ratio, ; And calculate the perturbation difference of the confidence dispersion within the time window : , where is the confidence dispersion of the nth modality participating in the fusion at time t + 1; Calculate the modality dispersion perturbation factor : , where is the perturbation difference of the mth modality participating in the fusion, ; Calculate the modal confidence imbalance coefficient: , where is the modal confidence imbalance coefficient of mode n.
3. The multi-modal perception fusion human-computer interaction intelligent recognition system according to claim 1, characterized in that: The acquisition logic of the modal missing detection delay coefficient is as follows: Construct a modal propagation graph , where is a set of nodes, each mode corresponds to a node, is a set of edges. The information propagation relationship between each pair of modes is represented by an edge . The weight of each edge is represented by the modal information offset intensity. The calculation expression of the modal information offset intensity is as follows: , where represents the gradient of the missing probability of mode n at time t, , where is the output of the missing state probability of mode n by the system at time t, represents the gradient of the missing probability of mode m at time t; Judging the actual missing moment of the mode based on the entropy mutation and the gradient of the missing probability of mode n at time t, the judgment logic is as follows: Denote finding the earliest moment on the time axis such that the entropy of mode n mutates exceeding the preset mutation threshold and the gradient of the missing probability of mode n at time t exceeds the preset gradient threshold ; where is the probability distribution entropy of mode n; Modal missing response detection moment Calculate: Indicates the missing state probability output of mode n at time t Exceeds the preset missing determination threshold Or there exists a mode m adjacent to mode n, and the modal information offset intensity between it and mode n Is greater than the preset offset intensity threshold And the missing state probability output of mode m Exceeds the preset missing determination threshold , where Is the neighborhood mode set of mode n; Calculate the modal missing detection delay coefficient: , where is the modal missing detection delay coefficient of mode n, is the maximum tolerable delay of the system.
4. The multimodal perception fusion-based human-computer interaction intelligent recognition system according to claim 1, wherein: Construct a modal fusion hysteresis drift failure model based on the modal confidence imbalance coefficient and the modal missing detection delay coefficient, and output the fusion hysteresis drift failure coupling index. The formula on which the modal fusion hysteresis drift failure model is based is as follows , where is the fusion hysteresis drift failure coupling index, is the modal confidence imbalance coefficient of mode n, is the modal missing detection delay coefficient of mode n, is the total number of modes currently participating in the fusion, respectively represent the preset proportionality coefficients of the modal confidence imbalance coefficient and the modal missing detection delay coefficient, and are both greater than 0.
5. The multi-modal perception fusion human-computer interaction intelligent recognition system according to claim 4, characterized in that: Compare the fusion hysteresis drift failure coupling index with a preset fusion hysteresis drift failure coupling index threshold to determine the potential drift failure risk of the current fusion strategy, specifically as follows: If the fusion hysteresis drift failure coupling index is greater than the fusion hysteresis drift failure coupling index threshold, there is a potential drift failure risk for the current fusion strategy; If the fusion hysteresis drift failure coupling index is less than or equal to the fusion hysteresis drift failure coupling index threshold, there is no obvious potential drift failure risk for the current fusion strategy, and the current strategy can be continued.
6. The multi-modal perception fusion human-computer interaction intelligent recognition system according to claim 5, characterized in that: When there is a potential risk of drift failure in the current fusion strategy, obtain the fusion lag drift failure coupling index and combine it with the current human-computer interaction duration to obtain the warning coefficient: , where is the warning coefficient, is the human-computer interaction duration.
7. The multi-modal perception fusion human-computer interaction intelligent recognition system according to claim 6, wherein: Compare the warning coefficient with a preset warning coefficient threshold to perform dynamic warning on the current human-computer interaction operation, specifically as follows: If the warning coefficient is greater than the warning coefficient threshold, generate a dynamic warning signal; If the warning coefficient is less than or equal to the warning coefficient threshold, there is no need to generate a dynamic warning signal.
Citation Information
Patent Citations
Confidence coefficient calculation method for multi-channel interaction system channel
CN112711392A
Dynamic multi-modal data fusion and real-time analysis method
CN120046119A
Intelligent agent construction method and system based on agentive workflow
CN120144577A
Weighted deep fusion architecture
US20220019867A1
Cited By
Man-machine interaction method based on multi-modal data
CN120561522A
Large model tuning method and system based on multi-modal information and AI
CN120910811A
Ecological bearing capacity overload risk real-time monitoring method based on big data analysis
CN122334985A