Man-machine interaction method based on multi-modal data

By introducing multi-dimensional performance feature calculation indicators and optimal performance compensation values, the existing performance evaluation methods are corrected, and the problem of modal imbalance in human-computer interaction is solved, more efficient multimodal data fusion and dynamic gradient modulation are achieved, and interaction accuracy and robustness are improved.

CN120561522AActive Publication Date: 2025-08-29SHANDONG DOLANG TECH EQUIP
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511061708.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-08-29
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

In the human-computer interaction scenario, there is a problem of modal imbalance in multimodal data fusion, resulting in poor interaction accuracy and robustness, and the existing dynamic gradient modulation methods cannot effectively solve this problem.

Method used

By introducing multi-dimensional performance feature calculation indicators, including data quality performance evaluation values ​​and real-time response performance evaluation values, we can build the optimal performance compensation values, correct the existing performance evaluation methods, realize dynamic gradient weighting and optimization, and solve the problem of modal imbalance.

Benefits of technology

It improves the accuracy and robustness of human-computer interaction, effectively solves the defects of modal imbalance, and improves interaction performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561522A_ABST
    Figure CN120561522A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a man-machine interaction method based on multi-modal data, and the method comprises the steps: obtaining multi-modal interaction data in real time; for any modal interaction data, obtaining a data quality performance evaluation value and a real-time response performance evaluation value of any modal interaction data, and obtaining an optimal performance compensation value of any modal interaction data; using the optimal performance compensation value to correct the traditionally calculated performance value to obtain a final performance value of any modal interaction data; obtaining a final performance value of each piece of modal interaction data, and obtaining a weight of each piece of modal interaction data; and according to the weight of each piece of modal interaction data, obtaining a weighted gradient of each piece of modal interaction data, and according to the weighted gradient of each piece of modal interaction data, carrying out gradient fusion and model updating to obtain an updated language model for outputting response information of the multi-modal interaction data, and the defect of modal imbalance in man-machine interaction is effectively solved through performance compensation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a human-computer interaction method based on multimodal data. Background Art

[0002] Multimodal human-computer interaction (MMHCI) integrates multiple sensory and interaction channels (such as vision, hearing, touch, gesture, and voice) to make human-computer interaction more natural and efficient. Processing multimodal data is a key component of achieving natural and efficient interaction. The core goal of this process is to extract complementary information and eliminate ambiguity by fusing data from different modalities (such as voice and vision), thereby improving the accuracy and robustness of the interaction.

[0003] In human-computer interaction scenarios, data from different modalities is affected by differences in acquisition cost, data quality, and learning speed. This can cause models in the human-computer interaction process to overly rely on data from one modality while ignoring data from other modalities. This is known as modal imbalance, which can severely impact interaction performance, user experience, and resource waste. Existing technologies address this issue in multimodal data fusion by combining data and model layers. At the data level, dynamic gradient modulation (DGM) is used to dynamically adjust gradient weights based on modal data performance to address modal imbalance. At the model level, attention mechanisms are used to measure modal importance, and gating mechanisms are used to dynamically control modal contributions.

[0004] Although existing technologies use a control and processing method that combines the data level and the model level to solve the modal imbalance problem in multimodal data fusion and analysis scenarios with good results and have been effectively verified in actual operations, they have serious defects and problems in human-computer interaction scenarios. The defects are mainly manifested in data-level processing. The reason is that the data types in human-computer interaction scenarios are mainly based on sensory acquisition data, namely auditory speech data and visual image data. The intensity of noise interference in the acquisition process of this type of data is theoretically higher (user behavior and environment); at the same time, this type of data is subject to response delays, resulting in obvious interference problems such as timeline asynchrony. When calculating and evaluating the performance indicators of multimodal data, DGM usually only relies on machine learning parameters such as accuracy, F1 score, and recall rate to construct performance vectors. It cannot effectively realize accurate performance evaluation and dynamic gradient modulation of multimodal data in human-computer interaction scenarios, resulting in poor accuracy and robustness of the interaction. Summary of the Invention

[0005] In view of this, an embodiment of the present invention provides a human-computer interaction method based on multimodal data to solve the problem that performance vectors are usually constructed only by relying on machine learning parameters such as accuracy, F1 score, and recall rate, and cannot effectively realize accurate performance evaluation and dynamic gradient modulation of multimodal data in human-computer interaction scenarios.

[0006] An embodiment of the present invention provides a human-computer interaction method based on multimodal data, the method comprising the following steps: Acquiring multimodal interaction data in real time, wherein the multimodal interaction data includes visual data, auditory data, and tactile data; For any modal interaction data, perform effective data and noise interference analysis on the any modal interaction data to obtain a data quality performance evaluation value of the any modal interaction data; obtain a real-time response performance evaluation value of the any modal interaction data based on the data volume and data response delay time contained in the any modal interaction data; and obtain an optimal performance compensation value for the any modal interaction data by combining the data quality performance evaluation value and the real-time response performance evaluation value; Obtaining a performance value of any modal interaction data calculated using dynamic gradient modulation, and correcting the performance value using the optimal performance compensation value to obtain a final performance value of any modal interaction data; obtaining the final performance value of each modal interaction data, and obtaining a weight of each modal interaction data; According to the weight of each modal interaction data, the weighted gradient of each modal interaction data is obtained, and gradient fusion and model update are performed according to the weighted gradient of each modal interaction data to obtain an updated language model, which is used to output response information of multimodal interaction data.

[0007] Preferably, performing effective data and noise interference analysis on the interaction data of any modality to obtain a data quality performance evaluation value of the interaction data of any modality includes: Obtain the ratio of the number of valid data to the total number of data in the interaction data of any modality, normalize the ratio to obtain a data validity index, calculate the signal-to-noise ratio of the interaction data of any modality, normalize the signal-to-noise ratio to obtain a data reliability index, and obtain a data quality performance evaluation value of the interaction data of any modality based on the average of the data validity index and the data reliability index.

[0008] Preferably, obtaining the real-time response performance evaluation value of the any modal interaction data according to the data volume and data response delay time contained in the any modal interaction data includes: Obtain the proportion of the data volume contained in the any modal interaction data in the total data volume of all modal interaction data, normalize the proportion to obtain a data volume evaluation value, obtain the data response delay time of the any modal interaction data according to the time length from input to output of the any modal interaction data, use an exponential function with a natural constant as the base to inversely normalize the data response delay time to obtain an interaction real-time index; perform weighted processing on the data volume evaluation value and the interaction real-time index to obtain a real-time response performance evaluation value of the any modal interaction data.

[0009] Preferably, the combining of the data quality performance evaluation value and the real-time response performance evaluation value to obtain the optimal performance compensation value of any modal interaction data includes: The average of the data quality performance evaluation value and the real-time response performance evaluation value is used as the optimal performance compensation value of any modal interaction data.

[0010] Preferably, the using the optimal performance compensation value to correct the performance value to obtain the final performance value of any modal interaction data includes: The performance value is normalized to obtain a normalized value, and the average of the optimal performance compensation value and the normalized value is calculated as the final performance value of the interaction data of any modality.

[0011] Preferably, the method for obtaining the weight of each modal interaction data includes: For any modal interaction data, the final performance value of the any modal interaction data is used as the independent variable of the softmax function to obtain the weight of the any modal interaction data.

[0012] Preferably, obtaining the weighted gradient of each modal interaction data according to the weight of each modal interaction data includes: For any modal interaction data, obtain the gradient of the any modal interaction data, calculate the sum of the weight of the any modal interaction data and the preset smoothing coefficient, and take the product of the inverse of the sum and the gradient as the weighted gradient of the any modal interaction data.

[0013] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: The present invention performs data analysis through multimodal data and introduces multi-dimensional performance feature calculation indicators, that is, the optimal performance compensation value obtained by the data quality performance evaluation value and the real-time response performance evaluation value. The optimal performance compensation value is then used to correct the optimal performance formula (performance value) constructed by the accuracy, F1 score, and recall rate under the existing method to obtain the final performance value. According to the final performance value, dynamic gradient weighted adjustment and optimization are implemented, effectively solving the modal imbalance defect in human-computer interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0015] Figure 1 This is a method flow chart of a human-computer interaction method based on multimodal data provided in Example 1 of the present invention. DETAILED DESCRIPTION

[0016] The embodiments of the present disclosure are described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to be used to explain the present disclosure, but should not be understood as limiting the present disclosure.

[0017] It should be noted that the terms "first," "second," and the like in the specification of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure.

[0018] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.

[0019] See also Figure 1 , is a method flow chart of a human-computer interaction method based on multimodal data provided by the first embodiment of the present invention, such as Figure 1 As shown, the method may include: Step S101 : acquiring multimodal interaction data in real time, where the multimodal interaction data includes visual data, auditory data, and tactile data.

[0020] Human-computer interaction is the process by which a terminal acquires data such as video, images, and / or audio from the outside world, as well as interactive information from the user, analyzes and processes the aforementioned data, and ultimately outputs information to the user. For example, consider the process of human-computer interaction between a home surveillance camera and a user: the home surveillance camera monitors the user's home by collecting video and audio content, and is able to receive and process the user's interactive information accordingly. Another example is the process of human-computer interaction between a smart car with autonomous driving capabilities and a user: the autonomous driving function identifies and analyzes roads, vehicles, pedestrians, etc. by collecting video and audio content, and is able to receive and process the user's interactive information accordingly.

[0021] The general steps of human-computer interaction are: 1. Data collection and preprocessing; 2. Multimodal data fusion; 3. Interaction decision-making: mapping data to intent; 4. Feedback generation: collaborative expression of multimodal outputs; 5. Optimization and iteration: continuous improvement of interaction performance. However, in the aforementioned scenarios of human-computer interaction using multimodal data, the types and sources of multimodal data are primarily sensory data, and the data volumes vary significantly across different dimensions. (For example, in voice interaction, users generate hundreds of hours of voice data daily, but the corresponding lip-reading videos may only be a few minutes.) Furthermore, the acquisition of common voice and visual data in these scenarios is theoretically more susceptible to noise interference (due to user behavior and the environment). Furthermore, such data is significantly affected by timeline desynchronization caused by response delays. For example, during in-car interaction, the voice modality may be mixed with wind and tire noise at high speeds, while the visual modality (such as driver expression recognition) may lose features due to sunglasses or strong sunlight, both of which can lead to modal imbalance. However, the existing DGM modulation relies on machine learning parameters to construct performance vectors. This processing method cannot effectively perform accurate performance evaluation and dynamic gradient modulation of multimodal data in human-computer interaction scenarios.

[0022] Therefore, the embodiment of the present invention introduces performance compensation based on scene data features during the model training process in the multimodal data fusion stage to accurately implement dynamic gradient adjustment (DGM), effectively solve the defects and problems of modal imbalance in the human-computer interaction process, and greatly improve the accuracy and robustness of the interactive behavior.

[0023] Specifically, electronic devices are first used to collect multimodal data in real time: visual data is collected using cameras (RGB / depth / infrared), 3D sensors (such as LiDAR), and eye trackers; auditory data is collected using microphone arrays (supporting noise reduction and sound source localization); and tactile data is collected using IMUs (accelerometers, gyroscopes), pressure sensors, and force feedback gloves. The collected multimodal data is then preprocessed to obtain multimodal interaction data. Preprocessing includes: (1) Noise reduction and enhancement processing Speech: Beamforming suppresses background noise, and speech enhancement algorithms (such as Wiener filtering) improve the signal-to-noise ratio. Image: Deblurring (such as DeblurGAN), super-resolution reconstruction (such as ESRGAN), and illumination normalization (such as Histogram Equalization). Haptics: Sliding average filtering eliminates sensor jitter noise.

[0024] (2) Data standardization Unify data from different modalities to the same scale (e.g. normalize to the range of [0, 1]) and try to avoid a certain modality dominating the fusion result due to its large numerical range in the early stage.

[0025] At this point, multimodal interaction data is obtained in real time, which is used to subsequently construct a performance compensation model to address the defects and problems of modal imbalance in the human-computer interaction process.

[0026] Step S102: For any modal interaction data, perform effective data and noise interference analysis on any modal interaction data to obtain a data quality performance evaluation value of any modal interaction data; obtain a real-time response performance evaluation value of any modal interaction data based on the amount of data contained in any modal interaction data and the data response delay time; and obtain the optimal performance compensation value of any modal interaction data by combining the data quality performance evaluation value and the real-time response performance evaluation value.

[0027] To improve the accuracy of modal imbalance adjustment using dynamic gradient modulation (DGM) under existing technologies, this embodiment of the present invention introduces multi-dimensional performance analysis. By measuring the quality of each modal interaction data during the interaction process, a performance compensation model is constructed. Therefore, for any modal interaction data, valid data and noise interference analysis are performed on any modal interaction data to obtain the data quality performance evaluation value of any modal interaction data: Obtain the ratio of the number of valid data to the total number of data in the interaction data of any modality, normalize the ratio to obtain a data validity index, calculate the signal-to-noise ratio of the interaction data of any modality, normalize the signal-to-noise ratio to obtain a data reliability index, and obtain a data quality performance evaluation value of the interaction data of any modality based on the average of the data validity index and the data reliability index.

[0028] In one embodiment, the calculation formula for the data quality performance evaluation value of any modal interaction data is:

[0029] in, Represents the data quality performance evaluation value of any modal interaction data, represents the normalization function, N represents the number of valid data in any modal interaction data, M represents the total number of data in any modal interaction data, Represents the signal-to-noise ratio of the interaction data of any modality.

[0030] It should be noted that the specific determination basis for the number of valid data N and the total number of data M in different modal interaction data is different. For example: in voice interaction, if the valid voice frame is 8 seconds in a 10-second conversation (excluding silence and noise), then is 0.8; in image interaction, The judgment can be based on the number of non-completely black frames / completely white frames (pixel value standard deviation>10); in video interaction, It is expressed as the percentage of optical flow changes between consecutive frames that are greater than the threshold. The larger the value of , the higher the data integrity of any modal interaction data, and the higher the corresponding data quality; It indicates the noise level in any modal interaction data and reflects the reliability of the data. The higher the signal-to-noise ratio, the less interference from noise and the better the data quality. On the contrary, the lower the signal-to-noise ratio, the greater the interference from noise and the worse the data quality. Therefore, The larger the value of , the larger the corresponding data quality performance evaluation value E. Norm is a normalization function that maps the result to the value range [0, 1].

[0031] Furthermore, in human-computer interaction scenarios, the types and sources of multimodal data are basically sensory data, and the amount of data in different modalities varies greatly (for example, in voice interaction, users generate hundreds of hours of voice data every day, but the corresponding lip reading video may only be a few minutes long). This phenomenon will also cause the problem of modal imbalance: some modalities (such as text) are data-rich, while other modalities (such as sensor signals) are data-scarce, causing the language model to over-rely on modalities with sufficient data and ignore modalities with insufficient data. At the same time, data in different modalities all have more or less response delay problems. For example, in intelligent assistants, the response delay of the voice modality is 300ms (including ASR and TTS time), and the visual modality (such as expression recognition) is 150ms. Therefore, according to the amount of data contained in any modal interaction data and the data response delay time, the real-time response performance evaluation value of any modal interaction data is obtained. The specific acquisition method is: Obtain the proportion of the data volume contained in the any modal interaction data in the total data volume of all modal interaction data, normalize the proportion to obtain a data volume evaluation value, obtain the data response delay time of the any modal interaction data according to the time length from input to output of the any modal interaction data, use an exponential function with a natural constant as the base to inversely normalize the data response delay time to obtain an interaction real-time index; perform weighted processing on the data volume evaluation value and the interaction real-time index to obtain a real-time response performance evaluation value of the any modal interaction data.

[0032] In one embodiment, the calculation formula for the real-time response performance evaluation value of any modal interaction data is:

[0033] in, represents the real-time response performance evaluation value of any modal interaction data, represents the first weight, represents the normalization function, Indicates the amount of data contained in any modal interaction data, Indicates the total amount of data for all modal interaction data, represents the second weight, represents an exponential function with a natural constant as base, Indicates the data response delay time of any modal interaction data.

[0034] It should be noted that The larger the value of is, the richer the data of any modal interaction data is. , used to measure the delay time from input to output feedback of any modal interaction data, The smaller the value, the smaller the delay, and the higher the corresponding real-time interaction performance, which is the real-time response performance evaluation value of any modal interaction data. It can be set according to specific needs. In this embodiment, it is believed that the modal response delay problem has a more serious impact on the interactive processing of the language model. A high degree of delay will directly lead to the model overfitting historical data, so it can be made , the reference value is: .

[0035] At this point, the data quality performance evaluation value and real-time response performance evaluation value of any modal interaction data are obtained from two aspects, and then the optimal performance compensation value of any modal interaction data is obtained by combining the data quality performance evaluation value and the real-time response performance evaluation value: the average between the data quality performance evaluation value and the real-time response performance evaluation value is taken as the optimal performance compensation value of any modal interaction data.

[0036] In one embodiment, the calculation formula for the optimal performance compensation value of any modal interaction data is:

[0037] in, represents the optimal performance compensation value of any modal interaction data, Represents the data quality performance evaluation value of any modal interaction data, Represents the real-time response performance evaluation value of any modal interaction data.

[0038] Step S103: obtain the performance value of any modal interaction data calculated using dynamic gradient modulation, correct the performance value using the optimal performance compensation value, and obtain the final performance value of any modal interaction data; obtain the final performance value of each modal interaction data, and obtain the weight of each modal interaction data.

[0039] Based on step S102, the optimal performance compensation value of any modal interaction data can be obtained, and then the performance evaluation results constructed by the existing method based on the machine learning level such as accuracy, F1 score, and recall rate are corrected, thereby realizing the adjustment and optimization of dynamic gradient weighting. First, the performance calculation result (performance value F) of any modal interaction data using DGM in the traditional calculation method is obtained, and then the performance value is corrected using the optimal performance compensation value of any modal interaction data to obtain the final performance value of the any modal interaction data: the performance value is normalized to obtain a normalized value, and the average of the optimal performance compensation value and the normalized value is calculated as the final performance value of the any modal interaction data.

[0040] In one embodiment, the calculation formula for the final performance value of any modal interaction data is:

[0041] in, Represents the final performance value of any modal interaction data, represents the optimal performance compensation value of any modal interaction data, represents the normalization function, and F represents the performance value of any modal interaction data.

[0042] Based on the above method for obtaining the final performance value of any modal interaction data, the final performance value of each modal interaction data is obtained respectively, thereby forming a performance vector , where j represents the number of modal interaction data types. Then, through the normalization method, the final performance value of each modal interaction data is used to obtain the weight of each modal interaction data: for any modal interaction data, the final performance value of any modal interaction data is used as the independent variable of the softmax function to obtain the weight of any modal interaction data. That is, the weight of the jth modal interaction data , the worse the performance, the smaller the weight of the corresponding modal interaction data.

[0043] In step S104, based on the weight of each modal interaction data, the weighted gradient of each modal interaction data is obtained, and gradient fusion and model update are performed based on the weighted gradient of each modal interaction data to obtain an updated language model for outputting response information of the multimodal interaction data.

[0044] After determining the weight of each modal interaction data, dynamic gradient weighting adjustment and optimization can be performed. Specifically: for any modal interaction data, the gradient of any modal interaction data is obtained. , calculate the sum of the weight of the any modal interaction data and a preset smoothing coefficient, and take the product of the inverse of the sum and the gradient as the weighted gradient of the any modal interaction data.

[0045] In one embodiment, the weighted gradient is calculated as follows:

[0046] in, represents the weighted gradient of the j-th modal interaction data, represents the weight of the j-th modal interaction data, represents the gradient of the j-th modal interaction data, Indicates the preset smoothing coefficient.

[0047] It should be noted that is the smoothing coefficient, an infinitesimal constant, such as 0.01; it prevents the denominator from being 0; the gradient of the modality with poor performance (such as vision) will be amplified, and the gradient of the modality with good performance (such as speech) will be suppressed, thereby balancing the learning speed of each modality.

[0048] Similarly, the weighted gradient of each modal interaction data is obtained, and then the weighted gradients of all modal interaction data are fused (such as weighted summation) to obtain a global gradient, and the global gradient is used to update the model parameters to obtain an updated language model for outputting the response information of the multimodal interaction data. It should be noted that in the human-computer interaction scenario, the processing link of using dynamic gradient modulation (DGM) to avoid modal imbalance is mainly concentrated in the multimodal data fusion and model training stage, and the calculation frequency (the time range referenced by all the above modal interaction data) needs to be dynamically adjusted according to real-time requirements, data dynamics, computing resources and other factors, and is usually calculated in each round of model training iteration or each time window of real-time interaction. It is known that the focus of the present invention is to optimize the performance evaluation of each modal interaction data in the human-computer interaction scenario, and then use the corrected performance to perform gradient fusion and model update, which belongs to the prior art and will not be described in detail here.

[0049] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A human-computer interaction method based on multimodal data, characterized in that: The method comprises: Acquiring multimodal interaction data in real time, wherein the multimodal interaction data includes visual data, auditory data, and tactile data; For any modal interaction data, perform effective data and noise interference analysis on the any modal interaction data to obtain a data quality performance evaluation value of the any modal interaction data; obtain a real-time response performance evaluation value of the any modal interaction data based on the data volume and data response delay time contained in the any modal interaction data; and obtain an optimal performance compensation value for the any modal interaction data by combining the data quality performance evaluation value and the real-time response performance evaluation value; Obtaining a performance value of any modal interaction data calculated using dynamic gradient modulation, and correcting the performance value using the optimal performance compensation value to obtain a final performance value of any modal interaction data; obtaining the final performance value of each modal interaction data, and obtaining a weight of each modal interaction data; According to the weight of each modal interaction data, the weighted gradient of each modal interaction data is obtained, and gradient fusion and model update are performed according to the weighted gradient of each modal interaction data to obtain an updated language model, which is used to output response information of the multimodal interaction data.

2. The human-computer interaction method based on multimodal data according to claim 1, characterized in that: The performing effective data and noise interference analysis on the any modal interaction data to obtain a data quality performance evaluation value of the any modal interaction data includes: Obtain the ratio of the number of valid data to the total number of data in the interaction data of any modality, normalize the ratio to obtain a data validity index, calculate the signal-to-noise ratio of the interaction data of any modality, normalize the signal-to-noise ratio to obtain a data reliability index, and obtain a data quality performance evaluation value of the interaction data of any modality based on the average of the data validity index and the data reliability index.

3. The human-computer interaction method based on multimodal data according to claim 1, characterized in that: The obtaining, based on the data volume and data response delay time contained in the any modal interaction data, a real-time response performance evaluation value of the any modal interaction data includes: Obtain the proportion of the data volume contained in the any modal interaction data in the total data volume of all modal interaction data, normalize the proportion to obtain a data volume evaluation value, obtain the data response delay time of the any modal interaction data according to the time length from input to output of the any modal interaction data, use an exponential function with a natural constant as the base to inversely normalize the data response delay time to obtain an interaction real-time index; perform weighted processing on the data volume evaluation value and the interaction real-time index to obtain a real-time response performance evaluation value of the any modal interaction data.

4. The human-computer interaction method based on multimodal data according to claim 1, characterized in that: Combining the data quality performance evaluation value and the real-time response performance evaluation value to obtain the optimal performance compensation value of any modal interaction data includes: The average of the data quality performance evaluation value and the real-time response performance evaluation value is used as the optimal performance compensation value of any modal interaction data.

5. The human-computer interaction method based on multimodal data according to claim 1, characterized in that: The using the optimal performance compensation value to correct the performance value to obtain the final performance value of any modal interaction data includes: The performance value is normalized to obtain a normalized value, and the average of the optimal performance compensation value and the normalized value is calculated as the final performance value of the interaction data of any modality.

6. The human-computer interaction method based on multimodal data according to claim 1, characterized in that: The method for obtaining the weight of each modal interaction data includes: For any modal interaction data, the final performance value of the any modal interaction data is used as the independent variable of the softmax function to obtain the weight of the any modal interaction data.

7. The human-computer interaction method based on multimodal data according to claim 1, characterized in that: The step of obtaining the weighted gradient of each modal interaction data according to the weight of each modal interaction data includes: For any modal interaction data, obtain the gradient of the any modal interaction data, calculate the sum of the weight of the any modal interaction data and the preset smoothing coefficient, and take the product of the inverse of the sum and the gradient as the weighted gradient of the any modal interaction data.

Citation Information

Patent Citations

  • Multi-modal sentiment analysis method based on dynamic gradient and multi-view collaborative attention

    CN116204850A

  • Multi-mode interactive intelligent control system

    CN118226967A

  • User intention recognition method based on multi-modal cross improvement

    CN118656733A

  • Multi-modal data processing method and device based on large language model and reinforcement learning

    CN119338011A

  • Man-machine interaction system efficiency evaluation method and system based on multi-modal data

    CN119537165A