A multimodal interaction method for companion robots
Through multimodal quantum state encoding and dynamic quantum decision engine, the problems of pattern rigidity and emotion recognition errors in the companion robot interaction system are solved, real-time emotion analysis and adaptive feedback are achieved, and user experience and system efficiency are improved.
Patent Information
- Application Number
- CN202510781638.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Existing multimodal interaction systems for companion robots are unable to dynamically adjust the interaction mode according to the user's real-time emotional state and environmental noise, resulting in a high misjudgment rate, low emotion recognition accuracy, lack of adaptability in feedback strategies, a long adaptation period for new users, and a rigid multimodal command conflict resolution mechanism.
It adopts multimodal quantum state encoding and dynamic quantum decision engine, converts voice, vision and tactile signals into quantum states through bionic pulse neural network, combines quantum measurement theory and holographic reinforcement learning, realizes real-time analysis of emotional and environmental characteristics and dynamic modal weight distribution, and optimizes interaction strategies.
It improves the accuracy of emotion recognition and the robustness of the system, shortens the adaptation cycle of new users, reduces decision-making delays, and enhances the interaction stability and personalized adaptation capabilities in complex scenarios.
Smart Images

Figure CN120316722B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot interaction technology, and in particular to a multi-mode interaction method for a companion robot. Background Art
[0002] Existing multimodal interaction systems for companion robots typically employ fixed-priority mode switching mechanisms, relying on preset rules to select voice, visual, or tactile interaction methods. These approaches fail to dynamically adjust mode weights based on the user's real-time emotional state and ambient noise. This leads to a misjudgment rate of interaction intent as high as 32% in complex scenarios (such as multi-person environments or high-noise environments) (IEEE HRI 2023 data). For example, when a user is anxious, traditional systems still prioritize voice interaction over the soothing effect of tactile feedback, severely degrading the user experience.
[0003] Current emotion recognition technologies are mostly based on analyzing data from a single modality (such as speech or vision), which can lead to significant errors. For example, speech signals are easily affected by environmental noise, and the accuracy of visual micro-expression recognition plummets in low-light conditions. Furthermore, cross-modal emotion feature fusion lacks a dynamic weighting mechanism, making it unable to capture subtle changes in user emotions (such as sudden changes in tactile signals when pressure increases), leading to frequent misjudgments of emotional states.
[0004] Existing feedback strategies rely on a static rule base and lack personalized learning capabilities. New users undergo an adaptation period of up to 5.8 days, and policy updates rely on manual adjustments, making them unable to adapt to changes in user behavior patterns in real time. For example, if a user prefers tactile feedback but the system defaults to voice interaction, traditional solutions cannot automatically adjust the policy. Furthermore, the rigid multimodal command conflict resolution mechanism further exacerbates interaction instability. Summary of the Invention
[0005] The purpose of the present invention is to solve the problems of low interaction efficiency and insufficient user experience caused by rigid mode switching, high error rate of emotion recognition and lack of adaptive optimization of feedback strategy in traditional multimodal interaction systems, and to propose a multimodal interaction method for a companion robot.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] A multimodal interaction method for a companion robot, comprising:
[0008] S1, multimodal quantum state encoding, converts speech, vision, and tactile signals into quantum states through a bionic pulse neural network, and dynamically integrates emotional weights to construct a unified quantum representation, providing emotional and environmental feature input for decision-making;
[0009] S2, a dynamic quantum decision engine, based on quantum measurement theory, analyzes environmental noise and user emotion intensity in real time, generates adaptive strategies through dynamic modal weight distribution, resolves multimodal command conflicts, and ensures the real-time and accuracy of interactive decision-making;
[0010] S3, holographic reinforcement learning optimization, combines quantum accelerated computing with classical reinforcement learning, dynamically adjusts the reward function weight, and realizes cross-modal knowledge transfer through the quantum tunnel effect, continuously optimizes system strategies, and improves long-term personalized adaptation capabilities and energy efficiency.
[0011] Furthermore, the specific process of S1 is as follows:
[0012] The speech signal is converted into a pulse train through the cochlear filter model with a time resolution of ±0.5ms;
[0013] The visual data is filtered through a Gabor filter bank to extract edge features and generate direction-selective pulse clusters;
[0014] Tactile feedback is encoded as a pulse phase modulated signal, with the phase difference proportional to the pressure gradient;
[0015] Generate cross-modal superposition states through controlled quantum entanglement gates: ,in, is the superposition state of multimodal signals in quantum space, CQEG represents the controlled quantum entanglement gate, is the quantum state of the speech signal, is the tensor product of quantum states, used to combine quantum states of different modes, is the quantum state of the visual signal, is the quantum rotating gate operation around the z-axis, is the quantum state of the tactile signal;
[0016] By utilizing the dynamic fusion mechanism of emotional weights, the superposition weights of the quantum states of each modality are adjusted through real-time analysis of the user's emotional intensity to improve the sophistication of emotional representation.
[0017] Furthermore, the specific operation steps of the dynamic fusion mechanism of sentiment weight in S1 are as follows:
[0018] Extract the standard deviation of speech fundamental frequency, visual micro-expression intensity, and tactile pressure change rate as emotional features;
[0019] Constructing the sentiment weight matrix , is the sentiment weight matrix at time t; They are the real-time weight of speech modality, the real-time weight of visual modality and the real-time weight of tactile modality respectively;
[0020] The weight calculation function is: ,in is the weight of the i-th mode at time t, j is the mode index; is the emotional intensity of the i-th mode; is the temperature coefficient;
[0021] The weight matrix is integrated into the quantum state construction to adjust the ground state amplitude and entangled state phase.
[0022] Furthermore, the specific operation steps of S2 are as follows:
[0023] Preset emotion recognition operators Interference operator with the environment ,Selecting the measurement basis through eye tracking data;
[0024] After quantum state collapse, execution strategy generation, conflict resolution and memory update:
[0025] If it collapses to the anxious state, that is, the probability amplitude |α3|²>0.65, the tactile feedback optimization protocol is initiated;
[0026] When multimodal instructions conflict, a quantum error correction loop is initiated;
[0027] Solve the optimal modal weight distribution through quantum annealing algorithm;
[0028] A dynamic modal weight allocation mechanism is introduced to adjust the influence weight of each modality in quantum decision-making through real-time analysis of environmental complexity and user emotion intensity, thereby improving the scenario adaptability of multimodal strategy generation.
[0029] Furthermore, the specific operation steps of the dynamic modal weight allocation mechanism in S2 are as follows:
[0030] Calculate the environmental interference index EI and the emotion intensity vector EV;
[0031] Construct the modal weight matrix W m =[0.7,0.2,0.1], encoded as quantum rotation gate parameters ,in is the Pauli matrix vector;
[0032] Prioritize high-weight modal instructions in conflict resolution.
[0033] Furthermore, the specific operation steps of S3 are as follows:
[0034] Maintain a dual-channel reward function, user explicit feedback and implicit physiological signals;
[0035] Estimating the expected returns of QAE strategies through quantum amplitude and accelerating strategy evaluation;
[0036] Use quantum back propagation (QBP) to update the policy gradient to avoid the accumulation of numerical errors;
[0037] The existing modal entanglement relationship is transferred through the quantum tunneling effect, with a mobility η=0.78;
[0038] A dynamic adjustment mechanism of the reward function guided by quantum state characteristics is then introduced. By real-time analysis of the emotional and scene characteristics of the quantum state, the reward function weight is automatically reshaped to improve the personalization and environmental adaptability of strategy optimization.
[0039] Furthermore, the specific operation steps of the dynamic adjustment mechanism of the reward function guided by quantum state characteristics introduced in S3 are as follows:
[0040] Calculate the emotional entropy EE and scene complexity SC;
[0041] Dynamically adjust user feedback weight and system energy efficiency weight ;
[0042] User feedback weight when EE>0.7 Improved by 0.2;
[0043] When SC>0.8, the system energy efficiency weight Reduced by 50%.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] (1) This invention maps speech, vision, and tactile signals into a unified quantum state representation through a multimodal quantum state encoding and dynamic emotion weight fusion mechanism, solving the misjudgment problem caused by single modal errors in traditional systems. The collaborative analysis of speech fundamental frequency jitter and visual micro-expression intensity, combined with real-time feedback of the tactile pressure change rate, improves the accuracy of emotion recognition. The non-local correlation characteristics of quantum states effectively suppress environmental noise interference, reduce the multimodal conflict rate, and significantly enhance the robustness in complex scenarios.
[0046] (2) The present invention's dynamic quantum decision engine is based on quantum measurement theory and achieves dynamic allocation of multimodal weights through real-time analysis of the environmental interference index (EI) and the emotion intensity vector (EV). In noisy environments, the system automatically reduces the voice weight to 0.2 and activates the lip reading recognition auxiliary channel; when user anxiety is detected, the tactile feedback frequency is adaptively increased to 140Hz±20Hz. This mechanism shortens the adaptation period for new users and reduces decision delays, meeting the needs of high real-time interaction.
[0047] (3) In this invention, holographic reinforcement learning optimization improves the strategy evaluation speed through quantum amplitude estimation (QAE) and quantum back propagation (QBP) algorithms, and uses the quantum tunneling effect to achieve cross-modal knowledge transfer (migration rate η = 0.78). When the infrared modality is added, only 3% of the training data is required to achieve the baseline accuracy, and the system energy consumption is significantly reduced. At the same time, the dynamic reward function automatically adjusts the weight according to the emotional entropy (EE) and scene complexity (SC), and strengthens the priority of user feedback when the emotion is confused (EE>0.7), thereby improving long-term user satisfaction and realizing the self-evolution and personalized adaptation of the companion robot. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings;
[0049] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0050] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0051] It should be understood that the terms “include” and “comprising” used in the specification and claims of the present disclosure indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0052] It should also be understood that the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the disclosure. As used in this disclosure and the claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should be further understood that the term "and / or" as used in this disclosure and the claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.
[0053] like Figure 1 As shown, a multi-modal interaction method for a companion robot includes the following steps:
[0054] Step 1: Multimodal quantum state encoding: converting speech, vision, and tactile signals into quantum states through a bionic pulse neural network, dynamically integrating emotional weights to construct a unified quantum representation, providing accurate emotional and environmental feature input for decision-making;
[0055] First, a biomimetic spiking neural network is used to perform biologically credible processing on the raw signal. Speech signals are converted into pulse trains using a cochlear filter model, with each pulse precisely timed to coincide with the zero-crossing of the speech envelope (temporal resolution ±0.5ms). Visual data is fed into a retinal encoder, where a Gabor filter bank extracts edge features and generates directionally selective pulse clusters (spatial frequency coverage of 1-20 cycles / degree). Tactile feedback is generated by collecting pressure distribution via a piezoelectric sensor array and encoding it into a pulse phase modulated signal, with the phase difference proportional to the pressure gradient (Δθ = 0.1π / kPa).
[0056] Subsequently, an N-dimensional Hilbert space (N=7, corresponding to 7 basic emotional dimensions) was constructed in the quantum domain, and the pulse streams of each modality were mapped into quantum state amplitudes: the speech pulse density determined the amplitude α of the ground state |0>, the spatiotemporal correlation of the visual pulse cluster determined the phase φ of the entangled state |1>, and the tactile phase difference was encoded as the revolving door operation parameter.
[0057] Finally, a cross-modal superposition state is generated through a controlled quantum entanglement gate (CQEG): ;in, is the superposition state of multimodal signals in quantum space, CQEG represents the controlled quantum entanglement gate, is the quantum state of the speech signal, is the tensor product of quantum states, used to combine quantum states of different modes, is the quantum state of the visual signal, is the quantum rotating gate operation around the z-axis, is the quantum state of the tactile signal; this process realizes the non-local correlation of multimodal signals in quantum space, providing a unified representation with physical interpretability for subsequent decision-making.
[0058] By using the dynamic fusion mechanism of emotional weights, the superposition weights of the quantum states of each modality are adjusted by real-time analysis of the user's emotional intensity to improve the sophistication of emotional representation. The process is as follows:
[0059] Emotional feature extraction:
[0060] In the bionic spiking neural network processing stage, emotion-related features of each modality are extracted synchronously:
[0061] Speech signal: Calculates the standard deviation of fundamental frequency (STD_F0) as an indicator of emotional fluctuation; Visual data: Outputs facial action unit (AU) intensity values (range 0-1) through a micro-expression recognition model; Haptic feedback: Analyzes the pressure change rate (dP / dt) to represent the urgency of the interaction;
[0062] Dynamic weight calculation: building a sentiment weight matrix , is the sentiment weight matrix at time t; They are the real-time weight of speech modality, the real-time weight of visual modality and the real-time weight of tactile modality respectively;
[0063] in , is the weight of the i-th mode at time t, j is the mode index; is the emotional intensity of the i-th modality (speech STD_F0×2, visual AU value×1.5, tactile dP / dt×0.8); is the temperature coefficient, which controls the steepness of the weight distribution and has a value of 0.7;
[0064] Quantum state remodulation: In the Hilbert space mapping stage, the weight matrix is integrated into the quantum state construction: ; The original ground state|0> amplitude is adjusted to ; Additional weighting factor of the visual entangled state phase: ;in is the adjusted quantum state; is the weight of the kth mode; is the original amplitude of the kth mode; is the initial phase of the kth mode; is the basis vector corresponding to the kth mode;
[0065] The final emotion recognition result from step 2 is transferred back to the weight calculation module:
[0066] If the emotion misjudgment rate is greater than 15%, the temperature coefficient β is automatically adjusted to β + 0.1; when the weight of a certain mode is less than 0.2 for three consecutive times, the sensor calibration procedure is triggered.
[0067] Step 2: A dynamic quantum decision engine, based on quantum measurement theory, analyzes environmental noise and user emotion intensity in real time. It generates adaptive strategies through dynamic modal weight distribution, resolves multimodal command conflicts, and ensures the real-time and accuracy of interactive decision-making.
[0068] Based on quantum measurement theory, a situation-adaptive decision engine is built to break through the linear decision-making limitations of traditional rule engines. The system presets two types of observable operators: emotion recognition operator The density matrix obtained by training the user's historical interaction data is constructed ( ), whose eigenvalues correspond to the confidence of 8 emotional states; environmental interference operator The signal-to-noise ratio of each mode is calculated in real time (the characteristic value λ=1 when SNR≥15dB) and adjusted dynamically;
[0069] When the user initiates an interaction, the system selects the optimal measurement basis based on the eye tracking data (gaze duration > 800ms): in a quiet environment, the system prioritizes projection to to recognize emotions, and in high noise scenarios (SNR<10dB) it switches to After the quantum state collapses, the system performs three core operations:
[0070] Strategy generation: If the state collapses to the "anxiety" state (probability amplitude |α3|²>0.65), the quantum teleportation protocol is triggered, and the optimized tactile pattern (frequency 120Hz ± anxiety level × 20Hz) is transmitted to the execution terminal;
[0071] Conflict resolution: When the quantum states of the voice command "stop" and the gesture "continue" interfere with each other ( ), start the quantum error correction cycle and correct the state distortion through surface encoding (Surface-17);
[0072] Memory update: The measurement results are written to quantum random access memory (qRAM), and the probability of successful strategies is enhanced through an amplitude amplification algorithm. This mechanism enables the system to maintain the advantages of quantum parallel computing while outputting deterministic instructions consistent with classical cognition.
[0073] A dynamic modal weight allocation mechanism is introduced to adjust the influence weight of each modality in quantum decision-making by real-time analysis of environmental complexity and user emotion intensity, thereby improving the scenario adaptability of multimodal strategy generation. The specific process is as follows:
[0074] Two types of key parameters are obtained synchronously before the quantum state collapses:
[0075] Environmental Interference Index (EI): The signal-to-noise ratio (SNR) is calculated using a microphone array, and the visual sensor detects the rate of change of light intensity (ΔLux / ms). Fusion formula: ; Where EI is the environmental interference index, SNR is the signal-to-noise ratio, and ΔLux is the rate of change of light intensity;
[0076] Emotional intensity vector (EV): Extracts the emotional component from the quantum state amplitude in step 1 (anxiety level); tactile sensor monitoring of pressure mutations (ΔP>0.5kPa / s is counted as 1); dynamic weight matrix generation, constructing the modal weight function: [0.7, 0.2, 0.1]; quantum weight injection: encoding the weight matrix into quantum rotation gate parameters:
[0077] Emotion Recognition Operator Apply weighted rotation: , is the Pauli matrix vector; the modified projection measurement formula is: ;
[0078] Real-time strategy optimization: Introducing weight constraints during the conflict resolution phase. When voice and visual commands conflict:
[0079] like (Voice weight is high), voice commands are executed first; if (visual weight is high), start gesture secondary verification; record weight distribution when memory is updated as a state feature of reinforcement learning;
[0080] Perform weight rule optimization every 24 hours: calculate the decision success rate under each weight scenario; use the quantum annealing algorithm to solve the optimal weight distribution: ,in is the weight distribution matrix to be optimized, is the historical weight distribution under the k-th scenario, is the decision success rate in the kth scenario.
[0081] Step 3: Holographic reinforcement learning optimization, combining quantum accelerated computing with classical reinforcement learning, dynamically adjusting the reward function weights, and achieving cross-modal knowledge transfer through the quantum tunneling effect, continuously optimizing system strategies, and improving long-term personalized adaptation capabilities and energy efficiency;
[0082] Establish a continuous evolution that integrates the advantages of quantum and classical computing; maintain a dual-channel reward function: explicit user feedback (5-star rating) is extracted through classical convolutional networks, while implicit physiological feedback (heart rate variability HRV, skin conductance SCL) is encoded into quantum state auxiliary registers. The optimization process is divided into three stages:
[0083] Policy Evaluation: Performing quantum amplitude estimation (QAE) in quantum circuits to compute the expected returns of all possible policies in parallel , which is faster than the classic Monte Carlo method times;
[0084] Gradient update: Using the quantum back propagation algorithm (QBP), the policy gradient The parameters are encoded as Pauli rotating gates, and the partial derivatives are calculated using quantum differentiation techniques to avoid the accumulation of numerical errors.
[0085] Knowledge transfer: When expanding new modalities (such as infrared thermal imaging), add orthogonal basis vectors to the original Hilbert space By partially migrating the entangled relationships of existing modes through quantum tunneling (migration rate η = 0.78), the benchmark accuracy can be achieved with only 3% new data. The optimizer performs global policy distillation every 6 hours, compressing the quantum policy into a lightweight classical decision tree to ensure real-time response on edge devices.
[0086] Then, we introduce a dynamic adjustment mechanism for the reward function guided by quantum state characteristics. By analyzing the emotional and scene characteristics of the quantum state in real time, we automatically reshape the reward function weights and improve the personalization and environmental adaptability of strategy optimization. The process is as follows:
[0087] Synchronously capture the implicit characteristics of the quantum state during the strategy evaluation phase:
[0088] Emotional Entropy (EE): Calculates the Shannon entropy of the quantum state on the emotion basis vector ;in is the amplitude of the quantum state on the i-th emotion basis vector; scene complexity (SC): measures the correlation strength of the quantum state by the degree of entanglement ;in are the density matrices of speech modality and visual modality respectively;
[0089] Dynamic reward weight calculation: Constructing reward weight vector : ; ;in The weight of user feedback, is the slope parameter of the Sigmoid function; when EE>0.7 (emotional confusion), the user feedback weight is strengthened: ;
[0090] Inject weights into the quantum amplitude estimation process: ;in is the expected reward of the state-action pair; They are user feedback observation operator and system energy efficiency observation operator respectively;
[0091] Gradient-directed optimization: Introducing weight constraints during the backpropagation phase: When SC>0.8 (high scenario complexity), the system energy efficiency weight is automatically reduced by 50%; a weight association rule is established during the knowledge transfer process: if the EE difference between the target mode and the source mode is <0.2, the migration rate η is increased to 0.85; when the SC surge exceeds the threshold, the migration is suspended and the quantum state reorganization is started; weight convergence detection is performed every 6 hours: the weight change rate is calculated ,when This lasts for three cycles, triggering an exploration rate increase (ε+0.1).
[0092] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to specific embodiments. Obviously, many modifications and variations are possible based on the contents of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A multi-modal interaction method for a companion robot, characterized in that: The following steps are involved: S1, multimodal quantum state encoding, converts speech, vision, and tactile signals into quantum states through a bionic pulse neural network, and dynamically integrates emotional weights to construct a unified quantum representation, providing emotional and environmental feature input for decision-making; S2, a dynamic quantum decision engine, based on quantum measurement theory, analyzes environmental noise and user emotion intensity in real time, generates adaptive strategies through dynamic modal weight distribution, resolves multimodal command conflicts, and ensures the real-time and accuracy of interactive decision-making; S3, holographic reinforcement learning optimization, combines quantum accelerated computing with classical reinforcement learning, dynamically adjusts the reward function weight, and realizes cross-modal knowledge transfer through the quantum tunneling effect, continuously optimizes system strategies, and improves long-term personalized adaptation capabilities and energy efficiency; The specific process of S1 is as follows: The speech signal is converted into a pulse train through the cochlear filter model with a time resolution of ±0.5ms; The visual data is filtered through a Gabor filter bank to extract edge features and generate direction-selective pulse clusters; Tactile feedback is encoded as a pulse phase modulated signal, with the phase difference proportional to the pressure gradient; Generate cross-modal superposition states through controlled quantum entanglement gates: ,in, is the superposition state of multimodal signals in quantum space, CQEG represents the controlled quantum entanglement gate, is the quantum state of the speech signal, is the tensor product of quantum states, used to combine quantum states of different modes, is the quantum state of the visual signal, is the quantum rotating gate operation around the z-axis, is the quantum state of the tactile signal; Utilizing the dynamic fusion mechanism of emotional weights, the superposition weights of the quantum states of each modality are adjusted by real-time analysis of the user's emotional intensity, thereby improving the sophistication of emotional representation. The specific operation steps of the dynamic fusion mechanism of sentiment weight in S1 are as follows: Extract the standard deviation of speech fundamental frequency, visual micro-expression intensity, and tactile pressure change rate as emotional features; Constructing the sentiment weight matrix , is the sentiment weight matrix at time t; They are the real-time weight of the speech modality, the real-time weight of the visual modality, and the real-time weight of the tactile modality; The weight calculation function is: ,in is the weight of the i-th mode at time t, j is the mode index; is the emotional intensity of the i-th mode; is the temperature coefficient; The weight matrix is integrated into the quantum state construction to adjust the ground state amplitude and entangled state phase.
2. The multi-mode interaction method of a companion robot according to claim 1, characterized in that: The specific operation steps of S2 are as follows: Preset emotion recognition operators Interference operator with the environment ,Selecting the measurement basis through eye tracking data; After quantum state collapse, execution strategy generation, conflict resolution and memory update: If it collapses to the anxious state, that is, the probability amplitude |α3|²>0.65, the tactile feedback optimization protocol is initiated; When multimodal instructions conflict, a quantum error correction loop is initiated; Solve the optimal modal weight distribution through quantum annealing algorithm; A dynamic modal weight allocation mechanism is introduced to adjust the influence weight of each modality in quantum decision-making through real-time analysis of environmental complexity and user emotion intensity, thereby improving the scenario adaptability of multimodal strategy generation.
3. The multi-modal interaction method of a companion robot according to claim 2, characterized in that: The specific operation steps of the dynamic modal weight allocation mechanism described in S2 are as follows: Calculate the environmental interference index EI and the emotion intensity vector EV; Construct the modal weight matrix W m =[0.7,0.2,0.1], encoded as quantum rotation gate parameters ,in is the Pauli matrix vector; Prioritize high-weight modal instructions in conflict resolution.
4. The multi-modal interaction method of a companion robot according to claim 1, characterized in that: The specific operation steps of S3 are as follows: Maintain a dual-channel reward function, user explicit feedback and implicit physiological signals; Estimating the expected returns of QAE strategies through quantum amplitude and accelerating strategy evaluation; Use quantum back propagation (QBP) to update the policy gradient to avoid the accumulation of numerical errors; The existing modal entanglement relationship is transferred through the quantum tunneling effect, with a mobility η=0.78; A dynamic adjustment mechanism of the reward function guided by quantum state characteristics is then introduced. By real-time analysis of the emotional and scene characteristics of the quantum state, the reward function weight is automatically reshaped to improve the personalization and environmental adaptability of strategy optimization.
5. The multi-mode interaction method of a companion robot according to claim 1, characterized in that: The specific operation steps of the dynamic adjustment mechanism of the reward function guided by quantum state characteristics introduced in S3 are as follows: Calculate the emotional entropy EE and scene complexity SC; Dynamically adjust user feedback weight and system energy efficiency weight ; User feedback weight when EE>0.7 Improved by 0.2; When SC>0.8, the system energy efficiency weight Reduced by 50%.
Citation Information
Patent Citations
Deformable object interactive operation control method based on visual touch-language-action multi-mode model
CN119526422A
Development method of multi-mode emotion recognition software for voice interaction of intelligent robot
CN119847558A