Accompanying robot interaction control method and device, electronic equipment and storage medium
By synchronously collecting and fusing touch, voice, and visual signals to generate collaborative feature vectors, the problem of mechanical and emotional disconnect in the feedback of companion robots is solved, realizing human-like and highly adaptable interactive control, which is suitable for home, bionic, and portable emotional companion robots.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN MINRRAY IND CORP LTD
- Filing Date
- 2026-04-10
- Publication Date
- 2026-05-12
AI Technical Summary
Existing interactive control technologies for companion robots cannot build a multi-sensory linkage mechanism that conforms to the interaction logic of real pets, resulting in mechanical feedback, emotional disconnect, and poor scene adaptability.
Simultaneously collect user touch, voice, and visual signals, extract features through preprocessing, and fuse them using a dynamic weight allocation model to generate collaborative feature vectors. These vectors then generate action and voice control commands to achieve human-like interaction.
It achieves human-like and highly adaptable interaction between the companion robot and the user, solves the problems of mechanical feedback and poor scene adaptability, and adapts to various terminals to ensure real-time performance and stability.
Smart Images

Figure CN122018702A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of interactive control technology for intelligent companion robots, and in particular to an interactive control method, device, electronic device, and storage medium for companion robots. Background Technology
[0002] With the iterative upgrade of emotional companionship technology, companion robots have gradually evolved from a single entertainment medium into interactive partners with emotional perception and feedback capabilities. Users' core needs for them have shifted from "passively responding to single commands" to a higher-level experience of "active adaptation, multi-sensory resonance, and scenario-based interaction." However, the shortcomings of existing technologies in multimodal collaborative decision-making have become the core bottleneck restricting the upgrade of interactive experience.
[0003] Existing interactive control technologies for companion robots typically suffer from an inherent deficiency in collaborative capabilities, making it impossible to construct a multi-sensory linkage mechanism that conforms to the interaction logic of real pets. Summary of the Invention
[0004] This application provides a method, device, electronic device, and storage medium for interactive control of companion robots, which can solve at least one of the technical problems in the background art to a certain extent.
[0005] To achieve the above objectives, this application adopts the following technical solution: Firstly, a method for interactive control of a companion robot is provided, the method comprising: Simultaneously collect user touch signals, voice signals, and visual signals from the companion robot to obtain trimodal signals; The three-modal signals are preprocessed and three-modal features are extracted. The three-modal features are then fused using a pre-constructed dynamic weight allocation model to generate a collaborative feature vector. Based on the collaborative feature vector, action control commands and voice control commands for the companion robot are generated; The companion robot is controlled based on the action control commands and voice control commands, so that the companion robot can interact with the user.
[0006] Secondly, a companion robot interactive control device is provided, comprising: The acquisition module is used to simultaneously acquire the user's touch signals, voice signals, and visual signals to the companion robot, and obtain trimodal signals; The fusion module is used to preprocess the three-modal signals and extract the three-modal features, and fuse the three-modal features through a pre-constructed dynamic weight allocation model to generate a collaborative feature vector; The generation module is used to generate motion control commands and voice control commands for the companion robot based on the collaborative feature vector. The control module is used to control the companion robot based on the action control commands and voice control commands, so that the companion robot can interact with the user.
[0007] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the companion robot interactive control method as described in any one of the first aspects above.
[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the companion robot interactive control method as described in any one of the first aspects above.
[0009] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the companion robot interactive control method described in any of the first aspects above.
[0010] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.
[0011] In this embodiment, the user's touch, voice, and visual signals to the companion robot are first collected synchronously to obtain trimodal signals. These signals are then preprocessed to extract trimodal features. These features are then fused using a pre-built dynamic weight allocation model to generate a collaborative feature vector. Based on this vector, action and voice control commands are generated for the companion robot. The robot is then controlled based on these commands to enable interaction between the robot and the user. This approach is applicable to real-time interactive scenarios for home companion robots, bionic companion robots, and portable emotional companion robots. It addresses the issues of mechanical, emotionally disconnected, and poor scene adaptability caused by isolated processing of touch, voice, and visual signals, achieving human-like, highly adaptable interactive control that is compatible with various companion robot terminals, ensuring real-time performance and stability.
[0012] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0013] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiments below. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a flowchart illustrating the interactive control method for a companion robot provided in an embodiment of this application; Figure 2 A hardware and software collaborative workflow diagram of the companion robot interactive control method based on touch-voice-vision multimodal collaborative decision-making provided in the embodiments of this application; Figure 3 This is a schematic diagram of a three-modal interactive control architecture provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of the companion robot interactive control device provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0014] The embodiments of the technical solutions of this application will now be described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of this application, and are therefore merely examples and should not be used to limit the scope of protection of this application. When the following description relates to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but can be changed as will become apparent upon understanding this disclosure, except for operations that must be performed in a specific order. Furthermore, for clarity and conciseness, descriptions of features known in the art may be omitted.
[0015] The embodiments described in the following examples of this disclosure are not representative of all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0016] It should be noted that some solutions collect external stimuli such as touch and sound through multiple sensors, extract feature parameters, and determine the user interaction mode based on similarity. They can output personalized and intimate actions without prior registration. However, they only cover the core features of touch and sound dual-modality and do not systematically integrate key visual information such as user facial emotions and scene environment. They lack trimodal collaborative decision-making logic and the device lacks a dedicated visual signal processing and synchronization unit. Therefore, they cannot generate anthropomorphic feedback that fits the user's emotions and scene, and the feedback is highly mechanical.
[0017] Other solutions use multi-sensor fusion to collect multimodal data and combine algorithms to optimize feature extraction and motion mapping to improve the naturalness of interaction. However, they focus on remote pose mapping and motion optimization, lack multimodal deep linkage rules, do not build a dynamic weight allocation mechanism, only achieve simple signal superposition rather than collaborative decision-making, and do not design adaptation strategies for the close-range and multi-scenario interaction characteristics of companion robots, resulting in weak scene adaptability and emotional resonance capabilities.
[0018] See Figure 1 This is a flowchart illustrating the interactive control method for a companion robot provided in an embodiment of this application. The specific implementation methods of this application will be described in detail below with reference to the accompanying drawings.
[0019] like Figure 1 As shown, the companion robot interactive control method provided in this embodiment includes the following steps: Step 101: Simultaneously collect the user's touch signals, voice signals, and visual signals to the companion robot to obtain trimodal signals.
[0020] Touch signals can be collected using tactile sensors, such as a distributed flexible tactile sensor array. Voice signals can be collected using microphone arrays, such as a 4-channel microphone array. Visual signals can be collected using a high-definition camera module.
[0021] As an example, during the signal acquisition phase, a distributed flexible tactile sensor array, a 4-channel microphone array, and a high-definition camera module are used to synchronously acquire the touch signals of the user acting on the companion robot, the voice signals emitted by the user, and the visual signals of the user and the surrounding environment, respectively. A unified sampling timing benchmark is set, with the sampling rate of the touch signal being 100Hz, the sampling rate of the voice signal being 16kHz, and the acquisition frame rate of the visual signal being 30fps, to ensure the synchronization of the three-modal signals in the time dimension, and the timing deviation is controlled within 5ms.
[0022] It should be noted that different denoising algorithms can be used to filter interference based on the noise characteristics of different modal signals. For example, a sliding window denoising algorithm can be used to filter random noise collected by the sensor for touch signals, a spectral subtraction denoising algorithm can be used to eliminate environmental background noise for speech signals, and an inter-frame difference denoising algorithm can be used to remove salt-and-pepper noise and motion artifacts from visual signals.
[0023] Specifically, the denoised trimodal signals are then processed through signal standardization to generate standardized original touch signals, standardized original speech signals, and standardized original visual signals with uniform format and dimensionality, i.e., standardized original trimodal signals.
[0024] In this embodiment, the system may include an 8-channel distributed flexible tactile sensor array, a 4-channel adaptive noise-reducing microphone array, a high-definition camera module with infrared illumination, and a corresponding signal preprocessing unit. The preprocessing unit integrates filtering, amplification, and A / D conversion circuits, and transmits the signal synchronously to the three-modal collaborative decision control device via a high-speed link with a transmission rate ≥1Mbps and a timing deviation ≤5ms.
[0025] Step 102: Preprocess the three-modal signal and extract the three-modal features, and fuse the three-modal features through a pre-constructed dynamic weight allocation model to generate a collaborative feature vector.
[0026] Optionally, the touch location, touch force, and touch timing of the touch signal can be extracted to obtain touch features; the semantic intent and voice emotion of the voice signal can be extracted to obtain voice features; and the user posture, user facial emotion, and scene type of the visual signal can be extracted to obtain visual features.
[0027] As one possible implementation, the touch location can be an interactive area such as the head, back, or paws; the touch intensity can be represented by a quantization score from 0 to 10; and the touch timing can be represented by the number of consecutive touches and the duration of the touch interval. These features, including touch location, touch intensity, and touch timing, can be quantized into multi-dimensional vectors to form touch features.
[0028] As one possible implementation, semantic intent can be a summoning, soothing, or command, represented by 0 / 1 binary representation, while vocal emotion can be pleasant, anxious, or neutral, quantized using 0-1 scores. These semantic intent and vocal emotion features of the aforementioned vocal signal are quantized into multi-dimensional vectors to form vocal features.
[0029] As one possible implementation, user postures, such as squatting, standing, or waving, can be categorized and encoded. User facial emotions, such as happiness, anger, or calmness, can be quantized using a 0-1 score. Scene types can be at home, outdoors, or in low light. The user posture, facial emotions, and scene type features of the above visual signals are quantized into multi-dimensional vectors to form visual features.
[0030] Furthermore, a dynamic weight allocation model can be constructed based on the optimized DS evidence theory and the signal credibility of the three-modal signals. Based on the dynamic weight allocation model, touch features, voice features and visual features can be dynamically fused to generate a collaborative feature vector.
[0031] Optionally, in the dynamic weight allocation model, the weight values for touch features range from 0.3 to 0.5, the weight values for voice features range from 0.2 to 0.4, and the weight values for visual features range from 0.2 to 0.4.
[0032] When the confidence level of visual features is ≥0.8, the weight of visual features is adjusted to 0.4, and the weights of touch features and voice features are both 0.3.
[0033] When the touch signal is continuously and stably collected ≥3 times, the weight of the touch feature is adjusted to 0.5, and the weights of the voice feature and the visual feature are both 0.25.
[0034] It should be noted that the dynamic weight allocation model is built on the optimized DS evidence theory. It dynamically allocates the fusion weights of each modality feature based on the signal confidence of the three-modal signals (such as the confidence score of the visual recognition algorithm in the 0-1 range for visual features), thereby solving the decision bias problem caused by traditional fixed weight fusion.
[0035] In the dynamic weight allocation model, the weight values of each modality feature are as follows: touch feature weight 0.3-0.5, voice feature weight 0.2-0.4, and visual feature weight 0.2-0.4, and the sum of the three modality weights is 1.
[0036] Optionally, when the confidence level of visual features is ≥0.8, it indicates that the visual signal recognition result is stable, and the weight of visual features is adjusted to 0.4, while the weights of touch features and voice features are both allocated to 0.3. When touch signals are continuously and stably collected ≥3 times, it indicates that the confidence level of touch features is high, and the weight of touch features is adjusted to 0.5, while the weights of voice features and visual features are both allocated to 0.25.
[0037] In addition to the two scenarios mentioned above, the default weights for touch features (0.4), voice features (0.3), and visual features (0.3) can also be assigned.
[0038] Furthermore, based on the weight values output by the dynamic weight allocation model, the touch features, voice features, and visual features are weighted and fused. The fused feature vectors are then concatenated into a collaborative feature vector. For example, if a user taps the robot's head three times consecutively, with a touch force score of 6, the touch location (head), force (6), and timing (three consecutive taps with a 0.5s interval) are extracted to generate a 6-dimensional touch feature: [1,6,3,0.5,0,0]. If the user says "come here," the semantic intent is "summon" (encoded 1), and the voice emotion is "pleasant" (score 0.7), generating a 5-dimensional voice feature: [1,0.7,0,0,0]. The user's squatting posture (encoded 1), facial emotion of happiness (score 0.85), and the scene of being at home (encoded 0) are recognized. The visual feature confidence score is 0.75, generating an 11-dimensional visual feature: [1,0.85,0,0,0,0,0,0,0,0,0.75]. Since the visual feature confidence score is 0.75 < 0.8, but the touch signal was continuously and stably collected three times, when the touch signal is continuously and stably collected ≥ 3 times, it indicates high touch feature confidence. Therefore, the touch feature weight is adjusted to 0.5. Weighting: Touch feature weight = 0.5, Voice feature weight = 0.25, Visual feature weight = 0.25.
[0039] Step 103: Generate motion control commands and voice control commands for the companion robot based on the collaborative feature vector.
[0040] Optionally, the collaborative feature vector is matched with a pre-built trimodal collaborative mapping rule library to obtain the action control command and voice control command corresponding to the current combination relationship. Based on the trimodal features and the distance parameters between the companion robot and the user, it is determined whether the companion robot's active approach action is triggered. If triggered, the action parameters of the active approach action are added to the action control command.
[0041] Among them, the three-modal collaborative mapping rule base is a dynamic and iterative rule base. The rule base pre-stores multi-dimensional combination relationships of touch features, voice features, and visual features, as well as the action control instructions and voice control instructions corresponding to each combination relationship.
[0042] Each rule consists of three parts: feature combination conditions, action parameters, and voice parameters. The feature combination conditions specify the touch, voice, and visual feature thresholds that trigger the rule. The action parameters include action type, action amplitude, and action sequence. The voice parameters include voice content, volume, speech rate, and emotional tone.
[0043] Specifically, the collaborative feature vector can be decomposed into touch, voice, and visual sub-features. All rules in the rule base are traversed, the similarity between the current feature and the feature combination conditions in each rule is calculated, the rule with the highest similarity is selected as the matching rule, and its corresponding action control command and voice control command are output.
[0044] Furthermore, based on the trimodal features and the distance parameters between the companion robot and the user, it can be determined whether to trigger an active approach action. This can be triggered if any of the following rules are met: Touch characteristics include: ≥3 consecutive touches + ≥6 touch pressure + user distance ≤1 meter; In speech features, semantic intent = call to action + speech emotion ≥ 0.5 + user distance ≤ 1.5 meters; In visual features, user posture = waving + facial emotion = pleasure + user distance ≤ 2 meters.
[0045] It should be noted that if an active approach action is triggered, active approach action parameters are inserted at the beginning of the execution sequence of the basic action control command, including approach speed (the closer the distance, the lower the speed), target distance (finally stopping 0.3-0.5 meters from the user), and action smoothness (to avoid sudden stops and turns).
[0046] Optionally, the system can also acquire current visual scene features and then optimize the motion parameters in motion control commands and the voice parameters in voice control commands based on these features. Specifically, it can acquire current visual scene features, such as those extracted from visual characteristics like home, outdoors, or low light, and then specifically optimize the parameters of motion control and voice control commands to ensure the commands are adapted to the scene characteristics. For example, in outdoor scenes, motion amplitude and voice volume are increased by 30% and 15% respectively compared to home scenes; in low light scenes, motion amplitude is reduced by 40%, voice volume is reduced by 20%, and speech rate is slowed by 15%. In confined spaces, motion amplitude is compressed by 50%.
[0047]
[0048] The example table above is a rule example table of a trimodal collaborative mapping rule base, which defines the specific combination relationship of trimodal features and the mapping relationship between them and the corresponding servo motor action commands and voice response strategies of the companion robot.
[0049] Step 104: Control the companion robot based on action control commands and voice control commands to enable the companion robot to interact with the user.
[0050] Optionally, interaction feedback data between the companion robot and the user can be obtained, and then the three-modal collaborative mapping rule base can be iteratively optimized based on the interaction feedback data. The interactive feedback data includes positive user feedback data, negative user feedback data, and neutral user feedback data.
[0051] Specifically, the generated motion control commands are sent to the servo motor motion execution device of the companion robot to drive the robot to complete limb movements. The servo motor motion execution device includes ≥6 multi-degree-of-freedom servos, a servo motor controller, a drive circuit, and a trajectory optimization unit. It uses a trapezoidal acceleration / deceleration algorithm to optimize the motion trajectory, with a motion response delay of ≤10ms.
[0052] Specifically, voice control commands are sent to a voice generation device, which simultaneously plays the voice content, enabling real-time interaction between the robot and the user. Actions and voice execution are strictly synchronized (voice sounds within 0.2 seconds of action initiation), simulating the interaction logic of a real pet. The voice generation device integrates a synchronization calibration module, which works in conjunction with the servo motor action execution device, ensuring a timing deviation of ≤5ms between action and voice initiation.
[0053] The following example illustrates this point: Case 1: Intimate and comforting interaction at home (users prefer gentle interactions) Touch "behind the ear + light touch + 4s", voice "good baby" (soothing + happy), visual "smile + still + home"; Visual confidence score is 0.85, with weights allocated as follows: touch 0.3, voice 0.3, and vision 0.4, generating a collaborative feature vector. Rule matching: Matched the rules for intimate interactions, but did not trigger an initiative to approach; The range of motion is increased by 30%, and intimate voice messages are output simultaneously with a time difference of 0.08 seconds; For every 12 seconds of continuous user interaction (positive feedback), the corresponding rule weight increases by 0.1.
[0054] Case 2: Outdoor Call-Along Reassurance Interaction (User Emotionally Exhausted) Touching the abdomen (3 times), speaking "come here" (calling + tired), or visually waving (calm + outdoors) triggers a move closer. When the touch signal is stable, the weighting is 0.5 for touch, 0.25 for voice, and 0.25 for vision. Matching soothing rules and overlaying outdoor scene parameters, the approach speed is 1.1cm / s; The range of motion increased by 30%, the volume of the voice increased by 15%, and the speaking speed slowed down.
[0055] Optionally, the companion robot's sensors (such as touch sensors to detect whether the user continues to interact and visual sensors to capture changes in the user's facial emotions) can be used to collect interaction feedback data, which can be divided into positive feedback (such as continued touching and happy facial expression), negative feedback (such as avoiding the robot and unhappy facial expression), and neutral feedback (no obvious emotional change and no subsequent interaction).
[0056] Optionally, positive feedback increases the corresponding rule weight by 0.1, while negative feedback decreases it by 0.15. The rule base weights are updated every 20 valid data points, and the weight allocation model and synchronization parameters are optimized every 50 data points. User manual calibration of preferences is supported. For positive feedback, the matching priority of the rule combination is increased, enhancing the corresponding action amplitude and voice volume. For negative feedback, the matching priority of the rule combination is decreased, adjusting the action type or voice content. If there are three consecutive negative feedbacks, the rule is temporarily disabled. For neutral feedback, the rule parameters remain unchanged, and the feedback data is recorded for batch optimization.
[0057] In this embodiment, the user's touch, voice, and visual signals to the companion robot are first collected synchronously to obtain trimodal signals. These signals are then preprocessed to extract trimodal features. These features are then fused using a pre-built dynamic weight allocation model to generate a collaborative feature vector. Based on this vector, action and voice control commands are generated for the companion robot. The robot is then controlled based on these commands to enable interaction between the robot and the user. This approach is applicable to real-time interactive scenarios for home companion robots, bionic companion robots, and portable emotional companion robots. It addresses the issues of mechanical, emotionally disconnected, and poor scene adaptability caused by isolated processing of touch, voice, and visual signals, achieving human-like, highly adaptable interactive control that is compatible with various companion robot terminals, ensuring real-time performance and stability.
[0058] Figure 2 This document presents a hardware and software collaborative workflow diagram for a companion robot interactive control method based on touch-voice-vision multimodal collaborative decision-making, as provided in an embodiment of this application. The flowchart comprises four stages: Stage 1 is signal synchronization acquisition, taking ≤8ms. The user / environment generates touch, voice, and visual signals. The signal acquisition device performs synchronous preprocessing on the three modal signals and then transmits the synchronized three-modal data to the collaborative decision-making control device. Stage 2 is dynamic fusion and decision-making, taking ≤12ms. The collaborative decision-making control device sequentially performs feature extraction and dynamic weight fusion, matching the collaborative rule base, generating action / voice control commands, and sending action commands to the servo motor action execution device and voice commands to the voice generation device, respectively. Stage 3 is scenario-based execution. The servo motor action execution device executes scenario-based actions with a response delay ≤10ms, and the voice generation device plays adapted voice, with a time difference between action and voice output ≤0.1s. Stage 4 is feedback iteration, i.e., closed-loop optimization. The user / environment generates feedback on the interactive behavior. The feedback acquisition device generates new feedback data and uploads it to the collaborative decision-making control device. The collaborative decision-making control device updates the rule base and parameters, completing the closed-loop optimization.
[0059] Figure 3This is a schematic diagram of a three-modal interactive control architecture provided in an embodiment of this application. The architecture is divided into a hardware device layer and a software algorithm layer. The hardware device layer includes a "touch-voice-visual signal acquisition device," a three-modal collaborative decision-making control device, a servo motor action execution device, a voice generation device, and a feedback iteration device. The software algorithm layer includes a signal synchronization acquisition module, a feature extraction and dynamic weight fusion module, a collaborative mapping rule base matching module, an action / voice parameter scenario adaptation module, and an interactive feedback iteration optimization module. The "touch-voice-visual signal acquisition device" generates raw interactive signals and synchronously transmits the raw data to the signal synchronization acquisition module of the software algorithm layer. After synchronization clock calibration, the signal synchronization acquisition module outputs standardized time-series data to the feature extraction and dynamic weight fusion module. The feature extraction and dynamic weight fusion module outputs the comprehensive feature vector after dynamic weight fusion to the collaborative mapping rule base. The same mapping rule base matching module outputs interaction strategies to the action / voice parameter scenario adaptation module. The action / voice parameter scenario adaptation module outputs the adapted action and voice commands to the three-modal collaborative decision control device. The three-modal collaborative decision control device outputs action control commands to the servo motion execution device and voice control commands to the voice generation device. The servo motion execution device executes mechanical actions, and the voice generation device plays voice responses. Both act on the user / environment (touch, voice, and visual signals). The user / environment generates new feedback behavior data to the feedback iteration device. The feedback iteration device collects and uploads the feedback data to the interactive feedback iteration optimization module of the software algorithm layer. The interactive feedback iteration optimization module outputs optimized rule base weights to update the adaptation parameters and sends them to the collaborative mapping rule base matching module, forming a complete interactive control closed loop.
[0060] The embodiments of this application have at least the following beneficial effects: 1. Resolves the ambiguity of multimodal signals, significantly improving decision-making accuracy. Employing a dynamic weight allocation model based on signal credibility, the system automatically adjusts the weights of each modality when visual emotion / posture is clear (credibility ≥ 0.8) or the touch signal is stable, effectively avoiding ambiguity of single-modal signals.
[0061] 2. Achieve anthropomorphic emotions and temporal synchronization, significantly enhancing the naturalness of interaction. Through a three-dimensional collaborative mapping rule base of touch-voice-vision, combined with a dual synchronization strategy of action-voice, the robot's feedback is highly consistent in terms of emotional expression, temporal connection, and behavioral logic.
[0062] 3. Enhanced adaptability to complex scenarios and significantly expanded scope of application: Through the scenario-based adaptation engine, the system automatically adjusts the range of motion and voice parameters based on visual information. The same interaction rules can perform reasonably in scenarios such as home, outdoors, and low light, solving the problem of weak scenario adaptability of existing technologies.
[0063] 4. Constructs a closed-loop software and hardware collaboration system with excellent feasibility and generalizability: Adopting a modular software and hardware collaboration architecture, it meets real-time interaction requirements. Its lightweight design allows deployment on low-computing-power terminals, is not bound to specific hardware forms, and is compatible with various companion robot terminals.
[0064] See Figure 4 This is a structural schematic diagram of the companion robot interactive control device provided in an embodiment of this application. Figure 4 As shown, the companion robot interactive control device provided in this embodiment includes: The acquisition module 210 is used to simultaneously acquire the user's touch signals, voice signals and visual signals to the companion robot to obtain trimodal signals; The fusion module 220 is used to preprocess the three-modal signal and extract the three-modal features, and fuse the three-modal features through a pre-constructed dynamic weight allocation model to generate a collaborative feature vector; The generation module 230 is used to generate motion control commands and voice control commands for the companion robot based on the collaborative feature vector. The control module 240 is used to control the companion robot based on the action control commands and voice control commands, so that the companion robot can interact with the user.
[0065] Optionally, the generation module 230 is specifically used for: The collaborative feature vector is matched with a pre-built trimodal collaborative mapping rule library to obtain the action control command and voice control command corresponding to the current combination relationship. The trimodal collaborative mapping rule library is a dynamic iterative rule library. The rule library pre-stores multi-dimensional combination relationships of touch features, voice features, and visual features, as well as the action control command and voice control command corresponding to each combination relationship. Based on the trimodal features and the distance parameters between the companion robot and the user, it is determined whether to trigger the companion robot's active approach action; If triggered, the action parameters for the active approach action are added to the action control command.
[0066] Optionally, the generation module 230 is also used for: Obtain the current visual scene features; The motion parameters in the motion control command and the voice parameters in the voice control command are optimized based on the visual scene features.
[0067] Optionally, the fusion module is specifically used for: The touch position, touch force, and touch timing of the touch signal are extracted to obtain touch features; Extract the semantic intent and emotional tone of the speech signal to obtain speech features; The user's posture, facial emotions, and scene type are extracted from the visual signals to obtain visual features; Based on the optimized DS evidence theory and the signal credibility of the three-mode signal, the dynamic weight allocation model is constructed. Based on the dynamic weight allocation model, the touch features, the voice features, and the visual features are dynamically fused to generate the collaborative feature vector.
[0068] Optionally, the device also includes a feedback optimization module for: Obtain the interaction feedback data between the companion robot and the user; Based on the interactive feedback data, the three-modal collaborative mapping rule base is iteratively optimized. The interactive feedback data includes positive user feedback data, negative user feedback data, and neutral user feedback data.
[0069] Optionally, in the dynamic weight allocation model, the weight of touch features ranges from 0.3 to 0.5, the weight of voice features ranges from 0.2 to 0.4, and the weight of visual features ranges from 0.2 to 0.4. When the confidence level of the visual feature is ≥0.8, the weight of the visual feature is adjusted to 0.4, and the weights of the touch feature and the voice feature are both 0.3. When the touch signal is continuously and stably collected ≥3 times, the touch feature weight is adjusted to 0.5, and the weights of voice features and visual features are both 0.25.
[0070] In this embodiment, the user's touch, voice, and visual signals to the companion robot are first collected synchronously to obtain trimodal signals. These signals are then preprocessed to extract trimodal features. These features are then fused using a pre-built dynamic weight allocation model to generate a collaborative feature vector. Based on this vector, action and voice control commands are generated for the companion robot. The robot is then controlled based on these commands to enable interaction between the robot and the user. This approach is applicable to real-time interactive scenarios for home companion robots, bionic companion robots, and portable emotional companion robots. It addresses the issues of mechanical, emotionally disconnected, and poor scene adaptability caused by isolated processing of touch, voice, and visual signals, achieving human-like, highly adaptable interactive control that is compatible with various companion robot terminals, ensuring real-time performance and stability.
[0071] in addition, Figure 4The interactive control device for the companion robot shown can be a software unit, a hardware unit, or a combination of software and hardware built into existing electronic devices. It can also be integrated into electronic devices as an independent accessory, or exist as an independent electronic device.
[0072] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0073] Figure 5 This is a schematic diagram of the structure of the electronic device provided in an embodiment of this application. For example... Figure 5 As shown, the electronic device 5 of this embodiment includes: at least one processor 50 ( Figure 5 (Only one is shown in the diagram) a processor, a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50, wherein the processor 50 executes the computer program 52 to implement the steps in any of the above embodiments of the companion robot interactive control method.
[0074] The electronic device may be a desktop computer, laptop, handheld computer, or cloud server, etc. This electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 5 This is merely an example of electronic device 5 and does not constitute a limitation on electronic device 5. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0075] The processor 50 may be a central processing unit, or it may be other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0076] In some embodiments, the memory 51 may be an internal storage unit of the electronic device 5, such as a hard disk or memory of the electronic device 5. In other embodiments, the memory 51 may be an external storage device of the electronic device 5, such as a plug-in hard disk, smart memory card, secure digital card, flash memory card, etc., equipped on the electronic device 5. Further, the memory 51 may include both internal storage units and external storage devices of the electronic device 5. The memory 51 is used to store operating systems, applications, boot loaders, data, and other programs, such as the program code of the computer program. The memory 51 can also be used to temporarily store data that has been output or will be output.
[0077] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps in the above-described method embodiments.
[0078] This application provides a computer program product that, when run on an electronic device, enables the electronic device to implement the steps described in the various method embodiments above.
[0079] If the integrated unit is implemented as a software functional unit and used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / electronic device, a recording medium, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0080] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0081] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0082] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0083] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0084] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for interactive control of a companion robot, characterized in that, include: Simultaneously collect user touch signals, voice signals, and visual signals from the companion robot to obtain trimodal signals; The three-modal signals are preprocessed and three-modal features are extracted. The three-modal features are then fused using a pre-constructed dynamic weight allocation model to generate a collaborative feature vector. Based on the collaborative feature vector, action control commands and voice control commands for the companion robot are generated; The companion robot is controlled based on the action control commands and voice control commands, so that the companion robot can interact with the user.
2. The method according to claim 1, characterized in that, The step of generating motion control commands and voice control commands for the companion robot based on the collaborative feature vector includes: The collaborative feature vector is matched with a pre-built trimodal collaborative mapping rule library to obtain the action control command and voice control command corresponding to the current combination relationship. The trimodal collaborative mapping rule library is a dynamic iterative rule library. The rule library pre-stores multi-dimensional combination relationships of touch features, voice features, and visual features, as well as the action control command and voice control command corresponding to each combination relationship. Based on the trimodal features and the distance parameters between the companion robot and the user, it is determined whether to trigger the companion robot's active approach action; If triggered, the action parameters for the active approach action are added to the action control command.
3. The method according to claim 2, characterized in that, After generating motion control commands and voice control commands for the companion robot based on the collaborative feature vector, the method further includes: Obtain the current visual scene features; The motion parameters in the motion control command and the voice parameters in the voice control command are optimized based on the visual scene features.
4. The method according to claim 2, characterized in that, The process of preprocessing the trimodal signal and extracting trimodal features, and then fusing the trimodal features using a pre-constructed dynamic weight allocation model to generate a collaborative feature vector, includes: The touch position, touch force, and touch timing of the touch signal are extracted to obtain touch features; Extract the semantic intent and emotional tone of the speech signal to obtain speech features; The user's posture, facial emotions, and scene type are extracted from the visual signals to obtain visual features; Based on the optimized DS evidence theory and the signal credibility of the three-mode signal, the dynamic weight allocation model is constructed. Based on the dynamic weight allocation model, the touch features, the voice features, and the visual features are dynamically fused to generate the collaborative feature vector.
5. The method according to claim 4, characterized in that, After controlling the companion robot based on the action control commands and voice control commands to enable the companion robot to interact with the user, the method further includes: Obtain the interaction feedback data between the companion robot and the user; Based on the interactive feedback data, the three-modal collaborative mapping rule base is iteratively optimized. The interactive feedback data includes positive user feedback data, negative user feedback data, and neutral user feedback data.
6. The companion robot interactive control method according to claim 5, characterized in that, In the dynamic weight allocation model, the weight values for touch features range from 0.3 to 0.5, the weight values for voice features range from 0.2 to 0.4, and the weight values for visual features range from 0.2 to 0.
4. When the confidence level of the visual feature is ≥0.8, the weight of the visual feature is adjusted to 0.4, and the weights of the touch feature and the voice feature are both 0.
3. When the touch signal is continuously and stably collected ≥3 times, the touch feature weight is adjusted to 0.5, and the weights of voice features and visual features are both 0.
25.
7. A companion robot interactive control device, characterized in that, include: The acquisition module is used to simultaneously acquire the user's touch signals, voice signals, and visual signals to the companion robot, and obtain trimodal signals; The fusion module is used to preprocess the three-modal signals and extract the three-modal features, and fuse the three-modal features through a pre-constructed dynamic weight allocation model to generate a collaborative feature vector; The generation module is used to generate motion control commands and voice control commands for the companion robot based on the collaborative feature vector. The control module is used to control the companion robot based on the action control commands and voice control commands, so that the companion robot can interact with the user.
8. The apparatus according to claim 7, characterized in that, The generation module is specifically used for: The collaborative feature vector is matched with a pre-built trimodal collaborative mapping rule library to obtain the action control command and voice control command corresponding to the current combination relationship. The trimodal collaborative mapping rule library is a dynamic iterative rule library. The rule library pre-stores multi-dimensional combination relationships of touch features, voice features, and visual features, as well as the action control command and voice control command corresponding to each combination relationship. Based on the trimodal features and the distance parameters between the companion robot and the user, it is determined whether to trigger the companion robot's active approach action; If triggered, the action parameters for the active approach action are added to the action control command.
9. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the method as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, which is loaded by a processor to perform the steps of the method according to any one of claims 1 to 6.