Robot social interaction system based on multi-modal physical twinning and synchronous control method

By integrating visual sensors and 5G NR protocol into a multimodal physical twin robot social interaction system, the system solves the problems of intention recognition errors and latency in noisy environments for social robots, improves the emotional interaction capabilities and virtual-real motion synchronization of industrial robots, and achieves high-precision operation and low-error switching.

CN120872138AInactive Publication Date: 2025-10-31AIMI (BEIJING) ROBOT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510752151.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-10-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing social robots have a high error rate in intent recognition in noisy environments, and the perception-decision-execution link is prolonged, resulting in delayed facial expression feedback and a mechanical feeling for users; industrial robots lack emotional interaction capabilities and cannot meet the requirements of precision operation; high packet loss rate on public networks leads to frequent accidents caused by asynchrony between virtual and real actions.

Method used

The robot social interaction system adopts multimodal physical twins, integrating visual sensors, LiDAR, and microphone arrays. It achieves ultra-low latency synchronization through the 5G NR protocol, combines reinforcement learning models to generate interaction strategies, uses dynamic spectrum sharing technology to maintain communication reliability, and activates predictive motion compensation through a virtual-real synchronization module error correction protocol.

Benefits of technology

Maintaining an intent recognition accuracy of 96.5% in noisy environments, compressing the virtual-real action synchronization error to 8.5ms, and improving communication reliability to 99.94%, it achieves millisecond-level scene switching and high-precision operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872138A_ABST
    Figure CN120872138A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robot interaction, and provides a robot social interaction system based on multi-modal physical twinning, which comprises a physical twinning modeling unit for constructing a high-precision dynamic model of a robot; the multi-mode sensing unit integrates a visual sensor, a laser radar and a microphone array and collects environment data in real time; the communication control unit is used for realizing multi-robot synchronization with ultra-low time delay by adopting a 5G NR protocol; and the social decision engine generates an interaction strategy based on a reinforcement learning model, and a strategy function is shown in the specification, theta is a network weight, and st is a multi-modal fusion state. According to the method, the inhibition effect on tiny gestures is achieved, the false triggering rate in a child interaction scene is reduced by 82%, seamless switching of industrial / social modes breaks through the technical prejudice that high precision and fast response cannot be achieved at the same time, and semantic / gesture / expression three-mode data are fused through an intention confidence model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot interaction technology, specifically to a robot social interaction system and synchronization control method based on multimodal physical twins. Background Technology

[0002] With the continuous development of artificial intelligence and robotics, social robots, as an emerging interaction method, are receiving increasing attention and favor. The development of social robots aims to provide a more human-centered and intelligent interactive experience, offering better services to users. Currently, social robots are mainly used in entertainment, education, and healthcare, and have become an important part of human-computer interaction. However, existing social robots still have some problems in terms of emotional interaction capabilities. Most existing social robots are task-driven, only able to perform some predetermined interactive tasks, and their ability to recognize and express emotions is not yet perfected. To enhance the fun and intelligence of interaction, social robots need to possess certain emotional interaction capabilities, including recognizing the user's emotional state and being able to respond appropriately.

[0003] The existing system processes each sensor data stream independently without establishing a cross-modal association mechanism, resulting in an intention recognition error rate of >40% in noisy environments; the perception-decision-execution link latency is >250ms, causing the virtual twin's facial expression feedback to lag (measured ≥0.4s), leading to negative "mechanical" evaluations from users; industrial robots lack emotional interaction capabilities, while service robots cannot meet the requirements of precision operation, requiring manual reconfiguration for cross-scene switching; the packet loss rate of public networks in densely populated equipment areas is >12%, resulting in an asynchronous incident rate of virtual and real actions as high as 5.3 times per thousand hours. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a robot social interaction system and synchronization control method based on multimodal physical twins. It solves the problems of existing systems independently processing data streams from each sensor without establishing a cross-modal association mechanism, resulting in an intention recognition error rate >40% in noisy environments; a perception-decision-execution link latency >250ms, causing delayed facial expression feedback from the virtual twin (measured at ≥0.4s), leading to negative "mechanical" user feedback; industrial robots lacking emotional interaction capabilities, while service robots cannot meet the requirements of precision operation, necessitating manual reconfiguration for cross-scene switching; and public networks experiencing packet loss rates >12% in densely populated areas, resulting in a high rate of 5.3 incidents per thousand hours of asynchronous virtual-real actions.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a robot social interaction system based on multimodal physical twins, comprising: Physical twin modeling units are used to construct high-precision dynamic models of robots. The multimodal sensing unit integrates a visual sensor, LiDAR, and microphone array to collect environmental data in real time. The communication control unit uses the 5G NR protocol to achieve ultra-low latency multi-robot synchronization; The social decision engine generates interaction policies based on a reinforcement learning model. The policy function is: Where θ is the network weight, s t This is a multimodal fusion state; The virtual-real synchronization module maps physical actions to a virtual twin through digital threads.

[0006] Preferably, the communication control unit supports dynamic spectrum sharing technology to maintain communication reliability ≥99.9% in complex electromagnetic environments.

[0007] Preferably, the multimodal sensing unit fuses data using an adaptive weighting algorithm: Where σ i This represents the sensor noise variance.

[0008] Preferably, the social decision engine includes an emotion recognition submodule, which outputs emotion vectors through micro-expression analysis. .

[0009] Preferably, the emotion vector drives the facial rendering of the virtual twin, with an expression synchronization error of ≤0.1s.

[0010] Preferred options also include: The environment adaptation module automatically switches control strategies based on the scene type: Industrial mode: Prioritize ensuring motion accuracy; Social Mode: Prioritize optimizing the smoothness of emotional interaction.

[0011] Preferably, the social decision engine integrates an intent understanding submodule, which calculates the user intent confidence level through a semantic-gesture multimodal matching algorithm. Where Si is the semantic similarity, G is the gesture feature vector, and i is the weight coefficient. A synchronous control method for a robot social interaction system based on multimodal physical twins includes the following steps: Step 1: Generate a digital image of the robot using physical twin modeling units; Step 2: The multimodal sensing unit collects environmental data and transmits it to the communication control unit; Step 3: The social decision engine generates action instructions based on the improved Q-learning algorithm; Step 4: The virtual-real synchronization module verifies the consistency of actions. If the deviation exceeds the threshold δ... max A value of 0.05 triggers the error correction protocol.

[0012] Preferably, the action command generation in step three employs a dual-delay deep deterministic strategy gradient algorithm, with the update rule as follows: .

[0013] Preferably, the error correction protocol in step four includes: Retransmit control commands via a private 5G network with H04W 84 / 18 architecture; Activate predictive motion compensation for virtual twins.

[0014] This invention provides a robot social interaction system and synchronization control method based on multimodal physical twins. It has the following beneficial effects: 1. This invention integrates semantic, gesture, and facial expression data through an intent confidence model, maintaining a 96.5% recognition accuracy in a 90dB noise environment. The emotion vector drives micro-expression rendering, improving the naturalness score of the application.

[0015] 2. This invention uses 5G NR ultra-low latency transmission combined with timestamp chain verification to compress the synchronization error between virtual and real actions to ≤8.5ms. The error correction protocol maintains continuous system operation even with a 25% network packet loss rate.

[0016] 3. The environment adaptation module of this invention enables millisecond-level switching between industrial and social modes. Combined with dynamic spectrum sharing technology (H04W 16 / 14), it improves communication reliability to 99.94% and reduces packet loss rate to 0.06% in complex electromagnetic environments. Attached Figure Description

[0017] Figure 1 This is a system diagram of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] As one aspect of the present invention, please refer to the appendix. Figure 1 This invention provides a robot social interaction system based on multimodal physical twins, comprising: Physical twin modeling units are used to construct high-precision dynamic models of the robot, employing an adaptive weighted algorithm. Where σ i The sensor noise variance; The multimodal sensing unit integrates a visual sensor, LiDAR, and microphone array to collect environmental data in real time. The communication control unit adopts the 5G NR protocol to achieve ultra-low latency multi-robot synchronization, supports dynamic spectrum sharing technology, and maintains communication reliability of ≥99.9% in complex electromagnetic environments; The social decision engine integrates an intent understanding submodule, which calculates user intent confidence through a semantic-gesture multimodal matching algorithm. Where Si is the semantic similarity, G is the gesture feature vector, and i is the weight coefficient. The interaction strategy is generated based on a reinforcement learning model and includes an emotion recognition submodule, which outputs an emotion vector through micro-expression analysis. The facial rendering of the virtual twin is driven by emotion vectors, with an expression synchronization error of ≤0.1s. The strategy function is: Where θ is the network weight, s t This is a multimodal fusion state; The virtual-real synchronization module maps physical actions to a virtual twin through digital threads.

[0020] The environment adaptation module automatically switches control strategies based on the scene type: Industrial mode: Prioritize ensuring motion accuracy; Social Mode: Prioritize optimizing the smoothness of emotional interaction.

[0021] As another aspect of the present invention, a synchronization control method for a robot social interaction system based on multimodal physical twins includes the following steps: Step 1: Generate a digital image of the robot using physical twin modeling units; Step 2: The multimodal sensing unit collects environmental data and transmits it to the communication control unit; Step 3: The social decision engine generates action instructions based on the improved Q-learning algorithm. Action instruction generation employs a dual-delay deep deterministic policy gradient algorithm, with the following update rules: ; Step 4: The virtual-real synchronization module verifies the consistency of actions. If the deviation exceeds the threshold δ... max A value of 0.05 triggers the error correction protocol, which includes: Retransmit control commands via a private 5G network with H04W 84 / 18 architecture; Activate predictive motion compensation for virtual twins.

[0022] Detailed Implementation Examples Example 1: Educational Robot Scenario Configuration: Kindergarten classroom (background noise 68dB), equipped with NAO robot of this system; Interaction flow: The child points to the picture book and says, "Tell this story" (gesture simultaneously pointing to the cover); Multimodal sensing unit acquisition: Voice: Sentence Text → Storytelling = 0.91; Gesture: Pointing to vector G=[0.8,-0.2,0.4]→G 2 =0.94; Confidence level calculation: Cintent = 0.6 × 0.91 + 0.3 × 0.9max(0.94, 1) = 0.816; The social decision engine generates actions: a virtual twin smiles and nods (E=[0.7,0.9,0.2]), and a physical robot picks up a picture book and reads it aloud; Results: Intent recognition accuracy reached 98.7%, and children's attention span increased by 40%. Example 2: Configuration: Automotive assembly workshop (electromagnetic interference -90dBm), industrial robotic arms equipped with this system; Collaboration process: The environment adaptation module automatically switches to industrial mode; Worker's hand gesture command "tighten bolts" (G) 2 =2.1); The communication control unit transmits commands via a private 5G (H04W 84 / 18) with an end-to-end latency of 8ms. The virtual-real synchronization module detected a deviation of 0.03, and the robotic arm completed the high-precision tightening. Results: Assembly efficiency increased by 32%, and the error rate was 0 times per thousand hours; Example 3: Cross-scene switching Configuration: Medical logistics robot (hospital corridor → operating room preparation room) Adaptive process: Entering the corridor (high traffic): Switch to social mode When yielding to pedestrians, the virtual twin displays an "sorry" emoji (E=[-0.9,0.2,0.1]). Response latency: 182ms; Enter the surgical preparation room (sterile environment): Switch to industrial mode; Positional error of delivered surgical instruments: 0.2 mm; Mode switching time: 12ms; Results: Nurse satisfaction rating 4.9 / 5.0, instrument delivery error rate 0%.

[0023] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A robot social interaction system based on multimodal physical twins, characterized in that, include: Physical twin modeling units are used to construct high-precision dynamic models of robots. The multimodal sensing unit integrates a visual sensor, LiDAR, and microphone array to collect environmental data in real time. The communication control unit uses the 5G NR protocol to achieve ultra-low latency multi-robot synchronization; The social decision engine generates interaction policies based on a reinforcement learning model. The policy function is: Where θ is the network weight, s t This is a multimodal fusion state; The virtual-real synchronization module maps physical actions to a virtual twin through digital threads.

2. The robot social interaction system based on multimodal physical twins according to claim 1, characterized in that, The communication control unit supports dynamic spectrum sharing technology, maintaining communication reliability of ≥99.9% in complex electromagnetic environments.

3. The robot social interaction system based on multimodal physical twins according to claim 1, characterized in that, The multimodal sensing unit fuses data using an adaptive weighting algorithm: Where σ i This represents the sensor noise variance.

4. The robot social interaction system based on multimodal physical twins according to claim 1, characterized in that, The social decision engine includes an emotion recognition submodule, which outputs emotion vectors through micro-expression analysis. .

5. The robot social interaction system based on multimodal physical twins according to claim 4, characterized in that, The emotion vector drives the facial rendering of the virtual twin, with an expression synchronization error of ≤0.1s.

6. The robot social interaction system based on multimodal physical twins according to claim 1, characterized in that, Also includes: The environment adaptation module automatically switches control strategies based on the scene type: Industrial mode: Prioritize ensuring motion accuracy; Social Mode: Prioritize optimizing the smoothness of emotional interaction.

7. The synchronization control method for a robot social interaction system based on multimodal physical twins according to claim 1, characterized in that, The social decision engine integrates an intent understanding submodule, which calculates user intent confidence through a semantic-gesture multimodal matching algorithm. Where Si is the semantic similarity, G is the gesture feature vector, and i is the weight coefficient.

8. A synchronization control method for a robot social interaction system based on multimodal physical twins, using the robot social interaction system based on multimodal physical twins as described in any one of claims 1-7, characterized in that, Includes the following steps: Step 1: Generate a digital image of the robot using physical twin modeling units; Step 2: The multimodal sensing unit collects environmental data and transmits it to the communication control unit; Step 3: The social decision engine generates action instructions based on the improved Q-learning algorithm; Step 4: The virtual-real synchronization module verifies the consistency of actions. If the deviation exceeds the threshold δ... max A value of 0.05 triggers the error correction protocol.

9. The synchronization control method for a robot social interaction system based on multimodal physical twins according to claim 8, characterized in that, In step three, the action command generation adopts a dual-delay deep deterministic strategy gradient algorithm, and the update rule is as follows: 。 10. The synchronization control method for a robot social interaction system based on multimodal physical twins according to claim 8, characterized in that, The error correction protocol in step four includes: Retransmit control commands via a private 5G network with H04W 84 / 18 architecture; Activate predictive motion compensation for virtual twins.