A digital human entity robot system and an interaction driving method thereof
By designing a digital human physical robot system, and using perception units and control modules to coordinate the generation of voice, facial expressions and body movements, the problem of lack of body movements and stiff expressions in digital human interaction is solved, thereby improving the human-computer interaction effect and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI UNIV
- Filing Date
- 2023-08-17
- Publication Date
- 2026-04-21
AI Technical Summary
Existing digital humans lack physical interaction and have stiff facial expressions when interacting with users, resulting in simple interaction dimensions and poor user experience.
Design a digital human physical robot system, comprising a sensing unit, a control module, a communication unit, a storage unit, and a mechanical unit. The sensing unit acquires user information, and the control module coordinates the robot to generate voice, facial expressions, and body movements to achieve rich contact-based interaction.
By using body movements and rich facial expressions, the human-computer interaction effect is enhanced, the user experience is improved, and an emotional resonance with the user is achieved.
Smart Images

Figure CN117021131B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robotics, specifically to a digital human physical robot system and its interactive driving method. Background Technology
[0002] A digital human is a virtual entity created through artificial intelligence technology that possesses human characteristics and abilities. Digital humans can simulate human language, emotions, and behavior, and even have the ability to learn and adapt to their environment. With the continuous advancement of artificial intelligence technology, digital humans have important applications in many fields.
[0003] Digital humans possess powerful computing capabilities and intelligent thinking, enabling them to perform complex calculations and reasoning. Through interaction with vast amounts of data, digital humans can continuously learn and evolve, thereby acquiring increasingly higher levels of intelligence. Compared to traditional robots, digital humans are closer to human intelligence, better able to understand and respond to human needs.
[0004] Current digital humans have the advantages of high realism and strong customizability, but the way they interact with people is mainly through non-contact methods such as voice interaction, text interaction, and visual interaction. They lack contact interaction methods such as body interaction and touch interaction to convey richer information and intentions, resulting in insufficient realism in the user experience.
[0005] In existing technologies, most digital humans are limited to presenting images, sounds, and facial expressions. When providing services, their facial expressions are stiff and lack variety during the interaction process. The interaction is simple and lacks physical interaction with the user, making it impossible to create emotional resonance with humans. This results in poor human-computer interaction and a poor user experience. Summary of the Invention
[0006] Therefore, the technical problem that this invention aims to solve is that most existing digital humans are limited to non-contact interaction that presents images, sounds, and facial expressions. The interaction is simple in dimension, without physical interaction with the user, and the facial expressions during the interaction are rigid and lack variety when providing services, resulting in poor human-computer interaction and a poor user experience.
[0007] To address the aforementioned technical problems, this invention provides a digital human physical robot system, comprising a sensing unit, a control module, a communication unit, a storage unit, and a power connection module, as well as a mechanical unit. The mechanical unit includes a mechanical structure comprising a main body, extendable legs, extendable feet, extendable robotic arms, and a rotatable humanoid face shell. The sensing unit, control module, communication unit, storage unit, and power connection module are all mounted on the main body.
[0008] The sensing unit is used by the robot to sense the external environment and obtain instruction information issued by the interactive object;
[0009] The communication unit receives the information sent by the sensing unit and sends the information to the control module;
[0010] The storage unit is used to store the operating system, control program, perception algorithm, data and code required for core functions, perception data, user interaction records and maps;
[0011] The power connection module is electrically connected to a power source to supply power to the sensing unit, the control module, the communication unit, and the storage unit.
[0012] The control module includes a calling component, a sensor input component, a control component, and an execution component. Based on the information sent by the communication unit, the control module selects to run or execute the software program stored in the storage unit, and selects to retrieve the data stored in the storage unit, and controls the execution component to perform corresponding actions, so that the robot makes one or more of the following reactions: making a sound, making a limb movement, and making a facial expression.
[0013] An interactive driving method for a digital human-like robot system comprises three main parts:
[0014] 1. The robot emits a voice.
[0015] When the perception unit detects speech, facial expressions, and body movements from an interactive object, it will convert the facial expression information and body movement information into facial expression-speech text and body movement-speech text, respectively.
[0016] 1. Call the component to access the voice and text content file of the sensor input component;
[0017] 2. The voice text content file is sent to the text-to-speech conversion component, which converts the voice text content file into the voice signal that the robot needs to broadcast;
[0018] 3. The voice signal is transmitted to the voice-facial expression / body movement conversion component, which converts the voice signal into facial expression signals and body movement signals that are consistent with the voice signal;
[0019] 4. The voice signal is also transmitted to the voice generation component to generate the voice that the robot needs to play;
[0020] 5. The voice is transmitted to the audio output component to play the robot's voice;
[0021] 6. Facial expression signals and body movement signals are transmitted to the facial expression generation component and the body movement generation component, respectively, so that the robot can make facial expressions and body movements that are consistent with the speech.
[0022] II. Robots making facial expressions
[0023] When the perception unit detects speech, facial expressions, and body movements from an interactive object, it will convert the speech information and body movement information into speech-facial expression description files and body movement-facial expression description files, respectively.
[0024] 1. Call the component to access the facial expression description file of the sensor input component;
[0025] 2. The facial expression description file is sent to the facial expression recognition component, which then converts the facial expression description file into the expression signals that the robot needs to make.
[0026] 3. The facial expression signal is transmitted to the facial expression-voice / body movement conversion component, which converts the facial expression signal into a voice signal and body movement signal that are consistent with the facial expression signal;
[0027] 4. The expression signal is also transmitted to the facial expression generation component, thereby generating the facial expressions that the robot needs to project;
[0028] 5. Facial expressions are transmitted to the facial expression projection component, which then projects the facial expressions onto the humanoid face shell.
[0029] 6. The voice signal and body movement signal are transmitted to the voice generation component and the body movement generation component respectively, so that the robot can play voice that is consistent with facial expressions and make body movements that are consistent with facial expressions.
[0030] III. The robot performs limb movements.
[0031] When the perception unit detects that the interactive object emits voice, facial expressions, and body movements, it will convert the voice information and facial expression information into voice-body movement data files and facial expression-body movement data files, respectively.
[0032] 1. Call the component to input the limb motion data file from the sensor input component;
[0033] 2. The limb motion data file is transmitted to the limb motion recognition component, which then converts the limb motion data file into the limb motion signals that the robot needs to perform.
[0034] 3. The body movement signal is transmitted to the body movement-speech / facial expression conversion component, which converts the body movement signal into a speech signal and facial expression signal that are consistent with the body movement signal;
[0035] 4. The limb motion signals are also transmitted to the limb motion generation component, thereby generating the motion control sequence that the robot needs to execute;
[0036] 5. The motion control sequence is transmitted to the motion execution component, which then sends the motion control sequence to the leg, the foot, and the robotic arm to execute the motion.
[0037] 6. The voice signal and facial expression signal are transmitted to the voice generation component and the facial expression generation component respectively, so that the robot can play voice that is coordinated with the body movements and make facial expressions that are coordinated with the body movements.
[0038] The above three parts are executed in parallel.
[0039] The technical solution of the present invention has the following beneficial effects:
[0040] This invention provides a digital human physical robot system and its interactive driving method. By converting the external environmental information perceived by the sensing unit and the command information issued by the interactive object—one or more of the following: actions, expressions, and voice—into voice signals, facial expression signals, and body movement signals that are consistent with the command information, the robot's feedback forms become richer and more three-dimensional. It is no longer limited to text, voice, and visual interactions, but adds a function of contact interaction—body movement interaction. By configuring the digital human with a physical mechanical structure, the robot can make body movements, such as extending its arm to shake hands or opening its arms to hug, enabling contact interaction with the interactive object and generating emotional resonance with humans. Moreover, the robot's expressions can vary richly according to the various forms of command information issued by the interactive object, avoiding the rigidity of digital human expressions and greatly improving the effect of human-computer interaction and user experience. Attached Figure Description
[0041] To more clearly illustrate the specific embodiments of the present invention, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0042] Figure 1 These are schematic diagrams of the various parts of the robot system of the present invention;
[0043] Figure 2 A schematic diagram of a digital human physical robot system and its interactive driving method provided by an embodiment of the present invention;
[0044] Figure 3 This is a schematic diagram of the limb movement recognition component proposed in an embodiment of the present invention.
[0045] 101. Mechanical Unit; 1013. Mechanical Structure; 102. Sensing Unit; 103. Communication Unit; 105. Storage Unit; 110. Control Module; 111. Power Connection Module; C11. Voice Text Content File; C12. Facial Expression Description File; C13. Body Movement Data File; C21. Text-to-Speech Conversion Component; C31. Facial Expression Recognition Component; C41. Body Movement Recognition Component; C22. Voice-to-Facial Expression / Body Movement Conversion Component; C32. Facial Expression-to-Voice / Body Movement Conversion Component; C42. Body Movement-to-Voice / Facial Expression Conversion Component; C23. Voice Generation Component; C33. Facial Expression... The components are: Emotion Generation Component; C43, Body Movement Generation Component; C50, Audio Output Component; C60, Facial Expression Projection Component; C70, Action Execution Component; D11, Data Stream 1; D12, Data Stream 2; D13, Data Stream 3; D21, Data Stream 4; D22, Data Stream 5; D23, Data Stream 6; D24, Data Stream 7; D31, Data Stream 8; D32, Data Stream 9; D33, Data Stream 10; D34, Data Stream 11; D41, Data Stream 12; D42, Data Stream 13; D43, Data Stream 14; D44, Data Stream 15; D51, Audio File; D61, Expression File; D71, Action Control Sequence Data Stream. Detailed Implementation
[0046] The technical solution of the present invention will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of, and not all of, the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0047] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or component referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0048] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0049] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0050] The present invention provides a digital human physical robot system, including a sensing unit 102, a control module 110, a communication unit 103, a storage unit 105, a power connection module 111, and a mechanical unit 101.
[0051] The mechanical unit 101 includes a mechanical structure 1013, which comprises a main body, extendable legs, extendable feet, extendable robotic arms, and a rotatable humanoid face shell. The sensing unit 102, control module 110, storage unit 105, communication unit 103, and power connection module 111 are all mounted on the main body. It should be noted that the various component modules of the mechanical unit 101 can be one or multiple, depending on the specific situation; for example, the legs can have four components.
[0052] The sensing unit 102 is used by the robot to perceive the external environment and acquire instruction information issued by the interactive object. When interacting with the user, the sensing unit 102 can perceive the user's actions, postures, and behaviors through various sensors. The sensing unit 102 includes, but is not limited to, cameras, microphones, depth sensors, inertial sensors, contact sensors, tactile sensors such as electronic skin, and other sensors.
[0053] The communication unit 103 receives information sent by the sensing unit 102 and sends the information to the control module 110. Instructions issued by the interactive object, such as through touch, voice commands or other gestures, are captured and recognized by the sensing unit 102, such as through a touch sensor or microphone, and are transmitted to the control module 110 of the robot system for processing through the communication unit 103.
[0054] The storage unit 105 mainly includes a program storage unit and a data storage unit. The program storage unit is used to store the operating system, control program, perception algorithm and other data and code required for core functions. The data storage unit is used to record perception data, store user interaction records, save maps, etc. When the robot interacts with the interactive object, the robot can walk autonomously based on the map information.
[0055] The power connection module 111 is electrically connected to a power source and is used to supply power to various components. The power connection module 111 may include a battery and a power control board, which controls battery charging, discharging, and power consumption management. The power connection module 111 is electrically connected to the control module 110, and can also be electrically connected to sensing units 102, such as cameras, radar, speakers, and motors. It should be noted that each component can be connected to a different power connection module 111, or be powered by the same power connection module 111.
[0056] The control module 110 is the control center of the robot. It connects all components of the robot through various interfaces and lines. By running or executing software programs stored in the storage unit 105 and calling data stored in the storage unit 105, it controls the interaction between the robot and the user. The control module 110 includes a calling component, a sensor input component, a control component, and an execution component. Based on the information sent by the communication unit 103, the control module 110 selects to run or execute software programs stored in the storage unit 105 and selects to call data stored in the storage unit 105, controlling the execution component to perform corresponding actions, causing the robot to make one or more of the following reactions: make a sound, make a body movement, and make a facial expression.
[0057] Specifically, the interactive driving methods for robots are as follows:
[0058] If the sensing unit 102 senses one or all of the information from the voice, facial expression and body movement of the interactive object, the robot will perform the corresponding interactive action.
[0059] The robot first calls the interactive object information recording file by calling the component. The interactive object information recording file includes voice text content file C11, facial expression description file C12, body movement data file C13, etc., and requests to send it to the control component. It can be a disk file, a memory file, or a special file such as a real-time data stream. This invention does not limit the specific form of the file.
[0060] The voice text content file C11 contains the user's voice text information received by the microphone; the facial expression description file C12 contains the user's facial expression information received by the camera, which may include the user's smiling, frowning, and other facial expressions; and C13 is the body movement data file, which contains the user's body movement data information, including both contact and non-contact movements, received by touch sensors, tactile sensors such as electronic skin, and other sensors.
[0061] The robot's control components can be divided into three parts: voice processing, facial expression processing, and body movement processing.
[0062] The speech component is responsible for generating the speech the robot needs to play, as well as facial expressions and body movements that are coordinated with the speech. This component takes a speech text file, facial expression-speech signals, and body movement-speech signals as input, and an audio file, speech-expression signals, and speech-body movement signals as output. This component includes a text-to-speech conversion component C21, a speech-to-facial expression / body movement conversion component C22, and a speech generation component C23. The text-to-speech conversion component C21 converts the speech text content to be played by the robot into a speech signal. This component takes a speech text file as input and outputs a speech signal. The speech-to-facial expression / body movement conversion component C22 generates facial expression signals and body movement signals that are coordinated with the speech based on the speech text content. This component takes a speech signal as input and outputs expression signals and body movement signals. The speech generation component C23 generates the speech the robot needs to play, taking a speech signal as input and an audio file as output.
[0063] The facial expression module is responsible for generating the facial expressions that the robot needs to project, as well as the corresponding speech and body movements. This module takes a facial expression description file, speech-expression signals, and body movement-expression signals as input, and outputs the expression file, facial expression-speech signals, and facial expression-body movement signals. This module includes a facial expression recognition component C31, a facial expression-speech / body movement conversion component C32, and a facial expression generation component C33. The facial expression recognition component C31 analyzes the user's facial expressions and generates corresponding expression signals, taking a facial expression description file as input and the expression signals as output. The facial expression-speech / body movement conversion component C32 is responsible for generating speech and body movement signals that are consistent with the expression signals. This component takes the expression signals as input and outputs speech and body movement signals. The facial expression generation component C33 is responsible for generating the facial expressions that the robot needs to project, taking the expression signals as input and the expression file as output.
[0064] The limb motion module is responsible for generating the actions the robot needs to perform. It takes limb motion data files, voice-limb motion signals, and facial expression-limb motion signals as input, and outputs motion control sequences, limb motion signal-voice signals, and limb motion signal-facial expression signals. This module includes a limb motion recognition component C41, a limb motion-voice / facial expression conversion component C42, and a limb motion generation component C43. The limb motion recognition component C41 is mainly responsible for analyzing and recognizing the user's intention in limb movements, converting the limb motion data file into limb motion signals. This component takes the limb motion data file as input and outputs limb motion signals. The limb motion-voice / facial expression conversion component C42 is responsible for generating voice and facial expression signals that are consistent with the limb motion data information. This component takes the limb motion signals as input and outputs expression signals and facial expression signals. The limb motion generation component C43 is responsible for generating the actions the robot needs to perform, taking limb motion signals as input and outputting motion control sequences.
[0065] The robot's execution components include an audio output component C50, a facial expression projection component C60, and a motion execution component C70. The audio output component C50 outputs the robot's voice signal, taking the voice signal as input and sound as output. This component contains an audio player C51, which parses the voice signal and plays it as sound. The facial expression projection component C60 outputs the robot's facial expression signal, taking the expression signal as input and projecting it as output. This component contains a projector C61, which projects the signal onto the humanoid face shell. The motion execution component C70 executes the actions the user wants the robot to perform.
[0066] like Figure 3 As shown, the limb movement recognition component can include three modules: a feature extraction module C412, a movement intention recognition module C413, and a movement generation module C414. The feature extraction module C412 is responsible for extracting key feature information from the limb movement data. This module takes the limb movement data file D13 as input and outputs the key movement information data structure D131. The movement intention recognition module C413 is responsible for comparing and classifying the extracted features with a pre-trained model to identify the user's movement intention. This module takes the key movement information data structure D131 as input and outputs the movement intention information structure D132. The movement generation module C414 is responsible for generating corresponding limb movement information from the identified movement intention. This module takes the movement intention information structure D132 as input and outputs the limb movement signal D41.
[0067] According to different needs, embodiments of the present invention propose a robot-user interaction-driven process. The order of certain steps in this process can be changed, wherein S25, S26, S35, S36, S45, and S46 can be omitted. The process may include the following steps:
[0068] S11. The robot reads, but is not limited to, the speech text content file C11, the facial expression description file C12, and the body movement data file C13. In order to execute the robot's speech generation process, facial expression projection process, and movement generation process concurrently, steps S21, S31, and S41 are executed concurrently after this step is completed.
[0069] S21. The robot transmits the voice and text content to the text-to-speech conversion component C21 through data stream D11. Data stream D11 is an encapsulation of the voice and text content file C11.
[0070] S22. Text-to-speech conversion component C21 executes the text-to-speech conversion logic, generating speech signal data stream 4D21 and data stream 5D22. Data stream 4D21 and data stream 5D22 contain the same content. In order to concurrently execute the robot's audio output process, facial expression projection process, and limb movement recognition and execution process, after this step is completed, it synchronously waits for the completion of steps S35 and S46, and concurrently executes steps S23 and S24.
[0071] S23, the speech generation component C23 encapsulates the speech signals D21, D33, and D44 generated by the text-to-speech conversion component C21, the facial expression-to-speech / body movement conversion component C32, and the body movement-to-speech / facial expression conversion component C42 into an audio file D51, and transmits it to the audio output component C50. After this step is completed, step S51 is then executed.
[0072] S24. Text-to-speech conversion component C21 transmits the generated speech signal data stream D22 to speech-to-facial expression / body movement conversion component C22. In order to concurrently execute the robot's audio output process, facial expression projection process, and body movement recognition and execution process, steps S25 and S26 are executed concurrently after this step is completed.
[0073] S25, Voice-to-facial expression / body movement conversion component C22 executes voice-to-facial expression conversion logic, generating data stream 6D23. Data stream 6D23 is the expression signal obtained by processing the voice text content file C11. After this step is completed, it waits for the completion of steps S32 and S45, and then executes step S33.
[0074] S26. The voice-to-facial expression / body movement conversion component C22 executes the voice-to-body movement conversion logic and generates data stream 7D24. Data stream 7D24 is the body movement signal obtained by processing the voice text content file C11. After this step is completed, it waits for the completion of steps S36 and S42, and then executes step S43.
[0075] S31. The robot transmits the facial expression description file C12 to the facial expression recognition component C31 through data stream 2D12. Data stream 2D12 is a wrapper around the facial expression description file C12.
[0076] S32, Facial Expression Recognition Component C31 executes facial expression recognition logic, generating expression signal data stream 8D31 and data stream 9D32. Data stream 8D31 and data stream 9D32 have the same content. In order to concurrently execute the robot's audio output process, facial expression projection process, and limb movement recognition and execution process, after this step is completed, it synchronously waits for the completion of steps S25 and S45, and concurrently executes steps S33 and S34.
[0077] S33, the facial expression generation component C33 encapsulates the expression signals D23, D31, and D43 generated by the speech-to-facial expression / body movement conversion component C22, the facial expression recognition component C31, and the body movement-to-speech / facial expression conversion component C42 into an expression file D61, and transmits it to the facial expression projection component C60. Then, step S61 is executed.
[0078] S34. The facial expression recognition component C31 transmits the generated expression signal data stream D32 to the facial expression-voice / body movement conversion component C32. In order to concurrently execute the robot's audio output process, facial expression projection process, and body movement recognition and execution process, steps S35 and S36 are executed concurrently after this step is completed.
[0079] S35. Facial expression-voice / body movement conversion component C32 executes facial expression-voice conversion logic and generates data stream + D33. Data stream + D33 is the expression signal obtained by processing the facial expression description file C12. After this step is completed, it waits for the completion of steps S22 and S46, and then executes step S23.
[0080] S36. Facial expression-voice / body movement conversion component C32 executes facial expression-body movement conversion logic and generates data stream 11D34. Data stream 11D34 is the body movement signal obtained by processing the facial expression description file C12. After this step is completed, it waits for the completion of steps S26 and S42, and then executes step S43.
[0081] S41. Transmit the limb motion data file C13 to the limb motion recognition component C41 via data stream 3D13. Here, data stream 3D13 is a wrapper around the limb motion data file C13.
[0082] S42, the limb movement recognition component C41 executes the motion intention recognition logic, generating limb movement signal data streams 12D41 and 13D42, with identical content. To concurrently execute the robot's audio output process, facial expression projection process, and limb movement recognition and execution process, this step waits synchronously for the completion of steps S26 and S36 before concurrently executing steps S43 and S44.
[0083] S43, the body movement generation component C43 encapsulates the body movement signal data stream D41 generated by the voice-to-facial expression / body movement conversion component C22, the facial expression-to-voice / body movement conversion component C32, and the body movement recognition component C41 into a motion control sequence data stream D71, and transmits it to the motion execution component C70. After this step is completed, step S71 is executed.
[0084] S44, the limb movement recognition component C41 transmits the generated limb movement signal data stream D42 to the limb movement-voice / facial expression conversion component C42. In order to concurrently execute the robot's audio output process, facial expression projection process, and limb movement recognition and execution process, steps S45 and S46 are executed concurrently after this step is completed.
[0085] S45, Body Movement-Voice / Facial Expression Conversion Component C42 executes the body movement-facial expression conversion logic, generating data stream 14D43. Data stream 14D43 is the facial expression signal obtained from generating body movement data file C13. After this step is completed, it synchronously waits for the completion of steps S25 and S32, and then executes step S33.
[0086] S46, Body movement-voice / facial expression conversion component C42 executes body movement-voice conversion logic, generating data stream 15D44. Data stream 15D44 is the voice signal obtained from generating body movement data file C13. After this step is completed, it waits for the completion of steps S22 and S35, and then executes step S23.
[0087] S51: The audio output component C50 parses the audio file D51 and plays the robot's voice through the audio player C51. At this point, the robot's voice generation process is complete.
[0088] S61, the facial expression projection component C60 parses the expression file D61 and projects it onto the robot's humanoid face shell via the projector C61. At this point, the robot's facial expression recognition and projection process is complete.
[0089] In this embodiment of the invention, the user, as the interactive object, conveys commands to the robot via voice. A camera is used to capture the user's image information, and a microphone is used to capture the user's voice information. The image and voice information are transmitted to the control module 110 for processing. After processing, the control module 110 transmits the corresponding image signal to a projector, which projects the robot's facial expression image onto the humanoid face shell. The corresponding sound signal is then transmitted to an audio player for playback.
[0090] S71, the motion execution component C70 parses the motion control sequence data stream D71 and realizes the robot's limb movement. At this point, the robot's limb motion recognition and execution process is completed.
[0091] In this embodiment of the invention, by integrating multiple sensing technologies, the robot can effectively receive the user's motion command information, then preprocess and extract features from the received motion command information, and use machine learning algorithms to compare and classify the extracted features with a pre-trained intention model to determine the user's intention. By controlling the movement and posture of the upper limbs, the robot can make corresponding feedback and perform actions, such as extending its arm to shake hands or opening its arms to hug.
[0092] It should be further explained that, in the above embodiments, the robot interacts with the user using the language and actions of a specific person by recognizing and processing input information, and possesses the ability to flexibly adapt to interactions with different users and non-specific persons. The robot analyzes the communication data between the user and the specific person, including text, voice, and video, to acquire the specific person's language and behavioral habits. It learns their body language, gestures, etc., and simulates unique movement characteristics, enabling the robot to more realistically represent the specific person's behavioral style. By integrating this information, the robot achieves a comprehensive understanding and simulation of the specific person's personality traits.
[0093] Using the above technical solution, the robot, through the recognition and processing of input information, achieves the simulation of a real person's language style and behavioral actions. Its functions include, but are not limited to, voice and motion recognition and processing, facial expression image analysis and projection, and the execution of limb movements. Users can interact with the robot through voice commands, facial expressions, gestures, and touch, allowing the robot to understand their intentions, control its upper and lower limbs to perform corresponding actions and movements, and achieve physical contact actions such as shaking hands, hugging, or patting. This multifunctional robot enhances emotional communication and provides users with an intelligent, personalized, and realistic interactive experience.
[0094] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. An interactive driving method for a digital human physical robot system, applied to a digital physical robot system, comprising a sensing unit (102), a control module (110), a communication unit (103), a storage unit (105), and a power connection module (111), characterized in that, It includes three main parts:
1. The robot emits a voice. If the perception unit (102) perceives that the interactive object emits speech, facial expression and body movement information, it will convert the facial expression information and body movement information into facial expression-speech text and body movement-speech text. 1) Call the component to call the voice text content file (C11) of the sensor input component. 2) The speech text content file (C11) is sent to the text-to-speech conversion component (C21) to convert the speech text content file (C11) into the speech signal that the robot needs to broadcast; 3) The speech signal is transmitted to the speech-facial expression / body movement conversion component (C22), which converts the speech signal into facial expression signals and body movement signals that are consistent with the speech signal; 4) The voice signal is also transmitted to the voice generation component (C23) to generate the voice that the robot needs to play; 5) The voice is transmitted to the audio output component (C50) to play the robot's voice; 6) Facial expression signals and body movement signals are transmitted to the facial expression generation component (C33) and body movement generation component (C43) respectively, so that the robot can make facial expressions and body movements that are consistent with the speech. II. Robots making facial expressions If the perception unit (102) perceives that the interactive object emits voice, facial expression and body movement information, it will convert the voice information and body movement information into voice-facial expression description file and body movement-facial expression description file respectively. 1) Call the component to call the facial expression description file (C12) of the sensor input component; 2) The facial expression description file (C12) is transmitted to the facial expression recognition component (C31), which converts the facial expression description file (C12) into the expression signal that the robot needs to make; 3) The facial expression signal is transmitted to the facial expression-voice / body movement conversion component (C32), which converts the facial expression signal into a voice signal and body movement signal that are consistent with the facial expression signal; 4) The facial expression signal is also transmitted to the facial expression generation component (C33) to generate the facial expression that the robot needs to project; 5) Facial expressions are transmitted to the facial expression projection component (C60), which projects the facial expressions onto the humanoid face shell; 6) The voice signal and the body movement signal are transmitted to the voice generation component (C23) and the body movement generation component (C43) respectively, so that the robot can play voices that are consistent with facial expressions and make body movements that are consistent with facial expressions. III. The robot performs limb movements. If the perception unit (102) perceives that the interactive object emits voice, facial expression and body movement information, it will convert the voice information and facial expression information into voice-body movement data file and facial expression-body movement data file respectively. 1) Call the component to call the limb motion data file (C13) of the sensor input component. 2) The limb motion data file (C13) is transmitted to the limb motion recognition component (C41), which converts the limb motion data file (C13) into the limb motion signals that the robot needs to perform. 3) The body movement signal is transmitted to the body movement-speech / facial expression conversion component (C42), which converts the body movement signal into a speech signal and facial expression signal that are consistent with the body movement signal; 4) The limb motion signal is also transmitted to the limb motion generation component (C43) to generate the motion control sequence that the robot needs to execute; 5) The motion control sequence is transmitted to the motion execution component (C70), which then sends the motion control sequence to the legs, feet, and robotic arm to execute the motion. 6) The voice signal and facial expression signal are transmitted to the voice generation component (C23) and the facial expression generation component (C33) respectively, so that the robot can play voice that is coordinated with the body movements and make facial expressions that are coordinated with the body movements. The above three parts are executed in parallel; Includes the following steps: S11. The robot reads, but is not limited to, voice text content files (C11), facial expression description files (C12), and body movement data files (C13). In order to execute the robot's voice generation process, facial expression projection process, and movement generation process concurrently, steps S21, S31, and S41 are executed concurrently after this step is completed. S21. The robot transmits the speech and text content to the text-to-speech conversion component (C21) through data stream one (D11). Data stream one (D11) is an encapsulation of the speech and text content file (C11). S22, the text-to-speech conversion component (C21) executes the text-to-speech conversion logic, generating speech signal data stream four (D21) and data stream five (D22). The content of data stream four (D21) and data stream five (D22) is the same. In order to concurrently execute the robot's audio output process, facial expression projection process and limb action recognition and execution process, after this step is completed, it waits for the completion of steps S35 and S46, and concurrently executes steps S23 and S24. S23, the speech generation component (C23) encapsulates the data stream four (D21), data stream ten (D33), and data stream fifteen (D44) generated by the text-to-speech conversion component (C21), facial expression-to-speech / body movement conversion component (C32), and body movement-to-speech / facial expression conversion component (C42) into an audio file (D51) and transmits it to the audio output component (C50); after this step is completed, step S51 is then executed; S24. The text-to-speech conversion component (C21) transmits the generated speech signal data stream five (D22) to the speech-to-facial expression / body movement conversion component (C22). In order to concurrently execute the robot's audio output process, facial expression projection process, and body movement recognition and execution process, steps S25 and S26 are executed concurrently after this step is completed. S25, the voice-to-facial expression / body movement conversion component (C22) executes the voice-to-facial expression conversion logic and generates data stream six (D23). Data stream six (D23) is the expression signal obtained by processing the voice text content file (C11). After this step is completed, it waits for the completion of steps S32 and S45, and then executes step S33. S26. The voice-to-facial expression / body movement conversion component (C22) executes the voice-to-body movement conversion logic and generates data stream seven (D24). Data stream seven (D24) is the body movement signal obtained by processing the voice text content file (C11). After this step is completed, it waits for the completion of steps S36 and S42, and then executes step S43. S31. The robot transmits the facial expression description file (C12) to the facial expression recognition component (C31) through data stream two (D12). Data stream two (D12) is a wrapper for the facial expression description file (C12). S32, the facial expression recognition component (C31) executes the facial expression recognition logic, generating expression signal data stream eight (D31) and data stream nine (D32). The content of data stream eight (D31) and data stream nine (D32) is the same. In order to concurrently execute the robot's audio output process, facial expression projection process and limb action recognition and execution process, after this step is completed, it waits for the completion of steps S25 and S45, and concurrently executes steps S33 and S34. S33, the facial expression generation component (C33) encapsulates the facial expression signal data streams six (D23), eight (D31), and fourteen (D43) generated by the speech-facial expression / body movement conversion component (C22), the facial expression recognition component (C31), and the body movement-speech / facial expression conversion component (C42) into an expression file (D61), and transmits it to the facial expression projection component (C60). Then, step S61 is executed. S34. The facial expression recognition component (C31) transmits the generated expression signal data stream nine (D32) to the facial expression-voice / body movement conversion component (C32). In order to concurrently execute the robot's audio output process, facial expression projection process, and body movement recognition and execution process, steps S35 and S36 are executed concurrently after this step is completed. S35. The facial expression-speech / body movement conversion component (C32) executes the facial expression-speech conversion logic and generates data stream 10 (D33). Data stream 10 (D33) is the expression signal obtained by processing the facial expression description file (C12). After this step is completed, it waits for the completion of steps S22 and S46, and then executes step S23. S36. The facial expression-voice / body movement conversion component (C32) executes the facial expression-body movement conversion logic and generates data stream eleven (D34). Data stream eleven (D34) is the body movement signal obtained by processing the facial expression description file (C12). After this step is completed, it waits for the completion of steps S26 and S42, and then executes step S43. S41. Transmit the limb motion data file (C13) to the limb motion recognition component (C41) via data stream three (D13); where data stream three (D13) is a wrapper around the limb motion data file (C13); S42, the limb movement recognition component (C41) executes the motion intention recognition logic, generating limb movement signal data stream twelve (D41) and data stream thirteen (D42). The contents of data stream twelve (D41) and data stream thirteen (D42) are the same. In order to concurrently execute the robot's audio output process, facial expression projection process and limb movement recognition and execution process, after this step is completed, it waits for the completion of steps S26 and S36, and concurrently executes steps S43 and S44. S43, the body movement generation component (C43) encapsulates the body movement signal data stream twelve (D41) generated by the speech-facial expression / body movement conversion component (C22), the facial expression-speech / body movement conversion component (C32), and the body movement recognition component (C41) into a movement control sequence data stream (D71), and transmits it to the movement execution component (C70); after this step is completed, step S71 is executed; S44, The limb movement recognition component (C41) transmits the generated limb movement signal data stream thirteen (D42) to the limb movement-voice / facial expression conversion component (C42); In order to concurrently execute the robot's audio output process, facial expression projection process and limb movement recognition and execution process, steps S45 and S46 are executed concurrently after this step is completed; S45, Body Movement - Voice / Facial Expression Conversion Component (C42) executes body movement - facial expression conversion logic and generates data stream fourteen (D43). Data stream fourteen (D43) is the facial expression signal obtained from generating body movement data file (C13). After this step is completed, it waits for the completion of steps S25 and S32, and then executes step S33. S46, Body Movement-Voice / Facial Expression Conversion Component (C42) executes body movement-voice conversion logic and generates data stream fifteen (D44). Data stream fifteen (D44) is the voice signal obtained from generating body movement data file (C13). After this step is completed, it waits for the completion of steps S22 and S35, and then executes step S23. S51: The audio output component (C50) parses the audio file (D51) and plays the robot's voice through the audio player (C51); at this point, the robot's voice generation process is complete. S61, The facial expression projection component (C60) parses the expression file (D61) and projects it onto the robot's humanoid face shell through the projector (C61); at this point, the robot's facial expression recognition and projection process is completed. S71, the motion execution component (C70) parses the motion control sequence data stream (D71) and realizes the robot's limb movement. At this point, the robot's limb motion recognition and execution process is completed.
2. The interactive driving method for a digital human physical robot system as described in claim 1, characterized in that, The digital human physical robot system also includes a mechanical unit (101), which includes a mechanical structure (1013). The mechanical structure (1013) includes a main body, extendable legs, extendable feet, extendable robotic arms, and a rotatable humanoid face shell. The sensing unit (102), the control module (110), the communication unit (103), the storage unit (105), and the power connection module (111) are all installed on the main body. The sensing unit (102) is used by the robot to sense the external environment and obtain instruction information issued by the interactive object; The communication unit (103) receives the information sent by the sensing unit (102) and sends the information to the control module (110). The storage unit (105) is used to store the operating system, control program, perception algorithm, data and code required for core functions, perception data, user interaction records and maps; The power connection module (111) is electrically connected to the power supply to provide power to the sensing unit (102), the control module (110), the communication unit (103) and the storage unit (105); The control module (110) includes a calling component, a sensor input component, a control component, and an execution component. The control module (110) selects to run or execute the software program stored in the storage unit (105) and selects to access the data stored in the storage unit (105) based on the information sent by the communication unit (103). It controls the execution component to perform corresponding actions, so that the robot makes a sound, performs a limb movement, and makes a facial expression or one or more of the following reactions.
Citation Information
Patent Citations
Home-based care robot capable of conducting outdoor movable accompanying
CN107322593A
Intelligent robot
CN109421044A
Virtual digital human driving method and system
CN116206023A