Multi-part coordinated voice ai programming expression robot control system and method

By using technologies such as multi-core heterogeneous processors and AI acceleration units, a fully automated control chain is constructed, which solves the problems of complex programming and difficult debugging in existing facial expression robot systems. It realizes the automatic generation and execution of multi-part collaborative facial expressions under natural voice interaction, improving system reliability and development efficiency.

CN120941428BActive Publication Date: 2025-12-09深圳市小全科技文化有限公司

Patent Information

Application Number
CN202511480653.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2025-12-09
Estimated Expiration
2045-10-16

AI Technical Summary

Technical Problem

Existing facial expression robot systems suffer from high programming complexity, long debugging cycles, slow system response, and a lack of intuitive interaction methods in achieving natural, vivid, and multi-part coordinated emotional interaction. They are unable to generate and safely execute multi-part coordinated actions in real time through natural voice commands.

Method used

Employing a multi-core heterogeneous processor, hardware audio preprocessing, AI acceleration unit, secure storage area, and actuator drive circuit, it constructs a fully automated control chain through speech recognition, intent reasoning, and digital twin simulation modules, enabling the automatic generation and execution of natural voice interaction and multi-part collaborative facial expressions.

Benefits of technology

It realizes the automatic generation and execution of vivid and reliable multi-part collaborative facial expressions and actions based on natural voice interaction, which significantly reduces the programming threshold and debugging cycle, and improves development efficiency and system reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120941428B_ABST
    Figure CN120941428B_ABST
Patent Text Reader

Abstract

The application discloses a multi-part coordinated voice AI programming expression robot control system and method, the system comprises a master control unit, a voice processing module, an AI acceleration unit, a secure storage area, a debugging and program injection interface, an actuator driving circuit and a multi-stage safety hardware; the system extracts a structured control intention vector through voice recognition, combines an emotional action mapping library and physical constraints, and automatically generates or modifies control code by an AI programming agent; after digital twin simulation verification and automatic debugging optimization, the control code is safely burned into hardware and executed, realizing automatic generation and reliable execution of multi-part coordinated expression actions, significantly reducing programming complexity, and improving interaction naturalness and system safety.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robot control, in particular to a multi-part coordinated voice AI programming expression robot control system and method. BACKGROUND

[0002] With the deep integration of artificial intelligence and robot technology, bionic robots with emotional expression capabilities have become a frontier research direction in the field of human-computer interaction. However, existing expression robot systems still face many challenges in realizing natural, lively, and multi-part coordinated emotional interaction, such as high programming complexity, long debugging period, slow system response, and lack of intuitive interaction methods, which seriously restrict their popularization and application depth.

[0003] In the aspect of emotional understanding and decision-making, existing technologies have attempted to build the association between user emotions and robot behavior. The invention patent CN120015062A proposes a method of identifying user voice emotions based on a random forest model and deciding robot emotional behavior accordingly. This method effectively improves the emotional recognition rate in complex environments through a separate microphone array, sound source positioning, and multi-feature fusion, and builds a mapping relationship between emotional categories and robot behavior. However, the behavior output described in this method is heavily dependent on predefined action libraries and voice library templates. The decision result is actually the selection and invocation of pre-set templates, rather than the real-time generation of new, refined action control instructions based on emotional intensity parameters. This means that the system lacks true creativity and parameterized behavior generation capability, making it difficult to implement instructions such as "increase smile amplitude by 50%" or "slow blinking", which continuously and accurately control the amplitude and speed of actions. Its expression richness and adaptability have inherent limitations.

[0004] In the actuator layer, to achieve fine facial expressions, a high-degree of freedom bionic head structure design is crucial. The invention patent CN120552120A provides an advanced hardware solution for a multi-expression driven bionic intelligent robot head. This solution uses a layered, modular mechanical structure, integrates 26 servos, and uses a traction line and track mechanism cooperative transmission method to accurately control the movement of multiple parts such as eyeballs, eyelids, eyebrows, cheekbones, lips, and chin, achieving high-simulation multi-expression driving capability. This design effectively solves the problem of traditional robot head expression stiffness and lack of freedom. However, the core contribution of this patent is its precise mechanical structure and driving scheme, and it does not involve how to intelligently control the coordinated movement of the 26 servos to express specific emotions. Its control method still relies on pre-programmed sequences, lacking an upper-level intelligent control system that can receive high-level semantic instructions such as "nervously smile" and automatically calculate specific motion parameters, trajectories, and timing for the 26 servos.

[0005] Both of them do not solve the core problem of "how to make users generate and safely execute control codes that drive such complex mechanisms dynamically, accurately and automatically through the most natural voice interaction method". The existing control method still needs professional personnel to manually write the underlying code, and the debugging process is tedious and prone to errors, which cannot meet the interactive needs of real-time, online and personalized programming of robot expressions.

[0006] Therefore, there is an urgent need for a full-process control system that integrates advanced intent understanding, AI parameterized programming, virtual simulation verification and automated security deployment, to bridge the gap between high-level emotional instructions and fine control of underlying multi-actuator, ultimately reduce the programming threshold, improve development efficiency and system reliability, and realize lively and natural multi-part coordinated emotional expression. SUMMARY

[0007] The first object of the present application is to provide a multi-part coordinated voice AI programming expression robot control system, which aims to solve the problems of existing expression robot programming complexity, debugging difficulty, slow response and lack of natural interaction means, and cannot realize lively, safe and multi-part coordinated emotional expression.

[0008] To solve the above technical problems, a multi-part coordinated voice AI programming expression robot control system is provided, which comprises a main control unit, a voice processing module, an AI acceleration unit, a secure storage area, an actuator driving circuit and a debugging interface. The main control unit is a multi-core heterogeneous processor, which integrates application processing cores for running AI algorithms and operating systems and real-time control cores for real-time parallel control of multi-channel actuators. The voice processing module integrates a hardware audio preprocessing circuit for noise reduction, echo cancellation and sound source positioning at the hardware level, and outputs a pure voice signal. The AI acceleration unit is used to accelerate the physical calculation in the voice recognition model, the key information extraction model, the intent reasoning module and the digital twin simulation module. The secure storage area is used to store verified system programs, user configuration files, emotional action mapping libraries and system recovery images. The actuator driving circuit is a multi-channel driving circuit that communicates with the main control unit through a high real-time bus and is used to drive the multi-part actuators of the robot. The output end of the voice processing module is connected to the input end of the AI acceleration unit. The output end of the AI acceleration unit is connected to the application processing core of the main control unit. The output end of the real-time control core of the main control unit is connected to the input end of the actuator driving circuit. The secure storage area is bidirectionally connected to the main control unit. The main control unit is connected to the outside through the debugging and program injection interface.

[0009] In one of the embodiments, the pure voice signal output by the voice processing module is converted into text by the voice recognition model running on the AI acceleration unit, and the structured control intention vector containing action type, part, intensity, speed, and frequency is extracted from the text by the key information extraction model; the intention reasoning module running on the AI acceleration unit is used to receive and process the structured control intention vector, and output the control instruction parameter for simulation verification by the digital twin simulation module.

[0010] In one of the embodiments, the emotional action mapping library stored in the secure storage area is used to map the structured control intention vector to specific actuator parameters; the emotional action mapping library is constructed by combining expert rules and data-driven methods, and the emotional action mapping library contains emotional instructions, expression parameters, associated body movements, and intensity coding.

[0011] In one of the embodiments, an AI programming agent runs on the master control unit, the AI programming agent adopts a retrieval enhancement generation architecture, receives the structured control intention vector and the current state information of the robot, and automatically generates or modifies executable control code by querying the emotional action mapping library and the physical constraint library; the digital twin simulation module running on the AI acceleration unit is used to simulate and verify the generated control code.

[0012] In one of the embodiments, the digital twin simulation module is connected to the AI debugging agent, if the simulation fails, the AI debugging agent starts a multi-objective optimization algorithm to fine-tune and optimize the parameters or logic of the control code, and starts iterative simulation until the verification is passed.

[0013] In one of the embodiments, the system adopts A / B dual firmware partitioning and a timer, the new program is burned to the B area for running after verification, and the A area backs up the old version of the system program; if the running is abnormal, the timer triggers the system to automatically roll back to the stable version in the A area.

[0014] In one of the embodiments, the multi-part actuator includes a facial actuator, a neck actuator, a torso actuator, and an arm actuator.

[0015] The second object of the present application is to provide a multi-part collaborative voice AI programming expression robot control method, which aims to solve the problem that the existing robot expression control method relies on pre-programming and cannot generate and safely execute multi-part collaborative actions in real time through natural voice instructions.

[0016] To solve the above technical problems, a multi-part collaborative voice AI programming expression robot control method is provided, which is applied to the above-mentioned multi-part collaborative voice AI programming expression robot control system, and includes the following steps:

[0017] S1, voice instruction capture and semantic understanding: preprocess the voice signal into the pure voice signal through the voice processing module, and process the pure voice signal through the voice recognition model and the key information extraction model to output a structured control intention vector;

[0018] S2, AI-driven parameterized programming and code generation: an AI programming agent queries an emotional action mapping library and a physical constraint library according to the structured control intention vector to automatically generate or modify executable control code;

[0019] S3, virtual simulation and automated debugging: the generated control code is simulated and verified through the digital twin simulation module, the control code is debugged and optimized through the AI debugging agent, and the control code is iterated until it passes;

[0020] S4, secure burning and execution: the control code that passes the verification is securely burned into the main control unit, and the state of the actuator is monitored in real time during the execution process and a closed-loop feedback is formed.

[0021] In one embodiment, the AI-driven parameterized programming and code generation further includes that the AI programming agent uses retrieval enhancement generation technology to retrieve a pre-stored code template according to an intention reasoning result, and combines parameters to modify and combine, the generation process is embedded with safety rules and the emotional restoration degree of the generated result is evaluated.

[0022] In one embodiment, the secure burning and execution further includes that a verification code comparison is performed after the burning to ensure the integrity, and the actuator state data collected in real time during the execution process is fed back to the AI programming agent for evaluating the execution effect and optimizing subsequent instructions.

[0023] Implementing the embodiments of the present application will have the following beneficial effects:

[0024] 1. In the multi-part collaborative voice AI programming facial expression robot control system of this embodiment, the user issues commands through natural speech. The speech processing module performs hardware-level preprocessing on the microphone signal and outputs clean speech to the AI ​​acceleration unit. The AI ​​acceleration unit converts the speech into text through a speech recognition model, and the key information extraction model extracts a structured control intent vector containing action parameters from the text. The intent reasoning module processes the vector and outputs control parameters. The AI ​​programming agent queries the emotion-action mapping library and the physical constraint library based on these parameters and automatically generates or modifies the control code. The digital twin simulation module performs pre-simulation verification of the code, and the AI ​​debugging agent automatically debugs and optimizes it in the simulation environment, forming a closed-loop iteration. After the verified security code is burned through the A / B partitioning mechanism, it is controlled in real time by the main control unit. The core drives multi-channel actuator circuits via a high real-time bus to coordinate the control of facial, neck, torso, and arm actuators to complete emotional expression. Compared with existing technologies, this solution combines multi-core heterogeneous computing, hardware AI acceleration, retrieval-enhanced generation (RAG) technology, digital twin simulation, and multiple security mechanisms to construct a fully automated control chain of "voice → intent understanding → AI programming → simulation verification → automatic debugging → secure burning → execution feedback." This enables the automatic generation and execution of vivid and reliable multi-part coordinated facial expressions based on natural voice interaction. By using an AI agent to understand the subtle differences in user commands and parameterize them to specific actuator actions, the barrier to entry and debugging cycle for robot facial expression programming are significantly reduced, fundamentally improving development efficiency, system reliability, and the naturalness of human-computer interaction.

[0025] 2. The multi-part collaborative voice AI programming facial expression robot control method in this embodiment captures and preprocesses voice commands into clean signals through the voice processing module. After being recognized and extracted by the AI ​​acceleration unit, a structured control intent vector is generated. The AI ​​programming agent automatically generates or modifies the control code based on the vector. The digital twin module performs simulation verification on the code, and the AI ​​debugging agent automatically debugs and optimizes until it passes. Finally, the safety code is burned into the main control unit for execution, forming a closed-loop feedback, realizing fully automatic and intelligent control from natural voice input to safe output of multi-part collaborative actions. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1The multi-part coordinated voice AI programming expression robot control system overall hardware architecture block diagram of embodiment one of the present application is described;

[0028] Figure 2 The multi-part coordinated voice AI programming expression robot control method flow chart of embodiment two of the present application is described;

[0029] Figure 3 The multi-part coordinated voice AI programming expression robot multi-part coordinated control schematic diagram of embodiment one of the present application is described;

[0030] Figure 4 The multi-part coordinated voice AI programming expression robot emotional action mapping library conceptual schematic diagram of embodiment one of the present application is described;

[0031] Figure 5 The multi-part coordinated voice AI programming expression robot control method AI programming agent workflow schematic diagram of embodiment two of the present application is described;

[0032] Figure 6 The multi-part coordinated voice AI programming expression robot control method digital twin simulation and debugging interface schematic diagram of embodiment two of the present application is described;

[0033] Figure 7 The multi-part coordinated voice AI programming expression robot control system safety burning and rollback mechanism flow chart of embodiment one of the present application is described. DETAILED DESCRIPTION

[0034] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the accompanying drawings. The preferred embodiments of the present application are given in the accompanying drawings. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.

[0035] It should be noted that when an element is referred to as being "fixed" to another element, it can be directly on the other element or there can be an intervening element. When an element is referred to as being "connected" to another element, it can be directly connected to the other element or intervening elements can be present. The terms "vertical", "horizontal", "left", "right", and the like as used herein are for purposes of illustration only.

[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terminology used in the description of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0037] Embodiment one

[0038] With reference to Figure 1 , Figure 3 , Figure 4 and Figure 7 , embodiment one of the present application provides a multi-part coordinated voice AI programming expression robot control system.

[0039] With reference to Figure 1 , a multi-part coordinated voice AI programming expression robot control system for an embodiment includes a master control unit, a voice processing module, an AI acceleration unit, a secure storage area, an actuator driving circuit, and a debugging and program injection interface.

[0040] The master control unit is a multi-core heterogeneous processor integrating application processing cores (such as ARM Cortex A series) and real-time control cores (such as ARM Cortex M series). The application processing cores are used to run high-level AI algorithms and operating systems (such as Linux); the real-time control cores are used to generate high-precision, multi-channel actuator real-time parallel control instructions to ensure strict synchronization timing of the movement of each actuator.

[0041] The voice processing module integrates hardware audio preprocessing circuit (such as dedicated DSP chip and supporting ADC) for hardware-level noise reduction (NR), echo cancellation (AEC) and sound source positioning (DOA) on the microphone array signal, outputting high signal-to-noise ratio pure voice signal.

[0042] The AI acceleration unit (such as integrated NPU or FPGA) is used to accelerate the running of voice recognition models, key information extraction models, intent reasoning modules, and digital twin simulation modules.

[0043] The secure storage area (adopting a protected area divided by eMMC or SPIFlash) is used to store verified system programs, user configuration files, emotional action mapping libraries, and system recovery images.

[0044] The actuator driving circuit is a multi-channel, high-precision driving circuit (such as a servo steering engine driving board) that communicates with the master control unit through a high real-time bus (such as EtherCAT or CAN / FD) and is used to drive the robot multi-part actuators (including face, neck, torso, arm actuators).

[0045] The output end of the voice processing module is connected to the input end of the AI acceleration unit; the output end of the AI acceleration unit is connected to the application processing core of the master control unit; the real-time control core output end of the master control unit is connected to the input end of the actuator driving circuit; the secure storage area is bidirectionally connected to the master control unit; the master control unit is connected to the external debugging device through the debugging and program injection interface (such as SWD / JTAG).

[0046] Implementing the first embodiment of the present application will have the following beneficial effects: when the multi-part coordinated voice AI programmed expression robot control system of the present scheme is working, the user issues instructions through natural voice (such as "exaggeratedly laughing and waving hands"), the voice processing module outputs pure voice signals to the AI acceleration unit after hardware-level preprocessing of the microphone signals; the AI acceleration unit first converts the pure voice signals into text through the lightweight voice recognition model (such as based on CNNCTC architecture) it runs; then, the key information extraction model (such as based on lightweight BERT architecture) it runs extracts a structured control intent vector from the text, which contains multi-dimensional parameters such as action type (laugh, wave hands), part (face, arm), intensity (exaggerated), speed (default), and number of times (1 time); the vector is then passed to the intent reasoning module for processing.

[0047] Please refer to Figure 4 After receiving the vector, the intent reasoning module queries the emotion action mapping library in the secure storage area. The mapping library is constructed by combining expert rules and motion capture data learning, which not only maps high-level emotional instructions (such as "exaggeratedly laughing") to specific actuator parameters (such as "mouth corner servo angle increased to 25°, eye muscle servo contraction 40%"), but also supports parsing and mapping the fine parameters contained in the structured control intent vector, such as travel, speed, execution time, angle direction, and return flag. For example, the user can achieve precise and parameterized control of the actuator through specific voice instructions such as "left eyelid raised 5mm at a rate of 2mm per second, maintain for 3 seconds and then fall back" or "nose down 2mm, last for 1 second". "Wave hands" is mapped to "shoulder servo up 30°, elbow servo reciprocating motion 2 times".

[0048] Please refer to Figure 5, the AI programming agent (using the Retrieval Augmentation Generation (RAG) architecture) running on the main control unit is activated. It receives the structured parameters from the intent reasoning module and the current state sensor data of the robot, queries the emotional action mapping library and the physical constraint library (such as joint angle limits, maximum speed), retrieves the action primitive code templates for "laughing" and "waving" from the local knowledge base, and modifies the displacement and speed values in the templates according to the "exaggeration" intensity parameter to automatically generate a new, executable multi-part coordinated control code. The multi-part actuators include facial (such as eyelids, nose, mouth), neck, torso, and arm actuators. Among them, the neck actuator can perform rotation (such as turning the head left and right), pitching (such as nodding, lifting the head), and other actions to express rich emotional scenarios, such as when receiving negative voice instructions, the head is controlled to make a low bow or turn away; the eye actuator supports independent or coordinated movement of the eyeball in eight directions (left, right, up, down, left up, left down, right up, right down) to simulate natural eye-to-eye interaction.

[0049] Please refer to Figure 6 Before burning to hardware, the generated code is simulated and verified by the digital twin simulation module (based on a simplified version of the embedded physics engine such as Bullet) running on the AI acceleration unit. The simulation monitors in real time whether collisions occur between virtual actuators, whether the motion trajectory is feasible, and whether any physical constraints are violated. If the simulation reports an error (such as detecting a hand waving trajectory interfering with the head), the connected AI debugging agent will start a multi-objective optimization algorithm (such as Bayesian optimization) to automatically fine-tune the code parameters (such as reducing the arm waving amplitude or modifying the torso tilt angle), and start iterative simulation until verification is passed.

[0050] Please refer to Figure 7 The final code that passes the verification is safely burned to a specific memory area (B zone) of the main control unit through the debugging interface. The system uses A / B dual firmware partitioning and a hardware watchdog timer. After burning is complete, the hardware calculates the checksum (such as CRC32) of the code and compares it with the value generated by the simulation end to ensure integrity. The new program runs in the B zone, and the old version is backed up in the A zone. If the hardware watchdog timer detects that the new program is running abnormally (such as actuator overcurrent or program deadlock), it will immediately trigger a hardware-level rollback to the stable version in the A zone to ensure system safety.

[0051] Finally, the real-time control core of the main control unit executes the program and sends precise parallel control instructions to the actuator drive circuit through a high real-time bus to drive various actuators (such as Figure 3 ) on the face, neck, torso, and arms to perform coordinated movements, thereby achieving lively and natural multi-expression emotional expression. At the same time, the real-time state data (position, current, temperature) of the actuators are continuously collected and fed back to the AI programming agent, forming a closed loop for evaluating execution effectiveness and optimizing subsequent instructions.

[0052] Compared with the prior art, the scheme combines multi-core heterogeneous computing, hardware AI acceleration, retrieval augmented generation (RAG) technology, digital twin simulation, and multiple security mechanisms, and constructs a full-process automatic control link of “voice→intention understanding→AI programming→simulation verification→automatic debugging→secure burning→execution feedback”, so as to realize the automatic generation and execution of natural voice interaction-based, lively and reliable multi-site coordinated expression actions, significantly reduce the threshold and debugging period of robot expression programming, and fundamentally improve the development efficiency, system reliability, and naturalness of human-computer interaction.

[0053] It should be noted that the emotional action mapping library is not fixed and can be updated and optimized through data-driven learning (such as motion capture data recording real person performances), so that the robot's expression is more rich and human-like.

[0054] In addition, it should also be noted that the simulation accuracy of the digital twin simulation module can be configured as needed, from fast rough simulation to high-precision detailed simulation, to balance between debugging speed and accuracy.

[0055] Referring to Figure 1 In an optional embodiment, the hardware audio preprocessing circuit integrated in the voice processing module includes a special DSP chip (such as a domestic CI1122), which has a built-in hardware-accelerated neural network processor for offline running of a lightweight speech recognition model, further reducing system latency and protecting user privacy.

[0056] According to actual needs, the NPU (Neural Processing Unit) used in the AI acceleration unit supports INT8 quantization calculation, which can significantly reduce the computing power consumption while ensuring the recognition accuracy, and is suitable for embedded mobile platforms.

[0057] In addition, in still another optional embodiment, the high real-time bus preferably uses the EtherCAT bus, which has a microsecond-level synchronization accuracy, can ensure the strict synchronization of the movements of dozens of actuators distributed in the robot's head, neck, torso, and arms, avoid dislocation and lag of actions, and is the key to realizing natural coordinated movement.

[0058] Embodiment Two

[0059] The multi-site coordinated voice AI programming expression robot control method in this embodiment two is different from the subject matter protected by the control system in embodiment one, and the specific differences are as follows:

[0060] Please refer to Figure 2 and Figure 5For another embodiment of the multi-part coordinated voice AI programming expression robot control method, applied to the system described in embodiment one, including the steps of:

[0061] S1, voice instruction capture and semantic understanding: capture the voice signal through the voice processing module and perform hardware level preprocessing to obtain a pure voice signal; the pure voice signal is converted into text through the voice recognition model running in the AI acceleration unit, and then processed through the key information extraction model to output a structured control intent vector;

[0062] S2, AI-driven parameterized programming and code generation: the AI programming agent queries the emotional action mapping library and the physical constraint library according to the control intent vector, automatically retrieves the relevant code templates using retrieval enhancement generation technology, and modifies and combines them with strength, speed and other parameters to generate new executable control codes. The generation process embeds safety rules and performs emotional restoration degree evaluation;

[0063] S3, virtual simulation and automatic debugging: the generated control code is simulated and verified through the digital twin simulation module to monitor collision, trajectory feasibility and boundary conditions; if the simulation reports an error, the AI debugging agent starts the optimization algorithm to fine-tune the code parameters or logic, and iteratively simulates until the verification is passed;

[0064] S4, safe burning and execution: the verified control code is safely burned into the main control unit, and the verification code is compared after burning to ensure integrity; the main control unit executes the program to drive the multi-part actuators to complete the coordinated action; the actuator state data is collected in real time during execution and fed back to the AI programming agent for evaluation of execution effect and optimization of subsequent instructions.

[0065] By implementing the second embodiment of the present application, the following beneficial effects are achieved: the voice instruction is captured and preprocessed by the voice processing module into a pure signal, and the structured control intent vector is generated after recognition and extraction by the AI acceleration unit; the AI programming agent automatically generates or modifies the control code according to the vector; the digital twin module simulates and verifies the code, and the AI debugging agent automatically debugs and optimizes until it passes; finally, the safe code is burned into the main control unit for execution, and a closed-loop feedback is formed, realizing full-automatic and intelligent control from natural voice input to multi-part coordinated action safe output.

[0066] Referring to Figure 2In an optional embodiment, in AI-driven parameterized programming and code generation, the AI programming agent adopts the Retrieval Augmentation Generation (RAG) technology, and a large number of pre-verified action primitives (Skill Primitives) code templates (such as C++ functions, Python script blocks) are stored in the local knowledge base of the agent. The agent first retrieves the most relevant template according to the intention, and then modifies and combines according to the parameters, greatly improving the efficiency and reliability of code generation.

[0067] In yet another embodiment, a multi-part coordinated voice AI programming expression robot control system further includes host computer debugging software connected with the master control unit through a debugging interface, which can be used for real-time monitoring of system status, visual display of digital twin simulation process, manual intervention debugging, updating of emotion action mapping library, and management of A / B partition firmware, providing convenience for development and maintenance.

Claims

1. A multi-part coordinated voice AI programmed expression robotic control system, characterized by, include: The main control unit is a multi-core heterogeneous processor that integrates an application processing core for running AI algorithms and operating systems and a real-time control core for real-time parallel control of multi-channel actuators. The voice processing module integrates a hardware audio preprocessing circuit for hardware-level noise reduction, echo cancellation, and sound source localization of microphone signals, outputting a clean voice signal. The AI ​​acceleration unit is used to accelerate the physical computation in the speech recognition model, key information extraction model, intent reasoning module, and digital twin simulation module. The secure storage area is used to store verified system programs, user configuration files, emotion and action mapping libraries, and system recovery images. An actuator drive circuit, which is a multi-channel drive circuit, communicates with the main control unit through a high real-time bus and is used to drive actuators in multiple parts of the robot. The output of the voice processing module is connected to the input of the AI ​​acceleration unit; the output of the AI ​​acceleration unit is connected to the application processing core of the main control unit; the output of the real-time control core of the main control unit is connected to the input of the actuator drive circuit; the secure storage area is bidirectionally connected to the main control unit; the main control unit is connected to the outside through the debugging and program injection interface.

2. The multi-partite synergic voice AI-programmed expressive robotic control system according to claim 1, wherein, The clean speech signal output by the speech processing module is converted into text by the speech recognition model running by the AI ​​acceleration unit, and the key information extraction model extracts a structured control intent vector containing action type, location, intensity, speed, and frequency from the text. The intent reasoning module running by the AI ​​acceleration unit receives and processes the structured control intent vector and outputs control command parameters for simulation verification by the digital twin simulation module.

3. The multi-partite synergic voice AI-programmed expressive robotic control system according to claim 2, wherein, The emotion-action mapping library stored in the secure storage area is used to map the structured control intent vector to specific actuator parameters. The emotion-action mapping library is constructed by combining expert rules and data-driven methods, and includes emotion commands, facial expression parameters, associated body movements, and intensity codes.

4. The multi-partite synergic voice AI-programmed expressive robotic control system according to claim 3, wherein, The main control unit runs an AI programming agent, which adopts a retrieval-enhanced generation architecture. It receives the structured control intent vector and the robot's current state information, and automatically generates or modifies executable control code by querying the emotion-action mapping library and the physical constraint library. The AI ​​acceleration unit runs a digital twin simulation module to simulate and verify the generated control code.

5. The multi-partite synergic voice AI-programmed expressive robotic control system according to claim 4, wherein, The digital twin simulation module is connected to the AI ​​debugging agent. If the simulation fails, the AI ​​debugging agent starts a multi-objective optimization algorithm to fine-tune and optimize the parameters or logic of the control code, and starts iterative simulation until the verification is successful.

6. The multi-partite, coordinated voice-AI-programmed, expressive robotic control system of claim 1, wherein, The multi-part collaborative voice AI programming facial expression robot control system adopts A / B dual firmware partitioning and timers. After the new program is verified, it is burned to partition B for operation, while partition A backs up the old version of the system program. If the operation is abnormal, the system will automatically roll back to the stable version in partition A through the timer.

7. The multi-partite synergic voice AI-programmed expressive robotic control system according to claim 1, wherein, The multi-site actuators include a face actuator, a neck actuator, a torso actuator and an arm actuator. 8.A method for controlling a multi-part coordinated voice AI programming expression robot, applied to the multi-part coordinated voice AI programming expression robot control system according to any one of claims 1-7, characterized in that, The method comprises the steps of: S1, voice instruction capture and semantic understanding: the voice processing module preprocesses the voice signal into a pure voice signal, the pure voice signal is processed by the voice recognition model, the feature extraction model and the key information extraction model, and a structured control intention vector is output; S2, AI-driven parameterized programming and code generation: the AI programming agent queries the emotion action mapping library and the physical constraint library according to the structured control intention vector, and automatically generates or modifies executable control code; S3, virtual simulation and automatic debugging: the control code is simulated and verified by the digital twin simulation module, the control code is debugged and optimized by the AI debugging agent, and the control code is iterated until it is passed; S4, secure burning and execution: the control code that passes the verification is securely burned into the main control unit, and the actuator state is monitored in real time during the execution process and a closed-loop feedback is formed.

9. The multi-partite synergic voice AI programming expression robot control method according to claim 8, wherein, The AI-driven parameterized programming and code generation further comprises that the AI programming agent adopts retrieval enhancement generation technology, retrieves the pre-stored code template according to the intention reasoning result, and modifies and combines the parameters, generates the process embedded with safety rules and evaluates the emotion restoration degree of the generated result.

10. The multi-partite synergic voice AI programming expression robot control method according to claim 8, wherein, The secure burning and execution further comprises that after the burning is completed, the check code is compared, and the actuator state data collected in real time during the execution process is fed back to the AI programming agent for evaluating the execution effect and optimizing the subsequent instructions.

Citation Information

Patent Citations

  • Emotional behavior expression method of humanoid robot

    CN120015062A

  • Multi-expression driving bionic intelligent robot head

    CN120552120A

  • Humanoid-head robot device with human-computer interaction function and behavior control method thereof

    CN101618280A

  • Semantic-based behavior generation method

    CN112052688A

Cited By

  • Bionic facial expression robot and a single-chip-based method for adapting to various refined facial expression control techniques

    CN122500754A