Methods, apparatus, media, and program product for synthesizing robotic works
By combining user stories and timelines, the robot's actions, voices, and skills can be freely combined, solving the problem of high complexity in humanoid robot development. It provides a low-cost and flexible method for generating robot creations, lowers the technical threshold for embodied applications, and supports the development and practice of ordinary users.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI MATRIX SUPER INTELLIGENT SYSTEM INTEGRATION CO LTD
- Filing Date
- 2025-12-09
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, humanoid robots have stiff joint movements, insufficient fine manipulation capabilities, are prone to losing balance, and have high development costs. There is a technical barrier between professional-level development and public exploration, and non-professional users cannot participate in the expansion of embodied applications.
By organizing user stories and combining them with a timeline, the robot's actions, sounds, and skills can be freely combined. The robot creations are generated based on the timeline selected by the user, which lowers the development threshold, provides a zero-code development platform, and supports ordinary users to freely combine and customize robot creations.
It enables low-cost and flexible development of robot creations, lowers the technical threshold for embodied applications, and allows ordinary users to easily explore and practice with robots. It simplifies traditional complex processes and builds a universal embodied application development platform.
Smart Images

Figure CN121340369B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics, and more particularly to a technique for creating synthetic robotic works. Background Technology
[0002] In existing technologies, compared to the high degree of freedom and flexibility of human joints, humanoid robot joints, which employ a "rigid actuator + reducer" technical solution, often suffer from stiff movements and insufficient fine manipulation capabilities. Furthermore, humanoid robots are prone to loss of balance under sudden external forces or in complex terrain, posing a risk of tipping over. Moreover, their bodies are mostly bulky, energy-intensive, and have limited range of motion, falling far short of the public's desire for general intelligence—a disconnect between robot hardware, algorithmic complexity, and user understanding. Current machine AI (Artificial Intelligence) model training is limited to simulations or specific scenarios, resulting in decreased performance when transferred to real-world, complex environments. Customized development for complex, changing customer scenarios is prohibitively expensive, highlighting a disconnect between AI maturity and complex scenarios. Therefore, developing embodied applications based on humanoid robots often requires expensive motion capture equipment and specialized developers, leading to a lack of interoperability between professional development tools and public exploration, creating technical barriers. This prevents non-professional users from participating in application expansion and hinders the formation of an open, collaborative ecosystem. Summary of the Invention
[0003] One object of this application is to provide a method, apparatus, medium, and program product for creating synthetic robotic works.
[0004] According to one aspect of this application, a method for creating synthetic robotic works is provided, the method comprising:
[0005] In response to a user's selection action on a timeline, obtain a combination of robot skills selected by the user for at least one point in time on the timeline, wherein the combination of robot skills includes at least one of robot actions, robot voices, and robot skills;
[0006] Based on the robot skill combination corresponding to the at least one time point, a corresponding robot creation is synthesized, such that the humanoid robot deployed with the robot creation executes the robot skill combination at the at least one time point.
[0007] According to one aspect of this application, a computer device for creating synthetic robotic works is provided, comprising a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of any of the methods described above.
[0008] According to one aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of any of the methods described above.
[0009] According to one aspect of this application, a computer program product is provided, comprising a computer program, characterized in that, when executed by a processor, the computer program implements the steps of any of the methods described above.
[0010] According to one aspect of this application, a computer device for creating synthetic robotic works is provided, the device comprising:
[0011] The module is used to respond to a user's selection operation on the timeline and obtain the robot skill combination selected by the user for at least one point in time on the timeline, wherein the robot skill combination includes at least one of robot actions, robot voice, and robot skills.
[0012] The first and second modules are used to synthesize a corresponding robot creation based on the robot skill combination corresponding to the at least one time point, so that the humanoid robot deployed with the robot creation executes the robot skill combination at the at least one time point.
[0013] Compared with existing technologies, this application enables the free combination of robot actions, robot voices, and robot skills through a user story organization method. The user stories are organized along a timeline, allowing for flexible customization of robot user stories and ultimately generating robot works. This allows for low-cost customization of robot work development and process control based on business scenarios, lowers the technical threshold for embodied application development, enables zero-code development of robot embodied applications, and builds a universal embodied application development platform for the general public. It also lowers the technical threshold for embodied application exploration, simplifies and visualizes the complex process of traditional robot application development, and allows ordinary users who do not need professional programming skills to explore and practice robots with maximum freedom and ease. Attached Figure Description
[0014] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0015] Figure 1 This diagram illustrates a method for synthesizing a robotic work according to one embodiment of the present application;
[0016] Figure 2 A schematic diagram of a robot work according to one embodiment of this application is shown;
[0017] Figure 3This diagram illustrates a computer device structure for creating a synthetic robot, according to one embodiment of this application.
[0018] Figure 4 Exemplary systems that can be used to implement the various embodiments described in this application are shown.
[0019] The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation
[0020] The present application will now be described in further detail with reference to the accompanying drawings.
[0021] In a typical configuration of this application, the terminal, the device of the service network, and the trusted party all include one or more processors (e.g., a central processing unit (CPU)), input / output interfaces, network interfaces, and memory.
[0022] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory. Memory is an example of computer-readable media.
[0023] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PCM), programmable random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0024] The devices referred to in this application include, but are not limited to, user equipment, network equipment, or devices composed of user equipment and network equipment integrated through a network. The user equipment includes, but is not limited to, any mobile electronic product capable of human-computer interaction (e.g., via a touchpad), such as smartphones and tablets. These mobile electronic products can use any operating system, such as Android or iOS. The network equipment includes an electronic device capable of automatically performing numerical calculations and information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and embedded devices. The network equipment includes, but is not limited to, computers, network hosts, single network servers, multiple network server clusters, or clouds composed of multiple servers. Here, a cloud consists of a large number of computers or network servers based on cloud computing, where cloud computing is a type of distributed computing, consisting of a virtual supercomputer composed of a group of loosely coupled computer clusters. The network includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, VPN network, wireless ad hoc network, etc. Preferably, the device can also be a program running on the user equipment, network device, or a device formed by integrating user equipment and network device, network device, touch terminal, or network device and touch terminal through a network.
[0025] Of course, those skilled in the art should understand that the above-described devices are merely examples, and other existing or future devices that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.
[0026] In the description of this application, "multiple" means two or more, unless otherwise expressly and specifically defined.
[0027] Figure 1The diagram illustrates a method for synthesizing a robotic creation according to an embodiment of this application. The method includes steps S11 and S12. In step S11, a computer device, in response to a user's selection operation on a timeline, obtains a combination of robot skills selected by the user for at least one point in time on the timeline. The combination of robot skills includes at least one of robot actions, robot voice, and robot skills. In step S12, the computer device synthesizes a corresponding robotic creation based on the robot skill combination corresponding to the at least one point in time, such that a humanoid robot deployed with the robotic creation performs the robot skill combination at the at least one point in time.
[0028] In step S11, the computer device, in response to a selection operation performed by the user on the timeline, obtains a combination of robot skills selected by the user for at least one point in time on the timeline, wherein the combination of robot skills includes at least one of robot actions, robot voices, and robot skills.
[0029] In some embodiments, a timeline is presented to the user on the user device. The user can select at least one time point on the timeline and then select the robot skill combination corresponding to each selected time point. That is, each time point corresponds to a robot skill combination. For example, the user can drag a robot skill combination to the position or near the selected time point on the timeline and use it as the robot skill combination corresponding to that time point.
[0030] In some embodiments, the robot skill set includes at least one of robot actions, robot voice, and robot skills. In some embodiments, robot actions are primarily achieved through video recording. By capturing videos of human actions, joint mapping and redirection from human actions to the robot body are performed in the cloud, ultimately achieving the imitation and capture of human actions to generate corresponding robot actions. In some embodiments, robot voices are primarily achieved through TTS (Text-to-Speech) methods. By selecting timbre, persona, and emotion, human voices are automatically synthesized based on a TTS model to generate corresponding robot voices. In some embodiments, robot skills include at least one of robot navigation, robot grasping, robot lighting, and robot facial expressions. In some embodiments, the user can select the specific content included in each robot skill set. For example, a robot skill set may contain only one of robot actions, robot voice, and robot skills, or it may contain multiple of these three elements, meaning the user can freely combine robot actions, robot voice, and robot skills.
[0031] This application enables an exploration method based on user stories. User stories are organized along a timeline, allowing users to freely combine and flexibly customize exploration materials, including robot actions and sounds, and exploration skills (i.e., robot skills) in a user story-based manner, ultimately enabling the development of exploration works (i.e. robot works) adapted to specific needs.
[0032] In step S12, the computer device synthesizes a corresponding robot creation based on the robot skill combination corresponding to the at least one time point, so that the humanoid robot to which the robot creation is deployed executes the robot skill combination at the at least one time point.
[0033] In some embodiments, a corresponding robot creation is synthesized based on at least one time point selected by the user on the timeline and a robot skill combination selected by the user for each time point. That is, the robot creation includes at least one time point on the timeline and a robot skill combination corresponding to each of the at least one time point. In some embodiments, the user can deploy the synthesized robot creation in a humanoid robot, causing the deployed humanoid robot to execute the robot skill combination corresponding to each of the at least one time point. Executing the robot skill combination includes at least one of performing robot actions, playing robot sounds, and releasing robot skills.
[0034] In some embodiments, before deploying a robot creation, it is necessary to first determine the motion control strategy used for the robot's actions in the robot creation. For simple robot actions, the general motion control strategy of the humanoid robot body can be directly reused to achieve the tracking and imitation of any action and ensure the balance of motion control of any action. For complex robot actions, it is necessary to train a special motion control strategy for the complex action to achieve the tracking and imitation of the complex action and ensure the balance of motion control of the complex action.
[0035] In some embodiments, before deploying a robot creation, a unified training and simulation platform is used to automatically plan GPU resources, train the model, conduct simulation tests, and evaluate the model's movements. This entire process is transparent to the user. Only after completing the full simulation test can the user deploy the robot creation on a humanoid robot. In some embodiments, GPU computing resources are automatically allocated for the robot's movements in the creation, and a model training container is automatically created. This container mainly includes a simulation environment for training the model (i.e., motion control strategy). Then, by using reinforcement learning and imitation learning algorithms, the robot's movements are imitated. Automatic model evaluation is then implemented. When the model performance meets predetermined requirements, the user is notified. The user can then check whether the model meets expectations through the simulation environment. If it does, the robot creation can be deployed on a real robot (humanoid robot).
[0036] This application utilizes a user story organization method to enable the free combination of robot actions, robot voices, and robot skills. The user stories are organized along a timeline, allowing for flexible customization of robot user stories and ultimately generating robot works. This approach allows for low-cost customization of robot work development and process control based on business scenarios, lowers the technical threshold for embodied application development, enables zero-code development of robot embodied applications, and builds a universal embodied application development platform for the general public. It also lowers the technical threshold for embodied application exploration, simplifies and visualizes the complex process of traditional robot application development, and allows ordinary users who do not require professional programming skills to explore and practice robots with maximum freedom and ease.
[0037] In some embodiments, the method further includes: obtaining the human motion recording video uploaded by the user; performing joint mapping and local redirection based on the human motion recording video to generate corresponding robot motions. In some embodiments, users can upload human motion recording videos they have filmed (e.g., human dance videos, body movement videos, etc.), automatically identify and track human movement based on AI vision algorithms, automatically perform joint mapping and local redirection from human motion to the robot body, ultimately achieving the imitation and acquisition of human motions, generating corresponding robot motions, and storing them in a motion material library. Joint mapping is a process of establishing a correspondence, defining which joint of the robot should correspond to each human joint. Body redirection, after establishing the joint mapping, calculates the final joint angles that the robot can safely, naturally, and effectively execute under its own physical and geometric constraints based on the human posture. This application provides a video-based human motion imitation technology, which significantly reduces the technical threshold and equipment cost compared to the original motion capture equipment-based data acquisition method for motion imitation.
[0038] In some embodiments, the method further includes: obtaining text information input by the user; obtaining sound configuration information selected by the user for the text information; and synthesizing a corresponding robot voice based on the text information and the sound configuration information. In some embodiments, the user needs to edit the text information first, and then select the sound configuration information corresponding to the edited text information. The sound configuration information is used to configure how to generate the sound corresponding to the text information. Then, based on the text information and the sound configuration information, a TTS (Text-To-Speech) method is used to automatically synthesize the human voice based on the TTS model, synthesizing the corresponding robot voice, and storing it in a sound material library.
[0039] In some embodiments, the voice configuration information includes at least one of timbre information, character design information, and emotion information. This application provides a text-based human voice synthesis technology that enables the automatic synthesis of robot voices with multiple timbres, character designs, and emotions, thereby allowing for low-cost customization of sound requirements for works based on business scenarios.
[0040] In some embodiments, the robot skills include at least one of robot navigation, robot grasping, robot lighting, and robot facial expressions. In some embodiments, robot navigation refers to the robot moving to a specific location; that is, robot navigation includes the destination location, or it may also include the specific path taken during the movement. For example, robot navigation could be "walk to the table." In some embodiments, robot grasping includes, but is not limited to, the identification information of the grasped item (e.g., item name, item image, etc.), the initial placement position of the item, the target object to be delivered to, or the delivery destination location of the item. This example embodiment does not impose any special limitations on these aspects. For example, robot grasping could be "delivering a bottle of mineral water." In some embodiments, robot lighting includes, but is not limited to, warm light, cool light, and neutral light. This example embodiment does not impose any special limitations on these aspects. In some embodiments, robot facial expressions include, but are not limited to, happiness, blinking, confusion, and thinking. This example embodiment does not impose any special limitations on these aspects.
[0041] In some embodiments, the method further includes: presenting text information corresponding to the robot voice to the user; and, in response to an insertion operation performed by the user on the text information, inserting the robot action at a specified position in the text information, so that the humanoid robot begins to execute the robot action when the robot voice is played at the specified position. In some embodiments, presenting the user with the text information corresponding to the robot voice selected by the user for a certain point in time on the timeline may be presenting the specific content of the text information, or it may be presenting a label corresponding to the text information (e.g., a welcome message, a farewell message, etc.), which is not specifically limited in this example embodiment. In some embodiments, the user may insert the robot action selected by the user for that point in time into the specified position in the text information. In some embodiments, if the specified position includes only one position, then that position corresponds to the start execution time of the robot action; if the specified position includes two positions, then the earlier position corresponds to the start execution time of the robot action, and the later position corresponds to the end execution time of the robot action. In some embodiments, the humanoid robot subsequently deployed with the robot creation will not immediately begin to execute the robot action corresponding to that point in time, but will only begin to execute the robot action corresponding to that point in time when the robot voice corresponding to that point in time is played at the specified position. In some embodiments, if the specified position includes only one position, the corresponding robot action is executed when the robot voice is played to that position; if the specified position includes two positions, the corresponding robot action is executed when the robot voice is played to the earlier of the two positions, and the robot action ends when the robot voice is played to the later of the two positions.
[0042] In some embodiments, the method further includes: setting playback configuration information for the robot sound and setting execution configuration information for the robot action based on first duration information corresponding to the robot sound at the designated location and second duration information corresponding to the robot action, so as to automatically align the robot sound and the robot action. In some embodiments, the first duration information corresponding to the robot sound at the designated location refers to the duration from when the robot sound is played to that location until the robot sound ends; if the designated location includes two locations, the first duration information refers to the duration from when the robot sound is played to the earlier of the two locations until the robot sound is played to the later of the two locations. In some embodiments, the second duration information corresponding to the robot action refers to the execution duration of the robot action, that is, the duration from the start to the end of the robot action. In some embodiments, the playback configuration information of the robot voice and the execution configuration information of the robot action can be set according to the first duration information and the second duration information. The playback configuration information is used to configure how to play the robot voice, and the execution configuration information is used to configure how to execute the robot action, thereby automatically aligning the robot voice and the robot action, so that the robot voice and the robot action maintain consistency and synchronization in time. That is, if the specified position includes only one position, it is guaranteed that the robot voice and the robot action will end playback and end execution at the same time, that is, it is guaranteed to end at the same time. If the specified position includes two positions, it is guaranteed that when the robot voice plays to the later position of the two positions, the robot action ends execution.
[0043] In some embodiments, the playback configuration information includes speech rate information, and the execution configuration information includes action execution speed information. In some embodiments, the robot's voice and actions can be aligned by setting the speech rate of the robot's voice and the action execution speed of the robot's actions, so that the robot's voice and actions achieve time consistency and synchronization.
[0044] In some embodiments, the method further includes: obtaining the logic control unit set by the user for the robot skill combination; wherein, synthesizing the corresponding robot creation based on the robot skill combination corresponding to the at least one time point includes: synthesizing the corresponding robot creation based on the robot skill combination corresponding to the at least one time point and the logic control unit corresponding to the robot skill combination, such that the humanoid robot executes the robot skill combination based on the logic control unit at the at least one time point. In some embodiments, after selecting a robot skill combination corresponding to a time point, the user can set the logic control unit corresponding to the robot skill combination. The logic control unit is used to control the execution logic of the robot skill combination. The logic control unit includes, but is not limited to, sequence units, concurrency units, branch units, loop units, etc., which are not specifically limited in this example embodiment. In some embodiments, the humanoid robot that has deployed the robot creation executes the robot skill combination according to the logic set by the logic control unit of the robot skill combination corresponding to the at least one time point at each time point. This application can realize the programming control of robot user stories. The user story supports control logic, including logic control units that support sequence, loop, pause, concurrency, branch, etc., to realize functional programming control of user stories.
[0045] In some embodiments, the logic control unit includes at least one of a sequence unit, a concurrency unit, a branch unit, a loop unit, and a pause unit. In some embodiments, a point in time may correspond to multiple robot skill combinations. The sequence unit controls the sequential execution of these multiple robot skill combinations and their corresponding execution order, i.e., which robot skill combination is executed first and which is executed next at that point in time. The concurrency unit controls the concurrent execution of these multiple robot skill combinations, i.e., multiple robot skill combinations are executed concurrently at that point in time. In some embodiments, the loop unit controls the cyclic execution of robot skill combinations and the corresponding number of loops and / or loop termination conditions, i.e., at a given point in time, the robot skill combination corresponding to that point in time is executed cyclically multiple times. In some embodiments, the pause unit controls the execution of robot skill combinations at a given point in time, specifying that execution is not immediate but paused for a certain duration before starting, and specifying the corresponding pause duration. That is, at a given point in time, the robot skill combination corresponding to that point in time is not executed immediately but paused for a certain duration before starting execution.
[0046] In some embodiments, the logic control unit includes a branch unit, which is configured to trigger a corresponding robot skill combination when a corresponding trigger condition is met, such that the humanoid robot executes the robot skill combination corresponding to the branch unit if the trigger condition corresponding to the branch unit is met at at least one time point. In some embodiments, the branch unit is configured to trigger the corresponding robot skill combination when a corresponding trigger condition is met; that is, at a time point, it is first determined whether the trigger condition set by the branch unit is met. Only if the trigger condition is met will the robot skill combination corresponding to the trigger condition (i.e., the robot skill combination set by the branch unit) be executed. If the condition is not met, the robot skill combination corresponding to that time point will not be executed. In some embodiments, a time point may correspond to multiple branch units. At that time point, it is determined which branch unit's trigger condition is met, and only the robot skill combination corresponding to the branch unit whose trigger condition is met will be executed.
[0047] In some embodiments, the triggering conditions include gesture triggering conditions or visual triggering conditions. In some embodiments, the triggering conditions include gesture triggering conditions (i.e., determining whether to trigger based on the user's gesture) or visual triggering conditions (i.e., determining whether to trigger based on whether a specific object exists in the real-time image of the robot's camera). The robot can identify the current gesture of the user located within a preset range in front of or near the robot to determine whether the current gesture is a specific gesture set by the branch unit, thereby determining whether the gesture triggering condition is met. Alternatively, it can also perform image content recognition on the real-time image of the robot's camera to determine whether a specific object (e.g., a child) set by the branch unit exists in the real-time image, thereby determining whether the gesture triggering condition is met. For example, the visual triggering condition is "seeing a child".
[0048] In some embodiments, the method further includes: obtaining a motion control strategy corresponding to the robot action, such that the humanoid robot performs the robot action based on the motion control strategy at at least one time point. In some embodiments, for robot actions in a robot skill set, a corresponding motion control strategy needs to be obtained so that the humanoid robot performs the robot action at each time point based on the motion control strategy corresponding to the robot action in the robot skill set at that time point. The motion control strategy is an integrated set of algorithms, rules, and computational methods, whose core objective is to calculate the required joint torque or position command in real time based on the robot's current state, sensor information, and the robot action to be performed, thereby enabling the robot to perform robot actions stably, balancedly, efficiently, and adaptively. In some embodiments, after obtaining the motion control strategy corresponding to the robot action, the robot action needs to be simulated and tested in a simulation environment based on the motion control strategy. Only if the simulation test passes can the user deploy the robot creation containing the robot action on the humanoid robot. In some embodiments, a robot model can be established based on physical formulas, and then a controller can be designed to calculate the force or torque required to achieve a specific robot action, thereby obtaining the motion control strategy corresponding to the robot action. In some embodiments, the robot's actions can be learned by imitation or reinforcement learning using an AI model to obtain the motion control strategy corresponding to the robot's actions.
[0049] In some embodiments, the method further includes: obtaining a motion control strategy corresponding to the robot action based on the complexity of the robot action. In some embodiments, it is necessary to determine the motion control strategy corresponding to the robot action based on the complexity of the robot action. If the complexity of the robot action meets the preset simple action conditions, the motion control strategy general to the humanoid robot body is used as the motion control strategy corresponding to the robot action. If the complexity of the robot action meets the preset complex action conditions, the motion control strategy corresponding to the robot action is obtained by training a dedicated motion control model corresponding to the robot action. In some embodiments, relevant information about the robot action can be input into a trained model to obtain the complexity of the robot action output by the model. Alternatively, the complexity of the robot action can be determined by content recognition of a video recording of human actions corresponding to the robot action, and the complexity of the robot action can be determined based on the recognition results. Alternatively, the complexity of the robot action can be set by the user. This example embodiment does not impose any special limitations on this.
[0050] In some embodiments, obtaining the motion control strategy corresponding to the robot action based on the complexity of the robot action includes: if the complexity of the robot action meets a preset simple action condition, using the motion control strategy common to the humanoid robot as the motion control strategy corresponding to the robot action. In some embodiments, if the complexity of the robot action meets the preset simple action condition, that is, if the robot action is a simple action, the motion control strategy common to the humanoid robot can be directly reused to achieve tracking and imitation of the simple action.
[0051] In some embodiments, obtaining the motion control strategy corresponding to the robot action based on the complexity of the robot action includes: if the complexity of the robot action meets a preset complex action condition, obtaining the motion control strategy corresponding to the robot action by training a dedicated motion control model corresponding to the robot action. In some embodiments, if the complexity of the robot action meets the preset complex action condition, that is, if the robot action is a complex action, then it is necessary to train a dedicated motion control model corresponding to the complex action, and use the trained dedicated motion control model as the motion control strategy corresponding to the robot action to achieve tracking and imitation of the complex action. This example embodiment does not specifically limit the specific model structure and model parameters.
[0052] Figure 2 A schematic diagram of a robot work according to one embodiment of this application is shown.
[0053] like Figure 2As shown, the robot project includes multiple time points (t1, t2, t3, t4, t5, t6) in the timeline. Time point t1 corresponds to one branch unit, which corresponds to the gesture trigger condition "gesture wake-up". Time point t2 corresponds to two branch units. One branch unit is used to trigger the execution of the corresponding robot skill combination (warm light, happy expression, "I love you" sound, and heart gesture) when the corresponding visual trigger condition "seeing a child" is met. The other branch unit is used to trigger the execution of the corresponding robot skill combination (normal light, blinking expression, and a welcome sound) when the corresponding visual trigger condition "seeing a leader" is met. The robot skill combination at time point t3 is navigation "walk to the table". The robot skill combination at time point t4 is grabbing "hand over mineral water". Time point t5 corresponds to a branch unit, which is used to trigger the execution of the corresponding robot skill combination (warm light, happy expression, hello voice, handshake action) when the corresponding gesture trigger condition "handshake posture" is met. Time point t6 corresponds to a branch unit, which is used to trigger the execution of the corresponding robot skill combination (normal light, blinking expression, farewell voice, goodbye action) when the corresponding gesture trigger condition "wave goodbye" is met.
[0054] Figure 3 The diagram illustrates a computer device structure for synthesizing a robotic work according to an embodiment of this application. The computer device includes a primary module 11 and a secondary module 12. The primary module 11 is configured to obtain a combination of robotic skills selected by the user for at least one time point in the timeline in response to a selection operation performed by the user on a timeline. The combination of robotic skills includes at least one of robotic actions, robotic voice, and robotic skills. The secondary module 12 is configured to synthesize a corresponding robotic work based on the robotic skill combination corresponding to the at least one time point, such that a humanoid robot deployed with the robotic work performs the robotic skill combination at the at least one time point.
[0055] In some embodiments, the computer device includes, but is not limited to, user equipment and network devices with information processing or computing capabilities, such as mobile phones, tablets, computers, servers, etc. This example embodiment does not impose any special limitations on this.
[0056] Module 11 is used to respond to a user's selection operation on a timeline and obtain a combination of robot skills selected by the user for at least one point in time on the timeline, wherein the combination of robot skills includes at least one of robot actions, robot voices, and robot skills.
[0057] In some embodiments, a timeline is presented to the user on the user device. The user can select at least one time point on the timeline and then select the robot skill combination corresponding to each selected time point. That is, each time point corresponds to a robot skill combination. For example, the user can drag a robot skill combination to the position or near the selected time point on the timeline and use it as the robot skill combination corresponding to that time point.
[0058] In some embodiments, the robot skill set includes at least one of robot actions, robot voice, and robot skills. In some embodiments, robot actions are primarily achieved through video recording. By capturing videos of human actions, joint mapping and redirection from human actions to the robot body are performed in the cloud, ultimately achieving the imitation and capture of human actions to generate corresponding robot actions. In some embodiments, robot voices are primarily achieved through TTS (Text-to-Speech) methods. By selecting timbre, persona, and emotion, human voices are automatically synthesized based on a TTS model to generate corresponding robot voices. In some embodiments, robot skills include at least one of robot navigation, robot grasping, robot lighting, and robot facial expressions. In some embodiments, the user can select the specific content included in each robot skill set. For example, a robot skill set may contain only one of robot actions, robot voice, and robot skills, or it may contain multiple of these three elements, meaning the user can freely combine robot actions, robot voice, and robot skills.
[0059] Module 12 is used to synthesize a corresponding robot creation based on the robot skill combination corresponding to the at least one time point, so that the humanoid robot deployed with the robot creation executes the robot skill combination at the at least one time point.
[0060] In some embodiments, a corresponding robot creation is synthesized based on at least one time point selected by the user on the timeline and a robot skill combination selected by the user for each time point. That is, the robot creation includes at least one time point on the timeline and a robot skill combination corresponding to each of the at least one time point. In some embodiments, the user can deploy the synthesized robot creation in a humanoid robot, causing the deployed humanoid robot to execute the robot skill combination corresponding to each of the at least one time point. Executing the robot skill combination includes at least one of performing robot actions, playing robot sounds, and releasing robot skills.
[0061] In some embodiments, before deploying a robot creation, it is necessary to first determine the motion control strategy used for the robot's actions in the robot creation. For simple robot actions, the general motion control strategy of the humanoid robot body can be directly reused to achieve the tracking and imitation of any action and ensure the balance of motion control of any action. For complex robot actions, it is necessary to train a special motion control strategy for the complex action to achieve the tracking and imitation of the complex action and ensure the balance of motion control of the complex action.
[0062] In some embodiments, before deploying a robot creation, a unified training and simulation platform is used to automatically plan GPU resources, train the model, conduct simulation tests, and evaluate the model's movements. This entire process is transparent to the user. Only after completing the full simulation test can the user deploy the robot creation on a humanoid robot. In some embodiments, GPU computing resources are automatically allocated for the robot's movements in the creation, and a model training container is automatically created. This container mainly includes a simulation environment for training the model (i.e., motion control strategy). Then, by using reinforcement learning and imitation learning algorithms, the robot's movements are imitated. Automatic model evaluation is then implemented. When the model performance meets predetermined requirements, the user is notified. The user can then check whether the model meets expectations through the simulation environment. If it does, the robot creation can be deployed on a real robot (humanoid robot).
[0063] In some embodiments, the computer device is further configured to: obtain a video recording of human movements uploaded by the user; perform joint mapping and local redirection based on the video recording of human movements to generate corresponding robot movements. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0064] In some embodiments, the computer device is further configured to: obtain text information input by the user; obtain sound configuration information selected by the user for the text information; and synthesize a corresponding robot voice based on the text information and the sound configuration information. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0065] In some embodiments, the sound configuration information includes at least one of timbre information, character information, and emotion information. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0066] In some embodiments, the robot skills include at least one of robot navigation, robot grasping, robot lighting, and robot facial expressions. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0067] In some embodiments, the computer device is further configured to: present text information corresponding to the robot voice to the user; and, in response to an insertion operation performed by the user on the text information, insert the robot action at a specified position in the text information, so that the humanoid robot begins to perform the robot action when the robot voice is played to the specified position. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0068] In some embodiments, the computer device is further configured to: set playback configuration information for the robot sound and set execution configuration information for the robot action based on first duration information corresponding to the robot sound at the designated location and second duration information corresponding to the robot action, so as to automatically align the robot sound and the robot action. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0069] In some embodiments, the playback configuration information includes speech rate information, and the execution configuration information includes action execution speed information. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0070] In some embodiments, the computer device is further configured to: obtain the logic control unit set by the user for the robot skill combination; wherein, synthesizing the corresponding robot creation based on the robot skill combination corresponding to the at least one time point includes: synthesizing the corresponding robot creation based on the robot skill combination corresponding to the at least one time point and the logic control unit corresponding to the robot skill combination, such that the humanoid robot executes the robot skill combination based on the logic control unit at the at least one time point. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0071] In some embodiments, the logic control unit includes at least one of a sequential unit, a concurrent unit, a branching unit, a looping unit, and a pause unit. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0072] In some embodiments, the logic control unit includes a branch unit, which is configured to trigger a corresponding robot skill combination when a corresponding triggering condition is met, such that the humanoid robot executes the robot skill combination corresponding to the branch unit if the triggering condition corresponding to the branch unit is met at at least one time point. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0073] In some embodiments, the triggering conditions include gesture triggering conditions or visual triggering conditions. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0074] In some embodiments, the computer device is further configured to: obtain a motion control strategy corresponding to the robot action, such that the humanoid robot performs the robot action based on the motion control strategy at at least one time point. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0075] In some embodiments, the computer device is further configured to: obtain a motion control strategy corresponding to the robot action based on the complexity of the robot action. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0076] In some embodiments, obtaining the motion control strategy corresponding to the robot action based on the complexity of the robot action includes: if the complexity of the robot action meets a preset simple action condition, using the motion control strategy common to the humanoid robot body as the motion control strategy corresponding to the robot action. Here, related operations and... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0077] In some embodiments, obtaining the motion control strategy corresponding to the robot action based on the complexity of the robot action includes: if the complexity of the robot action meets preset complex action conditions, obtaining the motion control strategy corresponding to the robot action by training a dedicated motion control model corresponding to the robot action. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0078] Figure 4Exemplary systems that can be used to implement the various embodiments described in this application are shown; such as Figure 4 As shown in some embodiments, system 300 can function as any of the devices described in each of the embodiments. In some embodiments, system 300 may include one or more computer-readable media having instructions (e.g., system memory or NVM / storage device 320) and one or more processors (e.g., one or more processors 305) coupled to the one or more computer-readable media and configured to execute the instructions to implement the module and thus perform the actions described in this application.
[0079] In one embodiment, the system control module 310 may include any suitable interface controller to provide any suitable interface to at least one of the processors 305 and / or any suitable device or component communicating with the system control module 310.
[0080] The system control module 310 may include a memory controller module 330 to provide an interface to the system memory 315. The memory controller module 330 may be a hardware module, a software module, and / or a firmware module.
[0081] System memory 315 can be used, for example, to load and store data and / or instructions for system 300. In one embodiment, system memory 315 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, system memory 315 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).
[0082] In one embodiment, the system control module 310 may include one or more input / output (I / O) controllers to provide interfaces to the NVM / storage device 320 and (one or more) communication interfaces 325.
[0083] For example, NVM / storage device 320 may be used to store data and / or instructions. NVM / storage device 320 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drive (HDD), one or more optical disc (CD) drives, and / or one or more digital universal optical disc (DVD) drives).
[0084] NVM / storage device 320 may include storage resources that are physically part of a device on which system 300 is mounted, or that can be accessed by the device without necessarily being part of it. For example, NVM / storage device 320 may be accessed via a network through one or more communication interfaces 325.
[0085] One or more communication interfaces 325 may provide the system 300 with an interface to communicate over one or more networks and / or with any other suitable device. The system 300 may wirelessly communicate with one or more components of a wireless network in accordance with any of one or more wireless network standards and / or protocols.
[0086] In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 (e.g., memory controller module 330). In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 to form a system-in-package (SiP). In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die. In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die to form a system-on-a-chip (SoC).
[0087] In various embodiments, system 300 may be, but is not limited to, a server, workstation, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, system 300 may have more or fewer components and / or different architectures. For example, in some embodiments, system 300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.
[0088] In addition to the methods and devices described in the above embodiments, this application also provides a computer-readable storage medium storing computer code that, when executed, performs the method described in any of the preceding embodiments.
[0089] This application also provides a computer program product that, when executed by a computer device, performs the method described in any of the preceding claims.
[0090] This application also provides a computer device, the computer device comprising:
[0091] One or more processors;
[0092] Memory, used to store one or more computer programs;
[0093] When the one or more computer programs are executed by the one or more processors, the one or more processors cause the one or more processors to perform the method as described in any of the preceding methods.
[0094] It should be noted that this application can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of this application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, a magnetic or optical drive, a floppy disk, or similar devices. Furthermore, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.
[0095] Furthermore, a portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0096] Communication media include media through which communication signals containing, for example, computer-readable instructions, data structures, program modules, or other data are transmitted from one system to another. Communication media can include guided transmission media (such as cables and wires (e.g., optical fibers, coaxial cables, etc.)) and wireless (unguided transmission) media capable of propagating energy waves, such as sound, electromagnetic, RF, microwave, and infrared. Computer-readable instructions, data structures, program modules, or other data can be embodied as modulated data signals in, for example, wireless media (such as carrier waves or similar mechanisms embodied as part of spread spectrum technology). The term "modulated data signal" refers to a signal whose one or more characteristics are altered or set in a manner that encodes information in the signal. Modulation can be analog, digital, or a hybrid modulation technique.
[0097] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. For example, computer-readable storage media include, but are not limited to, volatile memories such as random access memory (RAM, DRAM, SRAM); and non-volatile memories such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM); and magnetic and optical storage devices (hard disks, magnetic tapes, CDs, DVDs); or other media now known or hereafter developed capable of storing computer-readable information / data for use by a computer system.
[0098] Herein, one embodiment of this application includes an apparatus comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the apparatus is triggered to run a method and / or technical solution based on the foregoing embodiments of this application.
[0099] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within this application. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the apparatus claims may also be implemented by a single unit or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.
Claims
1. A method for creating synthetic robotic works, wherein, The method includes: In response to a user's selection action on a timeline, obtain a combination of robot skills selected by the user for at least one point in time on the timeline, wherein the combination of robot skills includes at least one of robot actions, robot voices, and robot skills; Based on the robot skill combination corresponding to the at least one time point, a corresponding robot creation is synthesized, such that the humanoid robot deployed with the robot creation executes the robot skill combination at the at least one time point; The method further includes: The system obtains a logic control unit set by the user for the robot skill combination, wherein the logic control unit includes a branch unit, which is used to trigger the corresponding robot skill combination when a corresponding trigger condition is met, such that the humanoid robot executes the robot skill combination corresponding to the branch unit if the trigger condition corresponding to the branch unit is met at at least one time point: The step of synthesizing a corresponding robot creation based on the robot skill combination corresponding to at least one time point includes: Based on the robot skill combination corresponding to the at least one time point and the logic control unit corresponding to the robot skill combination, a corresponding robot creation is synthesized, so that the humanoid robot executes the robot skill combination based on the logic control unit at the at least one time point.
2. The method according to claim 1, wherein, The method further includes: Obtain the recorded video of human actions uploaded by the user; Based on the recorded video of the human movements, joint mapping and local redirection are performed to generate corresponding robot movements.
3. The method according to claim 1, wherein, The method further includes: Obtain the text information input by the user; Obtain the sound configuration information selected by the user for the text information; Based on the text information and the sound configuration information, the corresponding robot voice is synthesized.
4. The method according to claim 3, wherein, The voice configuration information includes at least one of the following: timbre information, character information, and emotion information.
5. The method according to claim 1, wherein, The robot skills include at least one of robot navigation, robot grasping, robot lighting, and robot facial expressions.
6. The method according to claim 1, wherein, The method further includes: Present the user with the text information corresponding to the robot's voice; In response to the user's insertion operation on the text information, the robot action is inserted at a specified position in the text information so that the humanoid robot begins to perform the robot action when the robot voice is played at the specified position.
7. The method according to claim 6, wherein, The method further includes: Based on the first duration information corresponding to the robot voice at the designated location and the second duration information corresponding to the robot action, the playback configuration information of the robot voice and the execution configuration information of the robot action are set to automatically align the robot voice and the robot action.
8. The method according to claim 7, wherein, The playback configuration information includes speech rate information, and the execution configuration information includes action execution speed information.
9. The method according to claim 1, wherein, The logic control unit further includes at least one of a sequence unit, a concurrency unit, a loop unit, and a pause unit.
10. The method according to claim 1, wherein, The triggering conditions include gesture triggering conditions or visual triggering conditions.
11. The method according to claim 1, wherein, The method further includes: Obtain the motion control strategy corresponding to the robot action, so that the humanoid robot performs the robot action based on the motion control strategy at at least one time point.
12. The method according to claim 11, wherein, The method further includes: Based on the complexity of the robot's actions, a motion control strategy corresponding to the robot's actions is obtained.
13. The method according to claim 12, wherein, The step of obtaining the motion control strategy corresponding to the robot action based on the complexity of the robot action includes: If the complexity of the robot's action meets the preset simple action conditions, the motion control strategy common to the humanoid robot body will be used as the motion control strategy corresponding to the robot's action.
14. The method according to claim 13, wherein, The step of obtaining the motion control strategy corresponding to the robot action based on the complexity of the robot action includes: If the complexity of the robot's action meets the preset complex action conditions, the motion control strategy corresponding to the robot's action is obtained by training the dedicated motion control model corresponding to the robot's action.
15. A computer device for performing robot skill combinations, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 14.
16. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 14.
17. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method as described in any one of claims 1 to 14.