Robot optimization learning method, device, equipment and medium for demonstration learning

By collecting user action and voice data through intelligent robots to generate a multimodal feature information set, performing feature semantic alignment processing, and optimizing the autonomous learning model, the adaptability problem of robot systems in dynamic environments is solved, and the autonomous learning and execution capabilities are improved, making it suitable for home applications.

CN119347752BActive Publication Date: 2025-11-04ADDX (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411435122.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2025-11-04
Estimated Expiration
2044-10-15

AI Technical Summary

Technical Problem

Existing robot systems are difficult to adapt to dynamic and complex environmental changes, require programming knowledge, are difficult for users to operate, and have a cumbersome learning and training process, making them unsuitable for home applications.

Method used

The system collects user action and voice data using a depth camera and voice acquisition device of an intelligent robot, generates a multimodal feature information set, performs feature semantic alignment processing, inputs it into a multimodal demonstration learning model, records and optimizes the autonomous learning model to perform the target task.

Benefits of technology

It enhances the robot's autonomous learning and execution capabilities in complex tasks, enabling it to adapt to various dynamic environments, simplifying user operations, and making it suitable for home applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119347752B_ABST
    Figure CN119347752B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a robot optimization learning method, device, equipment and medium for demonstration learning. A specific embodiment of the method comprises: generating a multi-modal feature information set corresponding to an action image sequence, a point cloud data sequence and a voice; performing feature semantic alignment processing on the multi-modal feature information set to generate an aligned multi-modal feature information set; inputting the aligned multi-modal feature information set into a multi-modal demonstration learning model; recording the action image sequence, the point cloud data sequence, the voice, the aligned multi-modal feature information set and a multi-modal demonstration learning result into a memory of an intelligent robot, and controlling the intelligent robot to perform a target demonstration task; and in response to determining that the target demonstration task is performed, optimizing an autonomous learning model of the intelligent robot according to task data corresponding to the target demonstration task. The embodiment improves the autonomous learning and execution capability of the robot system in complex tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of computers, and more specifically to robot optimization learning methods, apparatus, devices, and media for demonstration learning. Background Technology

[0002] Current robotic systems primarily rely on programming and pre-set actions to complete tasks, making them ill-suited to adapting to dynamic and complex environmental changes. While machine learning and artificial intelligence have made progress, they often require vast amounts of data and complex training processes. Furthermore, existing robotic systems demand programming knowledge to set the robot's actions, making them difficult for ordinary users to operate; the robots cannot flexibly respond to changes, requiring human intervention and reprogramming; and the learning and training processes are cumbersome and time-consuming, making them unsuitable for home applications.

[0003] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0005] Some embodiments of this disclosure provide robot optimization learning methods, apparatuses, electronic devices, and computer-readable media for demonstration learning, in order to solve one or more of the technical problems mentioned in the background section above.

[0006] In a first aspect, some embodiments of this disclosure provide a robot optimization learning method for demonstration learning. The method includes: responding to receiving a user demonstration instruction, controlling the depth camera of the intelligent robot to acquire a sequence of motion images and a sequence of point cloud data of a target user; and controlling the voice acquisition device of the intelligent robot to acquire the voice of the target user in real time, wherein the motion images in the motion image sequence correspond to the point cloud data in the point cloud data sequence; generating a multimodal feature information set corresponding to the motion image sequence, the point cloud data sequence, and the voice, wherein the multimodal feature information set includes: motion image modal feature information corresponding to the motion images, point cloud data modal feature information corresponding to the point cloud data, and voice modal feature information corresponding to the voice. The system performs feature semantic alignment processing on the aforementioned multimodal feature information set to generate an aligned multimodal feature information set; inputs the aligned multimodal feature information set into a pre-trained multimodal demonstration learning model to obtain a multimodal demonstration learning result; records the aforementioned action image sequence, point cloud data sequence, the aforementioned speech, the aforementioned aligned multimodal feature information set, and the aforementioned multimodal demonstration learning result into the memory of the aforementioned intelligent robot, and controls the aforementioned intelligent robot to execute a target demonstration task, wherein the aforementioned target demonstration task is a demonstration task corresponding to the aforementioned multimodal demonstration learning result; in response to determining that the aforementioned target demonstration task has been completed, optimizes the autonomous learning model of the aforementioned intelligent robot based on the collected task data corresponding to the aforementioned target demonstration task.

[0007] Secondly, some embodiments of this disclosure provide a robot optimization learning device for demonstration learning. The device includes: a control unit configured to, in response to receiving a user demonstration command, control the depth camera of the intelligent robot to acquire a sequence of motion images and a sequence of point cloud data of a target user, and control the voice acquisition device of the intelligent robot to acquire the voice of the target user in real time, wherein the motion images in the motion image sequence correspond to the point cloud data in the point cloud data sequence; a generation unit configured to generate a multimodal feature information set corresponding to the motion image sequence, the point cloud data sequence, and the voice, wherein the multimodal feature information set includes: motion image modal feature information corresponding to the motion images, point cloud data modal feature information corresponding to the point cloud data, and voice modal feature information corresponding to the voice; and an alignment unit configured to... The system performs feature semantic alignment processing on the aforementioned multimodal feature information set to generate an aligned multimodal feature information set; the input unit is configured to input the aligned multimodal feature information set into a pre-trained multimodal demonstration learning model to obtain a multimodal demonstration learning result; the recording unit is configured to record the aforementioned action image sequence, point cloud data sequence, the aforementioned speech, the aforementioned aligned multimodal feature information set, and the aforementioned multimodal demonstration learning result into the memory of the intelligent robot, and to control the intelligent robot to execute a target demonstration task, wherein the aforementioned target demonstration task is a demonstration task corresponding to the aforementioned multimodal demonstration learning result; the optimization unit is configured to optimize the autonomous learning model of the intelligent robot based on the collected task data corresponding to the aforementioned target demonstration task in response to determining that the aforementioned target demonstration task has been completed.

[0008] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0009] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0010] The various embodiments disclosed above have the following beneficial effects: The robot optimization learning method for demonstration learning in some embodiments of this disclosure enhances the autonomous learning and execution capabilities of the robot system in complex tasks. It can not only efficiently imitate and execute user operations, but also self-optimize through continuous feedback and improvement, adapting to various complex and dynamic task environments. In the future, this system is expected to be widely applied in multiple fields, such as intelligent manufacturing, home services, and medical care, providing strong technical support for the automation and intelligent development of various industries. Firstly, in response to receiving a user demonstration command, the depth camera of the aforementioned intelligent robot is controlled to acquire the target user's motion image sequence and point cloud data sequence, and the voice acquisition device of the aforementioned intelligent robot is controlled to acquire the target user's voice in real time, wherein the motion images in the motion image sequence correspond to the point cloud data in the point cloud data sequence. Therefore, based on the user's specific operations, the dynamic changes in the environment during the operation and the user's voice commands are also covered, ensuring the comprehensiveness and accuracy of the recording. Next, a multimodal feature information set corresponding to the aforementioned action image sequence, point cloud data sequence, and speech is generated. This multimodal feature information set includes: action image modal feature information corresponding to the action images, point cloud data modal feature information corresponding to the point cloud data, and speech modal feature information corresponding to the speech. Then, feature semantic alignment processing is performed on the multimodal feature information set to generate an aligned multimodal feature information set. Subsequently, the aligned multimodal feature information set is input into a pre-trained multimodal demonstration learning model to obtain the multimodal demonstration learning result. This allows for the identification and generalization of commonalities and patterns in different demonstrations through a multimodal large-scale model algorithm, thereby providing a deeper understanding of the task's essence. Then, the aforementioned action image sequence, point cloud data sequence, speech, aligned multimodal feature information set, and multimodal demonstration learning result are recorded in the memory of the intelligent robot, and the intelligent robot is controlled to execute a target demonstration task, where the target demonstration task is a demonstration task corresponding to the multimodal demonstration learning result. Finally, in response to the confirmation that the aforementioned target demonstration task has been completed, the autonomous learning model of the intelligent robot is optimized based on the collected task data corresponding to the target demonstration task. This allows for dynamic adjustments and optimizations based on real-time feedback, ensuring the smooth completion and efficient execution of the task. This enhances the robot system's autonomous learning and execution capabilities in complex tasks. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0012] Figure 1 This is a flowchart of some embodiments of a robot optimization learning method for demonstration learning according to the present disclosure;

[0013] Figure 2 The example illustrates the steps and scenarios for an intelligent robot to optimize its driving route.

[0014] Figure 3 This is a flowchart of some embodiments of a robot-optimized learning device for demonstration learning according to the present disclosure;

[0015] Figure 4 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0016] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0017] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0018] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0019] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0020] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0021] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0022] Figure 1This is a flowchart of some embodiments of the robot optimization learning method for demonstration learning according to the present disclosure. A flowchart 100 of some embodiments of the robot optimization learning method for demonstration learning according to the present disclosure is shown. This robot optimization learning method for demonstration learning, applied to an intelligent robot, includes the following steps:

[0023] Step 101: In response to receiving a user demonstration command, control the depth camera of the intelligent robot to acquire motion image sequences and point cloud data sequences of the target user, and control the voice acquisition device of the intelligent robot to acquire the voice of the target user in real time.

[0024] In some embodiments, the execution entity (e.g., a computing device) of the robot optimization learning method for demonstration learning can, in response to receiving a user demonstration instruction, control the depth camera of the intelligent robot to acquire a sequence of motion images and a sequence of point cloud data of the target user, and control the voice acquisition device of the intelligent robot to acquire the voice of the target user in real time. The motion images in the motion image sequence correspond to the point cloud data in the point cloud data sequence. The user demonstration instruction can be an instruction for the user to demonstrate actions, and for the intelligent robot to record the user's action steps, operational details, and environmental changes. The intelligent robot can be an intelligent robot with a mechanical hand that can mimic user actions. The target user can be a demonstration user who demonstrates an operation task several times through manual operation, voice commands, or other means. For example, the voice acquisition device can be a microphone.

[0025] The intelligent robot employs advanced sensing technology, using a depth camera to record RGB and depth point cloud information of its operations, while simultaneously recording relevant voice information. This data includes not only the user's specific actions but also the dynamic changes in the environment during the operation and the user's voice commands, ensuring the comprehensiveness and accuracy of the recording.

[0026] Step 102: Generate the above-mentioned action image sequence, point cloud data sequence, and multimodal feature information set corresponding to the above-mentioned speech.

[0027] In some embodiments, the executing entity can generate a multimodal feature information set corresponding to the action image sequence, point cloud data sequence, and speech. The multimodal feature information set includes: action image modal feature information corresponding to the action images, point cloud data modal feature information corresponding to the point cloud data, and speech modal feature information corresponding to the speech. The multimodal feature information can characterize the semantic content of the object information in the corresponding modality. The multimodal feature information can be in vector form.

[0028] As an example, the aforementioned execution entity can utilize the corresponding multimodal feature extraction model to extract semantic content features from object information (action image sequences, point cloud data sequences, and speech) in each modality to generate multimodal feature information, thus obtaining a multimodal feature information set. There is a one-to-one correspondence between the modality feature extraction model and the modality in the multimodal model. The modality feature extraction model can be a neural network model that extracts semantic content features for the corresponding modality. The network structure of the modality feature extraction model can vary depending on the corresponding modality. Multimodality can represent action image modality, point cloud data modality, and speech modality. For example, multimodal feature extraction models can include, but are not limited to: image feature extraction models, point cloud feature extraction models, and speech feature extraction models.

[0029] Step 103: Perform feature semantic alignment processing on the above multimodal feature information set to generate an aligned multimodal feature information set.

[0030] In some embodiments, the aforementioned execution entity may perform feature semantic alignment processing on the aforementioned multimodal feature information set to generate an aligned multimodal feature information set. The LLaVA (Large Language and Vision Assistant) model can be used to perform feature semantic alignment processing on the aforementioned multimodal feature information set to generate an aligned multimodal feature information set.

[0031] Step 104: Input the above-mentioned aligned multimodal feature information set into the pre-trained multimodal demonstration learning model to obtain the multimodal demonstration learning result.

[0032] In some embodiments, the aforementioned execution entity can input the aligned multimodal feature information set into a pre-trained multimodal demonstration learning model to obtain multimodal demonstration learning results. The multimodal demonstration learning model can refer to a multi-model self-learning model, which may include: Type A multimodal models, Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), flow-based generative models, etc. Type A multimodal models are deep fusion models based on standard cross-attention. They use the standard Transformer model and add standard cross-attention layers to the internal layers of the model to achieve deep fusion of input multimodal information. This type of model requires a large amount of multimodal training data and has high computational resource requirements, but it has the advantage of fine-grained control over multimodal information. Typical Type A models include Flamingo and OpenFlamingo, which can process image and text data. Generative Adversarial Networks (GANs) are trained adversarially with a generative model and a discriminative model. The generative model attempts to deceive the discriminative model, making it unable to distinguish between real and fake data. Variational Autoencoders (VAEs) are generative models that map the input to a probability distribution in the hidden space, rather than a fixed encoding. Stream-based generative models progressively add noise to a real data distribution by defining forward and backward processes, and then learn to recover the real data from the noise. Multimodal learning is a method that utilizes data from different senses or interaction methods for learning; these data modalities may include text, images, audio, video, etc. Multimodal learning trains models by fusing multiple data modalities, thereby improving the model's perception and understanding capabilities and achieving cross-modal information interaction and fusion. Typical modal representation methods include the pre-trained text model BERT, convolutional neural networks (CNNs) for images, and deep neural networks (DNNs) for audio. Multimodal demonstration learning results can refer to the results of self-learning on an aligned set of multimodal feature information.

[0033] The multimodal demonstration learning model can be trained through the following steps:

[0034] The first step is to obtain a multimodal feature information training sample set. This multimodal feature information training sample set includes: sample action image feature information, sample point cloud data feature information, sample speech feature information, and sample labels.

[0035] The second step is to determine the initial multimodal demonstration learning model. This initial multimodal demonstration learning model includes: an initial action image demonstration learning network, an initial point cloud data demonstration learning network, and an initial speech demonstration learning network. The initial speech demonstration learning network includes: an initial speech feature extraction network, an initial multi-level feature autoencoder network, and an initial speech recognition network. The initial action image demonstration learning network can be a self-learning network model used to learn action features from action images. The initial point cloud data demonstration learning network can be a self-learning network model used to learn action features from point cloud data. The initial speech demonstration learning network can be a self-learning network model used to learn speech features from speech. The initial speech feature extraction network can be an untrained neural network used to extract speech features. For example, the initial speech feature extraction network can be a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory network (LSTM), etc. The aforementioned initial multi-level feature autoencoder network includes: an initial encoder network and an initial decoder network. The initial encoder network includes: an initial speech feature decoupling network (containing an autoencoder network and an LSTM network) and an initial speech feature fusion network (VAE (Variational AutoEncoder)). The initial speech feature decoupling network includes: an initial static speech feature decoupling network and an initial dynamic speech feature decoupling network. The initial speech feature fusion network includes: an initial fusion network and an initial reparameterization network. The initial static speech feature decoupling network can be an Auto-Encoder model. The initial dynamic speech feature decoupling network can be an LSTM (Long Short-Term Memory) model. The initial fusion network can be a Feature Pyramid Network (FPN) or a Self-Attention Mechanism network. The initial reparameterization network can include: DyRep (Dynamic Re-parameterization) and ACNet (Asymmetric Convolution). Initial speech recognition networks can include rule-based models, statistical models, and deep learning-based models.

[0036] The third step is to select a target multimodal feature information training sample from the aforementioned multimodal feature information training sample set. One multimodal feature information training sample can be randomly selected from the aforementioned multimodal feature information training sample set as the target multimodal feature information training sample.

[0037] The fourth step is to train the initial multimodal demonstration learning model based on the target multimodal feature information training samples, thereby obtaining the trained multimodal demonstration learning model.

[0038] The fourth step mentioned above may include the following sub-steps:

[0039] The first sub-step involves inputting the sample action image feature information, which is included in the target multimodal feature information training samples, into the aforementioned initial action image demonstration learning network to obtain the initial action image demonstration learning result.

[0040] The second sub-step involves inputting the feature information of the sample point cloud data, which is included in the target multimodal feature information training samples, into the initial point cloud data demonstration learning network to obtain the initial point cloud data demonstration learning result.

[0041] The third sub-step involves inputting the sample speech feature information included in the target multimodal feature information training samples into the aforementioned initial speech feature extraction network to obtain the initial speech feature extraction information.

[0042] The fourth sub-step involves inputting the initial speech feature extraction information into the initial multi-level feature autoencoder network to obtain the initial multi-level speech feature information.

[0043] The fifth sub-step involves inputting the aforementioned initial multi-level speech feature information into the initial speech recognition network to obtain the initial speech recognition information.

[0044] The sixth sub-step involves merging the initial motion image demonstration learning results, the initial point cloud data demonstration learning results, and the initial speech recognition information into initial multimodal demonstration learning information.

[0045] The seventh sub-step involves determining the loss value between the initial multimodal demonstration learning information and the corresponding sample labels. This loss value can be determined using either the cross-entropy loss function or the hinge loss function.

[0046] The eighth sub-step is to determine the initial multimodal demonstration learning model as the trained multimodal demonstration learning model in response to the determination that the above loss value is less than or equal to the preset loss value.

[0047] Step 105: Record the above-mentioned action image sequence, point cloud data sequence, above-mentioned speech, above-mentioned aligned multimodal feature information set and above-mentioned multimodal demonstration learning results into the memory of the above-mentioned intelligent robot, and control the above-mentioned intelligent robot to perform the target demonstration task.

[0048] In some embodiments, the executing entity may record the motion image sequence, point cloud data sequence, speech, aligned multimodal feature information set, and multimodal demonstration learning results into the memory of the intelligent robot, and control the intelligent robot to perform a target demonstration task. The target demonstration task is a demonstration task corresponding to the multimodal demonstration learning results. The target demonstration task may be a pre-defined task instructing the intelligent robot to learn motion and speech features based on the motion image sequence and point cloud data sequence of the target user.

[0049] Step 106: In response to determining that the above-mentioned target demonstration task has been completed, the autonomous learning model of the above-mentioned intelligent robot is optimized based on the collected task data corresponding to the above-mentioned target demonstration task.

[0050] In some embodiments, the aforementioned execution entity may, in response to determining that the target demonstration task has been completed, optimize the autonomous learning model of the intelligent robot based on the collected task data corresponding to the target demonstration task. Task data can refer to the task execution data generated by the intelligent robot during the execution of the target demonstration task. The autonomous learning model can refer to a self-learning model set in the intelligent robot, or it can be a multimodal self-learning model. The intelligent robot establishes a model of the operation task (autonomous learning model), including action sequences, conditions, and expected results. It can also identify and generalize commonalities and patterns in different demonstrations through multimodal large model algorithms, thereby gaining a deep understanding of the essence of the task.

[0051] As an example, firstly, the difference between the pre-defined standard task data and the actual task data corresponding to the aforementioned target demonstration task can be determined. Secondly, the model parameters of the self-learning model in the intelligent robot can be adjusted using gradient descent, allowing it to continuously learn and improve, thereby enhancing the robot's autonomy and accuracy. For instance, users can provide the robot with more sample data through additional demonstration operations, helping it better understand task details. The robot will integrate this new data with the existing model, gradually improving and optimizing its operational strategies. Furthermore, the robot's self-learning capabilities enable it to continuously accumulate experience during use, improving execution efficiency and accuracy.

[0052] Optionally, information on the driving learning task requirements of the intelligent robot in the target area can be obtained.

[0053] In some embodiments, the aforementioned execution entity can obtain driving learning task requirement information for the intelligent robot in the corresponding target area. This driving learning task requirement information includes: a driving start point, a driving end point, and environmental information. The environmental information includes at least one obstacle. The at least one obstacle can refer to obstacle information between the driving start point and the driving end point, and may include obstacle location / obstacle volume. The driving start point can be the starting point for the intelligent robot's driving. The driving end point can be the ending point for the intelligent robot's driving.

[0054] Optionally, based on the above driving learning task requirements information, an initial driving route is generated between the above driving start point and the above driving end point.

[0055] In some embodiments, the executing entity can generate an initial driving route between the driving start point and the driving end point based on the driving learning task requirements. This initial driving route can be generated using a target-oriented heuristic search algorithm. The initial driving route can be a path starting from the driving start point and ending at the driving end point. In practice, the target-oriented heuristic search algorithm can be the A* (AStar) algorithm.

[0056] Optionally, the node routes between each driving node in the initial driving route can be optimized to generate an optimized driving route.

[0057] In some embodiments, the execution entity may optimize the node routes between each driving node in the initial driving route to generate an optimized driving route.

[0058] In practice, the aforementioned implementing entity can optimize the node routes between each driving node in the initial driving route through the following steps:

[0059] The first step is to determine the distance matrix and route matrix for the corresponding starting and ending points of the journey using an interpolation method. The interpolation method can be the Floyd-Warshall algorithm. The distance matrix represents the distances between all nodes on the journey between the starting and ending points. The route matrix represents the route information between all nodes on the journey between the starting and ending points. Each node can be a pre-defined landmark node or a route node set based on travel time or distance.

[0060] The second step is to optimize the node routes between each driving node in the initial driving route based on the distance matrix and the route matrix mentioned above, so as to generate an optimized driving route.

[0061] The second step mentioned above may include the following sub-steps:

[0062] The first sub-step involves determining the sequence of driving nodes corresponding to the initial driving route. This sequence may include: the starting point, the ending point, and all driving nodes. The driving nodes in the sequence are arranged in the order of travel.

[0063] The second sub-step involves dividing the driving nodes in the above driving node sequence into driving node group sequences. For example, each pair of adjacent driving nodes in the above driving node sequence is determined as a driving node subgroup, resulting in a driving node group sequence.

[0064] The third sub-step involves performing the following processing steps for each travel node group in the above travel node group sequence:

[0065] 1. Based on the aforementioned distance matrix and route matrix, determine the first node distance and first node route corresponding to the aforementioned driving node group. The first node distance represents the distance between each driving node in the driving node group. The first node route represents the connection route between each driving node in the driving node group. All driving nodes in the driving node group are adjacent. The aforementioned execution entity can select the node distance corresponding to the driving node group from the distance matrix as the first node distance, and select the node route corresponding to the driving node group from the route matrix as the first node route.

[0066] 2. Determine the second node route for the aforementioned driving node group within the initial driving route, and determine the second node distance corresponding to the second node route. The second node distance represents the distance between each driving node in the initial driving route. The second node route represents the connection route between each driving node in the initial driving route. The node distances corresponding to the driving node group can be selected from the distance matrix as the second node distances, and the node routes corresponding to the driving node group can be selected from the route matrix as the second node routes.

[0067] 3. In response to determining that the distance to the first node is less than or equal to the distance to the second node, the route to the first node is determined as a candidate node route.

[0068] 4. In response to determining that the distance to the first node is greater than the distance to the second node, the route to the second node is determined as a candidate node route.

[0069] The fourth sub-step involves combining the obtained candidate node routes into an optimized driving route.

[0070] Further reference Figure 2The example illustrates the steps and scenarios for optimizing a driving route for an intelligent robot, including: 201 represents the intelligent robot, 202 represents the starting point A, 203 represents the destination B, 204 represents the initial driving route (A→C→D→B) after performing a search algorithm on the starting point A and the destination B; and 205 represents the optimized route (A→E→D→B).

[0071] Optionally, based on the optimized driving route described above, the intelligent robot can be controlled to perform a driving learning task for the corresponding target area.

[0072] In some embodiments, the aforementioned executing entity can control the aforementioned intelligent robot to perform a driving learning task in a corresponding target area based on the optimized driving route. For example, a driving learning instruction can be sent to the intelligent robot, enabling the intelligent robot to automatically drive within the target area according to the optimized driving route, thereby adapting to the environment of the target area.

[0073] This allows the robot to adjust its movement path when faced with unexpected obstacles, or to adjust its perception strategy according to changes in ambient lighting conditions. This adaptive capability significantly enhances the robot's operational performance in complex environments.

[0074] Further reference Figure 3 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a robot optimization learning device for demonstration learning. These embodiments of the robot optimization learning device for demonstration learning are similar to... Figure 1 Corresponding to the method embodiments shown, this robot optimization learning device for demonstration learning can be specifically applied to various electronic devices.

[0075] like Figure 3As shown, a robot optimization learning device 300 for demonstration learning in some embodiments includes: a control unit 301, a generation unit 302, an alignment unit 303, an input unit 304, a recording unit 305, and an optimization unit 306. The control unit 301 is configured to, in response to receiving a user demonstration command, control the depth camera of the intelligent robot to acquire a sequence of motion images and a sequence of point cloud data of the target user, and control the voice acquisition device of the intelligent robot to acquire the voice of the target user in real time, wherein the motion images in the motion image sequence correspond to the point cloud data in the point cloud data sequence; the generation unit 302 is configured to generate a multimodal feature information set corresponding to the motion image sequence, the point cloud data sequence, and the voice, wherein the multimodal feature information set includes: motion image modal feature information corresponding to the motion images, point cloud data modal feature information corresponding to the point cloud data, and voice modal feature information corresponding to the voice; the alignment unit 303 is configured to perform feature semantic alignment on the multimodal feature information set. The system processes data to generate an aligned multimodal feature information set; an input unit 304 is configured to input the aligned multimodal feature information set into a pre-trained multimodal demonstration learning model to obtain a multimodal demonstration learning result; a recording unit 305 is configured to record the action image sequence, point cloud data sequence, speech, aligned multimodal feature information set, and multimodal demonstration learning result into the memory of the intelligent robot, and to control the intelligent robot to execute a target demonstration task, wherein the target demonstration task is a demonstration task corresponding to the multimodal demonstration learning result; an optimization unit 306 is configured to optimize the autonomous learning model of the intelligent robot based on the collected task data corresponding to the target demonstration task in response to determining that the target demonstration task has been completed.

[0076] It is understandable that the units described in the robot optimization learning device 300 for demonstration learning are similar to those in the reference. Figure 1 The steps in the described method correspond accordingly. Therefore, the operations, features, and beneficial effects described above for the method also apply to the robot optimization learning device 300 used for demonstration learning and the units contained therein, and will not be repeated here.

[0077] The following is for reference. Figure 4 This illustration shows a structural schematic of an electronic device (e.g., a computing device) 400 suitable for implementing some embodiments of the present disclosure. The electronic devices in some embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0078] like Figure 4 As shown, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of electronic device 400. The processing device 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.

[0079] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 4 Each box shown can represent a device or multiple devices as needed.

[0080] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program is executed by processing device 401, it performs the functions defined above in the methods of some embodiments of this disclosure.

[0081] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0082] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0083] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: in response to receiving a user demonstration instruction, control the depth camera of the aforementioned intelligent robot to acquire a sequence of motion images and a sequence of point cloud data of the target user, and control the voice acquisition device of the aforementioned intelligent robot to acquire the voice of the target user in real time, wherein the motion images in the aforementioned motion image sequence correspond to the point cloud data in the aforementioned point cloud data sequence; generate a multimodal feature information set corresponding to the aforementioned motion image sequence, the point cloud data sequence, and the aforementioned voice, wherein the multimodal feature information set includes: motion image modal feature information corresponding to the motion images, point cloud data modal feature information corresponding to the point cloud data, and voice modal feature information corresponding to the voice. The system acquires multimodal feature information; performs feature semantic alignment processing on the aforementioned multimodal feature information set to generate an aligned multimodal feature information set; inputs the aligned multimodal feature information set into a pre-trained multimodal demonstration learning model to obtain multimodal demonstration learning results; records the aforementioned action image sequence, point cloud data sequence, the aforementioned speech, the aforementioned aligned multimodal feature information set, and the aforementioned multimodal demonstration learning results into the memory of the aforementioned intelligent robot, and controls the aforementioned intelligent robot to execute a target demonstration task, wherein the aforementioned target demonstration task is a demonstration task corresponding to the aforementioned multimodal demonstration learning results; in response to determining that the aforementioned target demonstration task has been completed, optimizes the autonomous learning model of the aforementioned intelligent robot based on the collected task data corresponding to the aforementioned target demonstration task.

[0084] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0085] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0086] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including: a control unit, a generation unit, an alignment unit, an input unit, a recording unit, and an optimization unit. The names of these units do not necessarily limit the specific unit; for example, the optimization unit may be described as "a unit that, in response to determining that the aforementioned target demonstration task has been completed, optimizes the autonomous learning model of the intelligent robot based on the collected task data corresponding to the aforementioned target demonstration task."

[0087] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0088] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A robot optimization learning method for demonstration learning, applied to intelligent robots, comprising: In response to receiving a user demonstration command, the robot controls its depth camera to acquire a sequence of motion images and a sequence of point cloud data of the target user, and controls its voice acquisition device to acquire the voice of the target user in real time, wherein the motion images in the sequence of motion images correspond to the point cloud data in the sequence of point cloud data. Generate a multimodal feature information set corresponding to the action image sequence, point cloud data sequence, and speech, wherein the multimodal feature information set includes: action image modal feature information corresponding to the action image, point cloud data modal feature information corresponding to the point cloud data, and speech modal feature information corresponding to the speech; The multimodal feature information set is subjected to feature semantic alignment processing to generate an aligned multimodal feature information set; The aligned multimodal feature information set is input into a pre-trained multimodal demonstration learning model to obtain the multimodal demonstration learning result; The motion image sequence, point cloud data sequence, speech, aligned multimodal feature information set, and multimodal demonstration learning results are recorded in the memory of the intelligent robot, and the intelligent robot is controlled to execute a target demonstration task, wherein the target demonstration task is a demonstration task corresponding to the multimodal demonstration learning results; In response to determining that the target demonstration task has been completed, the autonomous learning model of the intelligent robot is optimized based on the collected task data corresponding to the target demonstration task.

2. The method according to claim 1, wherein, Before inputting the aligned multimodal feature information set into the pre-trained multimodal demonstration learning model to obtain the multimodal demonstration learning result, the method further includes: A multimodal feature information training sample set is obtained, wherein the multimodal feature information training samples in the multimodal feature information training sample set include: sample action image feature information, sample point cloud data feature information, sample speech feature information, and sample labels; An initial multimodal demonstration learning model is determined, wherein the initial multimodal demonstration learning model includes: an initial action image demonstration learning network, an initial point cloud data demonstration learning network, and an initial speech demonstration learning network, wherein the initial speech demonstration learning network includes: an initial speech feature extraction network, an initial multi-level feature autoencoder network, and an initial speech recognition network. Select target multimodal feature information training samples from the multimodal feature information training sample set; Training samples based on target multimodal feature information are used to train the initial multimodal demonstration learning model, resulting in a trained multimodal demonstration learning model.

3. The method according to claim 2, wherein, The training samples based on the target multimodal feature information are used to train the initial multimodal demonstration learning model, resulting in a trained multimodal demonstration learning model, including: The sample action image feature information included in the target multimodal feature information training sample is input into the initial action image demonstration learning network to obtain the initial action image demonstration learning result; The sample point cloud data feature information included in the target multimodal feature information training samples is input into the initial point cloud data demonstration learning network to obtain the initial point cloud data demonstration learning result; The sample speech feature information included in the target multimodal feature information training samples is input into the initial speech feature extraction network to obtain the initial speech feature extraction information; The initial speech feature extraction information is input into the initial multi-level feature autoencoder network to obtain the initial multi-level speech feature information; The initial multi-level speech feature information is input into the initial speech recognition network to obtain the initial speech recognition information; The initial motion image demonstration learning results, the initial point cloud data demonstration learning results, and the initial speech recognition information are merged into initial multimodal demonstration learning information; Determine the loss value between the initial multimodal demonstration learning information and the corresponding sample labels; In response to determining that the loss value is less than or equal to a preset loss value, the initial multimodal demonstration learning model is determined as the trained multimodal demonstration learning model.

4. The method according to claim 1, wherein, The method further includes: Obtain the driving learning task requirement information of the intelligent robot in the target area. The driving learning task requirement information includes: driving start point, driving end point, and environmental information. The environmental information includes: information on at least one obstacle. Based on the driving learning task requirements, an initial driving route is generated between the driving start point and the driving end point; The node routes between each driving node in the initial driving route are optimized to generate an optimized driving route; Based on the optimized driving route, the intelligent robot is controlled to perform driving learning tasks in the corresponding target area.

5. The method according to claim 4, wherein, The step of optimizing the node routes between each driving node in the initial driving route to generate an optimized driving route includes: Using the interpolation method, determine the distance matrix and route matrix corresponding to the starting point and ending point of the journey; Based on the distance matrix and the route matrix, the node routes between each driving node in the initial driving route are optimized to generate an optimized driving route.

6. A robot optimization learning device for demonstration learning, applied to an intelligent robot, comprising: The control unit is configured to, in response to receiving a user demonstration command, control the depth camera of the intelligent robot to acquire a sequence of motion images and a sequence of point cloud data of the target user, and control the voice acquisition device of the intelligent robot to acquire the voice of the target user in real time, wherein the motion images in the sequence of motion images correspond to the point cloud data in the sequence of point cloud data. The generation unit is configured to generate the action image sequence, the point cloud data sequence, and the multimodal feature information set corresponding to the speech, wherein the multimodal feature information set includes: action image modal feature information corresponding to the action image, point cloud data modal feature information corresponding to the point cloud data, and speech modal feature information corresponding to the speech; The alignment unit is configured to perform feature semantic alignment processing on the multimodal feature information set to generate an aligned multimodal feature information set; The input unit is configured to input the aligned multimodal feature information set into a pre-trained multimodal demonstration learning model to obtain the multimodal demonstration learning result; The recording unit is configured to record the motion image sequence, point cloud data sequence, speech, aligned multimodal feature information set and multimodal demonstration learning result into the memory of the intelligent robot, and to control the intelligent robot to execute a target demonstration task, wherein the target demonstration task is a demonstration task corresponding to the multimodal demonstration learning result; The optimization unit is configured to optimize the autonomous learning model of the intelligent robot in response to determining that the target demonstration task has been completed, based on the collected task data corresponding to the target demonstration task.

7. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-5.

8. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Robot motion skill learning method fusing text instruction and motion information

    CN117428780A

  • Mechanical arm grabbing method driven by natural language

    CN117773920A