Robot motion control method and device, electronic equipment and readable storage medium
By combining channel attention mechanism and conditional variational autoencoder, a method is used to extract robot environmental features and joint states to generate motion control signals, which solves the problem of robot action mismatch with environment and improves the accuracy and adaptability of robot motion control in complex environments.
Patent Information
- Application Number
- CN202411804453.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2044-12-09
AI Technical Summary
In existing technologies, the fixed motion of robots leads to a mismatch between their movements and the actual environment, resulting in poor motion performance.
By using a pre-trained target feature extraction network based on channel attention mechanism to extract environmental features, and combining it with a joint position state input prediction network to generate motion control signals, the robot motion control is achieved by a combination of conditional variational autoencoder and channel attention mechanism.
It improves the accuracy and adaptability of robots' visual perception and motion control in complex environments, reduces mismatches between actions and the actual environment, and enables precise execution of motion tasks.
Smart Images

Figure CN119681873B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of robots, in particular to a robot motion control method and device, electronic equipment and readable storage medium. BACKGROUND
[0002] At present, the control of robots is generally realized by program instructions for indicating fixed actions. However, due to the fixed actions of robots, the actions of robots may not match the actual environment, resulting in poor robot motion effect. SUMMARY
[0003] The embodiments of the present application provide a robot motion control method and device, electronic equipment and readable storage medium, which can accurately extract important features in visual input, and then generate a signal for motion control of a robot in combination with joint position states, so as to reduce the situation that the actions of the robot do not match the actual environment.
[0004] The embodiments of the present application can be implemented as follows:
[0005] In a first aspect, the embodiments of the present application provide a robot motion control method, which comprises:
[0006] obtaining a first environment image currently collected by a target robot and a first joint position state of the target robot at present;
[0007] extracting environment features from the first environment image by using a pre-trained target feature extraction network based on a channel attention mechanism to obtain first environment features;
[0008] inputting the first environment features and the first joint position state into a prediction network to obtain an initial motion control signal, wherein the prediction network is trained according to sample environment features, sample joint position states and sample motion control signals corresponding to each time point, the sample environment features and the sample joint position states are samples, and the sample motion control signals are labels corresponding to the samples;
[0009] controlling the motion of the target robot according to the initial motion control signal.
[0010] In a second aspect, the embodiments of the present application provide a robot motion control device, which comprises:
[0011] a collection module configured to obtain a first environment image currently collected by a target robot and a first joint position state of the target robot at present;
[0012] a processing module configured to perform environment feature extraction on the first environment image by using a pre-trained target feature extraction network based on a channel attention mechanism to obtain first environment features;
[0013] The processing module is further configured to input the first environment features and the first joint position state into a prediction network to obtain an initial motion control signal, wherein the prediction network is trained according to sample environment features, sample joint position states and sample motion control signals corresponding to respective time points, the sample environment features and the sample joint position states are regarded as samples, and the sample motion control signals are regarded as labels corresponding to the samples.
[0014] A control module is configured to control the motion of the target robot according to the initial motion control signal.
[0015] In a third aspect, an electronic device is provided, which includes a processor and a memory. The memory stores machine executable instructions capable of being executed by the processor. The processor can execute the machine executable instructions to implement the robot motion control method according to the foregoing embodiments.
[0016] In a fourth aspect, a readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the robot motion control method according to the foregoing embodiments.
[0017] The robot motion control method, device, electronic device and readable storage medium provided by the embodiments of the present application first obtain a first environment image currently obtained by a target robot and a first joint position state of the target robot at present. Then, a target feature extraction network based on a pre-trained channel attention mechanism is used to perform environment feature extraction on the first environment image to obtain first environment features. Next, the first environment features and the first joint position state are input into a prediction network to obtain an initial motion control signal. The prediction network is trained according to sample environment features, sample joint position states and sample motion control signals corresponding to respective time points. The sample environment features and the sample joint position states are regarded as samples, and the sample motion control signals are regarded as labels corresponding to the samples. Finally, the motion of the target robot is controlled according to the initial motion control signal. In this way, important features in visual input can be accurately extracted by using a target feature extraction network, and then a signal for motion control of the robot is generated by a prediction network in combination with a joint position state, so as to reduce the situation that the motion of the robot does not match the actual environment, and facilitate driving the robot to perform an accurate motion task. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be regarded as limiting the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor.
[0019] Figure 1 The block schematic diagram of the electronic device provided by the embodiments of the present application is shown in the following figure.
[0020] Figure 2 The flow schematic diagram of the robot motion control method provided by the embodiments of the present application is shown in the following figure.
[0021] Figure 3 The principle schematic diagram of the robot motion control method provided by the embodiments of the present application is shown in the following figure.
[0022] Figure 4 The flow schematic diagram of the robot motion control method provided by the embodiments of the present application is shown in the following figure.
[0023] Figure 5 The block schematic diagram of the robot motion control device provided by the embodiments of the present application is shown in the following figure.
[0024] Figure 6 The block schematic diagram of the robot motion control device provided by the embodiments of the present application is shown in the following figure.
[0025] Figure: 100-electronic device; 110-memory; 120-processor; 130-communication unit; 200-robot motion control device; 201-training module; 210-acquisition module; 220-processing module; 230-control module. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the following will combine the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, not all the embodiments. The components of the embodiments of the present application described and shown in the drawings can be arranged and designed in various different configurations.
[0027] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of the present application.
[0028] It should be noted that the terms "first", "second" and the like, which are used to describe various elements, are not used to denote any inherent relation or order between the elements. Also, the terms "comprise", "include" or "contain" or any other variant thereof are intended to mean that the elements listed are not exhaustive, and that other elements are possible. The terms "comprise", "comprising", "include", "including", and "contain", "containing" should be construed to be inclusive - i.e., to mean "including, but not limited to", and, therefore, should not be read to be closed or limiting of any type. The use of "consisting" or "consisting of" is intended to mean that the listed elements are exhaustive, and that no other elements are present. The use of "consisting essentially of" or "consisting essentially of" is intended to mean that the listed elements are exhaustive, and that other elements are present, but that the other elements do not materially alter the basic and novel characteristics of the claimed composition, structure, or method. In the case of a process, "consisting essentially of" or "consisting essentially of" means that the process recited includes the steps listed, but also includes other steps that do not materially alter the basic and novel characteristics of the process.
[0029] Some embodiments of the present application will now be described in detail in connection with the accompanying drawings. The following embodiments and features are not mutually exclusive and can be combined with each other.
[0030] Reference will be made to Figure 1 , Figure 1 A block diagram of an electronic device 100 according to an embodiment of the present application is shown in FIG. 1. The electronic device 100 can be, but is not limited to, a server, a robot, etc. The electronic device 100 includes a memory 110, a processor 120, and a communication unit 130. The memory 110, the processor 120, and the communication unit 130 are electrically connected to each other directly or indirectly to enable transmission or interaction of data. For example, the elements can be electrically connected to each other through one or more communication buses or signal lines.
[0031] The memory 110 is configured to store programs or data. The memory 110 can be, but is not limited to, a Random Access Memory (RAM), a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electric Erasable Programmable Read-Only Memory (EEPROM), etc.
[0032] The processor 120 is used to read / write data or programs stored in the memory 110 and execute corresponding functions. For example, the memory 110 stores a robot motion control device 200, which includes at least one software function module that can be stored in the memory 110 in the form of software or firmware. The processor 120 executes various functional applications and data processing by running the software programs and modules stored in the memory 110, such as the robot motion control device 200 in this embodiment, thereby realizing the robot motion control method in this embodiment.
[0033] The communication unit 130 is used to establish a communication connection between the electronic device 100 and other communication terminals through the network, and to send and receive data through the network.
[0034] It should be understood that, Figure 1 The structure shown is only a schematic diagram of the electronic device 100. The electronic device 100 may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown. Figure 1 The components shown can be implemented using hardware, software, or a combination thereof.
[0035] Please refer to Figure 2 , Figure 2 This is one of the flowcharts illustrating a robot motion control method provided in this application. The method is applied to the aforementioned electronic device 100. The specific flow of the robot motion control method is described in detail below. In this embodiment, the method may include steps S110 to S140.
[0036] Step S110: Obtain the first environmental image currently acquired by the target robot and the current position state of the first joint of the target robot.
[0037] In this embodiment, the target robot is the robot for which the current robot motion control is applied. The target robot may include a camera; the specific installation location and number of cameras can be determined based on actual needs and are not specifically limited here. The image currently captured by the camera can be used as the first environmental image currently captured by the target robot. The current joint position state of the target robot can also be obtained as the first joint position state. The first joint position state is used to indicate the current position of each joint of the target robot (i.e., the angle of the motors of each joint). The first joint position state may include the current position of each joint. The first environmental image and the first joint position state are used to describe the environment and the joint state of the target robot at a given moment.
[0038] Step S120, using a pre-trained target feature extraction network based on a channel attention mechanism to perform environment feature extraction on the first environment image to obtain first environment features.
[0039] Step S130, inputting the first environment features and the first joint position state into a prediction network to obtain an initial motion control signal.
[0040] In this embodiment, the present inventors have found that if the first environment image and the first joint position state are directly input into the prediction network to predict the control signal, the subtle features in the complex environment may not be accurately extracted for prediction, which may result in poor prediction of the control signal. The channel attention mechanism adjusts the importance of each channel and improves the focusing ability on key features, so that the network can better process the channel level information of the visual input.
[0041] Based on the above considerations, in this embodiment, when the first environment image is obtained, the first environment image is input into a pre-trained target feature extraction network based on a channel attention mechanism, and the output result of the target feature extraction network based on the first environment image is taken as the first environment features. The target feature extraction network can be determined according to actual needs, and is not specifically limited here.
[0042] When the first environment features are obtained, the first environment features and the first joint position state can be input into a prediction network together, and an initial motion control signal can be obtained from the output result of the prediction network for the first environment features and the first joint position state. The prediction network is trained according to the sample environment features, sample joint position states and sample motion control signals corresponding to each time point. The sample environment features and joint position states are samples, and the sample motion control signals are labels corresponding to the samples. That is, based on the sample environment features and joint position states as samples, the sample motion control signals as labels corresponding to the samples, and based on the sample environment features, sample joint position states and sample motion control signals corresponding to each time point, the prediction network is trained.
[0043] Step S140, controlling the motion of the target robot according to the initial motion control signal.
[0044] In this embodiment, in the case that the initial motion control signal is obtained, the initial motion control signal can be directly used as a signal directly used in control to control the motion of the target robot, so that the joint position state of the target robot at the next moment is the joint position state represented by the initial motion control signal. The initial motion control signal can also be adjusted, and the adjusted signal can be directly used as a signal directly used in control to control the motion of the target robot, so that the joint position state of the target robot at the next moment is the joint position state represented by the adjusted initial motion control signal. The manner in which the motion of the target robot is controlled according to the initial motion control signal can be determined in combination with actual requirements, and is not specifically limited here.
[0045] Optionally, as a possible implementation manner, the head and / or the wrist of the target robot is provided with a camera, and optionally, the wrist camera can only obtain an image from one view angle, or can obtain an image from two different view angles respectively (i.e., the wrist camera can obtain two images at the same time). Correspondingly, the image currently obtained by the robot head camera of the target robot and / or the image currently obtained by the robot wrist camera can be used as the first environment image currently collected by the target robot, that is, the first environment image includes the image data obtained by the head camera and / or the image data obtained by the wrist camera. It can be understood that the image data corresponding to the sample environment features used in the training process of the prediction network is of the same type as the image data of the first environment image, for example, both are image data obtained by the head camera, or both are image data obtained by the wrist camera, or both include image data obtained by the head camera and image data obtained by the wrist camera. In this way, the image data obtained by the head camera and / or the image data obtained by the wrist camera can be used to control the motion of the robot arm of the target robot.
[0046] The current positions of the joints of the target robot can also be obtained in any manner to obtain the first joint position state. It can be understood that if the target robot only includes one joint, the first joint position state only includes the joint position of the joint; if the target robot includes multiple joints, the first joint position state includes the joint positions of the multiple joints. The first environment image and the first joint position state are used to describe the current environment and the state of the target robot. If it is desired to control the motion of the robot arm of the target robot, the first joint position state can include the joint positions of the joints of the robot arm.
[0047] Optionally, the target feature extraction network is an SEnet network or an ECA-Net network.
[0048] In an example implementation manner, as shown inFigure 3 As shown, the target feature extraction network is a SENet network, i.e., a SENet encoder. The SENet encoder uses a channel attention mechanism to process the input image data x (i.e., the first environment image mentioned above, i.e.) Figure 3 SENet encodes the image data (in the image data) to generate a feature vector U (i.e., the first environmental feature mentioned above), which captures important visual information in the environment. SENet adaptively enhances the weights of important channels through two steps: "Squeeze" and "Excitation," and extracts key visual features through "Re-weighting." Specifically, the Squeeze operation compresses the spatial information of channel features into a global description; the Excitation operation enhances important channel features through adaptive weight adjustment; and the Re-weighting operation uses the learned weights to reweight the feature map, resulting in a new feature map. The SENet encoding process can be represented as: U = f tr (x), where f tr This indicates that environmental features U are extracted using the SEnet encoder.
[0049] In this embodiment, the prediction network can be specifically determined based on actual needs. As one possible implementation, the prediction network is a generative model, and the specific generative model can be determined based on actual needs. The inventors of this application have discovered that a Conditional Variational Autoencoder (CVAE) can generate output data that meets specific requirements based on input conditional information, making it very suitable for robot state estimation and control command generation tasks. Therefore, the prediction network can be a CVAE.
[0050] In this embodiment, the prediction network may include an encoder and a decoder. The first environmental feature and the first joint position state can be input into the encoder to generate latent variables based on the first environmental feature and the first joint position state. Then, the latent variables can be input into the decoder to decode based on the generated latent variables, thereby generating the initial motion control signal.
[0051] As one implementation method, such as Figure 3 As shown, the prediction network is a CVAE network, which combines the first environmental feature U and the first joint position state y (i.e., Figure 3the joint (p) in the image U is input into a CVAE encoder to generate a latent variable z, which can be regarded as an abstract representation of the first environment image and the first joint position state, and the process can be represented by the following formula: z ~ q(z|U, y) = N(μ(U, y), σ 2 (U, y)), where μ(U, y) represents the mean value calculated by the encoder (i.e., the mean value of the latent variable), σ 2 (U, y) represents the variance calculated by the encoder (i.e., the variance of the latent variable), and μ(U, y) and σ 2 (U, y) reflect the latent distribution of the visual features and the joint position state. Then, the latent variable z is decoded by a CVAE decoder to generate the initial motion control signal The CVAE encoder and the CVAE decoder are both transform networks.
[0052] Optionally, after obtaining the initial motion control signal, to avoid jitter, a target motion control signal can be calculated by smoothing processing according to a plurality of historical motion control signals and the initial motion control signal, and then the target robot is controlled to move according to the target motion control signal. The joint position state of the target robot at the next moment is the state indicated by the target motion control signal. In this way, the generated initial motion control signal can be optimized to obtain the target motion control signal, so as to ensure the motion accuracy and stability of the target robot. Then, the target robot can be driven to perform an action according to the target motion control signal corresponding to the current environment and joint position state.
[0053] Optionally, as shown in Figure 3 , a moving average filter can be used for smoothing processing to obtain the target motion control signal. When the moving average filter is used for smoothing processing, the joint positions indicated by the latest preset number of historical motion control signals and the initial motion control signal can be obtained, and the average value of the joint positions is taken as the joint position indicated by the target motion control signal.
[0054] In the embodiment, the image data currently acquired by the camera can be encoded by the SEnet, important features in the image are extracted, and the important features are input into the CVAE model together with the current joint position state of the robot to generate a latent representation; then, the CVAE model decodes the latent representation to generate a preliminary motion control signal, the preliminary motion control signal is further processed to generate a final joint control signal, and the robot is driven to perform a precise motion task. The above method is a visual motion control method based on a conditional variational autoencoder (CVAE) and a channel attention mechanism (SEnet), and the combination of the channel attention mechanism of the SEnet and the generation capability of the CVAE can improve the accuracy and adaptability of visual perception and motion control of the robot in a complex environment, and can improve the motion control precision of the robot in the complex environment.
[0055] The visual motion control method based on the conditional variational autoencoder (CVAE) and the channel attention mechanism (SEnet) can generate high-precision control instructions by enhancing the extraction and analysis of important features of image channels in combination with the joint state data of the robot, and is suitable for robot motion control in a complex environment.
[0056] Please refer to Figure 4 , Figure 4 A flowchart of a robot motion control method provided in the embodiment is shown in FIG. 2. In the embodiment, before step S110, the method can further include steps S101-S102.
[0057] In step S101, a plurality of sample data are obtained.
[0058] In step S102, an initial model is trained according to the plurality of sample data to obtain a control strategy generation model.
[0059] In the embodiment, a large amount of original image data and corresponding original joint position states captured by the robot head camera can be collected, and then the data is preprocessed to obtain a plurality of sample data. Each sample data includes a sample environment image, a sample joint position state and a sample motion control signal corresponding to a time. That is, a sample data is used to represent a second environment image at time T, a second joint position state at time T, and a motion control signal at time T. The second environment image at time T is used to describe the environment at time T, the second joint position state at time T is used to describe the state of the robot at time T, and the motion control signal at time T is the motion control signal to be executed by the robot at time T. The above data can be used to train an initial model to obtain a control strategy generation model, so as to ensure that the control strategy generation model can generate high-precision motion control signals in different environments. In the training process, the second environment image and the second joint position state in a sample data are taken as a sample, and the sample motion control signal in the sample data is taken as a label corresponding to the sample. The control strategy generation model includes the target feature extraction network and the prediction network. That is, the control strategy generation model including the target feature extraction network and the prediction network is trained based on the above sample data in an end-to-end manner.
[0060] In the embodiment, the target feature extraction network is an SEnet network, which can be trained to extract important features in an image. SEnet compresses the spatial information of each channel into a global feature description through a "compression" operation, and then dynamically adjusts the weights according to the importance of different channels through an "excitation" operation to optimize the feature extraction capability. Finally, the feature vector U output by SEnet is a compressed feature representation of the image.
[0061] In the embodiment, the prediction network is a CVAE model. The CVAE model generates latent variables z by learning the joint distribution of visual features U and joint position states y, and generates motion control signals through a decoder. The training objective of CVAE is to minimize the reconstruction loss and the KL divergence loss, and the sum of the two losses is as follows:
[0062] Γ CVAE =E q(z|U,y) [logp(y|z)]-D KL ([q(z|U,y)||p(z)])
[0063] Wherein, the reconstruction loss is used to ensure that the joint control signal generated by the CVAE decoder is as close as possible to the real joint control signal; and the KL divergence loss is used to minimize the difference between the latent variable distribution q(z|U,y) and the prior distribution p(z).
[0064] The reconstruction loss and the KL divergence loss can be calculated during the training process in the training training, and the initial model at present is adjusted according to the obtained reconstruction loss and KL divergence loss, so as to obtain the control strategy generation model.
[0065] The joint training can be performed based on the sample data, parameters of the SE net and the CVAE model are optimized, so that the model can generate high-precision control signals in different environments. The SE net is responsible for extracting important features in image data, and the CVAE model generates motion control signals to ensure that the robot can make the best decision based on the current environment and joint state. The effectiveness and robustness of the method in complex environments can also be verified through multiple experiments. The accuracy and performance of the model are evaluated by measuring the error (such as mean square error) between the generated control signal and the real signal.
[0066] In order to perform the corresponding steps in the above-mentioned embodiments and various possible manners, an implementation of a robot motion control device 200 is given below, which can optionally adopt the device structure of the electronic device 100 shown in the above-mentioned Figure 1 Further, please refer to Figure 5 , Figure 5 A block diagram of a robot motion control device 200 provided by the embodiments of the present application is shown. It should be noted that the basic principle and technical effects of the robot motion control device 200 provided by the present embodiment are the same as those of the above-mentioned embodiments. For brief description, the part not mentioned in the present embodiment can refer to the corresponding content in the above-mentioned embodiments. In the present embodiment, the robot motion control device 200 can include a collection module 210, a processing module 220 and a control module 230.
[0067] The collection module 210 is configured to obtain a first environment image currently collected by a target robot and a first joint position state of the target robot.
[0068] The processing module 220 is configured to perform environment feature extraction on the first environment image by using a pre-trained target feature extraction network based on a channel attention mechanism to obtain a first environment feature.
[0069] The processing module 220 is further configured to input the first environment feature and the first joint position state into a prediction network to obtain an initial motion control signal. The prediction network is trained according to a plurality of sample environment features, sample joint position states and sample motion control signals corresponding to each time point. The sample environment features and the joint position states are samples, and the sample motion control signals are labels corresponding to the samples.
[0070] The control module 230 is configured to control the target robot motion according to the initial motion control signal.
[0071] Please refer to Figure 6 , Figure 6 A block diagram of a robot motion control device 200 according to an embodiment of the present application is shown in FIG. 2. In this embodiment, the robot motion control device 200 can further include a training module 201.
[0072] The training module 201 is configured to obtain a plurality of sample data, wherein each sample data includes a sample environment image, a sample joint position state and a sample motion control signal corresponding to a time point; and train an initial model according to the plurality of sample data to obtain a control strategy generation model, wherein the control strategy generation model includes the target feature extraction network and the prediction network.
[0073] Optionally, the above modules can be stored in the memory 110 shown in the form of software or firmware (Firmware) or solidified in the operating system (Operating System, OS) of the electronic device 100, and can be executed by the processor 120 in the electronic device 100. At the same time, the data, program code and the like required for executing the above modules can be stored in the memory 110. Figure 1 Figure 1
[0074] The embodiment of the present application further provides a readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the robot motion control method.
[0075] To sum up, the embodiment of the present application provides a robot motion control method and device, electronic equipment and readable storage medium. Firstly, a first environment image currently obtained by a target robot and a first joint position state of the target robot are obtained. Then, a target feature extraction network based on a channel attention mechanism is used to extract environment features from the first environment image to obtain first environment features. Next, the first environment features and the first joint position state are input into a prediction network to obtain an initial motion control signal. The prediction network is trained according to sample environment features, sample joint position states and sample motion control signals corresponding to each time point. The sample environment features and the joint position state are samples, and the sample motion control signal is a label corresponding to the sample. Finally, the target robot is controlled to move according to the initial motion control signal. In this way, important features in visual input can be accurately extracted by using the target feature extraction network, and then the joint position state is combined to generate a signal for motion control of the robot through the prediction network, thereby improving the adaptability of visual perception and motion control strategy of the robot in a complex environment, reducing the mismatch between the action of the robot and the actual environment, and facilitating the robot to perform accurate motion tasks.
[0076] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can also be implemented in other manners. The above described apparatus embodiments are only schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible implementation architectures, functions and operation of the apparatus, method and computer program product according to the embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a segment or a portion of code which comprises one or more executable instructions for implementing the specified logic function. It should also be noted that in some alternative implementations, the functions shown in the blocks can occur in different orders from those shown in the figures. For example, two blocks shown in succession can in fact be executed substantially concurrently or in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts and combinations of blocks in the block diagrams and / or flowcharts can be implemented by dedicated hardware-based systems which perform the specified functions or operations, or can be implemented by a combination of dedicated hardware-based systems and computer instructions.
[0077] In addition, each functional module in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0078] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts of the prior art that make contributions or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0079] The above only describes optional embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A robot motion control method characterized by, The method comprises: obtaining a first environment image currently collected by a target robot and a first joint position state of the target robot at present; extracting an environment feature of the first environment image by using a pre-trained target feature extraction network based on a channel attention mechanism to obtain a first environment feature; inputting the first environment feature and the first joint position state into a prediction network to obtain an initial motion control signal, wherein the prediction network is trained according to sample environment features, sample joint position states and sample motion control signals corresponding to multiple time points, the sample environment features and the sample joint position states are samples, and the sample motion control signals are labels corresponding to the samples; controlling motion of the target robot according to the initial motion control signal; wherein the prediction network comprises an encoder and a decoder, and the inputting of the first environment feature and the first joint position state into the prediction network to obtain the initial motion control signal comprises: generating a latent variable based on the first environment feature and the first joint position state by using the encoder; generating the initial motion control signal based on the generated latent variable by using the decoder.
2. The method of claim 1, wherein, The controlling of the motion of the target robot according to the initial motion control signal comprises: calculating a target motion control signal by smoothing processing according to multiple historical motion control signals and the initial motion control signal; controlling the motion of the target robot according to the target motion control signal, wherein a joint position state of the target robot at a next moment is a state indicated by the target motion control signal.
3. The method of claim 1, wherein, The method further comprises: obtaining a plurality of sample data, wherein each piece of sample data comprises a sample environment image, a sample joint position state and a sample motion control signal corresponding to a time point; training an initial model according to the plurality of sample data to obtain a control strategy generation model, wherein the control strategy generation model comprises the target feature extraction network and the prediction network.
4. The method of claim 3, wherein, The prediction network is a conditional variational autoencoder, and the training of the initial model according to the plurality of sample data to obtain the control strategy generation model comprises: in the training process, calculating a reconstruction loss and a KL divergence loss, and adjusting the current initial model according to the obtained reconstruction loss and the KL divergence loss to obtain the control strategy generation model.
5. The method of claim 1, wherein, The prediction network is a conditional variational autoencoder, and / or the target feature extraction network is an SEnet network or an ECA-Net network.
6. The method according to any one of claims 1 to 5, characterized in that, The first environment image and the image corresponding to the sample environment feature each comprise an image obtained by a robot head camera and / or an image obtained by a robot wrist camera.
7. A robot motion control apparatus characterized by comprising: The device comprises: a collection module configured to obtain a first environment image currently collected by a target robot and a first joint position state of the target robot at present; a processing module configured to extract an environment feature of the first environment image by using a pre-trained target feature extraction network based on a channel attention mechanism to obtain a first environment feature; The processing module is further configured to input the first environment feature and the first joint position state into a prediction network to obtain an initial motion control signal, wherein the prediction network is trained according to a plurality of time points each corresponding to sample environment features, sample joint position states and sample motion control signals, the sample environment features and the sample joint position states are samples, and the sample motion control signals are labels corresponding to the samples. A control module is configured to control motion of the target robot according to the initial motion control signal. The prediction network includes an encoder and a decoder, and the processing module is specifically configured to generate a latent variable based on the first environment feature and the first joint position state by using the encoder, and generate the initial motion control signal based on the generated latent variable by using the decoder.
8. An electronic device, comprising: The computer program is executed by a processor to implement the robot motion control method according to any one of claims 1-6.
9. A readable storage medium, having stored thereon a computer program, characterized in that, The computer program is executed by a processor to implement the robot motion control method according to any one of claims 1-6.
Citation Information
Patent Citations
Human body motion track prediction method and system based on multi-output space-time interaction
CN117474945A
System and method for navigating a vehicle using language instructions
WO2021058090A1