Method and device for controlling robot, medium and electronic equipment
By using trained vision models and converter models to process the robot's three-view pictures and target instructions, the control problem of the robot when understanding and performing tasks is solved, and more efficient and accurate robot control is achieved.
Patent Information
- Application Number
- CN202510194616.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has not yet solved the problem of robots' general artificial intelligence (AGI) thinking ability in robots in the field of robots, resulting in robots lacking efficient and accurate control when understanding target instructions and performing tasks.
By obtaining the robot's three-view picture and target instructions, using the trained visual model and converter model, the visual vector and text vector are converted into target values to accurately control the robot's movement.
It has achieved better understanding and more accurate execution of target instructions by the robot, significantly improving the intelligence and automation of the robot.
Smart Images

Figure CN119973990A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of robot control, and in particular, relates to a method for controlling a robot, an apparatus for controlling a robot, a non-transitory computer-readable storage medium, and an electronic device. Background Art
[0002] The success of large language models (LLMs) such as Chat GPT (Chat Generative Pre-trained Transformer, a chatbot program) has been a resounding success. Their demonstrated few-shot and even zero-shot learning capabilities seem to point to the dawn of artificial general intelligence (AGI). Cross-modal models such as CLIP (Contrastive Language-Image Pre-training, a multimodal model of images and language) have bridged the gap between natural language processing (NLP) and computer vision (CV). These models, such as CLIP's VLM (Vision-Language Model), are further advancing the development of AGI.
[0003] Therefore, considering that AGI can also bring significant potential for improvement in the field of robotics, if robots possess the thinking ability of LLM, the level of intelligence in the world will be greatly improved and advanced. However, at present, there is no relevant technology to solve this problem. Summary of the Invention
[0004] To overcome the problems existing in the related art, the present disclosure provides a method for controlling a robot, an apparatus for controlling a robot, a non-transitory computer-readable storage medium, and an electronic device.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a method for controlling a robot, the method comprising: Obtaining three-view images of the robot and obtaining target instructions for controlling the robot; Inputting the three-view images into a trained visual model so that the trained visual model outputs a visual vector; A text vector corresponding to the target instruction is obtained, and the visual vector and the text vector are input into a trained converter model so that the trained converter model outputs a target value, and the target value is used to control the movement of the robot.
[0006] Optionally, the three-view pictures include: a left view picture, a top view picture and a front view picture of the entire robot.
[0007] Optionally, the vision model includes: a Vision Transformer model; the converter model includes: a Transformer model.
[0008] Optionally, before inputting the three-view images into the trained visual model, the method further includes: Obtain three-view samples, instruction samples and corresponding numerical samples; The visual model to be trained and the converter model to be trained are trained using the three-view samples, the instruction samples and the numerical samples to obtain a trained visual model and a trained converter model.
[0009] Optionally, the using the three-view samples, the instruction samples, and the numerical samples to train the visual model to be trained and the converter model to be trained to obtain the trained visual model and the trained converter model includes: Obtaining a first vector corresponding to the instruction sample, and inputting the three-view sample into a visual model to be trained, so that the visual model to be trained outputs a second vector corresponding to the three-view sample; The first vector and the second vector are input into the converter model to be trained, so that the converter model to be trained outputs the numerical sample to obtain a trained visual model and a trained converter model.
[0010] Optionally, the target values include: a first group of values, a second group of values and a third group of values, the first group of values representing the posture of the robot, the second group of values representing the state of the robot, and the third group of values representing the process of the robot's movement.
[0011] Optionally, inputting the visual vector and the text vector into a trained converter model so that the trained converter model outputs a target value includes: Inputting the visual vector and the text vector into a trained converter model, so that an encoder in the trained converter model converts the visual vector and the text vector to obtain a sequence vector; The decoder in the trained converter model is used to perform vector prediction processing on the sequence vector to obtain a target value.
[0012] According to a second aspect of an embodiment of the present disclosure, there is provided a device for controlling a robot, comprising: a data acquisition module configured to acquire three-view images of the robot and obtain target instructions for controlling the robot; a visual vector module, configured to input the three-view image into a trained visual model so that the trained visual model outputs a visual vector; The numerical output module is configured to obtain a text vector corresponding to the target instruction, and input the visual vector and the text vector into a trained converter model so that the trained converter model outputs a target numerical value, which is used to control the movement of the robot.
[0013] According to a third aspect of an embodiment of the present disclosure, a non-temporary computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of any one of the methods described in the first aspect of the present disclosure are implemented.
[0014] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including: processor; a memory for storing processor-executable instructions; The processor is configured to: execute the executable instructions to implement the steps of any one of the methods described in the first aspect of the present disclosure.
[0015] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects: In the method and apparatus provided by the exemplary embodiments of the present disclosure, a trained visual model is used to process three-view images of a robot, enabling the visual model to more clearly and accurately understand the robot and its environment. Furthermore, a trained converter model is used to process target commands to determine the robot's actions. This large model is applied to the field of robot control, enabling accurate control of the robot, enabling it to better understand target commands and respond appropriately. This allows for more precise control of the robot's execution of tasks, significantly enhancing the robot's intelligence and automation levels, and possesses extremely important practical significance and application value.
[0016] Other features and advantages of the present disclosure will be described in detail in the following detailed description.
[0017] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following detailed description, they are used to explain the present disclosure but do not constitute a limitation of the present disclosure. In the accompanying drawings: Figure 1 The following schematically illustrates a flow chart of a method for controlling a robot in an exemplary embodiment of the present disclosure; Figure 2 A flowchart of a method for training a visual model and a converter model in an exemplary embodiment of the present disclosure is schematically shown; Figure 3 A flowchart of a method for further training a visual model and a converter model in an exemplary embodiment of the present disclosure is schematically shown; Figure 4 Schematically illustrates a flow chart of processing of a trained converter model in an exemplary embodiment of the present disclosure; Figure 5 A model framework diagram of a method for controlling a robot in an application scenario in an exemplary embodiment of the present disclosure is schematically shown; Figure 6 A schematic structural diagram of a device for controlling a robot in an exemplary embodiment of the present disclosure is shown schematically; Figure 7 Schematically illustrates an electronic device for implementing a method for controlling a robot in an exemplary embodiment of the present disclosure; Figure 8 Another electronic device for implementing a method for controlling a robot in an exemplary embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0019] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0020] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the corresponding data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0021] The present disclosure provides a method for controlling a robot. Figure 1 FIG. 1 is a flow chart of a method for controlling a robot according to an exemplary embodiment. Figure 1 As shown, the method may include at least the following steps: Step S110: Obtain three-view images of the robot and obtain target instructions for controlling the robot.
[0022] Step S120: Input the three-view images into the trained visual model, so that the trained visual model outputs a visual vector.
[0023] Step S130. Obtain a text vector corresponding to the target instruction, and input the visual vector and the text vector into the trained converter model so that the trained converter model outputs a target value, which is used to control the movement of the robot.
[0024] In the exemplary embodiments of the present disclosure, a trained visual model is used to process three-view images of a robot, enabling the visual model to more clearly and accurately understand the robot and its environment. Simultaneously, a trained converter model is used to process target commands to determine the robot's actions. This large model is applied to the field of robot control, enabling accurate control of the robot, enabling it to better understand target commands and respond correctly. This allows for more precise control of the robot's execution of tasks, significantly improving the robot's intelligence and automation levels, and possesses extremely important practical significance and application value.
[0025] The following describes in detail the various steps of the method for controlling the robot.
[0026] In step S110 , a three-view image of the robot is obtained, and a target instruction for controlling the robot is obtained.
[0027] In an exemplary embodiment of the present disclosure, the robot may be a robotic arm or a robot with a robotic arm device, or any other form of robot, which is not particularly limited in this exemplary embodiment.
[0028] In an optional embodiment, the three-view picture includes: a left view picture, a top view picture and a front view picture of the entire robot.
[0029] It is worth noting that the left view image, top view image and main view image can be collected by cameras and other equipment, and are all obtained by shooting the entire robot, and cannot be obtained by shooting a certain part of the robot.
[0030] The target instruction may be an instruction issued by a user to control the action of the robot, such as “Move red cube”.
[0031] In step S120 , the three-view images are input into the trained visual model, so that the trained visual model outputs a visual vector.
[0032] In an exemplary embodiment of the present disclosure, after obtaining the three-view images of the robot and the target instructions for controlling the robot, the three-view images can be input into the trained visual model, and the target instructions can be input into the trained converter model accordingly.
[0033] Before this, the corresponding vision model and converter model can be trained.
[0034] In an optional embodiment, the vision model includes: a Vision Transformer model; the converter model includes: a Transformer model.
[0035] Among them, Vision Transformer (ViT) is an image classification model.
[0036] The Transformer (an attention-based sequence model) was originally used primarily in the field of Natural Language Processing (NLP). In 2020, a paper (An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale) brought the Transformer from NLP to the field of Computer Vision (CV), achieving success in multiple visual tasks.
[0037] The main steps of the ViT model include the following: Patch Embedding (an embedding method used in natural language processing tasks): First, the original input image is sliced into pieces.
[0038] Assume the input image size is 224×224. Cut the image into fixed-size 16×16 blocks, with each block being a patch. Therefore, the number of patches in each image is (224×224) / (16×16) = 196. After cutting, we obtain 196 patches of [16, 16, 3]. These patches can be fed into the Linear Projection of Flattened Patches (Embedding layer). This layer flattens the input sequence, resulting in 196 tokens in the output. Each token has a flattened dimension of 16×16×3 = 768, resulting in an output dimension of [196,768]. It's easy to see that Patch Embedding transforms a CV problem into an NLP problem through cutting and flattening.
[0039] Position Embedding (embedding position information into word embedding vector): Each patch of the image, like the text, has a sequence and cannot be arbitrarily disrupted, so it is necessary to add position information to each token.
[0040] Similar to the BERT (Bidirectional Encoder Representation from Transformers) model, a pre-trained language representation model also requires the addition of a special character class token. Therefore, the final sequence dimension input to the Transformer Encoder is [197, 768]. Position Embedding adds position information.
[0041] Transformer Encoder: Input the sequence of shape [197, 768] into the standard Transformer Encoder.
[0042] MLP (Multilayer Perceptron) Head: The output of the Transformer Encoder is actually a sequence, but the ViT model only uses the class token output, which is fed into the MLP module and ultimately outputs the classification result. The MLP Head is used for the final classification.
[0043] In an alternative embodiment, Figure 2 A flowchart of a method for training a visual model and a converter model is shown, as Figure 2 As shown, the method may at least include the following steps: in step S210, three-view image samples, instruction samples and corresponding numerical samples are obtained.
[0044] The three-view video data before each action of the robot is recorded to obtain three-view samples, and the corresponding text instructions are recorded as instruction samples.
[0045] Correspondingly, the variables of the robot motion are recorded one by one as numerical samples to form a robot control data set.
[0046] In step S220 , the visual model to be trained and the converter model to be trained are trained using the three-view samples, the instruction samples and the numerical samples to obtain a trained visual model and a trained converter model.
[0047] In an alternative embodiment, Figure 3A flowchart of a method for further training the visual model and the converter model is shown, as Figure 3 As shown, the method may at least include the following steps: in step S310, a first vector corresponding to the instruction sample is obtained, and the three-view sample is input into the visual model to be trained, so that the visual model to be trained outputs a second vector corresponding to the three-view sample.
[0048] Generally, the first vector can be obtained by mapping the instruction sample through tokenization technology and / or embedding layer, etc. In practical applications, other technologies or methods can also be used to obtain the first vector, and this exemplary embodiment does not specifically limit this.
[0049] When the three-view samples, the instruction samples, and the corresponding numerical samples are used for model training, the three-view samples may be input into the visual model to be trained, so that the visual model to be trained outputs a second vector.
[0050] It is worth noting that the sizes of the second vector and the first vector should be consistent or matched and conform to the size that can be used by the converter model.
[0051] In step S320 , the first vector and the second vector are input into the converter model to be trained, so that the converter model to be trained outputs numerical samples to obtain a trained visual model and a trained converter model.
[0052] Furthermore, the first vector and the second vector are input into the converter model to be trained, and parameters such as weights and biases of the converter model to be trained are adjusted so that the converter model to be trained outputs numerical samples corresponding to the three-view samples and the instruction samples, and the training is terminated. At this point, a trained visual model and a trained converter model can be obtained.
[0053] In this exemplary embodiment, by training the visual model and the converter model, the trained visual model and converter model can accurately control the robot according to the three-view images and target instructions.
[0054] In order to verify the effectiveness of the trained visual model and the trained converter model, the three-view samples, instruction samples and corresponding numerical samples can be divided into two parts during the training process, namely the training set and the test set, and the division ratio of the training set and the test set can be 95% and 5%.
[0055] Through data verification on the test set, it can be found that the accuracy of the trained vision model and the trained converter model on the test set is 83%.
[0056] Therefore, the trained vision model and the trained converter model can correctly control the robot to perform tasks based on text instructions and combined with three-view data.
[0057] After the visual model and the converter model are trained, the three-view image can be input into the trained visual model so that the trained visual model can calculate and obtain the visual vector.
[0058] In step S130, a text vector corresponding to the target instruction is obtained, and the visual vector and the text vector are input into the trained converter model so that the trained converter model outputs a target value, which is used to control the movement of the robot.
[0059] In an exemplary embodiment of the present disclosure, a text vector may be obtained by mapping the target instruction using tokenization technology and / or an embedding layer. In practical applications, other technologies may also be used to obtain a text vector, and this exemplary embodiment does not specifically limit this.
[0060] In an alternative embodiment, Figure 4 A schematic diagram of the process of processing the trained converter model is shown in Figure 4 As shown, the method may at least include the following steps: in step S410, the visual vector and the text vector are input into the trained converter model, so that the encoder in the trained converter model converts the visual vector and the text vector to obtain a sequence vector.
[0061] The Transformer model has become a mainstream deep learning architecture. Its powerful ability to process complex language phenomena is due to its unique encoder and decoder structure.
[0062] Transformer is an "encoder-decoder" architecture consisting of an encoder and a decoder, both of which are superpositions of multi-head self-attention modules.
[0063] The encoder is a crucial component of the Transformer model, its primary task being to capture the semantic information of the input sequence. In the encoder, each input word is converted into a fixed-dimensional vector representation through an embedding layer. These vectors are then processed through multiple self-attention layers and feed-forward neural network layers to capture inter-word dependencies and semantic information.
[0064] The encoder is composed of multiple identical layers stacked together, each with two sublayers: the first is a multi-head self-attention convergence layer; the second is a position-based feedforward neural network. Each sublayer uses residual connections.
[0065] Therefore, by inputting the visual vector and text vector into the converter model, the encoder can perform word embedding representation on the input sequence and add position information.
[0066] After that, the input sequence of the current encoder layer is fed into the multi-head self-attention layer to generate a new vector. Specifically, when calculating the encoder’s self-attention, the query, key, and value all come from the output of the previous encoder layer.
[0067] The output of the multi-head self-attention layer is connected to the input of the current encoder layer through a residual connection, and the result is layer normalized.
[0068] The result of layer normalization is placed in the fully connected layer (feed forward). This layer transforms all position representations output by the self-attention layer, so it is called a position-based feedforward neural network.
[0069] After that, residual connection is performed and layer normalization is performed.
[0070] Furthermore, the result is sent to the next encoder layer and repeated N times to obtain a sequence vector.
[0071] The encoder can not only capture long-range dependencies but also perform efficient computation.
[0072] Specifically, the encoder can capture the long-range dependencies between words in the input sequence through the self-attention mechanism, which helps to understand the overall semantics of the sentence.
[0073] The encoder uses a self-attention mechanism for calculation. Compared with the traditional recurrent neural network (RNN), this calculation method is more efficient and can avoid the problem of gradient vanishing or gradient exploding when processing long sequences.
[0074] In step S420, the decoder in the trained converter model is used to perform vector prediction processing on the sequence vector to obtain a target value.
[0075] The decoder is the core component of the Transformer model. Its primary task is to generate a new output sequence based on the processed input sequence. The decoder receives the output sequence from the encoder and then performs multiple rounds of predictions through self-attention layers and feedforward neural network layers to generate a new output sequence. Each prediction step depends on all previous predictions, which enables the decoder to capture more complex language phenomena.
[0076] The decoder receives the final sequence vector from the encoder.
[0077] When the decoder receives the [BEGIN]Token, it generates the first character, then generates the second character from the first character, then generates the third character from the first two characters, then generates the first three characters, and then generates the token from the first four characters, until the sequence generation ends.
[0078] Therefore, after the sequence passes through the decoder, the generated vector will be classified using softmax (normalized exponential function), and the generated characters will be looked up in the vocabulary to obtain the target value.
[0079] The decoder can not only generate coherent output but also capture contextual information.
[0080] Specifically, since the decoder's prediction at each step depends on all previous predictions, it can generate coherent output sequences, which is very important in many NLP tasks.
[0081] The decoder can capture the impact of each word in the input sequence on the current output through the self-attention mechanism, thereby better understanding the contextual information.
[0082] In an optional embodiment, the target values include: a first group of values, a second group of values and a third group of values, the first group of values represents the posture of the robot, the second group of values represents the state of the robot, and the third group of values represents the process of the robot's movement.
[0083] The first set of values may include three vectors in the x, y, and z dimensions representing the position of the robot, and may also include three vectors in the flip, pitch, and sway dimensions representing the posture of the robot.
[0084] The second set of values may include a vector of the open and closed states of the clamp.
[0085] The third set of values may include a vector indicating whether the process has ended.
[0086] The following describes in detail the method for controlling a robot in an embodiment of the present disclosure in conjunction with an application scenario.
[0087] Figure 5 The model framework diagram of the method for controlling the robot in the application scenario is shown as follows: Figure 5 As shown, in step S510, three views.
[0088] Obtain three-view images of the robot. The three-view images may include a left view image, a top view image, and a front view image of the entire robot.
[0089] It is worth noting that the left view image, top view image and main view image can be collected by cameras and other equipment, and are all obtained by shooting the entire robot, and cannot be obtained by shooting a certain part of the robot.
[0090] In step S520, move the red cube.
[0091] Obtain a target instruction for controlling the robot. The target instruction may be an instruction issued by the user to control the robot's movement, such as "Move red cube".
[0092] In step S530, ViT.
[0093] After the visual model and the converter model are trained, the three-view image can be input into the trained visual model so that the trained visual model can calculate and obtain the visual vector.
[0094] In step S540, Transformer Encoder.
[0095] The target instruction is mapped and processed to obtain a text vector by using tokenization technology and / or embedding layer, etc. In practical applications, other technologies can also be used to obtain the text vector, which is not particularly limited in this exemplary embodiment.
[0096] The visual vector and the text vector are input into the trained converter model, so that the encoder in the trained converter model converts the visual vector and the text vector to obtain a sequence vector.
[0097] In step S550, Transformer Decoder.
[0098] The decoder in the trained converter model is used to perform vector prediction on the sequence vector to obtain the target value.
[0099] In order to verify the effectiveness of the trained visual model and the trained converter model, the three-view samples, instruction samples and corresponding numerical samples can be divided into two parts during the training process, namely the training set and the test set, and the division ratio of the training set and the test set can be 95% and 5%.
[0100] Through data verification on the test set, it can be found that the accuracy of the trained vision model and the trained converter model on the test set is 83%.
[0101] Therefore, the trained vision model and the trained converter model can correctly control the robot to perform tasks based on text instructions and combined with three-view data.
[0102] In step S560, Robot Action: 1.2, 2.3, 1.0, 0.5, 0.8, 1.1, closed.
[0103] Generally, the target values may include a first set of values, a second set of values, and a third set of values.
[0104] The first set of values represents the robot's posture, the second set of values represents the robot's state, and the third set of values represents the robot's movement process.
[0105] The first set of values may include three vectors in the x, y, and z dimensions representing the position of the robot, and may also include three vectors in the flip, pitch, and sway dimensions representing the posture of the robot.
[0106] The second set of values may include a vector of the open and closed states of the clamp.
[0107] The third set of values may include a vector indicating whether the process has ended.
[0108] It is worth noting that since the step process has not ended, the third set of values is not listed in step S560. Even if the step process ends, the third set of values may not be listed in step S560 and the entire process may end directly.
[0109] It can be seen that the third set of values may not necessarily be actually included in the target values, and the target values should represent the state that the third set of values are intended to represent.
[0110] Furthermore, the robot performs actions according to the above-mentioned Robot Action data.
[0111] After the robot completes the action, it uses the camera to take another photo to generate a new round of three-view data and pass it to the model.
[0112] The above steps are repeated until the model output process is completed.
[0113] In the exemplary embodiments of the present disclosure, a trained visual model is used to process three-view images of a robot, enabling the visual model to more clearly and accurately understand the robot and its environment. Simultaneously, a trained converter model is used to process target commands to determine the robot's actions. This large model is applied to the field of robot control, enabling accurate control of the robot, enabling it to better understand target commands and respond correctly. This allows for more precise control of the robot's execution of tasks, significantly improving the robot's intelligence and automation levels, and possesses extremely important practical significance and application value.
[0114] In addition, in an exemplary embodiment of the present disclosure, a device for controlling a robot is also provided. Figure 6 The schematic diagram of the structure of the device for controlling the robot is shown in FIG. Figure 6 As shown, the device 600 for controlling a robot may include: a data acquisition module 610, a visual vector module 620, and a numerical output module 630. Among them: The data acquisition module 610 is configured to acquire three-view images of the robot and obtain target instructions for controlling the robot; a visual vector module 620 configured to input the three-view image into a trained visual model so that the trained visual model outputs a visual vector; The numerical output module 630 is configured to obtain a text vector corresponding to the target instruction, and input the visual vector and the text vector into the trained converter model so that the trained converter model outputs a target numerical value, which is used to control the movement of the robot.
[0115] In an exemplary embodiment of the present invention, the three-view pictures include: a left view picture, a top view picture and a front view picture of the entire robot.
[0116] In an exemplary embodiment of the present invention, the vision model includes: a Vision Transformer model; the converter model includes: a Transformer model.
[0117] In an exemplary embodiment of the present invention, the device 600 for controlling a robot includes: A sample acquisition module is configured to acquire three-view samples, instruction samples, and corresponding numerical samples; The model training module is configured to train the visual model to be trained and the converter model to be trained using the three-view samples, the instruction samples and the numerical samples to obtain a trained visual model and a trained converter model.
[0118] In an exemplary embodiment of the present invention, the model training module includes: a visual model submodule configured to obtain a first vector corresponding to the instruction sample, and input the three-view sample into a visual model to be trained, so that the visual model to be trained outputs a second vector corresponding to the three-view sample; The language model submodule is configured to input the first vector and the second vector into the converter model to be trained, so that the converter model to be trained outputs the numerical sample to obtain a trained visual model and a trained converter model.
[0119] In an exemplary embodiment of the present invention, the target values include: a first group of values, a second group of values and a third group of values, the first group of values represents the posture of the robot, the second group of values represents the state of the robot, and the third group of values represents the process of the robot's movement.
[0120] In an exemplary embodiment of the present invention, the numerical output module 630 includes: an encoding processing submodule, configured to input the visual vector and the text vector into a trained converter model, so that an encoder in the trained converter model converts the visual vector and the text vector to obtain a sequence vector; The decoding processing submodule is configured to use the decoder in the trained converter model to perform vector prediction processing on the sequence vector to obtain a target value.
[0121] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0122] Figure 7 FIG. 7 is a block diagram of an electronic device 700 according to an exemplary embodiment. Figure 7 As shown, the electronic device 700 may include: a processor 701 , a memory 702 , and may further include one or more of a multimedia component 703 , an input / output (I / O) interface 704 , and a communication component 705 .
[0123] The processor 701 is used to control the overall operation of the electronic device 700 to complete all or part of the steps in the above-mentioned method for controlling a robot. The memory 702 is used to store various types of data to support the operation of the electronic device 700. This data may include, for example, instructions for any application or method operating on the electronic device 700, as well as application-related data such as contact information, sent and received messages, images, audio, video, etc. The memory 702 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 703 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 702 or transmitted via the communication component 705. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 704 provides an interface between the processor 701 and other interface modules, which may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 705 is used for wired or wireless communication between the electronic device 700 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more thereof, is not limited here. Therefore, the corresponding communication component 705 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.
[0124] In an exemplary embodiment, the electronic device 700 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the above-mentioned method of controlling a robot.
[0125] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When executed by a processor, the program instructions implement the steps of the above-described method for controlling a robot. For example, the computer-readable storage medium may be the aforementioned memory 702 including the program instructions. The program instructions may be executed by the processor 701 of the electronic device 700 to perform the above-described method for controlling a robot.
[0126] Figure 8 8 is a block diagram of an electronic device 800 according to an exemplary embodiment. For example, the electronic device 800 can be provided as a server. Figure 8 The electronic device 800 includes a processor 822, which may be one or more, and a memory 832 for storing a computer program executable by the processor 822. The computer program stored in the memory 832 may include one or more modules, each corresponding to a set of instructions. In addition, the processor 822 may be configured to execute the computer program to perform the above-mentioned method of controlling a robot.
[0127] In addition, the electronic device 800 may further include a power supply component 826 and a communication component 850. The power supply component 826 may be configured to perform power management of the electronic device 800, and the communication component 850 may be configured to enable communication, such as wired or wireless communication, of the electronic device 800. In addition, the electronic device 800 may further include an input / output (I / O) interface 858. The electronic device 800 may operate based on an operating system stored in the memory 832.
[0128] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When executed by a processor, the program instructions implement the steps of the above-described method for controlling a robot. For example, the non-transitory computer-readable storage medium may be the aforementioned memory 832 including the program instructions. The program instructions may be executed by the processor 822 of the electronic device 800 to perform the above-described method for controlling a robot.
[0129] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program executable by a programmable device, and has code portions for performing the above-mentioned method of controlling a robot when executed by the programmable device.
[0130] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details of the above embodiments. Within the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the scope of protection of the present disclosure.
[0131] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present disclosure will not further describe various possible combinations.
[0132] In addition, the various embodiments of the present disclosure may be arbitrarily combined, and as long as they do not violate the concept of the present disclosure, they should also be regarded as the contents disclosed by the present disclosure.
Claims
1. A method for controlling a robot, characterized in that: The method comprises: Obtaining three-view images of the robot and obtaining target instructions for controlling the robot; Inputting the three-view images into a trained visual model so that the trained visual model outputs a visual vector; A text vector corresponding to the target instruction is obtained, and the visual vector and the text vector are input into a trained converter model so that the trained converter model outputs a target value, and the target value is used to control the movement of the robot.
2. The method for controlling a robot according to claim 1, characterized in that: The three-view pictures include: a left view picture, a top view picture and a front view picture of the entire robot.
3. The method for controlling a robot according to claim 1, characterized in that: The visual model includes: a VisionTransformer model; the converter model includes: a Transformer model.
4. The method for controlling a robot according to claim 1, characterized in that: Before inputting the three-view images into the trained visual model, the method further includes: Obtain three-view image samples, instruction samples and corresponding numerical samples; The visual model to be trained and the converter model to be trained are trained by using the three-view samples, the instruction samples and the numerical samples to obtain a trained visual model and a trained converter model.
5. The method for controlling a robot according to claim 4, characterized in that: The method of using the three-view samples, the instruction samples and the numerical samples to train the visual model to be trained and the converter model to be trained to obtain the trained visual model and the trained converter model includes: Acquire a first vector corresponding to the instruction sample, and input the three-view sample into a visual model to be trained, so that the visual model to be trained outputs a second vector corresponding to the three-view sample; The first vector and the second vector are input into the converter model to be trained, so that the converter model to be trained outputs the numerical sample to obtain a trained visual model and a trained converter model.
6. The method for controlling a robot according to claim 1, characterized in that: The target values include: a first group of values, a second group of values and a third group of values, the first group of values represents the posture of the robot, the second group of values represents the state of the robot, and the third group of values represents the movement process of the robot.
7. The method for controlling a robot according to claim 1, characterized in that: The step of inputting the visual vector and the text vector into a trained converter model so that the trained converter model outputs a target value includes: Inputting the visual vector and the text vector into a trained converter model, so that an encoder in the trained converter model converts the visual vector and the text vector to obtain a sequence vector; The decoder in the trained converter model is used to perform vector prediction processing on the sequence vector to obtain a target value.
8. A device for controlling a robot, characterized in that: include: A data acquisition module is configured to acquire three-view images of the robot and acquire target instructions for controlling the robot; A visual vector module, configured to input the three-view image into a trained visual model so that the trained visual model outputs a visual vector; The numerical output module is configured to obtain a text vector corresponding to the target instruction, and input the visual vector and the text vector into a trained converter model so that the trained converter model outputs a target numerical value, and the target numerical value is used to control the movement of the robot.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.
10. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Robot control method and device, electronic equipment, storage medium and product
CN121989233A
A robot control method and device, electronic equipment, storage medium and product
CN121989233B