A manipulator control method, device, equipment and storage medium based on a large vision model
Through the robot arm control method based on the visual big model, the visual segmentation big model, multi-view attention network and action prediction network are used to process three-dimensional spatial information, which solves the problem of excessive computing resources consumption of existing methods and realizes efficient robot arm control.
Patent Information
- Application Number
- CN202410326154.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-03-21
AI Technical Summary
The existing three-dimensional space robotic arm control method has a large amount of data characterization, resulting in excessive computing resources consumption and low efficiency, making it difficult to widely use in large-scale embodied data sets.
Using a robot arm control method based on a visual big model, the description text of the target task and the scene images of multiple perspectives are processed by visual segmentation big model, multi-view attention network and action prediction network to obtain an action sequence, and the robot arm performs the task based on the action sequence.
It reduces the calculation amount of three-dimensional space robotic arm control, improves control efficiency, and is suitable for large-scale embodied data sets.
Smart Images

Figure CN118143940B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of embodied intelligence technology, and in particular, to a method, device, equipment and storage medium for controlling a robotic arm based on a large vision model. Background Art
[0002] In embodied intelligence, a robotic arm needs to understand objects in a three-dimensional space and make accurate action selections under language instructions during general object operations. Specifically, in complex three-dimensional operation tasks, the robotic arm needs to understand the physical structure of the working space, such as the position, posture and shape of objects, the blocking relationship between objects, the relationship between objects and the environment, etc., and then perform accurate robotic arm action modeling based on the environmental information in the three-dimensional space. Existing three-dimensional space robotic arm control methods mainly model spatial information by using 3D voxels as representations, and imitate and learn human expert data to learn action strategies, so as to achieve embodied robotic arm control. However, due to the fact that the data volume of 3D voxel representations increases cubically with the size of the space, this method requires a large amount of computing resources and has low efficiency, making it difficult to be widely applied to large-scale embodied datasets. Summary of the Invention
[0003] The embodiments of the present invention provide a method, device, equipment and storage medium for controlling a robotic arm based on a large vision model, which can reduce the computational amount of three-dimensional space robotic arm control and improve the control efficiency.
[0004] In a first aspect, the embodiments of the present invention provide a method for controlling a robotic arm based on a large vision model, including:
[0005] Obtaining a description text of a target task and scene images of multiple perspectives thereof;
[0006] Inputting the description text of the target task and the scene images of multiple perspectives thereof into an action prediction model to obtain an action sequence; wherein, the action prediction model includes: a large vision segmentation model, a multi-perspective attention network and an action prediction network; the action sequence includes multiple action pose information;
[0007] Controlling the robotic arm based on the action sequence to execute the target task.
[0008] In a second aspect, the embodiments of the present invention further provide a device for controlling a robotic arm based on a large vision model, including:
[0009] A scene image acquisition module, configured to obtain a description text of a target task and scene images of multiple perspectives thereof;
[0010] An action sequence acquisition module, configured to input the description text of the target task and the scene images of multiple perspectives thereof into an action prediction model to obtain an action sequence; wherein, the action prediction model includes: a visual segmentation large model, a multi-perspective attention network, and an action prediction network; the action sequence includes multiple action pose information;
[0011] A robotic arm control module, configured to control a robotic arm to execute the target task based on the action sequence.
[0012] In a third aspect, an embodiment of the present invention further provides an electronic device, where the electronic device includes:
[0013] At least one processor; and
[0014] A memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor can execute the robotic arm control method based on a visual large model according to the embodiment of the present invention.
[0016] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the robotic arm control method based on a visual large model according to the embodiment of the present invention is implemented.
[0017] An embodiment of the present invention discloses a robotic arm control method, device, equipment, and storage medium based on a visual large model. Obtain the description text of the target task and the scene images of multiple perspectives thereof; input the description text of the target task and the scene images of multiple perspectives thereof into an action prediction model to obtain an action sequence; wherein, the action prediction model includes: a visual segmentation large model, a multi-perspective attention network, and an action prediction network; the action sequence includes multiple action pose information; control a robotic arm to execute the target task based on the action sequence. The robotic arm control method based on a visual large model provided by the embodiment of the present invention processes the description text of the target task and the scene images of multiple perspectives thereof through an action prediction model including a visual segmentation large model, a multi-perspective attention network, and an action prediction network to obtain an action sequence, so as to control the robotic arm to execute the target task according to the action sequence, which can reduce the computational amount of three-dimensional space robotic arm control and improve the control efficiency. Description of the Drawings
[0018] Figure 1 is a flowchart of a robotic arm control method based on a visual large model in Embodiment 1 of the present invention;
[0019] Figure 2It is an example diagram for controlling a robotic arm based on a large vision model in Embodiment 1 of the present invention;
[0020] Figure 3 It is a schematic structural diagram of a multi - perspective attention network in Embodiment 1 of the present invention;
[0021] Figure 4 It is a schematic diagram of the principle of an action prediction network in Embodiment 1 of the present invention;
[0022] Figure 5 It is a schematic structural diagram of a robotic arm control device based on a large vision model in Embodiment 2 of the present invention;
[0023] Figure 6 It is a schematic structural diagram of an electronic device in Embodiment 3 of the present invention. Detailed implementation manners
[0024] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings, not all structures.
[0025] Embodiment 1
[0026] Figure 1 It is a flowchart of a method for controlling a robotic arm based on a large vision model provided in Embodiment 1 of the present invention. Among them, the large vision model can be a neural network model that can process visual information (such as images, videos, etc.) and has a large number of parameters (such as more than hundreds of millions). This embodiment is applicable to the situation of controlling a robotic arm. This method can be executed by a robotic arm control device based on a large vision model, and this device can be implemented in the form of software and / or hardware. Optionally, it is implemented through an electronic device, and this electronic device can be a mobile terminal, a PC, or a server, etc. Specifically, it includes the following steps:
[0027] S110, Obtain the description text of the target task and scene images of multiple perspectives.
[0028] Among them, the description text can be a description of the task to be executed by the robotic arm. For example: Move something from location A to location B. Scene images of multiple perspectives can be images obtained by photographing the scene where the target task is located from various angles.
[0029] S120, Input the description text of the target task and scene images of multiple perspectives into the action prediction model to obtain an action sequence.
[0030] Among them, the action prediction model can be a trained neural network model for predicting the action sequence of the robotic arm to perform the target task. The action prediction model includes: a large visual segmentation model, a multi-view attention network, and an action prediction network.
[0031] Among them, the large visual segmentation model can be obtained by retraining a pre-trained large image segmentation model (Segment Anything Model, SAM). Among them, the pre-trained large image segmentation model includes an image encoder for extracting features from images. In this embodiment, the method of retraining the pre-trained large image segmentation model can be to retrain the pre-trained large image segmentation model using a set fine-tuning algorithm. The set fine-tuning algorithm can be the Low-Rank Adaptation (LoRA) fine-tuning algorithm, and its principle can be to keep the parameters of the original large image segmentation model unchanged and adjust the parameters in a newly added network structure, where the newly added network structure is constructed based on the parameters in the pre-trained large image segmentation model and has a magnitude much smaller than that of the large image segmentation model. Exemplarily, assuming that the parameters in the large image segmentation model are represented as a matrix M*M, the parameters in the newly added network structure are represented as the dot product of matrices M*r and r*M, where r is much smaller than M. That is, the large visual segmentation model includes the pre-trained large image segmentation model and the newly added network structure, and the large visual segmentation model is obtained by adjusting the parameters in the newly added network structure based on the set fine-tuning algorithm. The multi-view attention network can be composed of a Transformer network, and the action prediction network can be composed of a convolutional neural network.
[0032] Among them, the action sequence includes multiple action pose information, and the action pose information includes position information and attitude information. The position information can be characterized by three-dimensional coordinates (x, y, z), and the attitude information can be represented by three attitude angles (pitch angle, yaw angle, roll angle). That is, the action pose information is a six-dimensional vector.
[0033] Specifically, the process of inputting the description text of the target task and the scene images of its multiple views into the action prediction model to obtain the action sequence can be: inputting the scene images of multiple views of the target task into the large visual segmentation model to output image features of multiple views; inputting the image features of multiple views and the description text into the multi-view attention network to output visual-text alignment features; inputting the visual-text alignment features into the action prediction network to obtain the action sequence.
[0034] In this embodiment, after inputting the scene images of multiple perspectives of the target task into the large visual segmentation model, the image encoders in the large visual segmentation model respectively extract features from the scene images of multiple perspectives and output image features of multiple perspectives. The multi-perspective attention network is used to perform cross-perspective and cross-modal processing on the multi-perspective image features and text features and output visual-text alignment features. The action prediction network is used to process the visual-text alignment features to obtain an action sequence. Exemplarily, Figure 2 is an example diagram of controlling a robotic arm based on a large visual model in this embodiment. As Figure 2 shown, the description text of the target task is "Use a stick to move the square onto the sky-blue color patch". The specific process is as follows: First, input the scene images of multiple perspectives into the large visual segmentation model to output image features of multiple perspectives; then input the description text and the image features of multiple perspectives into the multi-perspective attention network for cross-perspective and cross-modal processing to output visual-text alignment features. Then, input the visual-text alignment features into the action prediction network to obtain an action sequence including four action pose information. Finally, drive the end effector of the robotic arm to work based on the action sequence to execute the task of "Use a stick to move the square onto the sky-blue color patch".
[0035] Among them, the multi-perspective attention network includes a multi-modal pre-trained neural network, a single-perspective attention module, and a multi-perspective attention module. The multi-modal pre-trained neural network (Contrastive Language-Image Pre-Training, CLIP) is used to extract features from the text. The single-perspective attention module can be understood as a single-perspective Transformer module and is used to further process the image features of each perspective respectively. The multi-perspective attention module can be understood as a multi-perspective Transformer module and is used to fuse and align the text features and the image features of multiple perspectives to obtain visual-text alignment features. Exemplarily, Figure 3 is a schematic structural diagram of the multi-perspective attention network in this embodiment. As Figure 3 shown, the process of inputting the image features of multiple perspectives and the description text into the multi-perspective attention network and outputting visual-text alignment features can be: input the description text into the multi-modal pre-trained neural network to output text features; input the image features of multiple perspectives into the single-perspective attention module respectively to obtain intermediate features of multiple perspectives; input the text features and the intermediate features of multiple perspectives into the multi-perspective attention module to output visual-text alignment features.
[0036] Optionally, the way to input the visual text alignment feature into the action prediction network to obtain the action sequence can be: input the visual text alignment feature into the action prediction network and output the action sequence; or, input the visual text alignment feature into the action prediction network and output action sequence diagrams from multiple perspectives; fuse the action sequence diagrams from multiple perspectives to obtain the action sequence.
[0037] Among them, the action sequence diagram contains the pose sequence in the image coordinate system at this angle. In this embodiment, the action prediction network can directly output the action sequence, or first output action sequence diagrams from multiple perspectives, and then fuse the action sequence diagrams from multiple perspectives to obtain the action sequence. In this embodiment, the way to fuse the action sequence diagrams from multiple perspectives can be: perform coordinate system conversion (from the image coordinate system to the space coordinate system) on the pose sequences in each action sequence diagram, and then fuse the converted pose sequences to obtain the action sequence in the three-dimensional space. Exemplarily, Figure 4 is the schematic diagram of the action prediction network in this embodiment. As Figure 4 shown, after inputting the visual text alignment feature into the action prediction network, action sequence diagrams from multiple perspectives are output.
[0038] S130, control the robotic arm based on the action sequence to perform the target task.
[0039] In this embodiment, after obtaining the action sequence, control the robotic arm to move according to the action pose information in the action sequence to perform the target task.
[0040] Optionally, the training method of the action prediction model can be: obtain the description text of the task sample, the scene images from multiple perspectives, and the real action sequence; input the description text of the task sample and its scene images from multiple perspectives into the action prediction model to obtain the predicted action sequence; train the action prediction model based on the predicted action sequence and the real action sequence.
[0041] Among them, the number of action pose information included in the predicted action sequence is the same as the number of real action pose information included in the real action sequence, and this number can be pre-configured when training the action prediction model. To increase the diversity and quantity of samples, the way to obtain the scene images from multiple perspectives of the task sample and the real action sequence can be: obtain the task trajectory of the task sample, the task trajectory includes multiple trajectory points, and each trajectory point includes scene images and action pose information at multiple angles; split the task trajectory to obtain multiple sets of scene images from multiple perspectives of the task sample and the real action sequence. The description texts corresponding to the multiple sets of scene images from multiple perspectives of the task sample and the real action sequence are the same. Exemplarily, assume that the task trajectory of a certain task sample is: {(s 0 ,a 0),(s 1 ,a 1 ),…(s t ,a t )}, where s t is the scene image of multiple perspectives of the t-th trajectory point, and a t is the 6-dimensional action pose information. The result of splitting the task trajectory can be: {s 0 ; a 0 , a 2 , … a t}, {s 1 ; a 1 , a 2 , …, a t , a t}, {s 2 ; a 2 , a 3 , …, a t , a t , a t}, …… {s t ; a t , … a t , a t}.
[0042] In this embodiment, the processing process of the action prediction model for the description text of the task sample and the scene images of its multiple perspectives is similar to the processing process of the action prediction model for the description text of the target task and the scene images of its multiple perspectives in the above embodiment, and will not be elaborated here.
[0043] Specifically, the way to train the action prediction model based on the predicted action sequence and the true action sequence can be: determining the loss function according to the predicted action sequence and the true action sequence; performing backpropagation tuning on the action prediction network and the multi-perspective attention network based on the loss function, and performing backpropagation tuning on the newly added network structure in the visual segmentation large model.
[0044] Among them, the way to determine the loss function according to the predicted action sequence and the true action sequence can refer to the existing ways to determine various loss functions, which are not limited here. For example: mean square error loss, cross-entropy loss, etc. The way to perform backpropagation tuning on the newly added network structure in the visual segmentation large model can be: using the LoRA fine-tuning algorithm to perform backpropagation tuning on the newly added network structure in the visual segmentation large model. The LoRA fine-tuning algorithm can refer to the existing related technologies, which are not limited here.
[0045] The technical solution of this embodiment is to obtain the description text of the target task and the scene images from multiple perspectives; input the description text of the target task and the scene images from multiple perspectives into an action prediction model to obtain an action sequence; where the action prediction model includes: a large visual segmentation model, a multi-perspective attention network, and an action prediction network; the action sequence includes multiple action pose information; control the robotic arm based on the action sequence to execute the target task. The robotic arm control method based on the large visual model provided by the embodiments of the present invention processes the description text of the target task and the scene images from multiple perspectives through an action prediction model including a large visual segmentation model, a multi-perspective attention network, and an action prediction network to obtain an action sequence, so as to control the robotic arm to execute the target task according to the action sequence, which can reduce the computational complexity of the three-dimensional space robotic arm control and improve the control efficiency.
[0046] Embodiment 2
[0047] Figure 5 FIG. is a schematic structural diagram of a robotic arm control device based on a large visual model provided by Embodiment 2 of the present invention, as Figure 5 shown, the device includes:
[0048] A scene image acquisition module 510, configured to acquire the description text of the target task and the scene images from multiple perspectives;
[0049] An action sequence acquisition module 520, configured to input the description text of the target task and the scene images from multiple perspectives into an action prediction model to obtain an action sequence; where the action prediction model includes: a large visual segmentation model, a multi-perspective attention network, and an action prediction network; the action sequence includes multiple action pose information;
[0050] A robotic arm control module 530, configured to control the robotic arm based on the action sequence to execute the target task.
[0051] Optionally, the action sequence acquisition module 520 is further configured to:
[0052] Input the scene images from multiple perspectives of the target task into the large visual segmentation model, and output image features from multiple perspectives;
[0053] Input the image features from multiple perspectives and the description text into the multi-perspective attention network, and output visual-text alignment features;
[0054] Input the visual-text alignment features into the action prediction network to obtain an action sequence.
[0055] Optionally, the multi-perspective attention network includes a multi-modal pre-trained neural network, a single-perspective attention module, and a multi-perspective attention module; the action sequence acquisition module 520 is further configured to:
[0056] Input the description text into the multi-modal pre-trained neural network to output text features;
[0057] Input the image features from multiple perspectives into the single-perspective attention module respectively to obtain intermediate features from multiple perspectives;
[0058] Input the text features and the intermediate features from multiple perspectives into the multi-perspective attention module to output visual-text alignment features.
[0059] Optionally, the action sequence acquisition module 520 is further configured to:
[0060] Input the visual-text alignment features into the action prediction network to output an action sequence; or,
[0061] Input the visual-text alignment features into the action prediction network to output action sequence diagrams from multiple perspectives;
[0062] Fuse the action sequence diagrams from multiple perspectives to obtain an action sequence.
[0063] Optionally, the visual segmentation large model includes a pre-trained image segmentation large model and a newly added network structure, where the newly added network structure is constructed based on the parameters in the pre-trained image segmentation large model; the visual segmentation large model is obtained by adjusting the parameters in the newly added network structure based on a set fine-tuning algorithm.
[0064] Optionally, it further includes: an action prediction model training module, which is used for:
[0065] Obtain the description text of the task sample, the scene images from multiple perspectives, and the true action sequence;
[0066] Input the description text of the task sample and its scene images from multiple perspectives into the action prediction model to obtain a predicted action sequence;
[0067] Train the action prediction model based on the predicted action sequence and the true action sequence.
[0068] Optionally, the action prediction model training module is further configured to:
[0069] Determine the loss function according to the predicted action sequence and the true action sequence;
[0070] Based on the loss function, perform backpropagation to adjust the parameters of the action prediction network and the multi-perspective attention network, and perform backpropagation to adjust the parameters of the newly added network structure in the visual segmentation large model.
[0071] The above device can execute the methods provided in all the foregoing embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the above methods. For the technical details not described in detail in this embodiment, reference can be made to the methods provided in all the foregoing embodiments of the present invention.
[0072] Embodiment 3
[0073] Figure 6 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described herein and / or claimed.
[0074] As Figure 6 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0075] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0076] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the robotic arm control method based on a vision large model.
[0077] In some embodiments, the robotic arm control method based on a vision large model can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the robotic arm control method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the robotic arm control method by any other suitable means (e.g., by means of firmware).
[0078] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0079] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0080] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0081] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0082] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0083] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs that run on respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0084] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0085] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A robotic arm control method based on a visual large model, characterized in that: include: Obtain the description text of the target task and scene images from multiple perspectives; Inputting the description text of the target task and scene images of multiple perspectives into the action prediction model to obtain an action sequence; wherein the action prediction model includes: a large visual segmentation model, a multi-perspective attention network and an action prediction network; the action sequence includes multiple action posture information; Inputting the description text of the target task and the scene images of multiple perspectives thereof into the action prediction model to obtain the action sequence, including: inputting the scene images of multiple perspectives of the target task into the visual segmentation large model, and outputting the image features of multiple perspectives; Inputting the image features of the multiple perspectives and the description text into the multi-perspective attention network, and outputting visual text alignment features; Inputting the visual text alignment feature into the action prediction network to obtain an action sequence; The multi-view attention network includes a multimodal pre-trained neural network, a single-view attention module and a multi-view attention module; the image features of the multiple views and the description text are input into the multi-view attention network, and visual text alignment features are output, including: Inputting the description text into the multimodal pre-trained neural network and outputting text features; Inputting the image features of the multiple perspectives into the single-perspective attention module respectively to obtain intermediate features of the multiple perspectives; Inputting the text features and the intermediate features of the multiple perspectives into the multi-perspective attention module, and outputting visual text alignment features; The robot arm is controlled based on the action sequence to perform the target task.
2. The method according to claim 1, characterized in that Inputting the visual text alignment feature into the action prediction network to obtain an action sequence includes: Inputting the visual-text alignment feature into the action prediction network and outputting the action sequence; or, Inputting the visual text alignment feature into the action prediction network, and outputting action sequence graphs from multiple perspectives; The action sequence graphs of the multiple perspectives are fused to obtain an action sequence.
3. The method according to claim 1, characterized in that The visual segmentation model includes a pre-trained image segmentation model and a newly added network structure, wherein the newly added network structure is constructed based on the parameters in the pre-trained image segmentation model; the visual segmentation model is obtained by adjusting the parameters in the newly added network structure based on a set fine-tuning algorithm.
4. The method according to claim 3, characterized in that The training method of the action prediction model is: Obtain description text of task samples, scene images from multiple perspectives, and real action sequences; Inputting the description text of the task sample and scene images of multiple perspectives into the action prediction model to obtain a predicted action sequence; The action prediction model is trained based on the predicted action sequence and the actual action sequence.
5. The method according to claim 4, characterized in that Training the action prediction model based on the predicted action sequence and the real action sequence includes: Determine a loss function according to the predicted action sequence and the actual action sequence; Based on the loss function, the action prediction network and the multi-view attention network are reversely adjusted, and the newly added network structure in the visual segmentation model is reversely adjusted.
6. A robotic arm control device based on a visual large model, characterized in that: include: The scene image acquisition module is used to obtain the description text of the target task and scene images from multiple perspectives; An action sequence acquisition module is used to input the description text of the target task and the scene images of multiple perspectives into the action prediction model to obtain an action sequence; wherein the action prediction model includes: a large visual segmentation model, a multi-perspective attention network and an action prediction network; the action sequence includes multiple action posture information; The action sequence acquisition module is further used for: Input scene images of multiple perspectives of the target task into the large visual segmentation model and output image features of multiple perspectives; Input the image features and description texts from multiple perspectives into the multi-view attention network and output the visual text alignment features; Input the visual text alignment features into the action prediction network to obtain the action sequence; The multi-view attention network includes a multi-modal pre-trained neural network, a single-view attention module and a multi-view attention module; the action sequence acquisition module is also used to: Input the description text into the multimodal pre-trained neural network and output the text features; The image features of multiple views are input into the single-view attention module respectively to obtain the intermediate features of multiple views; Input text features and intermediate features of multiple views into the multi-view attention module and output visual-text alignment features; A robot arm control module is used to control the robot arm based on the action sequence to perform the target task.
7. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the robot arm control method based on a visual large model according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the robot arm control method based on a visual large model as described in any one of claims 1 to 5 when executed.
Citation Information
Patent Citations
Robot movement control method, device, system and equipment based on binocular vision
CN113601510A
Determining and utilizing corrections to robot actions
US20190001489A1