Motion reasoning method and device for immunization injection
By using the multimodal large model LLaVA to learn the mapping relationship between the posture information of the syringe and the target to be injected and the action sequence, the robotic arm is controlled to autonomously complete the vaccine injection, solving the problem that the robotic syringe cannot adapt to the environment and individual differences, and achieving efficient and safe vaccination.
Patent Information
- Application Number
- CN202411191812.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-08-28
AI Technical Summary
Existing vaccine injection robots lack the ability to adjust themselves autonomously and are unable to adapt to different environments and individual differences in livestock, which affects injection efficiency and accuracy.
The multimodal large model LLaVA is used to conduct multiple rounds of dialogue on the teaching injection video through visual question-answering technology to accumulate content understanding, learn the mapping relationship between the syringe posture information and the posture information of the target to be injected and the action sequence, control the robotic arm to execute the action sequence, and optimize the action sequence in combination with real-time posture.
It improves the efficiency and accuracy of vaccine injection, reduces the risk of manual operation, ensures that every livestock is vaccinated in time, and improves the production efficiency and safety of commercial farming.
Smart Images

Figure CN119319560B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, in particular to the field of intelligent animal husbandry technology, and more particularly to a motion reasoning method and device for immunization injection. Background Art
[0002] In modern animal husbandry, vaccinating livestock is a critical operation that is directly related to the health of livestock and the profitability of farming. Traditionally, this task is performed manually by veterinarians or experienced staff, but this method is labor-intensive and has low injection efficiency in large-scale farms. It is difficult to ensure that every livestock is vaccinated in a timely and accurate manner, which affects production efficiency and increases the safety risks of operators. With the advancement of robotics technology, it has become feasible and necessary to try to use automated robots to complete vaccination tasks. However, most of the current vaccination robots do not have the ability to adjust themselves. The operating parameters and movement modes of the injection robots are fixed, and they cannot adapt to different environments and individual differences of livestock, which affects the efficiency and accuracy of the injection. Summary of the Invention
[0003] The present disclosure aims to solve one of the technical problems in the related art at least to a certain extent.
[0004] To this end, the first embodiment of the present disclosure proposes a motion reasoning method for immune injection, which is applied to an injection robot that controls the movement of the syringe through a robotic arm. The method includes the following steps:
[0005] Acquire a current injection image, where the current injection image includes a syringe and a target to be injected;
[0006] Determining syringe posture information and target posture information to be injected according to the current injection image;
[0007] Inputting the syringe posture information and the posture information of the target to be injected into a pre-trained multimodal large model LLaVA to obtain an action sequence corresponding to the syringe posture information and the posture information of the target to be injected, wherein the action sequence includes multiple actions; the multimodal large model LLaVA uses visual question-answering technology to conduct multiple rounds of dialogue and accumulate content understanding on the teaching injection video, and learns the mapping relationship between the syringe posture information, the posture information of the target to be injected, and the action sequence;
[0008] Control the robotic arm to execute actions in the action sequence.
[0009] In some embodiments of the present disclosure, controlling the robotic arm to perform the actions in the action sequence includes: S1, controlling the robotic arm to perform the first N actions in the action sequence to obtain an injection image after performing the action; wherein, N is a positive integer; S2, using the posture information of the syringe in the injection image after performing the action as new syringe posture information, using the posture information of the target to be injected in the injection image after performing the action as new target posture information to be injected, inputting the new syringe posture information and the new target posture information to be injected into the multimodal large model LLaVA to obtain a new action sequence; S3, controlling the robotic arm to perform the first N actions in the new action sequence; S4, repeating steps S2-S3 until the robotic arm performs the preset injection end action.
[0010] In some embodiments of the present disclosure, the multimodal large model LLaVA uses visual question-answering technology to conduct multiple rounds of dialogue on the teaching injection video to accumulate content understanding, and learn to obtain the mapping relationship between the syringe posture information, the target posture information to be injected, and the action sequence, including: segmenting and sampling the teaching injection video to obtain continuous video frame images; determining the initial syringe posture information of the syringe and the initial target posture information to be injected in the teaching injection video; the multimodal large model LLaVA uses visual question-answering technology to select the initial action in the continuous video frame images from multiple injection actions in the action skill library through multiple rounds of dialogue tasks, and generates an initial action sequence corresponding to the initial syringe posture information and the initial target posture information to be injected based on the initial action; obtains the action sequence label in the teaching injection video, and the action sequence label is determined based on the robot vaccine injection operation skill knowledge; the multimodal large model LLaVA is trained based on the action sequence label and the initial action sequence, so that the multimodal large model LLaVA learns to obtain the mapping relationship between the syringe posture information, the target posture information to be injected, and the action sequence.
[0011] In some embodiments of the present disclosure, the syringe posture information includes the position information and spatial posture information of the syringe, and the spatial posture information of the syringe is obtained by: establishing a 3D model of the syringe; using a deep learning-based visual algorithm to extract the feature information of the syringe in the current injection image; matching the feature information with the 3D model to capture the spatial posture information of the syringe.
[0012] In some embodiments of the present disclosure, the action includes an action type and an action parameter.
[0013] A second aspect of the present disclosure provides a motion reasoning device for immunization injection, wherein the device is used to control the motion of a syringe through a robotic arm of an injection robot, and the device includes:
[0014] An acquisition module, configured to acquire a current injection image, wherein the current injection image includes a syringe and a target to be injected;
[0015] A first determining module is configured to determine syringe posture information and a target posture information to be injected according to the current injection image;
[0016] The second determination module is configured to input the syringe posture information and the target posture information into a pre-trained multimodal large model LLaVA to obtain an action sequence corresponding to the syringe posture information and the target posture information, wherein the action sequence includes a plurality of actions; the multimodal large model LLaVA uses visual question-answering technology to conduct multiple rounds of dialogue and accumulate content understanding of the teaching injection video, and learn the mapping relationship between the syringe posture information, the target posture information, and the action sequence;
[0017] A control module is used to control the robotic arm to perform actions in the action sequence.
[0018] In some embodiments of the present disclosure, the control module is specifically used to: S1, control the robotic arm to perform the first N actions in the action sequence to obtain an injection image after the action is performed; wherein, N is a positive integer; S2, use the posture information of the syringe in the injection image after the action is performed as new syringe posture information, use the posture information of the target to be injected in the injection image after the action is performed as new target posture information to be injected, input the new syringe posture information and the new target posture information to be injected into the multimodal large model LLaVA to obtain a new action sequence; S3, control the robotic arm to perform the first N actions in the new action sequence; S4, repeat steps S2-S3 until the robotic arm performs the preset injection end action.
[0019] In some embodiments of the present disclosure, the device also includes a training device; the training device is used to: segment and sample the teaching injection video to obtain continuous video frame images; determine the initial syringe posture information of the syringe and the initial target posture information of the target to be injected in the teaching injection video; the multimodal large model LLaVA adopts visual question-answering technology to select the initial action in the continuous video frame image from multiple injection actions in the action skill library through multiple rounds of dialogue tasks, and generates an initial action sequence corresponding to the initial syringe posture information and the initial target posture information based on the initial action; obtain the action sequence label in the teaching injection video, and the action sequence label is determined based on the robot vaccine injection operation skill knowledge; train the multimodal large model LLaVA based on the action sequence label and the initial action sequence, so that the multimodal large model LLaVA learns to obtain the mapping relationship between the syringe posture information and the target posture information to be injected and the action sequence.
[0020] A third embodiment of the present disclosure provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;
[0021] The memory stores computer-executable instructions;
[0022] The processor executes the computer-executable instructions stored in the memory to implement the method described in the first aspect.
[0023] The fourth aspect of the present disclosure provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer-executable instructions, which are used to implement the method described in the first aspect when executed by a processor.
[0024] The proposed motion inference method for immunization uses a pre-trained large multimodal LLaVA model to determine the motion sequence of the syringe pose and the target pose in the current scenario. This method then guides the robotic arm of the injection robot to autonomously complete the vaccination, adapting to different environments and individual differences in the target. This method significantly improves vaccination efficiency, reduces manual operation risks, ensures timely vaccination of every livestock, and enhances the overall production efficiency and safety of commercial farming.
[0025] Additional aspects and advantages of the present disclosure will be given in part in the description below and in part will be obvious from the description below, or will be learned through practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The above and / or additional aspects and advantages of the present disclosure will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0027] Figure 1 A flowchart of an action reasoning method for immune injection provided by an embodiment of the present disclosure;
[0028] Figure 2 A schematic diagram of an injection image provided by an embodiment of the present disclosure;
[0029] Figure 3 A schematic diagram of obtaining the spatial posture information of a syringe through a 3D model of the syringe provided in an embodiment of the present disclosure;
[0030] Figure 4 An implementation process of controlling a robotic arm to perform actions in an action sequence provided by an embodiment of the present disclosure;
[0031] Figure 5 A flowchart of a multimodal large model LLaVA training method provided in an embodiment of the present disclosure;
[0032] Figure 6 A multimodal large model LLaVA action parsing framework diagram provided by an embodiment of the present disclosure;
[0033] Figure 7 A schematic diagram of an action sequence template provided by an embodiment of the present disclosure;
[0034] Figure 8 A schematic diagram of an action sequence output by a multimodal large model LLaVA after topology optimization provided by an embodiment of the present disclosure;
[0035] Figure 9 A schematic diagram of topological optimization of actions based on an action sequence template provided in an embodiment of the present disclosure;
[0036] Figure 10 An embodiment of the present disclosure provides an action reasoning device for immunization injection. DETAILED DESCRIPTION
[0037] The following describes in detail embodiments of the present disclosure, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.
[0038] The present disclosure proposes a motion reasoning method and device for immunization injection. Specifically, the motion reasoning method and device for immunization injection according to an embodiment of the present disclosure are described below with reference to the accompanying drawings.
[0039] Figure 1This is a flow chart of a motion reasoning method for immune injection provided by an embodiment of the present disclosure. The method is applied to an injection robot that controls the movement of the syringe through a robotic arm. Figure 1 As shown, the action reasoning method for immune injection may include the following steps:
[0040] Step 101 : Acquire a current injection image, which includes a syringe and a target to be injected.
[0041] The injection image can be collected by the RGB-D camera on the injection robot. Taking pigs as an example, Figure 2 Schematic diagram of the injection image provided by the embodiment of the present disclosure. Figure 2 As shown, the injection image includes a syringe and a target to be injected (a live pig).
[0042] Step 102 : determining the syringe posture information and the object posture information to be injected according to the current injection image.
[0043] The posture information includes position information and spatial posture information. Optionally, the posture information of the syringe and the posture information of the target to be injected can be obtained by plane target detection technology.
[0044] In related technologies, the spatial posture information of the syringe is usually obtained by installing a posture detection sensor on the syringe. However, the learning cost of this method is relatively high. At the same time, the presence of the sensor will hinder the expert's teaching actions to a certain extent, affecting the accuracy of the operation and thus affecting the imitation learning effect. In order to improve the efficiency of posture detection, reduce the learning cost, and not affect the smoothness of the expert's teaching operation, in some embodiments of the present disclosure, the spatial posture information of the syringe can also be obtained by matching the 3D model of the syringe. Specifically, in one implementation method, such as Figure 3 As shown in the figure, the syringe can be modeled by using the NeRF-based visual 3D reconstruction algorithm of Solidworks to generate a digital format file and establish a 3D model of the syringe ( Figure 3 Part a). Using a deep learning-based visual algorithm, the current injection image ( Figure 3 Part b) performs feature extraction and analysis, and uses the learning ability of the neural network to identify the characteristic information of the syringe in the current injection image, such as the shape and contour of the syringe. Then, the characteristic information is matched with the 3D model of the syringe, thereby accurately obtaining the spatial posture information of the syringe in three-dimensional space ( Figure 3 (Part c).
[0045] Step 103 : Input the syringe posture information and the posture information of the target to be injected into the pre-trained multimodal large model LLaVA to obtain an action sequence corresponding to the syringe posture information and the posture information of the target to be injected, where the action sequence includes multiple actions.
[0046] The action includes an action type and action parameters. As an example, the action type may include positioning, tilting, releasing, pressing, etc., and the action parameters may include the position of the action in space, the thrust of the liquid pressing action during injection, etc.
[0047] It should be noted that the multimodal large-scale LLaVA model has already used visual question-answering technology to conduct multiple rounds of dialogue on the instructional injection video to accumulate content, imitate and learn the injection movements in the instructional injection video, and enable the model to learn the mapping relationship between the syringe posture information, the posture information of the target to be injected, and the action sequence. This makes the action sequence output by the model more adapted to the syringe posture information and the posture information of the target to be injected, thereby improving injection accuracy. For the training method of the multimodal large-scale LLaVA model in this embodiment, please refer to the description in the subsequent embodiments of this disclosure and will not be repeated here.
[0048] Step 104: Control the robotic arm to execute actions in the action sequence.
[0049] By implementing the disclosed embodiments, a pre-trained multimodal large-scale LLaVA model, combined with imitation learning, is used to determine the action sequence of the syringe pose and the target pose in the current scenario. This allows the robot's robotic arm to autonomously complete the vaccine injection, adapting to different environments and individual differences in the target. This significantly improves the efficiency of vaccination, reduces the risk of manual operation, ensures timely vaccination of every livestock, and enhances the overall production efficiency and safety of commercial farming.
[0050] In some embodiments of the present disclosure, in order to further improve the accuracy of the injection process, the action sequence can be dynamically updated and optimized based on the real-time posture during the injection process. In one implementation, Figure 4 The embodiment of the present disclosure provides a method for controlling a robotic arm to perform actions in an action sequence. Figure 4 As shown, the process of controlling the robot arm to perform actions in the action sequence may include the following steps:
[0051] Step 401: Control the robotic arm to execute the first N actions in the action sequence and obtain an injection image after the actions are executed, where N is a positive integer.
[0052] In step 402, the posture information of the syringe in the injection image after the action is performed is used as the new syringe posture information, and the posture information of the target to be injected in the injection image after the action is performed is used as the new target posture information. The new syringe posture information and the new target posture information are input into the multimodal large model LLaVA to obtain a new action sequence.
[0053] As an example, taking N as 1, after controlling the robotic arm to perform the first action in the action sequence, the target to be injected may move slightly, and the position of the target to be injected may change. In this case, continuing the unexecuted action in the action sequence to inject the target to be injected may result in poor injection. To improve the accuracy of the injection action, the position information of the syringe and the target to be injected after the first action is executed is used as the new syringe position information and the new target position information, respectively. The new syringe position information and the new target position information are re-input into the multimodal large model LLaVA to obtain a new action sequence.
[0054] Step 403: Control the robotic arm to execute the first N actions in the new action sequence.
[0055] Step 404: Repeat steps 402 to 403 until the robotic arm performs the preset injection end action.
[0056] Based on the new syringe posture information and the new target posture information, the first N actions in the new action sequence are continued to be executed, and steps 402-403 are repeated. After executing the first N actions in the new action sequence, the posture information of the syringe and the target to be injected after the actions are executed is re-determined as the new syringe posture information and the new target posture information to be injected. The action sequence is updated in real time until the robot arm executes the preset injection end action. Among them, the injection end action can be set as the syringe withdrawal action after the syringe has completed the liquid injection. When the robot arm executes the withdrawal action, it is considered that the injection of the target to be injected is completed, and the cycle ends.
[0057] Therefore, by implementing steps 401 to 404, the action sequence can be dynamically updated and optimized based on the real-time posture of the syringe and the target to be injected during the injection process, so that the actions performed by the robotic arm are more precise, further improving the injection accuracy of the target to be injected.
[0058] Figure 5 Schematic diagram of a training method for a multimodal large model LLaVA provided in an embodiment of the present disclosure. Figure 5 As shown, the training method of the multimodal large model LLaVA may include the following steps:
[0059] Step 501 : Segment and sample the teaching injection video to obtain continuous video frame images.
[0060] Figure 6 This is a multimodal large model LLaVA action parsing framework diagram provided by the embodiment of the present disclosure. Figure 6 As shown, the instructional injection video is segmented and sampled to obtain continuous video frame images. In one implementation, the instructional injection video can be segmented and sampled at preset intervals. Pairing adjacent frames creates frame combinations that preserve the temporal information of the action, laying the foundation for action recognition and analysis.
[0061] Step 502 : Determine the initial syringe posture information of the syringe and the initial target posture information of the target object to be injected in the teaching injection video.
[0062] Alternatively, the initial syringe pose information of the syringe in the teaching injection video and the initial target pose information of the target to be injected can be obtained by plane target detection technology. Alternatively, the initial syringe pose information of the syringe can be obtained by matching the 3D model of the syringe. For specific acquisition methods, please refer to Figure 1 The relevant description of the embodiment will not be repeated here.
[0063] In step 503, the multimodal large model LLaVA uses visual question-answering technology to select initial actions in continuous video frame images from multiple injection actions in the action skill library through multi-round dialogue tasks, and generates an initial action sequence corresponding to the initial syringe posture information and the initial target posture information to be injected based on the initial action.
[0064] like Figure 6 As shown, the multimodal large model LLaVA combines vision and language processing capabilities, uses visual question answering technology, responds to natural language questions, and deeply analyzes visual content. ChatGPT plays the role of questioner, guiding the model to focus on specific parts of the video frame, such as object contact or object movement. Each difference between video frame images represents the changes between consecutive frames, and these changes are the key basis for action detection and recognition. Through multiple rounds of dialogue, ChatGPT accumulates its understanding of the video content, selects corresponding actions from the predefined action skill library as the initial actions in the consecutive video frame images, and then constructs a coherent initial action sequence corresponding to the initial syringe posture information and the initial target posture information to be injected. In addition, ChatGPT can also provide a language description containing context for the selected initial action to learn the relationship between actions and guide the model to complete complex tasks.
[0065] Step 504 : Obtain action sequence labels in the teaching injection video. The action sequence labels are determined based on the robot vaccine injection operation skill knowledge.
[0066] It should be noted that the action sequence labels in the teaching injection video can be manually defined, and the robot vaccine injection operation skill knowledge includes the action sequence template in the robot's injection action. This action sequence template is different from the manual injection action. It is an action sequence suitable for injection robots to perform by imitating and learning human injection actions. Figure 7 This is a schematic diagram of an action sequence template provided by an embodiment of the present disclosure. Figure 7 As shown, the teaching staff can refer to the action sequence template to define the action label of each action in the teaching injection process, and generate the action sequence label in the teaching injection video according to the action labels of multiple actions.
[0067] Step 505 : training the multimodal large model LLaVA based on the action sequence labels and the initial action sequence, so that the multimodal large model LLaVA learns the mapping relationship between the syringe posture information, the posture information of the target to be injected, and the action sequence.
[0068] The multimodal large model LLaVA is trained based on action sequence labels and initial action sequences. The initial action sequence output by the multimodal large model LLaVA is trained and corrected using the action sequence labels, so that the initial action sequence output by the multimodal large model LLaVA is more consistent with the action sequence template.
[0069] It should be noted that the action sequence template is not fully applicable to all situations. In the case of different syringe postures and the postures of the target to be injected, the action sequence template may contain redundant or useless actions. Therefore, the multimodal large model LLaVA uses visual question-answering technology to conduct multiple rounds of dialogue on the teaching injection video to accumulate content understanding, so that the multimodal large model LLaVA learns the association or topological relationship between each action in the action sequence, and obtains the ability to reason about the actions before and after (for example, through training, the large model learns through knowledge reasoning that there is a before-and-after association between the syringe liquid pressing completion action and the syringe withdrawal action). Therefore, the multimodal large model LLaVA can further topologically optimize the action based on the output that is more in line with the action sequence template according to the syringe posture information and the posture information of the target to be injected, and output the action sequence under the syringe posture information and the posture information of the target to be injected based on the action reasoning ability. Figure 8 A schematic diagram of an action sequence output by a multimodal large model LLaVA after topology optimization provided by an embodiment of the present disclosure.
[0070] The multimodal large-scale LLaVA model, based on its ability to learn and reason about action sequence templates, learns the mapping relationship between syringe and target pose information and action sequences. This allows the model's output action sequences to better adapt to the syringe and target pose information, while also conforming to the action sequence templates. The action sequences output by the trained multimodal large-scale LLaVA model guide the injection robot to complete autonomous vaccine injections, significantly improving the accuracy and adaptability of vaccine injections, and significantly enhancing injection efficiency and safety.
[0071] Optionally, in some embodiments of the present disclosure, Figure 9 This is a schematic diagram of topological optimization of actions based on action sequence templates provided by the embodiment of the present disclosure. Figure 9 As shown in Figure 1, the key frame information, spatial posture information, force information and position information of the scene can be obtained through the vaccine injection scene information. Among them, the key frame information is used to represent the action stage of the syringe (such as Figure 8 The force information is the thrust of the syringe during injection. The multimodal large-scale LLaVA model uses the keyframe, spatial, force, and pose information of the current injection scenario to perform motion reasoning and topology optimization based on the action sequence template, ultimately outputting an action sequence tailored to the current scenario.
[0072] By implementing the disclosed embodiments, the multimodal large-scale model LLaVA uses visual question-answering technology to imitate and train the injection actions in the instructional injection video. The learned mapping relationship between the syringe posture information, the posture information of the target to be injected, and the action sequence is obtained. This allows the trained multimodal large-scale model LLaVA to infer the injection action sequence under different injection scenarios in actual application scenarios, autonomously adapt to different injection environments and individual differences in injections, improve the accuracy and adaptability of vaccine injections, significantly improve the efficiency of vaccine injections, make vaccination in large-scale commercial farming faster and more efficient, reduce the risk of exposure to operators, and improve workplace safety.
[0073] Figure 10 The embodiment of the present disclosure provides a motion inference device for immune injection, which is used to control the motion of the syringe through the robotic arm of the injection robot. Figure 10 As shown, the motion reasoning device for immunization injection includes: an acquisition module 1001 , a first determination module 1002 , a second determination module 1003 and a control module 1004 .
[0074] The acquisition module 1001 is used to acquire the current injection image, which includes the syringe and the target to be injected.
[0075] The first determining module 1002 is configured to determine the syringe posture information and the object posture information to be injected according to the current injection image.
[0076] The second determination module 1003 is configured to input the syringe pose information and the target pose information into a pre-trained multimodal large-scale model (LLaVA) to obtain an action sequence corresponding to the syringe pose information and the target pose information. The action sequence includes multiple actions. The multimodal large-scale model (LLaVA) uses visual question-answering technology to conduct multiple rounds of dialogue on the instructional injection video to understand the content and learn the mapping relationship between the syringe pose information, the target pose information, and the action sequence.
[0077] The control module 1004 is used to control the robotic arm to perform actions in the action sequence.
[0078] In some embodiments of the present disclosure, an action may include an action type and action parameters.
[0079] In some embodiments of the present disclosure, the control module 1004 is specifically configured to:
[0080] S1, controlling the robotic arm to execute the first N actions in the action sequence and obtaining the injection image after the actions are executed; where N is a positive integer;
[0081] S2, using the syringe pose information in the injection image after the action is performed as new syringe pose information, using the pose information of the target to be injected in the injection image after the action is performed as new target pose information, and inputting the new syringe pose information and the new target pose information into the multimodal large model LLaVA to obtain a new action sequence;
[0082] S3, control the robotic arm to execute the first N actions in the new action sequence;
[0083] S4, repeat steps S2-S3 until the robotic arm performs the preset injection end action.
[0084] In some embodiments of the present disclosure, Figure 10Based on the illustrated embodiment, the motion reasoning device for immunization injection may further include a training device. The training device is used to: segment and sample the teaching injection video to obtain continuous video frame images; determine the initial syringe pose information of the syringe and the initial target pose information of the target to be injected in the teaching injection video; the multimodal large model LLaVA uses visual question-answering technology to select the initial action in the continuous video frame images from multiple injection actions in the action skill library through multiple rounds of dialogue tasks, and generates an initial action sequence corresponding to the initial syringe pose information and the initial target pose information based on the initial action; obtain the action sequence label in the teaching injection video, the action sequence label is determined based on the robot vaccine injection operation skill knowledge; train the multimodal large model LLaVA based on the action sequence label and the initial action sequence, so that the multimodal large model LLaVA learns the mapping relationship between the syringe pose information, the target pose information and the action sequence.
[0085] In some embodiments of the present disclosure, the syringe posture information includes the position information and spatial posture information of the syringe. The spatial posture information of the syringe is obtained by: establishing a 3D model of the syringe; using a deep learning-based visual algorithm to extract the feature information of the syringe in the current injection image; matching the feature information with the 3D model to capture the spatial posture information of the syringe.
[0086] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0087] In order to implement the above embodiments, the present disclosure also proposes an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided by the above embodiments.
[0088] In order to implement the above embodiments, the present disclosure further proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the methods provided by the above embodiments.
[0089] In order to implement the above embodiments, the present disclosure further provides a computer program product, including a computer program, which implements the methods provided in the above embodiments when executed by a processor.
[0090] In the descriptions of the aforementioned embodiments, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are mutually inconsistent.
[0091] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0092] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.
[0093] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0094] It should be understood that various parts of the present disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0095] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0096] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.
[0097] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. A person of ordinary skill in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A motion reasoning method for immune injection, the method is applied to an injection robot, the injection robot controls the syringe movement through a robotic arm, characterized in that: The following steps are involved: Acquire a current injection image, where the current injection image includes a syringe and a target to be injected; Determining syringe posture information and target posture information to be injected according to the current injection image; Inputting the syringe posture information and the posture information of the target to be injected into a pre-trained multimodal large model LLaVA to obtain an action sequence corresponding to the syringe posture information and the posture information of the target to be injected, wherein the action sequence includes multiple actions; the multimodal large model LLaVA uses visual question-answering technology to conduct multiple rounds of dialogue and accumulate content understanding on the teaching injection video, and learns the mapping relationship between the syringe posture information, the posture information of the target to be injected, and the action sequence; Control the robotic arm to execute actions in the action sequence.
2. The method according to claim 1, characterized in that The controlling the robotic arm to perform the actions in the action sequence includes: S1, controlling the robotic arm to execute the first N actions in the action sequence to obtain an injection image after the actions are executed; wherein N is a positive integer; S2, using the posture information of the syringe in the injection image after the action is performed as new syringe posture information, using the posture information of the target to be injected in the injection image after the action is performed as new target posture information, and inputting the new syringe posture information and the new target posture information into the multimodal large model LLaVA to obtain a new action sequence; S3, controlling the robotic arm to execute the first N actions in the new action sequence; S4, repeat steps S2-S3 until the robotic arm performs the preset injection end action.
3. The method according to claim 1, characterized in that The multimodal large model LLaVA uses visual question-answering technology to conduct multiple rounds of dialogue on the teaching injection video to understand the content, and learn the mapping relationship between the syringe posture information, the posture information of the target to be injected, and the action sequence, including: Segmenting and sampling the teaching injection video to obtain continuous video frame images; Determining initial syringe pose information of the syringe and initial target pose information of the target object to be injected in the teaching injection video; The multimodal large model LLaVA uses visual question-answering technology to select an initial action in the continuous video frame images from multiple injection actions in the action skill library through a multi-round dialogue task, and generates an initial action sequence corresponding to the initial syringe posture information and the initial target posture information based on the initial action; Obtaining action sequence labels in the teaching injection video, wherein the action sequence labels are determined based on robot vaccine injection operation skill knowledge; The multimodal large model LLaVA is trained based on the action sequence label and the initial action sequence, so that the multimodal large model LLaVA learns and obtains the mapping relationship between the syringe posture information, the target posture information to be injected and the action sequence.
4. The method according to claim 1, wherein The syringe posture information includes the position information and spatial posture information of the syringe, and the spatial posture information of the syringe is obtained by: Establishing a 3D model of the syringe; Using a deep learning-based visual algorithm to extract feature information of the syringe in the current injection image; The feature information is matched with the 3D model to capture the spatial posture information of the syringe.
5. The method according to claim 1, wherein The action includes an action type and action parameters.
6. A motion inference device for immunization injection, the device is used to control the motion of the syringe through the mechanical arm of the injection robot, characterized in that: The device comprises: An acquisition module, configured to acquire a current injection image, wherein the current injection image includes a syringe and a target to be injected; A first determining module is configured to determine syringe posture information and a target posture information to be injected according to the current injection image; The second determination module is configured to input the syringe posture information and the target posture information into a pre-trained multimodal large model LLaVA to obtain an action sequence corresponding to the syringe posture information and the target posture information, wherein the action sequence includes a plurality of actions; the multimodal large model LLaVA uses visual question-answering technology to conduct multiple rounds of dialogue and accumulate content understanding of the teaching injection video, and learn the mapping relationship between the syringe posture information, the target posture information, and the action sequence; A control module is used to control the robotic arm to perform actions in the action sequence.
7. The device according to claim 6, characterized in that The control module is specifically used for: S1, controlling the robotic arm to execute the first N actions in the action sequence to obtain an injection image after the actions are executed; wherein N is a positive integer; S2, using the posture information of the syringe in the injection image after the action is performed as new syringe posture information, using the posture information of the target to be injected in the injection image after the action is performed as new target posture information, and inputting the new syringe posture information and the new target posture information into the multimodal large model LLaVA to obtain a new action sequence; S3, controlling the robotic arm to execute the first N actions in the new action sequence; S4, repeat steps S2-S3 until the robotic arm performs the preset injection end action.
8. The device according to claim 6, characterized in that The device further comprises a training device; the training device is used for: Segmenting and sampling the teaching injection video to obtain continuous video frame images; Determining initial syringe pose information of the syringe and initial target pose information of the target object to be injected in the teaching injection video; The multimodal large model LLaVA uses visual question-answering technology to select an initial action in the continuous video frame images from multiple injection actions in the action skill library through a multi-round dialogue task, and generates an initial action sequence corresponding to the initial syringe posture information and the initial target posture information based on the initial action; Obtaining action sequence labels in the teaching injection video, wherein the action sequence labels are determined based on robot vaccine injection operation skill knowledge; The multimodal large model LLaVA is trained based on the action sequence label and the initial action sequence, so that the multimodal large model LLaVA learns and obtains the mapping relationship between the syringe posture information, the target posture information to be injected and the action sequence.
9. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 5 when executed by a processor.
Citation Information
Patent Citations
Robot operation pose control method based on visual-touch multi-scale positioning
CN112060085A
Mechanical arm control method, device and equipment based on large visual model and storage medium
CN118143940A