Mirror-holding robot control method and system based on multimodal question-answering large model

By applying the control method of multimodal Q&A big model in the mirror-holding robot, the problem of insufficient tracking accuracy and stability of the mirror-holding robot is solved, and higher tracking accuracy and stability are achieved.

CN118664603BActive Publication Date: 2025-05-06SUN YAT SEN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410997464.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2025-05-06
Estimated Expiration
2044-07-24

AI Technical Summary

Technical Problem

The mirror-holding robot in the prior art has problems of low accuracy and unstableness in field of view tracking.

Method used

The control method based on the multimodal question and answer big model is adopted, and the images captured by the mirror-holding robot are obtained, and the image features are extracted using the pre-trained multimodal question and answer big model are generated, and the trajectory planning and movement are performed according to the current end position and coordinate error.

Benefits of technology

The tracking accuracy and stability of the mirror-holding robot are improved, and the current coordinates of the instrument tip in the image can be accurately obtained, and the target coordinates are predicted through the multimodal question and answer big model, thereby achieving accurate tracking and stable control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118664603B_ABST
    Figure CN118664603B_ABST
Patent Text Reader

Abstract

The present application discloses a control method and system for a mirror-holding robot based on a multimodal question-answering large model, which relates to the field of robot control technology. The method includes: obtaining an image captured by the mirror-holding robot; using a pre-trained multimodal question-answering large model to obtain the target coordinates corresponding to the image; determining the coordinates of the tip of the instrument in the image as the current coordinates, and determining the coordinate error between the current coordinates and the target coordinates; obtaining the current joint angle of the mirror-holding robot, and calculating the current terminal posture of the mirror-holding robot according to the current joint angle; determining the target terminal posture corresponding to the target coordinate according to the coordinate error and the current terminal posture; sending the target terminal posture to the mirror-holding robot, and the mirror-holding robot performs trajectory planning and movement according to the target terminal posture. The current coordinates of the tip of the instrument in the image can be accurately obtained, the target coordinates can be accurately predicted using the multimodal question-answering large model, and the mirror-holding robot can be driven according to the coordinate error to achieve precise tracking and stable control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of robot control technology, and in particular to a control method and system for a mirror-holding robot based on a multimodal question-answering large model. Background Art

[0002] The operator can observe the operation area through the arthroscope installed at the end of the mirror-holding robot. In some operation scenarios, the operation space is very small and multiple instruments and tools need to be used in coordination. Therefore, the position of the arthroscope needs to be adjusted continuously and accurately to adjust the field of view of the arthroscope in real time. However, the mirror-holding robot of the related technology still has the problem of inaccurate and unstable field of view tracking. Summary of the invention

[0003] The main purpose of the embodiments of the present application is to propose a control method and system for a mirror-holding robot based on a multimodal question-answering large model to improve the tracking accuracy and stability of the mirror-holding robot.

[0004] To achieve the above purpose, one aspect of an embodiment of the present application proposes a control method for a mirror-holding robot based on a multimodal question-answering large model, the method comprising the following steps:

[0005] Acquire images captured by the mirror-holding robot;

[0006] Obtaining target coordinates corresponding to the image using a pre-trained multimodal question-answering model;

[0007] determining the coordinates of the tip of the instrument in the image as current coordinates, and determining a coordinate error between the current coordinates and the target coordinates;

[0008] Acquire the current joint angle of the mirror-holding robot, and calculate the current end position of the mirror-holding robot according to the current joint angle;

[0009] Determine the target terminal posture corresponding to the target coordinates according to the coordinate error and the current terminal posture;

[0010] The target end position and posture are sent to the mirror-holding robot, and the mirror-holding robot performs trajectory planning and movement according to the target end position and posture.

[0011] In some embodiments, the step of obtaining the target coordinates corresponding to the image using a pre-trained multimodal question-answering model comprises the following steps:

[0012] Extracting image features of the image using a pre-trained visual encoding model in the multimodal question-answering large model;

[0013] Converting the image features into image embedding tags having the same dimension as the word embedding space in the language model using a projection matrix in the multimodal question-answering large model; wherein the multimodal question-answering large model includes the language model;

[0014] The expression of the image embedding mark is:

[0015] H v =W·Z v ,Z v =g(X v );

[0016] Among them, H v represents the image embedding label, W represents the trainable projection matrix, Z v represents the image features, X v representing said image;

[0017] Generate corresponding dialogue data according to the image embedding mark and the multiple rounds of language guidance text;

[0018] generating a question-answer sequence according to the dialogue data, and generating a training instruction according to the question-answer sequence;

[0019] The expression of the training instruction is:

[0020]

[0021] The conversation data is T is the total number of rounds, and t is the current round number; is the question text of the tth round corresponding to the image, is the answer text of the tth round corresponding to the image; is the training instruction corresponding to the tth round of the image;

[0022] Determine a prediction sequence on the original autoregressive training target according to the training instruction using the pre-trained multimodal question-answering large model, and determine a target answer according to the prediction sequence;

[0023] The expression of the probability of the target answer is:

[0024]

[0025] Among them, p(X A |X v ,X instruct ) is the probability of the target answer, L is the question-answer sequence, X A is the target answer, xi is the predicted sequence, θ is a trainable parameter, X instruct,<i and X A,<iare the current predicted sequence x i All previous training instructions and answer tokens;

[0026] The target coordinates corresponding to the image are determined according to the target answer.

[0027] In some embodiments, before using the pre-trained multimodal question-answering large model to obtain the target coordinates corresponding to the image, the method further includes the step of training the multimodal question-answering large model, and the step of training the multimodal question-answering large model includes:

[0028] Performing a forward propagation and a backward propagation on the projection matrix using the first training data;

[0029] The target data set and the multi-layer perceptron network are used to adjust the set parameters in the multimodal question-answering large model, and then the second training data is used to perform two forward propagations and two back propagations on the multimodal question-answering large model and the projection matrix to obtain the pre-trained multimodal question-answering large model; wherein the amount of data of the second training data is greater than that of the first training data.

[0030] In some embodiments, determining the coordinates of the tip of the instrument in the image as the current coordinates comprises the following steps:

[0031] Inputting the image into a deep learning-based target detection model to obtain a labeling box that labels the tip of the instrument in the image; wherein the tip of the instrument is at the geometric center of the labeling box;

[0032] Determine the pixel coordinates of the geometric center of the annotation box as the current coordinates;

[0033] The target detection model based on deep learning is obtained through pre-training, and the training steps of the target detection model based on deep learning include:

[0034] Acquire multiple simulated images in a simulated scene as an image data set;

[0035] Annotating each of the simulated images in the image data set with the detection frame, and the tip of the instrument in each of the simulated images is located at the geometric center of the corresponding detection frame;

[0036] Separating the image data set using a five-fold cross validation method to obtain a plurality of image sub-data sets;

[0037] The deep learning-based object detection model is trained using each of the image sub-datasets.

[0038] In some embodiments, the step of determining the target end pose corresponding to the target coordinates according to the coordinate error and the current end pose comprises the following steps:

[0039] Determining pixel coordinates of each pixel in the image according to the current end position;

[0040] Convert each pixel coordinate into a camera coordinate, and then convert each camera coordinate into a world coordinate;

[0041] The conversion relationship from the pixel coordinates to the camera coordinates includes:

[0042]

[0043] Among them, f u , f v , c u , c v is the internal parameter of the camera of the mirror holding robot, [X c ,Y c ,Z c ] is the camera coordinate, [u,v] is the pixel coordinate;

[0044] The conversion relationship between the camera coordinates and the world coordinates includes:

[0045]

[0046] Wherein, t is the translation vector of the camera, R is the rotation matrix of the camera, O 1×3 is a 1-row, 3-column all-zero vector; [X, Y, Z] is the world coordinate;

[0047] The conversion relationship between the pixel coordinates and the world coordinates includes:

[0048]

[0049] The target end position and posture corresponding to the target coordinates are determined according to the coordinate error and the world coordinates.

[0050] In some embodiments, the mirror-holding robot performs trajectory planning and movement according to the target end position, including the following steps:

[0051] The mirror-holding robot performs trajectory planning according to the target terminal posture to obtain a planned path;

[0052] The mirror-holding robot executes the movement corresponding to the planned path according to the control algorithm;

[0053] The calculation formula of the control algorithm includes:

[0054]

[0055] c px =||s mid -s tip ||2;

[0056]

[0057] in, is the output of the control algorithm, s mid is the current coordinate, s tip is the target coordinate; K track is the control function; K p,track and K d,track Respectively represent the proportional control coefficient and the differential control parameter; and are respectively the expected moving distance corresponding to the image of the current frame and the expected moving distance corresponding to the image of the previous frame; is the Jacobian matrix J of the camera of the mirror holding robot c The inverse of is the coordinate error; λ is an adjustable parameter.

[0058] In some embodiments, the method further comprises the following steps:

[0059] Get the question text;

[0060] The question text and the image are input into the pre-trained multimodal question-answering model to obtain corresponding answer data.

[0061] To achieve the above purpose, another aspect of the embodiment of the present application proposes a control system for a mirror-holding robot based on a multimodal question-answering large model, the system comprising:

[0062] An image acquisition unit, used to acquire images taken by the mirror-holding robot;

[0063] A coordinate acquisition unit, used to acquire the target coordinates corresponding to the image using a pre-trained multimodal question-answering large model;

[0064] a coordinate error confirmation unit, for determining the coordinates of the tip of the instrument in the image as the current coordinates, and determining the coordinate error between the current coordinates and the target coordinates;

[0065] A current posture acquisition unit, used to acquire the current joint angle of the mirror-holding robot, and calculate the current end posture of the mirror-holding robot according to the current joint angle;

[0066] A target posture acquisition unit, used to determine the target terminal posture corresponding to the target coordinates according to the coordinate error and the current terminal posture;

[0067] The control driving unit is used to send the target end position to the mirror holding robot, and the mirror holding robot performs trajectory planning and movement according to the target end position.

[0068] To achieve the above objective, another aspect of an embodiment of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above method when executing the computer program.

[0069] To achieve the above objective, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above method when executed by a processor.

[0070] The embodiments of the present application include at least the following beneficial effects:

[0071] The present application can obtain the image captured by the mirror-holding robot; use the pre-trained multimodal question-answering large model to obtain the target coordinates corresponding to the image; determine the coordinates of the tip of the instrument in the image as the current coordinates, and determine the coordinate error between the current coordinates and the target coordinates; obtain the current joint angle of the mirror-holding robot, and calculate the current terminal posture of the mirror-holding robot based on the current joint angle; determine the target terminal posture corresponding to the target coordinate based on the coordinate error and the current terminal posture; send the target terminal posture to the mirror-holding robot, and the mirror-holding robot performs trajectory planning and movement based on the target terminal posture. The present application can accurately obtain the current coordinates of the tip of the instrument in the image by identifying the image captured by the mirror-holding robot, and use the pre-trained multimodal question-answering large model to accurately predict the target coordinates based on the image, and then drive the mirror-holding robot to move to the target coordinate with the corresponding posture based on the coordinate error between the current coordinates and the target coordinates, so as to achieve precise tracking and stable control. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0073] Figure 1 A schematic flow chart of a method for controlling a mirror-holding robot based on a multimodal question-answering large model provided in an embodiment of the present application;

[0074] Figure 2 The overall architecture diagram of the multimodal question-answering model provided in the embodiment of the present application;

[0075] Figure 3 A schematic diagram of generating graphic data provided in an embodiment of the present application;

[0076] Figure 4 An example flow chart of training a multimodal question-answering model provided in an embodiment of the present application;

[0077] Figure 5 A working example diagram of a mirror-holding robot provided in an embodiment of the present application;

[0078] Figure 6 An example control flow chart of the mirror-holding robot provided in an embodiment of the present application;

[0079] Figure 7 A control algorithm flow chart of a mirror-holding robot provided in an embodiment of the present application;

[0080] Figure 8 An interactive flow chart of various nodes when controlling a mirror-holding robot provided in an embodiment of the present application;

[0081] Fig. 9 A schematic diagram of the structure of a mirror-holding robot control system based on a multimodal question-answering large model provided in an embodiment of the present application;

[0082] Fig.10 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0083] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of systems and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.

[0084] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".

[0085] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0086] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0087] Reference Figure 1 The embodiment of the present application provides a control method for a mirror-holding robot based on a multimodal question-answering large model. The method may include but is not limited to S100 to S150, which are as follows:

[0088] S100: Acquire an image captured by the mirror-holding robot.

[0089] It can be understood that the arthroscope (i.e., the camera of the mirror-holding robot) can be set at the end of the mirror-holding robot, and the image in the embodiment of the present application can be an image taken by the arthroscope.

[0090] S110: Obtain the target coordinates corresponding to the image using a pre-trained multimodal question-answering large model.

[0091] This embodiment can use the image-text alignment method to fine-tune the multimodal question-answering model based on a large-scale image-text dataset that widely covers operating scenarios. Specifically, the multimodal question-answering model first learns the operating scenario vocabulary using image-text alignment, and then uses the language model in the multimodal question-answering model to expand the obtained data title into the generated instruction tracking data to learn the open-ended conversation semantics, and widely imitate how laymen gradually acquire professional knowledge of operating scenarios. As an end-to-end model, the image-text dataset has an overall architecture of two parts. The first is the input alignment part, which includes the image vision processing alignment part and the text alignment processing part. The outputs of the two parts are spliced ​​to form the multimodal question-answering model input, and then the second part, the language model part, is used to predict the output decoding.

[0092] Furthermore, S110 may include S111 to S116:

[0093] S111: extracting image features of the image using a pre-trained visual encoding model in the multimodal question-answering large model;

[0094] S112: using the projection matrix in the multimodal question-answering large model to convert the image features into image embedding tags having the same dimension as the word embedding space in the language model; wherein the multimodal question-answering large model includes the language model;

[0095] The expression of the image embedding mark is:

[0096] H v =W·Z v ,Z v =g(X v );

[0097] Among them, H v represents the image embedding label, W represents the trainable projection matrix, Z v represents the image features, X v representing said image;

[0098] S113: Generate corresponding dialogue data according to the image embedding mark and the multiple rounds of language guidance texts;

[0099] S114: generating a question-answer sequence according to the dialogue data, and generating a training instruction according to the question-answer sequence;

[0100] The expression of the training instruction is:

[0101]

[0102] The conversation data is T is the total number of rounds, t is the current round number; is the question text of the tth round corresponding to the image, is the answer text of the tth round corresponding to the image; is the training instruction corresponding to the tth round of the image;

[0103] S115: using the pre-trained multimodal question-answering large model to determine a prediction sequence on the original autoregressive training target according to the training instruction, and determining a target answer according to the prediction sequence;

[0104] The expression of the probability of the target answer is:

[0105]

[0106] Among them, p(X A |X v ,X instruct ) is the probability of the target answer, L is the question-answer sequence, X A is the target answer, x i is the prediction sequence, θ is a trainable parameter, X instruct,<i and X A,<i are the current predicted sequence x i All previous training instructions and answer tokens;

[0107] S116: Determine the target coordinates corresponding to the image according to the target answer.

[0108] Specifically, for the input image X v , use the pre-trained visual encoding model to obtain the image feature Z v =g(X v ). Considering the speed of training the multimodal question answering model itself, this embodiment can use a simple linear layer to connect the image features to the word embedding space. Specifically, the trainable projection matrix W is used to transform Z v Converted to image embedding tags with the same dimensions as the word embedding space in the language model, that is:

[0109] H v =W·Z v ,with Z v =g(X v ); (1)

[0110] For each image X v , which can generate multi-round dialogue data Where T is the total number of rounds. The dialogue data consists of a question-answer sequence, all answers are considered as responses, and the training instructions are sent to the tth round. Defined as:

[0111]

[0112] The training instructions specify a uniform format for the question-answer sequence, where the image embedding tokens are input first, followed by the text question-answer pairs in the corresponding format. The language model is then instructed to predict the sequence on its original autoregressive training objective.

[0113] Specifically, for a question-answer sequence of length L, the target answer X A The probability can be expressed as:

[0114]

[0115] Where θ is a trainable parameter, X instruct,<i and X A,<i They are the current prediction sequence x i All previous instructions and answer tokens. In the condition of (3), the model explicitly adds X v , to emphasize that all answers are based on images, and to omit all prompts such as "The question is: The answer is:", etc., to improve readability. The overall architecture of the multimodal question answering model is as follows Figure 2 shown.

[0116] Furthermore, before S110, the embodiment of the present application may further include the step of training a multimodal question-answering large model, specifically including:

[0117] Performing a forward propagation and a backward propagation on the projection matrix using the first training data;

[0118] The target data set and the multi-layer perceptron network are used to adjust the set parameters in the multimodal question-answering large model, and then the second training data is used to perform two forward propagations and two back propagations on the multimodal question-answering large model and the projection matrix to obtain the pre-trained multimodal question-answering large model; wherein the amount of data of the second training data is greater than that of the first training data.

[0119] Specifically, when training a large multimodal question-answering model, it is first encoded through a visual encoding model, and a model processing method similar to CLIP is used to convert the image into an image embedding tag. The image is converted into an image embedding tag using multiple text interpretation dimensions. In order to improve the training effect, this embodiment can use the ViT-L / 14-336 model as a visual encoding model.

[0120] like Figure 3As shown, in order to align the image embedding tag with the corresponding image text description, the original LLaVA uses a simple linear layer to construct a projection layer, and projects the matrix information after image extraction to the dimension of text instruction fine-tuning through the projection layer, thereby achieving dimensional alignment. For the projection effect, this embodiment can use a two-layer multi-layer perceptron network (Multi-Layer Perceptron, MLP), considering that the continuous image only changes its features in certain small ranges, but different operations are required due to the changes in features in a small range during specific operations, so the multimodal question-answering large model needs to pay attention to the changes in features in a small range. The image embedding tag is simply embedded using the BERT model, and the image embedding tag is initialized using parameters. Then, in order to train a language model with a target output, the language model can be attached with a fixed prompt with a question, such as "The question is..., the answer is...", and the question is encoded using a language encoding model. The text dimension and the dimension after image alignment are then spliced ​​and embedded in the multimodal question-answering large model.

[0121] In view of the language model pre-trained with professional knowledge of operating scenarios and the effect of instruction tuning, and considering the insufficient amount of training data, both stages of training use the training data for full training, taking advantage of the pre-trained language model and sufficient amount of second-stage training data to effectively make up for the relatively small amount of training data.

[0122] In the first stage, by freezing the weights of the visual encoding model and the multimodal question-answering model, only the projection matrix is ​​updated, so that the multimodal question-answering model is aligned with the operation scene information, and the language model pre-trained with the professional knowledge of the operation scene can assist the training to achieve convergence faster. In order to prevent overfitting, the training of the projection matrix only trains one Epoch (i.e., one forward propagation and one back propagation). In the second stage, the weights of the visual encoding model are retained, and the pre-trained weights of the projection matrix and the multimodal question-answering model are continuously updated, and two Epochs are trained in the second stage. Considering that the first stage requires a large amount of data, a long training time, and the main function is pre-training alignment, the language model optimized by the data fine-tuning of the operation scene is used, and the multi-layer perceptron network is used at the same time to fully improve the performance of the multimodal question-answering model, thereby effectively reducing the defects caused by the insufficient number of training data and achieving a better training goal. In addition, the present embodiment can also increase the amount of training data in the second stage so that the language model can fully learn the alignment information of the operation scene.

[0123] The combination of the three can effectively improve the fine-tuning training effect of the multimodal question-answering model, and can achieve the effect of aligning the operation scene information while greatly reducing the cost of model training and data collection. The example flow chart of training a multimodal question-answering model can be referred to Figure 4 .

[0124] S120: Determine the coordinates of the tip of the instrument in the image as the current coordinates, and determine the coordinate error between the current coordinates and the target coordinates.

[0125] Further, S120 may include S121-S122:

[0126] S121: Inputting the image into a deep learning-based target detection model to obtain a labeling box that labels the tip of the instrument in the image; wherein the tip of the instrument is at the geometric center of the labeling box;

[0127] S122: Determine the pixel coordinates of the geometric center of the annotation box as the current coordinates;

[0128] The target detection model based on deep learning is obtained by pre-training before S120, and the training steps of the target detection model based on deep learning include:

[0129] Acquire multiple simulated images in a simulated scene as an image data set;

[0130] Annotating each of the simulated images in the image data set with the detection frame, and the tip of the instrument in each of the simulated images is located at the geometric center of the corresponding detection frame;

[0131] Separating the image data set using a five-fold cross validation method to obtain a plurality of image sub-data sets;

[0132] The deep learning-based object detection model is trained using each of the image sub-datasets.

[0133] S130: Acquire the current joint angle of the mirror-holding robot, and calculate the current end position and posture of the mirror-holding robot according to the current joint angle.

[0134] Specifically, assume that the joint angles of the mirror-holding robot are:

[0135] θ=[θ1,θ2,…,θ m ] T ; (4)

[0136] Its forward kinematic equation can be expressed as:

[0137]

[0138] in, It is the homogeneous transformation matrix from the m-1th coordinate system to the mth coordinate system.

[0139] S140: Determine the target terminal posture corresponding to the target coordinates according to the coordinate error and the current terminal posture.

[0140] Further, S140 may include S141 to S143:

[0141] S141: Determine pixel coordinates of each pixel in the image according to the current end position;

[0142] S142: converting each pixel coordinate into a camera coordinate, and then converting each camera coordinate into a world coordinate;

[0143] The conversion relationship from the pixel coordinates to the camera coordinates includes:

[0144]

[0145] Among them, f u , f v , c u , c v is the internal parameter of the camera of the mirror holding robot, [X c ,Y c ,Z c ] is the camera coordinate, [u,v] is the pixel coordinate;

[0146] The conversion relationship between the camera coordinates and the world coordinates includes:

[0147]

[0148] Wherein, t is the translation vector of the camera, R is the rotation matrix of the camera, O 1×3 is a 1-row, 3-column all-zero vector; [X, Y, Z] is the world coordinate;

[0149] The conversion relationship between the pixel coordinates and the world coordinates includes:

[0150]

[0151] S143: Determine the target end position and posture corresponding to the target coordinates according to the coordinate error and the world coordinates.

[0152] Specifically, the internal and external parameters of the camera can be obtained by calibrating the arthroscope (i.e., the camera of the mirror-holding robot). Combined with the known current end position, the target end position can be obtained through camera correction, which specifically includes: performing distortion correction based on the obtained pixel points, which is mainly used to remove the distortion error of the camera head, including radial distortion and tangential distortion. Subsequently, the current position of the camera is obtained, the pixel points captured by the camera at the current position are converted into corrected points, and the corrected points are mapped to the world coordinate system to determine the current position of the camera. Next, the error of the end effector is calculated to obtain the error between the current position and the target position of the camera, which can be used to correct the position of the end effector of the mirror-holding robot. Finally, the movement instruction is calculated, and the new target end position is calculated using the current end position and the error of the end effector.

[0153] According to the pinhole imaging principle and camera calibration parameters, the transformation relationship from the pixel coordinate system to the camera coordinate system can be expressed as:

[0154]

[0155] Among them, f u , f v , c u , c v is the camera’s internal parameter, [X c ,Y c ,Z c ] is the coordinate in the camera coordinate system, and [u,v] is the coordinate in the pixel coordinate system.

[0156] According to the camera's translation vector t and rotation matrix R, the conversion relationship between the world coordinate system and the camera coordinate system can be obtained as follows:

[0157]

[0158] Among them, O 1×3 is a 1-row, 3-column all-zero vector.

[0159] Substituting equation (6) into equation (7), we can obtain the transformation relationship from the pixel coordinate system to the world coordinate system:

[0160]

[0161] If the pixel coordinates [u, v] are known, the corresponding world coordinates can be calculated according to equation (8), and then the target end position of the mirror-holding robot can be obtained.

[0162] After obtaining the target end position, you can interact with the mirror robot. To achieve tracking of the mirror robot, you can first obtain the joint angle of the mirror robot, take an angle as the initial joint angle, and substitute the value of the joint angle into the forward kinematics equation fkine(Θ), and you can get the end position of the mirror robot at the initial moment. Then, combined with the camera's transformation matrix, you can calculate the target end position of the mirror robot. The mirror robot plans its own trajectory by calling the node and finally reaches the target end position.

[0163] S150: Sending the target end position to the mirror-holding robot, and the mirror-holding robot performs trajectory planning and movement according to the target end position.

[0164] Furthermore, the mirror-holding robot performs trajectory planning and movement according to the target terminal position and posture, which may include S151-S152:

[0165] S151: the mirror-holding robot performs trajectory planning according to the target terminal posture to obtain a planned path;

[0166] S152: the mirror-holding robot executes the movement corresponding to the planned path according to the control algorithm;

[0167] The calculation formula of the control algorithm includes:

[0168]

[0169] c px =||s mid -s tip ||2;

[0170]

[0171] in, is the output of the control algorithm, s mid is the current coordinate, s tip is the target coordinate; K track is the control function; K p,track and K d,track Respectively represent the proportional control coefficient and the differential control parameter; and are respectively the expected moving distance corresponding to the image of the current frame and the expected moving distance corresponding to the image of the previous frame; is the Jacobian matrix J of the camera of the mirror holding robot c The inverse of is the coordinate error; λ is an adjustable parameter.

[0172] Specifically, in order to improve the tracking stability, this embodiment designs a control algorithm for tracking pixel coordinates. According to the deep learning model, the target coordinates of the instrument tip in the current frame image can be calculated as s tip , set the center coordinates of the image (as the current coordinates of the instrument tip) to Among them, w s With h s are the width and height of the image obtained by the camera respectively. Since the pixel coordinate system is a two-dimensional image, the real-time tracking error of the frame image can be expressed by the Euclidean distance, that is, c px =||s mid -s tip ||2, where c px Indicates the expected moving distance.

[0173] The iterative criterion of the control algorithm is that the difference between the tracked target coordinates and the image field center coordinates is within the calculation threshold d err Because the threshold is small, it can be considered that the tip of the instrument in the center of the field of view coincides with the target coordinates, that is, the pixel distance is the smallest. Therefore, the calculation formula of the control algorithm includes:

[0174]

[0175] in, is the output of the control algorithm, K track is the control function; K p,track and K d,track Respectively represent the proportional control coefficient and the differential control parameter; and are the expected moving distance corresponding to the current frame image and the expected moving distance corresponding to the previous frame image respectively; is the Jacobian matrix J of the camera of the mirror holding robot c The inverse of is the coordinate error; λ is an adjustable parameter.

[0176] The above control algorithm is applied to the position control of pixel coordinates, that is, the pixel coordinate center obtained each time is subtracted from the target coordinates, and the difference is calculated through the control algorithm to calculate the pixel coordinate error, and then the current joint angle is read, and the current end position obtained by forward kinematics can be combined with the camera's coordinate transformation to convert the coordinate error into the target end position, and then the trajectory planning is performed through the robot's nodes, and a cycle ends. Since the goal is to minimize Considering the stability of instrument tracking and actual needs, a threshold δ is set here px , that is, the pixel error c when stable px <δ px , stop the tracking process, and the position and posture of the end of the mirror-holding robot remain unchanged.

[0177] In addition, in order to provide prompt information to the operator, the embodiment of the present application may also include:

[0178] Get the question text;

[0179] The question text and the image are input into the pre-trained multimodal question-answering model to obtain corresponding answer data.

[0180] It can be understood that the answer data may include text, image, audio and other data.

[0181] Next, the solution of the embodiment of the present application will be introduced and explained in conjunction with specific application examples:

[0182] Reference Figure 5 , this embodiment provides a working example diagram of a mirror-holding robot.

[0183] This embodiment adopts the "host computer-slave computer" working mode. The host computer includes a computer, and uses PyQt5 to build the front-end framework of the software interface to interact with the independently built slave computer (i.e., the mirror-holding robot). The mirror-holding robot includes a robotic arm, an arthroscope, connecting parts and operating instruments.

[0184] Figure 6 This is the control flow chart of the mirror-holding robot. The arthroscope provides images through the camera, performs image detection through the real-time target detection model, and calculates the center of the detection frame and feeds it back to the robotic arm. The robotic arm reads the current joint angle, calculates the current end position of the mirror-holding robot using forward kinematics, obtains the center of the detection frame to obtain the target coordinates, obtains the target end position through the camera coordinate transformation relationship, and then obtains the movement trajectory by calculating the inverse kinematics, and uses the trajectory planning function inside the robotic arm to guide the robotic arm to move to the target coordinates.

[0185] The multimodal question-answering model starts the Web service by deploying the model in the server, selecting to upload an image, asking questions, sending the information to Requests, and then getting the reply to the question and answer and returning it, and finally getting the corresponding analysis results. The local interactive interface is implemented using PyQt5, integrating modules such as camera tracking, image uploading, and interactive question-answering, and realizing multiple functions such as auxiliary operations based on images, assisting image detection and tracking of mirror-holding robots, etc.

[0186] The relationship between the joint angle and the target end position and posture can be determined through the forward and inverse kinematics of the mirror-holding robot. The target planning value can be transmitted to the mirror-holding robot, and it can communicate with the mirror-holding mechanical arm to realize operation tracking.

[0187] There can be a reasonable tracking center point to achieve the tracking task. In order to maintain field of view tracking within the small field of view of the arthroscope, this embodiment can track the field of view center point of the operating instrument or the target area in the operating area. Therefore, it is necessary to identify and segment the operating instrument and the operating environment. In order to effectively identify the operating instrument and the operating environment, this embodiment can collect a data set of corresponding operations of several operating instruments in a specific operating scenario, so as to train a target detection model based on deep learning and use the target detection model to detect the instrument tip in the image. When the arthroscope is turned on, the target detection model can detect the instrument tip in the field of view in real time and generate a detection frame for it. By analyzing multiple detection frames and then calculating the center, the tracking of the instrument tip can be achieved. Moreover, the image can be understood and analyzed through a large multimodal question and answer model and the text can be output. A control algorithm flow chart of a mirror-holding robot is shown below. Figure 7 shown.

[0188] This embodiment separates the motion control and command control of the mirror-holding robot. The visual control software of the mirror-holding robot is run on the host computer to implement command control. The host computer provides control commands to the motion control system of the mirror-holding robot. This embodiment can improve scalability and facilitate operation and demonstration.

[0189] The hardware system of the mirror-holding robot of this embodiment may include (1) a computer: connected to the mirror-holding robot through a network, obtaining joint angles, postures and images, generating motion instructions through visual servo control, and outputting the motion instructions to the mirror-holding robot; (2) an arthroscope: obtaining real-time images of the surgical environment, and connecting to a computer through an image acquisition card to obtain real-time images of the arthroscope in the host computer; (3) an arthroscope connector: connecting the end of the mirror-holding robot and the arthroscope. The drawings of the connectors are designed using 3D drawing software and then 3D printed, and the connectors are fixed by screws. (4) a robotic arm: providing the port of the arthroscope for the computer to read, and realizing joint control according to the end posture motion instructions sent by the computer, thereby adjusting the posture of the arthroscope in real time.

[0190] The command software is built on the host computer. After the mirror robot starts driving, to achieve the visual servo tracking task, the host computer needs to subscribe to the topic and read the mirror robot information. After target detection calculation and kinematic modeling, the target end position of the mirror robot is obtained, and then the target end position is returned to the mirror robot through the node. After receiving the command, the mirror robot performs trajectory planning and moves to the target position. The above process is executed once every cycle. The interaction flow chart of each node when controlling the mirror robot is as follows Figure 8As shown. The entire task deployment requires starting the arthroscope and publishing a topic, then calling the target detection algorithm to generate a detection frame for the tip of the instrument, then reading the joint angles of the mirror-holding robot, and calculating the target end position, and then publishing the target end position to the mirror-holding robot to achieve tracking. Considering that the visual servo tracking task requires the input of multiple start commands and requires higher operating efficiency, the host computer of this embodiment uses PyQt5 as the front-end framework, and arranges the arthroscope display frame in the upper left corner of the display interface to achieve the effect of tracking detection.

[0191] Reference Fig. 9 The embodiment of the present application also provides a control system for a mirror-holding robot based on a multimodal question-answering large model, which can implement the above control method. The system includes:

[0192] An image acquisition unit, used to acquire images taken by the mirror-holding robot;

[0193] A coordinate acquisition unit, used to acquire the target coordinates corresponding to the image using a pre-trained multimodal question-answering large model;

[0194] a coordinate error confirmation unit, for determining the coordinates of the tip of the instrument in the image as the current coordinates, and determining the coordinate error between the current coordinates and the target coordinates;

[0195] A current posture acquisition unit, used to acquire the current joint angle of the mirror-holding robot, and calculate the current end posture of the mirror-holding robot according to the current joint angle;

[0196] A target posture acquisition unit, used to determine the target terminal posture corresponding to the target coordinates according to the coordinate error and the current terminal posture;

[0197] The control driving unit is used to send the target end position to the mirror holding robot, and the mirror holding robot performs trajectory planning and movement according to the target end position.

[0198] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0199] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above control method when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.

[0200] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0201] See also Fig.10 , Fig.10 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:

[0202] The processor 1001 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0203] The memory 1002 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store an operating system and other application programs. When the technical solution provided in the embodiments of this specification is implemented by software or firmware, the relevant program code is stored in the memory 1002, and the processor 1001 calls and executes the control method of the embodiment of this application;

[0204] Input / output interface 1003, used to implement information input and output;

[0205] The communication interface 1004 is used to realize the communication interaction between the device and other devices. The communication can be realized through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);

[0206] A bus 1005 , which transmits information between various components of the device (e.g., the processor 1001 , the memory 1002 , the input / output interface 1003 , and the communication interface 1004 );

[0207] The processor 1001 , the memory 1002 , the input / output interface 1003 and the communication interface 1004 are connected to each other in communication within the device via the bus 1005 .

[0208] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, and the computer program implements the above-mentioned control method when executed by a processor.

[0209] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0210] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0211] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0212] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0213] The system embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.

[0214] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0215] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0216] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0217] In the several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or units, which can be electrical, mechanical or other forms.

[0218] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0219] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0220] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.

[0221] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. A control method for a mirror-holding robot based on a multimodal question-answering large model, characterized in that: The method comprises the following steps: Acquire images captured by the mirror-holding robot; Obtaining target coordinates corresponding to the image using a pre-trained multimodal question-answering model; determining the coordinates of the tip of the instrument in the image as current coordinates, and determining a coordinate error between the current coordinates and the target coordinates; Acquire the current joint angle of the mirror-holding robot, and calculate the current end position of the mirror-holding robot according to the current joint angle; Determine the target terminal posture corresponding to the target coordinates according to the coordinate error and the current terminal posture; The target end position and posture are sent to the mirror-holding robot, and the mirror-holding robot performs trajectory planning and movement according to the target end position and posture; The method of obtaining the target coordinates corresponding to the image by using the pre-trained multimodal question-answering model comprises the following steps: Extracting image features of the image using a pre-trained visual encoding model in the multimodal question-answering large model; Converting the image features into image embedding tags having the same dimension as the word embedding space in the language model using a projection matrix in the multimodal question-answering large model; wherein the multimodal question-answering large model includes the language model; The expression of the image embedding mark is: H v =W·Z v ,Z v =g(X v ); Among them, H v represents the image embedding label, W represents the trainable projection matrix, Z v represents the image features, X v representing said image; Generate corresponding dialogue data according to the image embedding mark and the multiple rounds of language guidance text; generating a question-answer sequence according to the dialogue data, and generating a training instruction according to the question-answer sequence; The expression of the training instruction is: The conversation data is T is the total number of rounds, and t is the current round number; is the question text of the tth round corresponding to the image, is the answer text of the tth round corresponding to the image; is the training instruction corresponding to the tth round of the image; Determine a prediction sequence on the original autoregressive training target according to the training instruction using the pre-trained multimodal question-answering large model, and determine a target answer according to the prediction sequence; The expression of the probability of the target answer is: Among them, p(X A |X v ,X instruct ) is the probability of the target answer, L is the question-answer sequence, X A is the target answer, x i is the prediction sequence, θ is a trainable parameter, X instruct,<i and X A,<i are the current predicted sequence x i All previous training instructions and answer tokens; The target coordinates corresponding to the image are determined according to the target answer.

2. The control method of the mirror-holding robot based on the multimodal question-answering large model according to claim 1 is characterized in that: Before using the pre-trained multimodal question-answering large model to obtain the target coordinates corresponding to the image, the method further includes the step of training the multimodal question-answering large model, and the step of training the multimodal question-answering large model includes: Performing a forward propagation and a backward propagation on the projection matrix using the first training data; The target data set and the multi-layer perceptron network are used to adjust the set parameters in the multimodal question-answering large model, and then the second training data is used to perform two forward propagations and two back propagations on the multimodal question-answering large model and the projection matrix to obtain the pre-trained multimodal question-answering large model; wherein the amount of data of the second training data is greater than that of the first training data.

3. The control method of the mirror-holding robot based on the multimodal question-answering large model according to claim 1 is characterized in that: Determining the coordinates of the tip of the instrument in the image as the current coordinates comprises the following steps: Inputting the image into a deep learning-based target detection model to obtain a labeling box of the tip of the instrument in the image; wherein the tip of the instrument is located at the geometric center of the labeling box; Determine the pixel coordinates of the geometric center of the annotation box as the current coordinates; The target detection model based on deep learning is obtained through pre-training, and the training steps of the target detection model based on deep learning include: Acquire multiple simulated images in a simulated scene as an image data set; Annotating each of the simulated images in the image data set with the annotation frame, and the tip of the instrument in each of the simulated images is located at the geometric center of the corresponding annotation frame; Separating the image data set using a five-fold cross validation method to obtain a plurality of image sub-data sets; The deep learning-based object detection model is trained using each of the image sub-datasets.

4. The control method of the mirror-holding robot based on the multimodal question-answering large model according to claim 1 is characterized in that: The step of determining the target terminal posture corresponding to the target coordinates according to the coordinate error and the current terminal posture comprises the following steps: Determining pixel coordinates of each pixel in the image according to the current terminal posture; Convert each pixel coordinate into a camera coordinate, and then convert each camera coordinate into a world coordinate; The conversion relationship from the pixel coordinates to the camera coordinates includes: Among them, f u , f v , c u , c v is the internal parameter of the camera of the mirror holding robot, [X c ,Y c ,Z c ] is the camera coordinate, [u,v] is the pixel coordinate; The conversion relationship between the camera coordinates and the world coordinates includes: Wherein, t is the translation vector of the camera, R is the rotation matrix of the camera, O 1×3 is a 1-row, 3-column all-zero vector; [X, Y, Z] is the world coordinate; The conversion relationship between the pixel coordinates and the world coordinates includes: The target end position and posture corresponding to the target coordinates are determined according to the coordinate error and the world coordinates.

5. The control method of the mirror-holding robot based on the multimodal question-answering large model according to claim 1 is characterized in that: The mirror-holding robot performs trajectory planning and movement according to the target end position, including the following steps: The mirror-holding robot performs trajectory planning according to the target terminal posture to obtain a planned path; The mirror-holding robot executes the movement corresponding to the planned path according to the control algorithm; The calculation formula of the control algorithm includes: c px =||s mid -s tip ||2; in, is the output of the control algorithm, s mid is the current coordinate, s tip is the target coordinate; K track is the control function; K p,track and K d,track Respectively represent the proportional control coefficient and the differential control parameter; and are respectively the expected moving distance corresponding to the image of the current frame and the expected moving distance corresponding to the image of the previous frame; is the Jacobian matrix j of the camera of the mirror holding robot c The inverse of is the coordinate error; λ is an adjustable parameter.

6. The control method of the mirror-holding robot based on the multimodal question-answering large model according to any one of claims 1 to 5, characterized in that: The method further comprises the following steps: Get the question text; The question text and the image are input into the pre-trained multimodal question-answering model to obtain corresponding answer data.

7. A mirror-holding robot control system based on a multimodal question-answering large model, characterized in that: The system comprises: An image acquisition unit, used to acquire images taken by the mirror-holding robot; A coordinate acquisition unit, used to acquire the target coordinates corresponding to the image using a pre-trained multimodal question-answering large model; a coordinate error confirmation unit, for determining the coordinates of the tip of the instrument in the image as the current coordinates, and determining the coordinate error between the current coordinates and the target coordinates; A current posture acquisition unit, used to acquire the current joint angle of the mirror-holding robot, and calculate the current end posture of the mirror-holding robot according to the current joint angle; A target posture acquisition unit, used to determine the target terminal posture corresponding to the target coordinates according to the coordinate error and the current terminal posture; A control driving unit is used to send the target end position to the mirror holding robot, and the mirror holding robot performs trajectory planning and movement according to the target end position; The method of obtaining the target coordinates corresponding to the image by using the pre-trained multimodal question-answering model comprises the following steps: Extracting image features of the image using a pre-trained visual encoding model in the multimodal question-answering large model; Converting the image features into image embedding tags having the same dimension as the word embedding space in the language model using a projection matrix in the multimodal question-answering large model; wherein the multimodal question-answering large model includes the language model; The expression of the image embedding mark is: H v =W·Z v ,Z v =g(X v ); Among them, H v represents the image embedding label, W represents the trainable projection matrix, Z v represents the image features, X v representing said image; Generate corresponding dialogue data according to the image embedding mark and the multiple rounds of language guidance text; generating a question-answer sequence according to the dialogue data, and generating a training instruction according to the question-answer sequence; The expression of the training instruction is: The conversation data is T is the total number of rounds, and t is the current round number; is the question text of the tth round corresponding to the image, is the answer text of the tth round corresponding to the image; is the training instruction corresponding to the tth round of the image; Determine a prediction sequence on the original autoregressive training target according to the training instruction using the pre-trained multimodal question-answering large model, and determine a target answer according to the prediction sequence; The expression of the probability of the target answer is: Among them, p(X A |X v ,X instruct ) is the probability of the target answer, L is the question-answer sequence, X A is the target answer, x i is the prediction sequence, θ is a trainable parameter, X instruct,<i and X A,<i are the current predicted sequence x i All previous training instructions and answer tokens; The target coordinates corresponding to the image are determined according to the target answer.

8. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Pathological image targeted cell detection method and labeling method

    CN117315653A

  • Visual servo method and system for lens-holding robot

    CN117901090A