Machine Learning-Based Industrial Robot Vision Positioning Method and System
Through the preset picture recognition model and target direction recognition model of machine learning, the actual position of the target object of industrial robots is directly identified, the computing steps are simplified, the recognition speed and efficiency are improved, and the shortcomings of visual positioning technology in the existing technology are solved.
Patent Information
- Application Number
- CN202411660012.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-11-20
AI Technical Summary
For industrial robots, existing visual positioning technology has redundant computing steps, which affects the recognition speed and efficiency, and is difficult to adapt to the needs of a single working scenario.
The preset image recognition model and target direction recognition model based on machine learning are adopted. The preset image recognition model assists in training the target direction recognition model, directly identifying the actual pose of the target object, simplifying the computing steps and improving training efficiency.
The calculation amount of visual positioning of industrial robots has been greatly simplified, the recognition speed and efficiency have been improved, and the problem of poor visual positioning experience in the prior art has been solved.
Smart Images

Figure CN119600104B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial vision image recognition, and particularly to an industrial robot vision positioning method and system based on machine learning. Background Art
[0002] Industrial robots are indispensable automation devices in modern industry. They perform highly repetitive tasks through programming to improve production efficiency and output, while reducing human errors and enhancing operational safety. Vision positioning technology is a crucial auxiliary technology for industrial robots.
[0003] Vision positioning technology is a technology that uses visual information to determine the position of an object in space. It usually involves computer vision and image processing technologies, and identifies the position, orientation, and motion state of an object by analyzing and interpreting visual data. In the prior art, vision positioning technology generally includes multiple steps such as image recognition, feature extraction, feature comparison, and coordinate transformation.
[0004] However, in practice, the working scenarios of industrial robots are relatively simple, generally for guiding and rectifying, simple assembly, etc., and the similarity of the images that need to be recognized during their work is also relatively high. Therefore, for industrial robots, the prior art seems to be relatively redundant. On the one hand, complex calculation steps may reduce the vision recognition speed, thereby affecting the production operation speed. On the other hand, the cumbersome recognition process also increases the difficulty of optimizing the algorithm strategy for specific scenarios, making the actual application efficiency of the prior art not ideal. Summary of the Invention
[0005] Therefore, the present invention provides an industrial robot vision positioning method and system based on machine learning to solve the problem that the vision positioning technology in the prior art has an unsatisfactory user experience for industrial robots.
[0006] The present invention provides an industrial robot vision positioning method based on machine learning, including:
[0007] Obtaining a preset picture recognition model and a target direction recognition model, wherein the preset picture recognition model is used to recognize the probability that the image content is a target object in a preset pose, and this probability is negatively correlated with the deviation between the actual pose of the target object and the preset pose, and the target direction recognition model is a neural network model used to recognize the actual pose of the target object in the image;
[0008] Training the target direction recognition model in combination with the preset picture recognition model, wherein the optimization intensity of each iteration of the target direction recognition model is positively correlated with the probability output by the preset picture recognition model;
[0009] Obtain the actual image of the target object, and based on the trained target orientation recognition model, obtain the actual pose of the target object;
[0010] Perform visual positioning on the target object according to the actual pose of the target object.
[0011] In a preferred implementation: the preset image recognition model is a visual bag-of-words model. The visual bag-of-words model includes a variety of visual words, and each visual word corresponds to a contour feature. The preset image recognition model is used to obtain the probability that the image content is the target object in the preset pose according to the number of various visual words contained in the image.
[0012] In a preferred implementation: training the target orientation recognition model in combination with the preset image recognition model includes:
[0013] Obtain the preset pose vector, the target training image, and the corresponding true label;
[0014] Preprocess the target training image to obtain the input image;
[0015] Input the input image into the preset image recognition model to obtain the object recognition probability;
[0016] Input the input image into the target orientation recognition model to obtain the orientation recognition vector;
[0017] Establish a loss function according to the preset pose vector, the true label, the orientation recognition vector, and the object recognition probability;
[0018] Optimize the target orientation recognition model according to the loss function.
[0019] In a preferred implementation: preprocessing the target training image to obtain the input image includes:
[0020] Scale the target training image to obtain the first initial image;
[0021] Offset the target training image to obtain the second initial image;
[0022] Preprocess the first initial image and the second initial image respectively to obtain multiple input images.
[0023] In a preferred implementation: establishing a loss function according to the preset pose vector, the true label, the orientation recognition vector, and the object recognition probability includes:
[0024] Establish a first loss function according to the orientation recognition vector and the true label;
[0025] Establish a difference eigenvalue according to the difference between the orientation recognition vector and the preset pose vector;
[0026] According to the numerical level gap between the difference eigenvalue and the object recognition probability, perform feature scaling on the direction recognition vector to obtain a correction vector;
[0027] According to the correction vector and the true label, establish a second loss function;
[0028] According to the first loss function and the second loss function, establish a loss function.
[0029] In a preferred implementation: According to the numerical level gap between the difference eigenvalue and the object recognition probability, performing feature scaling on the direction recognition vector to obtain a correction vector includes:
[0030] According to the numerical level gap between the difference eigenvalue and the object recognition probability, obtain a scaling coefficient;
[0031] According to the scaling coefficient, correct the direction recognition vector to increase the difference between different elements in the direction recognition vector to obtain a correction vector.
[0032] In a preferred implementation: The loss function is:
[0033] ;
[0034] Wherein, represents the loss function, represents the first loss function, represents the second loss function, is a preset weight;
[0035] The second loss function is specifically:
[0036] ;
[0037] Wherein, represents the number of elements, represents the total number of elements, represents the th element in the true label, represents the th element in the correction vector.
[0038] In a preferred implementation: Performing visual positioning on the target object according to the actual pose of the target object includes:
[0039] According to the actual pose of the target object, adjust the position of the target shooting device;
[0040] Based on the target shooting device, capture the target object at the adjusted position to obtain a secondary acquisition picture;
[0041] Based on the secondarily acquired image, perform visual positioning on the target object based on a preset image recognition model.
[0042] The present invention also provides an industrial robot vision positioning system based on machine learning, including:
[0043] A model preparation module, configured to obtain a preset image recognition model and a target direction recognition model, wherein the preset image recognition model is used to recognize the probability that the image content is a target object in a preset pose, and this probability is negatively correlated with the deviation between the actual pose of the target object and the preset pose, and the target direction recognition model is a neural network model, used to recognize the actual pose of the target object in the image;
[0044] A model optimization module, configured to train the target direction recognition model in combination with the preset image recognition model, wherein the optimization intensity of each iteration of the target direction recognition model is positively correlated with the probability output by the preset image recognition model;
[0045] A pose recognition module, configured to obtain the actual image of the target object, and based on the trained target direction recognition model, obtain the actual pose of the target object;
[0046] A vision positioning module, configured to perform visual positioning on the target object according to the actual pose of the target object.
[0047] The beneficial effects of adopting the above embodiments are:
[0048] The present invention provides an industrial robot vision positioning method and system based on machine learning. First, it obtains a preset image recognition model and a target direction recognition model, then trains the target direction recognition model in combination with the preset image recognition model, then obtains the actual image of the target object, and based on the trained target direction recognition model, obtains the actual pose of the target object, and finally performs visual positioning on the target object according to the actual pose of the target object. Compared with the prior art, because the application scenarios of industrial robots in practice are single, the present invention directly uses the target direction recognition model to recognize the actual pose of the target object, without performing steps such as feature recognition and extraction on the image itself, greatly simplifying the computational amount of actual applications. In addition, in practice, the objects to be recognized by industrial robots are often unchanged, only the viewing angles are different, so there is a situation where the probability of the preset image recognition model recognizing an object is negatively correlated with the deviation between the actual pose and the preset pose. The present invention utilizes this point and further assists in training the target direction recognition model through the preset image recognition model, improving the efficiency of the whole process from development to application, and solving the problem that the visual positioning technology in the prior art has an unsatisfactory use experience for industrial robots. Description of the Drawings
[0049] Figure 1It is the method flowchart of the industrial robot vision positioning method based on machine learning provided by the present invention;
[0050] Figure 2 For Figure 1 It is the specific step diagram of step S102 in
[0051] Figure 3 For Figure 2 It is the specific step diagram of step S205 in
[0052] Figure 4 It is the system structure diagram of the industrial robot vision positioning system based on machine learning provided by the present invention. Specific embodiments
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0054] Combined with Figure 1 As shown, a specific embodiment of the present invention discloses an industrial robot vision positioning method based on machine learning, including:
[0055] S101. Obtain a preset picture recognition model and a target direction recognition model. Among them, the preset picture recognition model is used to recognize the probability that the image content is a target object in a preset pose, and this probability is negatively correlated with the deviation between the actual pose of the target object and the preset pose. The target direction recognition model is a neural network model used to recognize the actual pose of the target object in the image;
[0056] S102. Train the target direction recognition model in combination with the preset picture recognition model, where the optimization intensity of each iteration of the target direction recognition model is positively correlated with the probability output by the preset picture recognition model;
[0057] S103. Obtain the actual picture of the target object, and based on the trained target direction recognition model, obtain the actual pose of the target object;
[0058] S104. Perform visual positioning on the target object according to the actual pose of the target object.
[0059] Among them, the target object is an object with a shape to be recognized in practice, such as a specific workpiece, a hole, a crack, or the flatness of an edge on a structure. The preset image recognition model is a preset model for recognizing the target object, but there are certain requirements for the pose of the target object in the recognized image, that is, the preset pose. In practice, the preset image recognition model can be trained and provided by the equipment supplier. The target direction recognition model is used to recognize the specific position of the target object in the image, that is, the actual pose. It can be understood that the preset pose and the actual pose are relative concepts. According to the actual specific situation, it can refer to the position and posture of the target object, or the shooting angle of the shooting device.
[0060] For industrial robots, they often need to perform a large number of simple and repetitive tasks. The images they recognize often have high similarity. And in most cases, the ultimate goal of industrial robots is not to identify the specific type of the target object, but to analyze the actual pose of the target object in order to guide the actions of other instruments. This premise provides the feasibility for the neural network model to directly recognize the pose. Therefore, in this embodiment, the target direction recognition model is used to directly recognize the actual pose of the target object to improve the operation efficiency of the algorithm during actual work. Similarly, this premise also makes the output result of the preset image recognition model more linear. Theoretically, when the target object is in the preset pose, the recognition result output by the preset image recognition model, that is, the probability that the image content is the target object is 100%. And the greater the deviation of the actual pose of the target object from the preset pose, the greater the deformation of the target object reflected in the image, and then the lower the probability output by the preset image recognition model. By regarding the output probability as a measure of the accuracy of the output result of the target direction recognition model to assist in the training of the target direction recognition model, the training speed and accuracy of the target direction recognition model can be greatly improved.
[0061] For example, if the probability output by the preset image recognition model is high, intuitively it indicates that the probability of an image containing the target object is high. And in this embodiment, it can be understood that the difference between the actual pose and the preset pose of the target object is small. At this time, if the result input to the target direction recognition is greatly different from the preset pose, the learning rate can be increased to accelerate the convergence speed. On the contrary, if the result input to the target direction recognition is slightly different from the preset pose, the learning rate can be decreased to prevent overfitting.
[0062] Therefore, compared with the prior art, the present invention directly uses the target direction recognition model to recognize the actual pose of the target object without performing steps such as feature recognition and extraction on the picture itself, greatly simplifying the computational complexity of practical applications. In addition, the present invention further assists in training the target direction recognition model through a preset picture recognition model, improving the efficiency of the whole process from development to application, and solving the problem that the visual positioning technology in the prior art has an unsatisfactory user experience for industrial robots.
[0063] Further, in a preferred embodiment, the preset picture recognition model is a bag-of-visual-words model. The bag-of-visual-words model includes a variety of visual words, and each visual word corresponds to a contour feature. The preset picture recognition model is used to obtain the probability that the image content is the target object in a preset pose according to the number of various visual words included in the picture.
[0064] The bag-of-visual-words technology (BoVW) is an image representation method widely used in the field of computer vision. It draws on the bag-of-words model in natural language processing. This method regards an image as a set composed of a series of "visual words", similar to the set of words in a text, without considering the spatial arrangement order of these words. For this embodiment, the bag-of-visual-words technology has a natural advantage in generating the probability with a linear relationship with the pose deviation. For example, assuming that a variety of tiny line segments of a certain length (straight lines with different extension angles, curves with different curvatures, broken lines with different included angles, etc.) are contour features and are used as visual words, then pictures of the target object in different poses can be represented by a variety of different visual words. By statistically analyzing the distribution of these visual words, the probability that the image content is the target object can be obtained.
[0065] Further, as shown in Figure 2 In a preferred embodiment, step S102, training the target direction recognition model in combination with the preset picture recognition model, specifically includes:
[0066] S201. Obtain a preset pose vector, a target training picture, and the corresponding true label;
[0067] S202. Preprocess the target training picture to obtain an input picture;
[0068] S203. Input the input picture into the preset picture recognition model to obtain an object recognition probability;
[0069] S204. Input the input picture into the target direction recognition model to obtain a direction recognition vector;
[0070] S205. Establish a loss function according to the preset pose vector, the true label, the direction recognition vector, and the object recognition probability;
[0071] S206. Optimize the target direction recognition model according to the loss function.
[0072] In the above process, the target training image is the image actually used for training the target direction model, the true label is the vector representing the actual pose of the target training image, the preset pose vector is the vector representing the preset pose, and the direction recognition vector is the vector representing the pose of the target object predicted by the target direction recognition model. The specific encoding method of these vectors representing poses can be flexibly set according to actual needs. For example, by using two unit vectors, the position angle of a part can be determined. At this time, the above vectors are all 6-dimensional vectors. The above process uses the preset image recognition model to assist learning by optimizing the loss function. Compared with the method of adjusting the learning rate mentioned above, its correction is more detailed and accurate.
[0073] In a preferred embodiment, the above step S202, preprocess the target training image to obtain the input image, specifically including:
[0074] Scale the target training image to obtain the first initial image;
[0075] Offset the target training image to obtain the second initial image;
[0076] Preprocess the first initial image and the second initial image respectively to obtain multiple input images.
[0077] For image recognition technology, when the image size is fixed and no position information is specified in advance, the algorithm will only obtain the analysis result based on the distribution of pixels themselves. This leads to the situation that for target images with the same pose, if they appear at different positions or sizes in the image, their recognition results may also be different. Therefore, in the above process of preprocessing the target training image, in addition to the conventional processes such as equalization and binarization, two additional steps of scaling and offset are added. The purpose is to simulate the situation where the size and position of the target object are different due to the pose change during shooting in the actual application scenario of industrial robots, and increase the number of samples to further eliminate errors.
[0078] Further, as shown in Figure 3 In a preferred embodiment, the above step S205, establish a loss function according to the preset pose vector, true label, direction recognition vector, and object recognition probability, specifically including:
[0079] S301. Establish a first loss function according to the direction recognition vector and the true label;
[0080] S302. Establish a difference eigenvalue according to the difference between the direction recognition vector and the preset pose vector;
[0081] S303. Scale the feature of the direction recognition vector according to the numerical level gap between the difference eigenvalue and the object recognition probability to obtain a correction vector;
[0082] S304. Establish a second loss function according to the correction vector and the true label;
[0083] S305. Establish a loss function according to the first loss function and the second loss function.
[0084] In the above process, the first loss function can be regarded as the definition of the loss function in the prior art, that is, the loss obtained by analyzing the direction recognition vector itself, while the second loss function is added in the present invention on the basis of the prior art. The second loss function is established based on the difference eigenvalue, and the difference eigenvalue is used to represent the difference between the direction recognition vector and the preset pose vector, such as the vector angle between two vectors, the cosine similarity between two vectors, the Euclidean distance, etc. The numerical level gap refers to the difference between the difference degree of the direction recognition vector and the preset pose vector represented by the difference eigenvalue and the difference degree of the actual pose and the preset pose of the target object represented by the object recognition probability, that is, the gap between the difference levels represented by the numerical values of these two data, namely the difference eigenvalue and the object recognition probability. How to calculate this difference can be flexibly set according to the actual situation.
[0085] The present invention scales the feature of the direction recognition vector according to the numerical level gap between the difference eigenvalue and the object recognition probability to increase or decrease the gap between different elements in the direction recognition vector, and establishes a second loss function based on this. In this way, the second loss function can be regarded as combining the probability output by the preset picture recognition model to further amplify or reduce the gap between different elements in the direction recognition vector and the true label, obtain a correction vector containing more correction information, and analyze the loss obtained from this correction vector.
[0086] For example, when the difference represented by the difference eigenvalue is large, the numerical value of each element in the direction recognition vector can be adjusted to make the difference between each element in the direction recognition vector and the true label larger, so as to increase the loss size, thereby increasing the correction strength of the target direction recognition model, and vice versa.
[0087] The present invention provides a more preferred method. Specifically, in a preferred embodiment, the above step S303. Scale the feature of the direction recognition vector according to the numerical level gap between the difference eigenvalue and the object recognition probability to obtain a correction vector, specifically includes:
[0088] Obtain a scaling coefficient according to the numerical level gap between the difference eigenvalue and the object recognition probability;
[0089] According to the scaling factor, the direction recognition vector is corrected to increase the difference between different elements in the direction recognition vector, and a corrected vector is obtained.
[0090] The main purpose of the above process is to increase the difference between different elements in the direction recognition vector to maintain the total scale invariance.
[0091] Substantially, the present invention is an improvement of the existing technology of knowledge distillation. In knowledge distillation, a loss function is also added to use the teacher model to assist the training of the student model. In this embodiment, a second loss function is actually added, so that the target direction recognition model can learn the knowledge about recognizing the target object and the preset pose in the preset picture recognition model, achieve faster convergence, and improve the flexibility of the application of machine learning technology.
[0092] Specifically, in a preferred embodiment, the loss function is:
[0093] ;
[0094] Wherein, represents the loss function, represents the first loss function, represents the second loss function, is a preset weight;
[0095] The second loss function is specifically:
[0096] ;
[0097] Wherein, represents the number of elements, represents the total number of elements, represents the th element in the true label, represents the th element in the corrected vector;
[0098] The corrected vector is expressed as:
[0099] ;
[0100] Wherein, represents the th element in the direction recognition vector, represents the scaling factor;
[0101] The calculation formula of the scaling factor is:
[0102] ;
[0103] Wherein, represents the direction recognition vector, Indicates the true label, Indicates the vector angle between the direction recognition vector and the preset pose vector, Indicates the object recognition probability.
[0104] Furthermore, in a preferred embodiment, the above step S104, visually positioning the target object according to the actual pose of the target object, specifically includes:
[0105] Adjust the position of the target shooting device according to the actual pose of the target object;
[0106] Based on the target shooting device, shoot the target object at the adjusted position to obtain a secondary acquisition picture;
[0107] According to the secondary acquisition picture, based on the preset picture recognition model, visually position the target object.
[0108] It can be understood that the above process is only a preferred visual positioning method. In practice, after knowing the actual pose, other methods can also be adopted according to specific situations. In this embodiment, the target shooting device is first adjusted to the position where the target object with the preset pose can be shot, so that the existing preset picture recognition model can be called for visual positioning, which improves the code reusability, reduces the bloatedness of the actual software, and at the same time reduces the deployment cost.
[0109] Combined with Figure 4 As shown, the present invention also provides an industrial robot visual positioning system based on machine learning, including:
[0110] A model preparation module 410, configured to obtain a preset picture recognition model and a target direction recognition model, wherein the preset picture recognition model is used to recognize the probability that the image content is the target object in the preset pose, and this probability is negatively correlated with the deviation between the actual pose and the preset pose of the target object, and the target direction recognition model is a neural network model for recognizing the actual pose of the target object in the image;
[0111] A model optimization module 420, configured to train the target direction recognition model in combination with the preset picture recognition model, wherein the optimization strength of each iteration of the target direction recognition model is positively correlated with the probability output by the preset picture recognition model;
[0112] A pose recognition module 430, configured to obtain the actual picture of the target object and obtain the actual pose of the target object based on the trained target direction recognition model;
[0113] A visual positioning module 440, configured to visually position the target object according to the actual pose of the target object.
[0114] It should be noted here that: The corresponding system provided in the above embodiments can implement the technical solutions described in the above method embodiments. For the specific implementation principles of the above modules or units, reference can be made to the corresponding content in the above method embodiments, which will not be elaborated here.
[0115] The present invention provides an industrial robot vision positioning method and system based on machine learning. First, a preset image recognition model and a target direction recognition model are obtained. Then, the target direction recognition model is trained in combination with the preset image recognition model. After that, the actual image of the target object is obtained, and based on the trained target direction recognition model, the actual pose of the target object is obtained. Finally, visual positioning of the target object is performed according to the actual pose of the target object. Compared with the prior art, since the application scenarios of industrial robots are single in practice, the present invention directly uses the target direction recognition model to identify the actual pose of the target object without performing steps such as feature recognition and extraction on the image itself, greatly simplifying the computational complexity of practical applications. In addition, in practice, the objects to be recognized by industrial robots are often unchanged, only the viewing angles are different, which results in a situation where the probability of the preset image recognition model recognizing an object is negatively correlated with the deviation between the actual pose and the preset pose. The present invention utilizes this point and further assists in training the target direction recognition model through the preset image recognition model, improving the efficiency of the whole process from development to application and solving the problem that the visual positioning technology has an unsatisfactory user experience for industrial robots in the prior art.
[0116] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0117] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An industrial robot vision positioning method based on machine learning, characterized in that, Including: Obtain a preset image recognition model and a target orientation recognition model. Among them, the preset image recognition model is used to recognize the probability that the image content is a target object in a preset pose, and this probability is negatively correlated with the deviation between the actual pose of the target object and the preset pose. The target orientation recognition model is a neural network model used to recognize the actual pose of the target object in the image; Train the target orientation recognition model in combination with the preset image recognition model; Obtain the actual image of the target object, and based on the trained target orientation recognition model, obtain the actual pose of the target object; Perform visual positioning on the target object according to the actual pose of the target object; Among them, training the target orientation recognition model in combination with the preset image recognition model includes: Obtain a preset pose vector, target training images, and corresponding true labels; Preprocess the target training images to obtain input images; Input the input images into the preset image recognition model to obtain object recognition probabilities; Input the input images into the target orientation recognition model to obtain orientation recognition vectors; Establish a loss function according to the preset pose vector, true label, orientation recognition vector, and object recognition probability; Optimize the target orientation recognition model according to the loss function; Among them, establishing a loss function according to the preset pose vector, true label, orientation recognition vector, and object recognition probability includes: Establish a first loss function according to the orientation recognition vector and the true label; Establish a difference eigenvalue according to the difference between the orientation recognition vector and the preset pose vector; Perform feature scaling on the orientation recognition vector according to the numerical level gap between the difference eigenvalue and the object recognition probability to obtain a corrected vector; Establish a second loss function according to the corrected vector and the true label; Establish a loss function according to the first loss function and the second loss function.
2. The machine learning-based industrial robot vision positioning method according to claim 1, wherein The preset image recognition model is a visual bag-of-words model. The visual bag-of-words model includes multiple visual words, and each visual word corresponds to a contour feature. The preset image recognition model is used to obtain the probability that the image content is a target object in a preset pose according to the number of multiple visual words included in the image.
3. The method for visual positioning of an industrial robot based on machine learning according to claim 1, wherein Preprocessing the target training images to obtain input images includes: Scale the target training images to obtain first initial images; Offset the target training images to obtain second initial images; Preprocess the first initial images and the second initial images respectively to obtain multiple input images.
4. The machine learning-based industrial robot vision positioning method according to claim 1, wherein Performing feature scaling on the orientation recognition vector according to the numerical level gap between the difference eigenvalue and the object recognition probability to obtain a corrected vector includes: Obtain a scaling coefficient according to the numerical level gap between the difference eigenvalue and the object recognition probability; Correct the orientation recognition vector according to the scaling coefficient to increase the difference between different elements in the orientation recognition vector to obtain a corrected vector.
5. The machine learning-based industrial robot vision positioning method according to claim 4, wherein, The loss function is: ; Among them, represents the loss function, represents the first loss function, represents the second loss function, is a preset weight; The second loss function is specifically: ; Among them, represents the number of elements, represents the total number of elements, represents the th element in the true label, represents the th element in the correction vector.
6. The method for visual positioning of an industrial robot based on machine learning according to claim 1, wherein Performing visual positioning on the target object according to the actual pose of the target object includes: Adjust the position of the target shooting device according to the actual pose of the target object; Based on the target shooting device, capture the target object at the adjusted position to obtain secondary acquisition images; Based on the secondarily acquired images, visually locate the target object based on a preset image recognition model.
7. An industrial robot vision positioning system based on machine learning, characterized in that, Including: A model preparation module, configured to obtain a preset image recognition model and a target orientation recognition model. The preset image recognition model is used to identify the probability that the image content is a target object in a preset pose, and this probability is negatively correlated with the deviation between the actual pose of the target object and the preset pose. The target orientation recognition model is a neural network model used to identify the actual pose of the target object in the image. A model optimization module, configured to train the target orientation recognition model in combination with the preset image recognition model. The optimization strength of each iteration of the target orientation recognition model is positively correlated with the probability output by the preset image recognition model. A pose recognition module, configured to obtain the actual image of the target object and obtain the actual pose of the target object based on the trained target orientation recognition model. A visual positioning module, configured to perform visual positioning on the target object according to the actual pose of the target object. Among them, training the target orientation recognition model in combination with the preset image recognition model includes: Obtain a preset pose vector, target training images, and corresponding ground truth labels. Preprocess the target training images to obtain input images. Input the input images into the preset image recognition model to obtain object recognition probabilities. Input the input images into the target orientation recognition model to obtain orientation recognition vectors. Establish a loss function according to the preset pose vector, ground truth label, orientation recognition vector, and object recognition probability. Optimize the target orientation recognition model according to the loss function. Among them, establishing a loss function according to the preset pose vector, ground truth label, orientation recognition vector, and object recognition probability includes: Establish a first loss function according to the orientation recognition vector and the ground truth label. Establish a difference eigenvalue according to the difference between the orientation recognition vector and the preset pose vector. Feature scale the orientation recognition vector according to the numerical level gap between the difference eigenvalue and the object recognition probability to obtain a corrected vector. Establish a second loss function according to the corrected vector and the ground truth label. Establish a loss function according to the first loss function and the second loss function.
Citation Information
Patent Citations
Visual robot grabbing method and system based on deep learning
CN111723782A
Industrial robot posture recognition method and device and storage medium
CN115270399A