A Visual Localization Method and System for an Automated Laboratory Based on the YOLO Algorithm

By using the YOLO algorithm in the automation laboratory for visual positioning and object recognition, combined with the method of visual field adjustment and supplementary image data acquisition, the problem of low visual positioning accuracy in the prior art is solved, and efficient object recognition and capture in complex scenarios is achieved.

CN119991784BActive Publication Date: 2025-07-01STATE GRID FUJIAN ELECTRIC POWER RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510483993.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-01
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing visual positioning methods have low accuracy in scenes where multiple targets are stored in a messy manner, and cannot effectively identify and locate target objects.

Method used

The visual positioning method of an automated laboratory based on the YOLO algorithm is adopted to collect initial image data through visual sensors, and target recognition is used to use the YOLO algorithm. The field of view is adjusted and supplementary image data is collected when there is no target object in the first recognition result, and the second target recognition and capture is performed.

Benefits of technology

It improves the accuracy of target recognition and ensures that the robot can complete the target recognition and grasping tasks when the target is partially blocked or the field of view is limited, which significantly improves the work efficiency and adaptability of the grabber robot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991784B_ABST
    Figure CN119991784B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of image analysis of neural networks, and discloses a visual positioning method and system for an automated laboratory based on the YOLO algorithm, including: First, precise environmental positioning is carried out using macroscopic and local field-of-view information, improving the accuracy of target recognition. Second, through dynamic field-of-view adjustment and supplementary image data acquisition, it is ensured that even when the target is partially occluded or the field of view is limited, the robot can still complete the target recognition and grasping tasks. The YOLO algorithm is used for target recognition, effectively improving the recognition speed and accuracy, and optimizing the recognition results through multiple iterations, thereby ensuring the success rate of the grasping tasks. Overall, the present invention can significantly improve the working efficiency and adaptability of the grasping robot, reduce the influence brought by environmental interference, and has high practical value and broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image analysis of neural networks, and specifically to a visual positioning method and system for an automated laboratory based on the YOLO algorithm. Background Art

[0002] With the continuous development of automation technology, grasping devices are increasingly widely used in industrial production, warehousing logistics, distribution centers and other fields. Especially in intelligent manufacturing and material handling, grasping robots can efficiently perform repetitive tasks, significantly improving work efficiency. However, the success of robot grasping tasks is usually affected by the positioning accuracy and grasping accuracy of the target object.

[0003] To achieve efficient automated grasping, the robot needs to accurately identify the position and posture of the target object, which requires the system to accurately acquire and analyze environmental information. Traditional grasping robot positioning methods usually rely on simple sensors or static models and are difficult to handle complex environments and dynamic changes. Especially in complex scenarios, such as when the target object is occluded, the environmental light changes, etc., the challenges of visual positioning are more prominent.

[0004] Therefore, how to use visual sensors and intelligent algorithms to achieve accurate grasping robot positioning and improve target recognition and grasping efficiency has become a current research hotspot. Solving this problem can not only improve the stability and accuracy of robot grasping, but also reduce manual intervention and improve the intelligence level of the automated system. Summary of the Invention

[0005] In view of the above existing problems, the present invention is proposed.

[0006] Therefore, the technical problem solved by the present invention is that the existing visual positioning methods have low accuracy and cannot perform recognition and positioning in the scenario of multi-target chaotic storage, etc.

[0007] To solve the above technical problem, the present invention provides the following technical solution: A visual positioning method for an automated laboratory based on the YOLO algorithm, including:

[0008] Collect initial image data of the target area through a visual sensor to perform field of view positioning on the grasping robot;

[0009] Perform target recognition on the initial image data to obtain the first recognition result, and perform target grasping according to the first recognition result;

[0010] When there is no target object in the first recognition result, adjust the field of view according to the first recognition result and collect supplementary image data of the target area;

[0011] Using the supplementary image data, perform target recognition again to obtain the recognition result of the second time, and perform target grasping according to the recognition result of the second time.

[0012] As a preferred solution of the visual positioning method of the automated laboratory based on the YOLO algorithm according to the present invention, wherein: the visual sensor includes, but is not limited to, a first type of visual sensor configured in the environment for obtaining macroscopic field-of-view information; a second type of visual sensor configured on the grasping robot for obtaining local field-of-view information;

[0013] The macroscopic field-of-view information is an image at the level of the target object storage space, including the position information of the grasping robot and the target area where the target object is located;

[0014] The local field-of-view information is an image at the level of the target area, including the target object in the target area.

[0015] As a preferred solution of the visual positioning method of the automated laboratory based on the YOLO algorithm according to the present invention, wherein: the field-of-view positioning includes constructing an environmental map of the current environment, and positioning the grasping robot and the target area on the environmental map according to the macroscopic field-of-view information;

[0016] After completing the positioning, analyze the current local field-of-view information; if the local field-of-view information contains a complete target area, directly collect the first frame of the image to be recognized;

[0017] If the target area is larger than the local field-of-view information, after marking the position of the grasping robot in the map, collect the first frame of the image to be recognized.

[0018] As a preferred solution of the visual positioning method of the automated laboratory based on the YOLO algorithm according to the present invention, wherein: the target recognition includes using the trained YOLO algorithm to recognize the local field-of-view information;

[0019] Input the grasping target and the first frame of the image into the YOLO algorithm to obtain the recognition result of the first time;

[0020] The recognition result of the first time includes the confidence that the i-th target object in the image is recognized as the grasping target ; preset two thresholds and , if , then determine that the i-th target object is recognized as the grasping target; if , then determine that the i-th target object is recognized as the object to be confirmed; if , then determine that the i-th target object is recognized as the non-graspable target;

[0021] When the number of the grasping targets in the recognition result reaches the requirement, directly perform grasping; when the number of the grasping targets in the recognition result does not reach the requirement, adjust the visual field of the grasping robot.

[0022] As a preferred solution of the visual positioning method of the automated laboratory based on the YOLO algorithm of the present invention, wherein: the visual field adjustment includes, if the local visual field information contains a complete target area, controlling the grasping robot to move to obtain the supplementary image data;

[0023] If the target area is larger than the local visual field information, make the grasping robot move from the current position to any edge on both sides, and collect the images of the target area in real time as the supplementary image data; after the first move, when the number of the grasping targets in the recognition result by the YOLO algorithm does not reach the requirement, return to the marked position and then move to the other edge, and perform recognition again after moving.

[0024] As a preferred solution of the visual positioning method of the automated laboratory based on the YOLO algorithm of the present invention, wherein: the second recognition result includes inputting the first frame of the supplementary image data into the YOLO algorithm to obtain the output result for each target object, and comparing the output result with the grasping target to obtain the final recognition result;

[0025] When performing target recognition again, assume the input image sequence , select m images from F to form ;

[0026] Use the YOLO algorithm to recognize the classification of the i-th target object in, if the classification result is a non-graspable target, then in , mark the i-th target object as a non-graspable target and stop the recognition of the i-th target object during the recognition process; if the classification result is a graspable target, then in , mark the i-th target object as a graspable target and stop the recognition of the i-th target object during the recognition process; if the classification result is an object to be confirmed, continue to recognize the i-th target object;

[0027] After completing the recognition of , perform the recognition of , and so on until the recognition is completed;

[0028] After completing m times of recognition, the unrecognized images only include objects to be confirmed. For the j-th object to be confirmed, obtain in The confidence of each image in the image set with respect to the grasping target; obtain k images with a confidence higher than g times the average value, and let any one of the images be , then and are the two ends of the acquisition interval;

[0029] In F, pair and , and the obtained image corresponds to and corresponding and ;

[0030] Utilize the image sequence between and to reselect m images for recognition until the maximum number of iterations is reached or the confidence no longer improves, and output all recognition results marked as the grasping target;

[0031] Among them, represents the m-th image in the image ; represents the n-th image in the image F; h represents the index of the image in M; e and u represent the indices of different images in F.

[0032] As a preferred solution of the vision positioning method of the automated laboratory based on the YOLO algorithm described in the present invention, wherein: the target grasping includes transmitting the pose information of the target object to the grasping robot control system, and the control system adjusts the motion trajectory of the robot arm according to the pose information and drives the grasping actuator to perform the grasping operation.

[0033] An automated laboratory vision positioning system adopting the method described in the present invention based on the YOLO algorithm, wherein:

[0034] A primary acquisition unit, through a vision sensor, acquires the initial image data of the target area and locates the field of view of the grasping robot;

[0035] A primary analysis unit, performs target recognition on the initial image data, obtains the first recognition result, and performs target grasping according to the first recognition result;

[0036] A secondary acquisition unit, when there is no target object in the first recognition result, adjusts the field of view according to the first recognition result and acquires the supplementary image data of the target area;

[0037] A secondary analysis unit, utilizes the supplementary image data to perform target recognition again, obtains the second recognition result, and performs target grasping according to the second recognition result.

[0038] A computer device, comprising: a memory and a processor; the memory stores a computer program, wherein: when the processor executes the computer program, the steps of any one of the methods of the present invention are implemented.

[0039] A computer-readable storage medium, on which a computer program is stored, wherein: when the computer program is executed by a processor, the steps of any one of the methods of the present invention are implemented.

[0040] Advantages of the present invention: The visual positioning method of the automated laboratory based on the YOLO algorithm provided by the present invention improves the accuracy of target recognition, ensuring that even when the target is partially occluded or the field of view is limited, the robot can still complete the target recognition and grasping tasks. Using the YOLO algorithm for target recognition effectively improves the recognition speed and accuracy, and optimizes the recognition results through multiple iterations, thus ensuring the success rate of the grasping task. Overall, the present invention can significantly improve the working efficiency and adaptability of the grasping robot, reduce the impact of environmental interference, and has high practical value and broad application prospects. Description of the Drawings

[0041] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0042] Figure 1 It is the overall flowchart of a visual positioning method of an automated laboratory based on the YOLO algorithm provided by the first embodiment of the present invention. Detailed Embodiments

[0043] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed embodiments of the present invention in conjunction with the drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0044] Example 1, referring to Figure 1 , which is an embodiment of the present invention, provides a visual positioning method of an automated laboratory based on the YOLO algorithm, including:

[0045] S1: Through a visual sensor, collect the initial image data of the target area and perform field of view positioning on the grasping robot.

[0046] Further, the first type of visual sensor is configured in the environment to obtain macroscopic field of view information; the second type of visual sensor is configured on the grasping robot to obtain local field of view information. The macroscopic field of view information is an image at the level of the storage space of the target object, including the position information of the grasping robot and the target area where the target object is located. The local field of view information is an image at the level of the target area, including the target object in the target area.

[0047] It should be noted that the first type of visual sensor configured in the environment is used to obtain an image of the entire target area, including the area where the target object is located and the position information of the grasping robot. This provides the system with global information about the environment, helping the robot understand the approximate location of the target area and ensuring that it can quickly determine the target area and the robot's current spatial position during target localization. The second type of visual sensor configured on the grasping robot is used to obtain a more detailed local image of the target area. By obtaining a detailed image within the target area, the robot can more accurately identify and locate the target object. This local field of view information especially helps the robot accurately identify the target in cases where the target object may be occluded or unevenly distributed. Through this hierarchical and complementary visual information acquisition method, the system can obtain the target position within a large range and perform precise target recognition and grasping in the local area, greatly improving the grasping accuracy and efficiency.

[0048] Furthermore, an environmental map of the current environment is constructed, and based on the macroscopic field of view information, the grasping robot and the target area are located on the environmental map. After the positioning is completed, the current local field of view information is analyzed; if the local field of view information contains a complete target area, the first frame of the image to be recognized is directly collected. If the target area is larger than the local field of view information, after marking the position of the grasping robot in the map, the first frame of the image to be recognized is collected.

[0049] By using the macroscopic field of view information to construct an environmental map and calibrate the positions of the grasping robot and the target area, it is mainly to clarify the edge positions of the field of view and the target area. This can help the system accurately determine the spatial range of the target area and provide accurate reference coordinates for subsequent local field of view acquisition and target recognition. Once the environmental positioning is completed, the system analyzes the local field of view information to determine the range of the target area. If the local field of view information already fully covers the entire target area, the system can directly collect the first frame of the image to be recognized for target recognition. This method reduces unnecessary calculations and improves processing efficiency. If the local field of view information cannot cover the entire target area, the system will mark the position of the grasping robot in the map according to the current positioning and then collect a new first frame of the image. In this way, the system can obtain more information about the target area by moving the robot's perspective to ensure that the target is fully recognized.

[0050] S2: Perform object recognition on the initial image data to obtain the first recognition result, and perform object grasping according to the first recognition result.

[0051] Object recognition includes using the trained YOLO algorithm to recognize the local field of view information. Input the grasping target and the first frame of the image into the YOLO algorithm to obtain the first recognition result.

[0052] The first recognition result includes the confidence that the i-th target object in the image is recognized as the grasping target. ; Preset two thresholds and , if , then determine that the i-th target object is recognized as the grasping target; if , then determine that the i-th target object is recognized as the object to be confirmed; if , then determine that the i-th target object is recognized as the non-graspable target.

[0053] When the number of the grasping targets in the recognition result reaches the requirement (the number of grasping targets in the laboratory may be multiple, that is, multiple grasps are required), then directly perform the grasping; when the number of the grasping targets in the recognition result does not reach the requirement, adjust the field of view of the grasping robot.

[0054] In YOLO, the algorithm divides the image into several grids, each grid predicts the bounding box of each object, and outputs the category and confidence of each object.

[0055] According to the image processing result, use the pose estimation algorithm to estimate the three-dimensional spatial position and pose of the target object relative to the grasping robot or the robot arm. Transmit the pose information of the target object to the grasping robot control system, and the control system adjusts the motion trajectory of the robot arm according to the pose information and drives the grasping actuator to perform the grasping operation. Perform precise grasping control according to the pose information to ensure that the grasping tool contacts the target object and successfully grasps it.

[0056] S3: When there is no target object in the first recognition result, adjust the field of view according to the first recognition result and collect the supplementary image data of the target area.

[0057] If the local field of view information contains a complete target area, control the grasping robot to move (it can move around the contour of the target area or move linearly forward and backward, depending on the specific settings), and obtain the supplementary image data. If the target area is larger than the local field of view information, move the grasping robot from its current position to any edge on both sides (the so-called edge is the edge of the field of view), and collect the images of the target area in real time as the supplementary image data; after the first move, when the number of grasping targets in the recognition result by the YOLO algorithm does not meet the requirements, return to the marked position and then move to the other edge, and perform recognition again after moving.

[0058] It should be noted that the above steps can ensure that even if the target cannot be successfully recognized in the first recognition result, the system can still adjust the field of view, continuously collect supplementary image data, expand the coverage of the target area, thereby improving the success rate of target recognition and finally completing the target grasping task. If the target area is larger than the current local field of view, the system will guide the grasping robot to move along the edge of the field of view to collect more image information of the target area in real time. By moving, the robot can continuously expand the field of view range to ensure that the complete target area can be captured and provide more data for the YOLO algorithm to process. After adjusting the field of view and collecting supplementary image data, the robot will perform target recognition again. If after the first move, the YOLO algorithm still fails to recognize enough grasping targets, the robot will return to the original position and move to the other edge to continue image collection and recognition until the requirements are met. This process ensures that the number of grasping targets can meet the requirements through iterative optimization, avoiding missing any potential grasping targets. At the same time, by directly moving back to the position, it can avoid wasting computing power due to repeated detection.

[0059] S4: Use the supplementary image data to perform target recognition again, obtain the second recognition result, and perform target grasping according to the second recognition result.

[0060] Input the first frame of the supplementary image data into the YOLO algorithm, obtain the output result for each target object, and compare the output result with the grasping target to obtain the final recognition result. Specifically, it includes:

[0061] When performing target recognition again, assume the input image sequence , select m images from F to form .

[0062] Use the YOLO algorithm to recognize the classification of the i-th target object in, if the classification result is a non-graspable target, then in Among them, the i-th target object is marked as a non-graspable target, and the recognition of the i-th target object is stopped during the recognition process; if the classification result is a grasp target, then in Among them, the i-th target object is marked as a grasp target, and the recognition of the i-th target object is stopped during the recognition process; if the classification result is an object to be confirmed, the recognition of the i-th target object continues.

[0063] After completing the recognition of , the recognition of is carried out, and so on until the recognition is completed.

[0064] After m times of recognition, the images that are not recognized only include objects to be confirmed. For the j-th object to be confirmed, obtain the confidence of each image in regarding the grasp target; obtain k images with a confidence higher than g times the average value. Let any one image be , then and are the two ends of the acquisition interval.

[0065] In F, and are corresponded, and the obtained image is corresponding to and ; and ;

[0066] Using the image sequence between and , re-select m images to perform recognition until the maximum number of iterations is reached or the confidence no longer improves, and output all the recognition results marked as grasp targets.

[0067] Among them, represents the m-th image in the image ; represents the n-th image in the image F; h represents the index of the image in M; e, u represent the indices of different images in F.

[0068] It should be noted that if there are multiple target regions, after completing the recognition of the first target region, the second target region is recognized until the target grasping is completed.

[0069] After collecting supplementary image data, the system performs object recognition again through the YOLO algorithm to obtain more visual information about the object. During this process, the YOLO algorithm classifies and judges each object, and determines whether it is a grasping target according to the preset classification rules. If it is recognized as a grasping target, it is immediately marked as a grasping target, and further recognition of this target is stopped. If it is recognized as a non-graspable target, it is marked as a non-graspable target, and the possibility of it being a grasping target is excluded. For those targets marked as objects to be confirmed, the system will continue with subsequent recognition until their grasping status is finally confirmed. Through iterative optimization, the system can gradually exclude targets that do not meet the grasping conditions, while continuing to recognize targets to be confirmed, and finally ensure that all eligible targets are accurately recognized. After each round of object recognition, the system calculates the confidence of each object, and filters out the images that are most likely to be grasping targets based on the confidence. Among all the images, the images with higher confidence will be used as the objects for priority recognition for further processing. Through multiple iterations, the system can continuously improve the recognition accuracy, ensuring that the recognized grasping targets are highly accurate. Finally, after multiple recognitions and iterations, the system will output all the recognition results marked as grasping targets, ensuring the success rate of the grasping operation and reducing the cases of misrecognition or missed recognition.

[0070] On the other hand, this embodiment also provides a visual positioning system for an automated laboratory based on the YOLO algorithm, which includes:

[0071] A primary acquisition unit, through a vision sensor, acquires the initial image data of the target area to position the field of view of the grasping robot.

[0072] A primary analysis unit, performs object recognition on the initial image data, obtains the first recognition result, and performs object grasping according to the first recognition result.

[0073] A secondary acquisition unit, when there is no object in the first recognition result, adjusts the field of view according to the first recognition result, and acquires the supplementary image data of the target area.

[0074] A secondary analysis unit, uses the supplementary image data to perform object recognition again, obtains the second recognition result, and performs object grasping according to the second recognition result.

[0075] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0076] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0077] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memories), optical fiber devices, and portable compact disc read-only memories (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0078] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0079] Embodiment 2 is an embodiment of the present invention, which provides a visual positioning method for an automated laboratory based on the YOLO algorithm. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0080] The test environment is mainly divided into three placement methods: placing the sampling objects in a single column, neatly arranging multiple columns, and arranging multiple columns randomly. These three environments simulate different scenarios that a grasping robot may face in reality. Under each environment, the effect of using the YOLO algorithm of the present invention for target recognition is tested respectively to verify its accuracy and efficiency.

[0081] Experimental process:

[0082] Placing the sampling objects in a single column: In this environment, all target objects are neatly arranged at a certain interval, and the space between the targets is large. This environment can minimize the interference of obstacles and facilitate recognition and grasping.

[0083] Neatly arranging multiple columns: The target objects are neatly arranged in several columns with appropriate spacing. This situation tests the recognition ability of the system when the target objects are slightly denser.

[0084] Randomly arranging multiple columns: The target objects are randomly arranged, the intervals between the targets are inconsistent, and some targets may be partially blocked or interlaced. This environment tests the performance of the system in dealing with complex backgrounds and overlapping objects.

[0085] Experimental data:

[0086] In the experiment, 100 target objects were used for testing, and 10 independent tests were carried out under each environment, recording the accuracy of target recognition and the grasping time.

[0087] Visiting the sampling objects in a single column: In this environment, due to the large distance between the target objects and no occlusion, the recognition accuracy of the YOLO algorithm reaches 98%. The average grasping time is 2.5 seconds, with a high accuracy and almost no misrecognition and incorrect grasping. This shows that in such an environment, the algorithm can efficiently and accurately perform target recognition and grasping, fully demonstrating its advantages in an ideal environment.

[0088] Neatly arranged in multiple columns: In this environment, the distance between the target objects is slightly smaller, and the recognition accuracy of the YOLO algorithm decreases to 95%. However, the grasping time only increases by 0.5 seconds, with an average of 3.0 seconds, a misrecognition rate of 3%, and an incorrect grasping rate of 2%. These results indicate that in the case of moderate distance between the target objects and no complex occlusion, the algorithm can still maintain a high accuracy and grasping efficiency, demonstrating its stability in a relatively complex environment.

[0089] Messily arranged in multiple columns: In this environment, the occlusion and overlap of the target objects increase significantly, resulting in a decrease in the recognition accuracy to 88%. The grasping time increases to 4.2 seconds, and the misrecognition rate and incorrect grasping rate are 8.5% and 5% respectively. Although the performance of the algorithm decreases in a complex environment, it can still maintain a relatively reasonable performance, which proves the strong adaptability of the present invention in dealing with complex scenarios.

[0090] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A visual positioning method for an automated laboratory based on the YOLO algorithm, characterized in that: include: Through the visual sensor, the initial image data of the target area is collected to locate the field of vision of the grasping robot; Performing target recognition on the initial image data to obtain a first recognition result, and capturing the target according to the first recognition result; When the target object does not exist in the first recognition result, adjusting the field of view according to the first recognition result, and collecting supplementary image data of the target area; Using the supplementary image data, target recognition is performed again to obtain a second recognition result, and target capture is performed based on the second recognition result; The target recognition includes using a trained YOLO algorithm to recognize local field of view information; Input the captured target and the first frame to the YOLO algorithm to obtain the first recognition result; The first recognition result includes the confidence level z of the i-th target object in the image being recognized as the grasping target. i ; Presets Two thresholds p1 and p2. If p1 < z i , then it is determined that the i-th target object is recognized as a grasping target; if p2 ≤ z i ≤ p1, then it is determined that the i-th target object is recognized as an object to be confirmed; if z i < p2, then it is determined that the i-th target object is recognized as a non-graspable target; When the number of grab targets described in the recognition result reaches the requirement, the grabbing is performed directly; When the number of grasping targets in the recognition result does not meet the requirement, adjusting the field of view of the grasping robot; The second recognition result includes inputting the supplementary image data and the first frame to the YOLO algorithm to obtain an output result for each of the target objects, and comparing the output result with the captured target to obtain a final recognition result; When performing target recognition again, let the input image sequence F = [f1, f2, ..., f n ], select m images from F to form M = [M1, M2, ..., M m ]; Use the YOLO algorithm to identify the classification of the i-th target object in M1. If the classification result is an ungraspable target, then in F, mark the i-th target object as an ungraspable target, and stop the recognition of the i-th target object during the recognition process; if the classification result is a graspable target, then in F, mark the i-th target object as a graspable target, and stop the recognition of the i-th target object during the recognition process; if the classification result is an object to be confirmed, continue to recognize the i-th target object; After completing the identification of M1, identify M2, and so on, until M m Complete identification; After completing m times of recognition, the images that have not been recognized only include the objects to be recognized. For the jth object to be recognized, obtain the confidence of each image in M ​​about the grasped target; obtain k images whose confidence is higher than g times the average value, and set any image as M h , then M h-1 and M h+1 are the two ends of the collection interval; In F, put M h-1 and M h+1 The corresponding image is h-1 and M h+1 The corresponding f e and f u ; Using f e and f u The image sequence between the two images is reselected to perform recognition until the maximum number of iterations is reached or the confidence level is no longer improved, and the recognition results of all images marked as grasped targets are output; Among them, M m represents the mth image in image M; f n represents the nth image in image F; h represents the index of the image in M; e and u represent the indexes of different images in F.

2. The visual positioning method for an automated laboratory based on the YOLO algorithm as claimed in claim 1, characterized in that: The visual sensors include but are not limited to: a first type of visual sensor configured in the environment to obtain macroscopic visual field information; a second type of visual sensor configured on a grasping robot to obtain local visual field information; The macroscopic field of view information is an image at the target object storage space level, including the position information of the grasping robot and the target area where the target object is located; The local field of view information is an image at the target area level, including the target object in the target area.

3. The visual positioning method of the automated laboratory based on the YOLO algorithm as claimed in claim 2, characterized in that: The field of view positioning includes constructing an environment map about the current environment, and positioning the grasping robot and the target area on the environment map according to the macroscopic field of view information; After completing the positioning, the current local field of view information is analyzed; if the local field of view information contains a complete target area, the first frame to be identified is directly acquired; If the target area is larger than the local field of view information, the position of the grasping robot in the map is marked, and then the first frame of the image to be identified is collected.

4. The visual positioning method of the automated laboratory based on the YOLO algorithm as claimed in claim 3, characterized in that: The field of view adjustment includes, if the local field of view information includes a complete target area, controlling the grasping robot to move and acquiring the supplementary image data; If the target area is larger than the local field of view information, the grasping robot moves from the current position to any edge on both sides, and collects images of the target area in real time as the supplementary image data; after the first movement, when the number of grasped targets in the recognition result of the YOLO algorithm does not meet the requirement, it returns to the marked position and moves to the other edge, and performs recognition again after moving.

5. The visual positioning method for an automated laboratory based on the YOLO algorithm as claimed in claim 4, characterized in that: The target grasping includes transmitting the position information of the target object to the grasping robot control system, the control system adjusts the motion trajectory of the robot arm according to the position information, and drives the grasping actuator to perform the grasping operation.

6. A visual positioning system for an automated laboratory based on the YOLO algorithm using the method as claimed in any one of claims 1 to 5, characterized in that: The primary acquisition unit collects the initial image data of the target area through the visual sensor and locates the field of view of the grasping robot; A primary analysis unit performs target recognition on the initial image data to obtain a first recognition result, and performs target capture according to the first recognition result; A secondary acquisition unit, when the target object does not exist in the first recognition result, adjusts the field of view according to the first recognition result and acquires supplementary image data of the target area; The secondary analysis unit uses the supplementary image data to perform target recognition again to obtain a second recognition result, and then captures the target based on the second recognition result.

7. A computer device comprising: Memory and processor; The memory stores a computer program, wherein the processor implements the steps of any one of the methods of claims 1-5 when executing the computer program.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Visual recognition and positioning method for robot intelligent capture application

    CN108171748A

  • Robot 3D visual guidance grabbing method

    CN119795178A