Automatic laboratory visual positioning method and system based on YOLO algorithm

By applying the YOLO algorithm in an automated laboratory for visual positioning, combined with the methods of visual field adjustment and iterative optimization, the problem of low visual positioning accuracy in the existing technology is solved, and efficient target recognition and capture in complex scenarios is achieved.

CN119991784AActive Publication Date: 2025-05-13STATE GRID FUJIAN ELECTRIC POWER RES INST +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510483993.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing visual positioning methods have low accuracy in scenes where multiple targets are stored in a messy manner, and cannot effectively identify and locate target objects.

Method used

The visual positioning method of an automated laboratory based on the YOLO algorithm is adopted to collect initial image data through vision sensors, and target recognition is used to use the trained YOLO algorithm, and the field of view is adjusted and supplementary image data acquisition is performed when the recognition results are insufficient, so as to iteratively optimize the recognition results.

Benefits of technology

It improves the accuracy of target recognition and ensures that the robot can complete the target recognition and grasping tasks when the target is partially blocked or the field of view is limited, which significantly improves the work efficiency and adaptability of the grab robot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991784A_ABST
    Figure CN119991784A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image analysis of neural networks, and discloses a YOLO algorithm-based visual positioning method and system for an automatic laboratory, and the method comprises the steps: firstly, carrying out the precise environment positioning through macroscopic and local view information, and improving the target recognition accuracy; and secondly, through dynamic view adjustment and supplementary image data acquisition, it is ensured that the robot can still complete target recognition and grabbing tasks even under the condition that the target is partially shielded or the view is limited. The YOLO algorithm is adopted for target recognition, the recognition speed and precision are effectively improved, and the recognition result is optimized through multiple iterations, so that the success rate of the grabbing task is ensured. On the whole, the working efficiency and adaptability of the grabbing robot can be remarkably improved, the influence caused by environment interference is reduced, and the high practical value and the wide application prospect are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image analysis of neural networks, and in particular to a visual positioning method and system for an automated laboratory based on a YOLO algorithm. Background Art

[0002] With the continuous development of automation technology, gripping equipment is increasingly used in industrial production, warehousing logistics, distribution centers and other fields. Especially in intelligent manufacturing and material handling, gripping robots can efficiently perform repetitive tasks and significantly improve work efficiency. However, the success of robot gripping tasks is usually affected by the positioning accuracy and gripping accuracy of the target object.

[0003] In order to achieve efficient automated grasping, the robot needs to accurately identify the position and posture of the target object, which requires the system to accurately obtain and analyze environmental information. Traditional grasping robot positioning methods usually rely on simple sensors or static models, which are difficult to cope with complex environments and dynamic changes. Especially in complex scenes, such as when the target object is blocked or the ambient light changes, the challenge of visual positioning is more prominent.

[0004] Therefore, how to use visual sensors and intelligent algorithms to achieve accurate grasping robot positioning and improve target recognition and grasping efficiency has become a hot topic in current research. Solving this problem can not only improve the stability and accuracy of robot grasping, but also reduce manual intervention and improve the intelligence level of the automation system. Summary of the invention

[0005] In view of the above-mentioned problems, the present invention is proposed.

[0006] Therefore, the technical problem solved by the present invention is that the existing visual positioning method has low accuracy and cannot perform recognition and positioning in scenes where multiple targets are cluttered.

[0007] In order to solve the above technical problems, the present invention provides the following technical solutions: a visual positioning method for an automated laboratory based on the YOLO algorithm, comprising:

[0008] Through the visual sensor, the initial image data of the target area is collected to locate the field of vision of the grasping robot;

[0009] Performing target recognition on the initial image data to obtain a first recognition result, and capturing the target according to the first recognition result;

[0010] When the target object does not exist in the first recognition result, adjusting the field of view according to the first recognition result, and collecting supplementary image data of the target area;

[0011] The target is recognized again using the supplementary image data to obtain a second recognition result, and the target is captured based on the second recognition result.

[0012] As a preferred solution of the visual positioning method of the automated laboratory based on the YOLO algorithm described in the present invention, the visual sensor includes but is not limited to a first type of visual sensor configured in the environment to obtain macroscopic visual field information; a second type of visual sensor configured on a grasping robot to obtain local visual field information;

[0013] The macroscopic field of view information is an image at the target object storage space level, including the position information of the grasping robot and the target area where the target object is located;

[0014] The local field of view information is an image at the target area level, including the target object in the target area.

[0015] As a preferred solution of the visual positioning method of the automated laboratory based on the YOLO algorithm described in the present invention, wherein: the field of view positioning includes constructing an environmental map of the current environment, and positioning the grasping robot and the target area on the environmental map according to the macroscopic field of view information;

[0016] After completing the positioning, the current local field of view information is analyzed; if the local field of view information contains a complete target area, the first frame to be identified is directly acquired;

[0017] If the target area is larger than the local field of view information, the position of the grasping robot in the map is marked, and then the first frame of the image to be identified is collected.

[0018] As a preferred solution of the visual positioning method of the automated laboratory based on the YOLO algorithm described in the present invention, wherein: the target recognition includes using the trained YOLO algorithm to recognize the local field of view information;

[0019] Inputting the captured target and the first frame into the YOLO algorithm to obtain a first recognition result;

[0020] The first recognition result includes the confidence level of the i-th target object in the image being recognized as the grasping target. ; Preset two thresholds and ,like , then the i-th target object is identified as the grasping target; if , then the i-th target object is identified as the object to be confirmed; if , then the i-th target object is identified as an ungraspable target;

[0021] When the number of grasping targets in the recognition result meets the requirement, grasping is directly performed; when the number of grasping targets in the recognition result does not meet the requirement, the field of view of the grasping robot is adjusted.

[0022] As a preferred solution of the visual positioning method of the automated laboratory based on the YOLO algorithm described in the present invention, wherein: the field of view adjustment includes, if the local field of view information contains a complete target area, controlling the grasping robot to move to obtain the supplementary image data;

[0023] If the target area is larger than the local field of view information, the grasping robot moves from the current position to any edge on both sides, and collects images of the target area in real time as the supplementary image data; after the first movement, when the number of grasped targets in the recognition result of the YOLO algorithm does not meet the requirement, it returns to the marked position and moves to the other edge, and performs recognition again after moving.

[0024] As a preferred solution of the visual positioning method of the automated laboratory based on the YOLO algorithm of the present invention, wherein: the second recognition result includes: inputting the first frame of the supplementary image data into the YOLO algorithm to obtain an output result for each of the target objects, and comparing the output result with the grasped target to obtain a final recognition result;

[0025] When performing target recognition again, let the input image sequence be , select m images from F to form ;

[0026] Using YOLO algorithm to identify The classification of the i-th target object in , if the classification result is an ungraspable target, then In the process, the i-th target object is marked as an ungraspable target, and the recognition of the i-th target object is stopped during the recognition process; if the classification result is a graspable target, then In the process, the i-th target object is marked as a grasping target, and the recognition of the i-th target object is stopped during the recognition process; if the classification result is an object to be confirmed, the recognition of the i-th target object continues;

[0027] Complete the pair After identification, Identify, and so on, until Complete identification;

[0028] After completing m times of recognition, the images that have not been recognized only include the objects to be confirmed. For the jth object to be confirmed, obtain The confidence of each image in the grasped target; obtain k images with confidence higher than g times the average value, and set any image as ,but and are the two ends of the collection interval;

[0029] In F and The corresponding image is and Corresponding and ;

[0030] use and The image sequence between the two images is reselected to perform recognition until the maximum number of iterations is reached or the confidence level is no longer improved, and the recognition results of all images marked as grasped targets are output;

[0031] in, Representing images , the mth image; represents the nth image in image F; h represents the index of the image in M; e and u represent the indexes of different images in F.

[0032] As a preferred solution of the visual positioning method of the automated laboratory based on the YOLO algorithm described in the present invention, the target grasping includes transmitting the posture information of the target object to the grasping robot control system, the control system adjusts the motion trajectory of the robot arm according to the posture information, and drives the grasping actuator to perform the grasping operation.

[0033] A visual positioning system for an automated laboratory based on the YOLO algorithm using the method of the present invention, wherein:

[0034] The primary acquisition unit collects the initial image data of the target area through the visual sensor and locates the field of view of the grasping robot;

[0035] A primary analysis unit performs target recognition on the initial image data to obtain a first recognition result, and performs target capture according to the first recognition result;

[0036] A secondary acquisition unit, when the target object does not exist in the first recognition result, adjusts the field of view according to the first recognition result and acquires supplementary image data of the target area;

[0037] The secondary analysis unit uses the supplementary image data to perform target recognition again to obtain a second recognition result, and then captures the target based on the second recognition result.

[0038] A computer device comprises: a memory and a processor; the memory stores a computer program, wherein: the processor implements the steps of any one of the methods of the present invention when executing the computer program.

[0039] A computer-readable storage medium stores a computer program, wherein: when the computer program is executed by a processor, the steps of any one of the methods of the present invention are implemented.

[0040] Beneficial effects of the present invention: The visual positioning method for automated laboratories based on the YOLO algorithm provided by the present invention improves the accuracy of target recognition and ensures that the robot can still complete the target recognition and grasping tasks even when the target is partially blocked or the field of view is limited. The use of the YOLO algorithm for target recognition effectively improves the recognition speed and accuracy, and optimizes the recognition results through multiple iterations, thereby ensuring the success rate of the grasping task. Overall, the present invention can significantly improve the work efficiency and adaptability of the grasping robot, reduce the impact of environmental interference, and has high practical value and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0042] Figure 1 An overall flow chart of a visual positioning method for an automated laboratory based on the YOLO algorithm provided in the first embodiment of the present invention. DETAILED DESCRIPTION

[0043] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0044] Example 1, reference Figure 1 , is an embodiment of the present invention, and provides a visual positioning method for an automated laboratory based on a YOLO algorithm, comprising:

[0045] S1: Collect the initial image data of the target area through the visual sensor and locate the field of view of the grasping robot.

[0046] Furthermore, the first type of visual sensor is configured in the environment to obtain macroscopic visual field information; the second type of visual sensor is configured on the grasping robot to obtain local visual field information. The macroscopic visual field information is an image at the target object storage space level, including the position information of the grasping robot and the target area where the target object is located. The local visual field information is an image at the target area level, including the target object in the target area.

[0047] It should be said that the first type of visual sensor configured in the environment is used to obtain images of the entire target area, including the area where the target object is located and the position information of the grasping robot. This provides the system with global information of the environment, helps the robot understand the approximate location of the target area, and ensures that the target area and the robot's current spatial position can be quickly determined when locating the target. The second type of visual sensor configured on the grasping robot is used to obtain a more detailed local image of the target area. By obtaining detailed images in the target area, the robot can more accurately identify and locate the target object. This local field of view information helps the robot to accurately identify the target, especially when the target object may be blocked or unevenly distributed. Through this hierarchical and complementary visual information acquisition method, the system can obtain the target position in a larger range, and perform accurate target identification and grasping in a local area, greatly improving the grasping accuracy and efficiency.

[0048] Furthermore, an environmental map of the current environment is constructed, and the grasping robot and the target area are positioned on the environmental map according to the macroscopic field of view information. After positioning, the current local field of view information is analyzed; if the local field of view information contains a complete target area, the first frame to be identified is directly collected. If the target area is larger than the local field of view information, the position of the grasping robot in the map is marked, and the first frame to be identified is collected.

[0049] By using macroscopic field of view information to construct an environmental map and calibrate the position of the grasping robot and the target area, the main purpose is to clarify the edge position of the field of view and the target area. This can help the system accurately determine the spatial range of the target area and provide accurate reference coordinates for subsequent local field of view acquisition and target recognition. Once the environmental positioning is completed, the system determines the range of the target area by analyzing the local field of view information. If the local field of view information is sufficient to cover the entire target area, the system can directly capture the first frame to be identified for target recognition. This method reduces unnecessary calculations and improves processing efficiency. If the local field of view information cannot cover the entire target area, the system will mark the position of the grasping robot in the map according to the current positioning, and then collect a new first frame. In this way, the system can obtain more target area information through the perspective of the mobile robot to ensure that the target is fully identified.

[0050] S2: Performing target recognition on the initial image data to obtain a first recognition result, and performing target capture according to the first recognition result.

[0051] Target recognition includes using a trained YOLO algorithm to recognize the local field of view information. The captured target and the first frame are input into the YOLO algorithm to obtain a first recognition result.

[0052] The first recognition result includes the confidence level of the i-th target object in the image being recognized as the grasping target. ; Preset two thresholds and ,like , then the i-th target object is identified as the grasping target; if , then the i-th target object is identified as the object to be confirmed; if , then the i-th target object is identified as an ungraspable target.

[0053] When the number of grasping targets described in the recognition result meets the requirement (the number of grasping targets in the laboratory may be multiple, that is, multiple grasping is required), the grasping is performed directly; when the number of grasping targets described in the recognition result does not meet the requirement, the field of view of the grasping robot is adjusted.

[0054] In YOLO, the algorithm divides the image into several grids, predicts the bounding box of each object in each grid, and outputs the category of each object and its confidence.

[0055] Based on the image processing results, a pose estimation algorithm is used to estimate the three-dimensional spatial position and pose of the target object relative to the grasping robot or robot arm. The pose information of the target object is transmitted to the grasping robot control system, which adjusts the motion trajectory of the robot arm according to the pose information and drives the grasping actuator to perform the grasping operation. Accurate grasping control is performed based on the pose information to ensure that the grasping tool contacts the target object and grasps it successfully.

[0056] S3: When the target object does not exist in the first recognition result, the field of view is adjusted according to the first recognition result, and supplementary image data of the target area is collected.

[0057] If the local field of view information contains a complete target area, the grasping robot is controlled to move (it can be moved around the outline of the target area or can be moved forward and backward in a straight line, depending on how it is set) to obtain the supplementary image data. If the target area is larger than the local field of view information, the grasping robot is moved from the current position to any edge on both sides (the edge is the edge of the field of view), and the image of the target area is collected in real time as the supplementary image data; after the first movement, when the number of grasped targets in the recognition result of the YOLO algorithm does not meet the requirement, it returns to the marked position and moves to the other edge, and performs recognition again after moving.

[0058] It should be noted that the above steps can ensure that even if the target is not successfully identified in the first recognition result, the system can still adjust the field of view, continue to collect supplementary image data, expand the coverage of the target area, thereby improving the success rate of target recognition and finally completing the target grasping task. If the target area is larger than the current local field of view, the system will guide the grasping robot to move along the edge of the field of view to collect more image information of the target area in real time. By moving, the robot can continuously expand the field of view to ensure that the complete target area can be captured and provide more data for the YOLO algorithm to process. After adjusting the field of view and collecting supplementary image data, the robot will perform target recognition again. If the YOLO algorithm still fails to identify enough grasping targets after the first move, the robot will return to its original position and move to the edge of the other side to continue image collection and recognition until the demand is met. This process is iteratively optimized to ensure that the number of grasping targets can meet the demand and avoid missing any potential grasping targets. At the same time, by directly moving and homing, repeated detection and waste of computing power can be avoided.

[0059] S4: using the supplementary image data, performing target recognition again to obtain a second recognition result, and capturing the target based on the second recognition result.

[0060] Input the first frame of the supplementary image data into the YOLO algorithm to obtain an output result for each target object, and compare the output result with the captured target to obtain a final recognition result. Specifically including:

[0061] When performing target recognition again, let the input image sequence be , select m images from F to form .

[0062] Using YOLO algorithm to identify The classification of the i-th target object in , if the classification result is an ungraspable target, then In the process, the i-th target object is marked as an ungraspable target, and the recognition of the i-th target object is stopped during the recognition process; if the classification result is a graspable target, then In the process, the i-th target object is marked as a grasping target, and the recognition of the i-th target object is stopped during the recognition process; if the classification result is an object to be confirmed, the recognition of the i-th target object continues.

[0063] Complete the pair After identification, Identify, and so on, until Complete identification.

[0064] After completing m times of recognition, the images that have not been recognized only include the objects to be confirmed. For the jth object to be confirmed, obtain The confidence of each image in the grasped target; obtain k images with confidence higher than g times the average value, and set any image as ,but and The two ends of the collection interval.

[0065] In F and The corresponding image is and Corresponding and ;

[0066] use and The image sequence between m images is reselected for recognition until the maximum number of iterations is reached or the confidence level is no longer improved, and the recognition results of all images marked as grasped targets are output.

[0067] in, Representing images , the mth image; represents the nth image in image F; h represents the index of the image in M; e and u represent the indexes of different images in F.

[0068] It should be noted that if there are multiple target areas, after the first target area is identified, the second target area is identified until the target is captured.

[0069] After collecting the supplementary image data, the system uses the YOLO algorithm to perform target recognition again in order to obtain more visual information about the target. In this process, the YOLO algorithm will classify and judge each target object, and determine whether it is a grasping target according to the preset classification rules. If it is recognized as a grasping target, it is immediately marked as a grasping target and further recognition of the target is stopped. If it is recognized as an ungraspable target, it is marked as an ungraspable target and excluded as a grasping target. For those targets marked as pending objects, the system will continue to perform subsequent recognition until its grasping status is finally confirmed. Through iterative optimization, the system can gradually exclude targets that do not meet the grasping conditions, while continuing to identify the pending objects, and ultimately ensure that all eligible targets are accurately identified. After each round of target recognition, the system will calculate the confidence of each target and filter out the images that are most likely to be grasping targets based on the confidence. Among all images, images with higher confidence will be further processed as priority objects. Through multiple iterations, the system can continuously improve the recognition accuracy and ensure that the recognized grasping targets have high accuracy. Finally, after multiple recognitions and iterations, the system will output all recognition results marked as grasping targets, ensuring the success rate of the grasping operation and reducing misidentification or missed recognition.

[0070] On the other hand, this embodiment also provides a visual positioning system for an automated laboratory based on the YOLO algorithm, which includes:

[0071] The primary acquisition unit collects the initial image data of the target area through the visual sensor and locates the field of view of the grasping robot.

[0072] The primary analysis unit performs target recognition on the initial image data to obtain a first recognition result, and captures the target according to the first recognition result.

[0073] The secondary acquisition unit adjusts the field of view according to the first recognition result and acquires supplementary image data of the target area when the first recognition result does not show the target object.

[0074] The secondary analysis unit uses the supplementary image data to perform target recognition again to obtain a second recognition result, and then captures the target based on the second recognition result.

[0075] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program code.

[0076] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0077] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0078] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0079] Example 2 is an embodiment of the present invention, which provides a visual positioning method for an automated laboratory based on the YOLO algorithm. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0080] The test environment is mainly divided into three placement modes: placing the sampled objects in one row, placing them in multiple rows neatly, and placing them in multiple rows in a disorderly manner. These three environments simulate different scenarios that a grasping robot may face in reality. In each environment, the effect of the target recognition using the YOLO algorithm in the present invention was tested to verify its accuracy and efficiency.

[0081] Experimental process:

[0082] Arrange the sampled objects in a row: In this environment, all the target objects are neatly arranged at a certain interval, and the space between the objects is large. This environment can minimize the interference of obstacles and facilitate identification and grasping.

[0083] Neatly arranged in multiple columns: The target objects are neatly arranged in several columns with moderate spacing. This situation tests the system's recognition ability when the target objects are slightly dense.

[0084] Cluttered Multi-column: Target objects are randomly placed, the spacing between targets is inconsistent, and some targets may be partially occluded or interlaced. This environment tests the system's performance in dealing with complex backgrounds and overlapping objects.

[0085] Experimental data:

[0086] In the experiment, 100 target objects were used for testing. Ten independent tests were conducted in each environment, and the target recognition accuracy and grasping time were recorded.

[0087] The sampled objects are visited in one column: In this environment, due to the large distance between the target objects and no occlusion, the recognition accuracy of the YOLO algorithm reaches 98%. The average grasping time is 2.5 seconds, with a high accuracy rate and almost no misidentification and wrong grasping. This shows that in this environment, the algorithm can efficiently and accurately recognize and grasp the target, fully demonstrating its advantages in a relatively ideal environment.

[0088] Arrange multiple columns neatly: In this environment, the distance between the target objects is slightly smaller, and the recognition accuracy of the YOLO algorithm has dropped to 95%. However, the grasping time has only increased by 0.5 seconds, averaging 3.0 seconds, with a false recognition rate of 3% and a false grasping rate of 2%. These results show that when the distance between the target objects is moderate and there is no complex occlusion, the algorithm can still maintain a high accuracy and grasping efficiency, showing its stability in relatively complex environments.

[0089] Multiple rows in a messy environment: In this environment, the occlusion and overlap of the target objects increased significantly, causing the recognition accuracy to drop to 88%. The capture time increased to 4.2 seconds, and the misrecognition rate and false capture rate were 8.5% and 5%, respectively. Although the algorithm's performance declined in complex environments, it still maintained a relatively reasonable performance, which proves the strong adaptability of the invention in dealing with complex scenes.

[0090] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A visual positioning method for an automated laboratory based on the YOLO algorithm, characterized in that: include: Through the visual sensor, the initial image data of the target area is collected to locate the field of vision of the grasping robot; Performing target recognition on the initial image data to obtain a first recognition result, and capturing the target according to the first recognition result; When the target object does not exist in the first recognition result, adjusting the field of view according to the first recognition result, and collecting supplementary image data of the target area; The target is recognized again using the supplementary image data to obtain a second recognition result, and the target is captured based on the second recognition result.

2. The visual positioning method for an automated laboratory based on the YOLO algorithm as claimed in claim 1, characterized in that: The visual sensors include but are not limited to: a first type of visual sensor configured in the environment to obtain macroscopic visual field information; a second type of visual sensor configured on a grasping robot to obtain local visual field information; The macroscopic field of view information is an image at the target object storage space level, including the position information of the grasping robot and the target area where the target object is located; The local field of view information is an image at the target area level, including the target object in the target area.

3. The visual positioning method of the automated laboratory based on the YOLO algorithm as claimed in claim 2, characterized in that: The field of view positioning includes constructing an environment map about the current environment, and positioning the grasping robot and the target area on the environment map according to the macroscopic field of view information; After completing the positioning, the current local field of view information is analyzed; if the local field of view information contains a complete target area, the first frame to be identified is directly acquired; If the target area is larger than the local field of view information, the position of the grasping robot in the map is marked, and then the first frame of the image to be identified is collected.

4. The visual positioning method of the automated laboratory based on the YOLO algorithm as claimed in claim 3, characterized in that: The target recognition includes using a trained YOLO algorithm to recognize the local field of view information; Inputting the captured target and the first frame into the YOLO algorithm to obtain a first recognition result; The first recognition result includes the confidence level of the i-th target object in the image being recognized as the grasping target. ; Preset two thresholds and ,like , then the i-th target object is identified as the grasping target; if , then the i-th target object is identified as the object to be confirmed; if , then the i-th target object is identified as an ungraspable target; When the number of grasping targets in the recognition result meets the requirement, grasping is directly performed; when the number of grasping targets in the recognition result does not meet the requirement, the field of view of the grasping robot is adjusted.

5. The visual positioning method of the automated laboratory based on the YOLO algorithm as claimed in claim 4, characterized in that: The field of view adjustment includes, if the local field of view information includes a complete target area, controlling the grasping robot to move and acquiring the supplementary image data; If the target area is larger than the local field of view information, the grasping robot moves from the current position to any edge on both sides, and collects images of the target area in real time as the supplementary image data; after the first movement, when the number of grasped targets in the recognition result of the YOLO algorithm does not meet the requirement, it returns to the marked position and moves to the other edge, and performs recognition again after moving.

6. The visual positioning method for an automated laboratory based on the YOLO algorithm as claimed in claim 5, characterized in that: The second recognition result includes inputting the first frame of the supplementary image data into the YOLO algorithm, obtaining an output result for each of the target objects, and comparing the output result with the captured target to obtain a final recognition result; When performing target recognition again, let the input image sequence be , select m images from F to form ; Using YOLO algorithm to identify The classification of the i-th target object in , if the classification result is an ungraspable target, then In the process, the i-th target object is marked as an ungraspable target, and the recognition of the i-th target object is stopped during the recognition process; if the classification result is a graspable target, then In the process, the i-th target object is marked as a grasping target, and the recognition of the i-th target object is stopped during the recognition process; if the classification result is an object to be confirmed, the recognition of the i-th target object continues; Complete the pair After identification, Identify, and so on, until Complete identification; After completing m times of recognition, the images that have not been recognized only include the objects to be confirmed. For the jth object to be confirmed, obtain The confidence of each image in the grasped target; obtain k images with confidence higher than g times the average value, and set any image as ,but and are the two ends of the collection interval; In F and The corresponding image is and Corresponding and ; use and The image sequence between the two images is reselected to perform recognition until the maximum number of iterations is reached or the confidence level is no longer improved, and the recognition results of all images marked as grasped targets are output; in, Representing images , the mth image; represents the nth image in image F; h represents the index of the image in M; e and u represent the indexes of different images in F.

7. The visual positioning method of the automated laboratory based on the YOLO algorithm as claimed in claim 6, characterized in that: The target grasping includes transmitting the position information of the target object to the grasping robot control system, the control system adjusts the motion trajectory of the robot arm according to the position information, and drives the grasping actuator to perform the grasping operation.

8. A visual positioning system for an automated laboratory based on the YOLO algorithm using the method as claimed in any one of claims 1 to 7, characterized in that: The primary acquisition unit collects the initial image data of the target area through the visual sensor and locates the field of view of the grasping robot; A primary analysis unit performs target recognition on the initial image data to obtain a first recognition result, and performs target capture according to the first recognition result; A secondary acquisition unit, when the target object does not exist in the first recognition result, adjusts the field of view according to the first recognition result and acquires supplementary image data of the target area; The secondary analysis unit uses the supplementary image data to perform target recognition again to obtain a second recognition result, and then captures the target based on the second recognition result.

9. A computer device comprising: A memory and a processor; the memory stores a computer program, wherein the processor implements the steps of any method as claimed in claim 1 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Visual recognition and positioning method for robot intelligent capture application

    CN108171748A

  • Projection interaction method based on pure machine vision positioning

    CN111354007A

  • Unmanned aerial vehicle target positioning method based on YOLO attitude estimation

    CN118189959A

  • Robot 3D visual guidance grabbing method

    CN119795178A

  • KR20240027395A