An object positioning and recognizing method, device, electronic equipment and storage medium
By integrating laser and image deep learning technologies, precise object localization and type recognition are achieved, solving the problem that existing 2D laser technology cannot identify object types and enhancing the robot's ability to react to different objects.
Patent Information
- Application Number
- CN202011588336.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-28
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2040-12-28
AI Technical Summary
Existing 2D laser technology can achieve object localization very well, but it cannot identify the type of object, which makes it impossible for robots to react differently to different objects.
By fusing laser and image deep learning results, object detection, angle calculation, and laser localization are performed to determine the position and category of objects.
It enables precise positioning and object type identification, allowing the robot to react differently to different objects and providing richer semantic information about the objects.
Smart Images

Figure CN114692712B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of object detection, and in particular to a method, apparatus, electronic device, and storage medium for object localization and recognition. Background Technology
[0002] Currently, in object localization solutions, existing 2D laser technology can effectively locate objects. However, while 2D laser technology can accurately pinpoint an object's location, it cannot identify the object's type, thus preventing the robot from responding differently to different objects. Summary of the Invention
[0003] In view of this, the present invention provides a method, apparatus, electronic device and storage medium for object localization and recognition, which can accurately locate objects and identify the corresponding object types by fusing laser and image deep learning results, thereby realizing the function of identifying object types and localization, enabling robots to make different responses to different objects.
[0004] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:
[0005] According to one aspect of the present invention, a method for object localization and recognition is provided, the method comprising:
[0006] Perform image detection on the object to be identified and obtain the object detection results;
[0007] The object angle is calculated based on the object detection results to obtain the object angle result.
[0008] Laser positioning is performed on the object to be identified to obtain the object positioning result;
[0009] Based on the object positioning results and the object angle results, the location and category of the object to be identified are determined.
[0010] In one possible design, the step of performing image detection on the object to be identified and obtaining object detection results includes:
[0011] Once the object to be identified is determined, data collection and annotation of the object are performed.
[0012] Train a deep learning model for object detection based on the collected data;
[0013] The trained object detection deep learning model is used to detect objects in images acquired in a robotic environment, and outputs the object's location and classification in the image.
[0014] In one possible design, the step of calculating the object angle based on the object detection result to obtain the object angle result includes:
[0015] The object detection results are then corrected for distortion.
[0016] The distortion-corrected object detection results are then transformed to obtain the image detection results of the object detection results in the robot coordinate system.
[0017] The angle of the object is obtained from the image detection results in the robot coordinate system based on the object detection results.
[0018] In one possible design, the distortion correction of the object detection results includes:
[0019] The camera distortion correction model is determined based on the camera's internal and external parameters.
[0020] The object detection results are then corrected using a camera distortion correction model.
[0021] In one possible design, obtaining the object's angle based on the image detection results in the robot coordinate system includes:
[0022] Obtain the vertical center points of the left and right edges of the image;
[0023] Calculate the left and right edge angles to find that the object is within the angular range between the left and right edge angles.
[0024] In one possible design, the laser positioning of the object to be identified to obtain the object positioning result includes: specifically including:
[0025] A point cloud is obtained by using a laser scanning robot to scan the current environment, and the angles and positions of all points are obtained.
[0026] By utilizing the transformation relationship between the laser coordinate system and the robot coordinate system, the position of the point scanned by the laser in the robot coordinate system is calculated, thus obtaining the object positioning result.
[0027] In one possible design, determining the location and category of the identified object based on the object's location result and its angle result includes:
[0028] Visit each point transformed into the robot coordinate system one by one. If the angle of a point is within the range of the left edge angle and the right edge angle, then retain the point.
[0029] The remaining points are processed to obtain the object's position and category.
[0030] According to another aspect of the present invention, an object positioning and recognition apparatus is provided, the apparatus comprising: a detection module, a calculation module, a positioning module, and a determination module; wherein:
[0031] The detection module is used to perform image detection on the object to be identified and obtain the object detection result;
[0032] The calculation module is used to calculate the object angle based on the object detection result and obtain the object angle result.
[0033] The positioning module is used to perform laser positioning on the object to be identified and obtain the object positioning result;
[0034] The determining module is used to determine the location and category of the object to be identified based on the object positioning result and the object angle result.
[0035] According to another aspect of the present invention, an electronic device is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the object localization and recognition method provided in the embodiments of the present invention.
[0036] According to another aspect of the present invention, a storage medium is provided, characterized in that the storage medium stores a method program for object localization and recognition, wherein when the method program for object localization and recognition is executed by a processor, it implements the steps of the method for object localization and recognition provided in the embodiments of the present invention.
[0037] Compared with related technologies, the embodiments of the present invention provide a method, apparatus, electronic device, and storage medium for object localization and recognition. The method includes: performing image detection on the object to be identified to obtain an object detection result; calculating the object angle based on the object detection result to obtain an object angle result; performing laser localization on the object to be identified to obtain an object localization result; and determining the location and category of the object to be identified based on the object localization result and the object angle result. Through the embodiments of the present invention, by performing image detection on the object to be identified to obtain an object detection result; calculating the object angle based on the object detection result to obtain an object angle result; performing laser localization on the object to be identified to obtain an object localization result; and determining the location and category of the object to be identified based on the object localization result and the object angle result, the present invention, by fusing laser and image deep learning results, can accurately locate objects and identify their corresponding object types, thereby realizing the function of identifying object types and localizing them. This enables robots to react differently to different objects and provides richer semantic information about objects compared to existing simple laser localization technologies. Attached Figure Description
[0038] Figure 1 This is a flowchart illustrating a method for object localization and recognition provided in an embodiment of the present invention.
[0039] Figure 2 This is a schematic diagram of the structure of an object positioning and recognition device provided in an embodiment of the present invention.
[0040] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0041] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0042] To make the technical problems, solutions, and beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0043] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0044] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0045] In one embodiment, such as Figure 1 As shown, the present invention provides a method for object localization and recognition, the method comprising:
[0046] S1. Perform image detection on the object to be identified and obtain the object detection result.
[0047] S2. Calculate the object angle based on the object detection results to obtain the object angle result.
[0048] S3. Perform laser positioning on the object to be identified to obtain the object positioning result.
[0049] S4. Based on the object positioning result and the object angle result, determine the location and category of the object to be identified.
[0050] In this embodiment, image detection is performed on the object to be identified to obtain object detection results; object angles are calculated based on the object detection results to obtain object angle results; laser localization is performed on the object to be identified to obtain object localization results; and the location and category of the object to be identified are determined based on the object localization results and the object angle results. Thus, by fusing laser and image deep learning results, objects can be accurately located and their corresponding types identified, thereby achieving the functions of object type recognition and localization. This allows the robot to react differently to different objects, providing richer semantic information about objects compared to existing laser-based localization technologies.
[0051] In one embodiment, step S1, which involves performing image detection on the object to be identified to obtain an object detection result, includes: performing image image detection on the object to be identified based on deep learning to obtain an object detection result. Specifically, this includes:
[0052] S11. Determine the object to be identified, and collect and label the data of the object.
[0053] S12. Train a deep learning model for object detection based on the collected data.
[0054] S13. The trained object detection deep learning model is used to perform object detection on the image acquired in the robot environment, and the object detection results, including the position and classification of the object in the image, are output.
[0055] In this embodiment, a deep learning model for object detection is first trained by collecting data. Then, the trained deep learning model is used to detect objects in images acquired in the robot environment, and the object detection results, including the position and classification of the objects in the image, are output, thereby accurately identifying the position and classification of the objects.
[0056] In one embodiment, step S2, which involves calculating the object angle based on the object detection result to obtain the object angle result, includes:
[0057] S21. The object detection results are then distorted.
[0058] When a camera performs image detection on an object to be identified, distortion can occur due to the precision and manufacturing process of the camera lens, resulting in distorted images. Therefore, a camera distortion correction model is needed to correct the distortion.
[0059] The camera distortion correction model is determined based on the camera's internal and external parameters.
[0060] The object detection results are then corrected using a camera distortion correction model.
[0061] S22. Perform coordinate transformation on the distorted object detection results to obtain the image detection results of the object detection results in the robot coordinate system.
[0062] Since the image is obtained in the camera coordinate system, it is necessary to obtain the image detection result of the object detection result in the robot coordinate system by using the transformation relationship between the robot coordinate system and the camera coordinate system.
[0063] S23. Based on the object detection results in the robot coordinate system, obtain the object's angle.
[0064] Obtain the vertical center points of the left and right edges of the image;
[0065] By calculating the left edge angle 'a' and the right edge angle 'b' using the tangent function tanh, we can determine that the object lies within the angular range from the left edge angle 'a' to the right edge angle 'b'.
[0066] In this embodiment, by performing distortion correction, coordinate transformation, and angle calculation on the object detection results, the angular range of the object can be accurately obtained.
[0067] In one embodiment, step S3, which involves performing laser localization on the object to be identified to obtain an object localization result, includes: performing laser-based object localization on the object to be identified. Specifically, this includes:
[0068] S31. Use a laser to scan the current environment of the robot to obtain a point cloud frame, and obtain the angle and position of all points.
[0069] S32. By using the transformation relationship between the laser coordinate system and the robot coordinate system, the position of the point scanned by the laser in the robot coordinate system is calculated to obtain the object positioning result.
[0070] In this embodiment, by using the current environment of the laser scanning robot and the transformation relationship between the laser coordinate system and the robot coordinate system, the position of the point scanned by the laser in the robot coordinate system is calculated, thereby accurately obtaining the object positioning result.
[0071] In one embodiment, step S4, determining the location and category of the identified object based on the object positioning result and the object angle result, includes:
[0072] Visit each point transformed into the robot coordinate system one by one. If the angle of a point is within the range of the left edge angle a and the right edge angle b, then retain the point.
[0073] The remaining points are processed to obtain the object's position and category.
[0074] In this embodiment, by fusing the object positioning results obtained from laser and the object angle results obtained from image deep learning, the object can be accurately located and the corresponding object type can be identified.
[0075] In one embodiment, such as Figure 2 As shown, the present invention provides an object positioning and recognition device, the device comprising: a detection module 10, a calculation module 20, a positioning module 30, and a determination module 40; wherein:
[0076] The detection module 10 is used to perform image detection on the object to be identified and obtain the object detection result.
[0077] The calculation module 20 is used to calculate the object angle based on the object detection result and obtain the object angle result.
[0078] The positioning module 30 is used to perform laser positioning on the object to be identified and obtain the object positioning result.
[0079] The determining module 40 is used to determine the location and category of the object to be identified based on the object positioning result and the object angle result.
[0080] In this embodiment, image detection is performed on the object to be identified to obtain object detection results; object angles are calculated based on the object detection results to obtain object angle results; laser localization is performed on the object to be identified to obtain object localization results; and the location and category of the object to be identified are determined based on the object localization results and the object angle results. Thus, by fusing laser and image deep learning results, objects can be accurately located and their corresponding types identified, thereby achieving the functions of object type recognition and localization. This allows the robot to react differently to different objects, providing richer semantic information about objects compared to existing laser-based localization technologies.
[0081] In one embodiment, the detection module 10 is specifically used for:
[0082] S11. Determine the object to be identified, and collect and label the data of the object.
[0083] S12. Train a deep learning model for object detection based on the collected data.
[0084] S13. The trained object detection deep learning model is used to perform object detection on the image acquired in the robot environment, and the object detection results, including the position and classification of the object in the image, are output.
[0085] In this embodiment, a deep learning model for object detection is first trained by collecting data. Then, the trained deep learning model is used to detect objects in images acquired in the robot environment, and the object detection results, including the position and classification of the objects in the image, are output, thereby accurately identifying the position and classification of the objects.
[0086] In one embodiment, the computing module 20 is specifically used for:
[0087] S21. The object detection results are then distorted.
[0088] When a camera performs image detection on an object to be identified, distortion can occur due to the precision and manufacturing process of the camera lens, resulting in distorted images. Therefore, a camera distortion correction model is needed to correct the distortion.
[0089] The camera distortion correction model is determined based on the camera's internal and external parameters.
[0090] The object detection results are then corrected using a camera distortion correction model.
[0091] S22. Perform coordinate transformation on the distorted object detection results to obtain the image detection results of the object detection results in the robot coordinate system.
[0092] Since the image is obtained in the camera coordinate system, it is necessary to obtain the image detection result of the object detection result in the robot coordinate system by using the transformation relationship between the robot coordinate system and the camera coordinate system.
[0093] S23. Based on the object detection results in the robot coordinate system, obtain the object's angle.
[0094] Obtain the vertical center points of the left and right edges of the image;
[0095] By calculating the left edge angle 'a' and the right edge angle 'b' using the tangent function tanh, we can determine that the object lies within the angular range from the left edge angle 'a' to the right edge angle 'b'.
[0096] In this embodiment, by performing distortion correction, coordinate transformation, and angle calculation on the object detection results, the angular range of the object can be accurately obtained.
[0097] In one embodiment, the positioning module 30 is specifically used for:
[0098] S31. Use a laser to scan the current environment of the robot to obtain a point cloud frame, and obtain the angle and position of all points.
[0099] S32. By using the transformation relationship between the laser coordinate system and the robot coordinate system, the position of the point scanned by the laser in the robot coordinate system is calculated to obtain the object positioning result.
[0100] In this embodiment, by using the current environment of the laser scanning robot and the transformation relationship between the laser coordinate system and the robot coordinate system, the position of the point scanned by the laser in the robot coordinate system is calculated, thereby accurately obtaining the object positioning result.
[0101] In one embodiment, the determining module 40 is specifically used for:
[0102] Visit each point transformed into the robot coordinate system one by one. If the angle of a point is within the range of the left edge angle a and the right edge angle b, then retain the point.
[0103] The remaining points are processed to obtain the object's position and category.
[0104] In this embodiment, by fusing the object positioning results obtained from laser and the object angle results obtained from image deep learning, the object can be accurately located and the corresponding object type can be identified.
[0105] It should be noted that the above-described device embodiments and method embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments. Furthermore, the technical features in the method embodiments are all applicable to the device embodiments, and will not be repeated here.
[0106] Furthermore, embodiments of the present invention also provide an electronic device, such as... Figure 3 As shown, it includes: a memory, a processor, and one or more computer programs stored in the memory and executable on the processor. When the one or more computer programs are executed by the processor, they implement the following steps of the object localization and recognition method provided in this embodiment of the invention:
[0107] S1. Perform image detection on the object to be identified and obtain the object detection result.
[0108] S2. Calculate the object angle based on the object detection results to obtain the object angle result.
[0109] S3. Perform laser positioning on the object to be identified to obtain the object positioning result.
[0110] S4. Based on the object positioning result and the object angle result, determine the location and category of the object to be identified.
[0111] The methods disclosed in the above embodiments of the present invention can be applied to the processor 901, or implemented by the processor 901. The processor 901 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 901 or by instructions in the form of software. The processor 901 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 901 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of the present invention can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in the memory 902. The processor 901 reads the information in the memory 902 and combines its hardware to complete the steps of the aforementioned method.
[0112] It is understood that the memory 902 in this embodiment of the invention can be a volatile memory or a non-volatile memory, or it can include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory or other memory technologies, compact disk read-only memory (CD-ROM), digital video disk (DVD) or other optical disc storage, magnetic cartridges, magnetic tapes, disk storage or other magnetic storage devices; the volatile memory can be random access memory (RAM). By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM). Memory), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memories described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable types of memories.
[0113] It should be noted that the above-described electronic device embodiments and method embodiments belong to the same concept. For details of their implementation process, please refer to the method embodiments. Furthermore, the technical features in the method embodiments are all applicable to the electronic device embodiments, and will not be repeated here.
[0114] In addition, embodiments of the present invention also provide a computer-readable storage medium storing a method program for object localization and recognition. When executed by a processor, the object localization and recognition method implements the following steps of the object localization and recognition method provided in the embodiments of the present invention:
[0115] S1. Perform image detection on the object to be identified and obtain the object detection result.
[0116] S2. Calculate the object angle based on the object detection results to obtain the object angle result.
[0117] S3. Perform laser positioning on the object to be identified to obtain the object positioning result.
[0118] S4. Based on the object positioning result and the object angle result, determine the location and category of the object to be identified.
[0119] It should be noted that the above-described embodiment of the object localization and recognition method on the computer-readable storage medium belongs to the same concept as the method embodiment. The specific implementation process is detailed in the method embodiment, and the technical features in the method embodiment are all applicable to the above-described embodiment of the computer-readable storage medium, which will not be repeated here.
[0120] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0121] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0122] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0123] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method for object localization and recognition, characterized in that, The method includes: Perform image detection on the object to be identified to obtain the detection result of the object to be identified; The process of calculating the object angle based on the object detection results to obtain the angle result of the object to be identified includes: removing distortion from the object detection results; performing coordinate transformation on the distorted object detection results to obtain the image detection result of the object detection results in the robot coordinate system; and obtaining the angle of the object based on the image detection result of the object detection results in the robot coordinate system. Laser localization of the object to be identified, and obtaining the localization result of the object to be identified, includes: obtaining a point cloud frame by scanning the current environment of the robot with a laser, and obtaining the angle and position of all points; calculating the position of the laser-scanned points in the robot coordinate system through the transformation relationship between the laser coordinate system and the robot coordinate system, and obtaining the object localization result. Based on the object localization result and the object angle result, the location and category of the object to be identified are determined, including: visiting each point transformed into the robot coordinate system one by one; if the angle of a point is within the range of the left edge angle and the right edge angle, the point is retained; all retained points are processed to obtain the object's position and object category.
2. The method as described in claim 1, characterized in that, The object to be identified is subjected to image detection to obtain object detection results; include: Once the object to be identified is determined, data collection and annotation of the object are performed. Train a deep learning model for object detection based on the collected data; The trained object detection deep learning model is used to detect objects in images acquired in a robotic environment, and outputs the object's location and classification in the image.
3. The method as described in claim 1, characterized in that, The distortion correction of the object detection results includes: The camera distortion correction model is determined based on the camera's internal and external parameters; The object detection results are then corrected using a camera distortion correction model.
4. The method as described in claim 1, characterized in that, The step of obtaining the object's angle based on the image detection results in the robot coordinate system from the object detection results includes: Obtain the vertical center points of the left and right edges of the image; Calculate the left and right edge angles to find that the object is within the angular range between the left and right edge angles.
5. An object localization and recognition device, applied to the object localization and recognition method as described in any one of claims 1 to 4, characterized in that, The device includes: a detection module, a calculation module, a positioning module, and a determination module; wherein: The detection module is used to perform image detection on the object to be identified and obtain the detection result of the object to be identified; The calculation module is used to calculate the object angle based on the object detection result to obtain the angle result of the object to be identified. The positioning module is used to perform laser positioning on the object to be identified and obtain the positioning result of the object to be identified. The determining module is used to determine the location and category of the object to be identified based on the object's location result and its angle result.
6. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of a method for object localization and recognition as described in any one of claims 1 to 4.
7. A storage medium, characterized in that, The storage medium stores a method program for object localization and recognition, which, when executed by a processor, implements the steps of the object localization and recognition method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Autonomous acquisition method and system of substation inspection robot
CN110614638A
Multi-sensor fusion and personnel positioning method based on image processing
CN110766170A