Garbage rapid identification method based on unmanned environmental sanitation vehicle-mounted around-view camera
The bird's-eye viewing image is generated through surround-view video stitching and OpenGL rendering technology, and garbage recognition is combined with the Lang-Segment-Anything depth model, which solves the problem of computing resource limitation of unmanned sanitation vehicles and realizes lightweight and highly real-time garbage recognition.
Patent Information
- Application Number
- CN202510485442.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Unmanned sanitation vehicles have problems such as computing resource limitations, single-camera field of view and target object occlusion in garbage identification, making it difficult to achieve lightweight, good generalization and strong real-time garbage recognition algorithms.
The bird's-eye view image is generated through the surround view video stitching technology, using OpenGL for rapid rendering, and combining the Lang-Segment-Anything depth model for garbage instance segmentation to reduce the memory usage and computing resource requirements.
It realizes more efficient garbage recognition performance, reduces dependence on training data, meets real-time requirements, and is suitable for matching with map information.
Smart Images

Figure CN120260013A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and particularly relates to a method for rapid garbage recognition based on an omnidirectional camera of an unmanned sanitation vehicle. Background Art
[0002] In the construction of urban intelligentization, unmanned sanitation intelligent vehicles are an essential and important part. In the operation scenarios of open roads, the perception of high-dynamic environmental elements has high requirements for unmanned sanitation. As one of the core services of unmanned sanitation vehicles, garbage inspection makes the garbage recognition technology play a key role in unmanned sanitation, which is of great significance for improving operation quality, energy conservation and emission reduction, etc.
[0003] At present, the garbage image recognition technology implemented in unmanned sanitation vehicles is mainly based on monocular camera images. For mobile devices, the target detection of a single camera has problems such as a single field of view, target objects crossing the camera, occlusion, etc., and it is difficult to fuse with map information to achieve better scene understanding. Currently, unmanned sanitation vehicles are basically equipped with fish-eye cameras in multiple directions. The bird's-eye view target recognition algorithm based on multi-camera fusion (including garbage target objects) can achieve better recognition effects, but its algorithm is often trained based on deep models and has high requirements for computing resources such as graphics cards and CPUs. However, the on-vehicle environment is restricted by computing resources, and it is difficult for the bird's-eye view garbage detection algorithm to be directly implemented in industrial scenarios. How to implement a more lightweight, generalization-friendly, and real-time garbage recognition algorithm for industrial scenarios is an urgent problem to be solved at present. Summary of the Invention
[0004] Aiming at the above deficiencies in the prior art, the present invention provides a method for rapid garbage recognition based on an omnidirectional camera of an unmanned sanitation vehicle, which solves the problems of large computing requirements for garbage recognition and high dependence on training data by using omnidirectional video stitching technology and a network model for garbage recognition in bird's-eye view images.
[0005] To achieve the above invention purpose, the technical solution adopted by the present invention is: A method for rapid garbage recognition based on an omnidirectional camera of an unmanned sanitation vehicle, comprising the following steps: S1: Obtain original images with synchronized timestamps according to multiple unmanned vehicle fish-eye cameras; S2: Process the original images with synchronized timestamps using calibrated internal and external parameters to obtain an internal parameter mapping table and an external parameter mapping table; S3: According to the original images with synchronized timestamps, manually select overlapping regions, and perform pixel-level hierarchical fusion on the overlapping regions to generate a mask with the weight of the fusion region; S4: According to the internal parameter mapping table, the external parameter mapping table, and the mask of the fusion region weights, use OpenGL for fast rendering to obtain a bird's-eye view image; S5: According to the bird's-eye view image, use the Lang-Segment-Anything depth model to perform garbage instance segmentation to obtain a segmentation mask of the garbage; S6: According to the segmentation mask of the garbage, obtain the garbage coordinates in the bird's-eye view image, and perform conversion according to the calibrated internal and external parameters to obtain the garbage recognition coordinate values in the vehicle body coordinate system.
[0006] The beneficial effects of the present invention are as follows: The present invention generates a bird's-eye view image through the panoramic video stitching technology, solves the problems of limited field of view of a single camera, objects spanning cameras and occlusion, improves the performance of the garbage target recognition algorithm for unmanned sanitation vehicles, and is suitable for subsequent matching with map information; The present invention performs garbage detection on the bird's-eye view image through the trained Lang-Segment-Anything depth model, has lower video memory occupancy, and based on the OpenGL technology, realizes the rapid generation of the bird's-eye view image, meeting the real-time requirements; The present invention, through the trained Lang-Segment-Anything depth model, does not require the construction of relevant data sets and subsequent model training, reducing the dependence of the garbage rapid recognition algorithm model training on data.
[0007] Furthermore: The specific steps of S2 are as follows: S201: According to the calibrated internal parameters, use the InitUndistortRectifyMap image distortion correction function for processing to obtain an internal parameter mapping table; S202: According to the internal parameter mapping table, perform pixel mapping on the original images with synchronized timestamps to obtain corrected images; S203: According to the corrected images, perform region cropping to obtain cropped images; S204: According to the calibrated external parameters, project the cropped images onto the vehicle body bird's-eye coordinate system to obtain an external parameter mapping table.
[0008] The beneficial effects of the above further solution are as follows: By processing the original images synchronously collected by the fisheye camera with internal and external parameters, an external parameter mapping table and an internal parameter mapping table are obtained, facilitating the subsequent generation of the bird's-eye view image.
[0009] Furthermore: The specific steps of S3 are as follows: S301: According to the original images with synchronized timestamps, obtain the overlapping regions between the respective original images by manually selecting the overlapping regions; S302: Define the basic variables for fusion and the weight function for fusion based on the original images synchronized by the timestamps and the selected overlapping regions. S303: Based on the basic variables for fusion and the weight function for fusion, calculate the distance of each point from the boundary of the fusion region pixel by pixel to generate the fusion weights of each point for the two overlapping images, and obtain the mask of the fusion region weights.
[0010] The beneficial effect of the above further solution is that by determining the mask of the fusion region weights, data can be provided for the generation of the bird's-eye view image, reducing the video memory occupancy.
[0011] Furthermore: The specific steps of S4 are as follows: S401: Normalize the coordinate system of the bird's-eye view image to obtain the normalized coordinate points. S402: Use the normalized coordinate points as vertex data and input them into the vertex shader of OpenGL for processing to obtain the vertex coordinates and input them into the fragment shader of OpenGL. S403: According to the external parameter mapping table, use the fragment shader of OpenGL to perform an inverse perspective transformation on the vertex coordinates, and according to the internal parameter mapping table, perform an inverse correction transformation on the vertex coordinates to obtain the texture coordinates. S404: According to the mask of the fusion region weights, use the fragment shader of OpenGL to perform weighted rendering of the fusion region on the texture coordinates to obtain the rendered coordinates. S405: Use the vertex coordinates for mapping and use the rendered coordinates for rendering to obtain the bird's-eye view image.
[0012] The beneficial effect of the above further solution is that by rendering with OpenGL to generate the bird's-eye view image, it has lower video memory occupancy during garbage detection, and OpenGL can achieve the rapid generation of the bird's-eye view image, meeting the real-time requirements for rapid garbage recognition.
[0013] Furthermore: The specific implementation method of S5 is as follows: Input the prompt words related to garbage into the Lang-Segment-Anything deep model, and use the Lang-Segment-Anything deep model to perform garbage instance segmentation on the bird's-eye view image to obtain the segmentation mask of the garbage; the Lang-Segment-Anything deep model includes: the Segment-Anything image segmentation model and the GroundingDINO object detection model.
[0014] The beneficial effects of the above further solution are as follows: By using the Lang-Segment-Anything deep segmentation model for garbage recognition, the publicly available model can be directly used without the need for model construction and training, reducing the dependence of garbage recognition on training data and simultaneously reducing the demand for computing resources of the garbage rapid recognition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a flowchart of a method for rapid garbage recognition based on an omnidirectional camera of an unmanned sanitation vehicle; Figure 2 It is a diagram of the form of fused basic variables; Figure 3 It is a flowchart of OpenGL rendering. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.
[0017] As Figure 1 shown, a flowchart of a method for rapid garbage recognition based on an omnidirectional camera of an unmanned sanitation vehicle includes the following steps: S1: Obtain original images with synchronized timestamps according to multiple unmanned vehicle fisheye cameras; S2: Process the original images with synchronized timestamps using calibrated internal and external parameters to obtain an internal parameter mapping table and an external parameter mapping table; wherein, the calibrated internal and external parameters are the calibrated internal and external parameters of the unmanned vehicle fisheye camera; S3: Manually select the overlapping area according to the original images with synchronized timestamps, and perform pixel-level hierarchical fusion on the overlapping area to generate a mask with the weight of the fusion area; S4: Use OpenGL for rapid rendering according to the internal parameter mapping table, the external parameter mapping table, and the mask with the weight of the fusion area to obtain a bird's-eye view image; S5: Use the Lang-Segment-Anything deep model to perform garbage instance segmentation on the bird's-eye view image to obtain a segmentation mask of the garbage; S6: Obtain the garbage coordinates in the bird's-eye view image according to the segmentation mask of the garbage, and perform conversion according to the calibrated internal and external parameters to obtain the garbage recognition coordinate values in the vehicle body coordinate system.
[0018] In one embodiment of the present invention, the specific steps of S2 are as follows: S201: According to the calibrated internal parameters, use the InitUndistortRectifyMap image distortion correction function to process and obtain the mapping table from the original camera pixel points to the corrected image pixel points, that is, the internal parameter mapping table; S202: According to the internal parameter mapping table, perform pixel mapping on the timestamp-synchronized original image to obtain the corrected image; S203: According to the corrected image, perform regional cropping to remove the unnecessary edge parts and obtain the cropped image; S204: According to the calibrated external parameters, project the cropped image onto the vehicle body bird's-eye coordinate system to obtain the mapping table from the corrected image pixel points to the vehicle body bird's-eye image pixel points, that is, the external parameter mapping table, where the vehicle body bird's-eye coordinate system is defined according to the usage requirements to determine the display range and resolution of the bird's-eye view image.
[0019] In one embodiment of the present invention, the specific steps of S3 are as follows: S301: According to the timestamp-synchronized original images, manually select the overlapping areas to obtain the overlapping areas between the original images; S302: According to the timestamp-synchronized original images and the selected overlapping areas, define the basic variables for fusion and the weight function for fusion; where, as Figure 2 shown, it is the form diagram of the basic variables for fusion. The basic variables for fusion are the distances of the pixel points from the boundaries of the overlapping areas, that is, Figure 2 the virtual line segments D1 and D2 in, area1 and area2 are the non-overlapping areas of the images, and various methods can be selected to calculate the distance. For example, Figure 2 in the left figure of, the distance obtained by directly calculating the perpendicular distance from the point to the boundary line is shown, Figure 2 in the right figure of, a straight line with a specified slope is drawn for this pixel point, and the distance from this point to the intersection of the line and the boundary is used to obtain the expression of the weight function for fusion as follows: , where, is the fusion weight of this pixel point on the i-th image, is the transformation function, is the distance of this pixel point from the boundary of the overlapping area of the i-th image, is the distance of this pixel point from the boundary of the overlapping area of the j-th image.
[0020] S303: Based on the fused basic variables and the fused weight function, calculate the distance of each point relative to the boundary of the fused region pixel by pixel to generate the fusion weight of each point for the two overlapping images, and obtain the mask of the fused region weight. The mask of the fused region weight is used for pixel-level hierarchical fusion of the overlapping region.
[0021] In an embodiment of the present invention, as Figure 3 shown, it is the OpenGL rendering flow chart. First, normalize the left points of the bird's-eye view image to obtain the GL vertex data of the target image and input it to the vertex shader of OpenGL for processing. Then, transfer it to the fragment shader for processing. Perform an inverse top-down transformation through the external parameter mapping table and an inverse correction transformation through the internal parameter mapping table to obtain the texture coordinate points. Perform fusion weights on the texture coordinate points through the fusion mask. Finally, map the output of the OpenGL vertex shader and combine it with the output of the OpenGL fragment shader for rendering to obtain the screen coordinates. The present invention uses OpenGL to generate the bird's-eye view perspective image. The specific steps of S4 are as follows: S401: Normalize the coordinate system of the bird's-eye view perspective image to obtain the normalized coordinate points; S402: Take the normalized coordinate points as vertex data and input them to the vertex shader of OpenGL for processing to obtain the vertex coordinates and input them to the fragment shader of OpenGL; S403: According to the external parameter mapping table, use the fragment shader of OpenGL to perform an inverse top-down transformation on the vertex coordinates, and according to the internal parameter mapping table, perform an inverse correction transformation on the vertex coordinates to obtain the texture coordinates; S404: According to the mask of the fused region weight, use the fragment shader of OpenGL to perform weighted rendering of the fused region on the texture coordinates to obtain the rendering coordinates; S405: Use the vertex coordinates for mapping and use the rendering coordinates for rendering to obtain the bird's-eye view perspective image. Among them, OpenGL is based on the GPU for rendering, and the parallel rendering process is implemented through the shader, which speeds up the generation speed of the bird's-eye view perspective image and can meet the real-time requirements of garbage detection and recognition.
[0022] In one embodiment of the present invention, the Lang-Segment-Anything deep model in S5 is an existing model that can input garbage-related prompt words, such as "Rubbish", into the Lang-Segment-Anything deep model, enabling the Lang-Segment-Anything deep model to generate a mask for the garbage in the EBV bird's-eye view image, perform garbage instance segmentation on the bird's-eye view image, and obtain the segmentation mask of the garbage; the Lang-Segment-Anything deep model includes: the Segment-Anything image segmentation model and the GroundingDINO object detection model. In this embodiment, the existing trained model can be directly used without additional data processing and training, reducing the dependence of the garbage rapid recognition algorithm model training on data and reducing the amount of data calculation at the same time.
[0023] In one embodiment of the present invention, the internal parameter mapping table, the external parameter mapping table, and the mask of the fusion region weight in the present invention do not need to be updated after being generated once. During the actual operation of the present invention, direct rendering through OpenGL can meet the real-time requirement of bird's-eye view image generation, and the image is in the video memory and can be directly used for calculation by the segmentation algorithm.
[0024] The beneficial effects of the present invention are as follows: The present invention generates a bird's-eye view image through the panoramic video stitching technology, solves the problems of the limited field of view of a single camera, the target object spanning cameras and occlusion, improves the performance of the garbage target recognition algorithm of the unmanned sanitation vehicle, and is suitable for subsequent matching with map information; The present invention performs garbage detection on the bird's-eye view image through the trained Lang-Segment-Anything deep model, has lower video memory occupancy, and realizes the rapid generation of the bird's-eye view image based on the OpenGL technology, meeting the real-time requirement; The present invention uses the trained Lang-Segment-Anything deep model, without the need to construct relevant data sets and subsequent model training, reducing the dependence of the garbage rapid recognition algorithm model training on data.
Claims
1. A method for rapid garbage recognition based on an omnidirectional camera of an unmanned sanitation vehicle, characterized in that, It includes the following steps: S1: Obtain the original images with timestamp synchronization based on a multi-channel unmanned vehicle-mounted fish-eye camera; S2: Process the original images with timestamp synchronization using the calibrated internal and external parameters to obtain the internal parameter mapping table and the external parameter mapping table; S3: Based on the original images with timestamp synchronization, manually select the overlapping areas and perform pixel-level hierarchical fusion on the overlapping areas to generate a mask of the fusion area weights; S4: Use OpenGL for fast rendering based on the internal parameter mapping table, the external parameter mapping table, and the mask of the fusion area weights to obtain the bird's-eye view image; S5: Use the Lang-Segment-Anything depth model to perform garbage instance segmentation on the bird's-eye view image to obtain the segmentation mask of the garbage; S6: Based on the segmentation mask of the garbage, obtain the garbage coordinates in the bird's-eye view image and perform conversion according to the calibrated internal and external parameters to obtain the garbage recognition coordinate values in the vehicle body coordinate system.
2. The garbage rapid recognition method based on the omnidirectional camera of the unmanned sanitation vehicle according to claim 1, wherein The specific steps of S2 are as follows: S201: Process using the InitUndistortRectifyMap image distortion correction function according to the calibrated internal parameters to obtain the internal parameter mapping table; S202: Map the pixels of the original images with timestamp synchronization according to the internal parameter mapping table to obtain the corrected images; S203: Perform regional cropping on the corrected images to obtain the cropped images; S204: Project the cropped images into the vehicle body bird's-eye coordinate system according to the calibrated external parameters to obtain the external parameter mapping table.
3. The method for rapid garbage recognition based on an omnidirectional camera of an unmanned sanitation vehicle according to claim 1, wherein, The specific steps of S3 are as follows: S301: Based on the original images with timestamp synchronization, manually select the overlapping areas to obtain the overlapping areas between the original images; S302: Define the basic variables for fusion and the weight function for fusion according to the original images with timestamp synchronization and the selected overlapping areas; S303: Generate the fusion weights for each point for the two overlapping images by calculating the distance of each point from the boundary of the fusion area pixel by pixel according to the basic variables for fusion and the weight function for fusion to obtain the mask of the fusion area weights.
4. The garbage rapid recognition method based on the omnidirectional camera of the unmanned sanitation vehicle according to claim 1, characterized in that, The specific steps of S4 are as follows: S401: Normalize the bird's-eye view image coordinate system to obtain the normalized coordinate points; S402: Use the normalized coordinate points as vertex data and input them into the vertex shader of OpenGL for processing to obtain the vertex coordinates and input them into the fragment shader of OpenGL; S403: According to the external parameter mapping table, perform an inverse top-down transformation on the vertex coordinates using the fragment shader of OpenGL, and according to the internal parameter mapping table, perform an inverse correction transformation on the vertex coordinates to obtain the texture coordinates; S404: According to the mask of the fusion area weights, perform weighted rendering of the fusion area on the texture coordinates using the fragment shader of OpenGL to obtain the rendered coordinates; S405: Perform mapping using the vertex coordinates and perform rendering using the rendered coordinates to obtain the bird's-eye view image.
5. The method for rapid garbage recognition based on an omnidirectional camera of an unmanned sanitation vehicle according to claim 1, characterized in that The specific implementation method of S5 is as follows: Input the prompt words related to garbage into the trained Lang-Segment-Anything deep model, and use the trained Lang-Segment-Anything deep model to perform garbage instance segmentation on the bird's-eye view image to obtain the segmentation mask of the garbage; the trained Lang-Segment-Anything deep model is an existing model, including: the Segment-Anything image segmentation model and the GroundingDINO object detection model.
Citation Information
Patent Citations
Pavement pit and pond detection system and method based on vehicle-mounted look-around
CN112348775A
Environment sensing system for intelligent sanitation vehicle
CN112896879A
Vehicle blind area anti-collision early warning system and method
CN113276769A
Visual laser-based active garbage cleaning method for garbage cleaning unmanned vehicle
CN117130370A
Panoramic aerial view image generating method
WO2021185284A1