Method for recording three-dimensional scene information by adopting two-dimensional vector diagram and application thereof
By projecting multiple pairs of parallel-line lasers and recording laser stripe vector diagrams, the problems of large calculation volume and high complexity in the prior art are solved, and the rapid acquisition and real-time perception of three-dimensional scene information is realized, and the construction of three-dimensional data sets suitable for autonomous driving and machine learning is implemented.
Patent Information
- Application Number
- CN202510540998.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has large calculations, complex calculations and is difficult to apply to dynamic scenarios when acquiring three-dimensional scene information, especially in the field of autonomous driving, especially in autonomous driving technology, and cannot respond quickly.
Optical methods are used to project multiple pairs of parallel-line lasers with fixed spacing to the scene, and the RGB dot matrix and laser stripe images are captured and separated by the camera. The perspective principle is used to compress it into a laser stripe vector image to record three-dimensional scene information. A new rule is used to record laser stripe vectorized lines to form a composite pixel coordinate recording method similar to three-dimensional coordinates.
It greatly reduces the consumption of computing resources, realizes the rapid acquisition and real-time perception of three-dimensional scene information, simplifies the computing needs of machine learning and training, reduces economic costs, and is suitable for the collection and reconstruction of three-dimensional information in dynamic and static scenarios.
Smart Images

Figure CN120343225A_ABST
Abstract
Description
Technical Field
[0002] The present invention relates to a method for recording three-dimensional scene information based on optical projection and vector compression and its application, and is particularly suitable for the efficient acquisition and reconstruction of three-dimensional information of dynamic or static scenes in the field of machine vision. Specific application scenarios include, but are not limited to: applications in the field of machine vision such as autonomous vehicles, drones, floor cleaning robots, mobile phones, remote sensing, engineering three-dimensional surveying and mapping, robots, manipulators, automated warehousing, automated production lines, assembly lines, conveyor lines, sorting lines, inspection lines, etc., especially for assisting human driving and / or autonomous vehicles, including vehicles for carrying people and goods, logistics distribution, etc. and methods for road and environment perception of engineering vehicles and machinery such as automatically cleaning roads, automatically plowing and harvesting fields, automatically excavating earth and stone, automatically mining and transporting minerals, etc., as well as applications in aspects such as constructing a lightweight three-dimensional scene dataset or database to support model machine learning and training. Background Art
[0003] The methods for recording three-dimensional scene information using two-dimensional images can be roughly divided into several categories: methods based on feature extraction and matching, deep learning-based models, technologies based on light fields and depth encoding, and methods combining multi-modal data, etc. The present invention belongs to the technology of depth encoding. It projects multiple pairs of parallel-line lasers with a fixed spacing onto the scene, uses the perspective principle naturally possessed by two-dimensional images to perform depth encoding on the scene image, and then uses data compression methods to achieve the recording of three-dimensional scene information with two-dimensional vector graphics. One of the techniques for obtaining three-dimensional information of the perceived scene using a camera, especially the method of inferring the distance of an object in the scene based on the perspective principle, is an indirect method based on computer image recognition: by comparing a large number of labeled physical pictures, identifying the target object, obtaining its true size information, and then estimating the distance of the target object by comparing the imaging size or apparent size of the target object captured by the camera on the photosensitive target of the camera based on the perspective principle. This software method based on computer image recognition of labeled images has low accuracy but high reliability in the presence of samples. However, in addition to the inherent defect that it is impossible to enumerate all the labeled sample images, this technology also brings disadvantages such as huge comparison calculation workload and the existence of long-tail problems. However, it is already the technical solution closest to the purpose and technical principle of the present invention. The difference is that it uses computer image processing methods to identify the labeled target object, and then obtains the distance information of the target object in the scene through the application of the perspective principle. It requires the storage and analysis processing capabilities of a large number of labeled image sample libraries or data sets. It is the method that Tesla's autonomous driving cars have used and still have not got rid of. In addition, the other main technical methods for a camera to obtain three-dimensional information of a scene are mainly binocular disparity technology, which has no inherent relationship with the perspective principle. Another widely used technology is to use structured light illumination to assist the camera in obtaining three-dimensional information: a structured light source is matched with a camera and maintains a strict positional relationship, or a structured light source is combined with one or more cameras that maintain a strict positional relationship with each other to obtain three-dimensional information by calculating fixed triangular relationships. It also belongs to depth encoding. These technologies also have no direct relationship with the recognition of the size of the object in the scene and the application of the perspective principle. Instead, the light source device of structured light is closest to the light source device used in the present invention. The main difference lies in the different derivative characteristics caused by different usage methods. And in the existing structured light technologies and methods, there has been no attempt to directly project multiple pairs of parallel-line lasers as depth encoding identifiers by optical methods. Generally, the perspective principle is not used either, because after three-dimensional information of the scene can be obtained through triangular calculation, the corresponding size information as a result can be obtained through calculation, rather than obtaining reference size information first in order to calculate the distance. Moreover, after obtaining the three-dimensional information of the scene, the distance information is also one of the size information. However, the structured light technology method requires anchoring the same target point and requires a large amount of triangular calculation work, which is quite complicated and not the strength of the processor.Before the results are obtained through trigonometric calculations, the data is completely incomprehensible and lacks corresponding rapid preview technology. Therefore, it can generally only be used for static scene measurement or 3D modeling and is difficult to be used in motion scenes that require rapid response. They have been proven to be less applicable in the field of unmanned driving technology.
[0004] No publicly available documents similar or close to the technical solution of the present invention have been retrieved. It can only be said that for the same technical purpose of obtaining 3D information of the scene and in the same optical technology field, the prior application comparison documents do not contain or even imply the basic principle of the present invention at all. The present invention is a major improvement in the data processing and recording method of the prior application: "Optical Method and Device for Projecting Reference Dimension Identifiers onto a Scene and Its Application" with the application date of April 22, 2024 and the application number of 202410485012.1. It expands the content of its technical capabilities, greatly simplifies the operations that consume huge computing resources in modern machine learning and training methods, such as extracting image features from video images and making 3D space occupancy predictions, etc. Especially in the transformation of huge machine learning video data sets or databases, it has great economic value. Summary of the Invention
[0005] The technical operation process of the present invention is as follows: First, an optical method is used to project multiple pairs of parallel line lasers perpendicular to the ground onto the scene, marking depth-encoded stripes on the scene. Then, a camera is used to take pictures. After obtaining the two-dimensional image of the scene, the two-dimensional dot matrix (RGB dot matrix) image composed of red, green, and blue colors and the laser stripe image formed by the laser stripes are separated from the two-dimensional image. The laser wavelength is not limited, but since red and infrared are easily compatible with color filters, red can be preferentially used. The stripe image only has positioning value, and the stripes themselves are very narrow and can be compressed into a vector image composed of straight line segments with theoretically no width, becoming a laser stripe vector image formed by parallel line lasers, and recorded according to the method disclosed in the present invention. This laser stripe vector image with a greatly compressed data volume directly uses the perspective principle that the spacing formed by the parallel line lasers for imaging is inversely proportional to the distance between the actual laser stripes and the camera to encode and decode depth information and restore the three-dimensional scene. The laser stripe vector image is completely spatiotemporally aligned with the two-dimensional dot matrix photo of red, green, and blue colors that details the scene details, and a collection of the two can be constructed to facilitate obtaining more information. For example, after using the laser stripe vector image formed by the parallel line lasers of the present invention to interpret the three-dimensional spatial information of the scene, it is more convenient to use the attention concentration mechanism to preferentially process the information of key targets identified in the scene from the three-dimensional information, such as traffic lights, lane lines, pedestrians, etc., and then further interpret the corresponding details in the aligned RGB dot matrix image. Compared with the previous feature extraction work of the convolutional neural pipeline for the entire image plane, it can greatly save the consumption of computing resources. Of course, the greatest value is that the three-dimensional spatial information of the scene interpreted from the laser stripe vector image formed by the parallel line lasers can completely replace the complex three-dimensional occupancy trend prediction technology and achieve direct three-dimensional real-time perception of the scene. Therefore, it has very important technical and economic value and actually becomes a brand-new lidar technology solution, and will create a brand-new three-dimensional vector image standard in the field of machine vision technology for constructing scene data sets or databases for machine learning and training for applications such as autonomous driving technology. Brief Description of the Drawings
[0006] Figure 1 It is a schematic axonometric view of the technical solution of multiple pairs of parallel line lasers distributed in a fan shape according to the present invention; Figure 2 It is a schematic axonometric view of the present invention using multiple pairs of parallel line lasers distributed in a fan shape projected onto a road scene; Figure 3 It is the laser stripe image separated from the photo of the present invention when multiple pairs of parallel line lasers distributed in a fan shape are projected onto the scene; Figure 4 It is an example of a pair of laser stripe segments generated by any pair of parallel line lasers in the laser stripe image of the present invention.
[0007] In the figure: 1. Parallel-line laser device; 2. Parallel-line laser; 3. Laser stripe; 4. Vehicle projecting parallel-line laser; 5. Laser stripe projected onto the road surface; 6. Laser stripe projected onto the obstacle; 7. Laser stripe projected onto the truck; 8. Segment A of the laser stripe; 9. Segment B of the laser stripe. Detailed implementation mode
[0008] As Figure 1 shown is an axonometric view of the technical solution of the multi-pair parallel-line lasers distributed in a fan shape according to the present invention. Among them, 1 is the parallel-line laser device, 2 is the parallel-line laser, and 3 is the laser stripe. The parallel-line laser device will generate multiple pairs of parallel-line lasers with a fixed spacing and project them into the scene to generate multiple pairs of parallel-line laser stripes. The laser stripe 3 among them is the situation generated when the parallel-line laser 2 projects onto an imaginary vertical circular screen to describe such a laser configuration. The detailed technical content is disclosed in the prior application with the application date of April 22, 2024 and the application number of 202410485012.1: "Optical Method and Device for Projecting Reference Dimension Markers onto a Scene and Its Applications".
[0009] As Figure 2 shown is an axonometric view of the present invention using multiple pairs of parallel-line lasers distributed in a fan shape and projected onto a road scene. 4 is the vehicle projecting parallel-line laser, that is, the parallel-line laser device 1 is installed on this vehicle, and the parallel-line laser is projected from it. 5 is the laser stripe projected onto the road surface, 6 is the laser stripe projected onto the obstacle, and 7 is the laser stripe projected onto the truck. Figure 2 It is drawn using computer three-dimensional modeling, vividly describing the scene that can be observed by the human eye and photographed by a camera when the parallel-line laser of the present invention is projected onto a simulated road traffic scene. In fact, the image of the two-dimensional photo taken by the camera is like this, and the only difference is that there is no color because patent drawings do not accept color drawings.
[0010] As Figure 3 shown is the laser stripe image separated from the photo of the present invention using multiple pairs of parallel-line lasers distributed in a fan shape and projected onto the scene. An ordinary color camera takes a color image, which is generally called an RGB image. When the parallel-line laser device 1 of the present invention uses red or infrared laser, the laser stripe in the scene will be a red or infrared image accordingly, and it can be separated from the color image alone. The separated red or infrared image is Figure 3 shown as follows. From the analytic geometry knowledge that the stripe is the intersection line generated by the multiple pairs of parallel-line lasers perpendicular to the ground intersecting with the scene surface, the geometric features of the scene can be recognized. When the vehicle moves, the multiple pairs of parallel-line lasers 2 will sweep across the scene, forming a comprehensive coverage of the scene. InFigure 2 , Figure 3 For convenience of viewing, the spacing between the parallel-line lasers is drawn much larger than the actual one.
[0011] Figure 4 The following shows an example of a pair of laser stripe segments generated by any pair of parallel-line lasers in the laser stripe image of the present invention. It is a small segment example of a pair of 2 stripe segments extracted from the Figure 3 laser stripe image in, and is used to introduce the vectorization recording method. First, using existing image recognition technologies such as performing differential calculations on adjacent pixels, find the middle position of the stripe line segment image with width and describe the stripe line segment only with the middle position, becoming a compressed stripe line segment without width; then record according to the existing vector graph recording method as Figure 4 shown a pair of laser stripe lines. The way of recording data is to record the pixel coordinates of the two endpoints of its straight line segments. The laser stripe segment A8 is recorded as: (x n , y n )(x n+1 , y n+1 ); the laser stripe segment B9 is decomposed into two straight line segments and recorded in sequence: (x n+2 , y n+2 )(x n+3 , y n+3 ) and (x n+3 , y n+3 )(x n+4 , y n+4 ). The 2 repeated (x n+3 , y n+3 ) coordinate values can be combined, becoming: (x n+2 , y n+2 )(x n+3 , y n+3 )(x n+4 , y n+4 ). In this way, an actually complete stripe line is disassembled into many straight line segments, and then the coordinate values of the endpoints of each straight line segment are recorded in sequence. When restoring the laser stripe line segment, connect all the endpoint coordinate positions in sequence with straight line segments.
[0012] Repeat this process for all the stripe segments in the photo image, and record all the endpoint coordinate data in sequence. Thus, the compression and recording tasks of the stripe segment image are completed, and a vector image composed of a series of endpoint pixel coordinate position data is formed. When restoring and displaying on the screen, simply restore and depict the straight line segments represented by the pixel coordinates. This is the result obtained after compressing the laser stripe image of the present invention using the existing vector image compression technology. It seemingly completes the vectorization compression task of the laser stripe image. However, a problem occurs when restoring the depth information: the straight line segments compressed from the laser stripe pairs formed by paired parallel lasers are not and are not easily paired and recorded. For example Figure 4 The line segments in n , y n , ) (x n+5 , y n+5 , ) and the line segments (x n+2 , y n+2 , ) (x n+3 , y n+3 , ) are paired straight line segments. The depth information can be directly calculated from the horizontal pixel distance between them. However, they are not recorded in adjacent positions because they belong to two laser stripe lines and are recorded independently and sequentially in order. This brings operational complexity to interpreting the imaging spacing of this pair of paired laser stripes because it is necessary to find the paired straight line segments, and their recording positions may be far apart due to the tortuousness of the laser stripes. If simply looking along the abscissa direction from the pixel coordinates of an endpoint, it may not be possible to directly find the inclined straight stripes, and even equations need to be solved to obtain the coordinate values of adjacent stripe segments, causing quite a lot of trouble in interpreting the stripe vector image. Therefore, obviously, a better way to record the laser stripe vector image is needed, and this becomes the task of the present invention.
[0013] To solve such a problem, the present invention adopts a brand-new rule to record the laser stripe vectorized lines: after each laser stripe is vectorized into a group of straight line segments connected end to end, the straight line segments in the paired laser stripes with the same vertical position (the ordinate values of the two endpoints are the same) are grouped into a pair of straight line segments and recorded in sequence. In addition, to maintain the paired property of the original parallel stripes, when any one of the paired laser stripes is bent, straight line segment parts are disassembled at the same ordinate position at the bending point as the two endpoints of the straight line segment. At the same time, the other straight line segment of the paired stripe is also disassembled into an endpoint of the straight line segment at the same pixel ordinate position whether it is bent or not, so that the straight line segments of the paired stripe vector image of the present invention all have the same pixel ordinate values at both ends. For example: 8 is the laser stripe segment A, and it can be represented by the pixel coordinates of its two endpoints (x n , yn )(x n+1 , y n+1 ) represents; 9 is the second segment of the laser stripe, and in the stripe image, it can be represented by the pixel coordinates of its two endpoints (x n+2 , y n+2 )(x n+3 , y n+3 )( and (x n+3 , y n+3 )(x n+4 , y n+4 ). However, the second segment of the laser stripe 9 is bent at (x n+3 , y n+3 ), while the first segment of the laser stripe 8 is not bent at this ordinate position. Now, an artificial endpoint is set at the position with the same ordinate as (x n+3 , y n+3 ): (x n+5 , y n+5 ), where y n+5 is equal to y n+3 . In this way, there is a pair of straight line segments of the stripe pair generated by parallel line laser projection (x n , y n )(x n+5 , y n+5 ) and the line segment (x n+2 , y n+2 )(x n+3 , y n+3 ). They are recorded adjacent to each other front and back and no longer need to be searched for. By the same principle, in the lower half of the second segment of the laser stripe 9, there is a section that is longer than the first segment of the laser stripe 8 for various reasons. Obviously, part of the first segment of the laser stripe 8 is at the edge of the object and the rest of the laser stripe is projected to other positions, or a small section of information is lost due to being blocked. A pair of straight line segments of the stripe pair generated by parallel line laser projection can also be obtained by the same processing method: (x n+5 , y n+5 )(x n+1 , y n+1 ) and (x n+3 , y n+3 )(x n+6 , y n+6 ). They are also recorded adjacent to each other front and back. The remaining straight line segment (x n+6 , y n+6 )(x n+4 , y n+4 ) has lost the positioning value of simply using the perspective principle and can be discarded. However, it is still useful in triangular calculation and can also be retained.
[0014] Summarize one of the vectorization methods for the stripe images generated by projecting multiple pairs of parallel line lasers perpendicular to the ground in the present invention. One way is to disassemble the paired laser stripe images into a series of pairs of line segments with the same endpoint pixel ordinate values corresponding pairwise, that is, the two paired line segments have the same pixel height, and the pixel ordinate values of the two corresponding endpoints are equal. Record the endpoint pixel coordinate values of such paired two line segments in adjacent order. The advantage of doing this is that the data recording is processed well for the first time, facilitating subsequent reading and parsing, especially when it needs to be read and parsed multiple times during machine learning and training.
[0015] The above vectorization method can already be used conveniently and reliably, but it does not yet reflect a standard format similar to a three-dimensional data structure with three-dimensional data characteristics in theory, that is, the endpoints of the vectorized line segments are recorded in a three-dimensional coordinate data format. It is both the ultimate in theory and the ideal of the inventor. The method to achieve this ideal is now disclosed by the present invention: According to the line segment data recording rule of the present invention, for example, for the stripe pair line segments: (x n ,y n )(x n+5 ,y n+5 ) and the line segment (x n+2 ,y n+2 )(x n+3 ,y n+3 ), Because there is y n = y n+2 , y n+5 = y n+3 And the pixel distances between x n , and x n+2 , x n+5 and x n+3 are very close, and their difference (offset value) Δx n+2 , Δx n+3 can be used as the recording method for the upper half of the laser stripe segment B9: That is, let x n+2 = x n + Δx n+2 , x n+3 = x n+5 + Δx n+3 Substitute into the pixel coordinates of the two endpoints (x n+2 ,y n+2 )(x n+3 ,y n+3 ) to represent the upper half of the stripe segment B9. The recording method changes from recording in sequence: (x n ,y n )(x n+5,y n+5 ), and the line segment (x n+2 ,y n+2 )(x n+3 ,y n+3 ) is changed to: (x n ,y n )(x n+5 ,y n+5 ) and (x n + Δx n+2 ,y n+2 )(x n+5 + Δx n+3 , y n+3 ). Note that there is redundant vertical coordinate data among the pixel coordinate data of these 4 pixels, which can be merged and processed. Therefore, the present invention uniquely creates a composite pixel coordinate recording method similar to a three-dimensional coordinate to record the pixel coordinates of the 4 endpoints of 2 line segments: (x n ,Δx n+2 ,y n )(x n+5 ,Δx n+3 , y n+5 ), where 2 coordinate values (x n ,y n )(x n+5 ,y n+5 ) are still the absolute coordinate positions of the 2 endpoints of stripe line segment A8, and the 2 pixel coordinates of stripe line segment B9 are relative coordinates recorded in the form of the difference Δx n , x n+5 , and are included in (x n+2 + Δx n+3 ,y n )(x n+2 + Δx n+2 )(x n+5 + Δx n+3 , y n+3 ). This unexpectedly brings a benefit: the meaning of the difference Δx n+2 ,Δx n+3 is originally that its numerical value is inversely proportional to the depth, and it is the depth information. Thus, the brand-new vector graphic data recording format of the present invention is very beautiful: a composite pixel coordinate value recording method similar to a three-dimensional coordinate is used to record the pixel coordinates of the 4 endpoints of 2 line segments, and it can be represented by 2 sets of three-dimensional data: (x n ,Δx n+2 ,y n )(x n+5 ,Δx n+3 ,y n+5 ), where, (x n ,y n )(xn+5 , y n+5 ) Partially records the pixel absolute coordinates of the two endpoints of the stripe segment A8. On this basis, add 2 offset values Δx of the abscissa n+2 and Δx n+3 and the common ordinate data y n and y n+5 That is the complete data of the two endpoint coordinates of the laser stripe segment B9. In fact, the two endpoint coordinates of the laser stripe segment B9 are no longer needed. What is needed is the depth information data at the pixel coordinate points representing the two endpoints of the stripe segment A8: Δx n+2 , Δx n+3 . Of course, changing (x n , Δx n+2 , y n ) (x n+5 , Δx n+3 , y n+5 ) to (x n , y n , Δx n+2 ) (x n+5 , y n+5 , Δx n+3 ) is also okay. Now it can be accurately known that the laser stripe vector diagram recorded in this three-dimensional data format constructed by the present invention is exactly a three-dimensional depth map, neither more nor less. And this straight line (x n , y n ) (x n+5 , y n+5 ) indicates that the scene is linear between these two pixel coordinate positions, and is completely equivalent to the depth data represented by a series of scan points near the position of this straight line of the dot matrix scanning lidar.
[0016] Furthermore: For example, using the two endpoint coordinates of the stripe segment A8 (x n , Δx n+2 , y n ) (x n+5 , Δx n+3 , y n+5)The pixel coordinates of the 4 endpoints of 2 line segments are recorded based on a single line segment format, overcoming the drawback of separate recording of paired laser stripe straight line segments. Therefore, when recording a complete continuous laser stripe line composed of multiple zigzag straight line segments connected end to end, by sequentially recording the start and end point coordinates of these straight line segments, there will be duplicate coordinate value data where the end point of the previous line segment is the start point of the next line segment for the intermediate straight line segments. Of course, the data volume can be compressed by merging the start and end points and only recording one coordinate value data, so that for a continuous laser stripe line, only the coordinate values starting from the beginning are recorded, and the coordinate values are recorded at the turning points of the straight line segments, and finally the recording of the completely continuous laser stripe line ends with the end point coordinates of the last straight line segment. Now a new problem arises: Do we need to isolate different stripe lines to distinguish them? Because if not, a "new line segment" will be fictitiously created in the scene during interpretation and restoration, connecting the end point of the previous laser stripe line to the start point of the next laser stripe line, which will represent the existence of this "fictitious linear space" in the scene, and it is very likely to be completely inconsistent with the actual scene space. Handling this problem will involve complex recording path topology issues and complicate the problem. The simplified processing method of the present invention is limited by a certain length value in pixels of this "new line segment" that will be fictitiously created. If it is shorter than the limited length value, it is retained and recorded as an acceptable "new line segment". Otherwise, it is interrupted, and the recording of the remaining laser stripe lines behind is restarted separately, so there are multiple laser stripe lines in a laser stripe vector map.
[0017] The present invention realizes the inventor's dream of directly recording three-dimensional information in a planar image! The patents of the inventor's two applications (application numbers are 2022106914970 and 2024104850121 respectively) are explorations driven by such an idea. Obviously, the basic geometric principle of the first application is to determine the depth by adding a laser plane constraint at a known position in planar imaging. When the second application was made, the best data recording method and three-dimensional data format had not been conceived yet.
[0018] When actually processing the stripe image, the situation where a single stripe is partially obscured will be encountered, that is, when there are unpaired stripes. In the data recording of the present invention, the offset value Δx n+4 can be directly recorded as 0, which is also self-consistent in mathematical logic and represents infinite depth, which has no meaning for environmental perception. Additionally, because Δx n+2 、Δx n+3The relative value is extremely small, so the required data bit width is also much smaller than that required for recording general pixel coordinates. Therefore, the data volume of the vector map recording format of the present invention can be compressed by nearly half again. This is not only very valuable for establishing a scene information library for machine learning and training, but also very convenient in system processing because it is not only easy to obtain, but also the three-dimensional information of the scene can be directly parsed. In contrast Figure 2 , it can be known that the preliminary interpretation of the stripe vector map of the present invention is also very simple: just look at the Δx n+2 and Δx n+3 values of the two endpoints. If the two values are the same, it means that the two stripe line segments are parallel and perpendicular to the ground, which must be the result of parallel line lasers projected on a vertical plane, that is, an obstacle; if they are different, it means that the parallel line lasers are projected on a horizontal plane or an inclined plane. The length of the two straight line segments, that is, the difference in the vertical pixel coordinates of the two endpoints, can be inferred. Of course, the influence of the error range is not discussed for the time being, but the general direction remains unchanged.
[0019] To summarize the second vectorization method of the present invention for recording the stripe image generated by projecting multiple pairs of parallel line lasers perpendicular to the ground onto the scene, it is to disassemble the paired laser stripe images into a series of straight line segment pairs with pairwise corresponding identical endpoint pixel ordinate values, that is, the two paired straight line segments have the same pixel height, and the pixel ordinate values of the 2 corresponding endpoints are equal. When recording the 4 endpoint coordinate values of such 2 straight line segment pairs, any one of the straight line segments is taken as the basis, and the 2 endpoint coordinates of the other straight line segment are combined and recorded in the way that the difference in the pixel abscissas of the two endpoints relative to the two endpoints of the first straight line segment is used as a new dimension, and combined into the class three-dimensional coordinate values of the 2 endpoints of a basic line segment, represented by three-dimensional format data: (x n , Δx n+2 , y n ) (x n+5 , Δx n+3 , y n+5 ). The 4 endpoints of a pair of laser stripe straight line segment pairs are recorded with such class three-dimensional coordinate values constructed by 2 groups of 3 values, and a pair of straight line segments is recorded. Since the relative abscissa offset value of the 2 endpoints is inherently inversely proportional to the depth, the class three-dimensional coordinate value record constructed by such 3 values is actually the three-dimensional description and record of the straight line segment endpoints.
[0020] Compare the technical feature of encoding the depth information of the scene in paired laser stripe images in the present invention with the technical feature that each laser scanning point in the point cloud image of the lidar for point cloud scanning contains depth information. It can be found that in the point cloud image of the lidar, if the spatial structure of multiple pairs of parallel-line lasers in the present invention is embedded into the point cloud image of the lidar, and the partial point cloud group with linear depth information of the lidar is connected by straight line segments according to the spatial structure mode of the multiple pairs of parallel-line lasers in the present invention, the point cloud image of the lidar can be converted into the laser stripe image in the present invention. The specific operation can be to convert the depth map of the scene point cloud of the lidar into a three-dimensional model of the scene, and then make an intersection with the three-dimensional model of the multiple pairs of parallel-line lasers in the present invention. The outer surface part of this intersection is the laser stripe. Displaying these laser stripes from the perspective of the vehicle-mounted camera or the original lidar can obtain a laser stripe image that is exactly the same as the laser stripe image obtained in the physical real scene in the present invention in terms of principle and structure. Then, just process and record this laser stripe image obtained from the lidar point cloud depth map using the vectorization method in the present invention. This innovative technical processing method has very important technical and economic value: Because the present invention creates a brand-new low-cost environmental perception technical route in the relatively mature lidar technical route. This itself will inevitably encounter the problem of the cumbersome old technology and the entry threshold of the new technology, as well as the problem of digital resources that can be used for machine learning and training. And being able to transplant and utilize the digital assets accumulated over the years in the old technical route has very important technical and economic value for the industry to transfer from the old technical route to the new technical route. The huge road scene dataset or database accumulated by the prior lidar industry can be utilized during the implementation of the present invention. After being processed and converted, it can be used for machine learning and training the autonomous driving system, which can greatly reduce the entry threshold, save the initial investment, and is conducive to lidar practitioners changing to the technical route in the present invention. In addition, Figure 2 、 Figure 3 It is also completely generated and drawn after computer three-dimensional modeling. Therefore, the method for obtaining the vector map of the real scene in the present invention is not limited to obtaining it by projecting physical multiple pairs of paired parallel-line lasers onto the real physical scene and then taking pictures of the real scene by the camera. It can also be obtained through computer simulation, and is used to produce and establish a scene information dataset or database for the purpose of machine learning and training for autonomous driving or assisted manual driving, etc. This also provides convenience for the initial promotion and use of the present invention. However, the three-dimensional model dataset or database of the real road scene is far less than the point cloud image data of the lidar.
[0021] The applicant declares: Designate Figure 4This is the abstract drawing. The applicant declares that the term "driverless car" or "autonomous vehicle" used in this application generally refers to all cars that use computer-assisted manual driving, fully automated, or semi-autonomous driving, and of course also includes all levels and all-purpose autonomous vehicles.
Claims
1. A method for recording three-dimensional scene information using two-dimensional vector graphics and its applications, applicable to machine vision applications, including but not limited to applications in the fields of autonomous vehicles, drones, floor sweepers, mobile phones, remote sensing and telemetry, engineering three-dimensional surveying and mapping, robots, automated production lines, assembly lines, conveyor lines, sorting lines, inspection lines, etc. in the field of machine vision. In particular, it is a method for road and environment perception of vehicles such as autonomous vehicles and assisted human-driven vehicles including but not limited to passenger and cargo transportation, logistics distribution, etc., and engineering vehicles and machinery such as automatic road sweeping, automatic plowing and harvesting, automatic earthwork excavation, automatic mining, etc., and its applications in aspects such as establishing a scene information dataset or database for the purpose of machine learning and training for autonomous driving or assisted human driving. It is characterized in that: The laser stripe image is obtained after multiple pairs of parallel line lasers are projected in the scene to form multiple pairs of laser stripes and are captured by a camera. The pixel positions at the bends of any pair of laser stripes are used as the endpoints of the line segments, and the laser stripe image is disassembled and segmented into many pairs of laser stripe line segments. The vertical coordinate values of the corresponding endpoints of each pair of line segments are the same. The pixel coordinate values of the endpoints of each pair of line segments are recorded successively before and after, and thus a vectorized line segment image of the paired laser stripes is formed. Or the vector diagram of the stripe image is recorded by taking any one of the two stripe line segments in a pair as the reference, and the other is recorded by the pixel horizontal coordinate difference between the corresponding endpoints of the reference stripe line segment. The way or data format for describing the endpoint coordinates of such paired stripe line segments after combination includes the pixel coordinate values in the following three dimensions: the pixel horizontal coordinate value of the endpoint of the line segment used as the reference, the pixel horizontal coordinate difference between the endpoint of the second line segment paired with it and the corresponding endpoint of the reference line segment, and the equal pixel vertical coordinate values of the corresponding endpoints of the two line segments.
2. The invention according to claim 1, characterized in that: The vector diagram of the laser stripe image is recorded by taking any one of the two stripe line segments in a pair as the reference, and the other is recorded by the pixel horizontal coordinate difference between the corresponding endpoints of the reference stripe line segment. The way or data format for describing the endpoint coordinates of such paired stripe line segments after combination includes the pixel coordinate values in the following three dimensions: the pixel horizontal coordinate value of the endpoint of the line segment used as the reference, the pixel horizontal coordinate difference between the endpoint of the second line segment paired with it and the corresponding endpoint of the reference line segment, and the equal pixel vertical coordinate values of the corresponding endpoints of the two line segments.
3. The invention according to claim 1, characterized in that: A three-dimensional model generated by computer three-dimensional modeling replaces the multiple pairs of parallel line lasers in the projection scene. After intersecting with the three-dimensional model generated by restoring the scene from the point cloud depth image of the lidar, the resulting intersection line image forms a laser stripe image, which is vectorized and used to construct a scene information dataset or database established for the purpose of machine learning and training for autonomous driving or assisted manual driving.
4. The invention according to claim 2, characterized in that: A three-dimensional model generated by computer three-dimensional modeling replaces the multiple pairs of parallel line lasers in the projection scene. After intersecting with the three-dimensional model generated by restoring the scene from the point cloud depth image of the lidar, the resulting intersection line image forms a laser stripe image, which is vectorized and used to construct a scene information dataset or database established for the purpose of machine learning and training for autonomous driving or assisted manual driving.
Citation Information
Patent Citations
Optical method and device for projecting reference size identification to scene and application of optical method and device
CN118540444A