Trajectory Image Alignment Method, Device, Computer Device, and Storage Medium
By using initial alignment, object detection, semantic segmentation and depth of field estimation techniques in track image alignment, we determine the relative distances of adjacent image pairs in the track image and calculate the track offset, the problem of poor positioning effect is solved, and the track image alignment with higher accuracy is achieved.
Patent Information
- Application Number
- CN202111536367.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-12-15
AI Technical Summary
The prior art cannot guarantee the alignment accuracy of the road trajectory image when the positioning effect is poor, resulting in inaccurate image alignment.
By acquiring the trajectory image to be processed, the road network circuit and positioning information are initially aligned, combined with the target detection, semantic segmentation and depth of field estimation results, the relative distance between adjacent image pairs is determined, and the offset between trajectories is calculated according to this calculation.
When the positioning effect is poor, the alignment accuracy of the track image can be improved to ensure the accuracy and effectiveness of image alignment.
Smart Images

Figure CN114332174B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular, to a method and apparatus for aligning trajectory images, a computer device, a storage medium, and a computer program product. Background Art
[0002] With the development of computer technology, computer vision technology has also been continuously developing and progressing. Computer vision technology (CV) is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, tracking, and measurement on targets, and further performing graphic processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. Road trajectory images are basic data in computer vision processing. Road trajectory images refer to pictures taken by in-vehicle cameras during vehicle driving. To reduce data redundancy, only one picture is taken and stored at each trajectory point on the road. The alignment of road trajectory images has extensive applications in fields such as map data collection tasks, high-precision map automatic generation, road data operations, and autonomous driving based on computer vision technology.
[0003] Currently, for the alignment of road trajectory images, it is generally necessary to first collect the position information corresponding to the trajectory images, and then project the points on one trajectory onto another trajectory, so as to perform the alignment of the trajectory images based on the projection results. However, this alignment method depends on the accuracy of the position information corresponding to the trajectory images, and the alignment accuracy cannot be guaranteed when the positioning effect is poor. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a method and apparatus for aligning trajectory images, a computer device, a computer-readable storage medium, and a computer program product that can improve the alignment accuracy of trajectory images.
[0005] In a first aspect, the present application provides a method for aligning trajectory images. The method includes:
[0006] Obtain a trajectory image to be processed;
[0007] Perform initial alignment on the trajectory image to be processed according to the road network line and positioning information corresponding to the trajectory image to be processed, and obtain the initial alignment position information corresponding to the trajectory image to be processed;
[0008] Perform content analysis processing on the trajectory image to be processed, and obtain the target detection result, semantic segmentation result, and depth of field estimation result of the reference object in the trajectory image to be processed;
[0009] Based on the initial alignment position information, object detection results, semantic segmentation results, and depth of field estimation results, determine the relative distance between adjacent image pairs in the to-be-processed trajectory image;
[0010] Determine the offset between the trajectories in the to-be-processed trajectory image according to the relative distance between the adjacent image pairs, and perform alignment processing on the to-be-processed trajectory image according to the offset.
[0011] In a second aspect, the present application also provides a trajectory image alignment device. The device includes:
[0012] An image acquisition module, configured to acquire a to-be-processed trajectory image;
[0013] An initial alignment module, configured to perform initial alignment on the to-be-processed trajectory image according to the road network line and positioning information corresponding to the to-be-processed trajectory image, and obtain the initial alignment position information corresponding to the to-be-processed trajectory image;
[0014] A content analysis module, configured to perform content analysis processing on the to-be-processed trajectory image, and obtain the object detection result, semantic segmentation result, and depth of field estimation result corresponding to the reference object in the to-be-processed trajectory image;
[0015] A relative distance calculation module, configured to determine the relative distance between adjacent image pairs in the to-be-processed trajectory image based on the initial alignment position information, object detection results, semantic segmentation results, and depth of field estimation results;
[0016] An image alignment module, configured to determine the offset between the trajectories in the to-be-processed trajectory image according to the relative distance between the adjacent image pairs, and perform alignment processing on the to-be-processed trajectory image according to the offset.
[0017] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0018] Acquire a to-be-processed trajectory image;
[0019] Perform initial alignment on the to-be-processed trajectory image according to the road network line and positioning information corresponding to the to-be-processed trajectory image, and obtain the initial alignment position information corresponding to the to-be-processed trajectory image;
[0020] Perform content analysis processing on the to-be-processed trajectory image, and obtain the object detection result, semantic segmentation result, and depth of field estimation result corresponding to the reference object in the to-be-processed trajectory image;
[0021] Based on the initial alignment position information, object detection results, semantic segmentation results, and depth of field estimation results, determine the relative distance between adjacent image pairs in the to-be-processed trajectory image;
[0022] Determine the offset between trajectories in the to-be-processed trajectory image according to the relative distance between adjacent image pairs, and perform alignment processing on the to-be-processed trajectory image according to the offset.
[0023] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0024] Obtain a to-be-processed trajectory image;
[0025] Perform initial alignment on the to-be-processed trajectory image according to the road network line and positioning information corresponding to the to-be-processed trajectory image, and obtain the initial alignment position information corresponding to the to-be-processed trajectory image;
[0026] Perform content analysis processing on the to-be-processed trajectory image, and obtain the object detection result, semantic segmentation result, and depth of field estimation result corresponding to the reference object in the to-be-processed trajectory image;
[0027] Based on the initial alignment position information, object detection results, semantic segmentation results, and depth of field estimation results, determine the relative distance between adjacent image pairs in the to-be-processed trajectory image;
[0028] Determine the offset between trajectories in the to-be-processed trajectory image according to the relative distance between adjacent image pairs, and perform alignment processing on the to-be-processed trajectory image according to the offset.
[0029] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0030] Obtain a to-be-processed trajectory image;
[0031] Perform initial alignment on the to-be-processed trajectory image according to the road network line and positioning information corresponding to the to-be-processed trajectory image, and obtain the initial alignment position information corresponding to the to-be-processed trajectory image;
[0032] Perform content analysis processing on the to-be-processed trajectory image, and obtain the object detection result, semantic segmentation result, and depth of field estimation result corresponding to the reference object in the to-be-processed trajectory image;
[0033] Based on the initial alignment position information, target detection results, semantic segmentation results, and depth of field estimation results, determine the relative distance between adjacent image pairs in the to-be-processed trajectory image;
[0034] Determine the offset between trajectories in the to-be-processed trajectory image according to the relative distance between adjacent image pairs, and perform alignment processing on the to-be-processed trajectory image according to the offset.
[0035] The above-mentioned trajectory image alignment method, device, computer device, storage medium, and computer program product, wherein the method, after obtaining the to-be-processed trajectory image, performs initial alignment on the to-be-processed trajectory image according to the road network line and positioning information corresponding to the to-be-processed trajectory image, so as to obtain the initial alignment position information corresponding to the to-be-processed trajectory image; perform content analysis processing on the to-be-processed trajectory image to obtain the target detection results, semantic segmentation results, and depth of field estimation results of the reference object in the to-be-processed trajectory image; based on the initial alignment position information, target detection results, semantic segmentation results, and depth of field estimation results, determine the relative distance between adjacent image pairs in the to-be-processed trajectory image; determine the offset between trajectories in the to-be-processed trajectory image according to the relative distance between adjacent image pairs, and perform alignment processing on the to-be-processed trajectory image according to the offset. In this application, after the initial alignment of the to-be-processed trajectory image, based on the initial alignment position information, as well as the target detection results, semantic segmentation results, and depth of field estimation results corresponding to adjacent road signs, the offset between trajectories is determined, so as to perform alignment processing of the trajectory image, which can ensure the alignment accuracy of the trajectory image when the positioning effect is poor. Description of the Drawings
[0036] Figure 1 It is a schematic diagram of the application environment of the trajectory image alignment method in an embodiment;
[0037] Figure 2 It is a schematic flowchart of the trajectory image alignment method in an embodiment;
[0038] Figure 3 It is a schematic flowchart of the step of performing initial alignment on the to-be-processed trajectory image in an embodiment;
[0039] Figure 4 It is a schematic diagram of the trajectory point projection step in an embodiment;
[0040] Figure 5 It is a schematic flowchart of the step of performing content analysis on the to-be-processed trajectory image in an embodiment;
[0041] Figure 6 It is a schematic diagram of the overall network structure of the neural network model in an embodiment;
[0042] Figure 7 Schematic diagram of the network structure of the target detection head in one embodiment;
[0043] Figure 8 Schematic diagram of the network structure of the semantic segmentation head in one embodiment;
[0044] Figure 9 Schematic diagram of the principle of pinhole imaging in the depth of field estimation process in one embodiment;
[0045] Figure 10 Schematic flowchart of the steps for obtaining the relative distance between adjacent image pairs in one embodiment;
[0046] Figure 11 Schematic flowchart of the steps for determining adjacent images in one embodiment;
[0047] Figure 12 Schematic flowchart of the steps for obtaining the relative distance between adjacent image pairs in another embodiment;
[0048] Figure 13 Schematic flowchart of the steps for aligning two trajectories in the trajectory image to be processed in one embodiment;
[0049] Figure 14 Schematic diagram of the structure of the matrix grid in one embodiment;
[0050] Figure 15 Schematic diagram of the alignment effect of the trajectory in one embodiment;
[0051] Figure 16 Schematic flowchart of the trajectory image alignment method in another embodiment;
[0052] Figure 17 Schematic block diagram of the trajectory image alignment device in one embodiment;
[0053] Figure 18 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners
[0054] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0055] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. The technical solution of this application mainly involves computer vision technology and machine learning technology in machine learning.
[0056] Among them, computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement in machine vision, and further performing graphic processing to make the computer-processed images more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0057] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0058] The trajectory image alignment method provided by the embodiments of this application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers. The terminal 102 can send the to-be-processed trajectory images to the server 104 to align these trajectory images through the server 104. The server 104 obtains the to-be-processed trajectory images; based on the road network lines and positioning information corresponding to the to-be-processed trajectory images, performs initial alignment on the to-be-processed trajectory images to obtain the initial alignment position information corresponding to the to-be-processed trajectory images; performs content analysis processing on the to-be-processed trajectory images to obtain the object detection results, semantic segmentation results, and depth of field estimation results corresponding to the reference objects in the to-be-processed trajectory images; based on the initial alignment position information, object detection results, semantic segmentation results, and depth of field estimation results, determines the relative distance between adjacent image pairs in the to-be-processed trajectory images; determines the offset between the trajectories in the to-be-processed trajectory images according to the relative distance between adjacent image pairs, and performs alignment processing on the to-be-processed trajectory images according to the offset. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0059] In one embodiment, as Figure 2 shown, a method for aligning trajectory images is provided. Taking the server 104 end in Figure 1 as an example, the method includes the following steps:
[0060] Step 201, obtain the to-be-processed trajectory images.
[0061] Among them, the trajectory image to be processed refers to the target image processed by the trajectory image alignment method of the present application. Specifically, the trajectory image can be an image collected by an in-vehicle camera during driving. These images are generally road images corresponding to the trajectory points of driving. In one embodiment, the in-vehicle camera can specifically be a monocular camera. In an automobile, due to cost limitations, generally only a single optical camera is equipped, so the parallax relationship between multiple cameras cannot be used to estimate the geometric characteristics of the image, etc. The trajectory point refers to the route along which the vehicle travels on the road. To reduce data redundancy, sequence information formed by collecting the current geographical location at regular distance intervals is generally used, and the geographical location is generally the GPS (Global Positioning System) position. The combination of multiple trajectory points collected during the vehicle's driving is the trajectory image, which specifically refers to the picture taken by the in-vehicle monocular optical camera during the vehicle's driving. Similarly, to reduce data redundancy, only one picture is taken and stored at each trajectory point. In one embodiment, for the convenience of processing, the trajectory image to be processed in the present application can process the trajectory images corresponding to only two trajectories at a time.
[0062] Specifically, when the staff on the terminal 102 side needs to perform tasks such as map data collection, high-precision map automatic generation, road data operation, and autonomous driving, generally, the trajectory images on the road need to be used as the basic data to complete these tasks. These trajectory images to be processed may be composed of multiple trajectory images collected by different vehicles during driving. Therefore, it may be necessary to perform alignment processing on these trajectory images before using them. Therefore, the trajectory image alignment method of the present application can be used to align the trajectory images corresponding to different vehicles pairwise.
[0063] Step 203: According to the road network line and positioning information corresponding to the trajectory image to be processed, perform initial alignment on the trajectory image to be processed to obtain the initial alignment position information corresponding to the trajectory image to be processed.
[0064] Among them, the road network information specifically refers to the road information corresponding to the collection location of the trajectory image to be processed. The positioning information refers to the position information corresponding to the trajectory points collected by each trajectory image to be processed. Limited by the accuracy of the positioning facility, accurate positioning may not be achieved. For example, the collection position of the trajectory image to be processed may not be located on the road. Performing initial alignment on the trajectory image to be processed specifically means aligning the road network line and positioning information corresponding to the trajectory image to be processed to make them unified. Specifically, it can mean projecting the trajectory points whose positioning information is not on the road onto the road network. The initial alignment position information is the position information after projecting the trajectory points not on the road onto the road.
[0065] Specifically, due to the accuracy of the positioning device on the vehicle, after collecting the trajectory information on the road, it may not be possible to accurately locate each trajectory point on the road at each trajectory point, and since the trajectory points must be taken on the road. Therefore, the initial alignment of the trajectory image to be processed can be performed first according to the road network line and the positioning information corresponding to the trajectory image to be processed, and the trajectory points not on the road are located on the road, obtaining the initial alignment position information corresponding to the trajectory image to be processed. The position of each point in the trajectory image to be processed does not necessarily correspond to the shooting location, but is projected onto the actual road obtained according to the road network information.
[0066] Step 205: Perform content analysis processing on the trajectory image to be processed to obtain the target detection result, semantic segmentation result, and depth of field estimation result corresponding to the reference object in the trajectory image to be processed.
[0067] Among them, the content analysis processing specifically refers to processing each trajectory image to be processed through computer vision technology. The content analysis processing can specifically include target detection processing and depth of field estimation. Target detection refers to identifying the target of interest in the image with a rectangular box. In this application, the processed image is the trajectory image on the road, and the targets of interest in the road data generally refer to road signs related to traffic elements. Therefore, these road signs can be used as reference objects for aligning the trajectory image. For example, if the target detection is interested in the road signs in the image, that is, these road signs are detected through a deep learning algorithm and the corresponding rectangular box positioning in the image is given. It can also be interested in the traffic lights in the image, that is, these traffic lights are detected through a deep learning algorithm and the corresponding rectangular box positioning in the image is given. The depth of field estimation corresponds to the target detection. When the camera collecting the proxy trajectory image is a monocular camera. The semantic segmentation result refers to obtaining the interested part in the trajectory image to be processed through semantic segmentation technology. The semantic segmentation technology is used to segment the key areas, and the interested areas and uninterested areas can be segmented from the image through the semantic segmentation technology. In the solution of this application, the uninterested areas specifically refer to the areas that will interfere with the distance estimation, such as the interior of the vehicle, other vehicles, pedestrians, and watermarks. The depth of field estimation is specifically monocular vision depth of field estimation. According to the pictures collected by the monocular camera, the depth of each object and each pixel in the image is estimated, so as to judge the distance of each object in the image from the camera. In the solution of this application, the reference object in the image needs to be detected in the target detection process, and the depth of field estimation result refers to the estimation of the depth of field corresponding to the reference object.
[0068] Specifically, in step 203, the to-be-processed trajectory image was roughly aligned through the road network line corresponding to the to-be-processed trajectory image and the positioning information, so that the trajectory points not on the road were projected onto the road. Since the position of the reference object on the trajectory image remains unchanged, the position of the camera at the time of shooting, that is, the position of the vehicle, can be estimated through the object detection result, semantic segmentation result, and depth of field estimation result corresponding to the reference object in the to-be-processed trajectory image, thereby realizing the alignment between different trajectories. Therefore, after obtaining the initial alignment position information, computer vision technology can be used to perform content analysis and processing on the to-be-processed trajectory image, determine the reference object in the to-be-processed trajectory image, determine which parts of the image will interfere with the relative distance estimation through semantic segmentation, and identify the position of the target from the camera in it to obtain the depth of field estimation information.
[0069] Step 207, based on the initial alignment position information, object detection result, semantic segmentation result, and depth of field estimation result, determine the relative distance between adjacent image pairs in the to-be-processed trajectory image.
[0070] Among them, the relative distance between adjacent image pairs in the to-be-processed trajectory image specifically includes the distance between adjacent different trajectory points in the same trajectory, and also includes the distance between adjacent trajectory points in two different trajectories. The relative distance between adjacent image pairs specifically refers to the distance between the actual trajectory points of adjacent images on the road.
[0071] Specifically, after obtaining the initial alignment position information, it is possible to roughly identify which images in all the to-be-processed trajectory images are images of adjacent trajectory points. These images of adjacent trajectory points may all have photographed reference objects on the road, and the same reference object remains unchanged in different to-be-processed trajectory images. Therefore, the object detection result, semantic segmentation result, and depth of field estimation result corresponding to the two to-be-processed trajectory images are used to estimate the distance between adjacent to-be-processed trajectory images. For example, the object detection result indicates that the trajectory point of the to-be-processed trajectory image A is in front of sign A, and through depth of field estimation, the camera that photographed the to-be-processed trajectory image A is 400m away from sign A, while the trajectory point of the to-be-processed trajectory image B is also in front of sign A, and through depth of field estimation, the camera that photographed the to-be-processed trajectory image B is 200m away from sign A. Then, it can be known by comparison that the relative distance between the to-be-processed trajectory image A and the to-be-processed trajectory image B is 400 - 200 = 200m.
[0072] Step 209, determine the offset between the trajectories in the to-be-processed trajectory image according to the relative distance between adjacent image pairs, and perform alignment processing on the to-be-processed trajectory image according to the offset.
[0073] Among them, trajectory alignment is specifically used to align two trajectories according to their actual shooting distances. The offset between trajectories specifically refers to the relative distance between two trajectories. After determining the offset between trajectories, the trajectory points with a relatively short relative distance in the two trajectories can be determined, so as to align two different trajectories.
[0074] Specifically, after obtaining the relative distance between each adjacent trajectory image in the trajectory image to be processed, the relatively close trajectory points in the trajectory image can be determined according to the relative distance between adjacent trajectory images, so as to align the trajectories in the trajectory image to be processed in pairs. For example, the trajectory image of a trajectory A is composed of the trajectory images corresponding to the trajectory points such as A1, A2, A3, A4, and A5, and the trajectory image of a trajectory B is composed of the trajectory images corresponding to the trajectory points such as B1, B2, B3, B4, B5, and B6. After calculating the relative distance between adjacent image pairs, the offset between the trajectories is obtained, so that it can be determined that the offset between A4 and B2 is the smallest, and the offset between A5 and B3 is the smallest. Therefore, these two trajectories can be aligned according to A4, A5, B2, and B3.
[0075] The above trajectory image alignment method includes: after obtaining the trajectory image to be processed; performing initial alignment on the trajectory image to be processed according to the road network line and positioning information corresponding to the trajectory image to be processed, so as to obtain the initial alignment position information corresponding to the trajectory image to be processed; performing content analysis processing on the trajectory image to be processed to obtain the target detection result, semantic segmentation result, and depth of field estimation result of the reference object in the trajectory image to be processed; based on the initial alignment position information, target detection result, semantic segmentation result, and depth of field estimation result, determining the relative distance between adjacent image pairs in the trajectory image to be processed; determining the offset between the trajectories in the trajectory image to be processed according to the relative distance between adjacent image pairs, and performing alignment processing on the trajectory image to be processed according to the offset. In this application, after initially aligning the trajectory image to be processed, based on the initial alignment position information, as well as the target detection result, semantic segmentation result, and depth of field estimation result corresponding to adjacent road signs, the offset between the trajectories is determined, so as to perform alignment processing on the trajectory image, which can ensure the alignment accuracy of the trajectory image when the positioning effect is poor.
[0076] In one embodiment, as Figure 3 shown, step 201 includes:
[0077] Step 302, determining the trajectory points corresponding to the trajectory image to be processed according to the positioning information corresponding to the trajectory image to be processed.
[0078] Step 304, projecting the trajectory points onto the tangent direction of the road network line to obtain the projection positions corresponding to the trajectory points.
[0079] Step 306: Obtain the initial alignment position information corresponding to the to-be-processed trajectory image according to the projection positions corresponding to the to-be-processed trajectory images.
[0080] Among them, the trajectory points corresponding to the processed trajectory image refer to the shooting points of the to-be-processed trajectory image identified through the positioning information. Due to the accuracy problem of the positioning technology, this trajectory point may not be accurately positioned and may not be on the road. The projection position corresponding to the trajectory point refers to the position point obtained by projecting the trajectory point not on the road onto the road.
[0081] Specifically, due to the accuracy problem of the positioning technology, the positioning points of the to-be-processed trajectory image obtained by positioning may not be accurately positioned and may not be on the road. Therefore, the trajectory points not on the road can be approximately projected onto the road through projection. After all the trajectory points in the to-be-processed trajectory image are projected onto the road, the task of the projection process is completed, and the initial alignment position information corresponding to the to-be-processed trajectory image is obtained. Among them, the projection process can specifically project the trajectory points onto the tangent direction of the road network line to obtain the projection positions corresponding to the trajectory points. As Figure 4 shown, the curve in the figure represents the road network line, and the points outside the road network line are the trajectory points corresponding to the to-be-processed trajectory image. When performing the projection process, the trajectory points can be projected onto the tangent direction of the road network line to obtain the projection positions corresponding to the trajectory points, and the projection positions of the trajectory points on the road can be determined. The point at the tangent position of the curve in the figure is the projection point corresponding to the trajectory point. After projecting all the points not on the road in the to-be-processed trajectory image onto the road network line, the initial alignment of the to-be-processed trajectory image can be completed. In this embodiment, by using the road network information and the positioning information corresponding to the to-be-processed trajectory image to perform the initial alignment of the to-be-processed trajectory image, the position points corresponding to the to-be-processed trajectory image can be effectively projected into the actual road, thereby effectively making an initial estimate of the relative positions between the to-be-processed trajectory images and ensuring the processing efficiency of the trajectory image alignment process.
[0082] In one embodiment, as Figure 5 shown, step 205 includes:
[0083] Step 502: Obtain the target detection result corresponding to the reference object in the to-be-processed trajectory image through the target detection technology.
[0084] Step 504: Determine the relative distance interference region in the to-be-processed trajectory image through the semantic segmentation technology, and obtain the semantic segmentation result corresponding to the to-be-processed trajectory image based on the determined relative distance interference region.
[0085] Step 506: Obtain the absolute depth map corresponding to the trajectory image to be processed based on the principle of pinhole imaging, and obtain the depth of field estimation result corresponding to the trajectory image to be processed through the absolute depth map.
[0086] Among them, the object detection result refers to the matrix frame containing the reference object recognized from the trajectory image to be processed through the object detection technology of computer vision. The semantic segmentation technology is used to segment the key areas, and the areas of interest and non-interest can be segmented from the image through the semantic segmentation technology. The relative distance interference area refers to the object that may interfere with the estimation of the relative distance. In a specific embodiment, the vehicle, the interior scene of the vehicle, the watermark part, etc. in the trajectory image to be processed can be separated from the original image and removed through the semantic segmentation technology. Pinhole imaging means that a board with a small hole is blocked between the wall and the object, and an inverted real image of the object will be formed on the wall. In this application, the camera is used as the small hole for imaging estimation, so as to obtain the absolute depth map corresponding to the depth estimation area.
[0087] Specifically, in this application, before aligning the trajectory images, the trajectory image to be processed can be subjected to content analysis processing, so as to obtain the object detection result, semantic segmentation result, and depth of field estimation result corresponding to the reference object contained therein. When performing content analysis, the trajectory image to be processed can be processed in parallel through a computer vision-related model to obtain the object detection result, semantic segmentation result, and depth of field estimation result corresponding to the reference object in the trajectory image to be processed respectively. In a specific embodiment, the process of performing content analysis processing on the trajectory image to be processed can be specifically implemented through a neural network model for multi-task learning. The neural network model includes a common backbone network and three different head sub-neural networks. One of the head networks is used for object detection, one head is used for semantic segmentation, and the other head network is used for depth of field estimation. The overall network structure can refer to Figure 6 as shown. Among them, for the object detection head network, a network structure similar to the YOLOv3 detection head can be used for object detection to extract the objects of interest in the image. In this application, it is mainly used to detect the objects related to the reference object, such as road signs, warning signs, danger signs, etc. These detected objects are used as the points for distance estimation in the subsequent relative distance estimation steps. The network structure of the object detection head can specifically refer to Figure 7。For the semantic segmentation head network, a semantic segmentation network head similar to DeepLabV3 can be used to remove the regions in the to-be-processed trajectory image that interfere with the relative distance estimation. By analyzing the actual content of the to-be-processed trajectory image, the semantic segmentation in this application mainly detects four categories: the interior of the vehicle, other vehicles, pedestrians, and watermarks. The pixel points falling in these four categories of regions do not participate in the subsequent relative distance estimation, and these regions are identified as relative distance interference regions. The network structure of the semantic segmentation head can specifically refer to Figure 8 。Finally, for the depth of field estimation head network, the model can be trained using the depth information annotated on the KITTI standard dataset to obtain a depth map without absolute scale. Then, by means of the positioning distance difference of the trajectory and based on the principle of pinhole imaging, an absolute depth map is obtained. Based on the obtained absolute depth map, the depth of field estimation result corresponding to the reference object identified in the object detection is determined. The schematic diagram of pinhole imaging can be referred to Figure 9 as shown. Combining Figure 9 , the final depth of field estimation result can specifically refer to the following formula:
[0088]
[0089] where d represents Figure 9 the distance from C to X in GPS , that is, the absolute distance from the shooting point of the first to-be-processed trajectory image to the reference object, which is the depth of field estimation result corresponding to the reference object in the first to-be-processed trajectory image. d GPS is the positioning distance difference between two trajectory images, and x 1 is the relative depth estimation of the reference object in the first to-be-processed trajectory image. x 2 is the relative depth estimation of the reference object in the second image. For example, in a specific embodiment, the estimated depth of the target reference object in Figure A is 0.2; the estimated depth of the same target reference object in Figure B is 0.1; and the GPS distance difference between Figure A and Figure B is 10 meters. Then the estimated absolute distance from Figure A to the target object is: (0.2 * 10) / (0.2 - 0.1) = 20 meters. In this embodiment, by sequentially performing object detection, semantic segmentation, and depth of field recognition processing on the to-be-processed trajectory image, the depth of field information corresponding to the reference object in the to-be-processed trajectory image can be effectively identified, thereby improving the accuracy of trajectory alignment.
[0090] In one embodiment, as Figure 10 shown, step 207 includes:
[0091] Step 1001, obtaining adjacent images in the to-be-processed trajectory image based on the initial alignment position information.
[0092] Step 1003, identifying the depth estimation regions in the adjacent images according to the semantic segmentation result, and obtaining the depth of field estimation results within the depth estimation regions.
[0093] Step 1005: Identify the object detection bounding boxes and the corresponding object types in the depth estimation region based on the object detection results.
[0094] Step 1007: Obtain the relative distance between adjacent image pairs based on the object detection bounding boxes, the corresponding object types of the object detection bounding boxes, and the depth of field estimation results within the depth estimation region.
[0095] Among them, adjacent images refer to two pictures with similar positioning distances and similar azimuth angles. Only when the positioning distances are similar and the azimuth angles are similar can the same reference objects appear in the images, so that the trajectory images can be aligned based on the same reference objects. The depth estimation region refers to other regions in the trajectory image to be processed except for the relative distance interference region. The rectangular box corresponding to the reference object is also located within this depth estimation region. Therefore, the relative distance can be estimated based on the depth of field estimation results corresponding to the pixel points within the depth estimation region, and the pixel points in other parts of the trajectory image to be processed can be ignored. The depth estimation region in adjacent images specifically refers to the region obtained by taking the intersection of two adjacent images. Object detection specifically detects the region where the reference object is located and the type of the reference object. Here, the type of the reference object can specifically refer to the types of road signs, including traffic lights, traffic signs, and traffic cameras, etc. The relative distance can be estimated by combining the object detection bounding boxes and the corresponding object types of the object detection bounding boxes.
[0096] Specifically, when calculating the relative distance, the distance can be estimated based on the same reference objects in adjacent images of the trajectory image to be processed. Therefore, it is necessary to first find the adjacent images in all the trajectory images to be processed through the initial alignment position information, and then determine which regions in these two adjacent images can be used to calculate the relative distance based on the semantic segmentation results. Identify the object detection bounding boxes and the corresponding object types in the depth estimation region according to the object detection results, and determine the coincidence degree of the detected objects in the two images. Obtain the relative distance between adjacent image pairs based on the object detection bounding boxes, the corresponding object types of the object detection bounding boxes, and the depth of field estimation results corresponding to each pixel point within the depth estimation region. In this embodiment, by combining the object detection bounding boxes corresponding to the reference objects in adjacent images and the corresponding depth of field estimation results to estimate the relative distance, the distance between adjacent images can be calculated more accurately, ensuring the effectiveness of the alignment process of the trajectory image to be processed.
[0097] In one embodiment, as Figure 11 shown, before step 1001, it further includes:
[0098] Step 1102: Obtain the azimuth angle information corresponding to the trajectory image to be processed.
[0099] Step 1104: Determine adjacent image pairs in the to-be-processed trajectory image according to the initial alignment position information and the azimuth angle information.
[0100] The azimuth angle specifically refers to the shooting angle corresponding to the to-be-processed trajectory image. Specifically, only the to-be-processed trajectory images with similar positions and shooting angles will capture the same reference object. Therefore, when identifying adjacent images, in addition to considering the initial alignment position information, the azimuth angle information also needs to be considered. Otherwise, for two images with completely opposite shooting angles, even if the shooting locations are the same, the detected objects identified in the images will not be the same. Therefore, when it is necessary to identify adjacent images that can detect the same reference object, the adjacent image pairs in the to-be-processed trajectory image can be determined according to the initial alignment position information and the azimuth angle information. Only when the difference between the initial alignment positions is less than the preset position difference threshold and the difference between the azimuth angles is less than the preset azimuth angle threshold, will the two to-be-processed trajectory images be identified as adjacent images. And when there are multiple trajectory images in a to-be-processed trajectory image whose differences between the initial alignment positions are less than the preset position difference threshold and the differences between the azimuth angles are less than the preset azimuth angle threshold, the differences can be normalized, and the trajectory image with the smallest sum of differences can be used as the adjacent image of the to-be-processed trajectory image. In this embodiment, by combining the azimuth angle information with the initial alignment position information to identify adjacent images in the to-be-processed trajectory image, the accuracy of adjacent image recognition can be effectively guaranteed, thus ensuring the effect of trajectory alignment.
[0101] In one embodiment, as Figure 12 shown, step 1007 includes:
[0102] Step 1201: Determine the background area and the detection frame area in the depth estimation area according to the target detection frame in the depth estimation area.
[0103] Step 1203: Obtain the background distance difference based on the difference in the depth of field estimation results corresponding to the background areas between adjacent image pairs.
[0104] Step 1205: Obtain the detection frame distance difference based on the difference in the depth of field estimation results corresponding to the detection frame areas between adjacent image pairs and the target type corresponding to the target detection frame.
[0105] Step 1207: Obtain the relative distance between adjacent image pairs according to the background distance difference and the detection frame distance difference.
[0106] Among them, the detection box area refers to the area covered by the detection boxes in two adjacent images, which here refers to the union of the areas where the detection boxes of the two images are located, while the background area is the other area in the depth estimation area except the detection box area. The background distance difference estimates the relative distance between the two images by combining the difference in the depth of field in the background area of two adjacent images, while the detection box distance difference estimates the relative distance between the two images by combining the difference in the depth of field in the detection box area of two adjacent images. When calculating, certain weights can be assigned to the two respectively, and then the relative distance is estimated by combining the background distance difference and the detection box distance difference to ensure the accuracy of the distance estimation.
[0107] Specifically, after identifying the same reference object, the relative distance between the two to-be-processed trajectory images can be estimated based on the depth of field of the reference object in different to-be-processed trajectory images and by combining the background content in the two images. When estimating, first determine the background area and the detection box area in the depth estimation area according to the target detection box in the depth estimation area, and then estimate the distance difference between two adjacent images by combining the depth of field distance differences corresponding to the two. The specific formula for the relative distance is as follows:
[0108]
[0109] Among them, D refers to the relative distance between the two trajectory images, where α is the weight of the background points, usually taking a very small value. β is the weight of the detection box, usually taking a relatively large value. The value of the detection box intersection refers to the number of the same type of detection boxes detected in the two trajectory images, that is, the number of the identified same reference objects; the value of the detection box union refers to the union of all detection boxes in the two trajectory images. x i refers to the depth value of the pixel point in the first image, and y i refers to the depth value of the pixel point in the second image. k j is the tolerance coefficient, and the pixel points with a depth estimation difference less than k j are regarded as having the same depth. h and w respectively refer to the height and width of the trajectory image. In this embodiment, by combining the depth estimation results of the detection box area and the background area in adjacent images to estimate the distance between adjacent images, the distance between adjacent images can be calculated more accurately, ensuring the effectiveness of the alignment processing of the to-be-processed trajectory images.
[0110] In one of the embodiments, as Figure 13 shown, step 209 includes:
[0111] Step 1302, according to the relative distance between adjacent image pairs, construct a matrix grid corresponding to two trajectories in the to-be-processed trajectory images through dynamic time warping.
[0112] Step 1304: Solve the shortest path corresponding to the matrix grid through the dynamic programming algorithm, and use the shortest path only as the offset between the two trajectories.
[0113] Step 1306: Align the two trajectories in the processed trajectory image according to the shortest path.
[0114] Among them, dynamic time warping, namely the Dynamic Time Warping algorithm, is a method for studying the alignment problem of sequence information. It is mainly used in template matching, such as in isolated word speech recognition (identifying whether two pieces of speech represent the same word), gesture recognition, data mining, and information retrieval. In this application, the dynamic time warping algorithm is applied to alignment estimation optimization to find the most matching image pairs between two trajectories, thereby achieving the alignment between the two trajectories. Through dynamic time warping, a grid between the two trajectories can be constructed. The dynamic programming algorithm is a branch of operations research, which is a process of solving the optimization of decision-making processes and is an algorithm with polynomial time complexity. After constructing the matrix grid corresponding to the two trajectories in the trajectory image to be processed through dynamic time warping, in order to solve the dynamic time warping problem, a recursive derivation formula can be constructed through the dynamic programming algorithm, and the optimal solution of the dynamic time warping problem can be obtained by solving the recursive derivation formula, so as to obtain the shortest path corresponding to the matrix grid for trajectory alignment.
[0115] Specifically, in this application, the dynamic time warping algorithm commonly used in speech recognition is referred to. It describes the time correspondence relationship between the test template and the reference template with a time warping function W(n) that meets certain conditions, and solves the warping function corresponding to the minimum cumulative distance when the two templates are matched. In speech recognition, dynamic time warping is used to judge the similarity of two pairs of speech sequences. We apply the dynamic time warping algorithm to alignment estimation optimization. The matrix grid corresponding to the trajectory is constructed through the dynamic time warping algorithm, and the matrix grid is solved through dynamic programming, so as to determine the shortest path corresponding to the matrix grid, and use the shortest path as the offset between the two trajectories, thereby achieving the alignment of the trajectories. In one embodiment, as Figure 14 shown, trajectory X contains nine trajectory images to be processed from x 1 to x 9 , and trajectory Y contains seven trajectory images to be processed from y 1 to y 7 . Based on the two trajectories, a 7*9 matrix grid can be constructed. The matrix element (i,j) in the matrix grid represents the distance d(x i ,y j ) between two points x i and y j)(That is, the relative distance between each trajectory point of trajectory X and each trajectory point of trajectory Y. The smaller the distance, the higher the similarity. Here, the order is not considered for now), generally using the Euclidean distance, d(x i ,y j ) = (x i -y j ) 2 (which can also be understood as the distortion). Each matrix element (i, j) represents the alignment of points x i ,y j . The dynamic programming algorithm can be reduced to finding a path through several grid points in this grid. The grid points passed by the path are the alignment points for calculating the two sequences. The recursive derivation formula of the dynamic programming algorithm is
[0116] r(i, j) = d(x i ,y j ) + min{(i - 1, j), (i - 1, j - 1), (i, j - 1)}
[0117] where r(i, j) represents the cumulative distance. Starting from the point (0, 0) to match the two sequences X and Y, at each point, the distances calculated for all previous points will be accumulated. After reaching the end point (9, 7), this cumulative distance is the final total distance mentioned above, that is, the offset between sequences X and Y. This dynamic programming can obtain the optimal solution and has a low time complexity, only O(n * m), where n and m refer to the lengths of the two trajectories to be aligned. The alignment effect of the two trajectories can be specifically referred to Figure 15 as shown. In this embodiment, by using the dynamic time warping algorithm to calculate the offset between trajectories and thus perform the alignment process of trajectories, the time complexity of the alignment calculation process can be effectively reduced, and the operation efficiency of the alignment process can be improved.
[0118] This application also provides an application scenario that applies the above trajectory image alignment method. Specifically, the application of the trajectory image alignment method in this application scenario is as follows:
[0119] When the user is performing the task of automatic generation of high-precision maps, some road trajectory images need to be collected as references. When collecting road trajectory images, alignment processing needs to be performed between different road trajectory images to ensure that different road trajectory images can be obtained at the same trajectory point. Before performing trajectory alignment, first determine whether the positioning of these road trajectory images is accurate enough. When the positioning accuracy is high, alignment can be directly performed. When the positioning accuracy is poor, the trajectory image alignment method of this application needs to be used for trajectory alignment. The overall process of trajectory alignment in this application can be referred to Figure 16As shown, when aligning road trajectory images through the trajectory image alignment method of the present application, it is necessary to first obtain two trajectory images on the same road, and then use the trajectory images on the two trajectories as the trajectory images to be processed for trajectory alignment. When performing trajectory alignment, first obtain the road network line and positioning information corresponding to the trajectory images to be processed. According to the positioning information corresponding to the trajectory images to be processed, the trajectory points corresponding to the trajectory images to be processed can be determined; project the trajectory points onto the tangent direction of the road network line to obtain the projection positions corresponding to the trajectory points; according to the projection positions corresponding to each trajectory image to be processed, obtain the initial alignment position information corresponding to the trajectory images to be processed. Complete the initial alignment of the trajectory images to be processed, and project the trajectory points not on the road onto the road trajectory. Then, it is necessary to perform content analysis and processing on each trajectory image to be processed through computer vision technology, specifically including: obtaining the target detection results corresponding to the reference objects in the trajectory images to be processed; determining the relative distance interference regions in the trajectory images to be processed through semantic segmentation technology; based on the determined relative distance interference regions, obtaining the depth estimation regions corresponding to the trajectory images to be processed; obtaining the absolute depth map corresponding to the depth estimation regions through the principle of pinhole imaging; obtaining the depth of field estimation results corresponding to the reference objects in the trajectory images to be processed through the absolute depth map. Then, based on the results of content analysis, estimate the relative distance between adjacent image pairs. First, obtain the adjacent images in the trajectory images to be processed based on the initial alignment position information; identify the same reference objects in the adjacent images according to the target detection results; obtain the relative distance between the adjacent image pairs according to the depth of field estimation results corresponding to the same reference objects. Among them, the adjacent images can be specifically determined according to the initial alignment position information and the azimuth information. Finally, according to the relative distance between the adjacent image pairs, construct a matrix grid corresponding to the two trajectories in the trajectory images to be processed through dynamic time warping; solve the shortest path corresponding to the matrix grid through the dynamic programming algorithm, and use the shortest path only as the offset between the two trajectories; align the two trajectories in the processed trajectory images according to the shortest path. Then, based on the aligned trajectory images, perform subsequent high-precision map automatic generation tasks.
[0120] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown sequentially according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0121] Based on the same inventive concept, an embodiment of the present application further provides a trajectory image alignment device for implementing the trajectory image alignment method involved above. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the trajectory image alignment device provided below can refer to the limitations on the trajectory image alignment method in the above text, and will not be repeated here.
[0122] In one embodiment, as Figure 17 shown, a trajectory image alignment device is provided, including:
[0123] An image acquisition module 1702, configured to acquire a trajectory image to be processed.
[0124] An initial alignment module 1704, configured to perform an initial alignment on the trajectory image to be processed according to the road network line and positioning information corresponding to the trajectory image to be processed, and obtain the initial alignment position information corresponding to the trajectory image to be processed.
[0125] A content analysis module 1706, configured to perform content analysis processing on the trajectory image to be processed, and obtain the target detection result, semantic segmentation result, and depth of field estimation result of the reference object in the trajectory image to be processed.
[0126] A relative distance calculation module 1708, configured to determine the relative distance between adjacent image pairs in the trajectory image to be processed based on the initial alignment position information, target detection result, semantic segmentation result, and depth of field estimation result.
[0127] An image alignment module 1710, configured to determine the offset between trajectories in the trajectory image to be processed according to the relative distance between adjacent image pairs, and perform alignment processing on the trajectory image to be processed according to the offset.
[0128] In one of the embodiments, the initial alignment module 1704 is specifically configured to: determine the trajectory points corresponding to the trajectory image to be processed according to the positioning information corresponding to the trajectory image to be processed; project the trajectory points onto the tangent direction of the road network line to obtain the projection positions corresponding to the trajectory points; and obtain the initial alignment position information corresponding to the trajectory image to be processed according to the projection positions corresponding to each trajectory image to be processed.
[0129] In one embodiment, the content parsing module 1706 is specifically configured to: obtain the object detection result corresponding to the reference object in the to-be-processed trajectory image through object detection technology; determine the relative distance interference region in the to-be-processed trajectory image through semantic segmentation technology, and obtain the semantic segmentation result corresponding to the to-be-processed trajectory image based on the determined relative distance interference region; obtain the absolute depth map corresponding to the to-be-processed trajectory image through the principle of pinhole imaging, and obtain the depth of field estimation result corresponding to the to-be-processed trajectory image through the absolute depth map.
[0130] In one embodiment, the relative distance calculation module 1708 is specifically configured to: obtain adjacent images in the to-be-processed trajectory image based on the initial alignment position information; identify the depth estimation region in the adjacent images according to the semantic segmentation result, and obtain the depth of field estimation result within the depth estimation region; identify the object detection frame and the object type corresponding to the object detection frame in the depth estimation region according to the object detection result; obtain the relative distance between adjacent image pairs according to the object detection frame, the object type corresponding to the object detection frame, and the depth of field estimation result within the depth estimation region.
[0131] In one embodiment, the relative distance calculation module 1708 is further configured to: obtain the azimuth information corresponding to the to-be-processed trajectory image; determine adjacent image pairs in the to-be-processed trajectory image according to the initial alignment position information and the azimuth information.
[0132] In one embodiment, the relative distance calculation module 1708 is further configured to: determine the background region and the detection frame region in the depth estimation region according to the object detection frame in the depth estimation region; obtain the background distance difference based on the difference in the depth of field estimation results corresponding to the background regions between adjacent image pairs; obtain the detection frame distance difference based on the difference in the depth of field estimation results corresponding to the detection frame regions between adjacent image pairs and the object type corresponding to the object detection frame; obtain the relative distance between adjacent image pairs according to the background distance difference and the detection frame distance difference.
[0133] In one embodiment, the image alignment module 1710 is specifically configured to: construct a matrix grid corresponding to two trajectories in the to-be-processed trajectory image through dynamic time warping according to the relative distance between adjacent image pairs; solve the shortest path corresponding to the matrix grid through a dynamic programming algorithm, and use the shortest path only as the offset between the two trajectories; align the two trajectories in the processed trajectory image according to the shortest path.
[0134] Each module in the above trajectory image alignment device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0135] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in Figure 18 . The computer device includes a processor, a memory, and a network interface connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store trajectory image data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for aligning trajectory images.
[0136] Those skilled in the art can understand that Figure 18 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.
[0137] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented.
[0138] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0139] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.
[0140] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties.
[0141] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0142] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.
[0143] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for aligning trajectory images, characterized in that, the method includes: Obtain the trajectory image to be processed; According to the road network line and positioning information corresponding to the trajectory image to be processed, perform initial alignment on the trajectory image to be processed, and obtain the initial alignment position information corresponding to the trajectory image to be processed; Perform content analysis processing on the trajectory image to be processed, and obtain the target detection result, semantic segmentation result, and depth of field estimation result corresponding to the reference object in the trajectory image to be processed; Based on the initial alignment position information, target detection result, semantic segmentation result, and depth of field estimation result, determine the relative distance between adjacent image pairs in the trajectory image to be processed. The relative distance between adjacent image pairs includes the distance between adjacent different trajectory points in the same trajectory and the distance between adjacent trajectory points in two different trajectories; Determine the offset between trajectories in the trajectory image to be processed according to the relative distance between adjacent image pairs, and perform alignment processing on the trajectory image to be processed according to the offset.
2. The method according to claim 1, characterized in that, the performing initial alignment on the trajectory image to be processed according to the road network information and positioning information corresponding to the trajectory image to be processed, and obtaining the initial alignment position information corresponding to the trajectory image to be processed includes: According to the positioning information corresponding to the trajectory image to be processed, determine the trajectory points corresponding to the trajectory image to be processed; Project the trajectory points onto the tangent direction of the road network line to obtain the projection positions corresponding to the trajectory points; According to the projection positions corresponding to each trajectory image to be processed, obtain the initial alignment position information corresponding to the trajectory image to be processed.
3. The method according to claim 1, characterized in that, the performing content analysis processing on the trajectory image to be processed, and obtaining the target detection result, semantic segmentation result, and depth of field estimation result corresponding to the reference object in the trajectory image to be processed includes: Through target detection technology, obtain the target detection result corresponding to the reference object in the trajectory image to be processed; Through semantic segmentation technology, determine the relative distance interference region in the trajectory image to be processed, and based on the determined relative distance interference region, obtain the semantic segmentation result corresponding to the trajectory image to be processed; Obtain the absolute depth map corresponding to the trajectory image to be processed through the principle of pinhole imaging, and through the absolute depth map, obtain the depth of field estimation result corresponding to the trajectory image to be processed.
4. The method according to claim 1, characterized in that, the determining the relative distance between adjacent image pairs in the trajectory image to be processed based on the initial alignment position information, target detection result, semantic segmentation result, and depth of field estimation result includes: Based on the initial alignment position information, obtain adjacent images in the trajectory image to be processed; According to the semantic segmentation result, identify the depth estimation region in the adjacent images, and obtain the depth of field estimation result within the depth estimation region; According to the target detection result, identify the target detection frame and the target type corresponding to the target detection frame in the depth estimation region; Obtain the relative distance between the adjacent image pairs according to the target detection frame, the target type corresponding to the target detection frame, and the depth of field estimation result within the depth estimation region.
5. The method according to claim 4, wherein, the obtaining of the adjacent images in the to-be-processed trajectory image based on the initial alignment position information includes: obtaining the azimuth information corresponding to the to-be-processed trajectory image; determining the adjacent images in the to-be-processed trajectory image according to the initial alignment position information and the azimuth information.
6. The method according to claim 4, wherein, the obtaining of the relative distance between the adjacent image pairs according to the target detection frame, the target type corresponding to the target detection frame, and the depth of field estimation result within the depth estimation region includes: determining the background region and the detection frame region in the depth estimation region according to the target detection frame in the depth estimation region; obtaining the background distance difference based on the difference in the depth of field estimation results corresponding to the background regions between the adjacent image pairs; obtaining the detection frame distance difference based on the difference in the depth of field estimation results corresponding to the detection frame regions between the adjacent image pairs and the target type corresponding to the target detection frame; obtaining the relative distance between the adjacent image pairs according to the background distance difference and the detection frame distance difference.
7. The method according to claim 1, wherein, the determining of the offset between the trajectories in the to-be-processed trajectory image according to the relative distance between the adjacent image pairs, and the alignment processing of the to-be-processed trajectory image according to the offset includes: constructing a matrix grid corresponding to two trajectories in the to-be-processed trajectory image by dynamic time warping according to the relative distance between the adjacent image pairs; solving the shortest path corresponding to the matrix grid by a dynamic programming algorithm, and taking the shortest path as the offset between the two trajectories only; performing alignment processing on the two trajectories in the processed trajectory image according to the shortest path.
8. A trajectory image alignment device, wherein, the device includes: an image acquisition module, configured to acquire a to-be-processed trajectory image; an initial alignment module, configured to perform initial alignment on the to-be-processed trajectory image according to the road network line and positioning information corresponding to the to-be-processed trajectory image, and acquire the initial alignment position information corresponding to the to-be-processed trajectory image; a content analysis module, configured to perform content analysis processing on the to-be-processed trajectory image, and acquire the target detection result, semantic segmentation result, and depth of field estimation result corresponding to the reference object in the to-be-processed trajectory image; a relative distance calculation module, configured to determine the relative distance between adjacent image pairs in the to-be-processed trajectory image based on the initial alignment position information, target detection result, semantic segmentation result, and depth of field estimation result, and the relative distance between the adjacent image pairs includes the distance between adjacent different trajectory points in the same trajectory and the distance between adjacent trajectory points in two different trajectories; An image alignment module, configured to determine the offset between trajectories in the to-be-processed trajectory image according to the relative distance between adjacent image pairs, and perform alignment processing on the to-be-processed trajectory image according to the offset.
9. The apparatus according to claim 8, wherein, the initial alignment module is specifically configured to: determine the trajectory points corresponding to the to-be-processed trajectory image according to the positioning information corresponding to the to-be-processed trajectory image; project the trajectory points onto the tangent direction of the road network line to obtain the projection positions corresponding to the trajectory points; obtain the initial alignment position information corresponding to the to-be-processed trajectory image according to the projection positions corresponding to each to-be-processed trajectory image.
10. The apparatus according to claim 8, wherein, the content analysis module is specifically configured to: obtain the target detection result of the reference object in the to-be-processed trajectory image through target detection technology; determine the relative distance interference region in the to-be-processed trajectory image through semantic segmentation technology, and obtain the semantic segmentation result corresponding to the to-be-processed trajectory image based on the determined relative distance interference region; obtain the absolute depth map corresponding to the to-be-processed trajectory image through the principle of pinhole imaging, and obtain the depth of field estimation result corresponding to the to-be-processed trajectory image through the absolute depth map.
11. The apparatus according to claim 8, wherein, the relative distance calculation module is specifically configured to: obtain adjacent images in the to-be-processed trajectory image based on the initial alignment position information; identify the depth estimation region in the adjacent images according to the semantic segmentation result, and obtain the depth of field estimation result within the depth estimation region; identify the target detection frame and the target type corresponding to the target detection frame in the depth estimation region according to the target detection result; obtain the relative distance between the adjacent image pairs according to the target detection frame, the target type corresponding to the target detection frame, and the depth of field estimation result within the depth estimation region.
12. The apparatus according to claim 11, wherein, the relative distance calculation module is further configured to: obtain the azimuth information corresponding to the to-be-processed trajectory image; determine adjacent images in the to-be-processed trajectory image according to the initial alignment position information and the azimuth information.
13. The apparatus according to claim 11, wherein, the relative distance calculation module is further configured to: determine the background region and the detection frame region in the depth estimation region according to the target detection frame within the depth estimation region; obtain the background distance difference based on the difference between the depth of field estimation results corresponding to the background regions between the adjacent image pairs; obtain the detection frame distance difference based on the difference between the depth of field estimation results corresponding to the detection frame regions between the adjacent image pairs and the target type corresponding to the target detection frame; obtain the relative distance between the adjacent image pairs according to the background distance difference and the detection frame distance difference.
14. The apparatus according to claim 8, wherein, The image alignment module is specifically configured to: according to the relative distance between adjacent image pairs, construct a matrix grid corresponding to two trajectories in the to-be-processed trajectory image through dynamic time warping; solve the shortest path corresponding to the matrix grid through a dynamic programming algorithm, and use the shortest path only as the offset between the two trajectories; Align the two trajectories in the processed trajectory image according to the shortest path.
15. A computer device, including a memory and a processor, where the memory stores a computer program, characterized in that, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
16. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
17. A computer program product, including a computer program, characterized in that, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Map data architecture for vehicle computer system
EP1111338A2
KR20210111052A