Image Processing Method, Apparatus, Device, Medium, and Computer Program Product
Through unsupervised training, the adjacent images are reconstructed using image sequences and vehicle pose information, solving the problem of high training costs of traditional deep learning networks and achieving efficient depth prediction.
Patent Information
- Application Number
- CN202111548086.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-12-17
AI Technical Summary
The traditional deep learning network training method requires a lot of manpower to mark image depth information, resulting in high training costs.
By acquiring sample on-vehicle images in the image sequence, predicting pixel point depth information using the depth prediction network, generating a three-dimensional point cloud, and reconstructing adjacent images based on vehicle position information, and unsupervised training is performed using image differences.
It reduces the training cost of deep prediction networks, improves training efficiency, and improves the accuracy of deep prediction.
Smart Images

Figure CN114119757B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to artificial intelligence technology, and more particularly to the field of vehicle networking technology, and in particular to an image processing method, apparatus, device, medium, and computer program product. Background Art
[0002] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0003] With the development of artificial intelligence technology, predicting the depth information of an image through a deep learning network has become the mainstream method for obtaining the depth information of an image. In traditional technologies, a deep learning network is usually trained in a supervised training manner, that is, an image labeled with depth information is used as training data to train the deep learning network. However, the current training method of the deep learning network requires a large amount of manpower to implement the annotation of the depth information of the image, resulting in a high training cost of the deep learning network. Summary of the Invention
[0004] Based on this, it is necessary to provide an image processing method, apparatus, device, medium, and computer program product that can reduce the training cost of the depth prediction network for the above technical problems.
[0005] An image processing method, the method includes:
[0006] Obtain an image sequence; the image sequence includes a plurality of sample vehicle-mounted images collected in sequence;
[0007] Respectively use each sample vehicle-mounted image in the image sequence as the current sample vehicle-mounted image, and input the current sample vehicle-mounted image into a depth prediction network to be trained, and predict the depth information of each pixel point in the current sample vehicle-mounted image;
[0008] Generate a current three-dimensional point cloud of the current sample vehicle-mounted image based on the depth information;
[0009] Obtain vehicle pose information determined based on the current sample vehicle-mounted image and the next adjacent sample vehicle-mounted image from a vehicle pose sensor;
[0010] Reconstruct the next adjacent vehicle-mounted image of the current sample vehicle-mounted image based on the current three-dimensional point cloud and vehicle pose information;
[0011] Train the depth prediction network based on the difference between the adjacent sample vehicle-mounted image and the reconstructed adjacent vehicle-mounted image; The trained depth prediction network is used to predict the pixel depth of the target vehicle-mounted image.
[0012] An image processing device, the device includes:
[0013] An acquisition module, configured to acquire an image sequence; The image sequence includes a plurality of sample vehicle-mounted images collected in sequence;
[0014] A prediction module, configured to use each sample vehicle-mounted image in the image sequence as the current sample vehicle-mounted image respectively, and input the current sample vehicle-mounted image into the depth prediction network to be trained, and predict the depth information of each pixel point in the current sample vehicle-mounted image;
[0015] A generation module, configured to generate the current three-dimensional point cloud of the current sample vehicle-mounted image based on the depth information;
[0016] The acquisition module is further configured to acquire vehicle pose information determined based on the current sample vehicle-mounted image and the next adjacent sample vehicle-mounted image from a vehicle pose sensor;
[0017] A reconstruction module, configured to reconstruct the next adjacent vehicle-mounted image of the current sample vehicle-mounted image based on the current three-dimensional point cloud and vehicle pose information;
[0018] A training module, configured to train the depth prediction network based on the difference between the adjacent sample vehicle-mounted image and the reconstructed adjacent vehicle-mounted image; The trained depth prediction network is used to predict the pixel depth of the target vehicle-mounted image.
[0019] In one embodiment, the current sample vehicle-mounted image is collected by a vehicle-mounted camera; The depth information includes the depth coordinates of each pixel point in the current sample vehicle-mounted image in the camera coordinate system; The generation module is further configured to acquire the plane coordinates of each pixel point in the current sample vehicle-mounted image in the camera coordinate system; The camera coordinate system is a coordinate system established with the optical center of the vehicle-mounted camera as the origin; Based on the depth coordinates and the plane coordinates of each pixel point, generate the current three-dimensional point cloud of the current sample vehicle-mounted image.
[0020] In one embodiment, the generating module is further configured to determine the focal coordinates of the focus of the vehicle-mounted camera in the camera coordinate system; determine the current pixel coordinates of each pixel point in the current sample vehicle-mounted image in the current pixel coordinate system; the current pixel coordinate system is a pixel coordinate system established based on the current sample vehicle-mounted image; and determine the planar coordinates of each pixel point in the camera coordinate system according to the focal coordinates and the current pixel coordinates of each pixel point in the current pixel coordinate system.
[0021] In one embodiment, the current sample vehicle-mounted image is acquired by a vehicle-mounted camera; the reconstructing module is further configured to determine the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image based on the current three-dimensional point cloud and the vehicle pose information; and reconstruct the next adjacent vehicle-mounted image of the current sample vehicle-mounted image according to the adjacent three-dimensional point cloud and the internal parameters of the vehicle-mounted camera.
[0022] In one embodiment, the vehicle pose information includes the offset distance and the rotation angle of the vehicle from the current position to the next adjacent position; the current position is the position where the vehicle is located when the current sample vehicle-mounted image is acquired; the next adjacent position is the position where the vehicle is located when the next adjacent sample vehicle-mounted image is acquired; and the reconstructing module is further configured to adjust each point in the current three-dimensional point cloud respectively according to the offset distance and the rotation angle to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image.
[0023] In one embodiment, the reconstructing module is further configured to construct an offset matrix based on the offset distance; construct a rotation matrix based on the rotation angle; offset each point in the current three-dimensional point cloud according to the offset matrix, and rotate each point in the current three-dimensional point cloud according to the rotation matrix to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image.
[0024] In one embodiment, the internal parameters include the focal coordinates of the focus of the vehicle-mounted camera in the camera coordinate system; the reconstructing module is further configured to construct a reconstruction matrix based on the focal coordinates; convert the three-dimensional coordinates of each point in the adjacent three-dimensional point cloud into two-dimensional coordinates based on the reconstruction matrix; and construct the next adjacent vehicle-mounted image of the current sample vehicle-mounted image based on the two-dimensional coordinates of each point.
[0025] In one embodiment, the sample vehicle-mounted image includes a sample vehicle-mounted monocular image; the obtaining module is further configured to acquire a plurality of sample vehicle-mounted monocular images through a vehicle-mounted monocular camera; and generate the image sequence according to the chronological order of the acquisition times corresponding to the plurality of sample vehicle-mounted monocular images.
[0026] In one embodiment, the apparatus further includes:
[0027] A navigation module, which is used to detect road elements in a target vehicle-mounted image, so as to identify road elements from the target vehicle-mounted image; based on the trained depth prediction network, predict the depth information of each pixel point in the road elements; based on the depth information of each pixel point in the road elements, generate navigation information of the target road elements.
[0028] In one embodiment, the target vehicle-mounted image is collected by a vehicle-mounted camera arranged on the current vehicle; the device further includes:
[0029] An early warning module, which is used to detect vehicles in the target vehicle-mounted image, so as to identify the image area corresponding to the vehicle in front from the target vehicle-mounted image; obtain the depth information of each pixel point in the image area; the depth information of each pixel point in the image area is predicted by the trained depth prediction network; based on the depth information of each pixel point in the image area, determine the relative distance between the current vehicle and the vehicle in front; perform collision warning based on the relative distance.
[0030] In one embodiment, the device further includes:
[0031] A writing module, which is used to detect road elements in the target vehicle-mounted image, so as to identify road elements from the target vehicle-mounted image; obtain the depth information of each pixel point in the road elements; the depth information of each pixel point in the road elements is predicted by the trained depth prediction network; based on the depth information of each pixel point in the road elements, write the target road elements into the vehicle-mounted navigation map.
[0032] A computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0033] Obtain an image sequence; the image sequence includes a plurality of sample vehicle-mounted images collected in sequence;
[0034] Respectively take each sample vehicle-mounted image in the image sequence as the current sample vehicle-mounted image, and input the current sample vehicle-mounted image into the depth prediction network to be trained, and predict the depth information of each pixel point in the current sample vehicle-mounted image;
[0035] Based on the depth information, generate the current three-dimensional point cloud of the current sample vehicle-mounted image;
[0036] Obtain vehicle pose information determined based on the current sample vehicle-mounted image and the next adjacent sample vehicle-mounted image from a vehicle-mounted pose sensor;
[0037] Reconstruct the next adjacent in-vehicle image of the current sample in-vehicle image based on the current 3D point cloud and vehicle pose information;
[0038] Train the depth prediction network based on the difference between the adjacent sample in-vehicle image and the reconstructed adjacent in-vehicle image; The trained depth prediction network is used to predict the pixel depth of the target in-vehicle image.
[0039] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0040] Obtain an image sequence; The image sequence includes a plurality of sequentially acquired sample in-vehicle images;
[0041] Respectively use each sample in-vehicle image in the image sequence as the current sample in-vehicle image, and input the current sample in-vehicle image into the depth prediction network to be trained, and predict the depth information of each pixel point in the current sample in-vehicle image;
[0042] Generate the current 3D point cloud of the current sample in-vehicle image based on the depth information;
[0043] Obtain vehicle pose information determined based on the current sample in-vehicle image and the next adjacent sample in-vehicle image from an in-vehicle pose sensor;
[0044] Reconstruct the next adjacent in-vehicle image of the current sample in-vehicle image based on the current 3D point cloud and vehicle pose information;
[0045] Train the depth prediction network based on the difference between the adjacent sample in-vehicle image and the reconstructed adjacent in-vehicle image; The trained depth prediction network is used to predict the pixel depth of the target in-vehicle image.
[0046] A computer program product, including a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0047] Obtain an image sequence; The image sequence includes a plurality of sequentially acquired sample in-vehicle images;
[0048] Respectively use each sample in-vehicle image in the image sequence as the current sample in-vehicle image, and input the current sample in-vehicle image into the depth prediction network to be trained, and predict the depth information of each pixel point in the current sample in-vehicle image;
[0049] Generate the current 3D point cloud of the current sample in-vehicle image based on the depth information;
[0050] Obtain the vehicle pose information determined based on the current sample vehicle-mounted image and the next adjacent sample vehicle-mounted image from the vehicle-mounted pose sensor;
[0051] Reconstruct the next adjacent vehicle-mounted image of the current sample vehicle-mounted image based on the current three-dimensional point cloud and the vehicle pose information;
[0052] Train the depth prediction network based on the difference between the adjacent sample vehicle-mounted image and the reconstructed adjacent vehicle-mounted image; The trained depth prediction network is used to predict the pixel depth of the target vehicle-mounted image.
[0053] The above image processing method, device, equipment, medium and computer program product obtain an image sequence including a plurality of sequentially acquired sample vehicle-mounted images, respectively use each sample vehicle-mounted image in the image sequence as the current sample vehicle-mounted image, and input the current sample vehicle-mounted image into the depth prediction network to be trained, and the depth information of each pixel point in the current sample vehicle-mounted image can be predicted. Based on the depth information, the current three-dimensional point cloud of the current sample vehicle-mounted image can be generated. From the vehicle-mounted pose sensor, the vehicle pose information determined based on the current sample vehicle-mounted image and the next adjacent sample vehicle-mounted image can be directly obtained. Based on the current three-dimensional point cloud and the vehicle pose information, the next adjacent vehicle-mounted image of the current sample vehicle-mounted image can be reconstructed. Based on the difference between the adjacent sample vehicle-mounted image and the reconstructed adjacent vehicle-mounted image, the depth prediction network can be directly trained without supervision. Compared with the traditional supervised training method, the present application can directly reconstruct the adjacent vehicle-mounted image corresponding to the adjacent sample vehicle-mounted image through the current sample vehicle-mounted image in the image sequence, the depth information of the current sample vehicle-mounted image, and the vehicle pose information determined based on the current sample vehicle-mounted image and the next adjacent sample vehicle-mounted image, and can directly realize the unsupervised training of the depth prediction network based on the difference between the adjacent sample vehicle-mounted image and the adjacent vehicle-mounted image, without pre-labeling the image depth information, which greatly reduces the training cost of the depth prediction network. Description of the Drawings
[0054] Figure 1 It is an application environment diagram of the image processing method in an embodiment;
[0055] Figure 2 It is a schematic flowchart of the image processing method in an embodiment;
[0056] Figure 3 It is a schematic diagram of the target vehicle-mounted image in an embodiment;
[0057] Figure 4 It is a schematic diagram of the depth image of the target vehicle-mounted image in an embodiment;
[0058] Figure 5Schematic diagram of the driving process of the current vehicle and the vehicle ahead in an embodiment;
[0059] Figure 6 Schematic flow chart of an image processing method in another embodiment;
[0060] Figure 7 Schematic flow chart of an image processing method in yet another embodiment;
[0061] Figure 8 Block diagram of the structure of an image processing device in an embodiment;
[0062] Figure 9 Block diagram of the structure of an image processing device in another embodiment;
[0063] Figure 10 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0064] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0065] The image processing method provided by the present application can be applied to an application environment as Figure 1 shown. Among them, the in-vehicle camera 104 can collect sample in-vehicle images, and the computer device 102 can communicate with the in-vehicle camera 104 to obtain the sample in-vehicle images. Among them, the computer device 102 can be a server or a terminal. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, portable wearable devices and in-vehicle terminals. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The computer device 102 and the in-vehicle camera 104 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here.
[0066] The computer device 102 can obtain an image sequence, which includes a plurality of sample vehicle-mounted images sequentially collected by the vehicle-mounted camera 104 in the vehicle-mounted terminal. Among them, the computer device 102 can directly obtain the sample vehicle-mounted images from the vehicle-mounted camera 104, or the vehicle-mounted camera 104 stores the collected sample vehicle-mounted images in the server, and then the computer device 102 obtains the sample vehicle-mounted images from the server. The computer device 102 can respectively use each sample vehicle-mounted image in the image sequence as the current sample vehicle-mounted image, and input the current sample vehicle-mounted image into the depth prediction network to be trained, and predict the depth information of each pixel point in the current sample vehicle-mounted image. The computer device 102 can generate the current three-dimensional point cloud of the current sample vehicle-mounted image based on the depth information, and obtain the vehicle pose information determined based on the current sample vehicle-mounted image and the next adjacent sample vehicle-mounted image from the vehicle pose sensor. The computer device 102 can reconstruct the next adjacent vehicle-mounted image of the current sample vehicle-mounted image based on the current three-dimensional point cloud and the vehicle pose information, and train the depth prediction network based on the difference between the adjacent sample vehicle-mounted image and the reconstructed adjacent vehicle-mounted image. Among them, the trained depth prediction network can be used to predict the pixel depth of the target vehicle-mounted image.
[0067] It can be understood that the computer device 102 can be a server or the vehicle-mounted terminal itself.
[0068] It should be noted that the image processing method in some embodiments of this application uses artificial intelligence technology. For example, the depth information of each pixel point in the current sample vehicle-mounted image belongs to the depth information predicted using artificial intelligence technology.
[0069] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0070] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement, and further performs graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition. It should be noted that the image processing methods in some embodiments of this application use computer vision technology. For example, the current three-dimensional point cloud of the current sample vehicle-mounted image belongs to the point cloud information generated using computer vision technology, and the next adjacent vehicle-mounted image of the current sample vehicle-mounted image also belongs to the image reconstructed using computer vision technology.
[0071] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. It should be noted that the image processing methods in some embodiments of this application use machine learning. For example, the trained depth prediction network belongs to the neural network trained using machine learning.
[0072] Autonomous driving technology usually includes technologies such as high-precision maps, environmental perception, behavior decision-making, path planning, and motion control. It should be noted that the image processing methods in some embodiments of this application use autonomous driving technology. For example, the vehicle pose sensor in this application can be a sensor installed in an autonomous driving vehicle. The target vehicle-mounted image in this application can be an image collected by a vehicle-mounted camera installed in an autonomous driving vehicle. The predicted pixel depth of the target vehicle-mounted image can be applied to obstacle warning during the autonomous driving process of an autonomous driving vehicle.
[0073] In one embodiment, as Figure 2As shown, an image processing method is provided. This method can be applied to a computer device 102 or to the interaction process between the computer device 102 and a server. In this embodiment, taking the application of this method to Figure 1 the computer device 102 in it as an example for illustration, it includes the following steps:
[0074] Step 202, obtain an image sequence; the image sequence includes a plurality of sample vehicle-mounted images collected in sequence.
[0075] Among them, the sample vehicle-mounted image is a vehicle-mounted image used as training data. It can be understood that the sample vehicle-mounted image is a vehicle-mounted image used to train a depth prediction network. A vehicle-mounted image refers to an image collected by a vehicle-mounted camera installed on a vehicle in a vehicle-mounted scenario. Collecting in sequence means that during the driving of the vehicle, the vehicle-mounted camera of the vehicle sequentially collects images in the order of driving time. An image sequence refers to a sequence including a plurality of sample vehicle-mounted images collected in sequence.
[0076] In one embodiment, an image sequence including a plurality of sample vehicle-mounted images collected in sequence is stored in the server. The computer device can communicate with the server and directly obtain the image sequence from the server.
[0077] In one embodiment, the computer device can be deployed on a vehicle including a vehicle-mounted camera. The computer device can, during the driving of the vehicle, collect a plurality of sample vehicle-mounted images in sequence through the vehicle-mounted camera on the vehicle, and directly generate an image sequence including a plurality of sample vehicle-mounted images collected in sequence based on the collected sample vehicle-mounted images.
[0078] Step 204, respectively use each sample vehicle-mounted image in the image sequence as the current sample vehicle-mounted image, and input the current sample vehicle-mounted image into a depth prediction network to be trained, and predict the depth information of each pixel point in the current sample vehicle-mounted image.
[0079] Among them, the current sample vehicle-mounted image is the sample vehicle-mounted image being currently processed. It can be understood that the image sequence includes a plurality of sample vehicle-mounted images, and the sample vehicle-mounted image being currently processed is the current sample vehicle-mounted image. A depth prediction network is a neural network used to predict the depth information of each pixel point in an image. The depth information of each pixel point in the current sample vehicle-mounted image refers to the distance information between each pixel point and the vehicle-mounted camera that collected the current sample vehicle-mounted image.
[0080] Specifically, the image sequence includes multiple sample vehicle-mounted images. The computer device can respectively use each sample vehicle-mounted image in the image sequence as the current sample vehicle-mounted image and input the current sample vehicle-mounted image into the depth prediction network to be trained. The computer device can perform depth prediction on the current sample vehicle-mounted image through the depth prediction network to be trained to obtain the depth information of each pixel point in the current sample vehicle-mounted image. It can be understood that the computer device can respectively input each sample vehicle-mounted image in the image sequence into the depth prediction network to be trained and perform depth prediction on each sample vehicle-mounted image in the sequentially input image sequence through the depth prediction network to be trained to obtain the depth information of each pixel point in each sample vehicle-mounted image in the image sequence.
[0081] In one embodiment, Figure 3 For the sample vehicle-mounted image, the sample vehicle-mounted image includes a vehicle 301, a person 302, and white clouds 303 in the sky. The computer device can, through the depth prediction network to be trained, predict the depth information of each pixel point in the sample vehicle-mounted image to obtain a depth image of the sample vehicle-mounted image as shown in Figure 4 As can be seen from Figure 4 it, the pixel points in the image area corresponding to the vehicle 301 are closest to the vehicle-mounted camera that captured the sample vehicle-mounted image, while the pixel points in the image areas corresponding to the person 302 and the white clouds 303 in the sky are farther from the vehicle-mounted camera that captured the sample vehicle-mounted image.
[0082] Step 206, generate the current three-dimensional point cloud of the current sample vehicle-mounted image based on the depth information.
[0083] Among them, the three-dimensional point cloud is a three-dimensional data representation method. It can be understood that the three-dimensional point cloud is the three-dimensional coordinate points of each pixel point in the image. The current three-dimensional point cloud is the three-dimensional point cloud corresponding to the current sample vehicle-mounted image.
[0084] Specifically, the computer device can, based on the depth information, convert the two-dimensional coordinates of each pixel point in the current sample vehicle-mounted image into three-dimensional coordinates to obtain the current three-dimensional point cloud of the current sample vehicle-mounted image.
[0085] In one embodiment, the computer device can perform coordinate conversion on the two-dimensional coordinates of each pixel point in the current sample vehicle-mounted image. The computer device can generate the current three-dimensional point cloud of the current sample vehicle-mounted image based on the converted two-dimensional coordinates of the pixel points and the depth information of each pixel point in the current sample vehicle-mounted image. It can be understood that the two-dimensional coordinates of the pixel points after coordinate conversion are the coordinates of two of the three dimensions in the three-dimensional coordinate system, and the depth information of each pixel point in the current sample vehicle-mounted image can be directly used as the coordinate of the third dimension in the three-dimensional coordinate system, thereby obtaining the current three-dimensional point cloud of the current sample vehicle-mounted image.
[0086] Step 208: Obtain vehicle pose information determined based on the current sample vehicle-mounted image and the next adjacent sample vehicle-mounted image from the vehicle-mounted pose sensor.
[0087] Among them, the vehicle-mounted pose sensor (IMU, Inertial Measurement Unit) is a sensor installed in the vehicle and used to obtain vehicle pose information. It can be understood that the vehicle-mounted pose sensor can automatically and accurately obtain the vehicle pose information during the driving process of the vehicle without relying on any pose prediction algorithm. The vehicle pose information is the information about the attitude change of the vehicle in space during the process from collecting the current sample vehicle-mounted image to collecting the next adjacent sample vehicle-mounted image. The next adjacent sample vehicle-mounted image is the next frame of sample vehicle-mounted image adjacent to the current sample vehicle-mounted image in the image sequence, that is, the sample vehicle-mounted image that is actually collected after the current sample vehicle-mounted image and adjacent to the current sample vehicle-mounted image.
[0088] Specifically, a vehicle-mounted pose sensor is deployed in the vehicle. The vehicle-mounted pose sensor deployed in the vehicle can automatically determine the vehicle pose information based on the current sample vehicle-mounted image and the next adjacent sample vehicle-mounted image, and send the vehicle pose information to the computer device. The computer device can receive the vehicle pose information sent by the vehicle-mounted pose sensor.
[0089] Step 210: Reconstruct the next adjacent vehicle-mounted image of the current sample vehicle-mounted image based on the current three-dimensional point cloud and the vehicle pose information.
[0090] Among them, the next adjacent vehicle-mounted image is a reconstructed vehicle-mounted image that has substantially the same image content as the next adjacent sample vehicle-mounted image. Substantially the same image content means that the main content of the image is the same, but there are slight differences in the image details.
[0091] Specifically, the computer device can adjust and process the current three-dimensional point cloud of the current sample vehicle-mounted image based on the vehicle pose information to obtain the processed three-dimensional point cloud. The computer device can reconstruct the next adjacent vehicle-mounted image of the current sample vehicle-mounted image based on the processed three-dimensional point cloud.
[0092] In one embodiment, the computer device can convert the current three-dimensional point cloud of the current sample vehicle-mounted image into a three-dimensional point cloud corresponding to the adjacent sample vehicle-mounted image based on the vehicle pose information. The computer device can reconstruct the next adjacent vehicle-mounted image of the current sample vehicle-mounted image based on the three-dimensional point cloud corresponding to the adjacent sample vehicle-mounted image.
[0093] Step 212: Train the depth prediction network based on the difference between the adjacent sample vehicle-mounted image and the reconstructed adjacent vehicle-mounted image; the trained depth prediction network is used to predict the pixel depth of the target vehicle-mounted image.
[0094] Among them, the target vehicle-mounted image is the vehicle-mounted image serving as the target. It can be understood that the target vehicle-mounted image is the vehicle-mounted image to be predicted collected by the vehicle-mounted camera deployed in the vehicle during the process of completing the training of the depth prediction network and putting it into actual application.
[0095] Specifically, the computer device can respectively extract the image features of the adjacent sample vehicle-mounted images and the image features of the reconstructed adjacent vehicle-mounted images, and determine the difference between the adjacent sample vehicle-mounted images and the reconstructed adjacent vehicle-mounted images based on the image features of the adjacent sample vehicle-mounted images and the image features of the reconstructed adjacent vehicle-mounted images. Thus, the computer device can iteratively train the depth prediction network in the direction of reducing the difference between the adjacent sample vehicle-mounted images and the reconstructed adjacent vehicle-mounted images until the training stop condition is reached, and obtain the trained depth prediction network.
[0096] In one embodiment, the training stop condition can be that the difference between the adjacent sample vehicle-mounted image and the reconstructed adjacent vehicle-mounted image is less than a preset difference threshold, or that the number of iterations reaches a preset number of learning times.
[0097] In the above image processing method, an image sequence including a plurality of sequentially collected sample vehicle-mounted images is obtained, each sample vehicle-mounted image in the image sequence is respectively used as the current sample vehicle-mounted image, and the current sample vehicle-mounted image is input into the depth prediction network to be trained, and the depth information of each pixel point in the current sample vehicle-mounted image can be predicted. Based on the depth information, the current 3D point cloud of the current sample vehicle-mounted image can be generated. From the vehicle pose sensor, the vehicle pose information determined based on the current sample vehicle-mounted image and the next adjacent sample vehicle-mounted image can be directly obtained. Based on the current 3D point cloud and the vehicle pose information, the next adjacent vehicle-mounted image of the current sample vehicle-mounted image can be reconstructed. Based on the difference between the adjacent sample vehicle-mounted image and the reconstructed adjacent vehicle-mounted image, the depth prediction network can be directly trained without supervision. Compared with the traditional supervised training method, the present application can directly reconstruct the adjacent vehicle-mounted image corresponding to the adjacent sample vehicle-mounted image through the current sample vehicle-mounted image in the image sequence, the depth information of the current sample vehicle-mounted image, and the vehicle pose information determined based on the current sample vehicle-mounted image and the next adjacent sample vehicle-mounted image, and can directly realize the unsupervised training of the depth prediction network based on the difference between the adjacent sample vehicle-mounted image and the adjacent vehicle-mounted image, without pre-labeling the image depth information, which greatly reduces the training cost of the depth prediction network.
[0098] At the same time, compared with the relatively expensive unsupervised training method based on binocular image pairs, the unsupervised training method based on the image sequence in the present application can further reduce the training cost of the depth prediction network.
[0099] In addition, based on the unsupervised training method of the image sequence in this application, since accurate vehicle pose information determined based on the current sample vehicle-mounted image and the next adjacent sample vehicle-mounted image can be directly obtained from the vehicle-mounted pose sensor, without the need to obtain it through an additional pose prediction network, that is, there is no need to train an additional pose prediction network, and only the depth prediction network needs to be trained, which reduces the complexity of network training and also improves the depth prediction accuracy of the trained depth prediction network.
[0100] In one embodiment, the current sample vehicle-mounted image is acquired by a vehicle-mounted camera; the depth information includes the depth coordinates of each pixel point in the current sample vehicle-mounted image in the camera coordinate system; based on the depth information, generating the current three-dimensional point cloud of the current sample vehicle-mounted image includes: obtaining the plane coordinates of each pixel point in the current sample vehicle-mounted image in the camera coordinate system; the camera coordinate system is a coordinate system established with the optical center of the vehicle-mounted camera as the origin; based on the depth coordinates and plane coordinates of each pixel point, generating the current three-dimensional point cloud of the current sample vehicle-mounted image.
[0101] Among them, the depth coordinates of each pixel point in the camera coordinate system are the coordinates representing the pixel depth information of each pixel point in the camera coordinate system. The plane coordinates of each pixel point in the camera coordinate system are the coordinates representing the pixel plane information of each pixel point in the camera coordinate system. For example, the camera coordinate system includes the X-axis, Y-axis, and Z-axis, then the plane coordinates of each pixel point in the camera coordinate system include the X value and Y value, which can be expressed as (x, y, 0), and the depth coordinates of each pixel point in the camera coordinate system include the Z value, which can be expressed as (0, 0, z). Thus, based on the depth coordinates (x, y, 0) and plane coordinates (0, 0, z) of each pixel point, the current three-dimensional point cloud (x, y, z) of the current sample vehicle-mounted image is generated.
[0102] Specifically, the computer device can perform coordinate transformation on the coordinates of each pixel point in the current sample vehicle-mounted image to obtain the plane coordinates of each pixel point in the current sample vehicle-mounted image in the camera coordinate system. Furthermore, the computing device can perform coordinate fusion on the depth coordinates and plane coordinates of each pixel point to obtain the current three-dimensional point cloud of the current sample vehicle-mounted image.
[0103] In one embodiment, the computer device can perform coordinate transformation on the coordinates of each pixel point in the current sample vehicle-mounted image based on the internal parameters of the vehicle-mounted camera to obtain the plane coordinates of each pixel point in the current sample vehicle-mounted image in the camera coordinate system.
[0104] In the above embodiment, by obtaining the plane coordinates of each pixel point in the current sample vehicle-mounted image in the camera coordinate system and based on the depth coordinates and plane coordinates of each pixel point, the current three-dimensional point cloud of the current sample vehicle-mounted image can be generated, realizing the conversion of each two-dimensional pixel point in the current sample vehicle-mounted image into three-dimensional point cloud information.
[0105] In one embodiment, obtaining the planar coordinates of each pixel point in the current sample vehicle-mounted image in the camera coordinate system includes: determining the focal coordinates of the focus of the vehicle-mounted camera in the camera coordinate system; determining the current pixel coordinates of each pixel point in the current sample vehicle-mounted image in the current pixel coordinate system; the current pixel coordinate system is a pixel coordinate system established based on the current sample vehicle-mounted image; and determining the planar coordinates of each pixel point in the camera coordinate system according to the focal coordinates and the current pixel coordinates of each pixel point in the current pixel coordinate system.
[0106] Wherein, the current pixel coordinate is the pixel coordinate of each pixel point in the current sample vehicle-mounted image in the current pixel coordinate system.
[0107] Specifically, the computer device can determine the focal coordinates of the focus of the vehicle-mounted camera in the camera coordinate system based on the internal parameters of the vehicle-mounted camera. It can be understood that the internal parameters of the vehicle-mounted camera include the focal coordinates of the focus of the vehicle-mounted camera in the camera coordinate system, and the computer device can directly obtain the focal coordinates of the focus of the vehicle-mounted camera in the camera coordinate system from the internal parameters of the vehicle-mounted camera. The computer device can determine the current pixel coordinates of each pixel point in the current sample vehicle-mounted image in the current pixel coordinate system. Furthermore, the computer device can determine the planar coordinates of each pixel point in the camera coordinate system according to the focal coordinates and the current pixel coordinates of each pixel point in the current pixel coordinate system.
[0108] In one embodiment, the computer device can construct a coordinate transformation matrix of the point cloud according to the focal coordinates. Furthermore, the computer device can perform coordinate transformation on the current pixel coordinates of each pixel point in the current pixel coordinate system based on the constructed coordinate transformation matrix of the point cloud to obtain the planar coordinates of each pixel point in the camera coordinate system.
[0109] In the above embodiment, by determining the focal coordinates of the focus of the vehicle-mounted camera in the camera coordinate system and determining the current pixel coordinates of each pixel point in the current sample vehicle-mounted image in the current pixel coordinate system, and then according to the focal coordinates and the current pixel coordinates of each pixel point in the current pixel coordinate system, the planar coordinates of each pixel point in the camera coordinate system can be determined, and the current pixel coordinates in the two-dimensional coordinate system can be converted into the planar coordinates in the three-dimensional coordinate system.
[0110] In one embodiment, the computer device can construct a pixel matrix based on the current pixel coordinates of each pixel point in the current sample vehicle-mounted image in the current pixel coordinate system. The computer device can generate the current three-dimensional point cloud of the current sample vehicle-mounted image based on the product of the coordinate transformation matrix of the point cloud and the pixel matrix and based on the depth information of each pixel point. Wherein, the pixel matrix is a matrix constructed based on each coordinate value in the current pixel coordinates of each pixel point in the current sample vehicle-mounted image.
[0111] In one embodiment, the current three-dimensional point cloud of the current sample vehicle-mounted image can be calculated by the following formula:
[0112]
[0113] Wherein, are the three-dimensional coordinates of each point in the current three-dimensional point cloud, d is the depth of each pixel point in the current sample vehicle-mounted image, and are the focal lengths of the focal points of the vehicle-mounted camera on the X-axis and Y-axis in the camera coordinate system respectively, are the current pixel coordinates of each pixel point in the current sample vehicle-mounted image in the current pixel coordinate system.
[0114] In one embodiment, the current sample vehicle-mounted image is acquired by a vehicle-mounted camera; based on the current three-dimensional point cloud and vehicle pose information, the next adjacent vehicle-mounted image of the current sample vehicle-mounted image is reconstructed, including: determining the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image based on the current three-dimensional point cloud and vehicle pose information; reconstructing the next adjacent vehicle-mounted image of the current sample vehicle-mounted image according to the adjacent three-dimensional point cloud and the internal parameters of the vehicle-mounted camera.
[0115] Wherein, the adjacent three-dimensional point cloud is the three-dimensional point cloud corresponding to the adjacent sample vehicle-mounted image.
[0116] In one embodiment, the computer device can adjust the current three-dimensional point cloud based on the vehicle pose information to generate the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image. Furthermore, the computer device can perform coordinate transformation on the adjacent three-dimensional point cloud based on the internal parameters of the vehicle-mounted camera, and reconstruct the next adjacent vehicle-mounted image of the current sample vehicle-mounted image based on the transformed coordinates.
[0117] In the above embodiment, based on the current three-dimensional point cloud and vehicle pose information, the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image can be determined, and then according to the adjacent three-dimensional point cloud and the internal parameters of the vehicle-mounted camera, the next adjacent vehicle-mounted image of the current sample vehicle-mounted image can be quickly reconstructed for subsequent training of the depth prediction network.
[0118] In one embodiment, the vehicle pose information includes the offset distance and rotation angle of the vehicle from the current position to the next adjacent position; the current position is the position where the vehicle is located when the current sample vehicle-mounted image is acquired; the next adjacent position is the position where the vehicle is located when the next adjacent sample vehicle-mounted image is acquired; determining the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image based on the current three-dimensional point cloud and vehicle pose information includes: adjusting each point in the current three-dimensional point cloud according to the offset distance and rotation angle respectively to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image.
[0119] Among them, the offset distance is the moving distance of the vehicle from the current position to the next adjacent position. The rotation angle is the angle between the angle corresponding to the vehicle at the current position and the angle corresponding to the vehicle at the next adjacent position.
[0120] In one embodiment, the computer device can offset each point in the current three-dimensional point cloud according to the offset distance, and at the same time, rotate each point in the current three-dimensional point cloud according to the rotation angle, so as to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image.
[0121] In one embodiment, the computer device can first offset each point in the current three-dimensional point cloud according to the offset distance, and then rotate each point in the offset three-dimensional point cloud according to the rotation angle, so as to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image.
[0122] In one embodiment, the computer device can first rotate each point in the current three-dimensional point cloud according to the rotation angle, and then offset each point in the rotated three-dimensional point cloud according to the offset distance, so as to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image.
[0123] In the above embodiments, by respectively adjusting and processing each point in the current three-dimensional point cloud according to the offset distance and the rotation angle, the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image can be obtained quickly and accurately.
[0124] In one embodiment, adjusting and processing each point in the current three-dimensional point cloud according to the offset distance and the rotation angle respectively to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image includes: constructing an offset matrix based on the offset distance; constructing a rotation matrix based on the rotation angle; offsetting each point in the current three-dimensional point cloud according to the offset matrix, and rotating each point in the current three-dimensional point cloud according to the rotation matrix, so as to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image.
[0125] Among them, the offset matrix is a matrix used to control the movement of each point in the current three-dimensional point cloud. The rotation matrix is a matrix used to control the rotation of each point in the current three-dimensional point cloud.
[0126] In one embodiment, the computer device can construct an offset matrix based on the offset distance and construct a rotation matrix based on the rotation angle. The computer device can perform coordinate transformation on the three-dimensional coordinates of each point in the current three-dimensional point cloud based on the constructed offset matrix, and perform coordinate transformation on the three-dimensional coordinates of each point in the current three-dimensional point cloud based on the constructed rotation matrix, so as to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image.
[0127] In one embodiment, the computer device may construct a first point cloud matrix based on the current three-dimensional point cloud of the current sample vehicle-mounted image, and calculate the offset three-dimensional point cloud based on the product of the offset matrix and the constructed first point cloud matrix. The first point cloud matrix is a matrix constructed based on the coordinate values of each point in the current three-dimensional point cloud and the supplemented dimensional coordinate values.
[0128] In one embodiment, the offset three-dimensional point cloud can be calculated by the following formula:
[0129]
[0130] Wherein, are the three-dimensional coordinates of each point in the offset three-dimensional point cloud, are the offsets of each point in the current three-dimensional point cloud on the X-axis, Y-axis, and Z-axis, respectively.
[0131] In one embodiment, the computer device may construct a second point cloud matrix based on the current three-dimensional point cloud of the current sample vehicle-mounted image, and calculate the rotated three-dimensional point cloud based on the product of each rotation matrix and the constructed second point cloud matrix. The second point cloud matrix is a matrix directly constructed based on the coordinate values of each point in the current three-dimensional point cloud.
[0132] In one embodiment, the rotated three-dimensional point cloud can be calculated by the following formula:
[0133]
[0134] Wherein, are the three-dimensional coordinates of each point in the rotated three-dimensional point cloud, are the rotation matrices of each point in the current three-dimensional point cloud for the X-axis, Y-axis, and Z-axis, respectively. Are respectively:
[0135]
[0136] Wherein, are the rotation angles of each point in the current three-dimensional point cloud on the X-axis, Y-axis, and Z-axis, respectively.
[0137] In one embodiment, the computer device may generate the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image based on the offset three-dimensional point cloud and the rotated three-dimensional point cloud.
[0138] In the above embodiment, the offset matrix can be constructed based on the offset distance, and the rotation matrix can be constructed based on the rotation angle. Furthermore, by offsetting each point in the current three-dimensional point cloud according to the offset matrix and rotating each point in the current three-dimensional point cloud according to the rotation matrix, the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image can be obtained quickly and accurately.
[0139] In one embodiment, the internal parameters include the focal coordinates of the on-vehicle camera's focus in the camera coordinate system; reconstructing the next adjacent on-vehicle image of the current sample on-vehicle image based on the adjacent three-dimensional point cloud and the internal parameters of the on-vehicle camera includes: constructing a reconstruction matrix based on the focal coordinates; based on the reconstruction matrix, converting the three-dimensional coordinates of each point in the adjacent three-dimensional point cloud into two-dimensional coordinates; constructing the next adjacent on-vehicle image of the current sample on-vehicle image based on the two-dimensional coordinates of each point.
[0140] The reconstruction matrix is a matrix used to reconstruct the next adjacent on-vehicle image of the current sample on-vehicle image.
[0141] In one embodiment, the computer device can construct a reconstruction matrix based on the focal coordinates, and based on the reconstruction matrix, convert the three-dimensional coordinates of each point in the adjacent three-dimensional point cloud into two-dimensional coordinates. Thus, the computer device can construct the next adjacent on-vehicle image of the current sample on-vehicle image based on the two-dimensional coordinates of each point. It can be understood that after the three-dimensional coordinates of each point in the adjacent three-dimensional point cloud are converted into two-dimensional coordinates, each point in the adjacent three-dimensional point cloud is converted into each pixel point in the next adjacent on-vehicle image of the current sample on-vehicle image.
[0142] In one embodiment, the computer device can construct a target point cloud matrix based on the adjacent three-dimensional point cloud of the adjacent sample on-vehicle image. The computer device can generate the next adjacent on-vehicle image of the current sample on-vehicle image based on the product of the reconstruction matrix and the target point cloud matrix. The target point cloud matrix is a matrix constructed based on the coordinate values of each point in the adjacent three-dimensional point cloud and the supplemented dimensional coordinate values.
[0143] In one embodiment, the coordinates of each pixel point in the adjacent on-vehicle image can be calculated by the following formula:
[0144]
[0145] where are the coordinates of each pixel point in the adjacent on-vehicle image, are the three-dimensional coordinates of each point in the adjacent three-dimensional point cloud.
[0146] In the above embodiment, a reconstruction matrix can be constructed based on the focal coordinates. Based on the reconstruction matrix, the three-dimensional coordinates of each point in the adjacent three-dimensional point cloud can be converted into two-dimensional coordinates. Furthermore, based on the two-dimensional coordinates of each point, the next adjacent on-vehicle image of the current sample on-vehicle image can be constructed quickly and accurately.
[0147] In one embodiment, the sample on-vehicle image includes a sample on-vehicle monocular image; obtaining an image sequence includes: collecting a plurality of sample on-vehicle monocular images through an on-vehicle monocular camera; generating an image sequence in the order of the acquisition times corresponding to the plurality of sample on-vehicle monocular images.
[0148] Among them, the sample vehicle-mounted monocular image is a sample vehicle image collected based on a vehicle-mounted monocular camera. It can be understood that each time a binocular camera collects images, a pair of binocular images can be obtained. The binocular images are a set of image pairs, which include two frames of images. Each time a vehicle-mounted monocular camera collects images, a single frame of sample vehicle image can be obtained, that is, the sample vehicle-mounted monocular image is a single-frame image.
[0149] In one embodiment, a vehicle-mounted monocular camera can be deployed in a vehicle, and a computer device can collect multiple sample vehicle-mounted monocular images through the vehicle-mounted monocular camera. The computer device can generate an image sequence according to the chronological order of the acquisition times corresponding to the multiple sample vehicle-mounted monocular images.
[0150] In the above embodiment, multiple sample vehicle-mounted monocular images are collected through a vehicle-mounted monocular camera with relatively low cost, and an image sequence can be generated according to the chronological order of the acquisition times corresponding to the multiple sample vehicle-mounted monocular images, thereby further reducing the training cost of the depth prediction network.
[0151] In one embodiment, the method further includes: performing road element detection on a target vehicle image to identify road elements from the target vehicle image; predicting the depth information of each pixel point in the road elements based on the trained depth prediction network; and generating navigation information of the target road elements based on the depth information of each pixel point in the road elements.
[0152] Among them, a road element is an entity object existing in a road, such as a road, a bridge, a road sign, and a tunnel entrance, etc. A target road element is a road element as a target. The navigation information of the target road element is information used to describe the target road element in a navigation scenario. For example, if the target road element is a tunnel, the navigation information of the target road element may include "enter the tunnel 50 meters ahead", etc.
[0153] In one embodiment, the computer device can perform feature extraction on the target vehicle image and perform road element detection based on the extracted features to identify road elements from the target vehicle image. The computer device can predict the depth information of each pixel point in the road elements based on the trained depth prediction network. Furthermore, the computer device can generate navigation information of the target road elements based on the depth information of each pixel point in the road elements.
[0154] In the above embodiment, performing road element detection on the target vehicle image can identify road elements from the target vehicle image. Based on the trained depth prediction network, the depth information of each pixel point in the road elements can be accurately predicted, and based on the depth information of each pixel point in the road elements, the navigation information of the target road elements can be generated, improving the navigation accuracy.
[0155] In one embodiment, the target vehicle-mounted image is collected by a vehicle-mounted camera disposed on the current vehicle; the method further includes: performing vehicle detection on the target vehicle-mounted image to identify an image area corresponding to the vehicle ahead from the target vehicle-mounted image; obtaining depth information of each pixel point in the image area; the depth information of each pixel point in the image area is predicted by a trained depth prediction network; determining a relative distance between the current vehicle and the vehicle ahead based on the depth information of each pixel point in the image area; and performing collision warning based on the relative distance.
[0156] Specifically, the computer device can extract features from the target vehicle-mounted image and perform vehicle detection based on the extracted features to identify an image area corresponding to the vehicle ahead from the target vehicle-mounted image. The computer device can predict the depth information of each pixel point in the image area through a trained depth prediction network. Furthermore, the computer device can determine a relative distance between the current vehicle and the vehicle ahead based on the depth information of each pixel point in the image area. The computer device can generate a collision warning message based on the relative distance and perform collision warning based on the generated collision warning message.
[0157] In one embodiment, performing collision warning based on the relative distance includes: determining a relative speed of the current vehicle relative to the vehicle ahead, determining a relative time for the current vehicle to catch up with the vehicle ahead based on the relative distance and the relative speed, generating a collision warning message when the relative time is less than a preset safety time, and performing collision warning based on the generated collision warning message.
[0158] In one embodiment, as Figure 5 shown, the current vehicle A and the vehicle B ahead of it are traveling on the road in the same driving direction. The computer device can predict the depth information of each pixel point in the image area corresponding to the vehicle B ahead through a trained depth prediction network. Furthermore, the computer device can determine a relative distance between the current vehicle A and the vehicle B ahead based on the depth information of each pixel point in the image area. At the same time, the computer device can determine a relative speed of the current vehicle A relative to the vehicle B ahead, determine a relative time for the current vehicle A to catch up with the vehicle B ahead based on the relative distance and the relative speed, generate a collision warning message when the relative time is less than a preset safety time, and send the collision warning message to the current vehicle A for collision warning.
[0159] In the above embodiment, performing vehicle detection on the target vehicle-mounted image can identify an image area corresponding to the vehicle ahead from the target vehicle-mounted image. Through a trained depth prediction network, the depth information of each pixel point in the image area can be accurately predicted. Furthermore, based on the depth information of each pixel point in the image area, the relative distance between the current vehicle and the vehicle ahead can be determined, and collision warning is performed based on the relative distance, improving driving safety.
[0160] In one embodiment, the method further includes: detecting road elements in the target vehicle-mounted image to identify road elements from the target vehicle-mounted image; obtaining depth information of each pixel point in the road elements; the depth information of each pixel point in the road elements is predicted by a trained depth prediction network; and writing the target road elements into the vehicle navigation map based on the depth information of each pixel point in the road elements.
[0161] The vehicle navigation map is a map that is set in the vehicle and used for vehicle navigation.
[0162] In one embodiment, the computer device can extract features from the target vehicle-mounted image and perform road element detection based on the extracted features to identify road elements from the target vehicle-mounted image. The computer device can predict the depth information of each pixel point in the road elements based on a trained depth prediction network. Further, the computer device can write the target road elements into the vehicle navigation map based on the depth information of each pixel point in the road elements to update the vehicle navigation map.
[0163] In the above embodiment, by detecting road elements in the target vehicle-mounted image, road elements can be identified from the target vehicle-mounted image. Through a trained depth prediction network, the depth information of each pixel point in the road elements can be accurately predicted. Thus, based on the depth information of each pixel point in the road elements, the target road elements can be written into the vehicle navigation map to accurately update the vehicle navigation map.
[0164] In one embodiment, as Figure 6 shown, the computer device can obtain an image sequence including a plurality of sequentially acquired sample vehicle-mounted images through a vehicle-mounted camera set on the vehicle, respectively use each sample vehicle-mounted image in the image sequence as the current sample vehicle-mounted image, and input the current sample vehicle-mounted image into the depth prediction network to be trained to predict the depth information of each pixel point in the current sample vehicle-mounted image. The computer device can generate the current three-dimensional point cloud of the current sample vehicle-mounted image based on the depth information. The computer device can obtain vehicle pose information determined based on the current sample vehicle-mounted image and the next adjacent sample vehicle-mounted image from a vehicle pose sensor set on the vehicle, and reconstruct the next adjacent vehicle-mounted image of the current sample vehicle-mounted image based on the current three-dimensional point cloud and the vehicle pose information. The computer device can perform iterative training on the depth prediction network through backpropagation based on the difference between the adjacent sample vehicle-mounted image and the reconstructed adjacent vehicle-mounted image until the training is completed when the iteration stop condition is met, and obtain a trained depth prediction network. The computer device can predict the pixel depth of the target vehicle-mounted image through the trained depth prediction network.
[0165] As Figure 7As shown, in one embodiment, an image processing method is provided, which specifically includes the following steps:
[0166] Step 702, obtain an image sequence; the image sequence includes a plurality of sampled vehicle images collected in sequence; the current sampled vehicle image is collected by a vehicle-mounted camera.
[0167] Step 704, respectively use each sampled vehicle image in the image sequence as the current sampled vehicle image, and input the current sampled vehicle image into a depth prediction network to be trained to predict the depth information of each pixel point in the current sampled vehicle image; the depth information includes the depth coordinates of each pixel point in the current sampled vehicle image in the camera coordinate system; the camera coordinate system is a coordinate system established with the optical center of the vehicle-mounted camera as the origin.
[0168] Step 706, determine the focus coordinates of the focus of the vehicle-mounted camera in the camera coordinate system.
[0169] Step 708, determine the current pixel coordinates of each pixel point in the current sampled vehicle image in the current pixel coordinate system; the current pixel coordinate system is a pixel coordinate system established based on the current sampled vehicle image.
[0170] Step 710, according to the focus coordinates and the current pixel coordinates of each pixel point in the current pixel coordinate system, determine the planar coordinates of each pixel point in the camera coordinate system.
[0171] Step 712, generate the current 3D point cloud of the current sampled vehicle image based on the depth coordinates and planar coordinates of each pixel point.
[0172] Step 714, obtain vehicle pose information determined based on the current sampled vehicle image and the next adjacent sampled vehicle image from the vehicle pose sensor; the vehicle pose information includes the offset distance and rotation angle of the vehicle from the current position to the next adjacent position.
[0173] Step 716, construct an offset matrix based on the offset distance and a rotation matrix based on the rotation angle.
[0174] Step 718, perform offset processing on each point in the current 3D point cloud according to the offset matrix, and perform rotation processing on each point in the current 3D point cloud according to the rotation matrix to obtain the adjacent 3D point cloud of the adjacent sampled vehicle image.
[0175] Step 720, construct a reconstruction matrix based on the focus coordinates of the focus of the vehicle-mounted camera in the camera coordinate system, and based on the reconstruction matrix, convert the 3D coordinates of each point in the adjacent 3D point cloud into 2D coordinates.
[0176] Step 722, construct the next adjacent vehicle image of the current sampled vehicle image based on the 2D coordinates of each point.
[0177] Step 724: Train the depth prediction network based on the difference between the adjacent sample vehicle-mounted images and the reconstructed adjacent vehicle-mounted images; the trained depth prediction network is used to predict the pixel depth of the target vehicle-mounted image.
[0178] In one embodiment, perform road element detection on the target vehicle-mounted image to identify road elements from the target vehicle-mounted image; based on the trained depth prediction network, predict the depth information of each pixel point in the road elements; based on the depth information of each pixel point in the road elements, generate navigation information for the target road elements.
[0179] In one embodiment, the target vehicle-mounted image is acquired by a vehicle-mounted camera disposed on the current vehicle; perform vehicle detection on the target vehicle-mounted image to identify the image area corresponding to the vehicle ahead from the target vehicle-mounted image; obtain the depth information of each pixel point in the image area; the depth information of each pixel point in the image area is predicted by the trained depth prediction network; based on the depth information of each pixel point in the image area, determine the relative distance between the current vehicle and the vehicle ahead; perform collision warning based on the relative distance.
[0180] In one embodiment, perform road element detection on the target vehicle-mounted image to identify road elements from the target vehicle-mounted image; obtain the depth information of each pixel point in the road elements; the depth information of each pixel point in the road elements is predicted by the trained depth prediction network; based on the depth information of each pixel point in the road elements, write the target road elements into the vehicle-mounted navigation map.
[0181] The present application also provides an application scenario that applies the above image processing method. Specifically, the image processing method can be applied to the scenario of on-vehicle monocular image processing. The computer device can obtain an image sequence; the image sequence includes a plurality of sampled on-vehicle monocular images collected in sequence; the current sampled on-vehicle monocular image is collected by an on-vehicle monocular camera. Each sampled on-vehicle monocular image in the image sequence is respectively used as the current sampled on-vehicle monocular image, and the current sampled on-vehicle monocular image is input into the depth prediction network to be trained, and the depth information of each pixel point in the current sampled on-vehicle monocular image is predicted; the depth information includes the depth coordinates of each pixel point in the current sampled on-vehicle monocular image in the camera coordinate system; the camera coordinate system is a coordinate system established with the optical center of the on-vehicle monocular camera as the origin. Determine the focal coordinates of the focus of the on-vehicle monocular camera in the camera coordinate system. Determine the current pixel coordinates of each pixel point in the current sampled on-vehicle monocular image in the current pixel coordinate system; the current pixel coordinate system is a pixel coordinate system established based on the current sampled on-vehicle monocular image. According to the focal coordinates and the current pixel coordinates of each pixel point in the current pixel coordinate system, determine the plane coordinates of each pixel point in the camera coordinate system. Based on the depth coordinates and plane coordinates of each pixel point, generate the current 3D point cloud of the current sampled on-vehicle monocular image.
[0182] The computer device can obtain the vehicle pose information determined based on the current sampled on-vehicle monocular image and the next adjacent sampled on-vehicle monocular image from the on-vehicle pose sensor; the vehicle pose information includes the offset distance and rotation angle of the vehicle from the current position to the next adjacent position. Construct an offset matrix based on the offset distance and a rotation matrix based on the rotation angle. Offset each point in the current 3D point cloud according to the offset matrix, and rotate each point in the current 3D point cloud according to the rotation matrix to obtain the adjacent 3D point cloud of the adjacent sampled on-vehicle monocular image.
[0183] The computer device can construct a reconstruction matrix based on the focal coordinates of the focus of the on-vehicle monocular camera in the camera coordinate system. Based on the reconstruction matrix, convert the 3D coordinates of each point in the adjacent 3D point cloud into 2D coordinates. Construct the next adjacent on-vehicle monocular image of the current sampled on-vehicle monocular image based on the 2D coordinates of each point. Train the depth prediction network based on the difference between the adjacent sampled on-vehicle monocular image and the reconstructed adjacent on-vehicle monocular image; the trained depth prediction network is used to predict the pixel depth of the target on-vehicle monocular image.
[0184] The present application further provides an application scenario that applies the above image processing method. Specifically, the image processing method can be applied to the scenario of on-vehicle binocular image processing. It can be understood that the scenario of on-vehicle binocular image processing in the present application is only used to provide more sample on-vehicle images, that is, to provide richer training data for training the depth prediction network, rather than realizing depth prediction based on the image differences between binocular images as in the traditional method.
[0185] It should be understood that although the steps in the flowcharts of the above embodiments are shown in sequence, these steps are not necessarily executed in sequence. Unless there is a clear indication in this article, the execution of these steps has no strict sequence limit, and these steps can be executed in other sequences. Moreover, at least a part of the steps in the above embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0186] In one embodiment, as Figure 8 shown, an image processing apparatus 800 is provided. The apparatus can be a software module, a hardware module, or a combination of both to form a part of a computer device. Specifically, the apparatus includes:
[0187] An acquisition module 801, configured to acquire an image sequence; the image sequence includes a plurality of sample on-vehicle images acquired in sequence.
[0188] A prediction module 802, configured to use each sample on-vehicle image in the image sequence as the current sample on-vehicle image, and input the current sample on-vehicle image into a depth prediction network to be trained, and predict the depth information of each pixel point in the current sample on-vehicle image.
[0189] A generation module 803, configured to generate the current three-dimensional point cloud of the current sample on-vehicle image based on the depth information;
[0190] The acquisition module 801 is further configured to acquire vehicle pose information determined based on the current sample on-vehicle image and the next adjacent sample on-vehicle image from an on-vehicle pose sensor.
[0191] A reconstruction module 804, configured to reconstruct the next adjacent on-vehicle image of the current sample on-vehicle image based on the current three-dimensional point cloud and the vehicle pose information.
[0192] A training module 805 is configured to train a depth prediction network based on the difference between an adjacent sample vehicle-mounted image and a reconstructed adjacent vehicle-mounted image; the trained depth prediction network is configured to predict the pixel depth of a target vehicle-mounted image.
[0193] In one embodiment, the current sample vehicle-mounted image is acquired by a vehicle-mounted camera; the depth information includes the depth coordinates of each pixel point in the current sample vehicle-mounted image in the camera coordinate system; the generation module 803 is further configured to obtain the planar coordinates of each pixel point in the current sample vehicle-mounted image in the camera coordinate system; the camera coordinate system is a coordinate system established with the optical center of the vehicle-mounted camera as the origin; based on the depth coordinates and planar coordinates of each pixel point, the current three-dimensional point cloud of the current sample vehicle-mounted image is generated.
[0194] In one embodiment, the generation module 803 is further configured to determine the focal coordinates of the focal point of the vehicle-mounted camera in the camera coordinate system; determine the current pixel coordinates of each pixel point in the current sample vehicle-mounted image in the current pixel coordinate system; the current pixel coordinate system is a pixel coordinate system established based on the current sample vehicle-mounted image; according to the focal coordinates and the current pixel coordinates of each pixel point in the current pixel coordinate system, the planar coordinates of each pixel point in the camera coordinate system are determined.
[0195] In one embodiment, the current sample vehicle-mounted image is acquired by a vehicle-mounted camera; the reconstruction module 804 is further configured to determine the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image based on the current three-dimensional point cloud and the vehicle pose information; according to the adjacent three-dimensional point cloud and the internal parameters of the vehicle-mounted camera, the next adjacent vehicle-mounted image of the current sample vehicle-mounted image is reconstructed.
[0196] In one embodiment, the vehicle pose information includes the offset distance and rotation angle of the vehicle from the current position to the next adjacent position; the current position is the position where the vehicle is located when the current sample vehicle-mounted image is acquired; the next adjacent position is the position where the vehicle is located when the next adjacent sample vehicle-mounted image is acquired; the reconstruction module 804 is further configured to adjust each point in the current three-dimensional point cloud according to the offset distance and rotation angle respectively to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image.
[0197] In one embodiment, the reconstruction module 804 is further configured to construct an offset matrix based on the offset distance; construct a rotation matrix based on the rotation angle; offset each point in the current three-dimensional point cloud according to the offset matrix, and rotate each point in the current three-dimensional point cloud according to the rotation matrix to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image.
[0198] In one embodiment, the internal parameters include the focal coordinates of the focus of the vehicle-mounted camera in the camera coordinate system; the reconstruction module 804 is further configured to construct a reconstruction matrix based on the focal coordinates; based on the reconstruction matrix, convert the three-dimensional coordinates of each point in the adjacent three-dimensional point cloud into two-dimensional coordinates; and construct the next adjacent vehicle-mounted image of the current sample vehicle-mounted image based on the two-dimensional coordinates of each point.
[0199] In one embodiment, the sample vehicle-mounted image includes a sample vehicle-mounted monocular image; the acquisition module 801 is further configured to collect a plurality of sample vehicle-mounted monocular images through the vehicle-mounted monocular camera; and generate an image sequence in the order of the acquisition times corresponding to the plurality of sample vehicle-mounted monocular images.
[0200] In one embodiment, the apparatus further includes:
[0201] A navigation module 806, configured to perform road element detection on a target vehicle-mounted image to identify road elements from the target vehicle-mounted image; predict depth information of each pixel point in the road elements based on the trained depth prediction network; and generate navigation information of the target road elements based on the depth information of each pixel point in the road elements.
[0202] In one embodiment, the target vehicle-mounted image is collected by a vehicle-mounted camera disposed on the current vehicle; the apparatus further includes:
[0203] An early warning module 807, configured to perform vehicle detection on the target vehicle-mounted image to identify an image area corresponding to a vehicle ahead from the target vehicle-mounted image; obtain depth information of each pixel point in the image area; the depth information of each pixel point in the image area is predicted by the trained depth prediction network; determine a relative distance between the current vehicle and the vehicle ahead based on the depth information of each pixel point in the image area; and perform collision warning based on the relative distance.
[0204] In one embodiment, the apparatus further includes:
[0205] A writing module 808, configured to perform road element detection on the target vehicle-mounted image to identify road elements from the target vehicle-mounted image; obtain depth information of each pixel point in the road elements; the depth information of each pixel point in the road elements is predicted by the trained depth prediction network; and write the target road elements into the vehicle-mounted navigation map based on the depth information of each pixel point in the road elements.
[0206] Reference Figure 9 , in one embodiment, the image processing apparatus 800 further includes a navigation module 806, an early warning module 807, and a writing module 808.
[0207] The above image processing device acquires an image sequence including a plurality of sequentially acquired sample vehicle images, takes each sample vehicle image in the image sequence as the current sample vehicle image respectively, and inputs the current sample vehicle image into a depth prediction network to be trained, so as to predict the depth information of each pixel point in the current sample vehicle image. Based on the depth information, the current 3D point cloud of the current sample vehicle image can be generated. From the vehicle pose sensor, the vehicle pose information determined based on the current sample vehicle image and the next adjacent sample vehicle image can be directly obtained. Based on the current 3D point cloud and the vehicle pose information, the next adjacent vehicle image of the current sample vehicle image can be reconstructed. Based on the difference between the adjacent sample vehicle image and the reconstructed adjacent vehicle image, the depth prediction network can be directly trained without supervision. Compared with the traditional supervised training method, the present application can directly reconstruct the adjacent vehicle image corresponding to the adjacent sample vehicle image through the current sample vehicle image in the image sequence, the depth information of the current sample vehicle image, and the vehicle pose information determined based on the current sample vehicle image and the next adjacent sample vehicle image, and can directly implement the unsupervised training of the depth prediction network based on the difference between the adjacent sample vehicle image and the adjacent vehicle image, without pre-labeling the image depth information, greatly reducing the training cost of the depth prediction network.
[0208] For the specific limitations of the image processing device, reference can be made to the limitations of the image processing method in the above text, which will not be elaborated here. Each module in the above image processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or be independent of it, or be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0209] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 10 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an image processing method is implemented.
[0210] Those skilled in the art can understand that Figure 10The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0211] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0212] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0213] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0214] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0215] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0216] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0217] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. An image processing method, characterized in that, The method includes: Obtaining an image sequence; the image sequence includes a plurality of sample vehicle images collected in sequence; Taking each sample vehicle image in the image sequence as the current sample vehicle image respectively, and inputting the current sample vehicle image into a depth prediction network to be trained to predict the depth information of each pixel point in the current sample vehicle image; Generating a current three-dimensional point cloud of the current sample vehicle image based on the depth information; Obtaining vehicle pose information determined based on the current sample vehicle image and the next adjacent sample vehicle image from a vehicle pose sensor; the vehicle pose information is information about the pose change of the vehicle in space during the process of the vehicle collecting the current sample vehicle image to collecting the next adjacent sample vehicle image; Reconstructing the next adjacent vehicle image of the current sample vehicle image based on the current three-dimensional point cloud and the vehicle pose information; the next adjacent vehicle image is a reconstructed vehicle image that has substantially the same image content as the next adjacent sample vehicle image; Training the depth prediction network based on the difference between the adjacent sample vehicle image and the reconstructed adjacent vehicle image; the trained depth prediction network is used to predict the pixel depth of the target vehicle image.
2. The method according to claim 1, characterized in that, The current sample vehicle image is collected by a vehicle-mounted camera; the depth information includes the depth coordinates of each pixel point in the current sample vehicle image in the camera coordinate system; the generating the current three-dimensional point cloud of the current sample vehicle image based on the depth information includes: Obtaining the planar coordinates of each pixel point in the current sample vehicle image in the camera coordinate system; the camera coordinate system is a coordinate system established with the optical center of the vehicle-mounted camera as the origin; Generating the current three-dimensional point cloud of the current sample vehicle image based on the depth coordinates and the planar coordinates of each pixel point.
3. The method according to claim 2, wherein The obtaining the planar coordinates of each pixel point in the current sample vehicle image in the camera coordinate system includes: Determining the focus coordinates of the focus of the vehicle-mounted camera in the camera coordinate system; Determining the current pixel coordinates of each pixel point in the current sample vehicle image in the current pixel coordinate system; the current pixel coordinate system is a pixel coordinate system established based on the current sample vehicle image; Determining the planar coordinates of each pixel point in the camera coordinate system according to the focus coordinates and the current pixel coordinates of each pixel point in the current pixel coordinate system.
4. The method according to claim 1, wherein The current sample vehicle image is collected by a vehicle-mounted camera; the reconstructing the next adjacent vehicle image of the current sample vehicle image based on the current three-dimensional point cloud and the vehicle pose information includes: Determining the adjacent three-dimensional point cloud of the adjacent sample vehicle image based on the current three-dimensional point cloud and the vehicle pose information; Reconstructing the next adjacent vehicle image of the current sample vehicle image according to the adjacent three-dimensional point cloud and the internal parameters of the vehicle-mounted camera.
5. The method according to claim 4, wherein The vehicle pose information includes the offset distance and rotation angle of the vehicle from the current position to the next adjacent position; the current position is the position where the vehicle is located when the current sample vehicle-mounted image is collected; the next adjacent position is the position where the vehicle is located when the next adjacent sample vehicle-mounted image is collected; Determining the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image based on the current three-dimensional point cloud and the vehicle pose information includes: Adjusting each point in the current three-dimensional point cloud according to the offset distance and the rotation angle respectively to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image.
6. The method according to claim 5, characterized in that, The adjusting each point in the current three-dimensional point cloud according to the offset distance and the rotation angle respectively to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image includes: Constructing an offset matrix based on the offset distance; constructing a rotation matrix based on the rotation angle; Offsetting each point in the current three-dimensional point cloud according to the offset matrix, and rotating each point in the current three-dimensional point cloud according to the rotation matrix to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image.
7. The method according to claim 4, characterized in that, The internal parameters include the focal coordinates of the focal point of the vehicle-mounted camera in the camera coordinate system; Reconstructing the next adjacent vehicle-mounted image of the current sample vehicle-mounted image according to the adjacent three-dimensional point cloud and the internal parameters of the vehicle-mounted camera includes: Constructing a reconstruction matrix based on the focal coordinates; Converting the three-dimensional coordinates of each point in the adjacent three-dimensional point cloud into two-dimensional coordinates based on the reconstruction matrix; Constructing the next adjacent vehicle-mounted image of the current sample vehicle-mounted image based on the two-dimensional coordinates of each point.
8. The method according to claim 1, characterized in that The sample vehicle-mounted image includes a sample vehicle-mounted monocular image; obtaining the image sequence includes: Collecting a plurality of sample vehicle-mounted monocular images through a vehicle-mounted monocular camera; Generating the image sequence according to the chronological order of the acquisition times corresponding to the plurality of sample vehicle-mounted monocular images.
9. The method according to any one of claims 1 to 8, characterized in that The method further includes: Detecting road elements in the target vehicle-mounted image to identify road elements from the target vehicle-mounted image; Predicting the depth information of each pixel point in the road elements based on the trained depth prediction network; Generating navigation information of the road elements based on the depth information of each pixel point in the road elements.
10. The method according to any one of claims 1 to 8, characterized in that, The target vehicle-mounted image is collected by a vehicle-mounted camera disposed on the current vehicle; the method further includes: Detecting a vehicle in the target vehicle-mounted image to identify the image area corresponding to the vehicle in front from the target vehicle-mounted image; Obtaining the depth information of each pixel point in the image area; the depth information of each pixel point in the image area is predicted by the trained depth prediction network; Determining the relative distance between the current vehicle and the vehicle in front based on the depth information of each pixel point in the image area; Performing a collision warning based on the relative distance.
11. The method according to any one of claims 1 to 8, characterized in that, The method further includes: Detecting road elements in the target vehicle-mounted image to identify road elements from the target vehicle-mounted image; Obtain the depth information of each pixel point in the road element; the depth information of each pixel point in the road element is predicted by the trained depth prediction network; Based on the depth information of each pixel point in the road element, write the road element into the in-vehicle navigation map.
12. An image processing apparatus, characterized in that, The device includes: An acquisition module, configured to acquire an image sequence; the image sequence includes a plurality of sample in-vehicle images collected in sequence; A prediction module, configured to respectively use each sample in-vehicle image in the image sequence as the current sample in-vehicle image, and input the current sample in-vehicle image into the depth prediction network to be trained, and predict the depth information of each pixel point in the current sample in-vehicle image; A generation module, configured to generate the current 3D point cloud of the current sample in-vehicle image based on the depth information; The acquisition module is further configured to obtain vehicle pose information determined based on the current sample in-vehicle image and the next adjacent sample in-vehicle image from an in-vehicle pose sensor; the vehicle pose information is the information about the attitude change of the vehicle in space during the process of the vehicle collecting the current sample in-vehicle image to collecting the next adjacent sample in-vehicle image; A reconstruction module, configured to reconstruct the next adjacent in-vehicle image of the current sample in-vehicle image based on the current 3D point cloud and the vehicle pose information; the next adjacent in-vehicle image is a reconstructed in-vehicle image that has substantially the same image content as the next adjacent sample in-vehicle image; A training module, configured to train the depth prediction network based on the difference between the adjacent sample in-vehicle image and the reconstructed adjacent in-vehicle image; the trained depth prediction network is used to predict the pixel depth of the target in-vehicle image.
13. The device according to claim 12, wherein The current sample in-vehicle image is collected by an in-vehicle camera; the depth information includes the depth coordinates of each pixel point in the current sample in-vehicle image in the camera coordinate system; the generation module is further configured to obtain the plane coordinates of each pixel point in the current sample in-vehicle image in the camera coordinate system; the camera coordinate system is a coordinate system established with the optical center of the in-vehicle camera as the origin; Generate the current 3D point cloud of the current sample in-vehicle image based on the depth coordinates and the plane coordinates of each pixel point.
14. The device according to claim 13, characterized in that, The generation module is further configured to determine the focal point coordinates of the focal point of the in-vehicle camera in the camera coordinate system; determine the current pixel coordinates of each pixel point in the current sample in-vehicle image in the current pixel coordinate system; the current pixel coordinate system is a pixel coordinate system established based on the current sample in-vehicle image; determine the plane coordinates of each pixel point in the camera coordinate system according to the focal point coordinates and the current pixel coordinates of each pixel point in the current pixel coordinate system.
15. The device according to claim 12, characterized in that, The current sample in-vehicle image is collected by an in-vehicle camera; the reconstruction module is further configured to determine the adjacent 3D point cloud of the adjacent sample in-vehicle image based on the current 3D point cloud and the vehicle pose information; Reconstruct the next adjacent in-vehicle image of the current sample in-vehicle image according to the adjacent 3D point cloud and the internal parameters of the in-vehicle camera.
16. The device according to claim 15, characterized in that, The vehicle pose information includes the offset distance and rotation angle of the vehicle from the current position to the next adjacent position; the current position is the position where the vehicle is located when the current sample vehicle-mounted image is acquired; the next adjacent position is the position where the vehicle is located when the next adjacent sample vehicle-mounted image is acquired; the reconstruction module is further configured to adjust each point in the current three-dimensional point cloud according to the offset distance and the rotation angle respectively to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image.
17. The device according to claim 16, characterized in that The reconstruction module is further configured to construct an offset matrix based on the offset distance; construct a rotation matrix based on the rotation angle; Perform offset processing on each point in the current three-dimensional point cloud according to the offset matrix, and perform rotation processing on each point in the current three-dimensional point cloud according to the rotation matrix to obtain the adjacent three-dimensional point cloud of the adjacent sample vehicle-mounted image.
18. The device according to claim 15, wherein The internal parameters include the focal coordinates of the focus of the vehicle-mounted camera in the camera coordinate system; the reconstruction module is further configured to construct a reconstruction matrix based on the focal coordinates; Based on the reconstruction matrix, convert the three-dimensional coordinates of each point in the adjacent three-dimensional point cloud into two-dimensional coordinates; construct the next adjacent vehicle-mounted image of the current sample vehicle-mounted image based on the two-dimensional coordinates of each point.
19. The device according to claim 12, characterized in that, The sample vehicle-mounted image includes a sample vehicle-mounted monocular image; the acquisition module is further configured to acquire a plurality of sample vehicle-mounted monocular images through a vehicle-mounted monocular camera; generate the image sequence according to the chronological order of the acquisition times corresponding to the plurality of sample vehicle-mounted monocular images.
20. The device according to any one of claims 12 to 19, characterized in that, The device further includes: A navigation module, configured to perform road element detection on a target vehicle-mounted image to identify road elements from the target vehicle-mounted image; predict the depth information of each pixel point in the road elements based on the trained depth prediction network; generate navigation information of the road elements based on the depth information of each pixel point in the road elements.
21. The device according to any one of claims 12 to 19, characterized in that, The target vehicle-mounted image is acquired by a vehicle-mounted camera disposed on the current vehicle; the device further includes: An early warning module, configured to perform vehicle detection on the target vehicle-mounted image to identify the image area corresponding to the vehicle ahead from the target vehicle-mounted image; obtain the depth information of each pixel point in the image area; the depth information of each pixel point in the image area is predicted by the trained depth prediction network; determine the relative distance between the current vehicle and the vehicle ahead based on the depth information of each pixel point in the image area; perform collision early warning based on the relative distance.
22. The device according to any one of claims 12 to 19, characterized in that The device further includes: A writing module, configured to perform road element detection on the target vehicle-mounted image to identify road elements from the target vehicle-mounted image; obtain the depth information of each pixel point in the road elements; the depth information of each pixel point in the road elements is predicted by the trained depth prediction network; write the road elements into the vehicle-mounted navigation map based on the depth information of each pixel point in the road elements.
23. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 11.
24. A computer-readable storage medium storing a computer program, characterized in that, The steps of the method according to any one of claims 1 to 11 are implemented when the computer program is executed by a processor.
25. A computer program product, comprising a computer program, characterized in that, The steps of the method according to any one of claims 1 to 11 are implemented when the computer program is executed by a processor.
Citation Information
Patent Citations
Depth and image information fused 3D enhanced panoramic look-around system and implementation method
CN111559314A
Obstacle detection method and device, computer equipment and storage medium
CN112417967A
Unsupervised learning of image depth and ego-motion prediction neural networks
IN202027016267A