Point cloud filling method and device based on deep neural network, terminal and medium

Through the point cloud filling method based on deep neural network, the images acquired by binocular cameras are processed and filled, which solves the problem that binocular stereo vision technology cannot obtain dense point cloud data, realizes high-precision point cloud data acquisition, and reduces costs.

CN119991513APending Publication Date: 2025-05-13SHENZHEN JUYUAN VIDEOCHIP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510132987.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing binocular stereo vision technology cannot obtain dense and accurate point cloud data.

Method used

Using a point cloud filling method based on deep neural network, the binocular images obtained by using a binocular camera are image correction and alignment processing, and the sparse depth images are calculated and the hollow parts are marked. Then, the sparse depth image with the hollow mark and the target reference image are input into the pre-trained deep neural network, and the hole portion is filled with the neural network to obtain the filled dense depth image and point cloud data.

Benefits of technology

It realizes the acquisition of point cloud data with dense and accurate information through binocular cameras, which solves the problem that binocular stereo vision technology cannot obtain dense point cloud data, and only using binocular cameras, significantly reducing the cost of use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991513A_ABST
    Figure CN119991513A_ABST
Patent Text Reader

Abstract

The invention provides a point cloud filling method and device based on a deep neural network, a terminal and a medium, and the method comprises the steps: carrying out the image correction and alignment processing of a binocular image obtained through a binocular camera, calculating a sparse depth image based on a target binocular image, marking a hole part in the sparse depth image, and carrying out the recognition of the hole part in the sparse depth image; inputting the sparse depth image with the hole mark and a target reference image corresponding to a reference sub-camera of the binocular camera in the target binocular image into a pre-trained target depth neural network, and the target deep neural network uses the image information provided by the target reference image to fill the cavity part in the sparse depth image, and dense point cloud data is acquired based on the dense depth image obtained through filling. According to the invention, dense point cloud data is obtained based on the binocular vision algorithm and in combination with deep neural network filling, only a binocular camera is used, and the use cost is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of point cloud technology, and in particular to a point cloud filling method, device, terminal and medium based on a deep neural network. Background Art

[0002] At present, point cloud data can be mainly obtained by using laser radar (LiDAR), depth camera and binocular stereo vision technology. The processing and analysis of point cloud data has a wide range of applications in many fields, such as architecture, urban planning, autonomous driving, robot navigation, cultural heritage protection, etc.

[0003] Among them, LiDAR measures distance by emitting laser beams and receiving reflected signals to obtain accurate three-dimensional coordinates of the surface of an object. LiDAR is usually used for large-scale, high-precision point cloud data acquisition, especially for geographic surveying, building scanning and autonomous driving. LiDAR has the advantages of high precision, fast acquisition, and adaptability to complex environments, but it is greatly affected by weather and environmental conditions, and the equipment is relatively expensive; while depth cameras generate point cloud data by capturing the depth information of each pixel in the scene (i.e., the distance to the camera). Common depth cameras include Kinect, RealSense, etc. These devices are relatively cheap, easy to use and integrate, but are usually not as accurate and range as LiDAR and are greatly affected by lighting conditions; binocular stereo vision technology is mostly used in robots and driverless cars. Two cameras are used to shoot the same scene from different angles, and image matching technology is used to calculate depth information. Binocular stereo vision does not require special hardware and the equipment is relatively cheap, but it requires high computing power, high computational complexity, and is easily affected by occlusion and lighting conditions. There is also the problem of sparse point cloud data obtained.

[0004] In summary, the above three methods of acquiring point cloud data cannot acquire point cloud data with dense and accurate information, but acquiring point cloud data using binocular stereo vision technology is a better method than acquiring point cloud data using lidar or depth camera. Therefore, how to provide a solution to the problem that binocular stereo vision technology cannot acquire point cloud data with dense and accurate information is a problem that technical personnel in this field currently need to solve. Summary of the invention

[0005] The technical problem to be solved by the present invention is that, in view of the above-mentioned defects of the prior art, a point cloud filling method, device, terminal and medium based on a deep neural network are provided, aiming to solve the problem that the existing binocular stereo vision technology cannot obtain point cloud data with dense and accurate information.

[0006] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0007] A point cloud filling method based on a deep neural network, wherein the method comprises:

[0008] Performing image correction and alignment processing on the binocular image acquired by the binocular camera to obtain a target binocular image, and calculating a corresponding sparse depth image based on the target binocular image;

[0009] Determine a hole portion existing in the sparse depth image, and mark the hole portion in the sparse depth image to obtain a sparse depth image with a hole mark;

[0010] Inputting the sparse depth image with hole marks and the target reference image in the target binocular image into a pre-trained target deep neural network, so that the target deep neural network fills the hole parts in the sparse depth image with image information provided by the target reference image to obtain a filled dense depth image; wherein the target reference image is an image corresponding to the reference sub-camera in the binocular camera;

[0011] Corresponding dense point cloud data is obtained based on the filled dense depth image.

[0012] In one implementation, the step of calculating a corresponding sparse depth image based on the target binocular image includes:

[0013] Determine the same features between the target binocular images using a preset semi-global matching algorithm to obtain corresponding left and right disparity images;

[0014] The left and right disparity images are converted into distances to generate corresponding sparse depth images.

[0015] In one implementation, the determining of a hole portion existing in the sparse depth image includes:

[0016] The left and right disparity images are compared to determine the corresponding pixels. Figure 1 The consistency check operation obtains the corresponding check result;

[0017] Based on the inspection result, a hole portion existing in the sparse depth image is determined.

[0018] In one implementation, determining the hole portion existing in the sparse depth image based on the inspection result includes:

[0019] If the inspection result indicates that there is a target area with inconsistent pixels in the left and right parallax images, the target area with inconsistent pixels is determined as a hole portion existing in the sparse depth image.

[0020] In one implementation, the preset semi-global matching algorithm is an algorithm that pre-optimizes and adjusts its cost aggregation step so that in the cost aggregation step, two adjacent pixels in different rows are processed per cycle based on a preset pipeline cost aggregation architecture, and aggregation in a single different direction is implemented.

[0021] In one implementation, the target deep neural network is a target convolutional neural network.

[0022] In one implementation, the step of inputting the sparse depth image with hole marks and the target reference image in the target binocular image into a pre-trained target deep neural network, so that the target deep neural network fills the hole portion in the sparse depth image using image information provided by the target reference image to obtain a filled dense depth image, including:

[0023] The sparse depth image with hole marks and the target reference image in the target binocular image are input into a pre-trained target convolutional neural network with a pixel-level pipeline architecture according to a preset pixel stream input method, so that the target convolutional neural network fills the hole parts in the sparse depth image based on an inter-layer fusion method and uses the image information provided by the target reference image to obtain a filled dense depth image.

[0024] The present invention also discloses a point cloud filling device based on a deep neural network, wherein the device comprises:

[0025] An image processing module is used to perform image correction and alignment processing on the binocular image acquired by the binocular camera to obtain a target binocular image;

[0026] A depth image calculation module, used to calculate a corresponding sparse depth image based on the target binocular image;

[0027] A hole determination module, used for determining a hole portion existing in the sparse depth image, and marking the hole portion in the sparse depth image to obtain a sparse depth image with a hole mark;

[0028] A point cloud filling module, used for inputting the sparse depth image with hole marks and the target reference image in the target binocular image into a pre-trained target deep neural network, so that the target deep neural network fills the hole parts in the sparse depth image with the image information provided by the target reference image to obtain a filled dense depth image; wherein the target reference image is an image corresponding to the reference sub-camera in the binocular camera;

[0029] The point cloud data acquisition module is used to acquire corresponding dense point cloud data based on the filled dense depth image.

[0030] The present invention also discloses a terminal, which includes: a memory, a processor, and a point cloud filling program based on a deep neural network stored in the memory and executable on the processor, wherein the point cloud filling program based on the deep neural network, when executed by the processor, implements the steps of the point cloud filling method based on the deep neural network as described above.

[0031] The present invention also discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the point cloud filling method based on a deep neural network as described above.

[0032] The present invention provides a point cloud filling method based on a deep neural network, a device, a terminal and a medium. The point cloud filling method based on a deep neural network includes: performing image correction and alignment processing on a binocular image acquired by a binocular camera to obtain a target binocular image, and calculating a corresponding sparse depth image based on the target binocular image; determining a hole part existing in the sparse depth image, and marking the hole part in the sparse depth image to obtain a sparse depth image with a hole mark; inputting the sparse depth image with the hole mark and a target reference image in the target binocular image into a pre-trained target deep neural network, so that the target deep neural network fills the hole part in the sparse depth image with image information provided by the target reference image to obtain a filled dense depth image; wherein the target reference image is an image corresponding to a reference sub-camera in the binocular camera; and obtaining corresponding dense point cloud data based on the filled dense depth image. It can be seen that the present invention obtains a binocular image through a binocular camera, and performs correction and alignment processing on it, and then calculates a sparse depth image based on the image, and marks the hole parts of the sparse depth image, and finally inputs the sparse depth image with hole marks and the target reference image in the target binocular image into a pre-trained target deep neural network, so as to achieve the acquisition of point cloud data with dense and accurate information through the filling of the deep neural network. That is, the technical solution of the present application is based on the binocular vision algorithm combined with the deep neural network filling to obtain dense point cloud data, thereby solving the problem that the existing binocular stereo vision technology cannot obtain point cloud data with dense and accurate information, and only uses a binocular camera, which greatly reduces the cost of use. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a flow chart of a preferred embodiment of a point cloud filling method based on a deep neural network in the present invention;

[0034] Figure 2 It is a schematic diagram of a specific point cloud filling method based on a deep neural network disclosed in the present invention;

[0035] Figure 3 It is a functional principle block diagram of a preferred embodiment of a point cloud filling device based on a deep neural network in the present invention;

[0036] Figure 4 It is a functional principle block diagram of a preferred embodiment of the terminal in the present invention. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0038] See also Figure 1 , Figure 1 is a flow chart of the point cloud filling method based on deep neural network in the present invention. Figure 1 As shown, the point cloud filling method based on deep neural network described in the embodiment of the present invention includes:

[0039] Step S11, performing image correction and alignment processing on the binocular image acquired by using the binocular camera to obtain a target binocular image, and calculating a corresponding sparse depth image based on the target binocular image.

[0040] In this embodiment, a binocular image is first acquired using a binocular camera, that is, a binocular image is acquired using a binocular stereo vision technology, such as obtaining a binocular image by shooting the same scene from different angles using two cameras, that is, an original left-eye RGB image and an original right-eye RGB image, wherein the binocular stereo vision technology can be mostly used in robots and driverless cars, and then image correction and alignment are performed to obtain a target binocular image, that is, a target left-eye RGB image and a target right-eye RGB image. In other words, the image acquired by the binocular camera may be distorted and needs to be rectified, such as by using an intrinsic reference matrix, a distortion coefficient, and a rotation matrix to perform dedistortion processing, and after determining the calibration parameters using a checkerboard calibration method, pixel-level hardware correction processing is performed to achieve binocular image acquisition and correction. Image alignment is to ensure that the images of the left and right cameras are on the same plane for subsequent stereo matching and three-dimensional reconstruction, such as by rotating and translating the images of the left and right cameras so that the imaging origin coordinates of the left and right views are consistent. It is understandable that before performing depth calculation, the binocular image acquired by the binocular camera needs to be corrected and aligned first, so as to ensure that the left and right images in the binocular image are on the same plane, thereby facilitating the subsequent stereo matching algorithm.

[0041] It should be pointed out that the binocular camera consists of two horizontally placed cameras, the left eye and the right eye, which are usually regarded as pinhole cameras. That is, the binocular camera consists of two side-by-side sensors, which can capture stereo images and thus achieve depth perception. The centers of their apertures are both located on the x-axis. The distance between the two cameras is called the baseline, which is a key parameter for calculating depth. The binocular camera acquires images based on the principle of parallax. The left and right cameras shoot the same scene from different angles. Due to the difference in viewing angles, the position of the same object in the left and right images will be different. This position difference (parallax) can be used to calculate information such as the distance from the object to the camera.

[0042] In this embodiment, the depth information of the target binocular image is calculated to obtain a sparse depth image. Specifically, the same features between the target binocular images are determined using a preset semi-global matching algorithm to obtain corresponding left and right disparity images, that is, the same features between the target left RGB image and the target right RGB image are determined to obtain left and right disparity images, and then the left and right disparity images are converted into distances to generate corresponding sparse depth images. It can be understood that the depth information is calculated using image matching technology, first finding the same features in the left and right camera images, outputting the disparity map, and then converting the disparity image into a distance to generate a depth map.

[0043] Among them, the preset semi-global matching algorithm is to optimize and adjust its cost aggregation step in advance, so that in the cost aggregation step, two adjacent pixels of different rows are processed in each cycle based on the preset pipeline cost aggregation architecture, and the algorithm for aggregation in a single different direction is realized. It should be pointed out that the matching process of the SGM algorithm includes four steps: preprocessing, cost aggregation, dynamic programming and post-processing. Among them, the cost aggregation step in the SGM (Semi-Global Matching) algorithm occupies more than 90% of the storage resources, and there is a large amount of redundancy in its cost path, so a two-directional cost aggregation strategy can be adopted, but due to the two directions (0° and 135°), aggregation will bring different degrees of circular dependence, that is, the former will bring pixel-level dependence, and the latter will lead to cross-row-level dependence. Therefore, in order to make full use of computing resources and improve speed, the present invention proposes a pipeline cost aggregation architecture, that is, the two-directional cost aggregation strategy of processing a single pixel per cycle is disassembled into processing two adjacent pixels of different rows per cycle, and aggregation in a single different direction. Through this strategy, it can be processed in the form of a pipeline, thereby reducing the computing bottleneck of stereo matching. Moreover, the SGM algorithm has higher matching accuracy than the local matching algorithm because it takes into account the global information of the image. At the same time, the computational efficiency of the SGM algorithm is relatively high and it is suitable for processing large-scale images.

[0044] Step S12: determine the hole parts existing in the sparse depth image, and mark the hole parts in the sparse depth image to obtain a sparse depth image with hole marks.

[0045] In this embodiment, binocular stereo vision technology is used to obtain binocular images and correct the alignment. Then, image matching technology is used to calculate the image depth information, and the hole parts existing in the sparse depth image are determined. The hole parts are marked in the sparse depth image to obtain a sparse depth image with hole marks. Specifically, the left and right disparity images are compared to perform the left and right disparity matching. Figure 1 The consistency check operation obtains the corresponding check result, and the hole part existing in the sparse depth image is determined based on the check result. If the check result shows that there is a target area with inconsistent pixels in the left and right parallax images, the target area with inconsistent pixels is determined as the hole part existing in the sparse depth image. In other words, the left and right parallax images are Figure 1 The consistency check is usually performed by comparing the corresponding pixels in the left and right disparity maps. If a significant difference is found between the left and right disparity maps, it may mean that the depth information of the area is unreliable, that is, the hole part of the image is composed of the position where the matching fails and is blocked and cannot be matched. It can be understood that the image depth information is calculated by using image matching technology and the left and right disparity maps are used to calculate the depth information of the image. Figure 1 The consistency check can obtain a more accurate sparse depth image with hole markers, and the calculation of image depth information and left and right can be realized through a pre-designed dedicated hardware architecture. Figure 1 Consistency check saves hardware resources.

[0046] Step S13: input the sparse depth image with hole marks and the target reference image in the target binocular image into a pre-trained target deep neural network, so that the target deep neural network fills the hole parts in the sparse depth image with the image information provided by the target reference image to obtain a filled dense depth image; wherein the target reference image is an image corresponding to the reference sub-camera in the binocular camera.

[0047] In this embodiment, in order to fill the holes in the sparse depth image, the sparse depth image with hole marks and the target reference image in the target binocular image are input into a pre-trained target deep neural network, so that the target deep neural network uses the image information provided by the target reference image to fill the holes in the sparse depth image to obtain a filled dense depth image.

[0048] Specifically, the target deep neural network can be a target convolutional neural network, that is, the sparse depth image with hole marks and the target reference image in the target binocular image are input into a pre-trained target convolutional neural network with a pixel-level pipeline architecture according to a preset pixel stream input method, so that the target convolutional neural network fills the hole parts in the sparse depth image based on the inter-layer fusion method and uses the image information provided by the target reference image to obtain a filled dense depth image.

[0049] It can be understood that the input of the target deep neural network is the target reference image corresponding to the reference sub-camera in the binocular camera and the sparse depth image with holes. The image is subjected to multi-layer convolution reasoning through a lightweight CNN network. Since the image is input by pixel stream, the use of storage resources can be reduced through the pixel-level pipeline architecture of CNN and the inter-layer fusion method, and the computing resources can be reused in the same layer to make it more real-time. It can also be further lightweighted through quantization to adapt to on-chip integration, that is, the chip is used to realize lightweight instant neural network reasoning that is fully integrated on-chip, and the use of lightweight networks saves computing resources and hardware implementation resources. It should be pointed out that in terms of hardware design, unlike the design of traditional neural network hardware accelerators, the present invention does not use external storage for reading.

[0050] Step S14: Acquire corresponding dense point cloud data based on the filled dense depth image.

[0051] In this embodiment, the sparse depth image with holes is filled through the target deep neural network, so that the corresponding dense point cloud data can be obtained based on the filled dense depth image.

[0052] The brightness value of each pixel in the dense depth image represents the distance from the point to the camera. In order to convert the dense depth image into a point cloud, the intrinsic and extrinsic parameters of the camera are required. For example, the intrinsic parameters include information such as focal length and principal point, while the extrinsic parameters involve the position and posture of the camera in the world coordinate system, which can be obtained through camera calibration. Then, the point cloud coordinates can be calculated based on the dense depth image and camera parameters. That is, for each pixel in the dense depth image, its three-dimensional coordinates in the world coordinate system can be calculated based on its depth value and camera parameters, and based on the principles of geometric transformation and triangulation.

[0053] It can be seen that in the embodiment of the present invention, a binocular image is acquired by a binocular camera, and correction and alignment processing are performed on it. Then, a sparse depth image is calculated based on the image, and the hole parts of the sparse depth image are marked. Finally, the sparse depth image with hole marks and the target reference image in the target binocular image are input into a pre-trained target deep neural network, so that the acquisition of point cloud data with dense and accurate information can be achieved through the filling of the deep neural network. That is, the technical solution of the present application is based on the binocular vision algorithm combined with the deep neural network filling to obtain dense point cloud data, thereby solving the problem that the existing binocular stereo vision technology cannot obtain point cloud data with dense and accurate information, and only using a binocular camera greatly reduces the use cost.

[0054] For example, see Figure 2 As shown in the figure, a binocular camera is first used to obtain binocular images, namely, the right RGB image and the left RGB image, and image correction and alignment are performed. Then, a preset semi-global matching algorithm is used to calculate the depth image. The semi-global matching algorithm is an efficient stereo matching algorithm. It sets a global energy function and combines the dynamic programming method to solve the optimal disparity value of each pixel, and uses the left and right disparities to calculate the depth image. Figure 1 The consistency check obtains a more accurate depth image with holes. In order to fill the holes in the depth image, the depth image with holes and the right-eye RGB image corresponding to the reference sub-camera of the binocular camera after correction and alignment are input into the deep neural network. The deep neural network obtains the filled dense depth image by learning the information provided by the right-eye RGB image, thereby obtaining the point cloud information, which is the technical solution of the present application. The binocular vision algorithm based on the preset semi-global stereo matching is combined with the deep neural network to realize the acquisition of point cloud data with dense and accurate information, which reduces the reasoning difficulty of the neural network while ensuring the accuracy, lightweights the network, and saves computing resources and hardware implementation resources.

[0055] In one embodiment, if Figure 3 As shown, based on the above-mentioned point cloud filling method based on deep neural network, the present invention also provides a point cloud filling device based on deep neural network, including:

[0056] The image processing module 11 is used to perform image correction and alignment processing on the binocular image acquired by the binocular camera to obtain a target binocular image;

[0057] A depth image calculation module 12, used to calculate a corresponding sparse depth image based on the target binocular image;

[0058] A hole determination module 13 is used to determine a hole portion existing in the sparse depth image, and mark the hole portion in the sparse depth image to obtain a sparse depth image with a hole mark;

[0059] The point cloud filling module 14 is used to input the sparse depth image with hole marks and the target reference image in the target binocular image into a pre-trained target deep neural network, so that the target deep neural network fills the hole part in the sparse depth image with the image information provided by the target reference image to obtain a filled dense depth image; wherein the target reference image is an image corresponding to the reference sub-camera in the binocular camera;

[0060] The point cloud data acquisition module 15 is used to acquire corresponding dense point cloud data based on the filled dense depth image.

[0061] Figure 4 A schematic diagram of the structure of a terminal provided in an embodiment of the present application. The terminal may include:

[0062] A memory 501 , a processor 502 , and a computer program stored in the memory 501 and executable on the processor 502 .

[0063] When the processor 502 executes the program, the point cloud filling method based on the deep neural network provided in the above embodiment is implemented.

[0064] Furthermore, the terminal further includes:

[0065] The communication interface 503 is used for communication between the memory 501 and the processor 502 .

[0066] The memory 501 is used to store computer programs that can be executed on the processor 502 .

[0067] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0068] If the memory 501, the processor 502 and the communication interface 503 are implemented independently, the communication interface 503, the memory 501 and the processor 502 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0069] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.

[0070] The processor 502 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0071] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned point cloud filling method based on a deep neural network.

[0072] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art that are not disclosed in this application. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present invention are indicated by the claims.

[0073] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0074] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor or other system that can read instructions from an instruction execution system, apparatus or device and execute instructions), or used in combination with these instruction execution systems, apparatuses or devices.

[0075] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA, Programmable Gate Array), a field programmable gate array (FPGA, Field-Programmable Gate Array), etc.

[0076] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A point cloud filling method based on deep neural network, characterized in that: The method comprises: Performing image correction and alignment processing on the binocular image acquired by the binocular camera to obtain a target binocular image, and calculating a corresponding sparse depth image based on the target binocular image; Determine a hole portion existing in the sparse depth image, and mark the hole portion in the sparse depth image to obtain a sparse depth image with a hole mark; Inputting the sparse depth image with hole marks and the target reference image in the target binocular image into a pre-trained target deep neural network, so that the target deep neural network fills the hole parts in the sparse depth image with image information provided by the target reference image to obtain a filled dense depth image; wherein the target reference image is an image corresponding to the reference sub-camera in the binocular camera; Corresponding dense point cloud data is obtained based on the filled dense depth image.

2. The point cloud filling method based on deep neural network according to claim 1 is characterized in that: The step of calculating a corresponding sparse depth image based on the target binocular image includes: Determine the same features between the target binocular images using a preset semi-global matching algorithm to obtain corresponding left and right disparity images; The left and right disparity images are converted into distances to generate corresponding sparse depth images.

3. The point cloud filling method based on deep neural network according to claim 2 is characterized in that: The determining of the hole portion existing in the sparse depth image comprises: Performing a left-right image consistency check operation by comparing corresponding pixel points in the left-right disparity images to obtain a corresponding check result; Based on the inspection result, a hole portion existing in the sparse depth image is determined.

4. The point cloud filling method based on deep neural network according to claim 3 is characterized in that: The determining, based on the inspection result, a hole portion existing in the sparse depth image comprises: If the inspection result indicates that there is a target area with inconsistent pixels in the left and right parallax images, the target area with inconsistent pixels is determined as a hole portion existing in the sparse depth image.

5. The point cloud filling method based on deep neural network according to claim 2 is characterized in that: The preset semi-global matching algorithm is an algorithm that optimizes and adjusts its cost aggregation step in advance, so that in the cost aggregation step, two adjacent pixels of different rows are processed per cycle based on a preset pipeline cost aggregation architecture, and aggregation in a single different direction is realized.

6. The point cloud filling method based on deep neural network according to any one of claims 1 to 5, characterized in that: The target deep neural network is a target convolutional neural network.

7. The point cloud filling method based on deep neural network according to claim 6 is characterized in that: The step of inputting the sparse depth image with hole marks and the target reference image in the target binocular image into a pre-trained target deep neural network, so that the target deep neural network fills the hole parts in the sparse depth image with image information provided by the target reference image to obtain a filled dense depth image, including: The sparse depth image with hole marks and the target reference image in the target binocular image are input into a pre-trained target convolutional neural network with a pixel-level pipeline architecture according to a preset pixel stream input method, so that the target convolutional neural network fills the hole parts in the sparse depth image based on an inter-layer fusion method and uses the image information provided by the target reference image to obtain a filled dense depth image.

8. A point cloud filling device based on deep neural network, characterized in that: The device comprises: An image processing module is used to perform image correction and alignment processing on the binocular image acquired by the binocular camera to obtain a target binocular image; A depth image calculation module, used to calculate a corresponding sparse depth image based on the target binocular image; A hole determination module, used for determining a hole portion existing in the sparse depth image, and marking the hole portion in the sparse depth image to obtain a sparse depth image with a hole mark; A point cloud filling module, used for inputting the sparse depth image with hole marks and the target reference image in the target binocular image into a pre-trained target deep neural network, so that the target deep neural network fills the hole parts in the sparse depth image with the image information provided by the target reference image to obtain a filled dense depth image; wherein the target reference image is an image corresponding to the reference sub-camera in the binocular camera; The point cloud data acquisition module is used to acquire corresponding dense point cloud data based on the filled dense depth image.

9. A terminal, characterized in that: include: A memory, a processor, and a point cloud filling program based on a deep neural network stored in the memory and executable on the processor, wherein the point cloud filling program based on a deep neural network, when executed by the processor, implements the steps of the point cloud filling method based on a deep neural network as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which can be executed to implement the steps of the point cloud filling method based on a deep neural network as described in any one of claims 1 to 7.