Spatial neural network depth completion method, system, device and storage medium
By adopting the spatial neural network depth completion method with adaptive diffusion cores in ToF technology, the challenges faced by the depth completion algorithm in the existing technology are solved, high-precision and high-efficiency depth completion are achieved, and high-quality dense depth images are generated.
Patent Information
- Application Number
- CN202011603338.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-29
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-12-29
AI Technical Summary
In the existing ToF technology, although the point scanning projection method can achieve high-precision depth measurement, it requires a complex projection scanning system and is cost-effective; while the surface light projection method is low in optical power density and reduced signal-to-noise ratio, and is only suitable for scenarios with reduced distance and lower accuracy. At the same time, the depth images obtained by the point scanning method have sparse characteristics, resulting in challenges in the depth completion algorithm.
The spatial neural network depth completion method based on adaptive diffusion core is adopted. By acquiring RGB images and sparse depth speckle images, the pre-trained depth completion model, including the U-shaped network model and the diffusion network model, is used to generate adjacent pixel similarity matrix, and select appropriate diffusion patterns according to the matrix for depth completion to generate dense depth speckle images.
It significantly improves the accuracy and calculation speed of depth completion, can effectively process sparse depth images and generate high-quality dense depth images.
Smart Images

Figure CN114693757B_ABST
Abstract
Description
Background Art
[0002] ToF (time of flight) technology is a 3D imaging technology that emits measurement light from a projector and reflects the measurement light back to the receiver through the target object, thereby obtaining the spatial distance from the object to the sensor based on the propagation time of the measurement light in this propagation path. Commonly used ToF technologies include point scanning projection method and surface light projection method.
[0003] The ToF method of point scanning projection uses a point projector to project multiple beams of collimated light. The projection direction of the multiple beams of collimated light is controlled by the scanning device so that they can be projected to different target positions. After a single beam of collimated light is reflected by the target object, part of the light is received by the light detector to obtain the depth measurement data of the current projection direction. This method can concentrate all the optical power on multiple target points, thereby achieving a high signal-to-noise ratio at the target point, and then achieving high-precision depth measurement. The scanning of the entire target object depends on scanning devices, such as mechanical motors, MEMS, optical phased radar, etc. The discrete point cloud data required for 3D imaging can be obtained by splicing the depth data points obtained by the scan. This method is conducive to long-distance 3D imaging, but it requires the use of a complex projection scanning system, which is costly.
[0004] The ToF method of surface light projection projects a surface light beam with a continuous energy distribution. The projected light continuously covers the surface of the target object. The light detector is a light detector array that can obtain the propagation time of the light beam. When the light signal reflected by the target object is imaged on the light detector through the optical imaging system, the depth obtained by each detector image point is the depth information of the object position corresponding to its object-image relationship. This method can get rid of the complex scanning system. However, since the light power density of surface light projection is much lower than that of a single collimated light, the signal-to-noise ratio is greatly reduced compared to the single-point scanning projection method, making this method only applicable to scenes with reduced distance and low accuracy.
[0005] However, the depth image obtained by the ToF method of point scanning projection is sparse, which brings algorithmic challenges to depth completion. Summary of the invention
[0006] In view of the defects in the prior art, the purpose of the present invention is to provide a spatial neural network depth completion method, system, device and storage medium based on adaptive diffusion kernel.
[0007] The spatial neural network depth completion method based on adaptive diffusion kernel provided by the present invention comprises the following steps:
[0008] Step S1: acquiring an RGB image and a sparse depth speckle image, wherein the RGB image and the sparse depth speckle image are acquired by an RGB camera and a depth camera respectively;
[0009] Step S2: obtaining a pre-trained depth completion model, wherein the depth completion model includes a U-shaped network model and a diffusion network model, and the diffusion network model includes a plurality of pre-set diffusion patterns;
[0010] Step S3: using the U-shaped network model to generate an adjacent pixel similarity matrix corresponding to each pixel in the sparse depth speckle image for the input RGB image and the sparse depth speckle image, and using the diffusion network model to select a corresponding diffusion pattern for each pixel of the sparse depth speckle image according to the adjacent pixel similarity matrix to perform depth completion to generate a dense depth speckle image.
[0011] Preferably, step S1 comprises the following steps:
[0012] Step S101: projecting a dot matrix light toward the target person through the beam projector end of the depth camera, and receiving the dot matrix light reflected by the target person through the detector end of the depth camera;
[0013] Step S102: the depth camera generates a sparse depth speckle image of the target person according to the dot matrix light received by the detector end;
[0014] Step S103: capturing an RGB image of the target person through an RGB camera.
[0015] Preferably, the depth completion model is trained and generated by the following method:
[0016] Step M101: obtaining an RGB image training set and a sparse depth speckle image training set, wherein the RGB image and the sparse depth speckle image are respectively generated by capturing a target person through an RGB camera and a depth camera;
[0017] Step M102: inputting the RGB image training set and the sparse depth speckle image training set into a depth completion model based on a convolutional neural network to generate a depth pre-completed speckle image;
[0018] Step M103: determining a loss function of the depth pre-complemented speckle image according to a preset standard depth speckle image, wherein the standard depth speckle image is a pre-collected dense depth speckle image of the target person;
[0019] Step M104: Repeat steps M101 to M103 until the loss function reaches a preset loss threshold range.
[0020] Preferably, step S3 comprises the following steps:
[0021] Step S301: the U-shaped network model generates an adjacent pixel similarity matrix corresponding to each pixel in the sparse deep speckle image according to the input RGB image and the sparse deep speckle image;
[0022] Step S302: selecting a corresponding diffusion pattern from the plurality of pre-set diffusion patterns according to the adjacent pixel similarity matrix corresponding to each pixel;
[0023] Step S303: for each pixel, calculating the depth value of the pixel located at the center according to the corresponding diffusion pattern, that is, achieving depth completion of the pixel;
[0024] Step S304: Repeat steps S302 to S303 to generate a dense depth speckle image.
[0025] Preferably, the diffusion pattern includes the following multiple patterns:
[0026] Diffusion pattern of central pixel depth calculation through eight neighboring pixels;
[0027] The diffusion pattern of the central pixel depth calculation is performed by two symmetrical pixels in the eight-neighborhood;
[0028] A diffusion pattern is obtained by calculating the depth of the central pixel from any three pixels in the eight-neighborhood;
[0029] A diffusion pattern for calculating the depth of a central pixel by any at least 8 pixels in a 5×5 pixel matrix;
[0030] The diffusion pattern of the center pixel depth calculation is performed by any at least 10 pixels in the 7×7 pixel matrix.
[0031] Preferably, the U-shaped network includes a convolutional network and a deconvolutional network, and the convolutional network and the deconvolutional network are connected to form a U-shaped structure;
[0032] The convolution network and the deconvolution network respectively include a plurality of layers of combined convolution blocks, and the combined convolution blocks are used to extract features from the input RGB image and the sparse depth speckle image;
[0033] The combined convolution block includes a plurality of convolution kernels of different sizes.
[0034] Preferably, step S301 includes the following steps:
[0035] Step S3011: traverse the similarity values between each pixel and the pixels in the surrounding eight neighborhoods. When there are at least two similar pixels between a pixel and the surrounding eight neighborhoods, generate a 3×3 adjacent pixel similarity matrix. Otherwise, execute step S3012. The similarity value is determined by the difference between the pixel values of the two pixels.
[0036] Step S3012: traverse each pixel in the 5×5 pixel matrix centered on each pixel to find similar pixels. When there are at least eight similar pixels, generate a 5×5 adjacent pixel similarity matrix, otherwise execute step S3013;
[0037] Step S3013: traverse each pixel in the 7×7 pixel matrix centered on each pixel to find similar pixels. When there are at least 10 similar pixels, generate a 7×7 adjacent pixel similarity matrix and remove the pixel as a noise point.
[0038] The spatial neural network depth completion system based on adaptive diffusion kernel provided by the present invention includes the following modules:
[0039] An image acquisition module, used to acquire an RGB image and a sparse depth speckle image, wherein the RGB image and the sparse depth speckle image are acquired by an RGB camera and a depth camera respectively;
[0040] A model acquisition module, used to acquire a pre-trained depth completion model, wherein the depth completion model includes a U-shaped network model and a diffusion network model, and the diffusion network model includes a plurality of pre-set diffusion patterns;
[0041] A depth completion module is used to generate an adjacent pixel similarity matrix corresponding to each pixel in the sparse depth speckle image for the input RGB image and the sparse depth speckle image through the U-shaped network model, and to perform depth completion on each pixel of the sparse depth speckle image according to the adjacent pixel similarity matrix by selecting a corresponding diffusion pattern through the diffusion network model to generate a dense depth speckle image.
[0042] The spatial neural network depth completion device based on adaptive diffusion kernel provided by the present invention includes:
[0043] processor;
[0044] a memory storing executable instructions of the processor;
[0045] Wherein, the processor is configured to perform the steps of the adaptive diffusion kernel-based spatial neural network depth completion method by executing the executable instructions.
[0046] The computer-readable storage medium provided according to the present invention is used to store a program, and when the program is executed, the steps of the spatial neural network depth completion method based on adaptive diffusion kernel are implemented.
[0047] Compared with the prior art, the present invention has the following beneficial effects:
[0048] The present invention calculates the adjacent pixel similarity matrix for each pixel in the sparse depth speckle image according to the RGB image of the target person and the sparse depth speckle image, and then selects a diffusion pattern according to the adjacent pixel similarity matrix, thereby calculating the depth of each pixel point according to the selected diffusion pattern to generate a dense depth speckle image. The present invention can not only significantly improve the accuracy of depth completion, but also significantly improve the calculation speed during depth completion. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings in the following descriptions are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without creative work. By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, purposes and advantages of the present invention will become more obvious:
[0050] Figure 1 It is a flowchart of the steps of the spatial neural network depth completion method based on the adaptive diffusion kernel in an embodiment of the present invention;
[0051] Figure 2 This is a flow chart of the steps of acquiring RGB images and sparse depth speckle images in an embodiment of the present invention;
[0052] Figure 3 A flowchart of the steps of performing depth completion to generate a dense depth speckle image in an embodiment of the present invention;
[0053] Figure 4 A flowchart of the steps of generating a similarity matrix of adjacent pixels in an embodiment of the present invention;
[0054] Figure 5 This is a flowchart of the steps of training a depth completion model in an embodiment of the present invention;
[0055] Figure 6 (a) is a schematic diagram of a first diffusion pattern in an embodiment of the present invention;
[0056] Figure 6 (b) is a schematic diagram of a second diffusion pattern in an embodiment of the present invention;
[0057] Figure 6 (c) is a schematic diagram of a third diffusion pattern in an embodiment of the present invention;
[0058] Figure 6 (d) is a schematic diagram of a fourth diffusion pattern in an embodiment of the present invention;
[0059] Figure 6(e) is a schematic diagram of a first diffusion pattern in an embodiment of the present invention;
[0060] Figure 6 (f) is a schematic diagram of a first diffusion pattern in an embodiment of the present invention;
[0061] Figure 6 (g) is a schematic diagram of a first diffusion pattern in an embodiment of the present invention;
[0062] Figure 6 (h) is a schematic diagram of a first diffusion pattern in an embodiment of the present invention;
[0063] Figure 6 (i) is a schematic diagram of a first diffusion pattern in an embodiment of the present invention;
[0064] Figure 6 (k) is a schematic diagram of a first diffusion pattern in an embodiment of the present invention;
[0065] Figure 6 (l) is a schematic diagram of a first diffusion pattern in an embodiment of the present invention;
[0066] Figure 7 It is a module schematic diagram of a spatial neural network deep completion system based on an adaptive diffusion kernel in an embodiment of the present invention;
[0067] Figure 8 is a schematic diagram of the structure of a spatial neural network depth completion device based on an adaptive diffusion kernel in an embodiment of the present invention; and
[0068] Fig. 9 Schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention. DETAILED DESCRIPTION
[0069] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements may be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0070] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0071] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0072] The spatial neural network depth completion method based on adaptive diffusion kernel provided by the present invention is intended to solve the problems existing in the prior art.
[0073] The following specific embodiments are used to describe in detail the technical solutions of the present invention and how the technical solutions of the present application solve the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.
[0074] Figure 1 FIG. 1 is a flowchart of the steps of the spatial neural network depth completion method based on the adaptive diffusion kernel in an embodiment of the present invention. Figure 1 As shown, the spatial neural network depth completion method based on adaptive diffusion kernel provided by the present invention includes the following steps:
[0075] Step S1: acquiring an RGB image and a sparse depth speckle image, wherein the RGB image and the sparse depth speckle image are acquired by an RGB camera and a depth camera respectively;
[0076] Figure 2 FIG. 1 is a flow chart of steps for collecting RGB images and sparse depth speckle images in an embodiment of the present invention. Figure 2 As shown, step S1 includes the following steps:
[0077] Step S101: projecting a dot matrix light toward the target person through the beam projector end of the depth camera, and receiving the dot matrix light reflected by the target person through the detector end of the depth camera;
[0078] Step S102: the depth camera generates a sparse depth speckle image of the target person according to the dot matrix light received by the detector end;
[0079] Step S103: capturing an RGB image of the target person through an RGB camera.
[0080] In the embodiment of the present invention, the beam projector includes a light source, a light source driver and a beam splitter; the light source driver is connected to the light source and is used to drive the light source to emit light; the beam splitter is used to project multiple discrete collimated light beams from the light source. The beam splitter 205 can be a diffraction grating (DOE), a spatial light modulator (SLM), etc.
[0081] In an embodiment of the present invention, the light receiving module adopts an infrared camera, including a lens, a filter and an image sensor arranged along the light path. The image sensor is used to receive the dot matrix light through at least four receiving windows; the at least four receiving windows are arranged in sequence at equal intervals in time sequence, and the sparse depth speckle image is calculated according to the infrared speckle images received by the four receiving windows.
[0082] Step S2: obtaining a pre-trained depth completion model, wherein the depth completion model includes a U-shaped network model and a diffusion network model, and the diffusion network model includes a plurality of pre-set diffusion patterns;
[0083] Figure 3 FIG. 1 is a flow chart of the steps of performing depth completion to generate a dense depth speckle image in an embodiment of the present invention. Figure 3 As shown, the depth completion model is trained and generated by the following method:
[0084] Step M101: obtaining an RGB image training set and a sparse depth speckle image training set, wherein the RGB image and the sparse depth speckle image are respectively generated by capturing a target person through an RGB camera and a depth camera;
[0085] Step M102: inputting the RGB image training set and the sparse depth speckle image training set into a depth completion model based on a convolutional neural network to generate a depth pre-completed speckle image;
[0086] Step M103: determining a loss function of the depth pre-complemented speckle image according to a preset standard depth speckle image, wherein the standard depth speckle image is a pre-collected dense depth speckle image of the target person;
[0087] Step M104: Repeat steps M101 to M103 until the loss function reaches a preset loss threshold range.
[0088] In the embodiment of the present invention, the loss threshold may be set to any value between 3% and 10%, such as 5%.
[0089] In an embodiment of the present invention, the U-shaped network includes a convolutional network and a deconvolutional network, and the convolutional network and the deconvolutional network are connected to form a U-shaped structure;
[0090] The convolution network and the deconvolution network respectively include a plurality of layers of combined convolution blocks, and the combined convolution blocks are used to extract features from the input RGB image and the sparse depth speckle image;
[0091] The combined convolution block includes a plurality of convolution kernels of different sizes.
[0092] Step S3: using the U-shaped network model to generate an adjacent pixel similarity matrix corresponding to each pixel in the sparse depth speckle image for the input RGB image and the sparse depth speckle image, and using the diffusion network model to select a corresponding diffusion pattern for each pixel of the sparse depth speckle image according to the adjacent pixel similarity matrix to perform depth completion to generate a dense depth speckle image.
[0093] In an embodiment of the present invention, the diffusion network model adopts a convolutional spatial propagation network module.
[0094] Figure 5 is a flowchart of the steps of training the depth completion model in an embodiment of the present invention, such as Figure 5 As shown, step S3 includes the following steps:
[0095] Step S301: the U-shaped network model generates an adjacent pixel similarity matrix corresponding to each pixel in the sparse deep speckle image according to the input RGB image and the sparse deep speckle image;
[0096] In the embodiment of the present invention, the adjacent pixel similarity matrix may be a 3×3 pixel matrix, or a 5×5 or 7×7 pixel matrix.
[0097] When the difference between the depth values of a pixel and the central pixel in the similarity matrix is within 10%, the pixel is retained in the adjacent pixel similarity matrix; otherwise, the pixel is deleted.
[0098] Figure 4 FIG. 1 is a flowchart of the steps of generating a similarity matrix of adjacent pixels in an embodiment of the present invention. Figure 4 As shown, step S301 includes the following steps:
[0099] Step S3011: traverse the similarity values between each pixel and the pixels in the surrounding eight neighborhoods. When there are at least two similar pixels between a pixel and the surrounding eight neighborhoods, generate a 3×3 adjacent pixel similarity matrix. Otherwise, execute step S3012. The similarity value is determined by the difference between the pixel values of the two pixels.
[0100] Step S3012: traverse each pixel in the 5×5 pixel matrix centered on each pixel to find similar pixels. When there are at least eight similar pixels, generate a 5×5 adjacent pixel similarity matrix, otherwise execute step S3013;
[0101] Step S3013: traverse each pixel in the 7×7 pixel matrix centered on each pixel to find similar pixels. When there are at least 10 similar pixels, generate a 7×7 adjacent pixel similarity matrix and remove the pixel as a noise point.
[0102] In the embodiment of the present invention, when the difference between the depth values of a pixel point and the central pixel point is within 10%, the two pixels are considered to be similar.
[0103] Step S302: selecting a corresponding diffusion pattern from the plurality of pre-set diffusion patterns according to the adjacent pixel similarity matrix corresponding to each pixel;
[0104] In the embodiment of the present invention, when selecting the corresponding diffusion pattern, the selection is made based on the morphological similarity between the diffusion pattern and the adjacent pixel similarity matrix. For example, when a pixel is similar to eight adjacent pixels, the diffusion pattern of the pixel is a 3×3 matrix; when a pixel is similar to the adjacent left, right, and upper pixels, the diffusion pattern of the pixel is Figure 6 For the matrix shown in (a), if a pixel is similar to two pixels on the diagonal line, the diffusion pattern of the pixel is Figure 6 (g) or Figure 6 (h) The matrix shown.
[0105] Step S303: for each pixel, calculating the depth value of the pixel located at the center according to the corresponding diffusion pattern, that is, achieving depth completion of the pixel;
[0106] In the embodiment of the present invention, when calculating the depth value according to the diffusion pattern, the pixel value of the central pixel is generated by performing weighted averaging according to the depth value corresponding to each pixel grid in the diffusion pattern.
[0107] Step S304: Repeat steps S302 to S303 to generate a dense depth speckle image.
[0108] Figure 6 is a schematic diagram of a diffusion pattern in an embodiment of the present invention, such as Figure 6As shown, in the embodiment of the present invention, the diffusion pattern includes the following multiple patterns:
[0109] Diffusion pattern of central pixel depth calculation through eight neighboring pixels;
[0110] The diffusion pattern of the central pixel depth calculation is performed by two symmetrical pixels in the eight-neighborhood, such as Figure 6 (g) Figure 6 (h)
[0111] The diffusion pattern of the central pixel depth calculation is performed by any at least three pixels in the eight-neighborhood, such as Figure 6 (a) Figure 6 (b) Figure 6 (c) Figure 6 (d) Figure 6 (e) Figure 6 (f)
[0112] The diffusion pattern of the center pixel depth calculation is performed by any at least 8 pixels in the 5×5 pixel matrix, such as Figure 6 (i) Figure 6 (j)
[0113] The diffusion pattern of the central pixel depth calculation is performed by any at least 10 pixels in the 7×7 pixel matrix, such as Figure 6 (k) Figure 6 (l) as shown.
[0114] Figure 7 Schematic diagram of a module of a spatial neural network depth completion system based on an adaptive diffusion kernel in an embodiment of the present invention. Figure 7 As shown, the spatial neural network depth completion system based on adaptive diffusion kernel provided by the present invention includes the following modules:
[0115] An image acquisition module, used to acquire an RGB image and a sparse depth speckle image, wherein the RGB image and the sparse depth speckle image are acquired by an RGB camera and a depth camera respectively;
[0116] A model acquisition module, used to acquire a pre-trained depth completion model, wherein the depth completion model includes a U-shaped network model and a diffusion network model, and the diffusion network model includes a plurality of pre-set diffusion patterns;
[0117] A depth completion module is used to generate an adjacent pixel similarity matrix corresponding to each pixel in the sparse depth speckle image for the input RGB image and the sparse depth speckle image through the U-shaped network model, and to perform depth completion on each pixel of the sparse depth speckle image according to the adjacent pixel similarity matrix by selecting a corresponding diffusion pattern through the diffusion network model to generate a dense depth speckle image.
[0118] The embodiment of the present invention also provides a liveness detection device based on a facial spot image, comprising a processor and a memory, wherein executable instructions of the processor are stored. The processor is configured to execute the steps of a liveness detection method based on a facial spot image by executing the executable instructions.
[0119] As mentioned above, in this embodiment, by collecting the spot image of the target person, the pixel area intercepted on the spot image, the spot clarity of the pixel area is calculated, and the distance information between the target person and the depth camera is determined according to the spot clarity and the preset spot clarity and distance information associated with the distance generation model. The depth information of the object can be obtained more quickly, and can be used in consumer products such as mobile phones, somatosensory games, and payments that obtain close-range facial depth information.
[0120] It will be appreciated by those skilled in the art that various aspects of the present invention may be implemented as systems, methods or program products. Therefore, various aspects of the present invention may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits", "modules" or "platforms".
[0121] Figure 8 Schematic diagram of the structure of the spatial neural network depth completion device based on adaptive diffusion kernel in the embodiment of the present invention. Figure 8 The electronic device 600 according to this embodiment of the present invention is described. Figure 8 The electronic device 600 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0122] like Figure 8 As shown, the electronic device 600 is in the form of a general computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.
[0123] The storage unit stores program codes, which can be executed by the processing unit 610, so that the processing unit 610 executes the steps of various exemplary embodiments of the present invention described in the above-mentioned method for detecting a living body based on a facial spot image. For example, the processing unit 610 can execute the following steps: Figure 1 Follow the steps shown in .
[0124] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .
[0125] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0126] Bus 630 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0127] The electronic device 600 may also communicate with one or more external devices 700 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device 600, and / or any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed through an input / output (I / O) interface 650. Furthermore, the electronic device 600 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 660. The network adapter 660 may communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although Figure 8 Not shown, other hardware and / or software modules may be used in conjunction with electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.
[0128] In an embodiment of the present invention, a computer-readable storage medium is also provided for storing a program, and the steps of the liveness detection method based on the face spot image are implemented when the program is executed. In some possible implementations, various aspects of the present invention can also be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of various exemplary embodiments of the present invention described in the above-mentioned liveness detection method based on the face spot image section of this specification.
[0129] As shown above, when the program of the computer-readable storage medium of this embodiment is executed, by collecting the spot image of the target person, the pixel area intercepted on the spot image, the spot clarity of the pixel area is calculated, and the distance information between the target person and the depth camera is determined according to the spot clarity and the preset spot clarity and distance information associated with the distance generation model. The depth information of the object can be obtained more quickly, and can be used in consumer products such as mobile phones, somatosensory games, and payments that obtain close-range facial depth information.
[0130] Fig. 9 Schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention. Fig. 9 As shown, a program product 800 for implementing the above method according to an embodiment of the present invention is described, which can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.
[0131] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0132] Computer readable storage media may include data signals propagated in baseband or as part of a carrier wave, wherein readable program codes are carried. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program codes contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.
[0133] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0134] In the embodiment of the present invention, the adjacent pixel similarity matrix is calculated for each pixel in the sparse depth speckle image according to the RGB image of the target person and the sparse depth speckle image, and then the diffusion pattern is selected according to the adjacent pixel similarity matrix, so that the depth of each pixel is calculated according to the selected diffusion pattern to generate a dense depth speckle image. The present invention can not only significantly improve the accuracy of depth completion, but also significantly improve the calculation speed during depth completion.
[0135] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same and similar parts between the embodiments can be referred to each other. The above description of the disclosed embodiments enables professionals and technicians in this field to implement or use the present invention. Various modifications to these embodiments will be obvious to professionals and technicians in this field, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown in this article, but will comply with the widest range consistent with the principles and novel features disclosed herein.
[0136] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art may make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.
Claims
1. A spatial neural network depth completion method based on adaptive diffusion kernel, characterized in that: The steps include: Step S1: acquiring an RGB image and a sparse depth speckle image, wherein the RGB image and the sparse depth speckle image are acquired by an RGB camera and a depth camera respectively; Step S2: obtaining a pre-trained depth completion model, wherein the depth completion model includes a U-shaped network model and a diffusion network model, and the diffusion network model includes a plurality of pre-set diffusion patterns; Step S3: generating an adjacent pixel similarity matrix corresponding to each pixel in the sparse depth speckle image for the input RGB image and the sparse depth speckle image through the U-shaped network model, and performing depth completion on each pixel of the sparse depth speckle image by selecting a corresponding diffusion pattern according to the adjacent pixel similarity matrix through the diffusion network model to generate a dense depth speckle image; The step S3 comprises the following steps: Step S301: the U-shaped network model generates an adjacent pixel similarity matrix corresponding to each pixel in the sparse deep speckle image according to the input RGB image and the sparse deep speckle image; Step S302: selecting a corresponding diffusion pattern from the plurality of pre-set diffusion patterns according to the adjacent pixel similarity matrix corresponding to each pixel; Step S303: for each pixel, calculating the depth value of the pixel located at the center according to the corresponding diffusion pattern, that is, achieving depth completion of the pixel; When calculating the depth value according to the diffusion pattern, a pixel value of a central pixel is generated by performing weighted averaging according to the depth value corresponding to each pixel grid in the diffusion pattern; Step S304: Repeat steps S302 to S303 to generate a dense depth speckle image; The step S301 includes the following steps: Step S3011: traverse the similarity values between each pixel and the pixels in the surrounding eight neighborhoods. When there are at least two similar pixels between a pixel and the surrounding eight neighborhoods, generate a 3×3 adjacent pixel similarity matrix, otherwise execute step S3012, the similarity value is determined by the difference in pixel values between the two pixels; the similar pixel refers to a pixel whose depth value difference with the central pixel is within 10%; Step S3012: traverse each pixel in the 5×5 pixel matrix centered on each pixel to find similar pixels. When there are at least eight similar pixels, generate a 5×5 adjacent pixel similarity matrix, otherwise execute step S3013; Step S3013: traverse each pixel in the 7×7 pixel matrix centered on each pixel to find similar pixels. When there are at least 10 similar pixels, generate a 7×7 adjacent pixel similarity matrix and remove the pixel as a noise point.
2. The method for spatial neural network depth completion based on adaptive diffusion kernel according to claim 1, characterized in that: The step S1 comprises the following steps: Step S101: projecting a dot matrix light toward a target person through a beam projector end of a depth camera, and receiving the dot matrix light reflected by the target person through a detector end of the depth camera; Step S102: the depth camera generates a sparse depth speckle image of the target person according to the dot matrix light received by the detector end; Step S103: capturing an RGB image of the target person through an RGB camera.
3. The method for spatial neural network depth completion based on adaptive diffusion kernel according to claim 1, characterized in that: The depth completion model is trained and generated by the following method: Step M101: obtaining an RGB image training set and a sparse depth speckle image training set, wherein the RGB image and the sparse depth speckle image are respectively generated by capturing a target person through an RGB camera and a depth camera; Step M102: inputting the RGB image training set and the sparse depth speckle image training set into a depth completion model based on a convolutional neural network to generate a depth pre-completed speckle image; Step M103: determining a loss function of the depth pre-complemented speckle image according to a preset standard depth speckle image, wherein the standard depth speckle image is a pre-collected dense depth speckle image of the target person; Step M104: Repeat steps M101 to M103 until the loss function reaches a preset loss threshold range.
4. The method for spatial neural network depth completion based on adaptive diffusion kernel according to claim 1, characterized in that: The diffusion pattern includes the following patterns: Diffusion pattern of central pixel depth calculation through eight neighboring pixels; The diffusion pattern of the central pixel depth calculation is performed by two symmetrical pixels in the eight-neighborhood; A diffusion pattern is obtained by calculating the depth of the central pixel from any three pixels in the eight-neighborhood; A diffusion pattern for calculating the depth of a central pixel by any at least 8 pixels in a 5×5 pixel matrix; The diffusion pattern of the center pixel depth calculation is performed by any at least 10 pixels in the 7×7 pixel matrix.
5. The method for spatial neural network depth completion based on adaptive diffusion kernel according to claim 1, characterized in that: The U-shaped network includes a convolutional network and a deconvolutional network, and the convolutional network and the deconvolutional network are connected to form a U-shaped structure; The convolution network and the deconvolution network respectively include a plurality of layers of combined convolution blocks, and the combined convolution blocks are used to extract features from the input RGB image and the sparse depth speckle image; The combined convolution block includes a plurality of convolution kernels of different sizes.
6. A spatial neural network depth completion system based on adaptive diffusion kernel, characterized in that: Includes the following modules: An image acquisition module, used to acquire an RGB image and a sparse depth speckle image, wherein the RGB image and the sparse depth speckle image are acquired by an RGB camera and a depth camera respectively; A model acquisition module, used to acquire a pre-trained depth completion model, wherein the depth completion model includes a U-shaped network model and a diffusion network model, and the diffusion network model includes a plurality of pre-set diffusion patterns; a depth completion module, configured to generate an adjacent pixel similarity matrix corresponding to each pixel in the sparse depth speckle image for the input RGB image and the sparse depth speckle image through the U-shaped network model, and to perform depth completion on each pixel of the sparse depth speckle image according to the adjacent pixel similarity matrix by selecting a corresponding diffusion pattern through the diffusion network model to generate a dense depth speckle image; The depth completion module includes the following steps during processing: Step S301: the U-shaped network model generates an adjacent pixel similarity matrix corresponding to each pixel in the sparse deep speckle image according to the input RGB image and the sparse deep speckle image; Step S302: selecting a corresponding diffusion pattern from the plurality of pre-set diffusion patterns according to the adjacent pixel similarity matrix corresponding to each pixel; Step S303: for each pixel, calculating the depth value of the pixel located at the center according to the corresponding diffusion pattern, that is, achieving depth completion of the pixel; When calculating the depth value according to the diffusion pattern, a pixel value of a central pixel is generated by performing weighted averaging according to the depth value corresponding to each pixel grid in the diffusion pattern; Step S304: Repeat steps S302 to S303 to generate a dense depth speckle image; The step S301 includes the following steps: Step S3011: traverse the similarity values between each pixel and the pixels in the surrounding eight neighborhoods. When there are at least two similar pixels between a pixel and the surrounding eight neighborhoods, generate a 3×3 adjacent pixel similarity matrix, otherwise execute step S3012, the similarity value is determined by the difference in pixel values between the two pixels; the similar pixel refers to a pixel whose depth value difference with the central pixel is within 10%; Step S3012: traverse each pixel in the 5×5 pixel matrix centered on each pixel to find similar pixels. When there are at least eight similar pixels, generate a 5×5 adjacent pixel similarity matrix, otherwise execute step S3013; Step S3013: traverse each pixel in the 7×7 pixel matrix centered on each pixel to find similar pixels. When there are at least 10 similar pixels, generate a 7×7 adjacent pixel similarity matrix and remove the pixel as a noise point.
7. A spatial neural network depth completion device based on adaptive diffusion kernel, characterized in that: include: processor; a memory storing executable instructions of the processor; Wherein, the processor is configured to execute the steps of the spatial neural network depth completion method based on adaptive diffusion kernel as described in any one of claims 1 to 5 by executing the executable instructions.
8. A computer-readable storage medium for storing a program, characterized in that: When the program is executed, the steps of the spatial neural network depth completion method based on adaptive diffusion kernel described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Monocular depth estimation method, apparatus, terminal, and storage medium
CN109087349A
A method for completing continuous missing data of dam deformation monitor
CN109101638A