Image Processing Method, Electronic Device, and Computer-Readable Storage Medium

By predicting and compensating the initial depth image of the depth camera, the problem of low depth image accuracy caused by different reflectivity of the depth camera is solved, and higher depth image accuracy and quality are achieved.

CN118590770BActive Publication Date: 2025-05-30DIANYUN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410747153.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-11
Publication Date
2025-05-30
Estimated Expiration
2044-06-11

AI Technical Summary

Technical Problem

When the depth camera measures the depth image, the accuracy of the depth image is low due to the different reflectivity of the captured object.

Method used

By obtaining the color image and the initial depth image of the target camera, the depth prediction calculation is performed to obtain the first predicted depth image, and the initial depth image is used to perform depth compensation on the initial depth image to obtain the target depth image.

Benefits of technology

Improves the accuracy of depth images under low reflectivity objects or high-light objects, and improves the quality and accuracy of depth images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118590770B_ABST
    Figure CN118590770B_ABST
Patent Text Reader

Abstract

The present application provides an image processing method, an electronic device, and a computer-readable storage medium. The method may include: obtaining a color image and an initial depth image of a target camera, where the color image and the initial depth image are images captured for the same region; performing depth prediction calculation on the color image to obtain a first predicted depth image; and using the first predicted depth image to perform depth compensation on the initial depth image to obtain a target depth image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technologies, and in particular, to an image processing method, an electronic device, and a computer-readable storage medium. Background Art

[0002] Currently, when a depth camera measures a depth image, due to different reflectivities of the photographed objects, the photographed results obtained at the same distance may be inconsistent, which may lead to the problem of low accuracy of the depth image obtained by the depth camera. Summary of the Invention

[0003] The purpose of the present application is to provide an image processing method, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the accuracy of the depth image obtained by a depth camera.

[0004] In a first aspect, the present invention provides an image processing method, including: obtaining a color image and an initial depth image of a target camera, where the color image and the initial depth image are images obtained by photographing the same area; performing depth prediction calculation on the color image to obtain a first predicted depth image; using the first predicted depth image to perform depth compensation on the initial depth image to obtain a target depth image.

[0005] In the above embodiment, by using the predicted depth image to compensate the initial depth image photographed by the target camera, the accuracy of the depth image under low-reflectivity objects or high-brightness objects can be improved.

[0006] In an optional embodiment, the target camera includes a depth sensor; the using the first predicted depth image to perform depth compensation on the initial depth image to obtain a target depth image includes: converting the first predicted depth image into a second predicted depth image in the coordinate system of the depth sensor; using the second predicted depth image to compensate the initial depth image to obtain a target depth image.

[0007] In the above embodiment, by converting the color image determined based on the color sensor into an image in the coordinate system of the depth sensor, it can more accurately compensate the depth image in the coordinate system of the depth sensor.

[0008] In an alternative embodiment, the target camera includes a color sensor; the step of converting the first predicted depth image into a second predicted depth image in the coordinate system of the depth sensor includes: converting the first predicted depth image into a first point cloud in the coordinate system of the color sensor; converting the first point cloud into a second point cloud in the coordinate system of the depth sensor; and converting the second point cloud into a second predicted depth image in the coordinate system of the depth sensor.

[0009] In the above embodiment, point cloud can be extracted from the first predicted depth image first, and the second predicted depth image can be obtained through calculation and conversion by using the point cloud, which is more convenient for calculation and can also improve the efficiency of image conversion.

[0010] In an alternative embodiment, the step of using the second predicted depth image to compensate the initial depth image to obtain a target depth image includes: calculating an absolute predicted depth image corresponding to the second predicted depth image according to the pixel size relationship between the second predicted depth image and the initial depth image; and using the absolute predicted depth image to compensate the initial depth image to obtain a target depth image.

[0011] In an alternative embodiment, the step of calculating an absolute predicted depth image corresponding to the second predicted depth image according to the pixel size relationship between the initial depth image and the second predicted depth image includes: extracting a first point set from the initial depth image; extracting a second point set from the second predicted depth image, where each point in the second point set corresponds to each point in the first point set; constructing the pixel size relationship between the initial depth image and the second predicted depth image according to the first point set and the second point set; and converting the depth values in the second predicted depth image according to the pixel size relationship to obtain an absolute predicted depth image.

[0012] In an alternative embodiment, the pixel size relationship includes a linear relationship; the step of constructing the pixel size relationship between the initial depth image and the second predicted depth image according to the first point set and the second point set includes: constructing a linear function; and calculating the slope and intercept of the linear function by using the depth values of the points in the first point set and the depth values of the points in the second point set to obtain the linear relationship between the initial depth image and the second predicted depth image.

[0013] In the above embodiments, for the images obtained from the same target area and the same camera, the pixel relationship more often presents a linear relationship. Based on this, the linear relationship between the initial depth image and the second predicted depth image can be obtained by constructing a linear function, which can more simply present the relationship between the initial depth image and the second predicted depth image and can also reduce the difficulty of the calculation process.

[0014] In an alternative embodiment, the using of the absolute predicted depth image to compensate the initial depth image to obtain a target depth image includes: calculating an aliasing parameter at each pixel position according to the depth values of each point in the absolute predicted depth image, the initial depth image, and the aliasing distance of the target camera; for a target pixel point, calculating a target depth value according to the aliasing parameter corresponding to the target pixel point and the depth value of the initial depth image, where the target pixel point is a pixel point in the initial depth image; and the target depth values of all pixel points in the initial depth image form the target depth image.

[0015] In the above embodiments, when performing compensation, the aliasing parameter can be determined in combination with the aliasing distance of the target camera, and the depth value of the initial depth image can be compensated and calculated based on the aliasing parameter, so as to achieve compensation suitable for the target camera and make the obtained target depth image more accurate.

[0016] In an alternative embodiment, the using of the absolute predicted depth image to compensate the initial depth image to obtain a target depth image includes: calculating an aliasing parameter at each pixel position according to the depth values of each point in the absolute predicted depth image, the initial depth image, and the aliasing distance of the target camera; for a first target pixel point, calculating a first target depth value according to the aliasing parameter corresponding to the first target pixel point and the depth value of the initial depth image, where the first target pixel point is a valid pixel point in the initial depth image; for a second target pixel point, determining a second target depth value using the depth value in the absolute predicted depth image corresponding to the second target pixel point; where the second target pixel point is an invalid pixel point in the initial depth image; and the first target depth values or the second depth values of all pixel points in the initial depth image form the target depth image.

[0017] In the above embodiments, different methods can be used to achieve compensation based on the valid pixel points and invalid pixel points in the initial depth image, so that the compensated target depth image can be more accurate.

[0018] In an alternative embodiment, determining the second target depth value by using the depth value in the absolute predicted depth image corresponding to the second target pixel point includes: if the difference between the depth value in the absolute predicted depth image corresponding to the second target pixel point and the depth value of the first pixel point in the initial depth image is less than a preset value, determining the depth value in the absolute predicted depth image corresponding to the second target pixel point as the second target depth value, where the first pixel point is the pixel point closest to the position of the second target pixel point; if the difference between the depth value in the absolute predicted depth image corresponding to the second target pixel point and the depth value of the first pixel point in the initial depth image is not less than the preset value, determining the second target depth value as a specified value.

[0019] In an alternative embodiment, performing depth prediction calculation on the color image to obtain a first predicted depth image includes: inputting the color image into a pre-trained model for calculation to obtain a first predicted depth image.

[0020] In the above embodiment, the determination of the predicted depth image is implemented through a pre-trained model, which can make the determination efficiency of the first predicted depth image higher.

[0021] In an alternative embodiment, the pre-trained model is obtained through the following training method: inputting a training data set into the current neural network model for calculation to obtain an initial predicted depth image; calculating a training loss according to the initial predicted depth image and the label of the training data set; when the training loss is greater than a preset loss value, updating the parameters of the current neural network model to obtain an updated current neural network model; repeating the above steps until the training loss is not greater than the preset loss value, or the number of training times is not less than a preset number value, to obtain a pre-trained model.

[0022] In a second aspect, the present invention provides an electronic device, including: a processor and a memory, where the memory stores machine-readable instructions executable by the processor, and when the electronic device runs, the machine-readable instructions are executed by the processor to perform the steps of the method according to any one of the foregoing embodiments.

[0023] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it performs the steps of the method according to any one of the foregoing embodiments.

[0024] In a fourth aspect, the present invention provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method according to any one of the foregoing embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation of the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0026] Figure 1 It is a block diagram of the electronic device provided by the embodiment of the present application;

[0027] Figure 2 It is a flowchart of the image processing method provided by the embodiment of the present application;

[0028] Figure 3 It is an alternative flowchart of step 230 of the image processing method provided by the embodiment of the present application;

[0029] Figure 4 It is an alternative flowchart of step 232 of the image processing method provided by the embodiment of the present application. Detailed implementation manners

[0030] The following will describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application.

[0031] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, terms such as "first" and "second" are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0032] An RGBD camera refers to a camera that can output RGB images and depth images. Among them, the depth image can be based on iToF (indirect Time of Flight) or dToF (direct Time of Flight) technology, or other technologies. ToF (Time of Flight) is a technology that uses the time of flight of light to measure distance, and measures the distance through the delay between the emitted light and the light emitted by the object.

[0033] Since the camera using the iToF technology is an active imaging device, it actively emits an illumination signal at a given frequency. The light is reflected back from the surface of the object to be photographed and, after being received by the sensor, is further calculated and processed to obtain the distance between the object to be photographed and the camera. In one case, when there is an object with a low reflectivity in the shooting environment, the emitted light will be absorbed by the object with a low reflectivity, making it impossible to perform effective ranging; in another case, when there is an object with a high reflectivity in the shooting environment, specular reflection may occur on the surface of the object with a high reflectivity, and the reflected light cannot be received by the sensor, so effective ranging cannot be performed either.

[0034] Based on the above research, an image processing method, an electronic device, a computer-readable storage medium, and a computer program product that can be provided in an embodiment of the present application can improve the ranging problem caused by different reflectivities of the object to be measured.

[0035] To facilitate the understanding of this embodiment, the electronic device that executes the image processing method disclosed in the embodiment of the present application will be introduced in detail first.

[0036] As Figure 1 shown, it is a block diagram of the electronic device. The electronic device 100 may include a memory 111 and a processor 113. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the electronic device 100. For example, the electronic device 100 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.

[0037] The above-mentioned memory 111 and processor 113 are directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these components may be electrically connected to each other through one or more communication buses or signal lines. The above-mentioned processor 113 is used to execute the executable module stored in the memory.

[0038] Among them, the memory 111 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. Among them, the memory 111 is used to store a program. After receiving an execution instruction, the processor 113 executes the program. The method executed by the electronic device 100 defined by any embodiment of the embodiments of the present application can be applied to or implemented by the processor 113.

[0039] The above-mentioned processor 113 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 113 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit

[0040] (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0041] In this embodiment, the electronic device 100 may be an RGBD camera. A memory and a processor are provided on the RGBD camera to process the depth image and the color image collected by the RGBD camera. The electronic device 100 may also be a computer device with image processing capabilities. The computer device can obtain the color image and the depth image collected by the RGBD camera to implement the compensation of the depth image based on the processing of the color image and the depth image.

[0042] The electronic device 100 in this embodiment can be used to execute each step in the various methods provided in the embodiments of the present application. The implementation process of the image processing method will be described in detail through several embodiments below.

[0043] Please refer to Figure 2 , which is a flowchart of the image processing method provided in the embodiments of the present application. The image processing method provided in the embodiments of the present application can be applied to an electronic device, and the steps in the image processing method are executed through the electronic device. The specific process shown below will be elaborated in detail. Figure 2 shown specific process will be elaborated in detail.

[0044] Step 210, obtain a color image and an initial depth image of the target camera.

[0045] Among them, the color image and the initial depth image are images taken for the same area.

[0046] The color image can be an RGB (Red Green Blue) image.

[0047] Exemplarily, the color image can be a color image of the target area taken by the target camera at a specified position, and the initial depth image can be a depth image of the target area taken by the target camera at the specified position.

[0048] Exemplarily, the target camera may include a color sensor and a depth sensor. Through the color sensor, an RGB image can be taken, and through the depth sensor, a depth camera can be taken.

[0049] Optionally, the execution subject of the image processing method can be the target camera, and the color image and the initial depth image can be obtained by taking pictures with the target camera. Optionally, the execution subject of the image processing method can also be a computer device with image processing functions. The computer device can obtain the color image and the initial depth image by reading the image data in the external device; it can also be communicatively connected to the target camera and receive the color image and the initial depth image transmitted by the target camera.

[0050] Step 220, perform depth prediction calculation on the color image to obtain a first predicted depth image.

[0051] Optionally, some neural network models can be used to infer the color image to obtain a first predicted depth image.

[0052] Exemplarily, a neural network model can be pre-trained with a training data set formed by color images in a supervised manner to obtain a pre-trained model. Then, the color image is input into the pre-trained model for calculation to obtain a first predicted depth image.

[0053] Step 230: Use the first predicted depth image to perform depth compensation on the initial depth image to obtain the target depth image.

[0054] Optionally, the first predicted depth image can be used to compensate for invalid pixel points in the initial depth image.

[0055] Optionally, the depth values of the first predicted depth image and the depth values of the initial depth image can also be used for weighted calculation to obtain the depth values of each pixel point in the target depth image.

[0056] Optionally, different calculation methods can be adopted for the valid pixel points and invalid pixel points of the initial depth image. For invalid pixel points, the first predicted depth image can be used to compensate for the invalid pixel points in the initial depth image; for valid pixel points, the depth values of the first predicted depth image and the depth values of the initial depth image can be used for weighted calculation to obtain the depth values of each pixel point in the target depth image.

[0057] Considering that the color image is an image acquired by a color sensor, and the initial depth image is an image acquired by a depth sensor. Therefore, before compensating the initial depth image, the depth values in the first predicted depth image can also be first converted into depth values in the coordinate system of the depth sensor, and then the initial depth image can be depth-compensated with the converted depth values.

[0058] Through the above method, the compensation of the depth image can be realized using the color image, which can improve the problem of inconsistency caused by different reflectivities of the target object when the depth sensor acquires the depth image, and improve the quality and accuracy of the depth image.

[0059] Considering the first predicted depth image predicted based on the color image, since the color image is determined based on the color sensor, and the initial depth image is determined based on the depth sensor. In some cases, the coordinate systems used by the color sensor and the depth sensor may be different. The target camera may include a color sensor and a depth sensor. Based on this, as Figure 3 shown, the above step 230 may include step 231 and step 232.

[0060] Step 231: Convert the first predicted depth image into a second predicted depth image in the depth sensor coordinate system.

[0061] Optionally, based on the conversion relationship between the coordinate systems of the color sensor and the depth sensor, the depth values of each pixel point in the first predicted depth image can be converted to obtain the second predicted depth image in the depth sensor coordinate system.

[0062] Exemplarily, the conversion relationship can be determined by the extrinsic parameters of the target camera. The conversion relationship can be a conversion matrix from the coordinate system of the color sensor to the coordinate system of the depth sensor, and the conversion matrix can be a rotation matrix, a translation matrix, or a rotation-translation matrix.

[0063] The depth values of the pixel points in the first predicted depth image are converted using the conversion matrix, so that the second predicted depth image in the coordinate system of the depth sensor can be obtained.

[0064] Step 232: Use the second predicted depth image to compensate the initial depth image to obtain the target depth image.

[0065] Optionally, the second predicted depth image can be used to compensate the invalid pixel points in the initial depth image.

[0066] Optionally, the depth values of the second predicted depth image and the depth values of the initial depth image can also be used for weighted calculation to obtain the depth values of the pixel points in the target depth image.

[0067] Optionally, different calculation methods can be adopted for the valid pixel points and the invalid pixel points of the initial depth image. For the invalid pixel points, the second predicted depth image can be used to compensate the invalid pixel points in the initial depth image; for the valid pixel points, the depth values of the second predicted depth image and the depth values of the initial depth image can be used for weighted calculation to obtain the depth values of the pixel points in the target depth image.

[0068] Through the above processing method, the depth image can be compensated based on the color image even when different coordinate systems are used for the color sensor and the depth sensor, increasing the applicability of the image processing method.

[0069] In one implementation, to facilitate the conversion of the image, the image can be first abstracted into a point cloud. Based on this, the above step 231 can include converting the first predicted depth image into a first point cloud in the coordinate system of the color sensor; converting the first point cloud into a second point cloud in the coordinate system of the depth sensor; and converting the second point cloud into a second predicted depth image in the coordinate system of the depth sensor.

[0070] Optionally, the coordinate system of the color sensor can be a three-dimensional coordinate system with the optical center of the color sensor as the origin.

[0071] The above conversion of the first point cloud into a second point cloud in the coordinate system of the depth sensor can include: converting the first point cloud into a second point cloud in the coordinate system of the depth sensor based on the conversion relationship between the coordinate systems of the color sensor and the depth sensor.

[0072] The conversion relationship can be a conversion matrix from the coordinate system of the color sensor to the coordinate system of the depth sensor. Then, multiplying the coordinates of each point in the second point cloud by this conversion matrix can obtain the coordinates of each point in the second point cloud in the coordinate system of the depth sensor.

[0073] Exemplarily, the ToF camera model of the depth sensor of the target camera can be used to convert the second point cloud into a second predicted depth image in the coordinate system of the depth sensor.

[0074] Optionally, the second predicted depth image can be a normalized depth image.

[0075] Through the above implementation method, by first extracting the point cloud in the image for image conversion, the efficiency of image conversion can be improved, and it is also convenient to implement the conversion calculation of each pixel point.

[0076] In one implementation, as Figure 4 shown, step 232 may include step 2321 and step 2322.

[0077] Step 2321, according to the pixel size relationship between the second predicted depth image and the initial depth image, calculate the absolute predicted depth image corresponding to the second predicted depth image.

[0078] In the case where the depth value of the second predicted depth image is a normalized value, the depth value of the second predicted depth image can be first converted into a depth value with the same size as the initial depth image.

[0079] Optionally, the pixel size relationship between the second predicted depth image and the initial depth image can be determined according to the size ratio of the depth value of the second predicted depth image to the depth value of the initial depth image at the corresponding position.

[0080] Optionally, a first point set can be extracted from the initial depth image; a second point set can be extracted from the second predicted depth image, where each point in the second point set corresponds one-to-one to each point in the first point set; according to the first point set and the second point set, the pixel size relationship between the initial depth image and the second predicted depth image is constructed; according to the pixel size relationship, the depth value in the second predicted depth image is converted to obtain the absolute predicted depth image.

[0081] Exemplarily, at least two pixel points can be selected from the second predicted depth image, and the pixel points at the corresponding positions can be selected from the initial depth image to determine the pixel size relationship between the second predicted depth image and the initial depth image.

[0082] Exemplarily, the pixel size relationship includes a linear relationship. The above-mentioned construction of the pixel size relationship between the initial depth image and the second predicted depth image based on the first point set and the second point set includes: constructing a linear function; using the depth values of the points in the first point set and the depth values of the points in the second point set to calculate the slope and intercept of the linear function, so as to obtain the image linear relationship between the initial depth image and the second predicted depth image.

[0083] Exemplarily, the linear function can be expressed as: y = kx + b. By substituting the depth values of the points in the first point set and the depth values of the points in the second point set into the linear function, the slope k and the intercept b are calculated to obtain the linear relationship between the second predicted depth image and the initial depth image.

[0084] After determining the linear relationship, the depth values of the second predicted depth image can be substituted into the linear relationship function to calculate the absolute predicted depth image.

[0085] Step 2322, use the absolute predicted depth image to compensate the initial depth image to obtain the target depth image.

[0086] Optionally, the absolute predicted depth image can be used to compensate the invalid pixel points in the initial depth image.

[0087] Optionally, the depth values of the absolute predicted depth image and the depth values of the initial depth image can also be used for weighted calculation to obtain the depth values of each pixel point in the target depth image.

[0088] Optionally, different calculation methods can be adopted for the valid pixel points and invalid pixel points of the initial depth image. For the invalid pixel points, the absolute predicted depth image can be used to compensate the invalid pixel points in the initial depth image; for the valid pixel points, the depth values of the absolute predicted depth image and the depth values of the initial depth image can be used for weighted calculation to obtain the depth values of each pixel point in the target depth image.

[0089] Through the above implementation method, while improving the accuracy of the target depth image, relevant calculations can also be realized in a more efficient way, improving the overall efficiency of image processing.

[0090] The iToF technology usually achieves the purpose of ranging by continuous wave modulation and demodulation, and indirectly obtaining time by calculating the phase difference Δφ. The distance can be determined by the following formula:

[0091] d = cΔt / 2 = cΔφ / 2ω = cΔφ / 4πf mod ;

[0092] where c represents the speed of light, and f mod represents the modulation frequency of the transmitted signal.

[0093] Among them, Δφ is determined by the following calculation method:

[0094] Δφ = arctan((Q 90 - Q 270 ) / (Q o - Q 180 ));

[0095] For each pixel, the time interval Δt between the signal emitted by the light source and the received signal can be calculated. It can be calculated from the signals obtained by sampling four phases in the reflected signal. The signals obtained by sampling these four phases are respectively denoted as Q 0 , Q 90 , Q 180 , Q 270 .

[0096] From the above calculation method, it can be known that the time interval Δt is actually indirectly calculated through the phase difference Δφ. The value range of the phase difference Δφ is (0, 2π). That is to say, when the distance exceeds 2π, there will be a phenomenon of spectral aliasing, that is, the measured phase 0 and phase 2π are the same. This phenomenon is called distance ambiguity (phase wrapping). Among them, the maximum distance that can be detected by the current modulation frequency can be expressed as d omb = c / 2f mod . If it is necessary to expand the measurement distance, the modulation frequency f mod can be reduced. However, this will increase the measurement error. Therefore, in order not to reduce the accuracy while measuring the distance, the current ToF all adopts multi-frequency technology. The multi-frequency technology is to add one or more frequency modulation waves for mixing. The real distance is the value measured jointly by multiple frequency modulation waves. The corresponding modulation frequency at this position is the greatest common divisor of multiple modulation frequencies, which is called the beat frequency. The beat frequency is generally lower than the modulation frequency of a single modulation wave, thus realizing the extension of a longer measurement distance. Then, even if multiple modulation waves are used to improve the problem of distance ambiguity, if the measurement data obtained by a single modulation wave is inaccurate, the final measurement result will also be affected. Based on this, the distance ambiguity can be corrected by compensating the initial depth image through the absolute predicted depth image. The above step 2322 may include calculating the aliasing parameter at each pixel position according to the depth value of each point of the absolute predicted depth image, the initial depth image, and the aliasing distance of the target camera; for the target pixel, calculating the target depth value according to the aliasing parameter corresponding to the target pixel and the depth value of the initial depth image.

[0097] Among them, the target pixel is a pixel in the initial depth image.

[0098] Among them, the target depth values of all pixel points in the initial depth image constitute the target depth image.

[0099] Exemplarily, the aliasing parameter can be the number of aliasing times.

[0100] The number of aliasing times is calculated by the following formula:

[0101]

[0102] Among them, the k i,j ∈ {0, 1, 2...} is a non - negative integer, representing the number of aliasing times; d i,j represents the depth value corresponding to the pixel at the i - th row and j - th column in the initial depth image; represents the depth value corresponding to the pixel at the i - th row and j - th column in the predicted depth image; d amb represents the maximum detectable distance of the target camera at the current modulation frequency.

[0103] The target depth value can be expressed as:

[0104] Among them, can represent the target depth value corresponding to the pixel at the i - th row and j - th column in the target depth image.

[0105] Considering that in the initial depth image, due to the reflectivity of the photographed object, some pixel points may be valid pixel points and some may be invalid pixel points. And since the depth values of invalid pixel points may have serious errors, the depth value of such a pixel point may not exist directly. Based on this, different methods can also be used to compensate for the valid pixel points and invalid pixel points in the initial depth image.

[0106] Optionally, for the valid pixel points in the initial depth image, the above - mentioned calculation method is used to compensate and obtain the corresponding target depth value, and for the invalid pixel points in the initial depth image, the depth values corresponding to the pixel points in this initial depth image may not be referred to.

[0107] The above - mentioned step 2322 may include calculating the aliasing parameter at each pixel point position according to the depth values of each point in the absolute predicted depth image, the initial depth image, and the aliasing distance of the target camera; for the first target pixel point, calculating the first target depth value according to the aliasing parameter corresponding to the first target pixel point and the depth value of the initial depth image.

[0108] Among them, the first target pixel point is a valid pixel point in the initial depth image.

[0109] Exemplarily, the calculation method of the first target depth value of the first target pixel point may be the same as the calculation method of the target pixel point described above. For specific details, reference may be made to the foregoing description method, which will not be elaborated here.

[0110] For the second target pixel point, the depth value in the absolute predicted depth image corresponding to the second target pixel point is used to determine the second target depth value.

[0111] The second target pixel point is an invalid pixel point in the initial depth image.

[0112] The first target depth value or the second depth value of all pixel points in the initial depth image constitutes the target depth image.

[0113] Exemplarily, based on the first target depth value of the valid pixel points in the initial depth image and the second target depth value of the invalid pixel points in the initial depth image, the target depth image can be constructed.

[0114] In one implementation, if the difference between the depth value in the absolute predicted depth image corresponding to the second target pixel point and the depth value of the first pixel point in the initial depth image is less than a preset value, the depth value in the absolute predicted depth image corresponding to the second target pixel point is determined as the second target depth value.

[0115] The first pixel point is the pixel point closest to the position of the second target pixel point.

[0116] The preset value can be a value set as needed. Exemplarily, it can be a value set by the user before executing this image processing method.

[0117] Optionally, the preset value can be determined based on the depth values of each pixel point in the initial depth image. For example, the preset value can be the difference between the maximum depth value and the minimum depth value in the initial depth image. For another example, the preset value can be half of the difference between the maximum depth value and the minimum depth value in the initial depth image.

[0118] In one implementation, if the difference between the depth value in the absolute predicted depth image corresponding to the second target pixel point and the depth value of the first pixel point in the initial depth image is not less than the preset value, the second target depth value is determined as a specified value.

[0119] The specified value can be zero, or it can be the depth value of the first pixel point. The specified value can also be a value calculated based on the depth values of the valid pixel points around the second target pixel point in the initial depth image. For example, the specified value can also be the average value of the valid pixel points around the second target pixel point in the initial depth image.

[0120] In the embodiments of the present application, the above-mentioned pre-trained model is trained in the following manner: Input the training data set into the current neural network model for calculation to obtain an initial predicted depth image; Calculate the training loss based on the initial predicted depth image and the label of the training data set; When the training loss is greater than the preset loss value, update the parameters of the current neural network model to obtain an updated current neural network model; Repeat the above steps until the training loss is not greater than the preset loss value, or the number of training times is not less than the preset number of times value, to obtain the pre-trained model.

[0121] Optionally, the training data set may be an RGB image. The current neural network model may be a CNN model.

[0122] The RGB image can be input into the current neural network model for forward propagation inference to obtain an initial predicted depth image; Calculate the training loss based on the initial predicted depth image and the true label of the RGB image.

[0123] Then update the parameters of the current neural network model through the backpropagation algorithm. Then repeat the above process until the training loss obtained by the current neural network model is less than the preset loss value, and take the latest obtained current neural network model as the pre-trained model.

[0124] In one case, when the number of training times has exceeded the preset number of times value, but the training loss is still not less than the preset loss value, the loop can also be exited, and the latest obtained current neural network model is taken as the pre-trained model.

[0125] When the pre-trained model needs to be used, the RGB image can be input into the pre-trained model, and the first predicted depth image can be obtained through forward propagation calculation.

[0126] In addition, the embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the image processing method described in the above method embodiments.

[0127] The computer program product of the image processing method provided by the embodiments of the present application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the steps of the image processing method described in the above method embodiments. For details, please refer to the above method embodiments and will not be elaborated here.

[0128] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0129] In addition, the functional modules in each embodiment of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0130] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes. It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. Without further limitations, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the said elements.

[0131] The above are only the preferred embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the protection scope of this application. It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0132] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or replacements, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

1. An image processing method, characterized in that: include: Obtaining a color image and an initial depth image of a target camera, wherein the color image and the initial depth image are images captured for the same area; the target camera includes a depth sensor; Performing depth prediction calculation on the color image to obtain a first predicted depth image; Converting the first predicted depth image into a second predicted depth image in the depth sensor coordinate system; According to a pixel size relationship between the initial depth image and the second predicted depth image, converting a depth value in the second predicted depth image according to the pixel size relationship to obtain an absolute predicted depth image; Compensating the initial depth image using the absolute predicted depth image to obtain a target depth image; Among them, the pixel size relationship between the initial depth image and the second predicted depth image includes: extracting a first point set from the initial depth image; extracting a second point set from the second predicted depth image, wherein each point in the second point set corresponds one-to-one to each point in the first point set; and constructing a pixel size relationship between the initial depth image and the second predicted depth image based on the first point set and the second point set.

2. The method according to claim 1, characterized in that The target camera further includes a color sensor; and converting the first predicted depth image into a second predicted depth image in the depth sensor coordinate system includes: Converting the first predicted depth image into a first point cloud in a coordinate system of the color sensor; Converting the first point cloud into a second point cloud in the coordinate system of the depth sensor; The second point cloud is converted into a second predicted depth image in the coordinate system of the depth sensor.

3. The method according to claim 1, characterized in that The pixel size relationship includes a linear relationship; and constructing the pixel size relationship between the initial depth image and the second predicted depth image according to the first point set and the second point set includes: Construct linear functions; The slope and intercept of the linear function are calculated using the depth values ​​of the points in the first point set and the depth values ​​of the points in the second point set to obtain a linear relationship between the initial depth image and the second predicted depth image.

4. The method according to claim 1, characterized in that: The using the absolute predicted depth image to compensate the initial depth image to obtain a target depth image includes: Calculate the aliasing parameter at each pixel position according to the depth value of each point of the absolute predicted depth image, the initial depth image and the aliasing distance of the target camera; For a target pixel, a target depth value is calculated according to the aliasing parameter corresponding to the target pixel and the depth value of the initial depth image, wherein the target pixel is a pixel in the initial depth image; The target depth values ​​of all pixels in the initial depth image constitute a target depth image.

5. The method according to claim 1, characterized in that The using the absolute predicted depth image to compensate the initial depth image to obtain a target depth image includes: Calculate the aliasing parameter at each pixel position according to the depth value of each point of the absolute predicted depth image, the initial depth image and the aliasing distance of the target camera; For a first target pixel, a first target depth value is calculated according to the aliasing parameter corresponding to the first target pixel and the depth value of the initial depth image, wherein the first target pixel is a valid pixel in the initial depth image; For a second target pixel, a second target depth value is determined using a depth value in the absolute predicted depth image corresponding to the second target pixel; wherein the second target pixel is an invalid pixel in the initial depth image; The first target depth values ​​or the second depth values ​​of all pixels in the initial depth image constitute a target depth image.

6. The method according to claim 5, characterized in that The using the depth value in the absolute predicted depth image corresponding to the second target pixel point to determine the second target depth value includes: If the difference between the depth value in the absolute predicted depth image corresponding to the second target pixel and the depth value of the first pixel in the initial depth image is less than a preset value, determine that the depth value in the absolute predicted depth image corresponding to the second target pixel is the second target depth value, wherein the first pixel is the pixel closest to the second target pixel; If the difference between the depth value in the absolute predicted depth image corresponding to the second target pixel and the depth value of the first pixel in the initial depth image is not less than a preset value, the second target depth value is determined to be a specified value.

7. The method according to claim 1, characterized in that The performing depth prediction calculation on the color image to obtain a first predicted depth image includes: The color image is input into a pre-trained model for calculation to obtain a first predicted depth image.

8. An electronic device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the machine-readable instructions are executed by the processor to perform the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Depth compensation method and device

    CN117294829A

  • Image processing method

    US20220414908A1