Image processing method and device, electronic device, and computer-readable storage medium
By using the estimated offset in monocular depth estimation to correct the depth value distribution on the depth map, the problems of insufficient clarity of the depth map and transition zone are solved, and the depth map clarity improvement without increasing the calculation amount is achieved.
Patent Information
- Application Number
- CN202210592041.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-05-27
AI Technical Summary
In the prior art, it is difficult to improve the clarity of the depth map without increasing the calculation amount in monocular depth estimation, especially in the boundary, there is a problem of transition zones and insufficient clarity at the boundary.
By performing depth estimation processing on the target image, the estimated depth value and estimated offset of each pixel point are obtained. The offset pixel point is determined using the estimated offset, and the corresponding estimated depth value is used as the corrected depth value to correct the depth value distribution on the depth map.
Eliminate transition zones on the depth map without increasing the calculation amount, improving the clarity of the depth map, especially significantly improving the clarity at the boundaries without causing false edges.
Smart Images

Figure CN114937072B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of deep learning technology, and in particular to an image processing method and device, an electronic device, and a computer-readable storage medium. Background Art
[0002] Monocular depth estimation refers to the prediction of the distance of each point on an RGB (Red Green Blue, a color system that obtains various colors by changing the three color channels of red, green, and blue and superimposing them on each other) image from the camera plane based on deep learning, that is, the depth of each point, thereby obtaining a depth map of the RGB image. However, the depth map obtained by the depth estimation method based on deep learning often has problems such as transition zones at the boundaries and insufficient clarity.
[0003] Related technologies There are two types of solutions to solve this problem. One is to use two-dimensional guided filtering to enhance depth details, but the depth map is a three-dimensional information, and this solution will destroy the three-dimensional geometric structure. For example, for an inclined wall, the depth is distributed from near to far, and this solution will change the inclination angle of the wall, and will process the two-dimensional image boundary as the boundary on the depth map, causing the problem of false edges. The other is to increase the input size of the depth estimation model to enhance depth details, but this will undoubtedly greatly increase the amount of calculation, which is not convenient for mobile applications. Summary of the invention
[0004] The present disclosure provides an image processing method and device, an electronic device, and a computer-readable storage medium to at least solve the problem in the related art of how to improve the clarity of the depth map without significantly increasing the amount of calculation, and may not solve any of the above problems.
[0005] According to a first aspect of the present disclosure, an image processing method is provided, which includes: performing depth estimation processing on a target image to obtain an estimated depth value and an estimated offset for each pixel of the target image; determining an offset pixel of a corresponding pixel according to the estimated offset for each pixel of the target image; using the estimated depth value corresponding to the offset pixel of each pixel in the target image as a corrected depth value of the corresponding pixel; and obtaining a depth map of the target image based on the corrected depth value of each pixel of the target image.
[0006] Optionally, determining the offset pixel point of the corresponding pixel point based on the estimated offset of each pixel point of the target image includes: for each pixel point of the target image, obtaining the coordinate value of the offset pixel point of the corresponding pixel point by adding the coordinate value of the corresponding pixel point to the estimated offset of the corresponding pixel point.
[0007] Optionally, each pixel point has a first direction coordinate value and a second direction coordinate value, and the estimated offset of each pixel point includes the first direction estimated offset and the second direction estimated offset, wherein, for each pixel point of the target image, the coordinate value of the offset pixel point of the corresponding pixel point is obtained by adding the coordinate value of the corresponding pixel point to the estimated offset of the corresponding pixel point, including: for each pixel point of the target image, the first direction coordinate value of the offset pixel point of the corresponding pixel point is obtained by adding the first direction coordinate value of the corresponding pixel point and the first direction estimated offset of the corresponding pixel point, and the second direction coordinate value of the offset pixel point of the corresponding pixel point is obtained by adding the second direction coordinate value of the corresponding pixel point and the second direction estimated offset of the corresponding pixel point.
[0008] Optionally, when the target image is in a rectangular coordinate system, the first direction is the x direction of the rectangular coordinate system, and the second direction is the y direction of the rectangular coordinate system; or when the target image is in a polar coordinate system, the first direction is the polar radial direction of the polar coordinate system, and the second direction is the polar angular direction of the polar coordinate system.
[0009] Optionally, performing depth estimation processing on the target image to obtain an estimated depth value and an estimated offset for each pixel of the target image includes: inputting the target image into a depth estimation model to obtain an output result of the depth estimation model; wherein the output result includes the estimated depth value and the estimated offset for each pixel of the target image.
[0010] Optionally, the depth estimation model is trained by the following steps: obtaining a sample image and a reference depth value for each pixel of the sample image; inputting the sample image into the depth estimation model to obtain a sample estimated depth value and a sample estimated offset for each pixel of the sample image; determining a reference offset for each pixel of the sample image based on the reference depth value and the sample estimated depth value for each pixel of the sample image; determining a loss value based on the reference depth value, the sample estimated depth value, the reference offset and the sample estimated offset for each pixel of the sample image; and adjusting the parameters of the depth estimation model based on the loss value to train the depth estimation model.
[0011] Optionally, determining the reference offset of each pixel of the sample image according to the reference depth value and the sample estimated depth value of each pixel of the sample image includes: traversing all pixels of the sample image to determine a number of reference points around a current pixel; determining the weight of a corresponding reference point of the current pixel according to the reference depth value of the current pixel and the sample estimated depth value of each reference point of the current pixel; determining the offset of each reference point of the current pixel relative to the current pixel; and obtaining a weighted average of the offsets of all reference points of the current pixel according to the weight of each reference point of the current pixel as the reference offset of the current pixel.
[0012] Optionally, determining the weight of the corresponding reference point of the current pixel based on the reference depth value of the current pixel and the sample estimated depth value of each reference point of the current pixel includes: determining the absolute value of the error between the sample estimated depth value of each reference point of the current pixel and the reference depth value of the current pixel as the error of the corresponding reference point; determining the weight of each reference point of the current pixel based on the error of each reference point of the current pixel, the weight of each reference point being negatively correlated with the error of the corresponding reference point.
[0013] Optionally, determining the weight of each reference point of the current pixel based on the error of each reference point of the current pixel includes: determining the sum of the error of each reference point of the current pixel and a set value, and determining a ratio of the set value to the sum as the weight of the corresponding reference point; wherein the set value is a positive number.
[0014] Optionally, determining a number of reference points around the current pixel point includes: determining, among all the pixel points of the sample image, pixel points whose distances from the current pixel point in the first direction and the second direction are less than corresponding reference values, as reference points of the current pixel point; or determining, among all the pixel points of the sample image, pixel points whose straight-line distances from the current pixel point are less than the reference value.
[0015] Optionally, determining the loss value based on the reference depth value of each pixel point of the sample image, the sample estimated depth value, the reference offset and the sample estimated offset includes: calculating a first loss value based on the reference depth value and the sample estimated depth value of each pixel point of the sample image; calculating a second loss value based on the reference offset and the sample estimated offset of each pixel point of the sample image; and determining the loss value based on the first loss value and the second loss value.
[0016] According to a second aspect of the present disclosure, an image processing device is provided, comprising: an estimation unit, configured to perform depth estimation processing on a target image to obtain an estimated depth value and an estimated offset for each pixel of the target image; an offset unit, configured to determine an offset pixel of a corresponding pixel according to the estimated offset for each pixel of the target image; a correction unit, configured to use the estimated depth value corresponding to the offset pixel of each pixel in the target image as a corrected depth value of the corresponding pixel; and a summary unit, configured to obtain a depth map of the target image based on the corrected depth value of each pixel of the target image.
[0017] Optionally, the offset unit is further configured to: for each pixel point of the target image, obtain the coordinate value of the offset pixel point of the corresponding pixel point by adding the coordinate value of the corresponding pixel point to the estimated offset amount of the corresponding pixel point.
[0018] Optionally, each pixel point has a first direction coordinate value and a second direction coordinate value, and the estimated offset of each pixel point includes the first direction estimated offset and the second direction estimated offset, and the offset unit is further configured to: for each pixel point of the target image, obtain the first direction coordinate value of the offset pixel point of the corresponding pixel point by adding the first direction coordinate value of the corresponding pixel point and the first direction estimated offset of the corresponding pixel point, and obtain the second direction coordinate value of the offset pixel point of the corresponding pixel point by adding the second direction coordinate value of the corresponding pixel point and the second direction estimated offset of the corresponding pixel point.
[0019] Optionally, when the target image is in a rectangular coordinate system, the first direction is the x direction of the rectangular coordinate system, and the second direction is the y direction of the rectangular coordinate system; or when the target image is in a polar coordinate system, the first direction is the polar radial direction of the polar coordinate system, and the second direction is the polar angular direction of the polar coordinate system.
[0020] Optionally, the estimation unit is further configured to: input the target image into a depth estimation model to obtain an output result of the depth estimation model; wherein the output result includes the estimated depth value and the estimated offset of each pixel point of the target image.
[0021] Optionally, the depth estimation model is trained by the following steps: obtaining a sample image and a reference depth value for each pixel of the sample image; inputting the sample image into the depth estimation model to obtain a sample estimated depth value and a sample estimated offset for each pixel of the sample image; determining a reference offset for each pixel of the sample image based on the reference depth value and the sample estimated depth value for each pixel of the sample image; determining a loss value based on the reference depth value, the sample estimated depth value, the reference offset and the sample estimated offset for each pixel of the sample image; and adjusting the parameters of the depth estimation model based on the loss value to train the depth estimation model.
[0022] Optionally, determining the reference offset of each pixel of the sample image according to the reference depth value and the sample estimated depth value of each pixel of the sample image includes: traversing all pixels of the sample image to determine a number of reference points around a current pixel; determining the weight of a corresponding reference point of the current pixel according to the reference depth value of the current pixel and the sample estimated depth value of each reference point of the current pixel; determining the offset of each reference point of the current pixel relative to the current pixel; and obtaining a weighted average of the offsets of all reference points of the current pixel according to the weight of each reference point of the current pixel as the reference offset of the current pixel.
[0023] Optionally, determining the weight of the corresponding reference point of the current pixel based on the reference depth value of the current pixel and the sample estimated depth value of each reference point of the current pixel includes: determining the absolute value of the error between the sample estimated depth value of each reference point of the current pixel and the reference depth value of the current pixel as the error of the corresponding reference point; determining the weight of each reference point of the current pixel based on the error of each reference point of the current pixel, the weight of each reference point being negatively correlated with the error of the corresponding reference point.
[0024] Optionally, determining the weight of each reference point of the current pixel based on the error of each reference point of the current pixel includes: determining the sum of the error of each reference point of the current pixel and a set value, and determining a ratio of the set value to the sum as the weight of the corresponding reference point; wherein the set value is a positive number.
[0025] Optionally, determining a number of reference points around the current pixel point includes: determining, among all the pixel points of the sample image, pixel points whose distances from the current pixel point in the first direction and the second direction are less than corresponding reference values, as reference points of the current pixel point; or determining, among all the pixel points of the sample image, pixel points whose straight-line distances from the current pixel point are less than the reference value.
[0026] Optionally, determining the loss value based on the reference depth value of each pixel point of the sample image, the sample estimated depth value, the reference offset and the sample estimated offset includes: calculating a first loss value based on the reference depth value and the sample estimated depth value of each pixel point of the sample image; calculating a second loss value based on the reference offset and the sample estimated offset of each pixel point of the sample image; and determining the loss value based on the first loss value and the second loss value.
[0027] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer executable instructions, wherein the computer executable instructions, when executed by the at least one processor, prompt the at least one processor to execute the image processing method according to the present disclosure.
[0028] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor is prompted to execute the image processing method according to the present disclosure.
[0029] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising computer instructions, which implement the image processing method according to the present disclosure when executed by at least one processor.
[0030] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:
[0031] According to the image processing method and image processing device of the embodiments of the present disclosure, the depth map and the offset of the depth map can be estimated simultaneously, and the distribution of depth values on the estimated depth map can be corrected by the estimated offset, thereby eliminating the transition band and improving the clarity of the depth map with very little computational effort.
[0032] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.
[0034] Figure 1a is a schematic diagram illustrating a target image according to an exemplary embodiment of the present disclosure.
[0035] Figure 1b It is a schematic diagram of a depth map of a target image obtained by applying an image processing method of related technology.
[0036] Figure 2 is a flowchart illustrating an image processing method according to an exemplary embodiment of the present disclosure.
[0037] Figure 3 is a flowchart illustrating an image processing method according to an exemplary embodiment of the present disclosure.
[0038] Figure 4a is a schematic diagram illustrating an estimated depth map according to an exemplary embodiment of the present disclosure.
[0039] Figure 4b is a schematic diagram showing an estimated offset in the x-direction according to an exemplary embodiment of the present disclosure.
[0040] Figure 4c is a schematic diagram showing an estimated offset in the y direction according to an exemplary embodiment of the present disclosure.
[0041] Figure 4d is a schematic diagram illustrating a depth value replacement operation according to an exemplary embodiment of the present disclosure.
[0042] Figure 4e is a schematic diagram illustrating a modified depth map according to an exemplary embodiment of the present disclosure.
[0043] Figure 5 is a flowchart illustrating a training method of a depth estimation model according to an exemplary embodiment of the present disclosure.
[0044] Figure 6 is a flowchart illustrating a method for training a depth estimation model according to an exemplary embodiment of the present disclosure.
[0045] Figure 7 is a block diagram illustrating an image processing apparatus according to an exemplary embodiment of the present disclosure.
[0046] Figure 8 is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0047] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.
[0048] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation methods described in the following examples do not represent all implementation methods consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the attached claims.
[0049] It should be noted that the phrase "at least one of the items" in the present disclosure includes three types of parallel situations: "any one of the items", "a combination of any number of the items", and "all of the items". For example, "including at least one of A and B" includes the following three types of parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "executing at least one of step 1 and step 2" which means the following three types of parallel situations: (1) executing step 1; (2) executing step 2; (3) executing step 1 and step 2.
[0050] Monocular depth estimation is to predict the distance of each point on the RGB image from the camera plane based on deep learning, that is, the depth of each point, so as to obtain the depth map of the RGB image. Figure 1a and Figure 1b It can be seen that the depth map obtained by the depth estimation method based on deep learning often has problems such as transition zones at the boundaries and insufficient clarity.
[0051] Related Art There are two types of solutions to solve this problem.
[0052] One is to use guided filtering post-processing to enhance depth details. Guided filtering uses a clear RGB image as a guide map to filter the depth map, so that the depth map corresponding to the area with relatively consistent texture on the RGB image is smoother, and the depth boundary is more closely aligned with the boundary of the RGB image. However, guided filtering is a two-dimensional operation, while the depth map is a three-dimensional information. For planar areas, the smoothing effect of guided filtering will destroy the three-dimensional geometric structure. For example, for an inclined wall, the depth is distributed from near to far, and smoothing filtering will destroy the geometric relationship, causing the inclination angle of the wall to change. In addition, at the depth boundary, guided filtering can enhance depth details, but it will also bring about the problem of false edges, such as the poster boundary on the wall and the zebra crossing on the road. The boundaries on the RGB of these areas will cause false boundaries to appear on the filtered depth map.
[0053] The other type is deep learning-based methods, which mainly improve depth details by increasing the input size of the depth estimation model. Large input will retain more detailed information in the image, making the predicted depth map clearer, but it will undoubtedly lead to a significant increase in the amount of calculation. For example, if the length and width of the input size become twice the original, the amount of calculation will increase by 4 times.
[0054] Next, we will refer to Figures 2 to 8 An image processing method and an image processing apparatus according to exemplary embodiments of the present disclosure are described in detail.
[0055] Figure 2 is a flowchart illustrating an image processing method according to an exemplary embodiment of the present disclosure. Figure 3 is a flowchart illustrating an image processing method according to an exemplary embodiment of the present disclosure. Figures 4a to 4e It is a series of images generated during the execution of the image processing method according to the exemplary embodiment of the present disclosure. It should be understood that the image processing method according to the exemplary embodiment of the present disclosure can be implemented in a terminal device such as a smart phone, a tablet computer, a personal computer (PC), or in a device such as a server.
[0056] Reference Figure 2 and Figure 3 In step 201, depth estimation is performed on the target image to obtain an estimated depth value and an estimated offset for each pixel of the target image.
[0057] The estimated depth values of each pixel of the target image are collected together to form an estimated depth map. A significant difference between the image processing method according to the exemplary embodiment of the present disclosure and the depth estimation method in the related art is that in addition to the estimated depth value, an estimated offset is also obtained, which can be used for the subsequent steps to correct the distribution of depth values on the estimated depth map.
[0058] As an example, see Figure 1a The target image shown, Figure 4a The estimated depth map with blurred boundaries obtained in step 201 is shown.
[0059] Optionally, step 201 specifically inputs the target image into a depth estimation model to obtain an output result of the depth estimation model; wherein the output result includes an estimated depth value and an estimated offset for each pixel of the target image. A depth estimation model can be used to output the estimated depth value and the estimated offset at the same time, without having to calculate the two separately, which helps to simplify the calculation process. The training process of the depth estimation model will be introduced later.
[0060] Return to reference Figure 2 In step 202, the offset pixel of the corresponding pixel is determined according to the estimated offset of each pixel of the target image. In combination with the estimated offset of each pixel, the corresponding pixel can be offset to the offset pixel, so that the depth value of the corresponding pixel can be corrected based on the offset pixel in the subsequent step.
[0061] Optionally, step 202 specifically includes: for each pixel point of the target image, by adding the coordinate value of the corresponding pixel point to the estimated offset of the corresponding pixel point, obtaining the coordinate value of the offset pixel point of the corresponding pixel point. By performing the coordinate value calculation, the accurate offset pixel point can be obtained, ensuring the simplicity and reliability of the calculation process.
[0062] Optionally, the target image is a two-dimensional plane image, each pixel has a first direction coordinate value and a second direction coordinate value, and the estimated offset of each pixel includes the first direction estimated offset and the second direction estimated offset, so that independent offsets in two directions can be achieved. Based on this, step 202 specifically includes: for each pixel of the target image, by adding the first direction coordinate value of the corresponding pixel and the first direction estimated offset of the corresponding pixel, the first direction coordinate value of the offset pixel of the corresponding pixel is obtained, and by adding the second direction coordinate value of the corresponding pixel and the second direction estimated offset of the corresponding pixel, the second direction coordinate value of the offset pixel of the corresponding pixel is obtained. By performing offset calculations on the two coordinate directions respectively and then summarizing the calculation results, the coordinate values of the offset pixel can be obtained, thereby ensuring the accuracy of the offset calculation.
[0063] Optionally, when the target image is in a rectangular coordinate system, the first direction is the x direction of the rectangular coordinate system, and the second direction is the y direction of the rectangular coordinate system, that is, the rectangular coordinate system is used to describe the coordinate values of the pixel points, which can be adapted to the calculation of conventional rectangular images. Figure 3 , the estimated offset in the x direction can be expressed as dx, and the estimated offset in the y direction can be expressed as dy. Figure 1a The target image shown, Figure 4bA schematic diagram showing the estimated offset in the x direction obtained in step 202 is shown. Figure 4c A schematic diagram showing the estimated offset in the y direction obtained in step 202 is shown.
[0064] Optionally, when the target image is in a polar coordinate system, the first direction is the polar radius direction of the polar coordinate system, and the second direction is the polar angle direction of the polar coordinate system, that is, the polar coordinate system is used to describe the coordinate values of the pixel points, which can be adapted to the calculation and description of circular images.
[0065] It should be understood that in actual calculations, a suitable coordinate system may be used as needed, and the present disclosure does not limit this.
[0066] Return to reference Figure 2 and Figure 3 In step 203, the estimated depth value corresponding to the offset pixel point of each pixel point in the target image is used as the corrected depth value of the corresponding pixel point. This step performs a depth value replacement operation. By replacing the estimated depth value of each pixel point of the estimated depth map obtained in step 201 with the estimated depth value corresponding to the offset pixel point of the pixel point, the distribution of depth values on the estimated depth map can be adjusted to obtain a corrected depth value without recalculating each depth value. There is no need to increase the input size of the depth estimation model, and the amount of calculation is very small, which can be applied to mobile terminals. At the same time, since the guided filtering method is not used, the feature extraction through the deep learning network can distinguish between real depth edges and false edges, which will improve the clarity without introducing false edges.
[0067] Reference Figure 3 For pixel point p(i,j), D1(i,j) represents its estimated depth value, dx(i,j) represents its estimated offset in the x direction, and dy(i,j) represents its estimated offset in the y direction. Then pixel point p'(i+dx(i,j), j+dy(i,j)) represents its offset pixel point, and the corrected depth value of pixel point p(i,j) is D2(i,j) = D1(i+dx(i,j), j+dy(i,j)). Figure 4d The following figure shows a schematic diagram of the depth value replacement operation. Figure 4b to Figure 4d It can be seen that the estimated offset at the boundary is more obvious, which can be understood as replacing the depth values of the pixels in the boundary transition zone with the depth values of the surrounding non-transition zone pixels, thereby improving the clarity at the boundary. For non-transition zone pixels, the estimated offset is basically 0, that is, no depth value replacement is required, which not only improves the boundary clarity in a targeted manner, but also reduces the computational complexity of the depth value replacement operation.
[0068] Based on this, optionally, step 203 may be performed on the pixel point only when the offset pixel point obtained in step 202 does not overlap with the corresponding pixel point. Furthermore, steps 202 and 203 may be performed on the pixel point only when the estimated offset obtained in step 201 is not 0. It should be understood that for the latter, since the estimated offset often includes the estimated offset in the first direction and the estimated offset in the second direction, steps 202 and 203 are specifically performed on the pixel point only when both the estimated offset in the first direction and the estimated offset in the second direction are not 0.
[0069] Return to reference Figure 2 In step 204, based on the corrected depth value of each pixel of the target image, a depth map of the target image is obtained. The corrected depth values of each pixel of the target image are collected together to obtain a depth map estimated by the image processing method according to the exemplary embodiment of the present disclosure. Figure 1a The target image shown, Figure 4e The depth map obtained in step 204 is shown. Figure 4a It can be clearly seen from the comparison that the transition zone has been eliminated and the boundary clarity has been significantly improved.
[0070] Next, we introduce the training method of the depth estimation model.
[0071] Figure 5 is a flowchart illustrating a training method of a depth estimation model according to an exemplary embodiment of the present disclosure. Figure 6 is a flowchart illustrating a method for training a depth estimation model according to an exemplary embodiment of the present disclosure.
[0072] Reference Figure 5 , the depth estimation model is trained through the following steps:
[0073] In step 501, a sample image and a reference depth value of each pixel of the sample image are obtained. This step is a training sample acquisition step, which can implement supervised training of the depth estimation model. The reference depth value can be collected by a depth sensor.
[0074] In step 502, the sample image is input into the depth estimation model to obtain the sample estimated depth value and sample estimated offset of each pixel of the sample image. The sample estimated offset is used to determine the offset pixel of the corresponding pixel, and the sample estimated depth value of the offset pixel of each pixel is used as the corrected depth value of the corresponding pixel. This step uses the sample image to be trained to run the depth estimation model, and the corresponding estimated value can be obtained for comparison with the reference value of the sample image to achieve supervised training.
[0075] In step 503, the reference offset of each pixel of the sample image is determined according to the reference depth value and the sample estimated depth value of each pixel of the sample image. Since the training sample only includes the sample image and the reference depth value, the output of the depth estimation model includes two, namely the sample estimated depth value and the sample estimated offset, so the reference offset is confirmed in this step. By combining the reference depth value and the sample estimated depth value of each pixel, it can be determined which specific pixel has the sample estimated depth value closest to the reference depth value of the current pixel, so that the closest pixel is used as the offset pixel of the current pixel, and the offset of the pixel relative to the current pixel is used as the reference offset, so that a more reliable reference offset can be obtained to guide model training.
[0076] Optionally, step 503 specifically includes the following four steps:
[0077] The first step is to traverse all the pixels of the sample image and determine several reference points around the current pixel. Since the offset pixel must be a pixel that is relatively close to the current pixel, the offset pixel can be determined by determining several reference points around the current pixel. It should be understood that the more reference points are selected, the greater the possibility of obtaining accurate offset pixels, but too many reference points will also cause a large computational burden, so the balance between accuracy and computational burden can be achieved by reasonably controlling the number of reference points to improve computational efficiency. In practice, the number of selected reference points can be determined according to the clarity of the sample image itself to improve the computational efficiency of step 503, or a number of reference points with higher universality can be selected, or a reference point determination standard with higher universality can be selected to reduce the computational amount of preliminary preparation work (here refers to the work of determining the number of reference points), and the present disclosure does not limit this.
[0078] Optionally, the operation of determining several reference points around the current pixel may specifically include: determining, among all the pixels of the sample image, pixels whose distances from the current pixel in the first direction and the second direction are both less than the corresponding reference values, as reference points for the current pixel; or determining, among all the pixels of the sample image, pixels whose straight-line distances from the current pixel are less than the reference value. By respectively configuring the reference values in the first direction and the second direction, or configuring the reference values in the straight-line distance, it is possible to provide a standard for determining the reference point, and to use the pixels in all directions around the current pixel as reference points, thereby increasing the possibility of obtaining accurate offset pixels. It should be understood that when the sample image is in a rectangular coordinate system, the first direction is the x direction of the rectangular coordinate system, and the second direction is the y direction of the rectangular coordinate system. The reference values in the x direction and the y direction can determine the rectangular neighborhood around the current pixel. When the sample image is in a polar coordinate system, the first direction is the polar diameter direction of the polar coordinate system, and the second direction is the polar angle direction of the polar coordinate system. The reference values in the polar diameter direction and the polar angle direction can determine the fan-shaped neighborhood or fan-shaped ring neighborhood around the current pixel; the reference value on the straight-line distance can determine the circular neighborhood around the current pixel. It can be selected as needed in practice. Of course, the neighborhood of the pixel near the boundary of the sample image is cut by the boundary and may not be a complete regular shape.
[0079] In the second step, the weight of the corresponding reference point of the current pixel is determined according to the reference depth value of the current pixel and the sample estimated depth value of each reference point of the current pixel.
[0080] The third step is to determine the offset of each reference point of the current pixel relative to the current pixel. The offset of each reference point relative to the current pixel is the difference between the coordinate value of the reference point and the coordinate value of the current pixel, which can be directly calculated.
[0081] Although in theory, it is also possible to directly take the reference point whose sample estimated depth value is closest to the reference depth value of the current pixel among all reference points as the offset pixel of the current pixel, it is found in the test that the reference offset obtained in this way is greatly affected by noise. According to an exemplary embodiment of the present disclosure, by determining the weight of each reference point and obtaining the weighted average of the offsets of all reference points of the current pixel in the fourth step as the reference offset, the influence of noise can be weakened and the accuracy of the obtained reference offset can be improved.
[0082] Optionally, the operation of determining the weight of each reference point of the current pixel specifically includes: determining the absolute value of the error between the sample estimated depth value of each reference point of the current pixel and the reference depth value of the current pixel as the error of the corresponding reference point; determining the weight of each reference point of the current pixel according to the error of each reference point of the current pixel, and the weight of each reference point is negatively correlated with the error of the corresponding reference point. By calculating the error of the reference point and determining the weight that is negatively correlated with the error, the weight can be used to intuitively reflect the similarity between the sample estimated depth value of each reference point and the reference depth value of the current pixel, and the influence of local noise can be reduced by weighting, which can effectively improve the accuracy of the obtained reference offset.
[0083] Specifically, the operation of determining the weight of each reference point of the current pixel point according to the error of each reference point of the current pixel point can be specifically performed as follows: determining the sum of the error of each reference point of the current pixel point and the set value, and determining the ratio of the set value to the sum as the weight of the corresponding reference point; wherein the set value is a positive number. That is to say, for each reference point, the weight w, the error diff, and the set value a satisfy w = a / (diff + a), which can make the weight negatively correlated with the error, and ensure that when the error diff is 0, the weight w is not only meaningful, but also just equal to 1, so that the error size can be intuitively reflected through the weight. It should be understood that the set value will also affect the weight value when the error is not 0. The larger the set value, the larger the weight corresponding to the same error. It can be reasonably selected according to the actual calculation requirements and calculation effects. As an example, the set value is 1, that is, the weight w = 1 / (diff + 1).
[0084] The fourth step is to calculate the weighted average of the offsets of all reference points of the current pixel according to the weight of each reference point of the current pixel, and use it as the reference offset of the current pixel. For a rectangular coordinate system, the reference offset can be expressed by the formula:
[0085] dx_gt(i,j) = ∑(mi) w(m,n) / ∑w(m,n)
[0086] dy_gt(i,j) = ∑(nj) w(m,n) / ∑w(m,n)
[0087] Among them, consistent with the previous text, the formula is for the current pixel point p(i,j), dx_gt(i,j) represents the x-direction reference offset of the current pixel point p(i,j), dy_gt(x,y) represents the y-direction reference offset of the current pixel point p(i,j), any reference point of the current pixel point p(i,j) is represented as p''(m,n), (mi) represents the x-direction offset of the reference point p''(m,n) relative to the current pixel point p(i,j), (nj) represents the y-direction offset of the reference point p''(m,n) relative to the current pixel point p(i,j), and w(m,n) represents the weight of the reference point p''(m,n).
[0088] In step 504, a loss value is determined based on the reference depth value of each pixel of the sample image, the sample estimated depth value, the reference offset and the sample estimated offset.
[0089] Optionally, in view of the fact that the depth estimation model has two output values, namely, an estimated depth value and an estimated offset, step 504 may specifically include: calculating a first loss value based on a reference depth value of each pixel of the sample image and a sample estimated depth value; calculating a second loss value based on a reference offset of each pixel of the sample image and a sample estimated offset; and determining a loss value based on the first loss value and the second loss value. By respectively determining the first loss value and the second loss value, the loss value for the depth value and the loss value for the offset can be clearly determined, thereby ensuring that the model training is carried out in an orderly and reliable manner.
[0090] In step 505, the parameters of the depth estimation model are adjusted based on the loss value to train the depth estimation model. Specifically, the parameters of the depth estimation model may be adjusted using a back propagation algorithm.
[0091] Reference Figure 6 In order to determine the second loss value, a reference offset generation module needs to be introduced to perform step 503 to determine the reference offset. In order to ensure that only the parameters of the depth estimation model are adjusted, gradient truncation can be performed on the reference offset generation module during training, that is, the parameters of the reference offset generation module are not adjusted.
[0092] Figure 7 is a block diagram showing an image processing apparatus according to an exemplary embodiment of the present disclosure. It should be understood that the image processing apparatus according to the exemplary embodiment of the present disclosure can be implemented in a terminal device such as a smart phone, a tablet computer, a personal computer (PC) in the form of software, hardware, or a combination of software and hardware, and can also be implemented in a device such as a server.
[0093] Reference Figure 7 The image processing device 700 includes an estimation unit 701, an offset unit 702, a correction unit 703, and a summary unit 704.
[0094] The estimation unit 701 may perform depth estimation processing on the target image to obtain an estimated depth value and an estimated offset of each pixel point of the target image.
[0095] The estimated depth values of each pixel of the target image are collected together to form an estimated depth map. A significant difference between the image processing method according to the exemplary embodiment of the present disclosure and the depth estimation method in the related art is that, in addition to the estimated depth value, an estimated offset is also obtained, which can be used by other units to correct the distribution of depth values on the estimated depth map.
[0096] As an example, see Figure 1a The target image shown, Figure 4a The estimated depth map with a blurred boundary obtained by the estimation unit 701 is shown.
[0097] Optionally, the estimation unit 701 performs image processing on the target image, specifically inputting the target image into a depth estimation model to obtain an output result of the depth estimation model; wherein the output result includes an estimated depth value and an estimated offset for each pixel of the target image. A depth estimation model can be used to output the estimated depth value and the estimated offset at the same time, without having to calculate the two separately, which helps to simplify the calculation process. The training process of the depth estimation model is described above and will not be repeated here.
[0098] The offset unit 702 can determine the offset pixel of the corresponding pixel according to the estimated offset of each pixel of the target image. Combined with the estimated offset of each pixel, the corresponding pixel can be offset to the offset pixel, so that other units can correct the depth value of the corresponding pixel based on the offset pixel.
[0099] Optionally, the offset unit 702 can obtain the coordinate value of the offset pixel point of the corresponding pixel point by adding the coordinate value of the corresponding pixel point to the estimated offset amount of the corresponding pixel point for each pixel point of the target image. By performing the coordinate value calculation, the accurate offset pixel point can be obtained, which ensures the simplicity and reliability of the calculation process.
[0100] Optionally, the target image is a two-dimensional plane image, each pixel has a first direction coordinate value and a second direction coordinate value, and the estimated offset of each pixel includes the first direction estimated offset and the second direction estimated offset, so that independent offsets in two directions can be achieved. Based on this, the offset unit 702 can obtain the first direction coordinate value of the offset pixel point of the corresponding pixel point by adding the first direction coordinate value of the corresponding pixel point and the first direction estimated offset of the corresponding pixel point for each pixel point of the target image, and obtain the second direction coordinate value of the offset pixel point of the corresponding pixel point by adding the second direction coordinate value of the corresponding pixel point and the second direction estimated offset of the corresponding pixel point. By performing offset calculations on the two coordinate directions respectively and summarizing the calculation results, the coordinate values of the offset pixel points can be obtained, thereby ensuring the accuracy of the offset calculation.
[0101] Optionally, when the target image is in a rectangular coordinate system, the first direction is the x direction of the rectangular coordinate system, and the second direction is the y direction of the rectangular coordinate system, that is, the rectangular coordinate system is used to describe the coordinate values of the pixel points, which can be adapted to the calculation of conventional rectangular images. Figure 3 , the estimated offset in the x direction can be expressed as dx, and the estimated offset in the y direction can be expressed as dy. Figure 1a The target image shown, Figure 4b A schematic diagram showing the estimated x-direction offset obtained by the offset unit 702 is shown. Figure 4c A schematic diagram showing the estimated offset in the y direction obtained by the offset unit 702 is shown.
[0102] Optionally, when the target image is in a polar coordinate system, the first direction is the polar radius direction of the polar coordinate system, and the second direction is the polar angle direction of the polar coordinate system, that is, the polar coordinate system is used to describe the coordinate values of the pixel points, which can be adapted to the calculation and description of circular images.
[0103] It should be understood that in actual calculations, a suitable coordinate system may be used as needed, and the present disclosure does not limit this.
[0104] The correction unit 703 can use the estimated depth value corresponding to the offset pixel point of each pixel point in the target image as the corrected depth value of the corresponding pixel point. The correction unit 703 performs a depth value replacement operation, by replacing the estimated depth value of each pixel point of the estimated depth map obtained by the estimation unit 701 with the estimated depth value corresponding to the offset pixel point of the pixel point, so as to adjust the distribution of the depth value on the estimated depth map to obtain a corrected depth value without recalculating each depth value, and thus without increasing the input size of the depth estimation model. The amount of calculation is very small and can be applied to mobile terminals. At the same time, since the guided filtering method is not adopted, the feature extraction through the deep learning network can distinguish between real depth edges and false edges, thereby improving clarity without introducing false edges.
[0105] Reference Figure 3 For pixel point p(i,j), D1(i,j) represents its estimated depth value, dx(i,j) represents its estimated offset in the x direction, and dy(i,j) represents its estimated offset in the y direction. Then pixel point p'(i+dx(i,j), j+dy(i,j)) represents its offset pixel point, and the corrected depth value of pixel point p(i,j) is D2(i,j) = D1(i+dx(i,j), j+dy(i,j)). Figure 4d The following figure shows a schematic diagram of the depth value replacement operation. Figure 4b to Figure 4d It can be seen that the estimated offset at the boundary is more obvious, which can be understood as replacing the depth values of the pixels in the boundary transition zone with the depth values of the surrounding non-transition zone pixels, thereby improving the clarity at the boundary. For non-transition zone pixels, the estimated offset is basically 0, that is, no depth value replacement is required, which not only improves the boundary clarity in a targeted manner, but also reduces the computational complexity of the depth value replacement operation.
[0106] Based on this, optionally, the correction unit 703 may be run on the pixel point only when the offset pixel point obtained by the offset unit 702 does not overlap with the corresponding pixel point. Furthermore, the offset unit 702 and the correction unit 703 may be run on the pixel point only when the estimated offset obtained by the estimation unit 701 is not 0. It should be understood that for the latter, since the estimated offset often includes the estimated offset in the first direction and the estimated offset in the second direction, the offset unit 702 and the correction unit 703 may be run on the pixel point only when both the estimated offset in the first direction and the estimated offset in the second direction are not 0.
[0107] The aggregation unit 704 may obtain a depth map of the target image based on the corrected depth value of each pixel of the target image. The corrected depth values of each pixel of the target image are aggregated together to obtain a depth map estimated by the image processing apparatus 700 according to an exemplary embodiment of the present disclosure. Figure 1a The target image shown, Figure 4e The depth map obtained by the summarization unit 704 is shown, and Figure 4a It can be clearly seen from the comparison that the transition zone has been eliminated and the boundary clarity has been significantly improved.
[0108] Figure 8 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure.
[0109] Reference Figure 8The electronic device 800 includes at least one memory 801 and at least one processor 802, wherein the at least one memory 801 stores a set of computer executable instructions, and when the computer executable instruction set is executed by the at least one processor 802, an image processing method according to an exemplary embodiment of the present disclosure is executed.
[0110] As an example, the electronic device 800 may be a PC, a tablet device, a personal digital assistant, a smart phone, or other device capable of executing the above instruction set. Here, the electronic device 800 is not necessarily a single electronic device, but may also be any device or circuit collection capable of executing the above instructions (or instruction sets) individually or in combination. The electronic device 800 may also be part of an integrated control system or system manager, or may be configured as a portable electronic device interconnected with a local or remote (e.g., via wireless transmission) interface.
[0111] In the electronic device 800, the processor 802 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller or a microprocessor. As an example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0112] The processor 802 may execute instructions or codes stored in the memory 801, wherein the memory 801 may also store data. Instructions and data may also be sent and received over a network via a network interface device, wherein the network interface device may employ any known transmission protocol.
[0113] The memory 801 may be integrated with the processor 802, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. In addition, the memory 801 may include a separate device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The memory 801 and the processor 802 may be operatively coupled, or may communicate with each other, such as through an I / O port, a network connection, etc., so that the processor 802 can read files stored in the memory.
[0114] In addition, the electronic device 800 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device 800 may be connected to each other via a bus and / or a network.
[0115] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided, and when the instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor is prompted to perform the image processing method according to the exemplary embodiment of the present disclosure. Examples of computer-readable storage media here include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device is configured to store computer programs and any associated data, data files and data structures in a non-transitory manner and provide the computer programs and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system, so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.
[0116] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided. The computer program product includes computer instructions. When the computer instructions are executed by at least one processor, the at least one processor is prompted to perform the image processing method according to the exemplary embodiment of the present disclosure.
[0117] According to the image processing method and device of the exemplary embodiment of the present disclosure, the depth map and the offset of the depth map can be estimated at the same time, and the distribution of the depth value on the estimated depth map can be corrected by the estimated offset, thereby eliminating the transition band and improving the clarity of the depth map. At the same time, the amount of calculation of the present disclosure using the offset to correct the depth map is very small, and it is only a depth value replacement operation, so it can be applied on the mobile terminal. In addition, compared with the method using RGB image guided filtering, the offset in the present disclosure is obtained through deep learning training. The feature extraction of the deep learning network can distinguish between real depth edges and false edges, and will not introduce false edges while clarifying the depth.
[0118] According to the training method of the depth estimation model of the exemplary embodiment of the present disclosure, in addition to the beneficial technical effects of the above-mentioned image processing method, a method for determining a reference offset is also proposed, which comprehensively considers the depth distribution trend in the local area, can effectively reduce the influence of noise, improve the accuracy of the determined reference offset, and ensure the estimation accuracy of the trained depth estimation model.
[0119] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0120] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An image processing method, characterized in that: The image processing method comprises: Inputting the target image into the depth estimation model to obtain an output result of the depth estimation model; wherein the output result includes an estimated depth value and an estimated offset for each pixel point of the target image; Determining an offset pixel point of a corresponding pixel point according to the estimated offset amount of each pixel point of the target image; Taking the estimated depth value corresponding to the offset pixel point of each pixel point in the target image as the corrected depth value of the corresponding pixel point; Based on the corrected depth value of each pixel of the target image, a depth map of the target image is obtained, The depth estimation model is trained by the following steps: Acquire a sample image and a reference depth value of each pixel of the sample image; Inputting the sample image into the depth estimation model to obtain a sample estimated depth value and a sample estimated offset for each pixel point of the sample image; Determining a reference offset of each pixel of the sample image according to the reference depth value of each pixel of the sample image and the sample estimated depth value; Determine a loss value based on the reference depth value, the sample estimated depth value, the reference offset and the sample estimated offset of each pixel of the sample image; Parameters of the depth estimation model are adjusted based on the loss value to train the depth estimation model.
2. The image processing method according to claim 1, characterized in that: The step of determining the offset pixel point of the corresponding pixel point according to the estimated offset amount of each pixel point of the target image comprises: For each pixel point of the target image, the coordinate value of the offset pixel point of the corresponding pixel point is obtained by adding the coordinate value of the corresponding pixel point to the estimated offset amount of the corresponding pixel point.
3. The image processing method according to claim 2, characterized in that: Each pixel point has a first direction coordinate value and a second direction coordinate value, and the estimated offset of each pixel point includes the first direction estimated offset and the second direction estimated offset, wherein for each pixel point of the target image, the coordinate value of the offset pixel point of the corresponding pixel point is obtained by adding the coordinate value of the corresponding pixel point to the estimated offset of the corresponding pixel point, including: For each pixel point of the target image, the first direction coordinate value of the offset pixel point of the corresponding pixel point is obtained by adding the first direction coordinate value of the corresponding pixel point and the first direction estimated offset of the corresponding pixel point, and the second direction coordinate value of the offset pixel point of the corresponding pixel point is obtained by adding the second direction coordinate value of the corresponding pixel point and the second direction estimated offset of the corresponding pixel point.
4. The image processing method according to claim 3, characterized in that: When the target image is in a rectangular coordinate system, the first direction is the x direction of the rectangular coordinate system, and the second direction is the y direction of the rectangular coordinate system; or When the target image is in a polar coordinate system, the first direction is the polar radius direction of the polar coordinate system, and the second direction is the polar angle direction of the polar coordinate system.
5. The image processing method according to any one of claims 1 to 4, characterized in that: The determining, according to the reference depth value of each pixel of the sample image and the sample estimated depth value, a reference offset of each pixel of the sample image comprises: Traversing all pixels of the sample image and determining a number of reference points around the current pixel; Determine a weight of a reference point corresponding to the current pixel according to the reference depth value of the current pixel and the sample estimated depth value of each reference point of the current pixel; Determine an offset of each reference point of the current pixel relative to the current pixel; According to the weight of each reference point of the current pixel, a weighted average of the offsets of all reference points of the current pixel is calculated as the reference offset of the current pixel.
6. The image processing method according to claim 5, characterized in that: The determining, according to the reference depth value of the current pixel and the sample estimated depth value of each reference point of the current pixel, the weight of the corresponding reference point of the current pixel comprises: Determine an absolute value of an error between the sample estimated depth value of each reference point of the current pixel and the reference depth value of the current pixel as an error of the corresponding reference point; According to the error of each reference point of the current pixel, the weight of each reference point of the current pixel is determined, and the weight of each reference point is negatively correlated with the error of the corresponding reference point.
7. The image processing method according to claim 6, characterized in that: The step of determining the weight of each reference point of the current pixel according to the error of each reference point of the current pixel comprises: Determine the sum of the error of each reference point of the current pixel and a set value, and determine the ratio of the set value to the sum as the weight of the corresponding reference point; wherein the set value is a positive number.
8. The image processing method according to claim 5, characterized in that: The determining of a plurality of reference points around the current pixel includes: Determine, among all the pixels of the sample image, a pixel whose distance from the current pixel in the first direction and the second direction is smaller than the corresponding reference value, as the reference point of the current pixel; or Among all the pixel points of the sample image, determine the pixel point whose straight-line distance to the current pixel point is less than a reference value.
9. The image processing method according to any one of claims 1 to 4, characterized in that: The determining of the loss value based on the reference depth value of each pixel of the sample image, the sample estimated depth value, the reference offset and the sample estimated offset comprises: Calculating a first loss value based on the reference depth value of each pixel of the sample image and the sample estimated depth value; Calculating a second loss value based on the reference offset and the sample estimated offset of each pixel of the sample image; The loss value is determined based on the first loss value and the second loss value.
10. An image processing device, characterized in that: The image processing device comprises: The estimation unit is configured to: input the target image into the depth estimation model to obtain an output result of the depth estimation model; wherein the output result includes an estimated depth value and an estimated offset of each pixel point of the target image; An offset unit is configured to: determine an offset pixel point of a corresponding pixel point according to the estimated offset amount of each pixel point of the target image; A correction unit is configured to: use the estimated depth value corresponding to the offset pixel point of each pixel point in the target image as a corrected depth value of the corresponding pixel point; a summarizing unit, configured to: obtain a depth map of the target image based on the corrected depth value of each pixel of the target image, The depth estimation model is trained by the following steps: Acquire a sample image and a reference depth value of each pixel of the sample image; Inputting the sample image into the depth estimation model to obtain a sample estimated depth value and a sample estimated offset for each pixel point of the sample image; Determining a reference offset of each pixel of the sample image according to the reference depth value of each pixel of the sample image and the sample estimated depth value; Determine a loss value based on the reference depth value, the sample estimated depth value, the reference offset and the sample estimated offset of each pixel of the sample image; Parameters of the depth estimation model are adjusted based on the loss value to train the depth estimation model.
11. The image processing device according to claim 10, wherein: The offset unit is further configured to: For each pixel point of the target image, the coordinate value of the offset pixel point of the corresponding pixel point is obtained by adding the coordinate value of the corresponding pixel point to the estimated offset amount of the corresponding pixel point.
12. The image processing device according to claim 11, wherein: Each pixel point has a first direction coordinate value and a second direction coordinate value, and the estimated offset of each pixel point includes the first direction estimated offset and the second direction estimated offset, and the offset unit is further configured as follows: For each pixel point of the target image, the first direction coordinate value of the offset pixel point of the corresponding pixel point is obtained by adding the first direction coordinate value of the corresponding pixel point and the first direction estimated offset of the corresponding pixel point, and the second direction coordinate value of the offset pixel point of the corresponding pixel point is obtained by adding the second direction coordinate value of the corresponding pixel point and the second direction estimated offset of the corresponding pixel point.
13. The image processing device according to claim 12, wherein: When the target image is in a rectangular coordinate system, the first direction is the x direction of the rectangular coordinate system, and the second direction is the y direction of the rectangular coordinate system; or When the target image is in a polar coordinate system, the first direction is the polar radius direction of the polar coordinate system, and the second direction is the polar angle direction of the polar coordinate system.
14. The image processing device according to any one of claims 10 to 13, characterized in that: The determining, according to the reference depth value of each pixel of the sample image and the sample estimated depth value, a reference offset of each pixel of the sample image comprises: Traversing all pixels of the sample image and determining a number of reference points around the current pixel; Determine a weight of a reference point corresponding to the current pixel according to the reference depth value of the current pixel and the sample estimated depth value of each reference point of the current pixel; Determine an offset of each reference point of the current pixel relative to the current pixel; According to the weight of each reference point of the current pixel, a weighted average of the offsets of all reference points of the current pixel is calculated as the reference offset of the current pixel.
15. The image processing device according to claim 14, wherein: The determining, according to the reference depth value of the current pixel and the sample estimated depth value of each reference point of the current pixel, the weight of the corresponding reference point of the current pixel comprises: Determine an absolute value of an error between the sample estimated depth value of each reference point of the current pixel and the reference depth value of the current pixel as an error of the corresponding reference point; According to the error of each reference point of the current pixel, the weight of each reference point of the current pixel is determined, and the weight of each reference point is negatively correlated with the error of the corresponding reference point.
16. The image processing device according to claim 15, characterized in that: The step of determining the weight of each reference point of the current pixel according to the error of each reference point of the current pixel comprises: Determine the sum of the error of each reference point of the current pixel and a set value, and determine the ratio of the set value to the sum as the weight of the corresponding reference point; wherein the set value is a positive number.
17. The image processing device according to claim 14, wherein: The determining of a plurality of reference points around the current pixel includes: Determine, among all the pixels of the sample image, a pixel whose distance from the current pixel in the first direction and the second direction is smaller than the corresponding reference value, as the reference point of the current pixel; or Among all the pixel points of the sample image, determine the pixel point whose straight-line distance to the current pixel point is less than a reference value.
18. The image processing apparatus according to any one of claims 10 to 13, characterized in that: The determining of the loss value based on the reference depth value of each pixel of the sample image, the sample estimated depth value, the reference offset and the sample estimated offset comprises: Calculating a first loss value based on the reference depth value of each pixel of the sample image and the sample estimated depth value; Calculating a second loss value based on the reference offset and the sample estimated offset of each pixel of the sample image; The loss value is determined based on the first loss value and the second loss value.
19. An electronic device, characterized in that: include: at least one processor; at least one memory storing computer executable instructions, Wherein, when the computer executable instructions are executed by the at least one processor, the at least one processor is prompted to perform the image processing method according to any one of claims 1 to 9.
20. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor is prompted to perform the image processing method according to any one of claims 1 to 9.
21. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by at least one processor, the image processing method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Depth map generation method and device, electronic equipment and storage medium
CN113012210A