Image Processing Method, Apparatus, Electronic Device, and Storage Medium
By obtaining the foreground and background edges in the depth map, flooding and information completion are performed, the distortion and tearing of the depth fault region of 3D photos are solved, and a clearer image completion effect is achieved.
Patent Information
- Application Number
- CN202210345732.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-02
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-04-02
AI Technical Summary
When completing images of 3D photos, the prior art is prone to introduce artificial traces such as distortion, deformation and tear into the depth fault area of the object, resulting in unsatisfactory completion results.
By obtaining the foreground and background edges in the depth map, flooding is performed separately, the occlusion area is determined, and the known background area is used to complete the depth information and RGB information to form a clear complete background image.
Accurately distinguishing the depth fault areas of the foreground and background achieves clearer and more reasonable image completion, reducing distortion and tearing, and improving image quality.
Smart Images

Figure CN114677426B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer vision technology, and in particular, to an image processing method, apparatus, electronic device, storage medium, and computer program product. Background Art
[0002] Currently, with the development of photography technology, in addition to taking 2D photos, people can also take 3D photos. Compared with 2D photos, 3D photos can reflect the depth of objects, thus providing more realism. However, in some shooting scenarios, the object to be photographed may be blocked by other objects, or the complete image of the object cannot be seen due to parallax. Therefore, it is necessary to perform post-processing on the photographed pictures, such as filling in the blocked parts of the pictures. However, existing image completion methods, such as deep learning-based distortion correction techniques, are prone to introducing artificial traces such as distortion, deformation, and tearing in the depth discontinuity region between two objects, resulting in unsatisfactory image completion results. Summary of the Invention
[0003] The present disclosure provides an image processing method, apparatus, electronic device, storage medium, and computer program product to at least solve the problem that the existing image completion methods are prone to introducing artificial traces such as distortion, deformation, and tearing in the depth discontinuity region between two objects. The technical solution of the present disclosure is as follows:
[0004] According to the first aspect of the embodiments of the present disclosure, an image processing method is provided, including:
[0005] Obtaining a depth map corresponding to the image to be processed, and determining a first edge of the foreground object and a second edge of the background object from the depth map;
[0006] Performing first flood filling on the first direction of the first edge and the second direction of the second edge in the depth map respectively to obtain a foreground region formed after the first edge is flood filled and a known background region formed after the second edge is flood filled, and determining an occlusion region based on the foreground region; wherein, the occlusion region represents the background region occluded by the foreground region; the first direction is opposite to the second direction;
[0007] Performing depth information completion and RGB information completion processing on the occlusion region through the known background region in the depth map to obtain a completed background image corresponding to the occlusion region.
[0008] In an exemplary embodiment, the determining a first edge of the foreground object and a second edge of the background object from the depth map includes:
[0009] Obtain the connectivity relationship between each pixel point in the depth map, and determine the number of connected pixel points of each pixel point based on the connectivity relationship; wherein, the depth difference between two connected pixel points is within a preset range;
[0010] Determine a target pixel point from the depth map; the number of connected pixel points corresponding to the target pixel point is less than a number threshold;
[0011] Based on the depth value of the target pixel point, divide the target pixel point into a pixel point of the first edge or a pixel point of the second edge, and obtain the first edge and the second edge based on the divided target pixel points.
[0012] In an exemplary embodiment, the first flood processing of the first direction of the first edge and the second direction of the second edge respectively includes:
[0013] Determine the direction of the first edge towards the center point of the foreground object as the first direction, and determine the direction of the second edge away from the center point of the foreground object as the second direction;
[0014] Based on the connectivity relationship between each pixel point in the depth map, perform first flood processing on the first direction of the first edge and the second direction of the second edge respectively.
[0015] In an exemplary embodiment, the to-be-processed image at least includes a first object, a second object, and a third object, wherein the first object is the foreground object of the second object, and the second object is the foreground object of the third object; the depth information and RGB information of the occluded area are complemented through the known background area to obtain the complemented background image corresponding to the occluded area, including:
[0016] Obtain the first edge of the first object and the second edge of the second object as a first pair of depth edges, and based on the first pair of depth edges, obtain the known background area of the second object and the occluded area occluded by the first object;
[0017] Obtain the first edge of the second object and the second edge of the third object as a second pair of depth edges, and based on the second pair of depth edges, obtain the known background area of the third object and the occluded area occluded by the second object;
[0018] If it is determined that the first object does not have edge occlusion on the second pair of depth edges, use the known background area of the second object to perform depth information complementation and RGB information complementation on the occluded area occluded by the first object of the second object to obtain the complemented background image corresponding to the second object;
[0019] Complete the depth information and RGB information of the occluded area of the third object occluded by the second object through the known background area of the third object to obtain the complementary background image corresponding to the third object.
[0020] In an exemplary embodiment, the process of completing the depth information and RGB information of the occluded area through the known background area to obtain the complementary background image corresponding to the occluded area further includes:
[0021] Obtain the first edge of the first object and the second edge of the third object as the third pair of depth edges, and based on the second pair of depth edges and the third pair of depth edges, obtain the known background area of the third object, the occluded area of the third object occluded by the first object, and the occluded area of the third object occluded by the second object;
[0022] If it is determined that the first object occludes a part of the second pair of depth edges, complete the occluded edge in the second pair of depth edges to obtain the completed second pair of depth edges;
[0023] Complete the depth information and RGB information of the occluded area of the second object occluded by the first object through the known background area of the second object to obtain the complementary background image corresponding to the second object; and complete the depth information and RGB information of the occluded area of the third object occluded by the first object through the known background area of the third object to obtain the initial complementary image corresponding to the third object;
[0024] Complete the depth information and RGB information of the occluded area of the third object occluded by the second object through the initial complementary image corresponding to the third object to obtain the complementary background image corresponding to the third object.
[0025] In an exemplary embodiment, the process of completing the depth information and RGB information of the occluded area of the third object occluded by the first object through the known background area of the third object to obtain the initial complementary image corresponding to the third object includes:
[0026] Complete the depth information and RGB information of the occluded area of the third object occluded by the first object through the known background area of the third object to obtain the partial complementary image of the occluded area of the third object occluded by the first object;
[0027] Based on the connectivity relationship between the pixel points in the locally completed image and the known background area of the third object, perform a second flooding process on the second direction of the second edge in the completed second pair of depth edges to obtain the initial completed image of the third object.
[0028] In an exemplary embodiment, the method further includes:
[0029] When it is detected that the second pair of depth edges is discontinuous, it is determined that the first object occludes the second pair of depth edges.
[0030] In an exemplary embodiment, before performing the first flooding process on the first direction of the first edge and the second direction of the second edge in the depth map respectively, it further includes:
[0031] Determine the pixels to be filtered from the pixel points of the background object in the depth map;
[0032] Obtain a sampling window corresponding to the background object, and perform median filtering on the pixels to be filtered based on the sampling window to obtain a filtered depth map;
[0033] The performing the first flooding process on the first direction of the first edge and the second direction of the second edge in the depth map respectively includes:
[0034] In the filtered depth map, perform the first flooding process on the first direction of the first edge and the second direction of the second edge respectively.
[0035] In an exemplary embodiment, the determining the pixels to be filtered from the pixel points of the background object in the depth map includes:
[0036] Determine candidate pixels with a depth value less than a depth threshold from the pixel points of the background object in the depth map;
[0037] Determine the range where the candidate pixels are located in the depth map and the gradient magnitude of the candidate pixels; the gradient magnitude represents the depth difference between the candidate pixels and adjacent pixels in a preset gradient direction;
[0038] Determine the pixels to be filtered from the candidate pixels according to the range where the candidate pixels are located and the gradient magnitude of the candidate pixels.
[0039] In an exemplary embodiment, there are multiple preset gradient directions, and the determining the pixels to be filtered from the candidate pixels according to the range where the candidate pixels are located and the gradient magnitude of the candidate pixels includes:
[0040] If the candidate pixel is within the first range, when the gradient magnitude of the candidate pixel in any gradient direction is greater than the first threshold, determine that the candidate pixel is a pixel to be filtered; the first range represents a range that is within the background object of the depth map and outside the second edge of the background object;
[0041] If the candidate pixel is within the second range, when the gradient magnitude of the candidate pixel in any gradient direction is greater than the second threshold, determine that the candidate pixel is a pixel to be filtered; wherein, the first threshold is greater than the second threshold, and the second range represents a range that is within the background object of the depth map and within the second edge of the background object.
[0042] In an exemplary embodiment, after obtaining the completed background image corresponding to the occluded area, it further includes:
[0043] According to the known background image and the completed background image, obtain the depth information and RGB information of each layer of the image to be processed;
[0044] According to the depth information and RGB information of each layer, obtain the reconstructed three-dimensional mesh information corresponding to the image to be processed;
[0045] Based on the set new view angle and the three-dimensional mesh information, construct a new view angle image of the image to be processed.
[0046] According to a second aspect of the embodiments of the present disclosure, there is provided an image processing apparatus, including:
[0047] A determination unit configured to execute obtaining a depth map corresponding to an image to be processed, and determining a first edge of a foreground object and a second edge of a background object from the depth map;
[0048] A flooding unit configured to execute respectively performing a first flooding process on a first direction of the first edge and a second direction of the second edge in the depth map, obtaining a foreground area formed after flooding the first edge and a known background area formed after flooding the second edge, and determining an occluded area based on the foreground area; wherein, the occluded area represents a background area occluded by the foreground area; the first direction is opposite to the second direction;
[0049] A completion unit configured to execute performing depth information completion and RGB information completion processing on the occluded area through the known background area in the depth map, obtaining a completed background image corresponding to the occluded area.
[0050] In an exemplary embodiment, the determining unit is further configured to obtain the connectivity relationship between pixel points in the depth map, and determine the number of connected pixel points of each pixel point based on the connectivity relationship; wherein, the depth difference between two connected pixel points is within a preset range; determine a target pixel point from the depth map; the number of connected pixel points corresponding to the target pixel point is less than a number threshold; based on the depth value of the target pixel point, divide the target pixel point into a pixel point of a first edge or a pixel point of a second edge, and obtain the first edge and the second edge based on the divided target pixel points.
[0051] In an exemplary embodiment, the flooding unit is further configured to determine that the direction in which the first edge tends to the center point of the foreground object is a first direction, and determine that the direction in which the second edge faces away from the center point of the foreground object is a second direction; based on the connectivity relationship between pixel points in the depth map, perform a first flooding process on the first direction of the first edge and the second direction of the second edge respectively.
[0052] In an exemplary embodiment, the image to be processed at least includes a first object, a second object, and a third object, wherein the first object is the foreground object of the second object, and the second object is the foreground object of the third object; the filling unit is further configured to obtain the first edge of the first object and the second edge of the second object as a first pair of depth edges, and based on the first pair of depth edges, obtain the known background area of the second object and the occluded area occluded by the first object; obtain the first edge of the second object and the second edge of the third object as a second pair of depth edges, and based on the second pair of depth edges, obtain the known background area of the third object and the occluded area occluded by the second object; if it is determined that the first object does not have edge occlusion on the second pair of depth edges, use the known background area of the second object to perform depth information filling and RGB information filling on the occluded area of the second object occluded by the first object to obtain a filled background image corresponding to the second object; use the known background area of the third object to perform depth information filling and RGB information filling on the occluded area of the third object occluded by the second object to obtain a filled background image corresponding to the third object.
[0053] In an exemplary embodiment, the completion unit is further configured to execute obtaining a first edge of the first object and a second edge of the third object as a third pair of depth edges, and obtaining a known background region of the third object, an occluded region of the third object occluded by the first object, and an occluded region of the third object occluded by the second object based on the second pair of depth edges and the third pair of depth edges; if it is determined that the first object occludes a part of the second pair of depth edges, performing edge completion on the occluded edge in the second pair of depth edges to obtain a completed second pair of depth edges; performing depth information completion and RGB information completion on the occluded region of the second object occluded by the first object through the known background region of the second object to obtain a completed background image corresponding to the second object; and performing depth information completion and RGB information completion on the occluded region of the third object occluded by the first object through the known background region of the third object to obtain an initial completed image corresponding to the third object; performing depth information completion and RGB information completion on the occluded region of the third object occluded by the second object through the initial completed image corresponding to the third object to obtain a completed background image corresponding to the third object.
[0054] In an exemplary embodiment, the completion unit is further configured to execute performing depth information completion and RGB information completion on the occluded region of the third object occluded by the first object through the known background region of the third object to obtain a partial completed image of the occluded region of the third object occluded by the first object; performing a second flooding process on a second direction of a second edge in the completed second pair of depth edges based on a connectivity relationship between pixel points in the partial completed image and pixel points in the known background region of the third object to obtain an initial completed image of the third object.
[0055] In an exemplary embodiment, the apparatus further includes a detection unit configured to execute determining that the first object occludes the second pair of depth edges when it is detected that the second pair of depth edges is discontinuous.
[0056] In an exemplary embodiment, the apparatus further includes a filtering unit configured to execute determining, from pixel points of a background object in the depth map, pixel points to be filtered; obtaining a sampling window corresponding to the background object, and performing median filtering on the pixel points to be filtered based on the sampling window to obtain a filtered depth map;
[0057] The flooding unit is further configured to execute performing a first flooding process on a first direction of the first edge and a second direction of the second edge in the filtered depth map, respectively.
[0058] In an exemplary embodiment, the filtering unit is further configured to determine candidate pixel points with depth values less than a depth threshold from the pixel points of the background object in the depth map; determine the range where the candidate pixel points are located in the depth map and the gradient amplitude of the candidate pixel points; the gradient amplitude represents the depth difference between the candidate pixel points and adjacent pixel points in a preset gradient direction; and determine to-be-filtered pixel points from the candidate pixel points according to the range where the candidate pixel points are located and the gradient amplitude of the candidate pixel points.
[0059] In an exemplary embodiment, there are multiple preset gradient directions, and the filtering unit is further configured to perform: if the candidate pixel points are in a first range, when the gradient amplitude of the candidate pixel points in any gradient direction is greater than a first threshold, determine that the candidate pixel points are to-be-filtered pixel points; the first range represents the range that is in the background object of the depth map and outside the second edge of the background object; if the candidate pixel points are in a second range, when the gradient amplitude of the candidate pixel points in any gradient direction is greater than a second threshold, determine that the candidate pixel points are to-be-filtered pixel points; wherein, the first threshold is greater than the second threshold, and the second range represents the range that is in the background object of the depth map and inside the second edge of the background object.
[0060] In an exemplary embodiment, the device further includes a new perspective image construction unit, which is configured to perform: obtain the depth information and RGB information of each layer of the to-be-processed image according to the known background image and the completed background image; obtain the reconstructed three-dimensional mesh information corresponding to the to-be-processed image according to the depth information and RGB information of each layer; and construct a new perspective image of the to-be-processed image based on a set new perspective and the three-dimensional mesh information.
[0061] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0062] a processor;
[0063] a memory for storing instructions executable by the processor;
[0064] wherein, the processor is configured to execute the instructions to implement the method as described in any one of the above.
[0065] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method as described in any one of the above.
[0066] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, which includes instructions that, when executed by a processor of an electronic device, enable the electronic device to execute the method described in any one of the above.
[0067] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0068] After obtaining the depth map corresponding to the image to be processed, the method separates the foreground object and the background object in the depth map through the first edge and the second edge, then performs the first flooding process on the first edge and the second edge respectively to obtain the foreground area formed after the first edge is flooded and the known background area formed after the second edge is flooded, determines the occlusion area based on the foreground area, and finally performs depth information completion and RGB information completion processing on the occlusion area through the known background area to obtain the completed background image corresponding to the occlusion area. This method of disconnecting the pixel points at the junction of the foreground object and the background object to form the first edge and the second edge, performing the flooding process on the first edge and the second edge respectively, and then performing completion can more accurately distinguish the pixel points in the depth fault area between the foreground object and the background object, so as to accurately complete the depth information and RGB information of the pixel points near the fault area and obtain a clearer and more reasonable completion result.
[0069] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation to the present disclosure.
[0071] Figure 1 is a flowchart showing a method for image processing according to an exemplary embodiment.
[0072] Figure 2 is a schematic diagram showing a process of image flooding according to an exemplary embodiment.
[0073] Figure 3 is a schematic diagram showing the visualization effect of the depth edge on the image and the depth map after 3 times of filtering according to an exemplary embodiment.
[0074] Figure 4 is a schematic diagram showing a new perspective image according to an exemplary embodiment.
[0075] Figure 5 is a complete flowchart showing a method for image processing according to an exemplary embodiment.
[0076] Figure 6 It is a schematic diagram showing the orientation completion result according to another exemplary embodiment.
[0077] Figure 7 It is a structural block diagram of an image processing apparatus shown according to an exemplary embodiment.
[0078] Figure 8 It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0079] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0080] It should be noted that the implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0081] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.
[0082] In an exemplary embodiment, as Figure 1 shown, an image processing method is provided. In this embodiment, it is exemplified that the method is applied to a terminal. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers. In this embodiment, the method includes the following steps:
[0083] In step S110, a depth map corresponding to the image to be processed is obtained, and a first edge of the foreground object and a second edge of the background object are determined from the depth map.
[0084] The depth map represents an image that uses the distance (depth) from the image collector to each point in the scene as a pixel value, and the depth map directly reflects the geometric shape of the visible surface of the scene.
[0085] The first edge and the second edge can be understood as two edges obtained by disconnecting the contour pixels at the depth edge formed at the junction of the foreground object and the background object. There is a large difference in depth value between the pixel points on the first edge and the pixel points on the second edge, that is, the pixel points on the two edges are not on the same plane. Figure 2 In the image, the contour area of the human body formed between the sea surface and the human body is a depth fault area. In this fault area, the pixel points closer to the human body constitute the first edge, and the pixel points closer to the sea surface constitute the second edge.
[0086] Among them, foreground and background are relative concepts. The foreground object can be understood as the object in the image that is closer to the image collector (or camera), showing a certain spatial relationship or human relationship. The background is behind the foreground object and is the scenery far away from the camera. For example, refer to Figure 2 , Figure 2 The human body in the image is a foreground object relative to the sea surface or the sky, the sea surface or the sky is a background object relative to the human body, and the sea surface is a foreground object relative to the sky.
[0087] In a specific implementation, the depth estimation of the image to be processed can be performed through the trained depth estimation network to obtain the depth map corresponding to the image to be processed. Among them, the depth estimation network can use a deep convolutional network as a feature extractor, and the network performs multiple pooling to reduce the resolution, and then predicts a high-resolution depth map through modules such as deconvolution layers, upsampling layers, multi-scale network structures, and skip connections. After obtaining the depth map, determine the position in the depth map where the change in depth value is greater than the threshold, and the position of the junction of the foreground object and the background object, disconnect the connected pixels at this position, and form the first edge of the foreground object and the second edge of the background object.
[0088] In step S120, a first flooding process is performed on the first direction of the first edge and the second direction of the second edge in the depth map, respectively, to obtain a foreground area formed after the first edge flooding and a known background area formed after the second edge flooding, and an occlusion area is determined based on the foreground area; wherein the occlusion area represents a background area occluded by the foreground area, and the first direction is opposite to the second direction.
[0089] Among them, the first direction can represent the direction in which the first edge tends towards the center point of the foreground object, and the second direction can represent the direction in which the second edge faces away from the center point of the foreground object.
[0090] Among them, flooding can represent the process of connecting connected pixel points in an image.
[0091] Among them, the occluded area represents the background area occluded by the foreground area. For example, referring to Figure 2 , is a schematic diagram of flood processing an image based on the detected first edge and second edge in an embodiment, Figure 2 from left to right in, are schematic diagrams of flood processing 0 times, 50 times, 100 times, 150 times, and 200 times respectively. In Figure 2 , the area 200b where the sea surface is occluded by the human body can represent the occluded area, and the area 202b where the sky is occluded by the sea surface can also represent the occluded area.
[0092] In specific implementation, after obtaining the depth map, the connectivity relationship between each pixel point can be determined based on the depth value of each pixel point in the depth map, so as to perform the first flood processing on the first direction of the first edge and the second direction of the second edge respectively based on this connectivity relationship, obtain the foreground area formed after the first edge is flooded and the known background area formed after the second edge is flooded, and determine the occluded area based on the foreground area.
[0093] More specifically, since the background object is farther from the camera than the foreground object, the depth value is larger. Therefore, the area with a larger depth value after the first flood processing is the known background area, and the area with a smaller depth value is the foreground area. For example, in Figure 2 , for the human body and the sea surface, the area 200b formed by flooding within the human body contour in the direction towards the center point of the human body is the foreground area of the human body, while the area 200a formed by flooding outside the human body contour in the direction away from the center point of the human body is the known background area of the sea surface. For the sky and the sea surface, the area 202b formed by flooding below the boundary line between the sky and the sea surface in the direction towards the center point of the sea surface is the foreground area of the sky, and the area 202a formed by flooding above the boundary line between the sky and the sea surface in the direction away from the center point of the sea surface is the known background area of the sky.
[0094] Furthermore, after obtaining the foreground area formed after the first edge is flooded and the known background area formed after the second edge is flooded, a consistency check can also be performed, that is, to check whether there are pixel points in the foreground area that are determined to be in the known background area, and whether there are pixel points in the known background area that are determined to be in the foreground area, and when this situation occurs, correction is performed.
[0095] In step S130, in the depth map, the depth information and RGB information of the occluded area are complemented through the known background area to obtain the complemented background image corresponding to the occluded area.
[0096] Among them, image completion is a technology that complements the missing area of the image to be repaired according to the information of the image itself or the image library, making the repaired image look more natural.
[0097] Among them, RGB represents the three primary colors of light, R represents Red (red), G represents Green (green), and B represents Blue (blue).
[0098] In specific implementation, when the image to be processed only includes a foreground object and a background object, the depth information and RGB information of the occluded area can be directly complemented through the known background area. When the image to be processed includes at least a first object, a second object, and a third object, the complementation of the occluded area can be divided into two cases. One is that the first object does not occlude the depth edge formed by the second object and the third object. At this time, the depth information and RGB information of their respective occluded areas are complemented through the known background areas of the second object and the third object respectively. The other is that the first object occludes the depth edge formed by the second object and the third object. At this time, the occluded edge needs to be complemented first, and then the depth information and RGB information of the occluded area are complemented according to the complemented edge and the known background area to obtain the complemented background image corresponding to the occluded area.
[0099] In the above image processing method, after obtaining the depth map corresponding to the image to be processed, the foreground object and the background object in the depth map are separated by the first edge and the second edge, and then the first edge and the second edge are respectively subjected to the first flood filling process to obtain the foreground area formed after the first edge is flooded and the known background area formed after the second edge is flooded, and the occluded area is determined based on the foreground area. Finally, the depth information and RGB information of the occluded area are complemented through the known background area to obtain the complemented background image corresponding to the occluded area. This method of disconnecting the pixel points at the junction of the foreground object and the background object to form the first edge and the second edge, respectively performing flood filling on the first edge and the second edge, and then performing complementation can more accurately distinguish the pixel points in the depth fault area between the foreground object and the background object, so that the pixel points near the fault area can be accurately complemented with depth information and RGB information to obtain a clearer and more reasonable complementation result.
[0100] In an exemplary embodiment, the above step S110 can be specifically determined in the following manner: Obtain the connectivity relationship between each pixel point in the depth map, and determine the number of connected pixel points of each pixel point based on the connectivity relationship; wherein, the depth difference between two connected pixel points is within a preset range; Determine the target pixel point from the depth map; the number of connected pixel points corresponding to the target pixel point is less than the number threshold; Based on the depth value of the target pixel point, divide the target pixel point into a pixel point of the first edge or a pixel point of the second edge, and obtain the first edge and the second edge based on the divided target pixel points.
[0101] Wherein, the number of connected pixel points represents the number of pixel points having a connectivity relationship with the pixel point.
[0102] In specific implementation, the first edge and the second edge are two edges generated by a depth value fault. Therefore, the depth difference between the pixel points on the first edge and the second edge and at least one adjacent pixel point is greater than or equal to the difference threshold, that is, there is no connectivity relationship between the pixel points on the first edge and the second edge and at least one adjacent pixel point. Therefore, the first edge of the foreground object and the second edge of the background object can be determined according to the number of connected pixel points of each pixel point.
[0103] More specifically, first, according to the connectivity relationship between each pixel point in the depth map, determine the number of connected pixel points of each pixel point, and determine the pixel points with the number of connected pixel points less than the number threshold from each pixel point as the target pixel points. Based on the depth value of the target pixel points, divide the target pixel points with larger depth values into the pixel points of the second edge of the background object, and divide the target pixel points with smaller depth values into the pixel points of the first edge of the foreground object. Thus, the first edge and the second edge are obtained according to the divided target pixel points. Wherein, the number threshold can be 4, that is, the boundary line formed by the pixel points with the number of connected pixel points less than 4 can be used as the pixel points of the first edge or the second edge.
[0104] In this embodiment, by determining the target pixel points through the number of connected pixel points of each pixel point in the depth map, and disconnecting the pixel points at the junction of the foreground object and the background object based on the depth value of the target pixel points, the first edge and the second edge are obtained, which can ensure the accuracy of the determined first edge and second edge, so as to facilitate obtaining accurate occlusion areas and foreground areas based on the first edge and the second edge in the subsequent process.
[0105] In an exemplary embodiment, in step S120, first flood processing is respectively performed on the first direction of the first edge and the second direction of the second edge, including: determining the direction in which the first edge tends to the center point of the foreground object as the first direction, and determining the direction in which the second edge is away from the center point of the foreground object as the second direction; based on the connectivity relationship between the pixel points in the depth map, first flood processing is respectively performed on the first direction of the first edge and the second direction of the second edge.
[0106] In specific implementation, before the flood processing, it is also necessary to determine the depth values of each pixel point in the depth map, and determine the connectivity relationship between each pixel point based on the depth values. Specifically, when the depth difference is less than the difference threshold, it is determined that two pixel points are connected, otherwise they are not connected. Further, flood processing is performed based on this connectivity relationship. Since the first edge is the edge of the foreground object, the flood direction of the first edge should be the direction tending to the center point of the foreground object, and the flooded area forms the foreground area of the foreground object. While the second edge is the edge of the background object, the flood direction of the second edge should be the direction away from the center point of the foreground object, and the flooded area is the known background area of the background object.
[0107] In this embodiment, based on the connectivity relationship between the pixel points in the depth map, first flood processing is respectively performed on the direction in which the first edge tends to the center point of the foreground object and the direction in which the second edge is away from the center point of the foreground object, so as to accurately determine the known background area and the occlusion area according to the flood result, and then perform image completion on the occlusion area according to the known background area.
[0108] In an exemplary embodiment, the image to be processed includes at least a first object, a second object, and a third object, where the first object is the foreground object of the second object, and the second object is the foreground object of the third object; in the above step S130, depth information completion and RGB information completion processing are performed on the occlusion area through the known background area to obtain a completed background image corresponding to the occlusion area, including two cases, where the first case is:
[0109] Step S130a, obtain the first edge of the first object and the second edge of the second object as the first pair of depth edges. Based on the first pair of depth edges, obtain the known background region of the second object and the occluded region occluded by the first object. Obtain the first edge of the second object and the second edge of the third object as the second pair of depth edges. Based on the second pair of depth edges, obtain the known background region of the third object and the occluded region occluded by the second object. If it is determined that the first object does not have edge occlusion on the second pair of depth edges, use the known background region of the second object to perform depth information completion and RGB information completion on the occluded region of the second object occluded by the first object, to obtain the completed background image corresponding to the second object. Use the known background region of the third object to perform depth information completion and RGB information completion on the occluded region of the third object occluded by the second object, to obtain the completed background image corresponding to the third object.
[0110] For example, referring to Figure 2 , take the human body as the first object, the sea surface as the second object, and the sky as the third object. Take the first edge of the human body and the second edge of the sea surface as the first pair of depth edges. Based on the first pair of depth edges, obtain the known background region 200a of the sea surface and the occluded region 200b of the sea surface occluded by the human body. Take the first edge of the sea surface and the second edge of the sky as the second pair of depth edges. Based on the second pair of depth edges, obtain the known background region 202a of the sky and the occluded region 202b of the sky occluded by the sea surface. When the human body does not occlude the second pair of depth edges of the sea surface and the sky, the depth information completion and RGB information completion can be performed on the occluded region 200b of the sea surface occluded by the human body according to the known background region 200a of the sea surface, to obtain the completed background image corresponding to the sea surface. Perform depth information completion and RGB information completion on the occluded region 202b of the sky occluded by the sea surface according to the known background region 202a of the sky, to obtain the completed background image corresponding to the sky.
[0111] In this embodiment, when there are at least three objects in the image to be processed and the first object does not have edge occlusion on the second pair of depth edges, the depth information completion and RGB information completion are respectively performed on the occluded regions of the second object through the known background region of the second object, and the depth information completion and RGB information completion are performed on the occluded regions of the third object through the known background region of the third object, to achieve directional completion of the occluded regions in sub-regions, so as to ensure clear and reasonable completion results.
[0112] In an exemplary embodiment, in the above step S130, the second case of performing depth information completion and RGB information completion processing on the occluded region through the known background region to obtain the completed background image corresponding to the occluded region is:
[0113] Step S130b: Obtain the first edge of the first object and the second edge of the third object as the third pair of depth edges. Based on the second pair of depth edges and the third pair of depth edges, obtain the known background area of the third object, the occluded area of the third object occluded by the first object, and the occluded area of the third object occluded by the second object. If it is determined that the first object occludes part of the second pair of depth edges, perform edge completion on the occluded edge in the second pair of depth edges to obtain the completed second pair of depth edges. Through the known background area of the second object, perform depth information completion and RGB information completion on the occluded area of the second object occluded by the first object to obtain the completed background image corresponding to the second object. Also, through the known background area of the third object, perform depth information completion and RGB information completion on the occluded area of the third object occluded by the first object to obtain the initial completed image corresponding to the third object. Through the initial completed image corresponding to the third object, perform depth information completion and RGB information completion on the occluded area of the third object occluded by the second object to obtain the completed background image corresponding to the third object.
[0114] In specific implementation, still taking Figure 2 as an example, when part of the first pair of depth edges formed by the sky and the sea surface is occluded, as shown in area 20 in the figure, in addition to the first pair of depth edges formed by the human body and the sea surface, and the second pair of depth edges formed by the sky and the sea surface, there is also a third pair of depth edges formed by the first edge of the human head area and the second edge of the sky. The occluded area of the sky includes not only the occluded area occluded by the sea surface but also the occluded area occluded by the human body. When performing image completion in this case, first perform edge completion on the edge of the second pair of depth edges formed by the sky and the sea surface occluded by the human head to obtain the completed second pair of depth edges. Then, through the known background area 200a of the sea surface, perform depth information completion and RGB information completion on the area of the sea surface occluded by the human body to obtain the completed background image of the sea surface. Also, through the known background area of the sky, perform depth information completion and RGB information completion on the occluded area of the sky occluded by the human body to obtain the initial completed image corresponding to the sky. Then, through the initial completed image corresponding to the sky, perform depth information completion and RGB information completion on the occluded area of the sky occluded by the sea surface to obtain the completed background image corresponding to the sky.
[0115] Further, in an exemplary embodiment, in the second scenario, the depth information and RGB information of the occluded area of the third object occluded by the first object are completed through the known background area of the third object to obtain the initial completed image corresponding to the third object, including: completing the depth information and RGB information of the occluded area of the third object occluded by the first object through the known background area of the third object to obtain the local completed image of the occluded area of the third object occluded by the first object; based on the connectivity relationship between the pixel points in the local completed image and the known background area of the third object, performing a second flooding process on the second direction of the second edge in the second pair of completed depth edges to obtain the initial completed image of the third object.
[0116] Specifically, the implementation process of completing the depth information and RGB information of the occluded area of the sky occluded by the human body through the known background area of the sky to obtain the initial completed image corresponding to the sky includes: first, completing the depth information and RGB information of the occluded area of the sky occluded by the human body through the known background area of the sky to obtain the local completed image of the occluded area of the sky occluded by the human body, and based on the connectivity relationship of the pixel points in the local completed image and the known background area of the sky, performing a second flooding process on the second direction of the second edge of the sky in the second pair of completed depth edges facing away from the sea surface, and taking the image corresponding to the area formed after flooding as the initial completed image of the sky.
[0117] In this embodiment, when there are at least three objects in the image to be processed and the first object occludes part of the second pair of depth edges, by performing edge completion, depth information completion, and RGB information completion in sequence, the problem of completing the occluded edge is solved, and the completed edge is subjected to secondary flooding, depth information, and RGB information completion to obtain multiple layers of completed results, ensuring the rationality of the spatial grid layering.
[0118] In an exemplary embodiment, the above method further includes: when it is detected that the second pair of depth edges is discontinuous, determining that the first object occludes the second pair of depth edges.
[0119] In this embodiment, by detecting the continuity of the second pair of depth edges, it is determined whether the first object occludes the second pair of depth edges according to the detection result, so that the corresponding completion method can be executed according to the detection result.
[0120] In an exemplary embodiment, in step S120, before performing the first flooding process on the first direction of the first edge and the second direction of the second edge in the depth map respectively, it further includes:
[0121] Step S10, determining the pixel points to be filtered from the pixel points of the background object in the depth map;
[0122] Step S11: Obtain a sampling window corresponding to the background object, and perform median filtering on the pixel points to be filtered based on the sampling window to obtain a filtered depth map.
[0123] Step S120: Perform first flood filling on the first direction of the first edge and the second direction of the second edge in the depth map, including: performing first flood filling on the first direction of the first edge and the second direction of the second edge in the filtered depth map.
[0124] Among them, median filtering is a non-linear smoothing technique used to set the gray value of each pixel point to the median of the gray values of all pixel points within a certain neighborhood window of this point.
[0125] Among them, the sampling window can be understood as the neighborhood window of the pixel points to be filtered. Different background objects can be set with different sizes of sampling windows. The sampling window can be set according to the characteristics of the background object. For example, for background objects with a larger range such as the sea surface and the sky, a larger sampling window can be set; for background objects with a smaller range such as trees and trash cans, a smaller sampling window can be set.
[0126] In specific implementation, in order to detect more obvious edges, before performing the first flood filling on the depth map, the depth map can also be filtered. Specifically, in this embodiment, the median filtering method is used to filter the depth map to effectively preserve the edge information in the depth map. However, since median filtering requires sorting inside each kernel, it consumes more time. Therefore, the present disclosure uses an improved sparse median filtering method to process the depth map.
[0127] More specifically, the protected area where the object that does not need to be filtered is located can be determined from the depth map first, and the protected area is masked. From outside the protected area, that is, from the pixel points of the background object, the pixel points to be filtered are determined. That is, the pixel points within the protected area are not filtered. Then, according to the characteristics of the background object, the corresponding sampling window is determined, and according to the size of the sampling window, the median filtering method is used to filter the pixel points to be filtered determined from the depth map, so as to realize the refinement processing of the depth edges in the depth map and obtain a depth map with a higher degree of fineness, that is, the filtered depth map. Further, the first flood filling is respectively performed on the first direction of the first edge and the second direction of the second edge in the filtered depth map with finer depth edges.
[0128] For example, referring to Figure 3 , it shows the visualization effect of the depth edges on the image and the schematic diagram of the depth map after being filtered 3 times. Figure 3In the first row of images, the visualization effect of the depth edges on the image is shown. In the second row of images, the depth maps corresponding to the first row of images are presented. From left to right, they are the depth maps after the first filtering, the depth maps after the second filtering, and the depth maps after the third filtering. Through the filtering process, the depth edges can be refined to a width of 2 pixels, with a higher level of fineness.
[0129] In this embodiment, before performing the first flooding process on the first direction of the first edge and the second flooding process on the second direction of the second edge respectively, filtering the depth map first can obtain a depth map with a higher level of fineness and more accurate edge information. And through the protected area masking process, this sparse median filtering method with masking can reduce the time consumption.
[0130] Further, in an exemplary embodiment, in the above step S10, to determine the pixel points to be filtered from the pixel points of the background objects in the depth map, it can be achieved in the following manner:
[0131] Step S10a: Determine candidate pixel points with depth values less than the depth threshold from the pixel points of the background objects in the depth map;
[0132] Step S10b: Determine the range where the candidate pixel points are located in the depth map and the gradient magnitude of the candidate pixel points; the gradient magnitude represents the depth difference between the candidate pixel points and the adjacent pixel points in the preset gradient direction;
[0133] Step S10c: Determine the pixel points to be filtered from the candidate pixel points according to the range where the candidate pixel points are located and the gradient magnitude of the candidate pixel points.
[0134] Among them, the preset gradient direction can be directions such as up, down, left, and right.
[0135] In the specific implementation, to reduce the time consumption, before determining the pixel points to be filtered, the pixel points of the background objects in the depth map can be preliminarily screened according to the depth values of each pixel point in the depth map, and the pixel points that are in the background objects and have depth values less than the depth threshold are selected as candidate pixel points. It can be understood that a depth value greater than the depth threshold indicates a greater distance from the camera. Therefore, filtering is not required, that is, the far - side area is not filtered, and only the near - side area with depth values less than the depth threshold is filtered to improve the filtering efficiency. Since the purpose of filtering is to obtain finer depth edges, after obtaining the candidate pixel points, the candidate pixel points can also be divided into candidate pixel points within the depth edge area and candidate pixel points not within the depth edge area based on the depth edge area. According to the range where the candidate pixel points are located and the gradient magnitude of the candidate pixel points in the preset gradient direction, it is determined whether the candidate pixel points are the pixel points to be filtered.
[0136] In this embodiment, by first screening out candidate pixel points with depth values less than the depth threshold from the depth map, and then determining the pixel points to be filtered from the candidate pixel points according to the range and gradient magnitude where the candidate pixel points are located, targeted filtering processing of the depth map is realized, and it is not necessary to perform filtering processing on all pixel points of the depth map, thereby improving the efficiency of filtering the depth map and reducing time consumption.
[0137] Furthermore, in an exemplary embodiment, a plurality of preset gradient directions are provided. In the above step S10c, according to the range where the candidate pixel points are located and the gradient magnitude of the candidate pixel points, determining the pixel points to be filtered from the candidate pixel points specifically includes: if the candidate pixel point is in the first range, when the gradient magnitude of the candidate pixel point in any gradient direction is greater than the first threshold, determining that the candidate pixel point is a pixel point to be filtered; the first range represents the range that is in the background object of the depth map and outside the second edge region of the background object; if the candidate pixel point is in the second range, when the gradient magnitude of the candidate pixel point in any gradient direction is greater than the second threshold, determining that the candidate pixel point is a pixel point to be filtered; wherein, the first threshold is greater than the second threshold, and the second range represents the range that is in the background object of the depth map and within the second edge of the background object.
[0138] It can be understood that since the first range is the range outside the second edge of the background object, it is uncertain whether there are other depth edges. Therefore, a relatively large first threshold can be used to determine the pixel points to be filtered to determine whether there are other depth edges and thus extract other depth edges for filtering. And the second range is the range within the second edge region of the background object, and it has been determined that there are depth edges. Therefore, a relatively small second threshold can be used to determine the pixel points to be filtered to improve the confidence level.
[0139] In specific implementation, 4 directions, namely up, down, left, and right, can be preset to calculate the gradient magnitude of each candidate pixel point. When the gradient magnitude of the candidate pixel point in any one gradient direction is greater than the threshold, it is determined that the candidate pixel point is a pixel point to be filtered. More specifically, the method for the pixel points to be filtered is different according to the different ranges where the candidate pixel points are located, and specifically includes the following situations: ① If the candidate pixel point is in the first range that is in the background object of the depth map and outside the second edge of the background object, when the gradient magnitude of the candidate pixel point in any gradient direction is greater than the first threshold, determining that the candidate pixel point is a pixel point to be filtered; ② If the candidate pixel point is in the second range that is in the background object of the depth map and within the second edge of the background object, when the gradient magnitude of the candidate pixel point in any gradient direction is greater than the second threshold, determining that the candidate pixel point is a pixel point to be filtered. After determining the pixel points to be filtered, mark the area where the pixel points to be filtered are located as the edge area that needs to be filtered, and perform median filtering.
[0140] In this embodiment, by adaptively selecting the filtering region and the protection region, and by setting corresponding thresholds for different ranges, the continuity and accuracy of the depth edge are improved through this dual-threshold method.
[0141] In an exemplary embodiment, after obtaining the completed background image corresponding to the occluded region in step S130, the method further includes: obtaining the depth information and RGB information of each layer of the image to be processed according to the known background image and the completed background image; obtaining the reconstructed three-dimensional mesh information corresponding to the image to be processed according to the depth information and RGB information of each layer; and constructing a new perspective image of the image to be processed based on the set new perspective and the three-dimensional mesh information.
[0142] Wherein, the new perspective represents a perspective different from the original shooting perspective of the image to be processed.
[0143] In specific implementation, after obtaining the completed background images corresponding to the respective occluded regions, since the completed background images corresponding to the second edges of each background object are at different depth distances in space, it is necessary to add the depth information and RGB information of the completed background images into the graph structure of the image to be processed to form a graph structure of the depth and RGB information of multiple layers. According to the connectivity relationship of the pixel points in the graph structure, the pixel points of the same layer are connected to generate point cloud, patch, and texture coordinate information, etc., for three-dimensional mesh reconstruction to obtain the complete three-dimensional mesh information of the scene of the image to be processed. Further, the three-dimensional mesh of the scene of the image to be processed can be rendered according to the set new perspective to obtain a new perspective image corresponding to the set new perspective. Alternatively, according to a preset camera trajectory, the three-dimensional mesh of the scene of the image to be processed can be rendered to obtain a series of new perspective images with large amplitudes, and further, the new perspective images of each perspective can be synthesized into a video, thereby obtaining a 3D image camera movement effect with a sense of space. In addition, after obtaining the new perspective image, face depth reconstruction, depth fusion, and adding spatial 3D particle special effects and other processes can be further performed.
[0144] For example, referring to Figure 4 , which is a schematic diagram of new perspective images of different perspectives, Figure 4 the first picture in is the original picture, and the second, third, and fourth pictures are the lower left view, upper right view, and middle view generated based on the original picture respectively.
[0145] In this embodiment, after obtaining a clear and reasonable completed background image corresponding to the occluded region, the depth information and RGB information of each layer are obtained according to the completed background image, and then the complete three-dimensional mesh information of the image to be processed is obtained, so that the construction of new perspective images with large amplitudes can be realized, overcoming the defect that the traditional method can only generate new perspectives through small-amplitude pose transformations.
[0146] The present disclosure finds occluded regions through deep edge extraction and flood-filling algorithms, and uses a directional image completion method to complete the edge, depth information, and RGB information of the occluded regions, obtaining a completed 3D scene mesh and rendering a new perspective image. In addition, this method can also enhance the three-dimensional sense of the human body by combining portrait segmentation and face reconstruction to achieve a more realistic 3D effect.
[0147] In an exemplary embodiment, for the convenience of those skilled in the art to understand the embodiments of the present disclosure, the following will be described with specific examples of the accompanying drawings. Referring to Figure 5 , which is a schematic diagram of the complete process of an image processing method in an application example. The main process can be divided into three parts: depth estimation model optimization, depth map post-processing technology, and image completion algorithm optimization. The overall process is as follows:
[0148] Step S510, predicting the depth map corresponding to the image to be processed through a depth estimation network model.
[0149] Among them, the depth estimation network uses a deep convolutional network as a feature extractor. The network performs multiple pooling operations to reduce the resolution, and then predicts a high-resolution depth map through modules such as deconvolution layers, upsampling layers, multi-scale network structures, and skip connections. It achieves a high degree of accuracy in terms of data optimization, network structure improvement, and loss function design, and uses auxiliary branches such as prediction segmentation and offset for optimization to obtain a depth map with higher prediction accuracy.
[0150] Step S520, performing sparse median filtering on the depth map to obtain obvious first edges and second edges. The specific steps are as follows:
[0151] (1) According to a preset depth threshold, the pixel points in the depth map are divided into pixel points with depth values less than the depth threshold and pixel points with depth values greater than the depth threshold. Among them, the pixel points with depth values greater than the depth threshold are not filtered, that is, the pixel points on the far side are not filtered and edge extracted.
[0152] (2) For the pixel points with depth values less than the depth threshold, calculate the gradient amplitude in the four directions of up, down, left, and right.
[0153] (3) According to the second edge mask corresponding to the input second edge region and the protection mask corresponding to the protection region (i.e., the region where the foreground object is located), for the region with depth values less than the depth threshold, perform a double-threshold edge extraction algorithm:
[0154] a) Inside the protection mask, no edge extraction and filtering are performed;
[0155] b) Outside the protection mask and outside the second edge mask, when the gradient amplitude in any direction is greater than the first threshold th, mark the corresponding region as the edge region to be filtered, otherwise, mark it as a non-edge region;
[0156] c) In the second edge mask outside the protection mask, when the gradient magnitude in any direction is greater than the second threshold th2, the corresponding area is marked as the edge area to be filtered, otherwise, it is marked as a non-edge area. Among them, th > th2.
[0157] (4) For the edge areas marked as to be filtered inside the kernel, median filtering is performed.
[0158] Through the above steps, the depth edge can be refined to a width of 2 pixels, with a high degree of fineness. As Figure 4 shown, the results of the image and depth map after 3 times of filtering are presented, and the edges are gradually refined.
[0159] Step S530: Perform flooding on the first direction of the first edge and the second direction of the second edge respectively to obtain an occlusion area mask and a known background area mask. The specific steps include:
[0160] (1) According to the coordinate position and depth of the pixel point, a 3D node is formed and added to the graph structure.
[0161] (2) According to the preset difference threshold, determine whether each node is connected. Specifically, if the depth difference between two nodes is less than the difference threshold, it is determined to be connected, otherwise it is not.
[0162] (3) Through the connected component analysis results, perform the first flooding process on the first direction of the first edge and the second direction of the second edge respectively to obtain the foreground area formed after flooding the first edge and the known background area formed after flooding the second edge. Based on the foreground area, determine the occlusion area, and perform a consistency check on the flooding results.
[0163] Step S540: Perform edge completion, depth completion, and RGB completion on the occlusion area successively through the known background area. Specifically, it includes:
[0164] (1) In the case of no edge occlusion, perform directional completion on each occlusion area separately. Each completion is based on the known background area and the occlusion area for bbox extraction. Within the bbox area, directional diffusion is performed from the known background area to the occlusion area to obtain a clear and reasonable completion result.
[0165] (2) For the case of edge occlusion (two edges intersect), first complete the occluded edge. After obtaining the completed edge, record the proximal area and the distal area of the completion for subsequent second flooding and the completion of depth information and RGB information. After performing depth information completion and RGB information completion, fuse the completion results into the known background image.
[0166] Through the above steps, a reasonable and clear completed RGB image can be obtained. For the completed background images corresponding to each edge, they are at different depth distances in space. Therefore, it is necessary to add the completion results to the graph structure to form a graph structure of multiple layers of depth and RGB images, and connect the nodes of the same layer according to the connection relationship.
[0167] Reference Figure 6 , which is a schematic diagram of the directional completion result. From left to right in the figure are the texture mask image, the completion mask image, the texture RGB image, the completed background image, and the completed depth image.
[0168] Step S550: Perform three-dimensional network information reconstruction based on the completed background image and the image to be processed.
[0169] Step S560: Render the three-dimensional mesh according to the camera trajectory preset by the user, and then the corresponding camera movement effect can be obtained, and a new perspective image can be obtained.
[0170] The image processing method proposed in this embodiment has the following beneficial effects:
[0171] (1) Improve the fine depth edges through masked sparse median filtering, adaptively select the area to be filtered and the protected area, and at the same time improve the continuity and accuracy of the depth edges through double thresholds.
[0172] (2) Extract depth edges through connected component analysis on the graph structure, and perform consistency checks to maintain the continuity and accuracy of the edges, and obtain the occluded area and the known background area through flooding to ensure the rationality of the occluded area and the known background area.
[0173] (3) Through regional directional completion, perform directional diffusion to the occluded completion area to ensure a clear and reasonable completion result.
[0174] (4) Solve the completion problem of occluded edges through edge completion, depth completion, and RGB image completion, and perform secondary flooding, depth, and RGB image completion on the completed edges to obtain multiple-layer completion results, ensuring the rationality of the spatial mesh layering.
[0175] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless specifically stated herein, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0176] It can be understood that the same / similar parts among the various embodiments of the above methods in this specification can be referred to each other. Each embodiment focuses on the differences from other embodiments. For the relevant parts, refer to the descriptions of other method embodiments.
[0177] Based on the same inventive concept, the embodiments of the present disclosure also provide an image processing apparatus for implementing the above-mentioned image processing method.
[0178] Figure 7 is a structural block diagram of an image processing apparatus shown according to an exemplary embodiment. Referring to Figure 7 , the apparatus includes: a determination unit 710, a flooding unit 720, and a completion unit 730, where
[0179] The determination unit 710 is configured to execute obtaining a depth map corresponding to the image to be processed, and determining a first edge of the foreground object and a second edge of the background object from the depth map;
[0180] The flooding unit 720 is configured to execute first flooding processing on the first direction of the first edge and the second direction of the second edge in the depth map respectively, to obtain a foreground area formed after flooding the first edge and a known background area formed after flooding the second edge, and determining an occlusion area based on the foreground area; where the occlusion area represents the background area occluded by the foreground area; the first direction is opposite to the second direction;
[0181] The completion unit 730 is configured to execute depth information completion and RGB information completion processing on the occlusion area through the known background area in the depth map, to obtain a completed background image corresponding to the occlusion area.
[0182] In an exemplary embodiment, the determining unit 710 is further configured to obtain the connectivity relationship between each pixel point in the depth map, and determine the number of connected pixel points of each pixel point based on the connectivity relationship; wherein, the depth difference between two connected pixel points is within a preset range; determine a target pixel point from the depth map; the number of connected pixel points corresponding to the target pixel point is less than a number threshold; based on the depth value of the target pixel point, divide the target pixel point into a pixel point of the first edge or a pixel point of the second edge, and obtain the first edge and the second edge based on the divided target pixel points.
[0183] In an exemplary embodiment, the flooding unit 720 is further configured to determine that the direction of the first edge towards the center point of the foreground object is the first direction, and determine that the direction of the second edge away from the center point of the foreground object is the second direction; based on the connectivity relationship between each pixel point in the depth map, perform a first flooding process on the first direction of the first edge and the second direction of the second edge respectively.
[0184] In an exemplary embodiment, the image to be processed at least includes a first object, a second object, and a third object, wherein the first object is the foreground object of the second object, and the second object is the foreground object of the third object; the complementing unit 730 is further configured to obtain the first edge of the first object and the second edge of the second object as the first pair of depth edges, and obtain the known background area of the second object and the occluded area occluded by the first object based on the first pair of depth edges; obtain the first edge of the second object and the second edge of the third object as the second pair of depth edges, and obtain the known background area of the third object and the occluded area occluded by the second object based on the second pair of depth edges; if it is determined that the first object does not have an edge occlusion on the second pair of depth edges, perform depth information complementation and RGB information complementation on the occluded area of the second object occluded by the first object through the known background area of the second object to obtain a complemented background image corresponding to the second object; perform depth information complementation and RGB information complementation on the occluded area of the third object occluded by the second object through the known background area of the third object to obtain a complemented background image corresponding to the third object.
[0185] In an exemplary embodiment, the completion unit 730 is further configured to execute obtaining a first edge of a first object and a second edge of a third object as a third pair of depth edges, and obtaining a known background region of the third object, an occluded region of the third object occluded by the first object, and an occluded region of the third object occluded by the second object based on the second pair of depth edges and the third pair of depth edges; if it is determined that the first object occludes a part of the second pair of depth edges, performing edge completion on the occluded edge in the second pair of depth edges to obtain a completed second pair of depth edges; performing depth information completion and RGB information completion on the occluded region of the second object occluded by the first object through the known background region of the second object to obtain a completed background image corresponding to the second object; and performing depth information completion and RGB information completion on the occluded region of the third object occluded by the first object through the known background region of the third object to obtain an initial completed image corresponding to the third object; performing depth information completion and RGB information completion on the occluded region of the third object occluded by the second object through the initial completed image corresponding to the third object to obtain a completed background image corresponding to the third object.
[0186] In an exemplary embodiment, the completion unit 730 is further configured to execute performing depth information completion and RGB information completion on the occluded region of the third object occluded by the first object through the known background region of the third object to obtain a partial completed image of the occluded region of the third object occluded by the first object; performing a second flooding process on the second direction of the second edge in the completed second pair of depth edges based on the connectivity relationship between the pixel points in the partial completed image and the pixel points in the known background region of the third object to obtain an initial completed image of the third object.
[0187] In an exemplary embodiment, the apparatus further includes a detection unit configured to execute determining that the first object occludes the second pair of depth edges when it is detected that the second pair of depth edges is discontinuous.
[0188] In an exemplary embodiment, the apparatus further includes a filtering unit configured to execute determining to-be-filtered pixel points from the pixel points of the background object in the depth map; obtaining a sampling window corresponding to the background object, and performing median filtering on the to-be-filtered pixel points based on the sampling window to obtain a filtered depth map.
[0189] The flooding unit 720 is further configured to execute performing a first flooding process on the first direction of the first edge and the second direction of the second edge respectively in the filtered depth map.
[0190] In an exemplary embodiment, the filtering unit is further configured to determine candidate pixel points with depth values less than a depth threshold from the pixel points of the background object in the depth map; determine the range where the candidate pixel points are located in the depth map and the gradient magnitude of the candidate pixel points; the gradient magnitude represents the depth difference between the candidate pixel points and adjacent pixel points in a preset gradient direction; and determine the pixel points to be filtered from the candidate pixel points according to the range where the candidate pixel points are located and the gradient magnitude of the candidate pixel points.
[0191] In an exemplary embodiment, there are multiple preset gradient directions, and the filtering unit is further configured to perform: if the candidate pixel points are in a first range, when the gradient magnitude of the candidate pixel points in any gradient direction is greater than a first threshold, determine that the candidate pixel points are the pixel points to be filtered; the first range represents the range that is in the background object of the depth map and outside the second edge of the background object; if the candidate pixel points are in a second range, when the gradient magnitude of the candidate pixel points in any gradient direction is greater than a second threshold, determine that the candidate pixel points are the pixel points to be filtered; wherein, the first threshold is greater than the second threshold, and the second range represents the range that is in the background object of the depth map and inside the second edge of the background object.
[0192] In an exemplary embodiment, the apparatus further includes a new perspective image construction unit, which is configured to perform: obtain the depth information and RGB information of each layer of the image to be processed according to the known background image and the complemented background image; obtain the reconstructed three-dimensional mesh information corresponding to the image to be processed according to the depth information and RGB information of each layer; and construct a new perspective image of the image to be processed based on the set new perspective and the three-dimensional mesh information.
[0193] Regarding the apparatus in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0194] Figure 8 FIG. is a block diagram of an electronic device 800 for implementing an image processing method according to an exemplary embodiment. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0195] Referring to Figure 8 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0196] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0197] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, optical disks, or graphene memory.
[0198] The power component 806 provides power to various components of the electronic device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0199] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may not only sense the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0200] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0201] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which may be a keyboard, a click wheel, buttons, etc. These buttons may include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0202] The sensor component 814 includes one or more sensors for providing status assessments of various aspects of the electronic device 800. For example, the sensor component 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor component 814 can also detect changes in the position of the electronic device 800 or components of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the device 800, and changes in the temperature of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0203] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0204] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0205] In an exemplary embodiment, there is also provided a computer-readable storage medium including instructions, such as a memory 804 including instructions, and the above instructions can be executed by a processor 820 of the electronic device 800 to complete the above method. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0206] In an exemplary embodiment, there is also provided a computer program product, and the computer program product includes instructions, and the above instructions can be executed by a processor 820 of the electronic device 800 to complete the above method.
[0207] It should be noted that the above-mentioned device, electronic device, computer-readable storage medium, computer program product, etc. may also include other implementation manners according to the description of the method embodiments. The specific implementation manners may refer to the description of the relevant method embodiments and will not be elaborated herein one by one.
[0208] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only to be regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0209] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An image processing method, characterized in that, Including: Obtain a depth map corresponding to the image to be processed, and determine a first edge of the foreground object and a second edge of the background object from the depth map; Perform a first flooding process on the first direction of the first edge and the second direction of the second edge in the depth map respectively to obtain a foreground area formed after flooding the first edge and a known background area formed after flooding the second edge, and determine an occlusion area based on the foreground area; wherein, the occlusion area represents the background area occluded by the foreground area; the first direction is opposite to the second direction; Perform depth information completion and RGB information completion processing on the occlusion area through the known background area in the depth map to obtain a completed background image corresponding to the occlusion area; When the image to be processed includes at least a first object, a second object, and a third object, if the first object does not occlude the depth edge formed by the second object and the third object, perform depth information completion and RGB information completion processing on the respective occlusion areas of the second object and the third object through their respective known background areas; If the first object occludes the depth edge formed by the second object and the third object, first complete the occluded depth edge, and then perform depth information completion and RGB information completion processing on the respective occlusion areas of the second object and the third object according to the completed edge and their respective known background areas; Wherein, the first object is the foreground object of the second object, and the second object is the foreground object of the third object.
2. The method according to claim 1, characterized in that, The determining a first edge of the foreground object and a second edge of the background object from the depth map includes: Obtain the connectivity relationship between each pixel point in the depth map, and determine the number of connected pixel points of each pixel point based on the connectivity relationship; wherein, the depth difference between two connected pixel points is within a preset range; Determine a target pixel point from the depth map; the number of connected pixel points corresponding to the target pixel point is less than a number threshold; Based on the depth value of the target pixel point, divide the target pixel point into a pixel point of the first edge or a pixel point of the second edge, and obtain the first edge and the second edge based on the divided target pixel points.
3. The method according to claim 2, wherein The performing a first flooding process on the first direction of the first edge and the second direction of the second edge respectively includes: Determine the direction towards the center point of the foreground object of the first edge as the first direction, and determine the direction away from the center point of the foreground object of the second edge as the second direction; Based on the connectivity relationship between each pixel point in the depth map, perform a first flooding process on the first direction of the first edge and the second direction of the second edge respectively.
4. The method according to claim 1, wherein The if the first object does not occlude the depth edge formed by the second object and the third object, perform depth information completion and RGB information completion processing on the respective occlusion areas of the second object and the third object through their respective known background areas includes: Obtain the first edge of the first object and the second edge of the second object as the first pair of depth edges, and based on the first pair of depth edges, obtain the known background region of the second object and the occluded region occluded by the first object; Obtain the first edge of the second object and the second edge of the third object as the second pair of depth edges, and based on the second pair of depth edges, obtain the known background region of the third object and the occluded region occluded by the second object; If it is determined that the first object does not have edge occlusion on the second pair of depth edges, use the known background region of the second object to perform depth information completion and RGB information completion on the occluded region of the second object occluded by the first object to obtain the completed background image corresponding to the second object; Use the known background region of the third object to perform depth information completion and RGB information completion on the occluded region of the third object occluded by the second object to obtain the completed background image corresponding to the third object.
5. The method according to claim 4, wherein If the first object occludes the depth edge formed by the second object and the third object, first complete the occluded depth edge, and then perform depth information completion and RGB information completion processing on their respective occluded regions according to the completed edge and the known background regions of the second object and the third object respectively, including: Obtain the first edge of the first object and the second edge of the third object as the third pair of depth edges, and based on the second pair of depth edges and the third pair of depth edges, obtain the known background region of the third object, the occluded region of the third object occluded by the first object, and the occluded region of the third object occluded by the second object; If it is determined that the first object occludes part of the second pair of depth edges, perform edge completion on the occluded edge in the second pair of depth edges to obtain the completed second pair of depth edges; Use the known background region of the second object to perform depth information completion and RGB information completion on the occluded region of the second object occluded by the first object to obtain the completed background image corresponding to the second object; and use the known background region of the third object to perform depth information completion and RGB information completion on the occluded region of the third object occluded by the first object to obtain the initial completed image corresponding to the third object; Use the initial completed image corresponding to the third object to perform depth information completion and RGB information completion on the occluded region of the third object occluded by the second object to obtain the completed background image corresponding to the third object.
6. The method according to claim 5, wherein The step of using the known background region of the third object to perform depth information completion and RGB information completion on the occluded region of the third object occluded by the first object to obtain the initial completed image corresponding to the third object includes: Use the known background region of the third object to perform depth information completion and RGB information completion on the occluded region of the third object occluded by the first object to obtain the partial completed image of the occluded region of the third object occluded by the first object; Based on the connectivity relationship between the pixel points in the locally completed image and the known background region of the third object, perform a second flooding process on the second direction of the second edge in the completed second pair of depth edges to obtain the initial completed image of the third object.
7. The method according to any one of claims 4 to 6, characterized in that The method further includes: When it is detected that the second pair of depth edges is discontinuous, it is determined that the first object occludes the second pair of depth edges.
8. The method according to claim 1, characterized in that Before performing the first flooding process on the first direction of the first edge and the second direction of the second edge in the depth map respectively, it further includes: Determine the pixels to be filtered from the pixel points of the background object in the depth map; Obtain a sampling window corresponding to the background object, and perform median filtering on the pixels to be filtered based on the sampling window to obtain a filtered depth map; The performing the first flooding process on the first direction of the first edge and the second direction of the second edge in the depth map respectively includes: In the filtered depth map, perform the first flooding process on the first direction of the first edge and the second direction of the second edge respectively.
9. The method according to claim 8, wherein The determining the pixels to be filtered from the pixel points of the background object in the depth map includes: Determine candidate pixels with a depth value less than a depth threshold from the pixel points of the background object in the depth map; Determine the range where the candidate pixels are located in the depth map and the gradient magnitude of the candidate pixels; the gradient magnitude represents the depth difference between the candidate pixels and adjacent pixels in a preset gradient direction; Determine the pixels to be filtered from the candidate pixels according to the range where the candidate pixels are located and the gradient magnitude of the candidate pixels.
10. The method according to claim 9, wherein There are multiple preset gradient directions, and the determining the pixels to be filtered from the candidate pixels according to the range where the candidate pixels are located and the gradient magnitude of the candidate pixels includes: If the candidate pixel is in the first range, when the gradient magnitude of the candidate pixel in any gradient direction is greater than the first threshold, determine that the candidate pixel is a pixel to be filtered; the first range represents the range that is in the background object of the depth map and outside the second edge of the background object; If the candidate pixel is in the second range, when the gradient magnitude of the candidate pixel in any gradient direction is greater than the second threshold, determine that the candidate pixel is a pixel to be filtered; wherein, the first threshold is greater than the second threshold, and the second range represents the range that is in the background object of the depth map and inside the second edge of the background object.
11. The method according to claim 1, wherein After obtaining the completed background image corresponding to the occluded area, it further includes: Obtain the depth information and RGB information of each layer of the image to be processed according to the known background region and the completed background image; Obtain the reconstructed three-dimensional mesh information corresponding to the image to be processed according to the depth information and RGB information of each layer; Based on the set new view angle and the three-dimensional mesh information, construct a new view angle image of the image to be processed.
12. An image processing apparatus, characterized in that, Includes: A determination unit, configured to execute obtaining a depth map corresponding to an image to be processed, and determining a first edge of a foreground object and a second edge of a background object from the depth map; A flooding unit, configured to execute performing a first flooding process on a first direction of the first edge and a second direction of the second edge in the depth map respectively, to obtain a foreground region formed after flooding the first edge and a known background region formed after flooding the second edge, and determining an occlusion region based on the foreground region; wherein, the occlusion region represents a background region occluded by the foreground region; the first direction is opposite to the second direction; A completion unit, configured to execute performing depth information completion and RGB information completion processing on the occlusion region through the known background region in the depth map, to obtain a completed background image corresponding to the occlusion region; The completion unit is further configured to execute when the image to be processed includes at least a first object, a second object, and a third object, if the first object does not occlude a depth edge formed by the second object and the third object, performing depth information completion and RGB information completion processing on respective occlusion regions of the second object and the third object through respective known background regions of the second object and the third object; if the first object occludes the depth edge formed by the second object and the third object, first completing the occluded depth edge, and then performing depth information completion and RGB information completion processing on respective occlusion regions of the second object and the third object according to the completed edge and respective known background regions of the second object and the third object; wherein, the first object is a foreground object of the second object, and the second object is a foreground object of the third object.
13. The device according to claim 12, characterized in that, The determination unit is further configured to execute obtaining a connectivity relationship between pixel points in the depth map, and determining the number of connected pixel points of each pixel point based on the connectivity relationship; wherein, the depth difference between two connected pixel points is within a preset range; determining a target pixel point from the depth map; the number of connected pixel points corresponding to the target pixel point is less than a number threshold; based on the depth value of the target pixel point, classifying the target pixel point as a pixel point of the first edge or a pixel point of the second edge, and obtaining the first edge and the second edge based on the classified target pixel points.
14. The device according to claim 13, characterized in that, The flooding unit is further configured to execute determining a direction of the first edge towards the center point of the foreground object as the first direction, and determining a direction of the second edge away from the center point of the foreground object as the second direction; based on the connectivity relationship between pixel points in the depth map, performing a first flooding process on the first direction of the first edge and the second direction of the second edge respectively.
15. The device according to claim 12, wherein The completion unit is further configured to execute obtaining a first edge of the first object and a second edge of the second object, as a first pair of depth edges, and obtaining a known background region of the second object and an occlusion region occluded by the first object based on the first pair of depth edges; Obtain the first edge of the second object and the second edge of the third object as the second pair of depth edges. Based on the second pair of depth edges, obtain the known background region of the third object and the occluded region occluded by the second object; if it is determined that the first object does not have edge occlusion on the second pair of depth edges, complete the depth information and RGB information of the occluded region of the second object occluded by the first object through the known background region of the second object, and obtain the complementary background image corresponding to the second object. Complete the depth information and RGB information of the occluded region of the third object occluded by the second object through the known background region of the third object, and obtain the complementary background image corresponding to the third object.
16. The device according to claim 15, wherein, The complementing unit is further configured to execute obtaining the first edge of the first object and the second edge of the third object as the third pair of depth edges, and based on the second pair of depth edges and the third pair of depth edges, obtain the known background region of the third object, the occluded region of the third object occluded by the first object, and the occluded region of the third object occluded by the second object. If it is determined that the first object occludes a part of the second pair of depth edges, complete the occluded edge in the second pair of depth edges to obtain the completed second pair of depth edges. Complete the depth information and RGB information of the occluded region of the second object occluded by the first object through the known background region of the second object, and obtain the complementary background image corresponding to the second object. And, complete the depth information and RGB information of the occluded region of the third object occluded by the first object through the known background region of the third object, and obtain the initial complementary image corresponding to the third object. Complete the depth information and RGB information of the occluded region of the third object occluded by the second object through the initial complementary image corresponding to the third object, and obtain the complementary background image corresponding to the third object.
17. The device according to claim 16, wherein The complementing unit is further configured to execute completing the depth information and RGB information of the occluded region of the third object occluded by the first object through the known background region of the third object, and obtain the partial complementary image of the occluded region of the third object occluded by the first object; based on the connectivity relationship between the pixel points in the partial complementary image and the known background region of the third object, perform a second flood fill on the second direction of the second edge in the completed second pair of depth edges to obtain the initial complementary image of the third object.
18. The device according to any one of claims 15-17, characterized in that The device further includes a detection unit configured to execute determining that the first object has occlusion on the second pair of depth edges when it is detected that the second pair of depth edges is discontinuous.
19. The device according to claim 12, characterized in that, The device further includes a filtering unit configured to determine to-be-filtered pixel points from the pixel points of the background object in the depth map; obtain a sampling window corresponding to the background object, and perform median filtering on the to-be-filtered pixel points based on the sampling window to obtain a filtered depth map; The flooding unit is further configured to perform a first flooding process on the first direction of the first edge and the second direction of the second edge respectively in the filtered depth map.
20. The device according to claim 19, characterized in that, The filtering unit is further configured to determine candidate pixel points with depth values less than a depth threshold from the pixel points of the background object in the depth map; determine the range where the candidate pixel points are located in the depth map and the gradient amplitude of the candidate pixel points; the gradient amplitude represents the depth difference between the candidate pixel points and adjacent pixel points in a preset gradient direction; and determine to-be-filtered pixel points from the candidate pixel points according to the range where the candidate pixel points are located and the gradient amplitude of the candidate pixel points.
21. The device according to claim 20, characterized in that, There are multiple preset gradient directions, and the filtering unit is further configured to perform: if the candidate pixel points are in a first range, when the gradient amplitude of the candidate pixel points in any gradient direction is greater than a first threshold, determine that the candidate pixel points are to-be-filtered pixel points; the first range represents the range that is in the background object of the depth map and outside the second edge of the background object; if the candidate pixel points are in a second range, when the gradient amplitude of the candidate pixel points in any gradient direction is greater than a second threshold, determine that the candidate pixel points are to-be-filtered pixel points; wherein, the first threshold is greater than the second threshold, and the second range represents the range that is in the background object of the depth map and inside the second edge of the background object.
22. The device according to claim 12, characterized in that, The device further includes a new perspective image construction unit configured to obtain the depth information and RGB information of each layer of the to-be-processed image according to the known background area and the complemented background image; Obtain the reconstructed three-dimensional mesh information corresponding to the to-be-processed image according to the depth information and RGB information of each layer; Construct a new perspective image of the to-be-processed image based on a set new perspective and the three-dimensional mesh information.
23. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the image processing method according to any one of claims 1 to 11.
24. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the image processing method according to any one of claims 1 to 11.
25. A computer program product, comprising instructions therein, characterized in that, When the instructions are executed by the processor of the electronic device, the electronic device is enabled to execute the image processing method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Image processing apparatus and method of generating a multi-view image
US20120114225A1
Technologies for improving the accuracy of depth cameras
US20160065930A1