Shielding area determination method and device and electronic equipment

By using motion vector and morphological dilation processing in the game temporal super-resolution algorithm, the occlusion area is accurately determined, solving the problem of inaccurate occlusion judgment and improving image quality.

CN121685272APending Publication Date: 2026-03-17VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511916507.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

In existing game temporal super-resolution algorithms, the occlusion region is not accurately identified, leading to ghosting and edge fragmentation problems during image processing.

Method used

The depth of the (N-1)th image frame is mapped to the Nth image frame using motion vectors to determine the first occlusion region and the background region. The second occlusion region is obtained through morphological dilation. Finally, the two regions are merged to obtain the target occlusion region, ensuring that the edge shape of the foreground region is clear.

Benefits of technology

It improves the accuracy of occlusion area detection, avoids ghosting and jagged edges, and optimizes image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685272A_ABST
    Figure CN121685272A_ABST
Patent Text Reader

Abstract

The invention discloses a shielding area determination method and device and electronic equipment. The method comprises the following steps: mapping the depth of an (N-1) th image frame into an Nth image frame according to a motion vector to obtain a first mapping depth; determining a first occlusion area and a first background area according to the first mapping depth and the depth of the Nth image frame; performing morphological expansion processing on the first shielding area to obtain a second shielding area; and carrying out merging processing on the second shielding region and the first background region, and determining a target shielding region in the Nth image frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of communication technology, specifically relating to a method, apparatus, and electronic device for determining obstructed areas. Background Technology

[0002] Existing game temporal super-resolution algorithms share a common design philosophy: they primarily improve the image quality of the current frame by fusing historical frames with the current frame and utilizing information from the historical frames. Specifically, historical frames need to be aligned to the current frame using motion vectors to avoid ghosting caused by the relative motion between the two frames. During this alignment process, classic object occlusion or leakage issues arise. Current mainstream methods typically fill in and repair occluded or leaked objects, but existing open-source algorithms are not accurate in identifying occluded regions. Summary of the Invention

[0003] The purpose of this application is to provide a method, apparatus, and electronic device for determining occlusion areas, in order to solve the problem of inaccurate occlusion area determination in existing image processing.

[0004] In a first aspect, embodiments of this application provide a method for determining an occlusion area, including:

[0005] Based on the motion vector, the depth of the (N-1)th image frame is mapped to the Nth image frame to obtain the first mapped depth;

[0006] Based on the first mapping depth and the depth of the Nth image frame, the first occlusion region and the first background region are determined;

[0007] The first occlusion region is morphologically expanded to obtain the second occlusion region.

[0008] The second occlusion region is merged with the background region to determine the target occlusion region in the Nth image frame.

[0009] Secondly, embodiments of this application provide an occlusion area determination device, including:

[0010] The first mapping module is used to map the depth of the (N-1)th image frame to the Nth image frame according to the motion vector, so as to obtain the first mapping depth;

[0011] The first processing module is used to determine the first occlusion region and the first background region based on the first mapping depth and the depth of the Nth image frame.

[0012] The second processing module is used to perform morphological dilation processing on the first occlusion area to obtain the second occlusion area;

[0013] The third processing module is used to merge the second occlusion area with the background area to determine the target occlusion area in the Nth image frame.

[0014] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0015] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0016] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0017] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0018] In the embodiments of this application, the depth of the (N-1)th image frame is mapped to the Nth image frame according to the motion vector to obtain a first mapped depth; a first occlusion region and a first background region are determined according to the first mapped depth and the depth of the Nth image frame; by performing morphological dilation processing on the obtained first occlusion region, the edge region of the foreground region in the obtained second occlusion region is clearer, thus the displayed foreground region is more accurate and complete; by merging the second occlusion region with the obtained first background region, the contours and regions of the foreground and background regions in the obtained image are clearer and more complete, thus the position of the occlusion region can be accurately obtained. This scheme can be applied to the occlusion relationship processing of temporal super-resolution algorithms, which can protect the edge region morphology of foreground objects in the image. In the processing, it retains sufficient image information from temporal fusion and avoids artifact problems caused by incorrect occlusion judgment, which can greatly optimize image quality. Attached Figure Description

[0019] Figure 1 This is one of the flowcharts illustrating the method for determining the occlusion area according to an embodiment of this application;

[0020] Figure 2 This is a schematic diagram of the depth-corresponding color of the Nth image frame in the embodiments of this application;

[0021] Figure 3This is a color diagram corresponding to the mapping depth in the embodiments of this application;

[0022] Figure 4 This is a schematic diagram of the first occlusion area according to an embodiment of this application;

[0023] Figure 5 This is a schematic diagram of the boundary of the expansion process in an embodiment of this application;

[0024] Figure 6 This is one of the schematic diagrams after the expansion process according to an embodiment of this application;

[0025] Figure 7 This is the second schematic diagram after the expansion process in an embodiment of this application;

[0026] Figure 8 This is a schematic diagram of the background area in an embodiment of this application;

[0027] Figure 9 This is a schematic diagram of the target occlusion area according to an embodiment of this application;

[0028] Figure 10 This is a schematic diagram of the occlusion area detected by the existing FSR algorithm;

[0029] Figure 11 This is a second schematic flowchart of the method for determining the occlusion area according to an embodiment of this application;

[0030] Figure 12 This is a schematic diagram of the structure of the occlusion area determination device according to an embodiment of this application;

[0031] Figure 13 This is one of the structural schematic diagrams of the electronic device according to an embodiment of this application;

[0032] Figure 14 This is a second schematic diagram of the structure of the electronic device according to an embodiment of this application. Detailed Implementation

[0033] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0034] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0035] The method for determining the occlusion area provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0036] like Figure 1 As shown in the figure, this application provides a method for determining an occlusion area, including:

[0037] Step 11: Based on the motion vector, map the depth of the (N-1)th image frame to the Nth image frame to be processed to obtain the first mapped depth;

[0038] Here, the (N-1)th image frame is the adjacent image of the previous image frame of the Nth image frame. For example, if the Nth image frame is the current image frame that needs to be processed, then the (N-1)th image frame is the previous image frame of the current image frame. It can also be understood that the (N-1)th image frame is a historical image frame.

[0039] In computer graphics, deferred rendering is a rendering method that postpones shading calculations until after depth testing. It can be understood as first drawing all objects into a screen-space geometric buffer (G-buffer), and then shading this buffer per light source. This avoids calculating the shading of fragments discarded by depth testing, thus avoiding unnecessary overhead. The G-buffer contains rich and diverse image information, such as depth and motion vectors (MV), which are strongly related to imaging. Depth can include the depth of the (N-1)th image frame, the depth of the Nth image frame, etc. The depth of the (N-1)th image frame can be denoted as `depth_pre`; the depth of the Nth image frame can be denoted as `depth_cur`.

[0040] In this embodiment, the depth of the (N-1)th image frame is aligned to the corresponding position of the Nth image frame by motion vectors to achieve depth mapping and obtain the first mapped depth, which can be denoted as depth_pre_warp.

[0041] Since depth and color have a complete positional correspondence, and directly visualizing depth is not intuitive, for ease of understanding, the color representation of the (N-1)th image frame corresponding to depth_pre is given, as follows: Figure 2 As shown; the mapped color representation corresponding to depth_pre_warp (denoted as color_pre_warp) is also given, such as Figure 3 As shown. Among them, Figure 2 In the (N-1)th image frame, the ball moves to the right, and the arrow indicates its direction of motion. Figure 3 In (the Nth image frame), the ball has moved to... Figure 2 The area pointed to by the arrow, due to occlusion, shows a black portion that is the shadow region caused by depth mapping (or depth alignment), which is the shadow portion present in `color_pre_warp`. According to... Figure 2 and Figure 3 It can be clearly seen that there is a trailing part in color_pre_warp. This trailing part is the occluded area, which is also the area that needs to be accurately detected by depth in this embodiment.

[0042] Step 12: Determine the first occlusion region and the first background region based on the first mapping depth and the depth of the Nth image frame;

[0043] In this embodiment, the first occlusion region is initially determined based on the first mapping depth, and the first background region is the background region determined by mapping the depth of the background region of the (N-1)th image frame to the Nth image frame, combined with the background region of the Nth image frame. The image may include a foreground image region (or foreground area) and a background image region (or background area). The foreground image region is the area of ​​focus in the image, such as the area where a moving subject is located. The moving subject may be a main person, animal, or other moving object in the image. The background image region is the background portion of the image, the remaining part of the image excluding the foreground, existing as a backdrop or environment, and is usually a secondary area. This step initially determines the occlusion region and the background region.

[0044] Step 13: Perform morphological dilation on the first occlusion region to obtain the second occlusion region;

[0045] In this embodiment, morphological dilation is an image processing operation. Performing morphological dilation on the first occluded region can be understood as: by moving the local boundary of the first occluded region, the size of the first occluded region changes, and the changed occluded region is called the second occluded region.

[0046] Step 14: Merge the second occlusion area with the first background area to determine the target occlusion area.

[0047] In this embodiment, merging the second occlusion area with the first background area can be understood as taking the union of the second occlusion area and the first background area. Based on the merged image, the final occlusion area, i.e. the target occlusion area, can be determined.

[0048] In the embodiments of this application, the depth of the (N-1)th image frame is mapped to the Nth image frame according to the motion vector to obtain a first mapped depth; a first occlusion region and a first background region are determined according to the first mapped depth and the depth of the Nth image frame; by performing morphological dilation processing on the obtained first occlusion region, the edge region of the foreground region in the obtained second occlusion region is clearer, thus the displayed foreground region is more accurate and complete; by merging the second occlusion region with the obtained first background region, the contours and regions of the foreground and background regions in the obtained image are clearer and more complete, thus the position of the occlusion region can be accurately obtained. This scheme can be applied to the occlusion relationship processing of temporal super-resolution algorithms, which can protect the edge region morphology of foreground objects in the image. In the processing, it retains sufficient image information from temporal fusion and avoids artifact problems caused by incorrect occlusion judgment, which can greatly optimize image quality.

[0049] In some embodiments, determining the first occlusion region based on the first mapping depth and the depth of the Nth image frame includes:

[0050] For each pixel, the difference between the depth of the Nth image frame and the first mapped depth is calculated to obtain the first depth difference value;

[0051] The first depth difference value is compared with the first threshold to determine the first occlusion area.

[0052] In this embodiment, the depth_cur of the Nth image frame is subtracted from the first mapped depth_pre_warped to obtain a first depth difference value, denoted as depth_diff. Using this first depth difference value and a set first threshold (denoted as thd0), the first occlusion region is initially determined.

[0053] In some embodiments, comparing the first depth difference value with a first threshold to determine the first occlusion region includes: comparing the first depth difference value with the first threshold to determine that the pixel corresponding to the first depth difference value that satisfies a first preset condition belongs to the first occlusion region; wherein the first preset condition includes: the first depth difference value is greater than or equal to the first threshold, or the first depth difference value is less than or equal to the first threshold.

[0054] In this embodiment, the first preset condition can be that the first depth difference value is greater than or equal to the first threshold, or that the first depth difference value is less than or equal to the first threshold. This can be understood as using the first threshold to filter pixels in the image when initially determining the occlusion region, and identifying pixels corresponding to the first depth difference value that meets the first preset condition as pixels in the first occlusion region. Comparing the first depth difference value with the first threshold can be understood as performing binarization processing. For example, setting the value to 0 when the first depth difference value corresponding to a pixel is greater than or equal to the first threshold, and setting it to 1 when the first depth difference value corresponding to a pixel is less than the first threshold, can obtain a binary image of the occlusion region (denoted as occ_mask0). Figure 4 As shown, this embodiment assumes the first threshold thd0 is 0.2. Figure 4 In the binary image shown, the white area is the initially determined first occlusion area, and the pixels in the white area can be considered as pixels in the first occlusion area.

[0055] In some embodiments, performing morphological dilation on the first occlusion region to obtain the second occlusion region includes:

[0056] For each pixel in the first occluded region, a second depth difference value is determined between the pixel and its neighboring pixels;

[0057] If the second depth difference value is greater than the second threshold, the pixel corresponding to the second depth difference value is determined to be the target boundary pixel of the first occlusion area, and the target boundary pixel is the boundary pixel of the adjacent foreground image.

[0058] The target boundary pixels are morphologically dilated in a direction away from the foreground image to obtain a second occlusion region.

[0059] In this embodiment, if the first occlusion region occ_mask0 is used directly, or the depth occlusion detection method provided in the ARM open-source FSR2.0 algorithm is used, the resulting image will have fragmented edges. In this embodiment, to solve this problem, the first occlusion region occ_mask0 is dilated. In this application, morphological dilation of the first occlusion region can be understood as: finding the boundary pixels of the adjacent foreground image in the first occlusion region, and moving the contour formed by the boundary pixels in a direction away from the foreground image.

[0060] Specifically, during the dilation process, referring to the depth of the Nth image frame, the dilation principle is as follows: for a given pixel, by comparing the depth of that pixel with the depth difference between it and other pixels in its neighborhood, it is determined whether that pixel is a boundary pixel near the foreground image. For example, for a given pixel, there are 8 neighboring pixels in total, arranged in a 3*3 pattern. The depth of the current pixel is compared with the depth of its 8 neighboring pixels. If the second depth difference between the current pixel and a neighboring pixel is greater than a preset second threshold thd1, then the current pixel is considered to be the boundary of an occluded area near the edge of the foreground image (such as a person). In this case, to protect the image quality of the edge of the foreground image (such as a person), the first occluded area occ_mask0 needs to be dilated outward from the foreground image (such as a person). Therefore, the neighboring pixel with the largest depth is selected for dilation processing.

[0061] If the second depth difference value between a certain pixel and its neighboring pixels is less than the preset second threshold thd1, then this pixel is considered to be the boundary of an occlusion region far from the edge of the foreground image (such as a person). In order to ensure that the calculated occlusion is not reduced, dilation should be avoided in this case.

[0062] After the above processing, the second occlusion region is obtained, denoted as occ_mask1. This process can be achieved through... Figure 5 exhibit, Figure 5 The content is the first occupancy region, occ_mask0. The purpose of this dilation process is to expand the edge region of the white area found in occ_mask0 that is closer to the foreground figure, moving it away from the figure. This region corresponds to... Figure 5 The pixels circled in red are considered to be in the nearby foreground image. Figure 5 The boundary pixels of the person in the image (i.e., the target boundary pixels). Since the person is irregularly shaped, the red circle only encloses a small portion; the actual algorithm processes the entire edge of the person. This step also needs to ensure that the outer boundary of the first occluded area, occ_mask0, is not shrunk by the dilation process. The outer area refers to the edge region far from the foreground person. Figure 5 The blue circle (similarly, the blue circle only circles a part of the area, and the actual processing is of the entire outer area). Figure 6 The result after the above operations is shown in the image viewing software. It is clear that the processed result achieves the effect of the edge area of ​​the figure expanding outwards while the outer area remains stationary. For ease of observation, this embodiment uses the difference between the results before and after processing for visualization. Figure 7 The red area represents the region affected by the expansion. As you can see, the expansion in this step only processes the edge areas of the foreground figures.

[0063] In some embodiments, determining a first background region based on the first mapping depth and the depth of the Nth image frame includes:

[0064] For each pixel, the first mapping depth and the third threshold are compared to determine that the pixels that meet the second preset condition belong to the second background region of the (N-1)th image frame.

[0065] Based on the motion vector, the depth of the second background region is mapped to the Nth image frame to obtain the mapped region of the depth of the second background region in the Nth image frame;

[0066] For each pixel, the depth of the Nth image frame is compared with a third threshold to determine that pixels that meet the second preset condition belong to the third background region of the Nth image frame.

[0067] The mapped region and the third background region are merged to obtain the first background region;

[0068] The second preset condition includes: the depth is greater than or equal to the third threshold, or the depth is less than or equal to the third threshold.

[0069] In this embodiment, since the depth difference of the background region in the image is small, the accuracy of using depth alone for occlusion judgment is very low. Therefore, this step needs to detect the background region first. A pre-set third threshold (denoted as thd2) can be used to compare the first mapping depth with the third threshold to first determine the background region of the (N-1)th image frame. The second preset condition can be greater than or equal to the third threshold, or less than or equal to the third threshold. For example, in the (N-1)th image frame, if the first mapping depth corresponding to a certain pixel is less than or equal to the third threshold, then the pixel is considered to belong to the background region of the (N-1)th image frame, which is denoted as the second background region; if the first mapping depth corresponding to a certain pixel is greater than the third threshold, then the pixel is considered not to belong to the background region of the (N-1)th image frame. The second background region of the (N-1)th image frame depth_pre can be obtained by this method, denoted as occ_mask2.

[0070] For example, in the example of this application embodiment, thd2 is 0.4. Pixels with a first mapping depth less than or equal to the third threshold are set to 0, and pixels with a first mapping depth greater than the third threshold are set to 1. The second background region of the N-1th image frame can be obtained through binarization processing.

[0071] Using motion vectors, the second background region occ_mask2 of the (N-1)th image frame is aligned to the Nth image frame. This yields the correspondence between the depth of the second background region in the (N-1)th image frame and its position in the Nth image frame. Specifically, the mapping region corresponding to the depth of the first background region in the Nth image frame is denoted as occ_mask3. Figure 8 As shown, the white area represents the position of the background area in the (N-1)th image frame corresponding to the Nth image frame.

[0072] To detect the background region, it is also necessary to filter out the background region in the Nth image frame. Using the third threshold (denoted as thd2), the depth of the Nth image frame is compared with the third threshold to determine the background region of the Nth image frame, which is denoted as the third background region.

[0073] For example, in the embodiments of this application, thd2 is set to 0.4. Pixels in the Nth image frame with a depth less than or equal to the third threshold are set to 0, and pixels in the Nth image frame with a depth greater than the third threshold are set to 1. Through binarization, the third background region of the Nth image frame can be obtained, denoted as occ_mask4. After obtaining the two background regions, the two background regions are merged. The merging process can be understood as taking the union of occ_mask4 and occ_mask3 to obtain the final first background region, denoted as occ_mask5. For example, it can be considered that the regions with a value of 1 in both occ_mask4 and occ_mask3 are the first background region occ_mask5.

[0074] After obtaining the second occlusion region occ_mask1 and the first background region occ_mask5, the union of occ_mask5 and occ_mask1 is taken to obtain the background region in occ_mask1 that is determined to be occluded. Further, the occlusion region occ_mask of the foreground image (such as a person) can be determined, such as... Figure 9 As shown, the white area is the final determined target occlusion area.

[0075] To more intuitively illustrate the differences between this application and the traditional FSR algorithm, the depth occlusion map of FSR is visualized, as shown below. Figure 10As shown, it can be seen that the method detected by this application is worse than the method in the background and the edge of the person in both the background and the occlusion areas. This is also the direct cause of jagged edges in the image and background areas.

[0076] To test the effectiveness of the occlusion region determination method of this application, the occlusion judgment method of this application and the occlusion judgment method of FSR2.0 were both applied to the same temporal super-resolution algorithm. The actual effects were tested in a static scene, motion scene and special effects scene of a certain game. It can be seen that the method of this application does not have problems such as ghosting and jagged edges at the edges of characters and background buildings.

[0077] This application proposes a method for determining occlusion regions in temporal video super-resolution algorithms. It improves existing occlusion detection methods by utilizing the engine's depth-domain motion vectors and designs a special dilation algorithm to solve the problem of foreground image edge fragmentation caused by inaccurate occlusion calculations, thereby enhancing the performance of temporal super-resolution algorithms. The method described in this application is as follows: Figure 11 As shown, the right-angled boxes represent the names of input parameters or intermediate variables, and the rounded boxes represent the specific calculation operations. The reasoning process of the steps is as follows:

[0078] 111. Depth Projection Alignment Processing. The rendering engine contains rich G-buffer information, including the depth of the (N-1)th image frame (depthPre), the depth of the Nth image frame (depthCur), and the motion vector (MV). The depth of the (N-1)th image frame is aligned to the position of the Nth image frame through the MV to obtain the projected depth, which is the first mapped depth depthPreWarp.

[0079] Calculate the depthPreWarp of the (N-1)th frame projected onto the position of the real frame: .

[0080] in, Represents the pixel of the Nth image frame ( The depth of ) This represents the motion vector corresponding to pixel i. Represents a function.

[0081] 112. Calculate the depth difference value. Subtract the depthPreWarp obtained in step 111 from the depthCur of the Nth image frame to obtain the first depth difference value, depthDiff.

[0082] Among them, the first depth difference between the projection depth and the depth of the Nth image frame is calculated. :

[0083] ;

[0084] in, This indicates the first depth difference.

[0085] 113. Binarize the depth difference value. Using the first depth difference value obtained in step 112 and the set first threshold thd0, the occlusion region template 0 (occ_Mask0) is initially obtained.

[0086] Specifically, the occlusion region template 0 (occ_Mask0) corresponding to the first depth difference value is calculated using the first threshold thd0:

[0087] .

[0088] 114. Dilate the occ_Mask0 obtained in step 113. During the dilation process, referencing the depth of the Nth image frame, the dilation principle is to compare the depth of the current pixel with the depth difference of other pixels in the neighborhood. When the difference is greater than the dilation threshold 1 (thd1), it is determined that this is the boundary of the occlusion area near the edge of the foreground person. In this case, to protect the image quality of the foreground person's edge, occ_Mask0 needs to be dilated outwards from the foreground person. Therefore, the nearest pixel with the largest depth is selected for dilation. When the difference is less than the set threshold thd1, it is determined that this is the boundary of the occlusion area far from the edge of the foreground person. To ensure that the occlusion calculated by occ_Mask is not reduced, no dilation is performed at this time. Finally, the occlusion area template 1 (occ_Mask1) is obtained.

[0089] Specifically, morphological dilation is performed on occ_Mask0 to determine the occlusion region template 1 (occ_Mask1) as follows:

[0090] .

[0091] 115. Binarize the background region of the (N-1)th image frame. Since the depth difference of the background region is small, the accuracy of occlusion judgment using depth alone is very low. Therefore, this step needs to detect the background region of the (N-1)th image frame and use the set dilation threshold 2 (thd2) to obtain the background region of depthPre, which is the occlusion template 2 (occ_Mask2) corresponding to the second background region.

[0092] Specifically, the background region and occ_Mask are determined using depth information and a set threshold thd2, and the occ_Mask2 occ_Mask2, which excludes the background region, is calculated.

[0093] .

[0094] 116. Align the projection of occ_Mask2. Using motion vectors, occ_Mask2 is aligned to the Nth image frame to obtain the depth of the second background region in the (N-1)th image frame and the corresponding positional relationship with the Nth image frame, thus obtaining the occlusion template 3 (occ_Mask3), which is the mapped region corresponding to occ_Mask2.

[0095] Specifically, using motion vector MV, occ_Mask2 is aligned to the Nth image frame, and the mapping region occ_Mask3 of the depth of the (N-1)th image frame in the Nth image frame is determined as follows:

[0096] .

[0097] 117. Binarize the background region of the Nth image frame. Since the depth difference of the background region is small, the accuracy of using depth alone for occlusion judgment is very low. Therefore, this step needs to detect the background region of the Nth image frame and use the set dilation threshold 2 (thd2) to obtain the background region of depthCur, which is the occlusion region template 4 (occ_Mask4) corresponding to the third background region.

[0098] Specifically, the occlusion template 4 (occ_Mask4) corresponding to the third background region is determined using the depth information of the Nth image frame and the set threshold thd2.

[0099] .

[0100] 118. Merging the mapped region occ_Mask3 and the third background region occ_Mask4 can be understood as taking the union of occ_Mask3 and occ_Mask4 to obtain the final background region. :

[0101] ;

[0102] 119. Finally, occMask5 and occMask1 are merged. This can be understood as taking the union of occMask5 and occMask1 to obtain the background area in occMask1 that is determined to be occluded, and then obtaining the final template occMask of the occluded area of ​​the foreground figure.

[0103] The final occlusion region template is calculated using the union of the two templates, as follows:

[0104] .

[0105] This method calculates the occlusion region, which can be used to guide the fusion of the (N-1)th and Nth image frames. The advantage of this approach is that it preserves sufficient image information from temporal fusion, avoids artifacts caused by incorrect occlusion judgment, and protects the jagged edges of foreground objects. The method in this application has lower resource overhead, higher compatibility, and can run in real time on mainstream commercial mobile devices.

[0106] In the embodiments of this application, the depth of the (N-1)th image frame is mapped to the Nth image frame according to the motion vector to obtain a first mapped depth; a first occlusion region and a first background region are determined according to the first mapped depth and the depth of the Nth image frame; by performing morphological dilation processing on the obtained first occlusion region, the edge region of the foreground region in the obtained second occlusion region is clearer, thus the displayed foreground region is more accurate and complete; by merging the second occlusion region with the obtained first background region, the contours and regions of the foreground and background regions in the obtained image are clearer and more complete, thus the position of the occlusion region can be accurately obtained. This scheme can be applied to the occlusion relationship processing of temporal super-resolution algorithms, which can protect the edge region morphology of foreground objects in the image. In the processing, it retains sufficient image information from temporal fusion and avoids artifact problems caused by incorrect occlusion judgment, which can greatly optimize image quality.

[0107] The occlusion area determination method provided in this application can be executed by an occlusion area determination device. This application uses an occlusion area determination device executing the occlusion area determination method as an example to illustrate the occlusion area determination device provided in this application.

[0108] The occlusion area determination device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0109] The occlusion area determination device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0110] The occlusion area determination device provided in this application embodiment can achieve... Figures 1 to 11 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0111] like Figure 12 As shown in the figure, this application embodiment also provides an occlusion area determination device 1200, including:

[0112] The first mapping module 1210 is used to map the depth of the (N-1)th image frame to the Nth image frame according to the motion vector to obtain the first mapping depth;

[0113] The first processing module 1220 is used to determine the first occlusion region and the first background region based on the first mapping depth and the depth of the Nth image frame.

[0114] The second processing module 1230 is used to perform morphological dilation processing on the first occlusion area to obtain the second occlusion area.

[0115] The third processing module 1240 is used to merge the second occlusion region with the first background region to determine the target occlusion region in the Nth image frame.

[0116] In some embodiments, the first processing module includes:

[0117] The first processing unit is used to calculate the difference between the depth of the Nth image frame and the first mapped depth for each pixel, and obtain a first depth difference value.

[0118] The second processing unit is used to compare the first depth difference value with the first threshold to determine the first occlusion area.

[0119] In some embodiments, the second processing unit is specifically used for:

[0120] Compare the first depth difference value with the first threshold to determine that the pixel corresponding to the first depth difference value that satisfies the first preset condition belongs to the first occlusion area.

[0121] The first preset condition includes: the first depth difference value is greater than or equal to the first threshold, or the first depth difference value is less than or equal to the first threshold.

[0122] In some embodiments, the second processing module is specifically used for:

[0123] For each pixel in the first occluded region, a second depth difference value is determined between the pixel and its neighboring pixels;

[0124] If the second depth difference value is greater than the second threshold, the pixel corresponding to the second depth difference value is determined to be the target boundary pixel of the first occlusion area, and the target boundary pixel is the boundary pixel of the adjacent foreground image.

[0125] The target boundary pixels are morphologically dilated in a direction away from the foreground image to obtain a second occlusion region.

[0126] In some embodiments, the first processing module includes:

[0127] The third processing unit is used to compare the first mapping depth and the third threshold for each pixel to determine that the pixel that meets the second preset condition belongs to the second background region of the (N-1)th image frame.

[0128] The fourth processing unit is used to map the depth of the second background region to the Nth image frame according to the motion vector, so as to obtain the mapping area of ​​the depth of the second background region in the Nth image frame;

[0129] The fifth processing unit is used to compare the depth of the Nth image frame with a third threshold for each pixel, and determine that the pixel that meets the second preset condition belongs to the third background region of the Nth image frame.

[0130] The sixth processing unit merges the mapped region and the third background region to obtain the first background region;

[0131] The second preset condition includes: the depth is greater than or equal to the third threshold, or the depth is less than or equal to the third threshold.

[0132] In the embodiments of this application, the depth of the (N-1)th image frame is mapped to the Nth image frame according to the motion vector to obtain a first mapped depth; a first occlusion region and a first background region are determined according to the first mapped depth and the depth of the Nth image frame; by performing morphological dilation processing on the obtained first occlusion region, the edge region of the foreground region in the obtained second occlusion region is clearer, thus the displayed foreground region is more accurate and complete; by merging the second occlusion region with the obtained first background region, the contours and regions of the foreground and background regions in the obtained image are clearer and more complete, thus the position of the occlusion region can be accurately obtained. This scheme can be applied to the occlusion relationship processing of temporal super-resolution algorithms, which can protect the edge region morphology of foreground objects in the image. In the processing, it retains sufficient image information from temporal fusion and avoids artifact problems caused by incorrect occlusion judgment, which can greatly optimize image quality.

[0133] Optionally, such as Figure 13 As shown, this application embodiment also provides an electronic device 1300, including a processor 1301 and a memory 1302. The memory 1302 stores a program or instructions that can run on the processor 1301. When the program or instructions are executed by the processor 1301, they implement the various steps of the above-described occlusion area determination method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0134] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0135] Figure 14 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0136] The electronic device 140 includes, but is not limited to, components such as: radio frequency unit 141, network module 142, audio output unit 143, input unit 144, sensor 145, display unit 146, user input unit 147, interface unit 148, memory 149, and processor 1410.

[0137] Those skilled in the art will understand that the electronic device 140 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1410 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 14 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0138] The processor 1410 is configured to: map the depth of the (N-1)th image frame to the Nth image frame based on the motion vector to obtain a first mapped depth; determine a first occlusion region and a first background region based on the first mapped depth and the depth of the Nth image frame; perform morphological dilation processing on the first occlusion region to obtain a second occlusion region; and merge the second occlusion region with the first background region to determine the target occlusion region in the Nth image frame.

[0139] In some embodiments, the processor 1410 determines the first occlusion region based on the first mapping depth and the depth of the Nth image frame, including:

[0140] For each pixel, the difference between the depth of the Nth image frame and the first mapped depth is calculated to obtain the first depth difference value;

[0141] The first depth difference value is compared with the first threshold to determine the first occlusion area.

[0142] In some embodiments, the processor 1410 compares the first depth difference value with a first threshold to determine a first occlusion region, including:

[0143] Compare the first depth difference value with the first threshold to determine that the pixel corresponding to the first depth difference value that satisfies the first preset condition belongs to the first occlusion area.

[0144] The first preset condition includes: the first depth difference value is greater than or equal to the first threshold, or the first depth difference value is less than or equal to the first threshold.

[0145] In some embodiments, the processor 1410 performs morphological dilation processing on the first occlusion region to obtain a second occlusion region, including:

[0146] For each pixel in the first occluded region, a second depth difference value is determined between the pixel and its neighboring pixels;

[0147] If the second depth difference value is greater than the second threshold, the pixel corresponding to the second depth difference value is determined to be the target boundary pixel of the first occlusion area, and the target boundary pixel is the boundary pixel of the adjacent foreground image.

[0148] The target boundary pixels are morphologically dilated in a direction away from the foreground image to obtain a second occlusion region.

[0149] In some embodiments, the processor 1410 determines a first background region based on the first mapping depth and the depth of the Nth image frame, including:

[0150] For each pixel, the first mapping depth and the third threshold are compared to determine that the pixels that meet the second preset condition belong to the second background region of the (N-1)th image frame.

[0151] Based on the motion vector, the depth of the second background region is mapped to the Nth image frame to obtain the mapped region of the depth of the second background region in the Nth image frame;

[0152] For each pixel, the depth of the Nth image frame is compared with a third threshold to determine that pixels that meet the second preset condition belong to the third background region of the Nth image frame.

[0153] The mapped region and the third background region are merged to obtain the first background region;

[0154] The second preset condition includes: the depth is greater than or equal to the third threshold, or the depth is less than or equal to the third threshold.

[0155] In the embodiments of this application, the depth of the (N-1)th image frame is mapped to the Nth image frame according to the motion vector to obtain a first mapped depth; a first occlusion region and a first background region are determined according to the first mapped depth and the depth of the Nth image frame; by performing morphological dilation processing on the obtained first occlusion region, the edge region of the foreground region in the obtained second occlusion region is clearer, thus the displayed foreground region is more accurate and complete; by merging the second occlusion region with the obtained first background region, the contours and regions of the foreground and background regions in the obtained image are clearer and more complete, thus the position of the occlusion region can be accurately obtained. This scheme can be applied to the occlusion relationship processing of temporal super-resolution algorithms, which can protect the edge region morphology of foreground objects in the image. In the processing, it retains sufficient image information from temporal fusion and avoids artifact problems caused by incorrect occlusion judgment, which can greatly optimize image quality.

[0156] It should be understood that, in this embodiment, the input unit 144 may include a graphics processing unit (GPU) 1441 and a microphone 1442. The GPU 1441 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 146 may include a display panel 1461, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 147 includes at least one of a touch panel 1471 and other input devices 1472. The touch panel 1471 is also called a touch screen. The touch panel 1471 may include a touch detection device and a touch controller. Other input devices 1472 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0157] The memory 149 can be used to store software programs and various data. The memory 149 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 149 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 149 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0158] Processor 1410 may include one or more processing units; optionally, processor 1410 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1410.

[0159] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described method for determining the occlusion area and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0160] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0161] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described occlusion area determination method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0162] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0163] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described occlusion area determination method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0164] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0165] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0166] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An occluded region determination method, characterized by, The method comprises the following steps: mapping the depth of the (N-1)th image frame into the Nth image frame according to the motion vector to obtain first mapping depth; determining a first occlusion region and a first background region according to the first mapping depth and the depth of the Nth image frame; performing morphological dilation processing on the first occlusion region to obtain a second occlusion region; merging the second occlusion region and the first background region to determine a target occlusion region in the Nth image frame.

2. The method of claim 1, wherein, The method comprises the following steps: calculating the difference between the depth of the Nth image frame and the first mapping depth for each pixel point to obtain a first depth difference value; comparing the first depth difference value with a first threshold to determine the first occlusion region.

3. The method of claim 2, wherein, The method comprises the following steps: comparing the first depth difference value with the first threshold to determine that the pixel point corresponding to the first depth difference value satisfying a first preset condition belongs to the first occlusion region; The first preset condition comprises that the first depth difference value is greater than or equal to the first threshold, or the first depth difference value is less than or equal to the first threshold.

4. The method of claim 2, wherein, The method comprises the following steps: determining the second depth difference value between each pixel point in the first occlusion region and adjacent pixel points; in the case that the second depth difference value is greater than a second threshold, determining that the pixel point corresponding to the second depth difference value is a target boundary pixel point of the first occlusion region, and the target boundary pixel point is a boundary pixel point adjacent to a foreground image; performing morphological dilation processing on the target boundary pixel point in a direction away from the foreground image to obtain a second occlusion region.

5. The method of claim 1, wherein, The method comprises the following steps: for each pixel point, comparing the first mapping depth with a third threshold to determine that the pixel point satisfying a second preset condition belongs to a second background region of the (N-1)th image frame; mapping the depth of the second background region into the Nth image frame according to the motion vector to obtain a mapping region of the depth of the second background region in the Nth image frame; for each pixel point, comparing the depth of the Nth image frame with the third threshold to determine that the pixel point satisfying the second preset condition belongs to a third background region of the Nth image frame; merging the mapping region and the third background region to obtain the first background region; The second preset condition comprises that the depth is greater than or equal to the third threshold, or the depth is less than or equal to the third threshold.

6. An occlusion area determination apparatus characterized by comprising: The method comprises the following steps: a first mapping module is configured to map the depth of the (N-1)th image frame into the Nth image frame according to the motion vector to obtain first mapping depth; a first processing module is configured to determine a first occlusion region and a first background region according to the first mapping depth and the depth of the Nth image frame; The second processing module is configured to perform morphological dilation processing on the first occlusion region to obtain a second occlusion region. The third processing module is configured to perform merging processing on the second occlusion region and the first background region to determine a target occlusion region in the Nth image frame.

7. The apparatus of claim 6, wherein, The first processing module comprises: The first processing unit is configured to calculate, for each pixel point, a difference between a depth of the Nth image frame and the first mapped depth to obtain a first depth difference value. The second processing unit is configured to compare the first depth difference value with a first threshold value to determine a first occlusion region.

8. The apparatus of claim 7, wherein, The second processing unit is specifically configured to: compare the first depth difference value with the first threshold value to determine that a pixel point corresponding to a first depth difference value satisfying a first preset condition belongs to the first occlusion region. The first preset condition comprises that the first depth difference value is greater than or equal to the first threshold value, or the first depth difference value is less than or equal to the first threshold value.

9. The apparatus of claim 7, wherein, The second processing module is specifically configured to: determine, for each pixel point in the first occlusion region, a second depth difference value between the pixel point and a neighboring pixel point; determine, in a case where the second depth difference value is greater than a second threshold value, that a pixel point corresponding to the second depth difference value is a target boundary pixel point of the first occlusion region, the target boundary pixel point being a boundary pixel point adjacent to a foreground image; perform morphological dilation processing on the target boundary pixel point in a direction away from the foreground image to obtain a second occlusion region.

10. The apparatus of claim 6, wherein, The first processing module comprises: The third processing unit is configured to compare, for each pixel point, the first mapped depth with a third threshold value to determine that a pixel point satisfying a second preset condition belongs to a second background region of the N-1th image frame. The fourth processing unit is configured to map, according to a motion vector, a depth of the second background region to the Nth image frame to obtain a mapped region of the depth of the second background region in the Nth image frame. The fifth processing unit is configured to compare, for each pixel point, a depth of the Nth image frame with a third threshold value to determine that a pixel point satisfying a second preset condition belongs to a third background region of the Nth image frame. The sixth processing unit is configured to perform merging processing on the mapped region and the third background region to obtain the first background region. The second preset condition comprises that the depth is greater than or equal to the third threshold value, or the depth is less than or equal to the third threshold value.

11. An electronic device, comprising: The device comprises a processor and a memory, the memory stores programs or instructions executable on the processor, and the programs or instructions are executed by the processor to implement the steps of the occlusion region determination method according to any one of claims 1 to 5.