Image processing method and device and electronic equipment

By acquiring images and depth maps, identifying significant areas and instance segmentation maps, and generating target map sequences for different scenarios, solving the problem of complex and time-consuming manual operations in the prior art, realizing automated video creation, and improving user creative efficiency and experience.

CN120219745APending Publication Date: 2025-06-27BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510344742.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, users need to convert still photos into dynamic videos through tedious manual operations, and the operation is complex, time-consuming and requires high requirements for users' operating skills, which reduces the enthusiasm for creation.

Method used

By acquiring the image and its depth map, identifying the prominent areas and instance segmentation maps of the image, generating target diagram sequences of different scenes based on the depth map and instance segmentation maps, realizing automated video creation.

Benefits of technology

It improves user creative efficiency, simplifies operational processes, reduces the requirements for user operation skills, and enhances user creative experience and enthusiasm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219745A_ABST
    Figure CN120219745A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and device and electronic equipment, and relates to the technical field of artificial intelligence, in particular to the technical fields of image processing, deep learning and the like, and the method comprises the steps: obtaining a first image and a depth map of the first image; obtaining a salient region of the first image according to the first image and the depth map; obtaining an instance segmentation image according to the first image and the salient region; according to the instance segmentation map and the depth map, target map sequences of different scenes are obtained, and the target map sequences comprise at least one instance segmentation map belonging to the scene, so that the target map sequences of the different scenes can be obtained more accurately and efficiently through the instance segmentation map and the depth map, and the user experience is improved. According to the technical scheme, the first image can be quickly divided into different levels of scenes, the creation efficiency of a user can be improved by obtaining the target image sequences of the different scenes, and the experience feeling of the user in the creation process is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as image processing and deep learning. In particular, it relates to an image processing method, apparatus, and electronic device. Background Art

[0002] With the rise of social media and short video platforms, more and more users hope to quickly convert static photos into dynamic videos. At the same time, they also hope that the converted videos are creative and can show personal characteristics. For example, generating a beat-matching video by combining music or adding filter effects. However, related technologies often rely on cumbersome manual operations, such as frame-by-frame matte extraction, background replacement, and video editing. The above processes are not only time-consuming and laborious, but also require relatively high operation skills from users, reducing the enthusiasm for creation. Summary of the Invention

[0003] The present disclosure provides an image processing method, apparatus, electronic device, storage medium, and computer program product.

[0004] According to a first aspect of the present disclosure, an image processing method is provided, including: obtaining a first image and a depth map of the first image; obtaining a salient region of the first image according to the first image and the depth map; obtaining an instance segmentation map according to the first image and the salient region; obtaining a sequence of target images for different scenes according to the instance segmentation map and the depth map, where the sequence of target images includes at least one instance segmentation map belonging to the scene.

[0005] According to a second aspect of the present disclosure, an image processing apparatus is provided, including: a first obtaining module for obtaining a first image and a depth map of the first image; a first processing module for obtaining a salient region of the first image according to the first image and the depth map; a second processing module for obtaining an instance segmentation map according to the first image and the salient region; a second obtaining module for obtaining a sequence of target images for different scenes according to the instance segmentation map and the depth map, where the sequence of target images includes at least one instance segmentation map belonging to the scene.

[0006] According to a third aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image processing method provided in the first aspect above.

[0007] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the image processing method proposed in the first aspect above.

[0008] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the image processing method proposed in the first aspect above.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 is a schematic flowchart of an image processing method according to an embodiment of the present disclosure;

[0012] Figure 2 is a schematic flowchart of an image processing method according to an embodiment of the present disclosure;

[0013] Figure 3 is a schematic flowchart of an image processing method according to an embodiment of the present disclosure;

[0014] FIG. 4 is a schematic diagram of adding a torn paper stroke special effect to an image according to an embodiment of the present disclosure;

[0015] Figure 5 is a schematic structural diagram of an image processing apparatus according to an embodiment of the present disclosure;

[0016] Figure 6 is a schematic block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0018] Artificial Intelligence (AI) is a new technical science that studies, develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence.

[0019] Image processing is a technology that uses a computer to analyze images to achieve the desired results. It is also known as image processing. Image processing generally refers to digital image processing. A digital image is a large two-dimensional array obtained by devices such as industrial cameras, video cameras, and scanners. The elements of this array are called pixels, and their values are called gray values. Image processing technology generally includes three parts: image compression, enhancement and restoration, and matching, description and recognition.

[0020] Deep Learning (DL for short) is a new research direction in the field of Machine Learning (ML for short). It is introduced into machine learning to make it closer to the original goal - artificial intelligence. Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained in these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability to analyze and learn like humans, and be able to recognize data such as text, images, and sounds.

[0021] Figure 1 The flowchart of the image processing method according to an embodiment of the present disclosure is shown as Figure 1 shown. The method includes:

[0022] S101, obtaining a first image and a depth map of the first image.

[0023] It should be noted that the execution subject of the image processing method in the embodiments of the present disclosure can be a hardware device with image processing capabilities and / or the necessary software to drive the hardware device to work. Optionally, the execution subject may include workstations, servers, computers, user terminals, and other intelligent devices. Among them, user terminals include but are not limited to mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle terminals, etc.

[0024] Among them, the first image can be an image collected in real time, an uploaded image, or an image extracted from a video.

[0025] It should be noted that the present disclosure does not limit the specific manner of obtaining the first image, and it can be selected according to the actual situation.

[0026] Optionally, an image can be collected in real time through an image acquisition device, and the image collected in real time can be used as the first image.

[0027] For example, an image can be collected in real time through a camera, and the image collected in real time can be used as the first image.

[0028] Optionally, the first image can be determined based on an image database.

[0029] For example, an image can be directly selected from an image database, and the selected image can be used as the first image, or operations such as cropping and splicing can be performed on the selected image to obtain the first image.

[0030] Optionally, an image can be extracted from a video, and any frame of the extracted image can be used as the first image.

[0031] In the embodiments of the present disclosure, after the first image is obtained, depth estimation can be performed on the first image to obtain the depth map of the first image.

[0032] Among them, each pixel in the depth map of the first image carries position information, and the depth map of the first image can reflect the spatial distance of the object in the first image, that is, it can reflect the distance between the object in the first image and the image acquisition device.

[0033] For example, for a street view image as the first image, the depth values of foreground objects (such as pedestrians and vehicles) are relatively small, and the depth values of background objects (such as buildings and the sky) are relatively large.

[0034] It should be noted that the present disclosure does not limit the specific method for obtaining the depth map of the first image, and can be selected according to the actual situation.

[0035] Optionally, a pre-trained depth map generation model can be called, and the first image is input into the depth map generation model to output the depth map of the first image.

[0036] S102. Obtain the salient region of the first image according to the first image and the depth map.

[0037] Among them, the salient region (salient area) of the first image refers to the eye-catching region in the first image and the region that can best represent the content of the first image.

[0038] It should be noted that the present disclosure does not limit the specific method for obtaining the salient region of the first image according to the first image and the depth map, and can be selected according to the actual situation.

[0039] Optionally, a pre-trained red-green-blue (RGB) image-depth (D) salient region detection model, that is, an RGB-D salient region detection model, can be called, and the first image and the depth map are input into the RGB-D salient region detection model to obtain the salient region of the first image.

[0040] Among them, the RGB-D salient region detection model is a salient detection technology that combines an RGB image and a depth map (D), and can more accurately identify and segment the salient region in the image.

[0041] It should be noted that the present disclosure does not limit the type of RGB-D saliency region detection model.

[0042] For example, the RGB-D saliency region detection model can be a Context Prior Network (CPNet) model.

[0043] In the embodiments of the present disclosure, by obtaining the salient region of the first image, the key content in the first image can be accurately identified, avoiding splitting the main body of the first image into multiple sub-images.

[0044] For example, when the first image is a portrait, the salient region of the first image needs to include the face, hair, upper body region, etc., so as to form a complete portrait main body.

[0045] S103. Obtain an instance segmentation map according to the first image and the salient region.

[0046] In the embodiments of the present disclosure, the first image can be subjected to instance recognition to obtain the full-scale instance segmentation map of the first image, determine the local instances included in the salient region, and perform occlusion processing on the segmentation maps of the local instances in the full-scale instance segmentation map to obtain the instance segmentation map.

[0047] Optionally, based on the instance segmentation model, the first image can be subjected to instance recognition to obtain the full-scale instance segmentation map of the first image.

[0048] For example, if the full-scale instance segmentation map of the first image includes the instance segmentation map of a person, the instance segmentation map of the sky, the instance segmentation map of a building, etc., and if the local instance included in the salient region is a person, perform occlusion processing on the instance segmentation map of the person in the full-scale instance segmentation map to obtain the instance segmentation map.

[0049] S104. Obtain a target map sequence of different scenes according to the instance segmentation map and the depth map, where the target map sequence includes at least one instance segmentation map belonging to the scene.

[0050] It should be noted that the present disclosure does not limit the specific manner of scene division, which can be divided according to the actual needs of the user.

[0051] For example, the scene can be divided into 3 levels, that is, including: foreground, middle ground, and background, and the scene can also be divided into 2 levels, that is, including foreground and background.

[0052] In the embodiments of the present disclosure, the instance segmentation map can be divided into scenes according to the coordinate values ​​of the pixels in the instance segmentation map in the vertical direction to obtain a first image sequence of different scenes. The first image sequence of different scenes can be optimized according to the depth map to obtain a second image sequence of different scenes. The instance segmentation maps in the second image sequence of each scene can be fused to obtain a target image sequence of different scenes.

[0053] It should be noted that in the related art, the video production method based on traditional cutouts and manual editing can be understood as users cutting out frames and editing videos through traditional image processing software. The specific steps include manually marking the foreground area in the image, that is, obtaining the foreground image sequence, using tools such as the lasso tool or the eraser to gradually and finely adjust the boundaries, and then adding the background image sequence and changing the layout as needed, and finally integrating the cutout results into the video clip for editing and rendering. The above scheme usually requires a lot of manual operations, and users need to be proficient in image processing and video editing software. It has a high learning threshold, complex operations and high time costs, which can easily lead to a decline in user creativity and affect sharing and creation frequency.

[0054] It should be noted that in the related art, the single model saliency detection or segmentation method based on deep learning can be understood as a salient region detection and instance segmentation model based on deep learning, which automatically identifies and segments salient regions or instances in the image. The above scheme is only based on RGB images for segmentation, and lacks consideration of image depth information, resulting in decreased segmentation accuracy in complex scenes. The above scheme only generates foreground segmentation or instance segmentation results, and cannot automatically divide the scene into hierarchical "multi-scenes" (for example: foreground, midground and background), which requires subsequent additional processing by the user.

[0055] It should be noted that, if the target image sequence includes multiple instance segmentation maps, the multiple instance segmentation maps in the target image sequence may be displayed according to average depth information or position information of the instance segmentation maps.

[0056] For example, for the foreground, if the target image sequence includes instance segmentation belonging to the foreground Figure 1 , instance segmentation Figure 2 and instance segmentation Figure 3 , respectively, to obtain instance segmentation Figure 1 , instance segmentation Figure 2 and instance segmentation Figure 3 The average depth information of each instance is segmented according to their respective average depth information. Figure 1 , instance segmentation Figure 2 and instance segmentation Figure 3 Sort from low to high, and use the sorting result as the display order of multiple instance segmentation maps in the target image sequence, that is, the instance segmentation maps at the front of the sorting result are displayed first.

[0057] For example, for the foreground, if the target image sequence includes instance segmentations belonging to the foreground Figure 1 and instance segmentations Figure 2 respectively obtain the coordinate values of the pixel points in the vertical direction in the instance segmentation Figure 1 and instance segmentations Figure 2 According to the coordinate values, sort the instance segmentation Figure 1 and instance segmentations Figure 2 from low to high, and use the sorting result as the display order of multiple instance segmentation maps in the target image sequence, that is, preferentially display the instance segmentation maps that are ranked higher in the sorting result.

[0058] In the embodiments of the present disclosure, after obtaining the target image sequences of different scenarios, they can be applied to various scenarios.

[0059] For example, if the scenario includes foreground, middle ground, and background, a torn paper stroke special effect can be added to the instance segmentation maps belonging to the foreground, middle ground, and background in the target image sequence.

[0060] For example, if the scenario includes foreground, middle ground, and background, after obtaining the instance segmentation maps belonging to the foreground, middle ground, and background in the target image sequence, the switching timing of each scenario can be intelligently adjusted according to the rhythm of the background music and the duration of the video that the user wants to generate, so as to automatically generate a music beat video.

[0061] The image processing method proposed by the present disclosure obtains a first image and a depth map of the first image, obtains a salient region of the first image according to the first image and the depth map, obtains an instance segmentation map according to the first image and the salient region, obtains a target image sequence of different scenarios according to the instance segmentation map and the depth map, and the target image sequence includes at least one instance segmentation map belonging to the scenario. Thus, the present disclosure can obtain the target image sequence of different scenarios more accurately and efficiently through the instance segmentation map and the depth map, can quickly divide the first image into different levels of scenarios, and by obtaining the target image sequence of different scenarios, it is beneficial to improve the user's creation efficiency and ensure the experience during the user's creation process.

[0062] Figure 2 is a schematic flowchart of the image processing method according to the second embodiment of the present disclosure.

[0063] As Figure 2 shown, on the basis of the embodiment shown in Figure 2 the image processing method of the embodiments of the present disclosure may specifically include the following steps:

[0064] S201, obtain a first image and a depth map of the first image.

[0065] S202. Obtain the salient region of the first image according to the first image and the depth map.

[0066] For the relevant content of S201 and S202, reference can be made to the above embodiments and will not be elaborated here.

[0067] Optionally, S103 in the above embodiments, "obtain the instance segmentation map according to the first image and the salient region", may specifically include the following S203 - S205.

[0068] S203. Perform instance recognition on the first image to obtain the full - scale instance segmentation map of the first image.

[0069] It should be noted that the present disclosure does not limit the specific manner of performing instance recognition on the first image to obtain the full - scale instance segmentation map of the first image.

[0070] Optionally, based on the instance segmentation model, instance recognition can be performed on the first image to obtain the bounding boxes (bboxes) of all instances and the corresponding segmentation masks (masks), and based on the bounding boxes (bboxes) of all instances and the corresponding segmentation masks (masks), the full - scale instance segmentation map of the first image can be obtained.

[0071] It should be noted that the present disclosure does not limit the type of the instance segmentation model. For example, the instance segmentation model can be the Efficient - Segment Anything Model (abbreviated as Efficient - SAM).

[0072] S204. Determine the local instances included in the salient region.

[0073] In the embodiments of the present disclosure, based on the full - scale instance segmentation map, the local instances included in the salient region can be determined.

[0074] S205. Perform occlusion processing on the segmentation maps of the local instances in the full - scale instance segmentation map to obtain the instance segmentation map.

[0075] In the embodiments of the present disclosure, occlusion processing can be performed on the local instances included in the salient region to display the instances included in the remaining region except the salient region in the full - scale instance segmentation map, so as to obtain the instance segmentation map.

[0076] Optionally, S104 in the above embodiments, "obtain the target image sequences of different scenarios according to the instance segmentation map and the depth map", may specifically include the following S206 - S205.

[0077] S206. According to the coordinate values of the pixel points in the vertical direction in the instance segmentation map, perform scene division on the instance segmentation map to obtain the first image sequence of different scenarios.

[0078] In the embodiments of the present disclosure, the coordinate values of pixel points in the vertical direction can be compared to determine the minimum coordinate value Y in the vertical direction in the instance segmentation map. min According to the minimum coordinate value Y min the instance segmentation map is scene-divided to obtain a first sequence of images.

[0079] Optionally, according to the minimum coordinate value, the instance segmentation maps are sorted from low to high to obtain a sequence of instance segmentation maps sorted from low to high according to the minimum coordinate value, and according to the height division thresholds preset for different scenes, the sorting results of the instance segmentation maps are scene-divided to obtain a first sequence of images.

[0080] For example, if the scenes include foreground, middle ground, and background, the preset height division threshold for the foreground can be [0 - y1], the preset height division threshold for the middle ground can be (y1 - y2], and the preset height division threshold for the background can be (y2 - y3].

[0081] It should be noted that after obtaining the height division thresholds preset for different scenes, the sorting results of the instance segmentation maps can be scene-divided, and the first sequence of images for different scenes can be quickly obtained, which is applicable to scenes with strong hierarchical features, such as natural landscapes, urban street scenes, etc. Without additional computing resources, only through the coordinate values in the vertical direction and the preset height division thresholds, the first sequence of images can be obtained.

[0082] For example, for the foreground, it can include the instance segmentation maps of the bottom area of the first image, such as the instance segmentation maps of the ground, grass, or other instance segmentation maps close to the camera; for the middle ground, it can include the instance segmentation maps of the middle area of the first image, such as the instance segmentation maps of buildings, trees, or people; for the background scene, it can include the instance segmentation maps of the top area of the first image, such as the instance segmentation maps of the sky, distant mountains, etc.

[0083] S207. According to the depth map, the first sequences of images for different scenes are optimized to obtain second sequences of images for different scenes.

[0084] It should be noted that by using the coordinate values of pixel points in the vertical direction in the instance segmentation map to perform scene division on the instance segmentation map, it can be understood as performing scene division on the instance segmentation map through geometric positions. However, for complex scenes, such as cases where there are occlusions between instances in the image or the hierarchy in the image is not obvious, the accuracy of the obtained first sequence of images will be limited. Therefore, according to the depth map, the first sequences of images for different scenes can be optimized to obtain second sequences of images for different scenes to improve the accuracy of the obtained second sequences of images.

[0085] In an embodiment of the present disclosure, the average depth information of the instance segmentation map can be obtained according to the depth map, and the first image sequence of different scenarios can be optimized according to the average depth information of the instance segmentation map to obtain the second image sequence of different scenarios.

[0086] It should be noted that the average depth information reflects the relative distance of the instance segmentation map in the three-dimensional space. Compared with the coordinate value (geometric position) of the pixel points in the instance segmentation map in the vertical direction, the average depth information provides an additional spatial dimension. By obtaining the average depth information of the instance segmentation map, the overall position of the instance segmentation map in the three-dimensional space can be determined.

[0087] It should be noted that the present disclosure does not limit the method for determining the average depth information of the instance segmentation map, and it can be selected according to the actual situation.

[0088] Optionally, the depth value of the pixel points in the instance segmentation map can be determined according to the depth map, the mask value of the pixel points in the instance segmentation map can be determined, and the average depth information of the instance segmentation map can be determined according to the depth value and the mask value of the pixel points.

[0089] In an embodiment of the present disclosure, the weight value of the pixel points can be determined according to the mask value of the pixel points, and the depth value of the pixel points can be weighted based on the weight value to obtain the average depth information of the instance segmentation map.

[0090] It should be noted that the weight value of the pixel points can be directly determined according to the mask value of the pixel points, that is, when the mask value of the pixel point is 1, the depth value of the pixel will participate in the subsequent weighted calculation, and when the mask value of the pixel point is 0, the depth value of the pixel will not participate in the subsequent weighted calculation.

[0091] For example, the following formula can be used to determine the average depth information of the instance segmentation map:

[0092]

[0093] where D(x,y) is the depth value of the pixel points in the instance segmentation map, S mask (x,y) is the mask value of the pixel (x,y) in the instance segmentation map, and D avg is the average depth information of the instance segmentation map.

[0094] In an embodiment of the present disclosure, the depth sorting within the sequence of the instance segmentation maps in the first image sequence can be performed according to the average depth information of the instance segmentation map, and the adjustment between the sequences of the instance segmentation maps in the first image sequence can be performed according to the preset depth division threshold of different scenarios to obtain the second image sequence of different scenarios.

[0095] For example, according to the average depth information of the instance segmentation maps, the instance segmentation maps in the first map sequence can be sorted in ascending depth within the sequence.

[0096] It should be noted that the depth division thresholds for different scenarios can be preset according to the actual situation.

[0097] For example, the foreground can be a shallow-depth area, and the preset depth division threshold for the foreground is [0 - d1]. The middle ground is a medium-depth area, and the preset depth division threshold for the middle ground is (d1 - d2]. The background can be a high-depth area, and the depth division threshold for the background is (d2 - d3].

[0098] In the embodiments of the present disclosure, after obtaining the depth division thresholds preset for different scenarios, the instance segmentation maps in the first map sequence can be adjusted between sequences to obtain the second map sequences for different scenarios.

[0099] For example, if the first image is an urban street view image, and the first map sequence of the background includes instance segmentation maps of high-rise buildings and the sky, but the average depth information of the instance segmentation map of the high-rise building is significantly less than that of the instance segmentation map of the sky, then the instance segmentation map of the high-rise building can be adjusted to the middle ground, so that the actual spatial hierarchy of the instance segmentation map can be more accurately recognized, and the division accuracy in complex scenarios is significantly improved.

[0100] S208, fuse the instance segmentation maps in the second map sequence for each scenario to obtain the target map sequence for different scenarios.

[0101] In the embodiments of the present disclosure, the area of the instance segmentation map can be determined, and based on the area and average depth information of the instance segmentation map, the instance segmentation maps are screened and merged to obtain the target instance segmentation maps in the second map sequence.

[0102] Among them, the first instance segmentation map may be caused by edge errors, noise interference, or detailed objects.

[0103] It should be noted that in the instance segmentation map, there are often some instance segmentation maps with small areas and incomplete semantics, that is, the first instance segmentation maps, such as handbags in urban street view images, small decorations of buildings, etc., and leaves in natural scene images.

[0104] Optionally, according to the area of the instance segmentation map, the first instance segmentation maps are screened from the instance segmentation maps. According to the average depth information, the second instance segmentation maps adjacent to the first instance segmentation maps are determined, and the first instance segmentation maps are merged into the corresponding adjacent second instance segmentation maps to obtain the target instance segmentation maps.

[0105] It should be noted that the present disclosure does not limit the specific method for screening the first instance segmentation map from the instance segmentation map according to the area of the instance segmentation map, which can be selected according to the actual situation.

[0106] Optionally, the area of the first image can be obtained, and an area threshold can be obtained according to the area of the first image. For example, the area threshold is set to 5% of the area of the first image, and the instance segmentation map with an area smaller than the area threshold in the instance segmentation map is used as the first instance segmentation map.

[0107] In the embodiments of the present disclosure, the position information of the first instance segmentation map can be obtained, and according to the position information of the first instance segmentation map, the third instance segmentation map adjacent to the first instance segmentation map can be determined. The difference between the average depth information of the first instance segmentation map and the average depth information of the third instance segmentation map is obtained, and the instance segmentation map with a difference less than a set value is selected as the adjacent second instance segmentation map.

[0108] In the embodiments of the present disclosure, in order to avoid the problem of semantic fragmentation, after the first instance segmentation map and the second instance segmentation map adjacent to the first instance segmentation map are obtained, the first instance segmentation map can be merged into the second instance segmentation map to obtain the target instance segmentation map.

[0109] For example, if the first image is an urban street view image, the handbag (the first instance segmentation map) in the urban street view image can be merged into the pedestrian (the second instance segmentation map) to obtain the target instance segmentation map.

[0110] In this embodiment, in order to optimize the region boundary to ensure the naturalness and visual consistency of the target instance segmentation map, boundary optimization and morphological processing can be performed on the target instance segmentation map obtained by merging.

[0111] Optionally, a region fusion operation can be performed on the target instance segmentation map to optimize the boundary of the target instance segmentation map.

[0112] For example, if the first image is a forest image, some leaves may be used as an instance segmentation map. Through the region fusion operation, the leaves can be merged into the associated instance segmentation map. For example, for the instance segmentation map with the trunk or the crown of the tree as the main body, through boundary optimization, the coherence and naturalness of the target instance segmentation map can be significantly improved.

[0113] It should be noted that the morphological processing includes but is not limited to dilation processing and erosion processing, etc.

[0114] Optionally, by performing dilation processing on the target instance segmentation map, the situation of boundary fracture of the target instance segmentation map can be repaired.

[0115] For example, if the first image is an architectural image, dilation processing can fill the gaps at the edges of the building roof, making the building (the target instance segmentation map) more complete.

[0116] Optionally, by performing erosion processing on the target instance segmentation map, small noise points and pseudo-segmentation regions in the target instance segmentation map can be removed.

[0117] For example, if the first image is a natural scene image, erosion processing can remove the fragmented miscellaneous points in the grass, making the grass (the target instance segmentation map) more coherent.

[0118] The specific process of the image processing method proposed in this disclosure will be explained below.

[0119] For example, as Figure 3 shown, (1) Obtain the depth map of the first image: Perform depth estimation on the original image (the first image) uploaded by the user to generate the depth map (Depth) of the first image, where the first depth map can be used for subsequent instance segmentation and salient region detection; (2) Obtain the salient region of the first image: Input the first image and the depth map into the RGB-D salient region detection model to generate the salient region segmentation result (salient mask), that is, obtain the salient region of the first image, where the first salient region can be used for subsequent scene segmentation and the generation of image special effects in the actual application scenario; (3) Obtain the instance segmentation map: Perform instance segmentation on the first image and the salient region through the instance segmentation model to generate the segmentation result to obtain the instance segmentation map; (4) According to the instance segmentation map and the depth map, obtain the target map sequence of different scenes; (5) Post-processing strategy: According to the specific needs of the user, design a flexible post-processing strategy to optimize the segmentation result, including: dividing the instance segmentation map into different scenes according to the coordinate values of the pixel points in the vertical direction of the instance segmentation map to obtain the first map sequence of different scenes (geometric sorting strategy), optimizing the first map sequence of different scenes according to the depth map to obtain the second map sequence of different scenes (depth sorting strategy), and fusing the instance segmentation maps in the second map sequence of each scene to obtain the target map sequence of different scenes (region fusion strategy) to obtain the target map sequence of different scenes that meets the creative requirements; (6) Output the target map sequence and special effect video of multiple sub-scenes (for example: 3 sub-scenes, including foreground, middle ground, and background): The target map sequence of different scenes generated by the post-processing strategy can be applied to various creative scenarios, such as: paper tearing filter special effects, music beat videos, etc. to meet the creative needs of users.

[0120] For example, for the creative requirement of the user to add a torn paper stroke effect to an image, if the image is as shown in Fig. 4(a), through the image processing method proposed by the present disclosure, a torn paper stroke effect can be added to the instance segmentation maps of the lighthouse, grassland, ground, and signboard belonging to the foreground in the target image sequence, a torn paper stroke effect can be added to the instance segmentation map of the sea area belonging to the middle ground in the target image sequence, and a torn paper stroke effect can be added to the instance segmentation maps of the sky, buildings, and mountains belonging to the background in the target image sequence, to generate an image that meets the creative requirements of the user as shown in Fig. 4(b).

[0121] The image processing method proposed by the present disclosure includes: obtaining a first image and a depth map of the first image, performing instance recognition on the first image to obtain a full-scale instance segmentation map of the first image, determining local instances included in the significant region, performing occlusion processing on the segmentation maps of the local instances in the full-scale instance segmentation map to obtain an instance segmentation map, dividing the instance segmentation map into different scene first image sequences according to the coordinate values of the pixel points in the vertical direction in the instance segmentation map, optimizing the first image sequences of different scenes according to the depth map to obtain second image sequences of different scenes, and fusing the instance segmentation maps in the second image sequences of each scene to obtain a target image sequence of different scenes. Thus, through the organic combination of the geometric sorting strategy, the depth sorting strategy, and the region fusion strategy, the present disclosure can obtain the target image sequences of different scenes more flexibly, accurately, and efficiently, realizing the hierarchical and coherent instance segmentation maps belonging to different scenes included in the target image sequence, realizing the automatic generation of the instance segmentation maps of multiple sub-scenes, being applicable to the requirements of various complex scenes, and through obtaining the target image sequences of different scenes, it can be applied to scenarios such as video editing and special effect synthesis, greatly improving the user's creation efficiency and experience, and being beneficial to improving the user's stickiness.

[0122] It should be noted that the information (including but not limited to question and answer information, user device information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in the present disclosure are all authorized by the user or fully authorized by all parties, and the sources, uses, and processing of the relevant data comply with relevant laws, regulations, and standards, and do not violate public order and good customs.

[0123] According to an embodiment of the present disclosure, the present disclosure also provides an image processing apparatus for implementing the above image processing method.

[0124] Figure 5 is a block diagram of an image processing apparatus according to an embodiment of the present disclosure.

[0125] As Figure 5As shown, the image processing apparatus 500 includes: a first acquisition module 501, a first processing module 502, a second processing module 503, and a second acquisition module 504.

[0126] The first acquisition module 501 is configured to acquire a first image and a depth map of the first image;

[0127] The first processing module 502 is configured to obtain a salient region of the first image according to the first image and the depth map;

[0128] The second processing module 503 is configured to obtain an instance segmentation map according to the first image and the salient region;

[0129] The second acquisition module 504 is configured to obtain a sequence of target maps for different scenarios according to the instance segmentation map and the depth map, where the sequence of target maps includes at least one instance segmentation map belonging to the scenario.

[0130] In an embodiment of the present disclosure, the second acquisition module 504 is configured to: divide the instance segmentation map into different scenarios according to the coordinate values of the pixel points in the vertical direction of the instance segmentation map to obtain a first sequence of maps for different scenarios; optimize the first sequence of maps for different scenarios according to the depth map to obtain a second sequence of maps for different scenarios; and fuse the instance segmentation maps in the second sequence of maps for each scenario to obtain the sequence of target maps for different scenarios.

[0131] In an embodiment of the present disclosure, the second acquisition module 504 is configured to: obtain average depth information of the instance segmentation map according to the depth map; and optimize the first sequence of maps for different scenarios according to the average depth information of the instance segmentation map to obtain a second sequence of maps for different scenarios.

[0132] In an embodiment of the present disclosure, the second acquisition module 504 is configured to: determine the depth value of the pixel points in the instance segmentation map according to the depth map; determine the mask value of the pixel points in the instance segmentation map; and determine the average depth information of the instance segmentation map according to the depth value and the mask value of the pixel points.

[0133] In an embodiment of the present disclosure, the second acquisition module 504 is configured to: determine the weight value of the pixel points according to the mask value of the pixel points; and perform weighted calculation on the depth value of the pixel points based on the weight value to obtain the average depth information of the instance segmentation map.

[0134] In one embodiment of the present disclosure, the second acquisition module 504 is configured to: compare the coordinate values of the pixel points in the vertical direction to determine the minimum coordinate value in the vertical direction in the instance segmentation map; according to the minimum coordinate value, perform scene division on the instance segmentation map to obtain the first graph sequence.

[0135] In one embodiment of the present disclosure, the second acquisition module 504 is configured to: sort the instance segmentation map from low to high according to the minimum coordinate value; according to the height division threshold preset for different scenes, perform scene division on the sorting result of the instance segmentation map to obtain the first graph sequence.

[0136] In one embodiment of the present disclosure, the second acquisition module 504 is configured to: perform in-sequence depth sorting on the instance segmentation maps in the first graph sequence according to the average depth information of the instance segmentation map; according to the depth division threshold preset for different scenes, perform inter-sequence adjustment on the instance segmentation maps in the first graph sequence to obtain the second graph sequence of different scenes.

[0137] In one embodiment of the present disclosure, the second acquisition module 504 is configured to: determine the area of the instance segmentation map; screen and merge the instance segmentation map according to the area and average depth information of the instance segmentation map to obtain the target instance segmentation map in the second graph sequence.

[0138] In one embodiment of the present disclosure, the second acquisition module 504 is configured to: screen the first instance segmentation map from the instance segmentation map; determine the second instance segmentation map adjacent to the first instance segmentation map according to the average depth information; merge the first instance segmentation map into the corresponding adjacent second instance segmentation map to obtain the target instance segmentation map.

[0139] In one embodiment of the present disclosure, the second acquisition module 504 is configured to: obtain the position information of the first instance segmentation map, and according to the position information of the first instance segmentation map, determine the third instance segmentation map adjacent to the first instance segmentation map; obtain the difference between the average depth information of the first instance segmentation map and the average depth information of the third instance segmentation map, and select the instance segmentation map with the difference less than the set value as the adjacent second instance segmentation map.

[0140] In one embodiment of the present disclosure, the apparatus 500 is configured to: perform boundary optimization and morphological processing on the target instance segmentation map obtained by merging.

[0141] In one embodiment of the present disclosure, the second processing module 503 is configured to: perform instance recognition on the first image to obtain a full-scale instance segmentation map of the first image; determine local instances included in the significant region; and perform occlusion processing on the segmentation maps of the local instances in the full-scale instance segmentation map to obtain the instance segmentation map.

[0142] The image processing apparatus proposed by the present disclosure obtains a first image and a depth map of the first image, obtains a significant region of the first image according to the first image and the depth map, obtains an instance segmentation map according to the first image and the significant region, and obtains a target map sequence of different scenarios according to the instance segmentation map and the depth map. The target map sequence includes at least one instance segmentation map belonging to the scenario. Thus, the present disclosure can more accurately and efficiently obtain the target map sequence of different scenarios through the instance segmentation map and the depth map, can quickly divide the first image into different levels of scenarios, and is beneficial to improving the user's creation efficiency and ensuring the experience during the user's creation process by obtaining the target map sequence of different scenarios.

[0143] According to an embodiment of the present disclosure, the present disclosure also proposes an electronic device, a readable storage medium, and a computer program product.

[0144] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0145] As Figure 6 shown, the device 600 includes a computing unit 601 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0146] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as a keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as a disk, optical disc, etc.; and communication unit 609, such as a network card, modem, wireless communication transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0147] Computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 601 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 601 executes the various methods and processes described above, such as an image processing method. For example, in some embodiments, the image processing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by computing unit 601, one or more steps of the image processing method described above can be executed. Alternatively, in other embodiments, computing unit 601 can be configured to execute the image processing method in any other suitable manner (e.g., by means of firmware).

[0148] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0149] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0150] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0151] To provide an interaction with a user account, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user account; and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user account can provide input to the computer. Other kinds of devices can also be used to provide an interaction with the user account; for example, the feedback provided to the user account can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user account can be received in any form (including acoustic input, voice input, or tactile input).

[0152] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user account computer having a graphical user account interface or a web browser through which the user account can interact with embodiments of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0153] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating blockchain.

[0154] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, it implements the steps of the image processing method described in the above embodiments of the present disclosure.

[0155] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitation is imposed herein.

[0156] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. An image processing method, wherein: The method comprises: Acquire a first image and a depth map of the first image; Acquire a salient area of ​​the first image according to the first image and the depth map; Acquire an instance segmentation map according to the first image and the salient area; According to the instance segmentation map and the depth map, a target map sequence of different scenes is acquired, where the target map sequence includes at least one instance segmentation map belonging to the scene.

2. The method according to claim 1, wherein: The step of acquiring target image sequences of different scenes according to the instance segmentation image and the depth image comprises: According to the coordinate values ​​of the pixels in the instance segmentation map in the vertical direction, the instance segmentation map is divided into scenes to obtain a first image sequence of different scenes; According to the depth map, optimizing the first image sequence of the different scenes to obtain a second image sequence of the different scenes; The instance segmentation images in the second image sequence of each scene are fused to obtain target image sequences of the different scenes.

3. The method according to claim 2, wherein: The step of optimizing the first image sequence of the different scenes according to the depth map to obtain the second image sequence of the different scenes includes: According to the depth map, obtaining average depth information of the instance segmentation map; The first image sequence of the different scenes is optimized according to the average depth information of the instance segmentation image to obtain the second image sequence of the different scenes.

4. The method according to claim 3, wherein: The obtaining, according to the depth map, average depth information of the instance segmentation map includes: Determine, according to the depth map, a depth value of a pixel point in the instance segmentation map; Determine a mask value of a pixel in the instance segmentation map; Determine average depth information of the instance segmentation map according to the depth value and the mask value of the pixel point.

5. The method according to claim 4, wherein: The step of determining average depth information of the instance segmentation map according to the depth value and the mask value of the pixel point includes: Determining the weight of the pixel point according to the mask value of the pixel point; Based on the weight value, the depth values ​​of the pixels are weightedly calculated to obtain average depth information of the instance segmentation map.

6. The method according to claim 2, wherein: The step of dividing the instance segmentation map into scenes according to the coordinate values ​​of the pixels in the instance segmentation map in the vertical direction to obtain a first image sequence of different scenes includes: Comparing the coordinate values ​​of the pixel points in the vertical direction to determine the minimum coordinate value in the vertical direction of the instance segmentation map; According to the minimum coordinate value, the instance segmentation map is divided into scenes to obtain the first scene division result.

7. The method according to claim 2, wherein: The performing scene division on the instance segmentation map according to the minimum coordinate value to obtain the first scene division result includes: According to the minimum coordinate value, sorting the instance segmentation map from low to high; According to the preset height division information of different scenes, the sorting results of the instance segmentation images are divided into scenes to obtain the first image sequences of the different scenes.

8. The method according to any one of claims 3 to 7, wherein: The step of optimizing the first image sequence of the different scenes according to the average depth information of the instance segmentation image to obtain the second image sequence of the different scenes includes: Determining, according to the average depth information of the instance segmentation maps, to perform depth sorting within a sequence on the instance segmentation maps in the first image sequence; According to the preset depth division ranges of different scenes, the instance segmentation images in the first image sequence are adjusted between sequences to obtain the second image sequence of the different scenes.

9. The method according to any one of claims 2 to 7, wherein: The step of fusing the instance segmentation images in the second image sequence of each scene to obtain the target image sequence of the different scenes includes: Determining the area of ​​the instance segmentation map; The instance segmentation maps are screened and merged according to the area and average depth information of the instance segmentation maps to obtain target instance segmentation maps in the second image sequence.

10. The method according to claim 9, wherein: The step of screening and merging the instance segmentation maps according to the area and average depth information of the instance segmentation maps to obtain the target instance segmentation maps in the second image sequence includes: Filtering a first instance segmentation map from the instance segmentation maps according to the area of ​​the instance segmentation map; Determine, according to the average depth information, a second instance segmentation map adjacent to the first instance segmentation map; The first instance segmentation map is merged with the corresponding adjacent second instance segmentation map to obtain the target instance segmentation map.

11. The method according to claim 10, wherein: The determining, according to the average depth information, a second instance segmentation map adjacent to the first instance segmentation map comprises: Acquire the position information of the first instance segmentation map, and determine a third instance segmentation map adjacent to the first instance segmentation map according to the position information of the first instance segmentation map; The difference between the average depth information of the first instance segmentation map and the average depth information of the third instance segmentation map is obtained, and the instance segmentation map whose difference is smaller than a set value is selected as the adjacent second instance segmentation map.

12. The method according to claim 10, wherein: The method further comprises: The target instance segmentation map obtained by merging is subjected to boundary optimization and morphological processing.

13. The method according to any one of claims 1 to 7, wherein: The acquiring of an instance segmentation map according to the first image and the salient area includes: Performing instance recognition on the first image to obtain a full instance segmentation map of the first image; determining a local instance included in the salient region; The segmentation map of the local instance in the full instance segmentation map is subjected to occlusion processing to obtain the instance segmentation map.

14. An image processing device, wherein: The device comprises: A first acquisition module, used to acquire a first image and a depth map of the first image; A first processing module, configured to obtain a salient area of ​​the first image according to the first image and the depth map; A second processing module, configured to obtain an instance segmentation map according to the first image and the salient area; The second acquisition module is used to acquire a target image sequence of different scenes according to the instance segmentation map and the depth map, wherein the target image sequence includes at least one instance segmentation map belonging to the scene.

15. An electronic device, characterized in that: including a processor and a memory; The processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to implement the method according to any one of claims 1 to 13.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 13 is implemented.

17. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 13.