Image processing method, electronic device, chip system and computer program product

By segmenting the target portrait and the first portrait in image processing, and combining the detection box and shadow segmentation map, the problem of poor performance of intelligent removal function in complex environments is solved, and more efficient and accurate image processing results are achieved.

CN120430930BActive Publication Date: 2026-05-22HONOR DEVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-10-22
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing intelligent elimination functions struggle to accurately separate the main character from the object to be eliminated in complex environments, resulting in poor elimination performance and negatively impacting user experience.

Method used

By segmenting the target human image and the first human image from the image to be processed, and combining the number of detection boxes and the human image segmentation map with the shadow segmentation map, non-target human images are accurately identified and eliminated, reducing the computational load of the algorithm and improving processing efficiency.

Benefits of technology

It improves the accuracy and processing efficiency of intelligent elimination, reduces misjudgment of passersby and false elimination of target person areas, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430930B_ABST
    Figure CN120430930B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image processing, and particularly to an image processing method, an electronic device, a chip system and a computer program product. In the image processing method of the present application, the main character and all characters in the image to be processed are detected, and the passers-by in the image to be detected can be determined by removing the main character from all characters. Since it is difficult to directly detect the passers-by, especially in the case of a large number of passers-by, the passers-by cannot be accurately detected. However, by using the method in the present application, the passers-by can be indirectly determined by detecting the main character, and the main character and the object to be eliminated can be accurately segmented, thereby helping to improve the effect of intelligent elimination and improving the user experience of using the intelligent elimination function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and more particularly to an image processing method, electronic device, chip system, and computer program product. Background Technology

[0002] With the development of image processing technology, electronic devices are becoming increasingly powerful in processing images. For example, electronic devices can provide intelligent removal (or AI removal) functions, which are used to remove specific elements from images, such as watermarks, background objects, or passersby.

[0003] When users take photos in complex environments, such as scenic spots, the surrounding environment and numerous tourists often interfere with the main subject or scenery being photographed, resulting in a large number of unwanted objects in the photos. In such cases, the situation of the objects to be removed can be quite complex. Existing intelligent removal functions cannot accurately separate the main subject from the objects to be removed, leading to poor intelligent removal results and affecting the user experience. Summary of the Invention

[0004] This application provides an image processing method, electronic device, chip system, and computer program product, which effectively improves the effect of intelligent image removal.

[0005] To achieve the above objectives, this application adopts the following technical solution:

[0006] Firstly, an image processing method is provided, including:

[0007] The target portrait and the first portrait are segmented from the image to be processed to obtain a first segmentation image of the target portrait and a second segmentation image of the first portrait; wherein, the first portrait includes the target portrait and non-target portraits;

[0008] A third segmentation map of the non-target human figure in the image to be processed is determined based on the first segmentation map and the second segmentation map;

[0009] The non-target human figures in the image to be processed are eliminated based on the third segmentation map.

[0010] In this embodiment, it is equivalent to detecting the main character and all other characters in the image to be processed. By removing the main character from all other characters, the passersby in the image to be detected can be identified. Since directly detecting passersby is difficult, especially when there are many passersby, it is impossible to detect them accurately. However, by using the method in this embodiment, passersby can be indirectly identified by detecting the main character, and the main character and the object to be eliminated can be segmented more accurately. This helps to improve the effect of intelligent elimination and enhances the user experience of using the intelligent elimination function.

[0011] In one implementation of the first aspect, segmenting the target portrait and the first portrait from the image to be processed to obtain a first segmentation image of the target portrait and a second segmentation image of the first portrait includes:

[0012] Detect the target portrait and the first portrait in the image to be processed;

[0013] If the number of first detection boxes is less than the number of second detection boxes and the number of first detection boxes is greater than 1, then the largest human image is segmented from the local image corresponding to the first detection box to obtain the first segmentation image.

[0014] The first portrait is segmented from the image to be processed based on the second detection box to obtain the second segmentation image; wherein, the first detection box is the detection box of the target portrait, and the second detection box is the detection box of the first portrait.

[0015] In one implementation of the first aspect, the method further includes:

[0016] If the number of the first detection boxes is greater than or equal to the number of the second detection boxes, and the number of the first detection boxes is greater than 1, then it is determined that the non-target human image does not exist in the image to be processed.

[0017] In one implementation of the first aspect, the method further includes:

[0018] If the number of the first detection boxes is 0 and the number of the second detection boxes is greater than 0, then the first portrait is segmented from the image to be processed based on the second detection boxes to obtain the second segmentation image;

[0019] The non-target human figure in the image to be processed is eliminated based on the second segmentation map.

[0020] In one implementation of the first aspect, the method further includes:

[0021] If the number of the first detection boxes is 0 and the number of the second detection boxes is 0, then it is determined that there is no human image in the image to be processed.

[0022] In one implementation of the first aspect, the method further includes:

[0023] If the number of the first detection boxes is 1 and there is no second detection box whose overlapping area with the first detection box reaches the first ratio, then the local image corresponding to the first detection box is subject segmented to obtain the fourth segmentation image.

[0024] If the number of human figures in the fourth segmentation image is less than the number of human figures in the second detection box, then the second segmentation image of the first human figure is segmented from the image to be processed based on the second detection box;

[0025] A fifth segmentation image of the non-target human figure in the first detection frame is determined based on the second segmentation image and the fourth segmentation image;

[0026] The non-target human figures in the image to be processed are eliminated based on the fifth segmentation map.

[0027] In one implementation of the first aspect, after performing subject segmentation on the local image corresponding to the first detection box to obtain a fourth segmentation image, the method further includes:

[0028] If the number of human figures in the fourth segmentation image is not less than the number of human figures in the second detection frame, then it is determined that the non-target human figure does not exist in the first detection frame.

[0029] In one implementation of the first aspect, segmenting the target portrait and the first portrait from the image to be processed to obtain a first segmentation image of the target portrait and a second segmentation image of the first portrait includes:

[0030] If the number of the first detection boxes is 1, and there is a second detection box whose overlapping area with the first detection box reaches a first ratio, then the human image with the largest area is segmented from the local image corresponding to the first detection box to obtain the first segmentation image.

[0031] The first human image is segmented from the image to be processed based on the second detection box to obtain the second segmentation image.

[0032] In the above embodiments, the target human figure in the image to be processed can be further determined based on the number of detection boxes for the main character and the number of detection boxes for all human figures. This helps to reduce the situation of misjudging passersby as main characters or vice versa, thereby improving the effect of intelligent elimination.

[0033] In one implementation of the first aspect, segmenting the target portrait and the first portrait from the image to be processed to obtain a first segmentation image of the target portrait and a second segmentation image of the first portrait includes:

[0034] The first human figure is segmented from the image to be processed to obtain the second segmentation image;

[0035] Detect the target human image in the image to be processed;

[0036] If the number of first detection boxes is less than the number of human figures in the second segmentation image, and the number of first detection boxes is greater than 1, then the segmentation image of the target human figure is obtained from the second segmentation image based on the first detection boxes, and the first segmentation image is obtained; wherein, the first detection box is the detection box of the target human figure.

[0037] In one implementation of the first aspect, the method further includes:

[0038] If the number of the first detection boxes is greater than or equal to the number of human figures in the second segmentation image, and the number of the first detection boxes is greater than 1, then it is determined that there is no non-target human figure in the image to be processed.

[0039] In one implementation of the first aspect, the method further includes:

[0040] If the number of the first detection boxes is 0 and the number of human figures in the second segmentation image is greater than 0, then the non-target human figures in the image to be processed are eliminated according to the second segmentation image.

[0041] In one implementation of the first aspect, the method further includes:

[0042] If the number of the first detection boxes is 0 and the number of the second detection boxes is 0, then it is determined that there is no human image in the image to be processed; wherein, the second detection box is the detection box of the human image in the second segmentation image.

[0043] In one implementation of the first aspect, the method further includes:

[0044] If the number of the first detection boxes is 1 and there is no second detection box whose overlapping area with the first detection box reaches the first ratio, then the local image corresponding to the first detection box is subject segmented to obtain a fourth segmentation image; wherein, the second detection box is the detection box of the human image in the second segmentation image;

[0045] If the number of human figures in the fourth segmentation image is less than the number of human figures in the first detection frame, then the fifth segmentation image of the non-target human figure in the first detection frame is determined based on the second segmentation image and the fourth segmentation image.

[0046] The non-target human figures in the image to be processed are eliminated based on the fifth segmentation map.

[0047] In one implementation of the first aspect, after performing subject segmentation on the local image corresponding to the first detection box to obtain a fourth segmentation image, the method further includes:

[0048] If the number of human figures in the fourth segmentation image is not less than the number of human figures in the second detection frame, then it is determined that the non-target human figure does not exist in the first detection frame.

[0049] In one implementation of the first aspect, the method further includes:

[0050] If the number of the first detection boxes is 1, and there is a second detection box whose overlapping area with the first detection box reaches a first ratio, then the human image with the largest area is segmented from the local image corresponding to the first detection box to obtain the first segmentation image; wherein, the second detection box is the detection box of the human image in the second segmentation image.

[0051] The above embodiments reduce the process of human face detection. Instead, they determine the passersby in the image based on the results of human face segmentation and the detection results of the main character. This approach helps to reduce the computational load of the algorithm, thereby improving the processing efficiency of intelligent removal.

[0052] In one implementation of the first aspect, segmenting the target portrait and the first portrait from the image to be processed to obtain a first segmentation image of the target portrait and a second segmentation image of the first portrait includes:

[0053] Detect the target portrait and the first portrait in the image to be processed;

[0054] The first human figure is segmented from the image to be processed to obtain the sixth segmentation image;

[0055] Based on the sixth segmentation image and the second detection box, the portrait is deduplicated to obtain the deduplicated third detection box; wherein, the second detection box is the detection box of the first portrait;

[0056] If the number of first detection boxes is less than the number of third detection boxes and the number of first detection boxes is greater than 1, then the segmentation map of the target human image is obtained from the sixth segmentation map based on the first detection boxes, and the first segmentation map is obtained.

[0057] The first portrait is segmented from the image to be processed based on the third detection box to obtain the second segmentation image.

[0058] In one implementation of the first aspect, the method further includes:

[0059] If the number of the first detection boxes is greater than or equal to the number of the third detection boxes, and the number of the first detection boxes is greater than 1, then it is determined that the non-target human image does not exist in the image to be processed.

[0060] In one implementation of the first aspect, the method further includes:

[0061] If the number of the first detection boxes is 0 and the number of the third detection boxes is greater than 0, then the first portrait is segmented from the image to be processed based on the third detection boxes to obtain the second segmentation image;

[0062] The non-target human figure in the image to be processed is eliminated based on the second segmentation map.

[0063] In one implementation of the first aspect, the method further includes:

[0064] If the number of the first detection boxes is 0 and the number of the third detection boxes is 0, then it is determined that there is no human image in the image to be processed.

[0065] In one implementation of the first aspect, the method further includes:

[0066] If the number of the first detection boxes is 1 and there is no third detection box whose overlapping area with the first detection box reaches the first ratio, then the local image corresponding to the first detection box is subject segmented to obtain the fourth segmentation image.

[0067] If the number of human figures in the fourth segmentation image is less than the number of human figures in the third detection frame, then the second segmentation image of the first human figure is segmented from the image to be processed based on the third detection frame;

[0068] A fifth segmentation image of the non-target human figure in the first detection frame is determined based on the second segmentation image and the fourth segmentation image;

[0069] The non-target human figures in the image to be processed are eliminated based on the fifth segmentation map.

[0070] In one implementation of the first aspect, after performing subject segmentation on the local image corresponding to the first detection box to obtain a fourth segmentation image, the method further includes:

[0071] If the number of human figures in the fourth segmentation image is not less than the number of human figures in the third detection frame, then it is determined that the non-target human figure does not exist in the first detection frame.

[0072] In one implementation of the first aspect, segmenting the target portrait and the first portrait from the image to be processed to obtain a first segmentation image of the target portrait and a second segmentation image of the first portrait includes:

[0073] If the number of the first detection boxes is 1, and there is a third detection box whose overlapping area with the first detection box reaches a first ratio, then the human image with the largest area is segmented from the local image corresponding to the first detection box to obtain the first segmentation image.

[0074] The first portrait is segmented from the image to be processed based on the third detection box to obtain the second segmentation image.

[0075] In the above embodiments, the position of the human face is determined by combining the segmentation and detection results. This can reduce the impact of detection errors of a single model on the subsequent elimination results. Especially for some small target human faces, when a single model fails to detect small target human faces, another model can supplement the detection results, thereby improving the detection accuracy of the human face and helping to improve the elimination effect.

[0076] In one implementation of the first aspect, the method further includes:

[0077] The shadows of the non-target human figures in the image to be processed are detected to obtain a shadow segmentation map; wherein, the shadow segmentation map includes the region of the non-target human figures and the region of the shadows of the non-target human figures in the third segmentation map;

[0078] The non-target human figure in the image to be processed is eliminated based on the shadow segmentation map.

[0079] In one implementation of the first aspect, the method further includes:

[0080] Determine whether the first region and the second region overlap; wherein, the first region is the shadowed region in the shadow segmentation image, and the second region is the target human image region in the first segmentation image;

[0081] Through the above embodiments, when there are shadows around a pedestrian, the mask image of the pedestrian's shadow can be effectively detected, and the pedestrian's shadow can also be eliminated in the subsequent elimination process, thereby improving the effect of intelligent elimination.

[0082] If the first region and the second region overlap, delete the overlapping portion of the first region to obtain the processed shadow segmentation map;

[0083] The non-target human figures in the image to be processed are eliminated based on the processed shadow segmentation map.

[0084] Through the above embodiments, when obtaining the segmentation map of the pedestrian's shadow, the target person's image can be effectively avoided. In the subsequent elimination process, the area of ​​the target person's image can be avoided, thereby improving the effect of intelligent elimination.

[0085] It should be noted that the process of segmenting human figures from the image to be processed can all adopt the human figure segmentation method in the following embodiments.

[0086] In one embodiment, the human face segmentation method may include the following steps:

[0087] The backbone network extracts image features from the image to be processed. A dot-mapping model identifies the human figure in the image and marks the human figure region with dots, obtaining detection points. The first head network segments the human figure from the image based on image features and detection points, obtaining a segmentation map of the human figure. The second head network segments the appendages of the human figure from the image based on image features and detection points, obtaining a segmentation map of the appendages. The segmentation maps of the human figure and the appendages are superimposed to obtain the final segmentation map.

[0088] Optionally, the center point of the human image area can be used as the detection point.

[0089] Optionally, multiple points within the human image area can be selected as detection points. Understandably, the more detection points there are, the more accurate the detection of attachments will be, but the computational load may be greater.

[0090] It should be noted that in some other implementations, detection boxes of the human figures in the image to be processed can be obtained, and the human figures and their surrounding objects can be segmented based on the detection boxes. Compared with this approach, the use of detection points in this embodiment helps to distinguish different human figures, especially when there are many human figures in the image and their positions are close together. Using detection points can more clearly distinguish different human figures, thereby reducing the situation where the segmentation effect is poor due to overlapping detection boxes.

[0091] Compared to traditional image segmentation models, the portrait segmentation model in this application adds a dot-mapping model and a second head network, such as... Figure 12 The portion shown is within the dashed box. The portrait segmentation model of this application can effectively detect objects surrounding a person, enabling the subsequent removal process to also eliminate these objects, thereby improving the effectiveness of intelligent removal.

[0092] In another embodiment, the human face segmentation method may include the following steps:

[0093] An initial segmentation map is obtained by segmenting human figures from the image to be processed using an image segmentation model. A semantic segmentation model is then used to identify the label of each human figure in the image. Based on the label of each human figure, the human figures in the initial segmentation map are filtered to obtain the final segmentation map.

[0094] The portrait segmentation model in this embodiment of the application incorporates a semantic segmentation model, such as... Figure 14 The portion shown is within the dashed box. Semantic segmentation models can capture the semantic differences between real and non-real images. The semantic segmentation results are then used to assist image segmentation models, effectively filtering out non-real images and allowing them to be retained during subsequent removal processes, thus improving the effectiveness of intelligent removal.

[0095] In another embodiment, the human face segmentation method may include the following steps:

[0096] Image segmentation model 1 is used to segment the human image from the image to be processed, resulting in an initial segmentation map. A small object model is then used to identify the human image in the image to be processed, generating human image detection boxes. These boxes are then filtered based on the initial segmentation map to obtain small object detection boxes. Image segmentation model 2 is used to segment the human image corresponding to each small object detection box from the image to be processed, generating a small object segmentation map. Finally, the initial segmentation map and the small object segmentation map are superimposed to obtain the final segmentation map.

[0097] The portrait segmentation model in this embodiment incorporates a small object detection model, such as... Figure 16 The area shown is within the dashed box. Small targets in the image are detected using a small target detection model, enabling subsequent removal of these small targets and thus improving the effectiveness of intelligent image removal.

[0098] Secondly, an electronic device is provided, comprising:

[0099] One or more processors;

[0100] One or more memory units;

[0101] The memory stores a computer program that, when executed by the processor, causes the electronic device to perform the method as described in any of the first aspects.

[0102] Thirdly, a chip system is provided, the chip system including a processor coupled to a memory, the processor being configured to run a computer program stored in the memory to implement the method as described in any of the first aspects.

[0103] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by one or more processors, implements the method as described in any of the first aspects.

[0104] Fifthly, a computer program product is provided, the computer program product including computer program code, which, when run on an electronic device, causes the electronic device to perform the method as described in any of the first aspects.

[0105] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0106] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0107] Figure 2 This is a software structure block diagram of the electronic device 100 according to an embodiment of this application;

[0108] Figure 3 A schematic diagram illustrating an application scenario of the intelligent elimination function in a gallery application provided in this application embodiment;

[0109] Figure 4 A schematic diagram illustrating an application scenario of the intelligent elimination function in a camera application provided in this application embodiment;

[0110] Figure 5 A schematic flowchart of the image processing method provided in the embodiments of this application;

[0111] Figure 6 A schematic diagram of the mask image provided in the embodiments of this application;

[0112] Figure 7 A schematic diagram illustrating the detection results of the target human image provided in an embodiment of this application;

[0113] Figure 8 A schematic flowchart of the image processing method provided in the embodiments of this application;

[0114] Figure 9 A schematic flowchart illustrating an image processing method provided in another embodiment of this application;

[0115] Figure 10 A schematic diagram of a shadow segmentation map provided in an embodiment of this application;

[0116] Figure 11 A schematic diagram illustrating target human figure area avoidance provided in an embodiment of this application;

[0117] Figure 12 A schematic diagram of a human face segmentation model provided in an embodiment of this application;

[0118] Figure 13 A schematic diagram of the detection of appendages provided in the embodiments of this application;

[0119] Figure 14 A schematic diagram of a human face segmentation model provided in an embodiment of this application;

[0120] Figure 15 This is a schematic diagram illustrating the detection of non-realistic human images provided in an embodiment of this application.

[0121] Figure 16 A schematic diagram of a human face segmentation model provided in an embodiment of this application;

[0122] Figure 17 A schematic diagram of small target detection provided for an embodiment of this application;

[0123] Figure 18 A schematic diagram of a human image segmentation model provided in an embodiment of this application. Detailed Implementation

[0124] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limiting purposes, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details.

[0125] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0126] It should also be understood that in the embodiments of this application, "one or more" refers to one, two, or more; "and / or" describes the relationship between the associated objects, indicating that three relationships can exist; for example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following associated objects have an "or" relationship.

[0127] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," "fourth," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0128] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0129] The image processing method provided in this application can be applied to electronic devices. Electronic devices include terminal devices, which can also be called terminals, user equipment (UE), mobile stations (MS), mobile terminals (MT), etc. Terminal devices can be mobile phones, smart TVs, wearable devices, tablets, smart screens, computers with wireless transceiver capabilities, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, and so on. The embodiments of this application do not limit the specific technologies or device forms used in the electronic devices.

[0130] See Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a touch sensor 180K, an ambient light sensor 180L, etc.

[0131] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0132] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0133] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.

[0134] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0135] External memory 120 generally refers to external storage. In the embodiments of this application, external storage refers to storage other than the memory of electronic devices and the cache of processors. This storage is generally non-volatile memory.

[0136] Internal memory 121, also known as "RAM," can be used to store executable program code for a computer, including instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a given function (such as sound playback, image playback, etc.).

[0137] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays screens 194, where N is a positive integer greater than 1. In some embodiments, electronic device 100 displays a user interface through the displays screens 194.

[0138] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0139] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0140] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0141] Electronic device 100 also includes various sensors that can convert different physical signals into electrical signals. For example, pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. Gyroscope sensor 180B can be used to determine the motion posture of electronic device 100. Barometric pressure sensor 180C is used to measure air pressure. Magnetic sensor 180D includes a Hall sensor. Accelerometer sensor 180E can detect the magnitude of acceleration of electronic device 100 in various directions (generally three axes). Distance sensor 180F is used to measure distance. Electronic device 100 can measure distance using infrared or laser. Proximity sensor 180G may include, for example, a light-emitting diode (LED) and a photodetector, such as a photodiode. Ambient light sensor 180L is used to sense ambient light brightness. Electronic device 100 can adaptively adjust the brightness of display screen 194 according to the sensed ambient light brightness. Fingerprint sensor 180H is used to collect fingerprints. Electronic device 100 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application lock, fingerprint photography, fingerprint answering of calls, etc. Temperature sensor 180J is used to detect temperature. In some embodiments, the electronic device 100 uses the temperature detected by the temperature sensor 180J to execute a temperature processing strategy. The bone conduction sensor 180M can acquire vibration signals.

[0142] Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of electronic device 100, in a different position than display screen 194.

[0143] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0144] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0145] The above is a detailed description of the embodiments of this application using electronic device 100 as an example. It should be understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on electronic device 100. Electronic device 100 may have more or fewer components than shown in the figures, may combine two or more components, or may have different component configurations. The various components shown in the figures can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0146] In addition, an operating system runs on top of the aforementioned components. Examples include iOS, Android (an open-source operating system), and Windows. Applications can be installed and run on this operating system.

[0147] The operating system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100.

[0148] Figure 2 This is a software structure block diagram of the electronic device 100 according to an embodiment of this application.

[0149] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the system libraries and Android runtime, and the kernel layer.

[0150] The application layer can include a series of application packages.

[0151] like Figure 2 As shown, the application package may include applications such as camera, gallery, calendar, call, map, e-book, shopping, Bluetooth, music, video, and SMS.

[0152] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0153] like Figure 3 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.

[0154] The window manager is used to manage windowed applications. It can obtain the screen size, determine if a status bar is present, lock the screen, and capture the screen, among other things.

[0155] The content provider is used to store and retrieve data, and to make this data accessible to applications. The data may include monitoring data acquired by the sensor module 180 (such as acceleration data acquired by the accelerometer 180E), video, images, audio, phone calls made and received, browsing history and bookmarks, phone books, etc.

[0156] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0157] The phone manager is used to provide communication functions for electronic device 100. For example, it manages call status (including connection and disconnection).

[0158] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and so on.

[0159] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of download completion or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0160] In this embodiment of the application, the application framework layer may include a camera access interface, wherein the camera access interface is used to provide an application programming interface and a programming framework for camera applications.

[0161] The Android runtime consists of core libraries and a virtual machine. The Android runtime is responsible for scheduling and managing the Android system.

[0162] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0163] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0164] The system layer can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0165] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0166] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0167] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0168] A 2D graphics engine is a graphics engine for 2D drawing.

[0169] The kernel layer is the layer between hardware and software. It contains at least the display driver, camera driver, audio driver, and sensor driver. The camera driver is used to drive the camera to capture images.

[0170] With the development of image processing technology, electronic devices are becoming increasingly powerful in processing images. For example, electronic devices can provide intelligent removal (or AI removal) functions, which are used to remove specific elements from images, such as watermarks, background objects, or passersby.

[0171] Understandable Figure 2 The application layer in the illustrated software architecture can provide intelligent removal functionality. For example, an application providing intelligent removal functionality may include program segments for image processing methods that implement intelligent removal functionality, enabling intelligent removal when the electronic device runs the application.

[0172] For example, a camera app or gallery app may offer a Smart Elimination feature. When a camera app offers Smart Elimination, users can use it to remove specific elements from a captured image, resulting in a cleaner image including the subject. Similarly, when a gallery app offers Smart Elimination, users can use it to remove specific elements from gallery photos, resulting in a cleaner gallery photo including the subject.

[0173] In one application scenario, see Figure 3 This is a schematic diagram illustrating an application scenario of the intelligent elimination function in a gallery application provided in this application embodiment. For example... Figure 3 The desktop shown in (a) includes icons for multiple applications, which the user can click to open. When the user clicks the Gallery app icon 301 on the desktop, the electronic device displays the following in response to the user's action: Figure 3 The image gallery application interface is shown in (b) above.

[0174] like Figure 3 The gallery application interface shown in (b) includes multiple image thumbnails, and users can click on different thumbnails to view the corresponding image. When a user clicks on thumbnail 302 in the gallery application interface, in response to the user's action, the electronic device displays as shown in (b). Figure 3 The image interface shown in (c) is shown in the image.

[0175] like Figure 3 The image interface shown in (c) includes multiple operation controls, such as a "Share" control, a "Favorite" control, an "Edit" control, a "Delete" control, and a "More" control. Users can click on different operation controls to perform operations on the image. When a user clicks the "Edit" control in the image interface, in response to the user's operation, the electronic device displays as shown in the image. Figure 3 The editing interface shown in (d) is shown in the image.

[0176] like Figure 3 The editing interface shown in (d) includes multiple editing controls, such as a "doodle" control, a "text" control, an "AI remove" control, a "mosaic" control, and a "watermark" control. Users can click on different editing controls to perform corresponding edits on the image. When a user clicks the "AI remove" control in the editing interface, in response to the user's action, the electronic device displays as shown in the image. Figure 3 The elimination interface is shown in (e) in the figure.

[0177] like Figure 3The elimination interface shown in (e) includes multiple operation controls, such as the "Smart Selection" control, the "Manual Smear" control, and the "One-Click Removal" control. When the user clicks the "Smart Selection" control, the electronic device enters the smart selection mode (also known as the elimination pen). In this mode, the user can circle the object to be eliminated on the screen, and the electronic device can intelligently identify and eliminate the circled object. When the user clicks the "Manual Smear" control, the electronic device enters the manual smear mode. The user can use their finger to smear back and forth on the screen, and the electronic device can eliminate the smeared area. When the user clicks the "One-Click Removal" control, the electronic device enters the smart elimination mode. This mode automatically identifies the main character (or target character) and passersby in the image and automatically eliminates passersby and related elements (such as shadows and accessories). After elimination, the electronic device displays as shown below. Figure 3 The elimination interface is shown in (f) in the figure.

[0178] like Figure 3 As shown in (f), only the main character is retained in the image, and passersby are eliminated.

[0179] Users can also... Figure 3 In the elimination interface shown in (f), clicking the back control cancels the previous operation. After the user completes the selection, they can also click the "Save" control to save the newly generated image after the objects are eliminated.

[0180] In another application scenario, see Figure 4 This is a schematic diagram illustrating an application scenario of the intelligent elimination function in a camera application provided in this application embodiment.

[0181] like Figure 4 The desktop shown in (a) includes icons for multiple applications, which the user can click to open. When the user clicks the camera application icon 401 on the desktop, in response to the user's action, the electronic device displays as shown in [image 1]. Figure 4 The camera application interface is shown in (b) above.

[0182] like Figure 4 The camera application interface shown in (b) includes multiple shooting modes, such as "Aperture," "Night Scene," "Portrait," "Photo," "Video," "Pro," and "More." When the user selects the "Photo" mode and clicks the shooting control 402, in response to the user's operation, the electronic device takes a picture of the scene in the preview frame, obtains the captured image, and displays it as shown in the image. Figure 4 The preview interface shown in (c) is shown in the image.

[0183] like Figure 4The preview interface shown in (c) includes a thumbnail display box 403, which displays thumbnails of the captured images. When the user clicks on the display box 403, in response to the user's action, the electronic device opens the gallery application and displays images such as... Figure 4 The image (d) shows the interface of the gallery application.

[0184] Figure 4 The interface shown in (d) is similar to Figure 3 The interface shown in (c) is the same as the one shown in the previous example. Users can edit the captured image by clicking the "Edit" control on this interface. For details on the intelligent image removal process, please refer to [link / reference]. Figure 3 The embodiments described in (c) to (f) are not repeated here.

[0185] In other application scenarios, users can edit images using voice assistants on electronic devices. Specifically, a user can open a voice assistant with a voice command, send an image to it, and in response, the voice assistant automatically recognizes the image. If the image meets the criteria for intelligent removal, the voice assistant can display a prompt message to the user through the interactive interface, suggesting "one-click removal." The user can then interact with the controls in the prompt message, and in response, the voice assistant will process the image to intelligently remove elements from it.

[0186] In other application scenarios, users can view images through applications other than the camera and gallery. When a user views an image in another application, the electronic device's smart assistant can display prompt A, suggesting that the user can edit the image using the smart assistant. The user can interact with the controls in prompt A. In response to this interaction, the smart assistant automatically identifies the image. If the image meets the criteria for intelligent removal, the smart assistant can display prompt B through the interactive interface, prompting the user to "remove with one click." The user can then interact with the controls in prompt B. In response to this interaction, the smart assistant performs image processing on the image to intelligently remove elements from it.

[0187] It should be noted that the above application scenarios are merely examples of entry points for the intelligent elimination function; the intelligent elimination function can also be accessed through other entry points. This application does not specifically limit the entry point for the intelligent elimination function.

[0188] Additionally, in some application scenarios, electronic devices may only provide a "one-click removal" control, that is, only a smart removal mode, without a smart selection mode and / or a manual smearing mode. In other application scenarios, the "one-click removal" control may also be set... Figure 3The interface shown in (d) is displayed simultaneously with the "AI Elimination" control. Of course, in some examples, the "One-Click Removal" control may also use other names, such as "Remove Passersby," "Person Elimination," "Eliminate Passersby," "Remove Passersby," etc. This application embodiment does not specifically limit the above situations.

[0189] When users take photos in complex environments, such as scenic spots, the surrounding environment and numerous tourists often interfere with the main people or scenery they want to photograph, resulting in a large number of unwanted objects in the photos. In such cases, the objects to be removed can be quite complex. For example, the image may include multiple passersby (extra people besides the main subject), which are not easily distinguishable from the main subject (also known as the target person); passersby in the image may be small and easily overlooked; passersby in the image may have shadows or appendages; the image may include portraits and sculptures, which are easily confused with real people.

[0190] In response to the aforementioned complexities, existing intelligent elimination functions are unable to accurately separate the main character from the object to be eliminated, resulting in poor intelligent elimination performance and impacting user experience.

[0191] Based on this, embodiments of this application provide an image processing method that can determine passersby in an image to be detected based on the main character and all other characters in the image to be processed. Since directly detecting passersby is difficult, especially when there are many passersby, it is impossible to detect them accurately. However, using the method in this application embodiment, passersby can be indirectly determined by detecting the main character, more accurately segmenting the main character and the object to be eliminated, thereby helping to improve the effect of intelligent elimination and enhancing the user experience of using the intelligent elimination function.

[0192] It should be noted that the image processing method provided in this application embodiment is mainly aimed at the process of electronic devices segmenting and eliminating passersby in the image in intelligent elimination mode after the user clicks the "one-click removal" control.

[0193] For ease of description, in the embodiments of this application, the subject being photographed (i.e., the person who does not need to be eliminated) is called the main subject, and the portrait of the main subject in the image to be processed is called the target portrait; the non-subject being photographed (i.e., the person other than the main subject) is called a passerby, and the portrait of the passerby in the image to be processed is called the non-target portrait.

[0194] The image processing method provided in the embodiments of this application is described below in conjunction with the above hardware and software structures.

[0195] In one embodiment, see Figure 5 This is a schematic flowchart of the image processing method provided in an embodiment of this application. It is intended as an example and not a limitation. Figure 5As shown, the image processing method may include the following steps:

[0196] S501, subject localization, detects the target human figure in the image to be processed.

[0197] In one implementation, a first model trained on the target image can be used to detect the target human figure in the image to be processed, and a detection box (first detection box) of the target human figure can be obtained.

[0198] Optionally, a first model can be pre-trained. For example, the process of training the first model may include: acquiring a large number of sample images and the labeling information of the target human image in each sample image; wherein the labeling information may include the coordinate information of the pixels in the image region to which the target human image belongs; inputting the sample images into the first model and outputting the detection box of the target human image; calculating a first loss value based on the labeling information and the detection box of the sample image; wherein the first loss value is used to represent the degree of difference between the coordinate information of the pixels in the detection box and the labeling information; if the first loss value is less than or equal to a preset threshold, then the training ends and the current first model is recorded as the trained first model; if the first loss value is greater than the preset threshold, then the model parameters of the first model are adjusted according to the first loss value, and the first model continues to be trained until the calculated first loss value is less than the preset threshold or the preset number of training iterations is reached.

[0199] In another implementation, a trained second model can be used to detect target human figures and non-realistic human figures in the image to be processed, obtaining detection boxes for the target human figure (first detection box) and non-realistic human figures. The second model is used to detect the target human figure and non-realistic human figures from the image. Non-realistic human figures can include portraits and human-shaped statues, etc.

[0200] Optionally, the second model can be based on the first model with the addition of a detection head. The first model can include a first module and a second module. The first module is used to extract image features from the image to be processed; the second module is used to detect the target human image based on the extracted image features, obtaining a first detection box. The second model can include the first module, the second module, and a third module, where the third module is used to detect non-realistic human images based on the extracted image features.

[0201] For example, the process of training the second model may include: acquiring a large number of sample images and the labeling information of the target human image in each sample image; wherein, the labeling information may include the first coordinate information of the pixels in the image region to which the target human image belongs and the second coordinate information of the pixels in the region to which the non-real human image belongs; inputting the sample images into the second model, and outputting the detection boxes of the target human image and the non-real human image; calculating a second loss value based on the first coordinate information in the labeling information of the sample image and the detection box of the target human image, and calculating a third loss value based on the second coordinate information in the labeling information and the detection box of the non-real human image; wherein, the second loss value is used to represent The third loss value represents the degree of difference between the coordinate information of the pixels in the detection box of the target human image and the first coordinate information. The third loss value is used to represent the degree of difference between the coordinate information of the pixels in the detection block of the non-real human image and the second coordinate information. The total loss is calculated based on the second loss value and the third loss value (e.g., the first loss value and the second loss value are weighted and summed). If the total loss is less than or equal to the preset threshold, the training ends and the current second model is recorded as the trained first model. If the total loss is greater than the preset threshold, the model parameters of the second model are adjusted according to the loss value, and the second model continues to be trained until the calculated total loss is less than the preset threshold or the preset number of training times is reached.

[0202] This method can detect non-real human figures in the image to be processed, providing a basis for subsequent filtering out of non-real human figures. It helps to reduce the situation of misidentifying non-real human figures as passersby, thereby helping to improve the accuracy of passerby removal.

[0203] S502, detect the first human figure in the image to be processed.

[0204] The first portrait includes both the target portrait and non-target portraits.

[0205] In one implementation, the image to be processed can first be input into a trained third model to detect the location of the person, resulting in a person detection box (second detection box). The third model is used to detect the person from the image.

[0206] It is understandable that the first detection box refers to the detection box obtained by detecting the target human image in the image to be processed, while the second detection box refers to the detection box obtained by detecting all human images in the image to be processed.

[0207] Optionally, the third model can be a human face detection model.

[0208] It should be noted that the training process of the third model is based on the same principle as that of the first model. For details, please refer to the training process of the first model, which will not be repeated here.

[0209] In some embodiments, S502 can be executed first. If the first human image is not detected, it means that there is no human image in the image to be processed, and S501 does not need to be executed. If the first human image is detected, then S501 is executed. In this way, when there is no human image in the image to be processed, the computational load of the algorithm can be reduced, thereby improving the efficiency of intelligent removal.

[0210] S503, determine whether the number of the first detection box is greater than 0, that is, determine whether the number of the main character box is greater than 0.

[0211] The first detection box is the detection box for the target human image.

[0212] If the number of the first detection boxes is 0, execute S504; if the number of the first detection boxes is greater than 0, execute S506.

[0213] S504, determine whether the number of the second detection box is greater than 0, that is, determine whether the number of human figures in the image to be processed is greater than 0.

[0214] The second detection box is the detection box for the first portrait.

[0215] If the number of the second detection boxes is equal to 0, it is determined that there are no pedestrians in the image to be processed.

[0216] S505, if the number of second detection boxes is greater than 0, it means that the images in the second detection boxes are all passersby. Then, the first image is segmented from the image to be processed based on the second detection boxes to obtain the segmentation map mask1. Non-target images in the image to be processed are eliminated based on the segmentation map mask1.

[0217] In one implementation, the second detection box and the image to be processed can be input into a fourth model, which outputs a segmentation map mask1. The fourth model is used to segment the human figure within the second detection box from the image to be processed.

[0218] For example, the process of training the fourth model may include: acquiring a large number of sample images and the detection boxes of the target human image in each sample image; inputting the sample images into the third model and outputting the segmentation map of the target human image; calculating the fourth loss value based on the detection boxes and segmentation map of the sample images; wherein, the fourth loss value is used to represent the degree of difference between the detection boxes and the segmentation map; if the fourth loss value is less than or equal to a preset threshold, the training ends and the current fourth model is recorded as the trained fourth model; if the fourth loss value is greater than the preset threshold, the model parameters of the fourth model are adjusted according to the fourth loss value, and the fourth model continues to be trained until the calculated fourth loss value is less than the preset threshold or the preset number of training iterations is reached.

[0219] The segmentation image can be a mask image (or mask map). The mask image is used to cover or selectively process specific regions of an image. Typically, the mask image is a binary or Boolean image of the same size as the original image, where the pixel values ​​of pixels in the specific region are different from the pixel values ​​of pixels in the non-specific region.

[0220] For example, see Figure 6 This is a schematic diagram of the mask image provided in the embodiments of this application. It is intended as an example and not a limitation. Figure 6 Image (a) shows the original image, as shown below. Figure 6 (b) in the image is the mask image of people 601 and 602 in the image to be processed. Figure 6 Image (c) shows the mask image of person 601 in the image to be processed. Figure 6 As shown, the mask image can be a black and white image, where the pixel value of the pixels in the black area is 0, and the pixel value of the pixels in the white area is 255.

[0221] Optionally, in some cases, the pixel value of pixels in the black area of ​​the mask image can be 0, and the pixel value of pixels in the white area can also be 1.

[0222] It is understandable that the image area of ​​the object to be removed corresponds to a black area in the mask image, and the corresponding image area that does not need to be removed corresponds to a white area in the mask image. Alternatively, the image area of ​​the object to be removed can also correspond to a white area in the mask image, and the corresponding image area that does not need to be removed corresponds to a black area in the mask image.

[0223] like Figure 6 In the example shown, if Figure 6 In the image (a), person 601 is the main character, and person 602 is a passerby. Figure 6 The mask image shown in (b) is a segmentation diagram of the first portrait. Figure 6 The mask image shown in (c) is a segmentation image of the target human image.

[0224] In one implementation, if the black area in the mask image is the area to be eliminated, during the elimination process, the mask image and its corresponding original image can be multiplied element-wise to eliminate the local image in the original image corresponding to the black area in the mask image.

[0225] In another implementation, if the white area in the mask image is the area to be eliminated, during the elimination process, the mask image and its corresponding original image can be multiplied element-wise to obtain the local image in the original image corresponding to the white area in the mask image, and then the local image is eliminated in the original image.

[0226] In one implementation, if a detection box for a non-realistic human image is detected in S501, the segmentation map mask1 can be checked for correctness based on the detection box. Specifically, it checks whether the region to be eliminated in the segmentation map mask1 overlaps with the detection box for the non-realistic human image; if they overlap, it indicates that the human image corresponding to the region to be eliminated is a non-realistic human image, and the region to be eliminated is then deleted.

[0227] It is understandable that the region to be eliminated in the segmentation image mask1 coincides with the detection box of the non-realistic human image. Alternatively, it can mean that the proportion of the overlapping area of ​​the region to be eliminated in the segmentation image mask1 and the detection box of the non-realistic human image to the total area of ​​the region to be eliminated is greater than a preset proportion.

[0228] S506, if the number of the first detection boxes is greater than 0, then determine whether the number of the first detection boxes is 1.

[0229] If the number of the first detection boxes is 1, then execute S507; if the number of the first detection boxes is greater than 1, then execute S516.

[0230] S507, if the number of first detection boxes is 1, then determine whether there is a second detection box whose overlapping area with the first detection box reaches the first ratio.

[0231] Understandably, if the area of ​​the second detection box is greater than half the area of ​​the first detection box, then the person in the second detection box is considered as the target person in the first detection box (i.e., the main character).

[0232] In one implementation, if there is an overlapping area between the first detection box and the second detection box, then a second detection box is determined to exist with an overlapping area of ​​a first proportion with the first detection box. However, this method is prone to classifying all images within the first detection box as meeting the criteria. If a passerby is present in the first detection box, the passerby will be identified as the main character, and may be retained subsequently, thus affecting the elimination effect.

[0233] In another implementation, if the area of ​​the overlapping region between the first and second detection boxes accounts for a first proportion (such as 0.5, 0.6, 0.8, etc.) of the total area of ​​the first detection box, then it is determined that a second detection box exists with an overlapping area of ​​the first detection box that also accounts for the first proportion. This implementation helps to reduce false positives.

[0234] S508, if there is no second detection box whose overlapping area with the first detection box reaches the first ratio, then perform subject segmentation on the local image corresponding to the first detection box to obtain the segmentation map mask2.

[0235] It should be noted that subject segmentation is the process of separating the subject from the background in an image. In the embodiments of this application, the subject object within the first detection box can be understood as the target human image. Therefore, by performing subject segmentation on the local image within the first detection box, the target human image within the first detection box can be obtained.

[0236] In one implementation, a first detection bounding box and the image to be processed can be input into a fifth model, which outputs a segmentation map mask2. The fifth model is used to segment the main object within the first detection bounding box from the image to be processed.

[0237] It should be noted that the training process of the fifth model is based on the same principle as that of the fourth model. For details, please refer to the training process of the fourth model, which will not be repeated here.

[0238] Understandably, the segmentation image mask2 only includes the target person image and not non-target person images.

[0239] S509, determine whether the number of human figures in the segmentation image mask2 (i.e. the number of human figures in the main character frame) is less than the number of human figures in the second detection frame (i.e. the total number of human figures in the image to be processed).

[0240] As described in S508, if there is no second detection box whose overlapping area with the first detection box reaches the first ratio, it cannot be determined whether there is a target human image in the first detection box, so further judgment is required.

[0241] If the number of human figures in the segmentation image mask2 is equal to the number of human figures in the second detection box, then it is determined that there are no pedestrians in the image to be processed.

[0242] S510, if the number of human images in the segmentation image mask2 is less than the number of human images in the second detection box, then the segmentation image mask1 of the first human image is segmented from the image to be processed according to the second detection box.

[0243] Step S510 is the same as step S505, and the details can be found in the description of the embodiment in S505.

[0244] S511, obtain the segmentation image mask2 from step S508.

[0245] S512, determine the segmentation map mask3 for non-target human figures in the image to be processed based on segmentation map mask2 and segmentation map mask1. Eliminate non-target human figures in the image to be processed based on segmentation map mask3.

[0246] In one implementation, the segmentation image mask2 and segmentation image mask1 are subtracted pixel by pixel to obtain the segmentation image mask3 of the non-target human image in the image to be processed.

[0247] The method of eliminating non-target human figures in the image to be processed based on segmentation map mask3 is the same as the principle of eliminating non-target human figures in the image to be processed based on segmentation map mask1. For details, please refer to the description in embodiment S505, which will not be repeated here.

[0248] In some cases, when there are many target human figures in the image to be processed, the execution result of S501 may output a detection box containing multiple target human figures, that is, a first detection box includes multiple target human figures. In this case, because the area of ​​the first detection box is relatively large, some smaller non-target human figures may be included in the box.

[0249] For example, see Figure 7 This is a schematic diagram illustrating the detection results of the target human image provided in an embodiment of this application. Figure 7 In the image to be processed shown, because the non-target human image 702 is small and the area of ​​the first detection box 701 is large, the non-target human image 703 is included within the first detection box 701. In other words, the first detection box 701 includes two target human images 702 and one non-target human image 703. In this case, if all human images in the image region corresponding to the first detection box are segmented, the non-target human image 703 may not be eliminated subsequently, thus affecting the elimination effect.

[0250] To address the aforementioned issues, in this embodiment of the application, steps S507-S512 enable the identification of non-target human figures from the first detection frame and the effective filtering of non-target human figures, thereby helping to improve the effectiveness of intelligent removal.

[0251] S513, if there is a second detection box whose overlapping area with the first detection box reaches the first ratio, then the first portrait is segmented from the image to be processed according to the second detection box to obtain the segmentation map mask1.

[0252] If there is a second detection box whose overlapping area with the first detection box reaches a first ratio, it can be determined that there is a target human image in the first detection box, and then image segmentation can be performed.

[0253] Step S513 is the same as step S505, and the details can be found in the description of the embodiment in S505.

[0254] S514, the human figure with the largest area is segmented from the local image of the first detection box to obtain the segmentation map mask4.

[0255] If there is a second detection frame whose overlapping area with the first detection frame reaches a first ratio, then the human image with the largest area in the first detection frame can be considered as the target human image.

[0256] In one implementation, the first detection box can be input into the trained sixth model, which outputs a segmentation map mask4. The sixth model is used to segment the target human image within the first detection box from the image to be processed.

[0257] It should be noted that the training process of the sixth model is based on the same principle as that of the fourth model. For details, please refer to the training process of the fourth model, which will not be repeated here.

[0258] S515, determine the segmentation map mask5 for non-target human figures in the image to be processed based on segmentation map mask1 and segmentation map mask4. Eliminate non-target human figures in the image to be processed based on segmentation map mask5.

[0259] Step S515 is the same in principle as step S512, and the specific details can be found in the description of the embodiment of S512.

[0260] S516, if the number of the first detection boxes is greater than 1, then determine whether the number of the first detection boxes is less than the number of the second detection boxes.

[0261] If the number of first detection boxes is greater than or equal to the number of second detection boxes, it means that all the human images in the second detection boxes are target human images, and it is determined that there are no non-target human images in the image to be processed.

[0262] S517, if the number of first detection boxes is less than the number of second detection boxes, it means that there is a non-target human image in the image to be processed. Then, the first human image is segmented from the image to be processed according to the second detection boxes to obtain the segmentation map mask1.

[0263] S518, the largest human figure is segmented from the local image corresponding to the first detection box to obtain the segmentation map mask6.

[0264] It should be noted that the "largest" portrait here can be the portrait with the largest area. For example, the area of ​​the smallest bounding rectangle of the portrait accounts for a certain proportion of the area in the first detection box; or, the proportion of the number of pixels of the portrait to the total number of pixels in the first detection box accounts for a certain proportion.

[0265] S519, determine the segmentation map mask7 for non-target human figures in the image to be processed based on segmentation map mask1 and segmentation map mask4. Eliminate non-target human figures in the image to be processed based on segmentation map mask7.

[0266] Steps S517-S519 are the same as steps S513-S515, and can be found in the description of steps S513-S515 in the embodiments.

[0267] It should be noted that, Figure 5In this embodiment, during processes S516-S519, the main character mask 6 can be recorded as the first segmentation image, the portrait mask 1 as the second segmentation image, and the mask 7 as the third segmentation image; during processes S513-S515, the mask 4 can be recorded as the first segmentation image, the portrait mask 1 as the second segmentation image, and the mask 5 as the third segmentation image. During processes S510-S511, the main subject cutout mask 2 can be recorded as the fourth segmentation image, and the mask 3 as the fifth segmentation image.

[0268] Figure 5 In the aforementioned embodiment, it is equivalent to detecting the main character and all other characters in the image to be processed. Removing the main character from all other characters identifies passersby in the image. Directly detecting passersby is difficult, especially when there are many passersby, making accurate detection impossible. However, the method described in this embodiment indirectly identifies passersby by detecting the main character, more accurately segmenting the main character and the object to be eliminated. This helps improve the effectiveness of intelligent elimination and enhances the user experience of the intelligent elimination function. Furthermore, Figure 5 In this embodiment, the target human figure in the image to be processed can be further determined based on the number of detection boxes for the main character and the number of detection boxes for all human figures. This helps to reduce the situation of misidentifying passersby as main characters or vice versa, thereby improving the effect of intelligent elimination.

[0269] In another embodiment, see Figure 8 This is a schematic flowchart of the image processing method provided in an embodiment of this application. It is intended as an example and not a limitation. Figure 8 As shown, the image processing method may include the following steps:

[0270] S801, subject localization, detects target human figures in the image to be processed.

[0271] S802, the first human image is segmented from the image to be processed to obtain the segmentation image mask1.

[0272] In one implementation, a segmentation map mask1 can be obtained through a trained seventh model, where the seventh model segments the human figure from the image. Specifically, the image to be processed is input into the seventh model, and the output is the segmentation map mask1.

[0273] Understandably, the difference between the seventh model and the fourth model in S505 is that the fourth model is only used to segment the human figures in the second detection box, while the seventh model can segment all human figures in the image to be processed.

[0274] Optionally, the seventh model can be an image segmentation model, which can employ semantic segmentation or instance segmentation models, such as the SAM model. Semantic segmentation predicts the category of each pixel in the image, while instance segmentation distinguishes different instances based on the predicted pixel category. Therefore, using an instance segmentation model helps obtain more accurate image segmentation results.

[0275] Optionally, a seventh model can be pre-trained. For example, the process of training the seventh model may include: acquiring a large number of sample images and the labeling information of the target human image in each sample image; wherein, the labeling information may include the coordinate information of the pixels in the image region to which the target human image belongs; inputting the sample images into the seventh model and outputting a segmentation map of the target human image; calculating a fifth loss value based on the labeling information and the segmentation map of the sample images; wherein, the fifth loss value is used to represent the degree of difference between the coordinate information of the pixels in the detection box and the labeling information; if the fifth loss value is less than or equal to a preset threshold, then the training ends, and the current seventh model is recorded as the trained seventh model; if the fifth loss value is greater than the preset threshold, then the model parameters of the seventh model are adjusted according to the fifth loss value, and the seventh model continues to be trained until the calculated fifth loss value is less than the preset threshold or the preset number of training iterations is reached.

[0276] S803, determine whether the number of the first detection box is greater than 0.

[0277] S804, determine whether the number of human figures in the image to be processed is greater than 0.

[0278] It should be noted that, unlike S504, S804 can identify the number of human figures in the segmentation image mask1, which is recorded as the number of human figures in the image to be processed.

[0279] S805, if the number of the second detection boxes is greater than 0, it means that the people in the second detection boxes are all passersby, then obtain the segmentation map mask1. Eliminate non-target people in the image to be processed based on the segmentation map mask1.

[0280] Unlike S505, since the human portrait in the image to be processed has already been segmented in S802, the segmentation result in S805, namely the segmentation image mask1, can be directly obtained.

[0281] In one implementation, S805 can refine the segmentation image mask1 based on the segmentation result of S802, thereby improving the accuracy of portrait segmentation. In other words, S802 can perform coarse-grained portrait segmentation processing on the image to be processed, while S805 can perform fine-grained portrait segmentation processing on the image to be processed.

[0282] The scheme for refining the segmentation image mask1 can be found in the description of the following embodiments.

[0283] S806, if the number of the first detection boxes is greater than 0, then determine whether the number of the first detection boxes is 1.

[0284] If the number of the first detection boxes is 1, then execute S807; if the number of the first detection boxes is greater than 1, then execute S816.

[0285] S807, if the number of first detection frames is 1, then determine whether there is a second detection frame whose overlapping area with the first detection frame reaches a preset ratio.

[0286] S808, if there is no second detection box whose overlapping area with the first detection box reaches the first ratio, then perform subject segmentation on the local image corresponding to the first detection box to obtain the segmentation map mask2.

[0287] S809, determine whether the number of human figures in the segmentation image mask2 (i.e. the number of human figures in the main character frame) is less than the number of human figures in the second detection frame (i.e. the total number of human figures in the image to be processed).

[0288] If the number of human figures in the segmentation image mask2 is equal to the number of human figures in the second detection box, then it is determined that there are no pedestrians in the image to be processed.

[0289] S810, if the number of human images in the segmentation image mask2 is less than the number of human images in the second detection box, then the segmentation image mask1 of the first human image is segmented from the image to be processed according to the second detection box.

[0290] S811, obtain the segmentation image mask2 from step S808.

[0291] S812, determine the segmentation map mask3 (fifth segmentation map) for non-target human figures in the image to be processed based on segmentation map mask2 and segmentation map mask1. Eliminate non-target human figures in the image to be processed based on segmentation map mask3.

[0292] S813, if there is a second detection box whose overlapping area with the first detection box reaches the first ratio, then the first portrait is segmented from the image to be processed according to the second detection box to obtain the segmentation map mask1.

[0293] S814, the human figure with the largest area is segmented from the local image of the first detection box to obtain the segmentation map mask4.

[0294] S815, determine the segmentation map mask5 for non-target human figures in the image to be processed based on segmentation map mask1 and segmentation map mask4. Eliminate non-target human figures in the image to be processed based on segmentation map mask5.

[0295] S816, if the number of the first detection boxes is greater than 1, then determine whether the number of the first detection boxes is less than the number of the second detection boxes.

[0296] If the number of first detection boxes is greater than or equal to the number of second detection boxes, it means that all the human images in the second detection boxes are target human images, and it is determined that there are no non-target human images in the image to be processed.

[0297] S817, if the number of first detection boxes is less than the number of second detection boxes, it means that there is a non-target human image in the image to be processed. Then, the first human image is segmented from the image to be processed according to the second detection boxes to obtain the segmentation map mask1.

[0298] S818, based on the first detection box, obtain the segmentation map of the target human image from the segmentation map mask1 to obtain the segmentation map mask6.

[0299] Understandably, since portrait segmentation processing has already been performed on the image to be processed in S802, i.e., the segmentation results of all portraits in the image to be processed are obtained, in this case, the part corresponding to the first detection box can be obtained from the segmentation map mask1 in S818 to obtain the segmentation map of the main character. For example, the overlapping area between the first detection box and each segmentation instance in the segmentation map mask1 can be calculated, and the segmentation instance corresponding to the largest overlapping area can be taken as the segmentation map mask6 corresponding to the first detection box.

[0300] S819, determine the segmentation map mask7 for non-target human figures in the image to be processed based on segmentation map mask1 and segmentation map mask6. Eliminate non-target human figures in the image to be processed based on segmentation map mask7.

[0301] Steps S806-S819 are in the same principle as steps S506-S519. For details, please refer to the description in the embodiment of steps S506-S519, which will not be repeated here.

[0302] It should be noted that, Figure 8 In this embodiment, during processes S816-S819, the main character mask 6 can be recorded as the first segmentation image, the portrait mask 1 as the second segmentation image, and the mask 7 as the third segmentation image; during processes S813-S815, the mask 4 can be recorded as the first segmentation image, the portrait mask 1 as the second segmentation image, and the mask 5 as the third segmentation image. During processes S810-S811, the main subject cutout mask 2 can be recorded as the fourth segmentation image, and the mask 3 as the fifth segmentation image.

[0303] Figure 8In the aforementioned embodiment, it is equivalent to detecting the main character and all other characters in the image to be processed. Removing the main character from all other characters identifies passersby in the image. Directly detecting passersby is difficult, especially when there are many passersby, making accurate detection impossible. However, the method described in this embodiment indirectly identifies passersby by detecting the main character, more accurately segmenting the main character and the object to be eliminated. This helps improve the effectiveness of intelligent elimination and enhances the user experience of the intelligent elimination function. Furthermore, Figure 8 In this embodiment, the target human figure in the image to be processed can be further determined based on the number of main character detection boxes and the total number of human figures. This helps to reduce the situation of misidentifying passersby as main characters or vice versa, thereby improving the effect of intelligent elimination.

[0304] and Figure 5 Compared to the previous examples, Figure 8 The embodiment reduces the human face detection process and instead determines passersby in the image based on the results of human face segmentation and the detection results of the main character. This approach helps reduce the computational load of the algorithm, thereby improving the processing efficiency of intelligent removal.

[0305] In another embodiment, see Figure 9 This is a schematic flowchart of an image processing method provided in another embodiment of this application. It is intended as an example and not a limitation. Figure 9 As shown, the image processing method may include the following steps:

[0306] S801, subject localization, detects target human figures in the image to be processed.

[0307] S802, the first human image is segmented from the image to be processed to obtain the segmentation image mask1.

[0308] S901, detect the first human figure in the image to be processed and obtain the second detection box.

[0309] Step S901 is the same as step S502, and the details can be found in the description of the embodiment in S502, which will not be repeated here.

[0310] S902, perform deduplication of the human image based on the segmentation map mask1 and the second detection box to obtain the deduplicated third detection box, and update the segmentation map mask1.

[0311] It should be noted that, in this embodiment, the segmentation image mask1 before the update (output result of S802) can be denoted as the sixth segmentation image, and the segmentation image mask1 after the update (output result of S902) can be denoted as the second segmentation image.

[0312] Optionally, deduplication of human images may include: calculating the overlap between the second detection box and the human image region (e.g., the white region) in the segmentation image mask1; if the overlap is greater than a preset overlap, it means that the human image in the second detection box and the human image region in the segmentation image mask1 are the same human image, and the second detection box is deleted; if the overlap is less than or equal to the preset overlap, the second detection box is retained; and the remaining second detection box is determined as the third detection box after deduplication.

[0313] This approach is equivalent to determining the location of a person based on both the segmentation and detection results. This reduces the impact of detection errors from a single model on subsequent removal results. Especially for small target portraits, if a single model fails to detect them, another model can supplement the detection results, thereby improving the detection accuracy of the portrait and enhancing the removal effect.

[0314] Updating the segmentation map mask1 can be achieved by inputting the third detection box and the image to be processed into the trained fourth model, and outputting the updated segmentation map mask1.

[0315] After S902, execute S803-S819.

[0316] It should be noted that, in this embodiment, when S805 is executed, the obtained segmentation image mask1 is the updated mask1 in S902. In this embodiment, when S804 is executed, it is based on whether the number of third detection boxes is greater than 0. In this embodiment, when S807 is executed, it is determined whether there is a third detection box that overlaps with the first detection box. In this embodiment, when S816 is executed, it is determined whether the number of first detection boxes is less than the number of third detection boxes.

[0317] Figure 9 In the aforementioned embodiment, it is equivalent to detecting the main character and all other characters in the image to be processed. Removing the main character from all other characters identifies passersby in the image. Directly detecting passersby is difficult, especially when there are many passersby, making accurate detection impossible. However, the method described in this embodiment indirectly identifies passersby by detecting the main character, more accurately segmenting the main character and the object to be eliminated. This helps improve the effectiveness of intelligent elimination and enhances the user experience of the intelligent elimination function. Furthermore, Figure 8 In this embodiment, the target human figure in the image to be processed can be further determined based on the number of main character detection boxes and the total number of human figures. This helps to reduce the situation of misidentifying passersby as main characters or vice versa, thereby improving the effect of intelligent elimination.

[0318] and Figure 5 and Figure 8 Compared to the previous examples, Figure 9In this embodiment, the position of the human face is determined by combining the segmentation and detection results. This reduces the impact of detection errors from a single model on subsequent elimination results. Especially for small target human faces, if a single model fails to detect a small target human face, another model can supplement the detection results, thereby improving the detection accuracy of the human face and helping to improve the elimination effect.

[0319] In some embodiments, such as Figure 5 In the processing flow, if the execution result of S501 outputs only one first detection box regardless of how many main characters are detected, meaning that the first detection box may include multiple main characters, then S506-S507 and S513-S519 can be deleted, while S508-S512 can be retained. That is, the solution is: if the number of first detection boxes is 0, then execute S504-S505; if the number of first detection boxes is not 0, then execute S508-S512.

[0320] Similarly, such as Figure 8 or Figure 9 In the processing flow, if the execution result of S801 outputs only one first detection box regardless of how many main characters are detected, meaning that the first detection box may include multiple main characters, then S806-S807 and S813-S819 can be deleted, while S808-S812 can be retained. That is, the solution is: if the number of first detection boxes is 0, then execute S804-S805; if the number of first detection boxes is not 0, then execute S808-S812.

[0321] In other words, if the result of the main character positioning only outputs a first detection box, the target portrait is determined by default through the process of subject cutout on the first detection box.

[0322] In other embodiments, such as Figure 5 In the processing flow, if the execution result of S501 outputs the first detection box corresponding to each protagonist, that is, the first detection box includes one protagonist, then S506-S515 can be deleted, and S516-S519 can be retained. That is, the solution is: if the number of first detection boxes is 0, then execute S504-S505; if the number of first detection boxes is not 0, then execute S516-S519.

[0323] Similarly, such as Figure 8 or Figure 9 In the processing flow, if the execution result of S801 outputs the first detection box corresponding to each protagonist, that is, the first detection box includes one protagonist, then S806-S815 can be deleted, and S816-S819 can be retained. That is, the solution is: if the number of first detection boxes is 0, then execute S804-S805; if the number of first detection boxes is not 0, then execute S816-S819.

[0324] In other words, if the result of the protagonist localization outputs a corresponding first detection box for each detected protagonist, the target portrait confirmation process does not need to be executed by default.

[0325] In some embodiments, shadow segmentation can be performed after obtaining the segmentation map of the image to be removed.

[0326] In one implementation, the segmentation map of the image to be removed can be input into the trained eighth model, which outputs a shadow segmentation map.

[0327] For example, see Figure 10 This is a schematic diagram of the shadow segmentation map provided in an embodiment of this application. For example... Figure 10 As shown in (a) of the image, the image to be processed includes passerby 1001 and the shadow 1002 of passerby 1001. The segmentation image of passerby 1001 segmented from the image to be processed is shown below. Figure 10 As shown in (b) of the image, the segmentation diagram includes the area to be eliminated (white area) for pedestrian 1001. Figure 10 The segmentation map shown in (b) and the image to be processed are input into the trained eighth model, and the output is as follows: Figure 10 The shadow segmentation diagram shown in (c) includes the area to be eliminated for pedestrian 1001 and the area to be eliminated for the shadow 1002 corresponding to pedestrian 1001.

[0328] Optionally, an eighth model can be pre-trained. For example, the process of training the eighth model may include: acquiring a large number of sample images and the labeling information of each sample image; wherein, the labeling information may include the coordinate information of pixels in the portrait region and the coordinate information of pixels in the shadow region corresponding to the portrait; inputting the sample images into the eighth model and outputting a shadow segmentation map; calculating a sixth loss value based on the labeling information of the sample images and the shadow segmentation map; wherein, the sixth loss value is used to represent the degree of difference between the coordinate information of pixels in the shadow segmentation map and the labeling information; if the sixth loss value is less than or equal to a preset threshold, then training ends, and the current eighth model is recorded as the trained eighth model; if the sixth loss value is greater than the preset threshold, then the model parameters of the eighth model are adjusted according to the sixth loss value, and the eighth model continues to be trained until the calculated sixth loss value is less than the preset threshold or the preset number of training iterations is reached.

[0329] Through the above embodiments, when there are shadows around a pedestrian, the mask image of the pedestrian's shadow can be effectively detected, and the pedestrian's shadow can also be eliminated in the subsequent elimination process, thereby improving the effect of intelligent elimination.

[0330] In another implementation, after obtaining the shadow segmentation map, avoidance of the target portrait region can be performed. Specifically, it is determined whether there is regional overlap between the shadow segmentation map and the target portrait segmentation map; if there is regional overlap, the part in the first region that overlaps with the second region is deleted, resulting in the processed shadow segmentation map. Here, the first region is the shadow to be eliminated in the shadow segmentation map, and the second region is the target portrait to be eliminated in the target portrait segmentation map.

[0331] It should be noted that the above processing can be performed on the segmentation map of the human face, such as on the segmentation map mask1 obtained in S505.

[0332] For example, see Figure 11 This is a schematic diagram illustrating target human image area avoidance provided in an embodiment of this application. It is intended as an example and not a limitation. Figure 11 The shadow segmentation diagram shown in (a) includes the region of the target person 1102, the region of the passerby 1103, and the region of the shadow 1101 of the passerby 1103. It can be seen that the region of shadow 1101 overlaps with the region of the target person 1102. Therefore, deleting the overlapping portion of the region of shadow 1101 with the region of the target person 1102 yields the following result: Figure 11 The shaded segmentation diagram shown in (b) is as follows. Figure 11 In the shadow segmentation diagram shown in (b), the area of ​​shadow 1101 has changed and no longer overlaps with the area of ​​target portrait 1102.

[0333] Through the above embodiments, when obtaining the segmentation map of the pedestrian's shadow, the target person's image can be effectively avoided. In the subsequent elimination process, the area of ​​the target person's image can be avoided, thereby improving the effect of intelligent elimination.

[0334] The following section introduces methods for portrait segmentation.

[0335] It should be noted that the process of segmenting the human image from the image to be processed in steps S505, S510, S513, S517 and steps S802, S805, S810, S813, S817 can all adopt the human image segmentation method in the following embodiments.

[0336] In some application scenarios, the attached information of a person's image, such as handbags, mobile phones, shadows, etc., may not be detected by the image segmentation model because they are not part of the human body. This may result in the elimination of the person but not their attachments, thus affecting the effectiveness of intelligent elimination.

[0337] Based on this, the portrait segmentation method in the embodiments of this application can provide attachment detection.

[0338] In one embodiment, see Figure 12This is a schematic diagram of a portrait segmentation model provided in an embodiment of this application. It is intended as an example and not a limitation. Figure 12 As shown, the human image segmentation model can include a dot matrix model, a backbone network, a first head network, and a second head network.

[0339] The process involves several key components: a dot-mapping model for identifying human figures in the image and marking them with dots; a backbone network for extracting image features; a first head network for segmenting the human figure based on the extracted features; and a second head network for segmenting the appendages of the human figure based on the extracted features.

[0340] Optionally, the dot model can adopt an instance segmentation model.

[0341] based on Figure 12 The portrait segmentation model shown can be described using a portrait segmentation method that includes the following steps:

[0342] S1201, the backbone network extracts image features from the image to be processed.

[0343] S1202, the dot-mapping model identifies human figures in the image to be processed and marks the human figure region to obtain the detection points of the human figure.

[0344] Optionally, the center point of the human image area can be used as the detection point.

[0345] Optionally, multiple points within the human image area can be selected as detection points. Understandably, the more detection points there are, the more accurate the detection of attachments will be, but the computational load may be greater.

[0346] It should be noted that in some other implementations, detection boxes of the human figures in the image to be processed can be obtained, and the human figures and their surrounding objects can be segmented based on the detection boxes. Compared with this approach, the use of detection points in this embodiment helps to distinguish different human figures, especially when there are many human figures in the image and their positions are close together. Using detection points can more clearly distinguish different human figures, thereby reducing the situation where the segmentation effect is poor due to overlapping detection boxes.

[0347] S1203, the first head network segments the human image from the image to be processed based on image features and detection points, and obtains the segmentation map of the human image.

[0348] S1204, the second head network segments the human figure's appendages from the image to be processed based on image features and detection points, and obtains a segmentation map of the appendages.

[0349] S1205, superimpose the segmentation map of the human figure and the segmentation map of the appendages to obtain the final segmentation map.

[0350] For example, see Figure 13 This is a schematic diagram of appendage detection provided in an embodiment of this application. It is intended as an example and not a limitation. Figure 13 As shown, inputting the image to be processed 1301 into the backbone network yields image features. Inputting the image to be processed 1301 into the dot matrix model yields detection points 1302. The first head network segments the human figure from the image to be processed based on the image features and the positions of the detection points 1302, resulting in segmentation image 1303. The second head network segments the appendages of the human figure from the image to be processed based on the image features and the positions of the detection points 1302, resulting in segmentation image 1304. Then, segmentation images 1303 and 1304 are superimposed to obtain the final segmentation image 1305. Figure 13 As can be seen, segmentation diagram 1305 includes not only the area of ​​the human figure, but also the area of ​​the appendages.

[0351] It should be noted that traditional image segmentation models can include a backbone network and a first head network. Compared to traditional image segmentation models, the portrait segmentation model in this application adds a dot-matrix model and a second head network, such as... Figure 12 The portion shown is within the dashed box. The portrait segmentation model of this application can effectively detect objects surrounding a person, enabling the subsequent removal process to also eliminate these objects, thereby improving the effectiveness of intelligent removal.

[0352] The attachment detection model can be pre-trained. In practical applications, the image to be processed is input into the trained attachment detection model, which outputs the final segmentation map.

[0353] In one implementation, the entire human face segmentation model can be trained.

[0354] In another implementation, training can be performed on a traditional image segmentation model. That is, a pre-trained traditional image segmentation model (i.e., only including the backbone network and the first head network) is obtained, and then the portrait segmentation model of this application embodiment (i.e., adding a dot matrix model and a second head network) is constructed on the pre-trained traditional portrait segmentation model, and the newly added part of the portrait segmentation model is trained.

[0355] Alternatively, the traditional image segmentation model can employ the SAM model.

[0356] For example, the training process for adding a new part in the portrait segmentation model may include: acquiring a large number of sample images and the labeling information of the portrait in each sample image; wherein, the labeling information may include the coordinate information of multiple first position points in the sample image; inputting the sample images into the portrait segmentation model and outputting a segmentation map of the samples, which includes multiple second position points; calculating the positional difference between the first position points and the second position points to obtain a seventh loss value; if the seventh loss value is less than or equal to a preset threshold, the training ends and the current portrait segmentation model is recorded as the trained portrait segmentation model; if the seventh loss value is greater than the preset threshold, the model parameters of the second head network are adjusted according to the seventh loss value, and the portrait segmentation model continues to be trained until the calculated seventh loss value is less than the preset threshold or the preset number of training iterations is reached.

[0357] In the training method described above, only the newly added part of the network used for segmenting attachments (i.e., the second head network) in the portrait segmentation model is trained. Compared with training the entire network, this method can significantly reduce training time. Furthermore, in this training method, the loss value is calculated based on the positional differences of points in the image. Compared with calculating the loss value based on the mask image, this not only significantly reduces the computational load but also lowers the learning difficulty, facilitating rapid model training.

[0358] In some application scenarios, sculptures, portraits, and other non-realistic human images may be detected as real human images by the image segmentation model. This may lead to the subsequent removal of these non-realistic human images, thus affecting the effectiveness of intelligent removal.

[0359] Based on this, the portrait segmentation method in the embodiments of this application can provide non-realistic portrait detection.

[0360] In one embodiment, see Figure 14 This is a schematic diagram of a portrait segmentation model provided in an embodiment of this application. It is intended as an example and not a limitation. Figure 14 As shown, portrait segmentation models can include image segmentation models and semantic segmentation models.

[0361] Image segmentation models can employ traditional methods, such as instance segmentation models or SAM models. These models are used to segment human figures from the image being processed. Semantic segmentation models are used to perform semantic recognition on objects in the image to obtain labels (i.e., categories) for each object.

[0362] Optionally, the image segmentation model may include an instance segmentation model and a SAM model. The instance segmentation model is used to identify bounding boxes for human figures in the image to be processed. The SAM model is used to segment the human figure from the image to be processed based on the bounding boxes for human figures output by the instance segmentation model.

[0363] based on Figure 14 The portrait segmentation model shown can be described using a portrait segmentation method that includes the following steps:

[0364] S1401, the human image is segmented from the image to be processed using an image segmentation model to obtain an initial segmentation map.

[0365] S1402, using a semantic segmentation model to identify the label of each portrait in the image to be processed.

[0366] S1403, filter the portraits in the initial segmentation image according to the label of each portrait to obtain the final segmentation image.

[0367] In one implementation, for each portrait region in the initial segmentation image, the first number of pixels belonging to the first label in the portrait region is counted based on the semantic segmentation result; if the first number accounts for a preset proportion (first proportion) of the total number of pixels in the portrait region, then the portrait region is determined to belong to the first label; wherein, the first label can represent a non-real portrait; the portrait regions belonging to the first label in the initial segmentation image are deleted.

[0368] For example, see Figure 15 This is a schematic diagram of non-realistic human image detection provided in an embodiment of this application. It is intended as an example and not a limitation. Figure 15 As shown, the image to be processed 1501 includes both human portraits and images. Inputting image 1501 into an image segmentation model yields an initial segmentation map 1502. Inputting image 1501 into a semantic segmentation model yields a semantic segmentation result 1503. The semantic segmentation result 1503 is labeled with tags corresponding to each human portrait, such as "human portrait" and "image". Semantic filtering is performed on the initial segmentation map 1502 based on the tags of each human portrait, resulting in a final segmentation map 1504. The final segmentation map 1504 includes only human portraits and excludes images.

[0369] Both the image segmentation model and the semantic segmentation model can be pre-trained. In practical applications, the image to be processed is input into the trained image segmentation model and semantic segmentation model respectively, and the final segmentation map is output.

[0370] It is understandable that image segmentation models and semantic segmentation models can be trained separately. The training process for the image segmentation model can be referenced from the training process of the seventh model, and will not be repeated here.

[0371] For example, the training process of a semantic segmentation model may include: acquiring a large number of sample images and the real labels of each portrait in each sample image; inputting the sample images into the semantic segmentation model and outputting semantic segmentation results, which include the predicted labels of each portrait in the sample images; calculating the difference between the real labels and the predicted labels to obtain the eighth loss value; if the eighth loss value is less than or equal to a preset threshold, the training ends and the current semantic segmentation model is recorded as the trained semantic segmentation model; if the eighth loss value is greater than the preset threshold, the model parameters of the semantic segmentation model are adjusted according to the eighth loss value, and the semantic segmentation model continues to be trained until the calculated eighth loss value is less than the preset threshold or the preset number of training iterations is reached.

[0372] Compared to traditional image segmentation models, the portrait segmentation model in this application embodiment incorporates a semantic segmentation model, such as... Figure 14 The portion shown is within the dashed box. Semantic segmentation models can capture the semantic differences between real and non-real images. The semantic segmentation results are then used to assist image segmentation models, effectively filtering out non-real images and allowing them to be retained during subsequent removal processes, thus improving the effectiveness of intelligent removal.

[0373] In some application scenarios, small human figures are difficult to detect by image segmentation models, which may result in the inability to remove these small human figures in the subsequent process, thus affecting the effectiveness of intelligent removal.

[0374] Based on this, the portrait segmentation method in the embodiments of this application can provide small target portrait detection.

[0375] In one embodiment, see Figure 16 This is a schematic diagram of a portrait segmentation model provided in an embodiment of this application. It is intended as an example and not a limitation. Figure 16 As shown, the portrait segmentation model can include image segmentation model 1, image segmentation model 2, and small object detection model.

[0376] Image segmentation model 1 and image segmentation model 2 can employ traditional image segmentation models, such as instance segmentation models and SAM models. The image segmentation model is used to segment human figures from the image to be processed. The small target model is used to detect small human figures from the image to be processed.

[0377] Optionally, image segmentation model 1 may include instance segmentation model 1 and SAM1 model. Instance segmentation model 1 is used to identify bounding boxes for human figures in the image to be processed. SAM1 model is used to segment a human figure from the image to be processed based on the bounding boxes for human figures output by the instance segmentation model. Image segmentation model 2 may include SAM1 model, used to segment a human figure from the image to be processed based on the bounding boxes for human figures output by the small object detection model.

[0378] based on Figure 16 The portrait segmentation model shown can be described using a portrait segmentation method that includes the following steps:

[0379] S1601, the human image is segmented from the image to be processed using image segmentation model 1 to obtain the initial segmentation image.

[0380] S1602, the human figure in the image to be processed is identified by the small target model, and the human figure detection box is obtained.

[0381] S1603, filter the human detection boxes based on the initial segmentation image to obtain small target detection boxes (sixth detection box).

[0382] In one implementation, the human detection bounding box can be first size-filtered, and then the size-filtered human detection bounding box can be filtered according to the initial segmentation image to obtain the small target detection bounding box.

[0383] Optionally, the size filtering process may include: calculating the area ratio of the human detection bounding box in the image to be processed; if the area ratio is greater than a preset ratio (third ratio), then the human detection bounding box is deleted. In other words, only smaller human detection bounding boxes are retained, that is, only the detection boxes of small targets are retained.

[0384] In one implementation, the process of filtering the human detection boxes based on the initial segmentation image may include: the instance segmentation model in image segmentation model 1 can output human detection boxes; the human detection boxes output by the instance segmentation model (fourth detection box) are compared with the human detection boxes output by the small object detection model (fifth detection box); if the area ratio of the overlapping part of the fourth and fifth detection boxes in the fourth or fifth detection boxes reaches a preset ratio (second ratio), then the fifth detection box is deleted. This method can exclude already segmented human detection boxes, thereby reducing duplicate elimination and improving the effectiveness of intelligent elimination.

[0385] S1604, the human image corresponding to the small target detection box is segmented from the image to be processed by image segmentation model 2 to obtain the small target segmentation map.

[0386] S1605, the initial segmentation image and the small target segmentation image are superimposed to obtain the final segmentation image.

[0387] Both the image segmentation model and the small object detection model can be pre-trained. In practical applications, the image to be processed is input into the trained image segmentation model and the small object detection model respectively, and the final segmentation map is output.

[0388] The training method for the small object detection model can be found in the training method of the first model mentioned above, and will not be repeated here.

[0389] Understandably, the sample images used to train the small target model are images that include small target human figures.

[0390] For example, see Figure 17 This is a schematic diagram of small target detection provided in an embodiment of this application. It is intended as an example and not a limitation. Figure 17 As shown, the image to be processed 1701 includes a target human figure and small targets. Inputting the image to be processed 1701 into image segmentation model 1 yields detection map 1702 and initial segmentation map 1703. Detection map 1702 includes the detection bounding box 17021 of the target human figure but excludes the detection bounding box of the small targets. Correspondingly, image segmentation model 1 segments the image based on detection map 1702 to obtain initial segmentation map 1703, which includes only the target human figure and excludes small targets. Inputting the image to be processed 1701 into a small target detection model outputs detection map 1704, which includes the detection bounding box 17041 of the target human figure and the detection bounding box 17042 of the small targets. Size filtering is applied to the detection bounding boxes in detection map 1704 to obtain detection map 1705, which filters out the larger target human figure detection bounding box 17042. Comparing detection maps 1702 and 1705, no duplicate detection boxes are found, resulting in the filtered detection map 1706. Image segmentation model 2 uses the detection boxes in detection map 1706 to segment the target into a small target segmentation map 1707, which contains only small targets. Finally, segmentation map 1707 is overlaid with the initial segmentation map 1703 to obtain the final segmentation map 1708. The final segmentation map 1708 includes not only the target human image but also small targets.

[0391] Compared to traditional image segmentation models, the portrait segmentation model in this application embodiment adds a small object detection model, such as... Figure 16 The area shown is within the dashed box. Small targets in the image are detected using a small target detection model, enabling subsequent removal of these small targets and thus improving the effectiveness of intelligent image removal.

[0392] For example, see Figure 18 This is a schematic diagram of a portrait segmentation model provided in an embodiment of this application. It is intended as an example and not a limitation. Figure 18 As shown, human image segmentation models can include SAM1 model, SAM2 model, instance segmentation model, semantic segmentation model, and small object detection model.

[0393] Among them, the SAM1 model and the instance segmentation model constitute Figure 16 Image segmentation model 1 in the embodiment. The SAM2 model constitutes... Figure 16 Image segmentation model 2 in the embodiment.

[0394] based on Figure 18 The portrait segmentation model shown can be described using a portrait segmentation method that includes the following steps:

[0395] S1801: Input the image to be processed into the backbone network of the SAM1 model to obtain the image features in the image to be processed.

[0396] S1802, input the image to be processed into the instance segmentation model to obtain the detection points and detection boxes of the human image.

[0397] S1803, the first head network segments the human image from the image to be processed based on image features and detection points, and obtains the segmentation map of the human image.

[0398] S1804, the second head network segments the human figure's appendages from the image to be processed based on image features and detection points, and obtains a segmentation map of the appendages.

[0399] S1805, superimpose the segmentation map of the human figure and the segmentation map of the appendages to obtain the initial segmentation map.

[0400] Steps S1801-S1805 are the same as steps S1201-S1205, and for details, please refer to the description in the embodiment of steps S1201-S1205.

[0401] S1806: Input the image to be processed into the small target detection model to obtain the human image detection box.

[0402] S1807, Filter the human body detection box output by the small target detection model based on the detection box output by the instance segmentation model to obtain the small target detection box.

[0403] S1808 inputs small target detection boxes into the SAM2 model and outputs small target segmentation maps.

[0404] S1809, the small target segmentation map and the initial segmentation map are superimposed to obtain the superimposed segmentation map.

[0405] Steps S1806-S1809 are the same as steps S1602-S1605. For details, please refer to the description in the embodiment of steps S1602-S1605, which will not be repeated here.

[0406] S1810: Input the image to be processed into the semantic segmentation model to obtain the label of each portrait in the image to be processed.

[0407] S1811, filter the portraits in the overlay segmentation image according to the label of each portrait to obtain the final segmentation image.

[0408] Steps S1810-S1811 are the same as steps S1402-S1403. For details, please refer to the description in the embodiment of steps S1402-S1403, which will not be repeated here.

[0409] Figure 18 In the portrait segmentation method shown, by adding a second head network, the attachments of the portrait can be effectively detected, enabling the subsequent elimination process to also eliminate the attachments of passersby, thereby improving the effect of intelligent elimination. By adding a semantic segmentation model, the semantic difference features between real and non-real portraits can be captured. Then, the semantic segmentation results are used to assist the image segmentation model, which can effectively filter out non-real portraits, enabling the subsequent elimination process to retain non-real portraits, thereby improving the effect of intelligent elimination. By adding a small target detection model, small targets in the image can be detected, enabling the subsequent elimination of these small targets, thereby improving the effect of intelligent elimination.

[0410] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0411] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the above-described method embodiments.

[0412] This application also provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.

[0413] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to the first device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0414] This application also provides a chip system, which includes a processor coupled to a memory. The processor executes a computer program stored in the memory to implement the steps of any method embodiment of this application. The chip system can be a single chip or a chip module composed of multiple chips.

[0415] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0416] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0417] Finally, it should be noted that the above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image processing method, characterized in that, include: The target portrait and the first portrait are segmented from the image to be processed to obtain a first segmentation image of the target portrait and a second segmentation image of the first portrait; wherein, the first portrait includes the target portrait and non-target portraits; A third segmentation map of the non-target human figure in the image to be processed is determined based on the first segmentation map and the second segmentation map; Eliminate the non-target human image in the image to be processed based on the third segmentation map; The step of eliminating the non-target human image in the image to be processed based on the third segmentation map includes: The shadows of the non-target human figures in the image to be processed are detected to obtain a shadow segmentation map; wherein, the shadow segmentation map includes the region of the non-target human figures and the region of the shadows of the non-target human figures in the third segmentation map; The non-target human figure in the image to be processed is eliminated based on the shadow segmentation map.

2. The method according to claim 1, characterized in that, The step of segmenting the target portrait and the first portrait from the image to be processed to obtain a first segmentation image of the target portrait and a second segmentation image of the first portrait includes: Detect the target portrait and the first portrait in the image to be processed; If the number of first detection boxes is less than the number of second detection boxes and the number of first detection boxes is greater than 1, then the largest human image is segmented from the local image corresponding to the first detection box to obtain the first segmentation image. The first portrait is segmented from the image to be processed based on the second detection box to obtain the second segmentation image; wherein, the first detection box is the detection box of the target portrait, and the second detection box is the detection box of the first portrait.

3. The method according to claim 2, characterized in that, The method further includes: If the number of the first detection boxes is greater than or equal to the number of the second detection boxes, and the number of the first detection boxes is greater than 1, then it is determined that the non-target human image does not exist in the image to be processed.

4. The method according to claim 2 or 3, characterized in that, The method further includes: If the number of the first detection boxes is 0 and the number of the second detection boxes is greater than 0, then the first portrait is segmented from the image to be processed based on the second detection boxes to obtain the second segmentation image; The non-target human figure in the image to be processed is eliminated based on the second segmentation map.

5. The method according to claim 2 or 3, characterized in that, The method further includes: If the number of the first detection boxes is 0 and the number of the second detection boxes is 0, then it is determined that there is no human image in the image to be processed.

6. The method according to claim 2 or 3, characterized in that, The method further includes: If the number of the first detection boxes is 1 and there is no second detection box whose overlapping area with the first detection box reaches the first ratio, then the local image corresponding to the first detection box is subject segmented to obtain the fourth segmentation image. If the number of human figures in the fourth segmentation image is less than the number of human figures in the second detection box, then the second segmentation image of the first human figure is segmented from the image to be processed based on the second detection box; A fifth segmentation image of the non-target human figure in the first detection frame is determined based on the second segmentation image and the fourth segmentation image; The non-target human figures in the image to be processed are eliminated based on the fifth segmentation map.

7. The method according to claim 6, characterized in that, After performing subject segmentation on the local image corresponding to the first detection box to obtain the fourth segmentation image, the method further includes: If the number of human figures in the fourth segmentation image is not less than the number of human figures in the second detection frame, then it is determined that the non-target human figure does not exist in the first detection frame.

8. The method according to claim 2 or 3, characterized in that, The step of segmenting the target portrait and the first portrait from the image to be processed to obtain a first segmentation image of the target portrait and a second segmentation image of the first portrait includes: If the number of the first detection boxes is 1, and there is a second detection box whose overlapping area with the first detection box reaches a first ratio, then the human image with the largest area is segmented from the local image corresponding to the first detection box to obtain the first segmentation image. The first human image is segmented from the image to be processed based on the second detection box to obtain the second segmentation image.

9. The method according to claim 1, characterized in that, The step of segmenting the target portrait and the first portrait from the image to be processed to obtain a first segmentation image of the target portrait and a second segmentation image of the first portrait includes: The first human figure is segmented from the image to be processed to obtain the second segmentation image; Detect the target human image in the image to be processed; If the number of first detection boxes is less than the number of human figures in the second segmentation image, and the number of first detection boxes is greater than 1, then the segmentation image of the target human figure is obtained from the second segmentation image based on the first detection boxes, and the first segmentation image is obtained; wherein, the first detection box is the detection box of the target human figure.

10. The method according to claim 9, characterized in that, The method further includes: If the number of the first detection boxes is greater than or equal to the number of human figures in the second segmentation image, and the number of the first detection boxes is greater than 1, then it is determined that there is no non-target human figure in the image to be processed.

11. The method according to claim 9 or 10, characterized in that, The method further includes: If the number of the first detection boxes is 0 and the number of human figures in the second segmentation image is greater than 0, then the non-target human figures in the image to be processed are eliminated according to the second segmentation image.

12. The method according to claim 9 or 10, characterized in that, The method further includes: If the number of the first detection boxes is 0 and the number of the second detection boxes is 0, then it is determined that there is no human image in the image to be processed; wherein, the second detection box is the detection box of the human image in the second segmentation image.

13. The method according to claim 9 or 10, characterized in that, The method further includes: If the number of the first detection boxes is 1 and there is no second detection box whose overlapping area with the first detection box reaches the first ratio, then the local image corresponding to the first detection box is subject segmented to obtain a fourth segmentation image; wherein, the second detection box is the detection box of the human image in the second segmentation image; If the number of human figures in the fourth segmentation image is less than the number of human figures in the first detection frame, then the fifth segmentation image of the non-target human figure in the first detection frame is determined based on the second segmentation image and the fourth segmentation image. The non-target human figures in the image to be processed are eliminated based on the fifth segmentation map.

14. The method according to claim 13, characterized in that, After performing subject segmentation on the local image corresponding to the first detection box to obtain the fourth segmentation image, the method further includes: If the number of human figures in the fourth segmentation image is not less than the number of human figures in the second detection frame, then it is determined that the non-target human figure does not exist in the first detection frame.

15. The method according to claim 9 or 10, characterized in that, The method further includes: If the number of the first detection boxes is 1, and there is a second detection box whose overlapping area with the first detection box reaches a first ratio, then the human image with the largest area is segmented from the local image corresponding to the first detection box to obtain the first segmentation image; wherein, the second detection box is the detection box of the human image in the second segmentation image.

16. The method according to claim 1, characterized in that, The step of segmenting the target portrait and the first portrait from the image to be processed to obtain a first segmentation image of the target portrait and a second segmentation image of the first portrait includes: Detect the target portrait and the first portrait in the image to be processed; The first human figure is segmented from the image to be processed to obtain the sixth segmentation image; Based on the sixth segmentation image and the second detection box, the portrait is deduplicated to obtain the deduplicated third detection box; wherein, the second detection box is the detection box of the first portrait; If the number of first detection boxes is less than the number of third detection boxes and the number of first detection boxes is greater than 1, then the segmentation map of the target human image is obtained from the sixth segmentation map based on the first detection boxes, and the first segmentation map is obtained. The first portrait is segmented from the image to be processed based on the third detection box to obtain the second segmentation image.

17. The method according to claim 16, characterized in that, The method further includes: If the number of the first detection boxes is greater than or equal to the number of the third detection boxes, and the number of the first detection boxes is greater than 1, then it is determined that the non-target human image does not exist in the image to be processed.

18. The method according to claim 16 or 17, characterized in that, The method further includes: If the number of the first detection boxes is 0 and the number of the third detection boxes is greater than 0, then the first portrait is segmented from the image to be processed based on the third detection boxes to obtain the second segmentation image; The non-target human figure in the image to be processed is eliminated based on the second segmentation map.

19. The method according to claim 16 or 17, characterized in that, The method further includes: If the number of the first detection boxes is 0 and the number of the third detection boxes is 0, then it is determined that there is no human image in the image to be processed.

20. The method according to claim 16 or 17, characterized in that, The method further includes: If the number of the first detection boxes is 1 and there is no third detection box whose overlapping area with the first detection box reaches the first ratio, then the local image corresponding to the first detection box is subject segmented to obtain the fourth segmentation image. If the number of human figures in the fourth segmentation image is less than the number of human figures in the third detection frame, then the second segmentation image of the first human figure is segmented from the image to be processed based on the third detection frame; A fifth segmentation image of the non-target human figure in the first detection frame is determined based on the second segmentation image and the fourth segmentation image; The non-target human figures in the image to be processed are eliminated based on the fifth segmentation map.

21. The method according to claim 20, characterized in that, After performing subject segmentation on the local image corresponding to the first detection box to obtain the fourth segmentation image, the method further includes: If the number of human figures in the fourth segmentation image is not less than the number of human figures in the third detection frame, then it is determined that the non-target human figure does not exist in the first detection frame.

22. The method according to claim 16 or 17, characterized in that, The step of segmenting the target portrait and the first portrait from the image to be processed to obtain a first segmentation image of the target portrait and a second segmentation image of the first portrait includes: If the number of the first detection boxes is 1, and there is a third detection box whose overlapping area with the first detection box reaches a first ratio, then the human image with the largest area is segmented from the local image corresponding to the first detection box to obtain the first segmentation image. The first portrait is segmented from the image to be processed based on the third detection box to obtain the second segmentation image.

23. The method according to claim 1, characterized in that, The method further includes: Determine whether the first region and the second region overlap; wherein, the first region is the shadowed region in the shadow segmentation image, and the second region is the target human image region in the first segmentation image; If the first region and the second region overlap, delete the overlapping portion of the first region to obtain the processed shadow segmentation map; The non-target human figures in the image to be processed are eliminated based on the processed shadow segmentation map.

24. An electronic device, characterized in that, include: One or more processors; One or more memory units; The memory stores a computer program that, when executed by the processor, causes the electronic device to perform the method as described in any one of claims 1 to 23.

25. A chip system, characterized in that, The chip system includes a processor coupled to a memory, the processor being configured to run a computer program stored in the memory to implement the method as described in any one of claims 1 to 23.

26. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by one or more processors, implements the method as described in any one of claims 1 to 23.

27. A computer program product, characterized in that, The computer program product includes computer program code that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1 to 23.