Portrait instance segmentation method and system based on face detection frame
By using a face instance segmentation model based on face detection boxes and leveraging the bidirectional interaction and fusion of image and detection box features, the problem of low accuracy in face instance segmentation under occlusion and large overlap conditions is solved, achieving higher segmentation accuracy.
Patent Information
- Application Number
- CN202511006553.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-07-22
AI Technical Summary
Existing technologies have poor accuracy in handling face instance segmentation under conditions of large overlap or occlusion, especially in the case of occlusion.
A face instance segmentation model based on face detection bounding boxes is adopted. By combining an image feature extraction network and a face detection bounding box feature extraction network with the bidirectional interactive fusion module of TwowayTransformer, the bidirectional interactive fusion of image and detection box features is performed to generate more accurate instance segmentation results.
It improves the accuracy of human instance segmentation in complex scenes, especially in cases of occlusion or multiple overlapping people. It can better utilize the auxiliary information of face detection boxes for segmentation and generate more accurate instance segmentation results.
Smart Images

Figure CN120510390B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and in particular to a method, system, computer device, and computer-readable storage medium for human face instance segmentation based on face detection boxes. Background Technology
[0002] In recent years, with the development of deep learning technology, attention mechanisms have received widespread attention in the field of deep learning and have been introduced into instance segmentation tasks. Currently, the most accurate instance segmentation methods are based on powerful object detection benchmark models, such as region-based convolutional neural networks (Fast / Faster R-CNN) and the YOLO series. They mostly follow a basic rule: first, a large number of candidate regions are generated, and then non-maximum suppression (NMS) is used to remove redundant regions. However, when two objects of the same class have significant overlap, NMS will treat one of the objects as a redundant proposal region and eliminate it. This means that almost all object detection methods cannot handle cases of large overlap. Furthermore, even if the detection method sometimes successfully detects two instances, bounding boxes are not suitable for instance segmentation under occlusion. If two instances severely overlap, they will appear in the same bounding box, making it difficult for the segmentation network to determine which instance should be the target in the Region of Interest (RoI).
[0003] In practice, target instances often severely occlude each other, significantly reducing segmentation accuracy. This is especially true in human face segmentation, where the segmentation of a specific person is of greater concern. "Human" is a special category that can be defined using face bounding boxes.
[0004] Currently, there is no effective solution to the problem of poor accuracy in human image instance segmentation methods in related technologies. Summary of the Invention
[0005] This application provides a method, system, computer device, and computer-readable storage medium for training a portrait instance segmentation model based on a face detection bounding box, in order to at least solve the problem of poor accuracy in related art portrait instance segmentation methods.
[0006] In a first aspect, embodiments of this application provide a method for training a portrait instance segmentation model based on a face detection bounding box. The instance segmentation model includes an image feature extraction network, a face detection bounding box feature extraction network, and an interaction fusion module. Any training step in the method includes:
[0007] Receive a training dataset, wherein the training dataset includes multiple portrait images, and each portrait image includes portrait instance segmentation results and face bounding box coordinates;
[0008] Image features are extracted based on the training dataset using the image feature extraction network, and detection box features are extracted based on the training dataset using the face detection box feature extraction network.
[0009] The image features and the detection box features are bidirectionally fused using an interactive fusion module built on TwowayTransformer to obtain the human image instance segmentation result.
[0010] In some embodiments, the method further includes:
[0011] Based on the training dataset, the Adam optimizer is used to iteratively run the training steps and optimize the model parameters. When the model converges, the trained instance segmentation model is obtained.
[0012] In some embodiments, extracting image features based on the training dataset through the image feature extraction network includes:
[0013] By using a backbone network built on CSPDarknet, information is extracted from each portrait image in the training dataset to obtain feature maps at different levels.
[0014] The feature maps at different levels are input into a fusion network built on PAFPN. The feature maps are upsampled and concatted to fuse and optimize the features at different levels, thereby obtaining enhanced features.
[0015] The enhanced features are output according to a preset dimension to obtain the image features.
[0016] In some embodiments, the detection box features are obtained by extracting detection box features based on the training dataset and outputting them according to the preset dimensions through a face detection box feature extraction network built on the PromptEncoder module of the standard MobileSAM network.
[0017] In some embodiments, the bidirectional interactive fusion of the image features and the detection box features, through an interactive fusion module built on TwowayTransformer, includes:
[0018] By using multiple blocks, based on the Q-KV interchangeable attention mechanism, bidirectional information flow fusion is performed on the image features and the detection box features to obtain intermediate feature representations;
[0019] In the output stage, the intermediate feature representations are further processed using the standard KQV attention mechanism to generate fused prediction results.
[0020] The fusion prediction result is multiplied by the upsampled image features to obtain the human image instance segmentation result.
[0021] In some embodiments, through multiple blocks, based on the Q-KV interchangeable attention mechanism, bidirectional information flow fusion is performed on the image features and the detection box features to obtain intermediate feature representations, including:
[0022] By performing self-attention encoding on the image features, the image features are integrated and enhanced.
[0023] The result of the first self-attention encoding is processed by a multilayer perceptron, and the transformed feature is then cross-attention encoded with the detection box feature to obtain the intermediate feature representation.
[0024] Secondly, embodiments of this application provide a method for human face instance segmentation based on a face detection bounding box, the method comprising:
[0025] Receive portrait images to be processed;
[0026] Based on the portrait image to be processed, the portrait instance segmentation model trained by the method described in the first aspect is used to perform instance segmentation on the portrait image to be processed, and the portrait instance segmentation result is obtained.
[0027] Thirdly, embodiments of this application provide a portrait instance segmentation system based on a face detection bounding box, including an acquisition module and a segmentation module, wherein:
[0028] The acquisition module is used to receive portrait images to be processed;
[0029] The segmentation module is used to perform instance segmentation on the portrait image to be processed based on the portrait image to be processed, using the portrait instance segmentation model trained by the method described in the first aspect, to obtain the portrait instance segmentation result.
[0030] Fourthly, embodiments of this application provide a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect above.
[0031] Fifthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method described in the first aspect above.
[0032] Compared to related technologies, the portrait instance segmentation method based on face detection boxes provided in this application acquires image features and face detection box features separately, and then fuses them for model segmentation. During feature fusion, a bidirectional interactive fusion update method is adopted, which makes full use of the relationship between image information and face detection box information. Compared with traditional methods based on a single attention mechanism or simple fusion, it can more effectively integrate the two types of information, improve the accuracy and reliability in portrait instance segmentation tasks, especially in handling portrait segmentation in complex scenes, such as when there is occlusion or multiple overlapping characters. It can better utilize the auxiliary information of face detection boxes to guide the processing and segmentation of image features, thereby obtaining more accurate instance segmentation results. Attached Figure Description
[0033] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0034] Figure 1 This is a method for training a portrait instance segmentation model based on a face detection bounding box, according to an embodiment of this application;
[0035] Figure 2 This is a schematic diagram of the backbone network in the image feature extraction network according to an embodiment of this application;
[0036] Figure 3 This is a detailed schematic diagram of the various components in the backbone network;
[0037] Figure 4 This is a schematic diagram of a feature fusion network according to an embodiment of this application;
[0038] Figure 5 This is a schematic diagram of a two-way interactive fusion module according to an embodiment of this application;
[0039] Figure 6 This is a schematic diagram of the original portrait image to be processed according to an embodiment of this application;
[0040] Figure 7 This is a schematic diagram of the face detection bounding box according to an embodiment of this application;
[0041] Figure 8 This is a schematic diagram of instance segmentation results according to an embodiment of this application;
[0042] Figure 9 This is a structural block diagram of a human image instance segmentation system according to an embodiment of this application;
[0043] Figure 10This is a schematic diagram of the internal structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.
[0045] Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without any inventive effort. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, any changes to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.
[0046] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0047] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following objects are in an "or" relationship. The terms "first," "second," and "third" used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.
[0048] Image instance segmentation is a technique that separates each object instance in an image and labels it with its specific outline and category. It combines the features of object detection and semantic segmentation, aiming not only to identify different categories of objects in the image but also to distinguish different instances of the same category, achieving accurate segmentation of each individual object in the image.
[0049] In recent years, with the development of deep learning technology, attention mechanisms have received widespread attention in the field of deep learning and have been introduced into instance segmentation tasks. Currently, the most accurate instance segmentation methods are based on strong object detection baselines, such as Fast / Faster R-CNN (Towards Real-Time Object Detection with Region Proposal Networks) and the YOLO series. These techniques generally follow a basic rule: first, a large number of proposal regions are generated, and then non-maximum suppression (NMS) is used to remove redundant regions. However, when two objects of the same category have significant overlap, NMS will treat one of the objects as a redundant proposal region and eliminate it. This means that almost all object detection methods cannot handle cases of large overlap. Furthermore, even if the detection method sometimes successfully detects two instances, the bounding box is not suitable for instance segmentation under occlusion conditions. If two instances severely overlap, they will appear in the same bounding box, making it difficult for the segmentation network to determine which instance should be the target in the Region of Interest (RoI).
[0050] In reality, target instances often severely occlude each other, significantly reducing segmentation accuracy. This is especially true in human image segmentation, where the focus is on segmenting instances of a specific target person. Currently, most instance segmentation networks are based on the original image and output segmentation maps for all instances. This approach fails to address situations where instances are missed in the image or when the goal is to identify a specific instance.
[0051] In view of this, in order to obtain a specified high-quality human instance segmentation result, the inventors of this case, considering that "human" is a special category that can be defined by face bounding boxes, proposed a human instance segmentation method based on face detection boxes. Using face detection boxes as auxiliary information, it can correspond one-to-one with the final instance segmentation result, and at the same time improves the accuracy in cases of severe occlusion.
[0052] Figure 1 This application discloses a method for training a portrait instance segmentation model based on a face detection bounding box, according to an embodiment of the present application. The instance segmentation model includes an image feature extraction network, a face detection bounding box feature extraction network, and an interaction fusion module, such as... Figure 1 As shown, the process includes the following steps:
[0053] S101, Receive training dataset, wherein the training dataset includes multiple portrait images, and each portrait image includes portrait instance segmentation results and face bounding box coordinates;
[0054] The training dataset is constructed through the following detailed steps:
[0055] S1. Select N portrait images as the original images for the training dataset. These images cover different scenes, lighting conditions, shooting angles, and human poses to reflect the diverse portrait situations in reality as comprehensively as possible.
[0056] S2, for each image Obtain the corresponding instance segmentation image , For each image The number of people in The range of values is within Integers between these values, where different values represent different instances; this step can be achieved using Photoshop or other automated tools.
[0057] S3, each image The coordinates of the bounding boxes for all faces in the image are obtained using the Labelme annotation tool. , The format is ;
[0058] Each sublist This represents the coordinate information of a face detection bounding box. The specific order can be the top left horizontal coordinate, the top left vertical coordinate, the bottom right horizontal coordinate, the bottom right vertical coordinate, etc. (The coordinate order can be determined according to the settings of the Labelme tool and subsequent usage requirements).
[0059] After the above step S101, a portrait segmentation training dataset with face detection boxes is successfully obtained. For example, this training dataset can be represented in the following form:
[0060]
[0061] in, Represents the original set of images. This represents the set of instance segments corresponding to the original image. This represents the set of face detection boxes corresponding to the original image. Indicates the size of the training dataset; optional, in this embodiment... .
[0062] S102, Image features are extracted based on the training dataset through an image feature extraction network, and detection box features are extracted based on the training dataset through a face detection box feature extraction network;
[0063] The image feature extraction network (ImageEncoder) in this application is built based on the YOLOv8 network. Specifically, the backbone network of this network is an improved CSPDarknet.
[0064] Figure 2 This is a schematic diagram of the backbone network in the image feature extraction network according to an embodiment of this application, as shown below. Figure 2 As shown, CSPDarknet consists of a series of convolutional modules, CSP modules, SPPF modules, etc. Figure 3 This is a detailed schematic diagram of the various components in the backbone network, such as... Figure 2 and Figure 3 As shown, these modules work together to extract features and perform preliminary processing on the input image, effectively capturing various information in the image and providing a rich feature base for subsequent processing.
[0065] Specifically, extracting image features based on the training dataset using an image feature extraction network includes the following steps:
[0066] S1, through the backbone network built on CSPDarknet, information is extracted from each portrait image in the training dataset to obtain feature maps at different levels;
[0067] The CSP (Cross-Stage Partial Network) module is primarily used to optimize the network structure, improve computational efficiency, and enhance feature representation capabilities. By segmenting and recombining feature maps along the channel dimension, one part is directly connected to subsequent layers, while the other part undergoes a series of convolutional operations before being merged with the former. This reduces computational cost while enhancing the network's ability to learn features, enabling it to better capture key information in images and improve the model's generalization performance across different scenarios. In the portrait instance segmentation task of this embodiment, it can more effectively extract and utilize features from the training dataset, providing strong support for subsequent accurate segmentation.
[0068] Specifically, the processing flow of the backbone network is illustrated below:
[0069] The input image enters CSPDarknet and is processed by multiple convolutional modules. These convolutional modules use different convolutional kernels and strides to perform feature extraction and preliminary feature transformation on the image, extracting low-level features of the image.
[0070] Next, the feature map enters the CSP module, which splits and reassembles the feature map along the channel dimension. One part is directly connected to the subsequent layers, while the other part is merged with the former after a series of convolution operations, thereby optimizing feature representation and reducing computation.
[0071] Finally, the feature maps processed by the CSP module are fed into the SPPF (Spatial Pyramid Pooling Fast) module. The SPPF module performs pooling operations on features at different scales, converting feature maps of different sizes into fixed-size outputs, further enhancing the feature fusion and expressive capabilities. This provides the subsequent feature fusion network PAFPN (Path Aggregation Network) with pre-processed and optimized feature maps, enabling it to play a better role in the entire image information extraction branch and laying the foundation for the final human instance segmentation task.
[0072] S2, input feature maps of different levels into the fusion network built based on PAFPN, and fuse and optimize the features of different levels by upsampling and concat operations to obtain enhanced features;
[0073] The feature fusion network in this embodiment is obtained by optimizing and improving the classic Feature Pyramid Network (FPN). Figure 4 This is a schematic diagram of a feature fusion network according to an embodiment of this application.
[0074] Furthermore, the SPPF module performs pooling operations on features at different scales, sampling and aggregating the input feature maps at different spatial scales, and converting feature maps of different sizes into fixed-size outputs. This allows the network to adapt to input images of different sizes and better fuse multi-scale information.
[0075] In this embodiment, for human image instance segmentation, the SPPF module can help the model obtain more comprehensive and representative features for human images of different sizes and for possible occlusion situations, thereby improving the model's ability to cope with complex scenes and thus improving the accuracy of instance segmentation.
[0076] S3 outputs the enhanced features according to the preset dimensions to obtain image features.
[0077] In this embodiment, the original complete prediction function is not directly utilized. The main function of the image feature extraction network is to extract the feature information of the image to prepare for subsequent fusion with the face detection box features, rather than directly outputting the complete detection result like the original YOLOv8.
[0078] Therefore, when constructing the ImageEncoder, the Head module in the original YOLOv8 structure was removed, and the final output was set to a preset dimension, such as 32. This operation simplifies the model structure, reduces unnecessary computation, and setting the output dimension to 32 allows for more efficient integration and information transfer with subsequent modules while ensuring the extraction of effective features.
[0079] As can be understood, the image feature extraction network in this embodiment uses CSPDarknet as the backbone network to perform preliminary feature extraction and processing on the input image, efficiently capturing key information in the image. Furthermore, PAFPN further fuses and optimizes these features based on those extracted by CSPDarknet. Through operations such as upsampling and concat, it integrates features at different levels, ensuring that the image's feature information is fully utilized at different scales, thereby enhancing the model's adaptability to targets of different sizes.
[0080] In addition, a face detection box feature extraction network is constructed based on the PromptEncoder module of the standard MobileSAM (Lightweight Segment Anything Model) network. Based on the training set, the detection box features are extracted and output with a preset dimension to obtain the detection box features.
[0081] Specifically, the BoxEncoder operates based on the PromptEncoder module in the MobileSAM network. During processing, its internal mechanism focuses on identifying face regions in the image and determining the location information of the face bounding boxes. This process involves detecting and judging the pixel features of the image, and through a series of calculations and processing steps, ultimately extracting key data such as the coordinates of the face bounding boxes. This information is then processed and transformed according to a predefined method, outputting a final 32-dimensional format. This provides accurate face bounding box information for the portrait instance segmentation model, enabling subsequent fusion with image feature information to achieve more precise instance segmentation.
[0082] S103 uses a bidirectional interactive fusion module built on TwowayTransformer to perform bidirectional interactive fusion of image features and detection box features to obtain fused prediction results and image features;
[0083] This embodiment uses TwoWayTransformer to fuse the image information features and face detection box information obtained in step S102. It should be noted that the TwoWayTransformer structure includes both standard KQV attention calculation and Q-KV interchange attention, corresponding to the bidirectional interactive fusion and update of prediction results (query) and image features (key / value).
[0084] Specifically, Figure 5 This is a schematic diagram of a bidirectional interactive fusion module according to an embodiment of this application, as shown below. Figure 5 The process includes the following steps:
[0085] S1, through multiple blocks, performs bidirectional information flow fusion on image features and detection box features based on the Q-KV interchangeable attention mechanism to obtain intermediate feature representation;
[0086] Specifically, it includes:
[0087] In the initial stage, image features undergo self-attention encoding. Through self-attention encoding, the image features themselves can learn and strengthen the correlation between features in different regions, and mine the feature structure information inside the image.
[0088] Furthermore, the features after self-attention encoding undergo a first cross-attention encoding with the face bounding box coordinate features. This process uses image features as the query vector and face bounding box coordinate features as the key and value vectors. This setup allows image features to adjust their feature representation based on the coordinate information of the face detection box, focusing on the image region related to the detection box and extracting more targeted features.
[0089] The result after the first cross-attention encoding enters the multilayer perceptron for further feature transformation and processing to improve the expressive power and adaptability of the features.
[0090] Furthermore, the output of the multilayer perceptron is subjected to a second cross-attention encoding with the coordinate features. In this case, the query vector is set as the coordinate features, and the key and value vectors are set as image features. The output is an intermediate feature representation. This step effectively readjusts the image features based on the face detection bounding box information, further strengthening the connection and fusion between the two.
[0091] Figure 5Although only one block structure is shown, it can be understood that in this TwoWayTransformer implementation, multiple blocks process the input image features and coordinate features sequentially. Within each block, a predefined attention calculation process is followed: first, self-attention encoding of the image features is performed; then, two cross-attention encodings are performed using different combinations of query vectors, key vectors, and value vectors; finally, the final fusion result is obtained through calculations including standard KQV. These blocks are stacked layer by layer, gradually deepening the fusion of image features and bounding box features, enabling the model to fully explore the correlation information between them, thereby improving the instance segmentation effect.
[0092] S2, in the output stage, further processes the intermediate feature representations through the standard KQV attention mechanism to generate fused prediction results and image features.
[0093] The output stage also includes a standard KQV calculation process to obtain the final Q and K, which correspond to the fused prediction result and image features, respectively.
[0094] The fused prediction result refers to the feature representation fused by the TwoWayTransformer module. The TwoWayTransformer uses self-attention and cross-attention mechanisms to interact with image features and face detection box features to generate the fused intermediate feature representation, i.e., the prediction result, which is used for subsequent mask generation.
[0095] S3 performs matrix multiplication between the fused prediction results and the upsampled image features to obtain the human image instance segmentation results.
[0096] Furthermore, the transposed convolutional image features are multiplied (inner product) with the upsampled image features to output masks, which precisely define the pixel regions of target instances, achieving pixel-level instance segmentation. It should be noted that masks are the result of instance segmentation, used to precisely define the pixel region of each target instance in the image. In portrait instance segmentation, they are used to depict the outline of the portrait, achieving pixel-level segmentation by labeling each pixel as foreground or background.
[0097] In addition, this module also outputs IOU scores, which measure the degree of overlap between the predicted detection box and the ground truth detection box, and help to judge the accuracy of the prediction.
[0098] The above steps S101 to S103 describe the complete operation process of each model. It can be understood that in order to obtain a model that can be practically deployed, it is necessary to use a large amount of data, repeatedly execute the above process, and continuously adjust the model for multiple iterations of training to obtain a trained model.
[0099] Specifically, in this embodiment, the training is implemented and performed based on the open-source deep learning framework PyTorch version 1.12. The training is carried out on four NVIDIA GeForce RTX 3090 graphics cards. The Adam optimizer is used for training, and the learning rate is set to 1e-4. The data augmentation methods used on the original images during training include random horizontal flipping, random brightness adjustment, and random contrast adjustment.
[0100] Furthermore, the instance segmentation loss function is based on binary cross-entropy (BCE).
[0101]
[0102] Where N is the total number of pixels in the image. It is the true label (0 or 1) of the i-th pixel. is the probability (values between 0 and 1) that the i-th pixel is predicted to belong to the foreground, and log is the natural logarithm.
[0103] This embodiment employs the BCE loss function, which directly measures the difference between the model's predicted probability of each pixel belonging to the foreground (instance) and the true label. In the portrait instance segmentation of this embodiment, by comparing the prediction and the true situation of each pixel, the model is effectively guided to learn how to accurately distinguish between foreground and background pixels, allowing the model to gradually optimize the prediction results and adjust towards the correct segmentation direction during training.
[0104] Furthermore, the goal of instance segmentation is to generate an accurate pixel-level mask for each target instance. The BCE loss function focuses on pixel-level classification loss calculation, which aligns closely with this task objective. By minimizing this loss function, the model can continuously improve its ability to recognize instance boundaries and internal pixels, thereby generating more accurate instance segmentation maps and contributing to the ultimate goal of accurately segmenting human images.
[0105] Through the steps S101 to S103 described above, compared with traditional instance segmentation methods, the face instance segmentation method based on face detection boxes acquires image features and face detection box features separately, and then fuses them for model segmentation. During feature fusion, a bidirectional interactive fusion update method is adopted, which makes full use of the relationship between image information and face detection box information. Compared with traditional methods that rely solely on a single attention mechanism or simple fusion, this method can more effectively integrate the two types of information, improving the accuracy and reliability in face instance segmentation tasks. Especially when dealing with face segmentation in complex scenes, such as when there is occlusion or multiple people overlapping, the auxiliary information of the face detection box can be better utilized to guide the processing and segmentation of image features, thereby obtaining more accurate instance segmentation results.
[0106] After obtaining the trained model, a method for human instance segmentation based on face detection bounding boxes is also provided, which includes the following steps:
[0107] S201, Receive the portrait image to be processed;
[0108] S202, Based on the portrait image to be processed, the instance segmentation model trained through the above steps is used to perform instance segmentation on the portrait image to be processed, and the portrait instance segmentation result is obtained.
[0109] Figure 6 This is a schematic diagram of the original portrait image to be processed according to an embodiment of this application. Figure 7 This is a schematic diagram of the face detection bounding box according to an embodiment of this application. Figure 8 This is a schematic diagram of the instance segmentation result according to an embodiment of this application.
[0110] according to Figure 6 , Figure 7 and Figure 8 As shown, the instance segmentation method provided in this application can output instance segmentation maps that correspond one-to-one with face detection boxes. Compared with the traditional method of outputting all instance segmentation maps based on the original image, it can more accurately segment the portrait instances corresponding to the specified face boxes, effectively solving the problem of missed instance detection or difficulty in identifying a specific instance in the image.
[0111] In human face instance segmentation, when target instances severely occlude each other, this method utilizes face detection bounding boxes as auxiliary information and combines them with a cross-attention mechanism to achieve feature fusion. This improves segmentation accuracy to a certain extent and overcomes the shortcomings of existing technologies in segmentation under occlusion conditions. For example, traditional object detection methods are prone to misclassification or difficulty in identifying targets when dealing with large overlaps or occlusions.
[0112] This embodiment also provides a portrait instance segmentation system based on face detection bounding boxes. This system deploys the portrait instance segmentation model trained in the above steps. Figure 9 This is a structural block diagram of a human image instance segmentation system according to an embodiment of this application, such as... Figure 9 As shown, the system includes: an acquisition module 90 and a segmentation module, wherein:
[0113] The acquisition module 90 is used to receive the portrait image to be processed;
[0114] The segmentation module 91 is used to perform instance segmentation on the portrait image to be processed based on the portrait instance segmentation model trained in the above embodiments, and obtain the portrait instance segmentation result.
[0115] The instance segmentation system proposed in this application acquires image features and face detection box features respectively, and then fuses them for model segmentation. During the feature fusion process, a bidirectional interactive fusion update method is adopted, which makes full use of the relationship between image information and face detection box information to output instance segmentation maps that correspond one-to-one with face detection boxes. Compared with traditional systems that output all instance segmentation maps based on the original image, this system can more accurately segment portrait instances corresponding to specified face boxes, effectively solving the problem of missed instance detection or difficulty in identifying a specific instance in the image.
[0116] Furthermore, in human face instance segmentation, when target instances severely occlude each other, this system utilizes face detection bounding boxes as auxiliary information and combines them with a cross-attention mechanism to achieve feature fusion. This improves segmentation accuracy to a certain extent and overcomes the shortcomings of existing technologies in segmentation under occlusion conditions. For example, traditional object detection methods are prone to instance misjudgment or difficulty in identifying targets when dealing with large overlaps or occlusions.
[0117] In one embodiment, Figure 10 This is a schematic diagram of the internal structure of an electronic device according to an embodiment of this application, such as... Figure 10 As shown, an electronic device is provided, which can be a server, and its internal structure diagram can be as follows. Figure 10 As shown, the electronic device includes a processor, a network interface, internal memory, and non-volatile memory connected via an internal bus. The non-volatile memory stores the operating system, computer programs, and a database. The processor provides computing and control capabilities, the network interface communicates with external terminals via a network connection, the internal memory provides the environment for the operating system to run, the computer program is executed by the processor to implement a face instance segmentation method based on face detection bounding boxes, and the database stores data.
[0118] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. Specifically, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0119] Furthermore, in conjunction with the face instance segmentation method based on face detection boxes in the above embodiments, this application embodiment can provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the face instance segmentation methods based on face detection boxes in the above embodiments.
[0120] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the face instance segmentation methods based on face detection boxes in the above embodiments.
[0121] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0122] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0123] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for training a portrait instance segmentation model based on face detection bounding boxes, characterized in that, The instance segmentation model includes an image feature extraction network, a face detection bounding box feature extraction network, and an interaction fusion module. Any training step in the method includes: Receive a training dataset, wherein the training dataset includes multiple portrait images, and each portrait image includes portrait instance segmentation results and face bounding box coordinates; Image features are extracted based on the training dataset using the image feature extraction network, and detection box features are extracted based on the training dataset using the face detection box feature extraction network. The image features and the detection box features are bidirectionally fused using an interactive fusion module built on TwowayTransformer to obtain the human image instance segmentation result.
2. The method according to claim 1, characterized in that, The method further includes: Based on the training dataset, the Adam optimizer is used to iteratively run the training steps and optimize the model parameters. When the model converges, the trained instance segmentation model is obtained.
3. The method according to claim 2, characterized in that, Extracting image features from the training dataset using the image feature extraction network includes: By using a backbone network built on CSPDarknet, information is extracted from each portrait image in the training dataset to obtain feature maps at different levels. The feature maps at different levels are input into a fusion network built on PAFPN. The feature maps are upsampled and concatted to fuse and optimize the features at different levels, thereby obtaining enhanced features. The enhanced features are output according to a preset dimension to obtain the image features.
4. The method according to claim 3, characterized in that, The face detection bounding box feature extraction network is constructed based on the PromptEncoder module in the standard MobileSAM network. The detection bounding box features are extracted based on the training dataset and output according to the preset dimensions to obtain the detection bounding box features.
5. The method according to claim 1, characterized in that, The interactive fusion module, built on the TwowayTransformer, performs bidirectional interactive fusion of the image features and the detection box features, including: By using multiple blocks, based on the Q-KV interchangeable attention mechanism, bidirectional information flow fusion is performed on the image features and the detection box features to obtain intermediate feature representations; In the output stage, the intermediate feature representations are further processed using the standard KQV attention mechanism to generate fused prediction results. The fusion prediction result is multiplied by the upsampled image features to obtain the human image instance segmentation result.
6. The method according to claim 5, characterized in that, Through multiple blocks, based on the Q-KV interchangeable attention mechanism, bidirectional information flow fusion is performed on the image features and the detection box features to obtain the intermediate feature representation, including: By performing self-attention encoding on the image features, the image features are integrated and enhanced. The result of the first self-attention encoding is processed by a multilayer perceptron, and the transformed feature is then cross-attention encoded with the detection box feature to obtain the intermediate feature representation.
7. A method for human face instance segmentation based on face detection bounding boxes, characterized in that, The method includes: Receive portrait images to be processed; Based on the portrait image to be processed, the portrait instance segmentation model trained by the method described in any one of claims 1-6 is used to perform instance segmentation on the portrait image to be processed, thereby obtaining the portrait instance segmentation result.
8. A human face instance segmentation system based on face detection bounding boxes, characterized in that, This includes an acquisition module and a segmentation module, wherein: The acquisition module is used to receive portrait images to be processed; The segmentation module is used to perform instance segmentation on the portrait image to be processed based on the portrait image to be processed, using a portrait instance segmentation model trained by the method described in any one of claims 1-6, to obtain a portrait instance segmentation result.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Embedded platform-oriented dual-mode target detection method and system
CN113255521A
Instance segmentation method and device and storage medium
CN115984304A