Foot region segmentation method and device and storage medium
By combining the pre-trained Florence-2 model and the BiRefNet model, accurate segmentation of the foot region can be achieved without training, solving the problems of inaccurate semantic understanding and high segmentation error rate in existing technologies, and improving the user experience of virtual try-on and the generalization ability of the system.
Patent Information
- Application Number
- CN202510872099.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-31
AI Technical Summary
In existing technologies, deep learning-based image segmentation methods cannot accurately respond to complex semantic instructions, resulting in inaccurate segmentation of the foot region. Furthermore, their generalization ability is insufficient, making it difficult to achieve high-precision segmentation and natural language interaction, which affects the effect of virtual try-on and user experience.
A pre-trained Florence-2 model is used for semantic parsing, combined with reverse NMS operation and BiRefNet model for fine segmentation. The Florence-2 model achieves accurate semantic parsing, and BiRefNet is used for fine segmentation to optimize the detection box localization, so as to achieve accurate segmentation of the foot region without training.
It achieves accurate segmentation of the foot region without retraining the model, improves the system's generalization ability and interactive flexibility, maintains the consistency of shoe shape, provides a reliable foundation for virtual shoe changing, and solves the problems of inaccurate semantic understanding and high segmentation error rate in existing technologies.
Smart Images

Figure CN120876846A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer vision and multimodal artificial intelligence, and in particular to a method, device and storage medium for foot region segmentation. Background Technology
[0002] In e-commerce virtual shoe-trying scenarios, accurate segmentation of the foot area is a key technology for enabling virtual shoe try-on. Among related technologies, deep learning-based image segmentation methods mainly rely on specific training data, but the traditional models used lack the ability to understand natural semantic instructions, resulting in an inability to accurately respond to diverse foot segmentation requests from users. Existing solutions typically employ fixed-pattern detection algorithms, which struggle to understand complex semantic instructions such as "please segment the left foot wearing socks" or "segment only the area above the ankle," severely impacting user experience. Furthermore, the generalization ability of existing solutions is severely limited. When segmenting new product categories (such as socks, anklets, etc.), it is necessary to re-collect data, label samples, and train dedicated models, failing to achieve rapid zero-shot / one-shot adaptation and incurring extremely high technical implementation costs. Traditional methods typically employ a two-stage processing flow: object detection followed by semantic segmentation. This fragmented processing leads to the loss of semantic information during transmission. Therefore, current mainstream solutions have failed to effectively address the coordination problem between multimodal semantic understanding and visual segmentation, making it difficult for the system to simultaneously meet the two core requirements of "high-precision segmentation" and "natural language interaction." This has become a key obstacle restricting the large-scale commercialization of virtual try-on technology. In addition, existing technologies often use simple non-maximum suppression algorithms when processing overlapping foot areas, which can easily lead to the loss of key areas and affect the realism of subsequent virtual try-on effects.
[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide a method, device and storage medium for foot region segmentation, which aims to solve the technical problem that the lack of semantic understanding in the existing segmentation model leads to segmentation defects in the foot region.
[0005] To achieve the above objectives, this application proposes a foot region segmentation method, the method comprising:
[0006] Obtain the detection object, which includes the target detection image and the detection semantics;
[0007] The detection semantics are parsed using a pre-trained Florence-2 model, and the foot region detection results are marked in the target detection image based on the parsing results of the detection semantics.
[0008] Perform a reverse NMS operation on the foot region detection results to obtain the target detection box;
[0009] The BiRefNet model extracts the foreground image from the object detection image, and captures the target location as the foot region based on the position of the object detection box in the foreground image.
[0010] In one embodiment, the step of using a pre-trained Florence-2 model to parse the detection semantics and marking the foot region detection result in the target detection image based on the parsing result of the detection semantics includes:
[0011] The detected semantics are input into the pre-trained Florence-2 model for semantic understanding, and a semantic parsing vector is generated based on the semantic understanding result;
[0012] Based on the semantic parsing vector, candidate bounding boxes for the foot region are located in the target detection image;
[0013] The candidate bounding boxes for the foot region are filtered by confidence, and the candidate bounding boxes for the foot region that meet the preset conditions are taken as the detection results of the foot region.
[0014] In one embodiment, the step of performing confidence screening on the candidate bounding boxes of the foot region and using the candidate bounding boxes of the foot region that meet preset conditions as the detection result of the foot region includes:
[0015] Calculate the semantic matching score of each foot region candidate box, and remove the foot region candidate boxes whose semantic matching score is lower than a preset threshold;
[0016] Non-maximum suppression is applied to the remaining candidate bounding boxes of the foot region to remove redundant detection boxes and obtain the detection result of the foot region.
[0017] In one embodiment, the step of performing a reverse NMS operation on the foot region detection results to obtain the target detection box includes:
[0018] Iterate through the detection results of each foot region and calculate the overlapping area between the detection results of each foot region;
[0019] The foot region detection result containing the most overlapping area is used as the target detection box. When there are multiple foot region detection results with the same overlapping area, the foot region detection result with the largest area is used as the target detection box.
[0020] In one embodiment, the step of using the detection result of the foot region containing the most overlapping area as the target detection box includes:
[0021] Calculate the cross-union ratio (CUC) among the detection results of each foot region;
[0022] The target detection box is determined from the foot region detection results based on the sum of the intersection-union ratios.
[0023] In one embodiment, the step of extracting the foreground image from the target detection image using the BiRefNet model and capturing the target location as the foot region based on the position of the target detection bounding box in the foreground image includes:
[0024] The target detection image is input into the BiRefNet model to obtain the multi-scale feature map output by the BiRefNet model;
[0025] Foreground separation is performed based on the multi-scale feature map to generate a preliminary foreground mask;
[0026] The initial foreground mask is optimized to obtain the foreground image.
[0027] In one embodiment, the step of extracting the foreground image from the target detection image using the BiRefNet model and capturing the target location as the foot region based on the position of the target detection bounding box in the foreground image includes:
[0028] The coordinates of the target detection box are mapped to the foreground image, and candidate foreground for the foot region are cropped based on the coordinate mapping;
[0029] Edge optimization is performed on the candidate foreground of the foot region to obtain the foot region.
[0030] In one embodiment, the step of edge optimization of the candidate foreground region of the foot region as the foot region includes:
[0031] Morphological closing operations are used to fill the internal holes of the candidate foreground of the foot region, and non-foot noise regions in the candidate foreground of the foot region are removed by edge detection algorithms;
[0032] Output a smooth and continuous foot region segmentation mask to optimize the edges of the foot region candidate foreground.
[0033] In addition, to achieve the above objectives, this application also proposes a foot region segmentation device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the foot region segmentation method as described above.
[0034] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the foot region segmentation method described above.
[0035] One or more technical solutions proposed in this application have at least the following technical effects:
[0036] In this application's technical solution, a detection object is obtained, comprising a target detection image and detection semantics. A pre-trained Florence-2 model is used to parse the detection semantics, and the foot region detection result is marked in the target detection image based on the parsing result of the detection semantics. A reverse NMS operation is performed on the foot region detection result to obtain a target detection box. A BiRefNet model is used to extract the foreground image from the target detection image, and the target position is captured as the foot region based on the position of the target detection box in the foreground image. This application achieves accurate semantic parsing through the Florence-2 model, optimizes the detection box localization by combining it with reverse NMS, and uses BiRefNet for refined segmentation, achieving accurate segmentation of the foot region without training. This effectively solves the problems of inaccurate semantic understanding and high segmentation error rate in existing technologies, while maintaining the consistency of shoe shape, providing a reliable foundation for virtual shoe changing. Attached Figure Description
[0037] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0038] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 A flowchart illustrating an embodiment of the foot region segmentation method of this application;
[0040] Figure 2 This is a schematic diagram of the device structure of the hardware operating environment involved in the foot region segmentation method in this application embodiment;
[0041] Figure 3 This is a schematic diagram illustrating the working principle of the Florence-2 model.
[0042] Figure 4 This is a schematic diagram illustrating the working principle of the BiRefNet model;
[0043] Figure 5 This is a schematic diagram illustrating the process of capturing the footstep region from an image based on target detection.
[0044] Figure 6 This is a schematic diagram confirming the target detection bounding box;
[0045] Figure 7 This is a schematic diagram for filling holes.
[0046] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0047] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0048] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0049] The main solution of this application embodiment is as follows: A detection object is obtained, which includes a target detection image and detection semantics; a pre-trained Florence-2 model is used to parse the detection semantics, and the foot region detection result is marked in the target detection image according to the parsing result of the detection semantics; a reverse NMS operation is performed on the foot region detection result to obtain the target detection box; a foreground image is extracted from the target detection image using a BiRefNet model, and the target position is captured as the foot region according to the position of the target detection box in the foreground image.
[0050] Since the mainstream solutions of existing technologies have not effectively solved the problem of coordination between multimodal semantic understanding and visual segmentation, the system is unable to simultaneously meet the two core requirements of "high-precision segmentation" and "natural language interaction". This has become a key obstacle restricting the large-scale commercialization of virtual try-on technology.
[0051] This application provides a solution that achieves accurate semantic parsing through the Florence-2 model, optimizes the detection box localization by combining reverse NMS, and uses BiRefNet for fine segmentation. This enables accurate segmentation of the foot region without training, effectively solving the problems of inaccurate semantic understanding and high segmentation error rate in existing technologies, while maintaining the consistency of shoe shape, providing a reliable foundation for virtual shoe changing.
[0052] Based on this, embodiments of this application provide a method for segmenting the foot region, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the foot region segmentation method of this application. In this embodiment, the foot region segmentation method includes steps S10 to S40:
[0053] Step S10: Obtain the detection object, which includes the target detection image and the detection semantics;
[0054] Step S20: The pre-trained Florence-2 model is used to parse the detection semantics, and the foot region detection results are marked in the target detection image according to the parsing results of the detection semantics.
[0055] Step S30: Perform a reverse NMS operation on the foot region detection results to obtain the target detection box;
[0056] Step S40: Extract the foreground image from the target detection image using the BiRefNet model, and capture the target location as the foot region based on the position of the target detection box in the foreground image.
[0057] In this embodiment, a detection object comprising a target detection image and detection semantics is acquired. The detection semantics of the detection object are input into a pre-trained Florence-2 model for semantic understanding, and a semantic parsing vector is generated based on the semantic understanding result. The target detection image can be acquired from a business platform or a locally stored image requiring foot region segmentation. Furthermore, the detection semantics are foot region segmentation instructions initiated based on the target detection image. The detection semantics can be directly input into the Florence-2 model for parsing; therefore, the detection semantics can also be defined as instruction information that the Florence-2 model can recognize. Specifically, it can be set based on the instruction object defined by the Florence-2 model; this embodiment does not limit this.
[0058] The Florence-2 model parses the detection semantics to obtain semantic parsing results, which are then used to label the foot region detection results in the target detection image. In practical applications, both the target detection image and the detection semantics can be input into the Florence-2 model. After parsing the detection semantics, the Florence-2 model labels the foot region detection results in the target detection image based on the parsing results and outputs a target detection image labeled with the foot region detection results. The pre-trained Florence-2 model refers to a visual language model trained on large-scale multimodal data, specifically implemented using a cross-modal encoder based on the Transformer architecture, used to map natural language instructions into visual positioning information. (See attached image.) Figure 3 , Figure 3 This is a schematic diagram illustrating the working principle of the Florence-2 model. Based on... Figure 3The schematic diagram illustrating the working principle of the Florence-2 model indicates that its data processing workflow may include the following steps:
[0059] 1. Use DaViT to extract the visual features of the target detection object;
[0060] 2. Encode the text containing the detected semantics using a BART decoder:
[0061] 3. The visual features of the target detection object and the textual features of the detection semantics are fused using a multimodal encoder;
[0062] 4. Generate bounding box coordinates from the fusion result using a decoder;
[0063] 5. Filter the bounding box coordinates using NMS;
[0064] 6. Return the detection box (bbox) based on the filtering results, which is the indicator box of the detection area.
[0065] Based on the target detection image with foot region detection results output by the Florence-2 model, the foreground image is extracted from the target detection image using the BiRefNet model. Considering the application in the specific implementation process, the foreground image can be extracted from the target detection image without foot region detection results, or from the target detection image with foot region detection results already labeled. The specific settings can be made according to the specific data structure of the BiRefNet model.
[0066] After parsing the detection semantics according to the Florence-2 model and returning the target detection image labeled with the foot region detection results, since the foot region detection results labeled by the Florence-2 model are represented as detection boxes, and the detection boxes may have large detection boxes surrounding small detection box regions based on the parsing of the detection semantics, that is, there are one or more foot detection results, the large detection boxes in the target detection image are removed by the reverse NMS algorithm, leaving only the foot region as the target detection box. The reverse NMS algorithm is an improved non-maximum suppression strategy algorithm, defined as Reverse NMS or Inverse NMS, used to retain potential effective boxes filtered by conventional NMS to improve recall or handle dense target scenes.
[0067] Next, the detected object image is input into the BiRefNet model to extract the foreground image. The detected object image input into the BiRefNet model can be either an image with the object detection bounding box labeled or an image without the bounding box labeled, depending on the relevant settings. The BiRefNet model is a powerful foreground segmentation model that achieves progressive segmentation through a localization module (LM) and a reconstruction module (RM). The working principle of its data structure can be found in [link to documentation]. Figure 4 , Figure 4 This is a schematic diagram illustrating the working principle of the BiRefNet model. Specifically, the working principle steps of the BiRefNet model are as follows:
[0068] 1. Perform image preprocessing on the input target detection image;
[0069] 2. Use the LM algorithm to locate the segmentation module in the preprocessed target detection image;
[0070] 3. The segmentation model is reconstructed using the RM algorithm;
[0071] 4. Output the reconstructed segmentation module and process it to obtain the foreground image.
[0072] Based on the foreground image of the obtained target detection image, the position of the target detection box in the target detection image is mapped onto the foreground image, and the target position in the foreground image is cropped as the foot region, which is the segmentation result.
[0073] In this embodiment, the Florence-2 model is used to achieve accurate semantic parsing, combined with reverse NMS to optimize the detection box localization, and BiRefNet is used for fine segmentation. This enables accurate segmentation of the foot region without training, while maintaining the consistency of shoe shape, providing a reliable foundation for virtual shoe changing.
[0074] Specifically, the technical content described in step S10 above is further analyzed, namely, the step of using a pre-trained Florence-2 model to parse the detection semantics and marking the foot region detection result in the target detection image according to the parsing result of the detection semantics includes:
[0075] The detected semantics are input into the pre-trained Florence-2 model for semantic understanding, and a semantic parsing vector is generated based on the semantic understanding result;
[0076] Based on the semantic parsing vector, candidate bounding boxes for the foot region are located in the target detection image;
[0077] The candidate bounding boxes for the foot region are filtered by confidence, and the candidate bounding boxes for the foot region that meet the preset conditions are taken as the detection results of the foot region.
[0078] In this embodiment, the pre-trained Florence-2 model is used to parse the detection semantics of the detected object to obtain a semantic understanding result, and the semantic understanding result is generated into a semantic parsing vector. The semantic understanding result is essentially the object region cropped from the foot region, which is obtained after semantic parsing by the Florence-2 model. When generating the semantic parsing vector from the semantic understanding result, the output result of the Florence-2 model can be used as the standard, that is, the output result of the Florence-2 model can be represented as a semantic parsing vector. The semantic parsing vector is a high-dimensional feature representation output by the Florence-2 model, implemented in the form of an embedding space vector, and is used to describe the correlation between the detection semantics and the image region.
[0079] Specifically, when a user inputs a natural language command containing a requirement for foot region segmentation (i.e., detection semantics), the Florence-2 model first performs semantic parsing on the detection semantics to generate a semantic parsing vector representing the semantic intent. This semantic parsing vector interacts with the image features of the target detection image through a cross-modal attention mechanism. Based on the interaction result, multiple candidate boxes that may contain the foot region are located in the target detection image, i.e., foot region candidate boxes. Since the semantic parsing vector is obtained based on the detection semantics, and the Florence-2 model outputs multiple semantic detection vectors during the parsing process, multiple foot region candidate boxes are located in the target detection image based on these semantic detection vectors. Figure 5 , Figure 5 This is a schematic diagram illustrating the process of capturing the footstep region from an image based on target detection.
[0080] For a target detection image with multiple labeled candidate bounding boxes for the foot region, a confidence screening process is performed on each candidate bounding box. Confidence screening involves filtering based on the quantified degree of matching between the candidate bounding box and the semantic instruction. Specifically, cosine similarity or probability thresholding can be used to eliminate erroneous localizations caused by semantic understanding biases. That is, the step of performing confidence screening on the candidate bounding boxes for the foot region, and selecting those meeting preset conditions as the foot region detection result, includes:
[0081] Calculate the semantic matching score of each foot region candidate box, and remove the foot region candidate boxes whose semantic matching score is lower than a preset threshold;
[0082] Non-maximum suppression is applied to the remaining candidate bounding boxes of the foot region to remove redundant detection boxes and obtain the detection result of the foot region.
[0083] When performing confidence screening on multiple labeled foot region candidate boxes in the target detection image, foot region candidate boxes that meet preset conditions are taken as foot region detection results. The specific process can be defined as follows: calculate the semantic matching score of each foot region candidate box, and remove the foot region candidate boxes whose semantic matching score is lower than a preset threshold; perform non-maximum suppression processing on the remaining foot region candidate boxes to remove redundant detection boxes and obtain the foot region detection results.
[0084] The process of calculating the matching degree between the foot region candidate boxes and the semantic vector can be characterized by using vector dot product or similarity matrix operations to filter out foot region candidate boxes with scores higher than a preset threshold. The semantic matching degree score refers to the quantified value of the correlation between the foot region candidate boxes and the detected semantics. Specifically, it can be achieved by calculating the cosine similarity between the image features corresponding to the foot region candidate boxes and the text features of the detected semantics, used to measure whether the foot region candidate boxes accurately reflect the user's semantic needs. The non-maximum suppression processing refers to an algorithm that filters overlapping candidate boxes using the intersection-union ratio (IU) threshold, specifically implemented using the standard NMS algorithm, to eliminate redundant results of the same target being detected multiple times.
[0085] This processing is based on directly parsing the semantic details in the natural language instructions represented by the detected semantics. In this embodiment, the semantic parsing capability of the Florence-2 model is used to transform abstract instructions into quantifiable visual positioning criteria. Combined with a confidence-based filtering mechanism, positioning errors caused by semantic ambiguity are effectively eliminated. This solves the problem of insufficient understanding of natural language instructions by traditional segmentation models, enabling accurate positioning of the foot region based on diverse semantic needs of user input. No model retraining is required; only adjustments to the semantic instructions are needed for rapid adaptation, significantly improving the system's generalization ability and interactive flexibility.
[0086] Specifically, step S30 above, namely the step of performing a reverse NMS operation on the foot region detection results to obtain the target detection box, includes:
[0087] Iterate through the detection results of each foot region and calculate the overlapping area between the detection results of each foot region;
[0088] The foot region detection result containing the most overlapping area is used as the target detection box. When there are multiple foot region detection results with the same overlapping area, the foot region detection result with the largest area is used as the target detection box.
[0089] In this embodiment, the reverse NMS operation is further defined, specifically including traversing the detection results of each foot region and calculating the overlap area between the detection results of each foot region. The detection result of the foot region containing the most overlap area is used as the target detection box. Alternatively, when there are multiple detection results of foot regions with the same overlap area, the detection result of the foot region with the largest overlap area is used as the target detection box. The target detection box refers to the bounding box of the foot region that is finally determined for subsequent segmentation. It is generated by comparing the sum of the overlap areas or directly selecting the largest area box. Its function is to ensure that the selected region can cover the complete outline of the foot.
[0090] The reverse NMS operation refers to a process that enhances region aggregation by retaining detection boxes with more overlapping areas. Specifically, this can be achieved by traversing the detection boxes represented by the foot region detection results and calculating the sum of the intersection-union ratios (IUR) between each detection box. This method avoids the region omission problem caused by a single threshold selection in traditional non-maximum suppression operations. Based on the traversal results, the overlap area between the detection boxes represented by the foot region detection results can be obtained. The overlap area refers to the proportion of the intersection area to the union area between two detection boxes, which can be calculated using coordinate differences. This quantifies the spatial correlation between detection boxes and selects the most representative region as the overlapping region. Specifically, the step of using the foot region detection result containing the most overlapping areas as the target detection box includes:
[0091] Calculate the cross-union ratio (CUC) among the detection results of each foot region;
[0092] The target detection box is determined from the foot region detection results based on the sum of the intersection-union ratios.
[0093] The intersection-union ratio (IUR) is calculated for the detection boxes represented by the detection results of each foot region. Specifically, the IUR is the ratio of the intersection area to the union area between two detection boxes represented by the detection results of the foot regions. This can be achieved by calculating the intersection and union areas using the coordinate parameters of the detection boxes, thereby quantifying the degree of overlap between the detection boxes. The maximum sum of IURs refers to the sum of the IURs calculated for each detection box with all other detection boxes. Based on the sum of IURs, the overlap of the detection boxes in each foot detection region is determined, and the size of the detection boxes in each foot detection region is determined based on the overlap. Specifically, based on the determination of the detection boxes, the detection boxes corresponding to the larger foot detection regions are removed, and the smaller detection boxes are retained as the target detection boxes. By comprehensively evaluating the overall overlap between the detection boxes and the surrounding areas, misselection caused by local overlap interference can be effectively avoided. Figure 6 , Figure 6This is a schematic diagram confirming the target detection bounding box.
[0094] Specifically, in the reverse NMS operation, all the pre-screened foot region detection results are first traversed, and the intersection-union ratio (IU) of each detection box with other detection boxes is calculated. Then, the detection box with the largest IU is selected as a candidate target; when multiple candidate targets exist, the detection box with the largest area is prioritized. This operation effectively integrates potentially scattered local detection results by retaining detection boxes with high spatial correlation, thereby forming a bounding box covering the entire foot region. For example, when multiple partially overlapping foot region candidate boxes are detected, this method can determine the most representative main region by calculating the IU, avoiding the loss of other effective regions due to the traditional NMS operation of retaining only a single high-resolution box.
[0095] In this embodiment, the reverse NMS operation actively retains and aggregates highly overlapping detection boxes, adapting to complex situations such as foot pose changes and occlusion, thus improving the robustness of region localization. It effectively solves the problem of missing regions caused by excessive suppression in traditional methods for foot region detection, ensuring that subsequent segmentation steps obtain a complete foot contour. Simultaneously, it reduces reliance on manually set thresholds, improving the algorithm's adaptability to diverse scenarios.
[0096] Specifically, according to the technical content of step S40 above, the step of extracting the foreground image from the target detection image using the BiRefNet model, and capturing the target position as the foot region based on the position of the target detection box in the foreground image, includes:
[0097] The target detection image is input into the BiRefNet model to obtain the multi-scale feature map output by the BiRefNet model;
[0098] Foreground separation is performed based on the multi-scale feature map to generate a preliminary foreground mask;
[0099] The initial foreground mask is optimized to obtain the foreground image.
[0100] In this embodiment, the multi-scale feature map is an image obtained by processing the target detection image based on the BiRefNet model. The multi-scale feature map is essentially a set of features containing local details and global semantic information extracted from the target detection image by convolutional neural networks of different levels. It is implemented using a pyramid structure or a cross-layer connection structure, which can capture the morphological features of the foot region at different resolutions.
[0101] Based on the multi-scale feature map output by the BiRefNet model, foreground separation is performed on the multi-scale feature map. The process of distinguishing the foreground and background regions in the target detection image using an attention mechanism or pixel-level classification algorithm can be specifically implemented using a joint optimization of channel attention and spatial attention modules to enhance the contrast between the foreground and background images in the target detection image. Foreground separation is performed using this contrast to obtain a preliminary foreground mask. This preliminary foreground mask is a probability distribution map of the foreground image generated through binarization, implemented using thresholding or soft masking, and can be used to initially locate the boundaries of the foreground image. Then, based on these boundaries, the foreground image is segmented from the target detection image to obtain the foreground image. The optimization process for the preliminary foreground mask essentially refines the foreground image of the preliminary segmentation result, eliminating jagged noise at the mask edges and filling internal holes. This optimization process can be implemented using a conditional random field algorithm or an edge-refining convolutional layer.
[0102] As shown above, based on the workflow of the BiRefNet model, it can be seen that the BiRefNet model essentially extracts multi-scale features from the target detection image through a parallel encoder structure. This parallel encoder structure contains shallow and deep networks. The shallow network captures detailed features such as foot edges and textures, while the deep network extracts the overall shape and spatial position information of the foot. The multi-scale feature maps are then processed by a feature fusion module through channel weighting and spatial stitching, and input to the foreground separation branch to generate a preliminary foreground mask. This mask is iteratively corrected by an optimization module, for example, by using a bidirectional propagation mechanism to interact deep semantic information with shallow detailed features, gradually eliminating misclassified pixels in background interference areas. In the final output foreground image, the foot region has smooth boundaries and continuous internal regions, providing a high-precision spatial localization basis for the coordinate mapping of the target detection box.
[0103] Compared to existing technologies, traditional methods typically use single-scale features for foreground segmentation, resulting in insufficient segmentation accuracy for small-scale foot regions (such as bare toes when wearing sandals) or large-scale regions (such as the leg extension when wearing boots). Our proposed solution, however, utilizes multi-scale feature fusion to simultaneously adapt to foot morphological variations at different scales. Combined with the dynamic correction capabilities of the optimization module, it significantly reduces the interference of complex backgrounds (such as carpet textures and shadows) on the segmentation results.
[0104] Through the above technical solution, this application solves the problem of blurred segmentation boundaries in the foot region caused by single feature extraction in existing technologies, and achieves high-precision extraction of foot contours in virtual try-on scenarios. This solution can handle the segmentation needs of new product categories such as socks and anklets without relying on specific training data, while reducing the cost of manual annotation and repeated model training.
[0105] Based on the technical content of step S40 in the above embodiment, the step of extracting the foreground image from the target detection image using the BiRefNet model and capturing the target position as the foot region according to the position of the target detection box in the foreground image includes:
[0106] The coordinates of the target detection box are mapped to the foreground image, and candidate foreground for the foot region are cropped based on the coordinate mapping;
[0107] Edge optimization is performed on the candidate foreground of the foot region to obtain the foot region.
[0108] In this embodiment, after extracting the foreground image from the target detection image using the BiRefNet model, the coordinates of the determined target detection box are mapped onto the foreground image, and candidate foreground regions of the foot area are cropped. Specifically, the coordinate mapping based on the coordinates involves transforming the geometric position information of the target detection box from the feedforward network output layer to the pixel coordinate system using affine transformation or bilinear interpolation, thereby ensuring spatial alignment between the target detection box and the foreground image. The candidate foreground region of the foot area is a set of pixels in the foreground image that overlap with the coordinates of the target detection box. This set is generated by matrix cropping or region masking and then superimposing the pixel set, thus initially filtering out foreground regions that may contain the foot area. For details on the process of extracting and filtering foreground images containing foot areas based on the foreground image, please refer to [link to documentation]. Figure 5 , Figure 5 This is a schematic diagram illustrating the process of capturing the footstep region from an image based on object detection.
[0109] Based on the initially selected candidate foreground for the foot region, morphological operations combined with edge detection filters are used to refine the contours of the candidate foreground for the foot region in order to eliminate burrs or broken areas in the segmentation results, thereby optimizing the edges of the panoramic view of the foot region. That is, the step of edge optimization of the candidate foreground for the foot region as the foot region includes:
[0110] Morphological closing operations are used to fill the internal holes of the candidate foreground of the foot region, and non-foot noise regions in the candidate foreground of the foot region are removed by edge detection algorithms;
[0111] Output a smooth and continuous foot region segmentation mask to optimize the edges of the foot region candidate foreground.
[0112] Since the foot area in the target detection region is likely to be covered by shoes, socks, or other accessories, there may be holes in the candidate foreground of the foot area cropped from the foreground image. Therefore, a morphological closing algorithm is used to process these holes. When processing the internal holes of the foot area using this morphological closing operation, the image is essentially processed through a dilation-erosion operation, which can be implemented based on rectangular or elliptical structural elements. For example, a 5x5 pixel rectangular kernel can be used to perform the closing operation to eliminate small holes inside the foot area caused by occlusion or uneven lighting. Figure 7 As shown, Figure 7 This is a schematic diagram for filling holes.
[0113] After filling the holes in the candidate foreground of the foot region, the edges of the candidate foreground of the foot region are processed by an edge detection algorithm. The purpose of this processing is to sharpen the boundaries of regions with abrupt gray-level changes in the image represented by the candidate foreground of the foot region. This can be achieved by using the Canny operator or the Sobel operator to process the boundaries of these regions. For example, the Canny algorithm with dual threshold parameters of 50 and 150 can be used to distinguish the edges of the real candidate foreground of the foot region from background image noise. The sharpening process outputs a smooth and continuous foot region segmentation mask, thereby optimizing the edges of the candidate foreground of the foot region. The foot region segmentation mask is a set of pixels marking the target region in the binarized image. Interpolation algorithms or morphological thinning algorithms are used to optimize the edge continuity. For example, cubic spline interpolation can be used to transform jagged edges into smooth curves.
[0114] Specifically, after obtaining candidate foreground regions of the foot, morphological closing operations are first performed on these regions. When internal holes formed by gaps in shoelaces or shadows exist in the candidate foreground, the closing operation fills these discontinuous areas, forming a complete and closed foot contour. Then, an edge detection algorithm scans the boundaries of the candidate foreground, identifying the difference between the real foot edges and background noise through gradient changes. For example, the transition area between the sock edge and skin texture is preserved, while clothing wrinkles or background blemishes are filtered out. Finally, the optimized edge coordinates are converted into a segmentation mask, and an interpolation algorithm is used to eliminate pixel-level jaggedness, generating smooth and continuous segmentation boundaries.
[0115] This embodiment improves the accuracy and continuity of foot region segmentation edges, especially when processing feet with decorations or complex textures, effectively eliminating internal holes and external noise interference. This solution allows the virtual try-on system to obtain accurate segmentation results without retraining the model when capturing new categories of items such as socks and anklets, while the generated smooth edges avoid pixel-level misalignment problems when virtual footwear is worn.
[0116] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the foot region segmentation method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0117] This application provides a foot region segmentation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the foot region segmentation method in the first embodiment described above.
[0118] The following is for reference. Figure 2 The diagram illustrates a structural schematic suitable for implementing a foot region segmentation device according to embodiments of this application. The foot region segmentation device in embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 2 The foot region segmentation device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0119] like Figure 2As shown, the foot region segmentation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the foot region segmentation device. The processing unit 1001, the ROM 1002, and the RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the foot segmentation device to communicate wirelessly or wiredly with other devices to exchange data. Although foot segmentation devices with various systems are shown in the figures, it should be understood that it is not required to implement or possess all of the systems shown. More or fewer systems may be implemented alternatively.
[0120] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0121] The foot region segmentation device provided in this application, employing the foot region segmentation method described in the above embodiments, can solve the technical problem of segmentation defects in the foot region caused by inadequate semantic understanding in existing segmentation models. Compared with the prior art, the beneficial effects of the foot region segmentation device provided in this application are the same as those of the foot region segmentation method provided in the above embodiments, and other technical features of this foot region segmentation device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0122] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0123] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0124] This application provides a storage medium, which is a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the foot region segmentation method in the above embodiments.
[0125] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0126] The aforementioned computer-readable storage medium may be included in the foot region segmentation device; or it may exist independently and not assembled into the foot region segmentation device.
[0127] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the foot region segmentation device, enable the foot region segmentation device to implement the technical content of the foot region segmentation method embodiment shown above.
[0128] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0130] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0131] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described foot region segmentation method. This solves the technical problem of segmentation defects in the foot region caused by inadequate semantic understanding in existing segmentation models. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the foot region segmentation method provided in the above embodiments, and will not be elaborated upon here.
Claims
1. A method for segmenting the foot region, characterized in that, The foot region segmentation method includes the following steps: Obtain the detection object, which includes the target detection image and the detection semantics; The detection semantics are parsed using a pre-trained Florence-2 model, and the foot region detection results are marked in the target detection image based on the parsing results of the detection semantics. Perform a reverse NMS operation on the foot region detection results to obtain the target detection box; The BiRefNet model extracts the foreground image from the object detection image, and captures the target location as the foot region based on the position of the object detection box in the foreground image.
2. The foot region segmentation method as described in claim 1, characterized in that, The step of using a pre-trained Florence-2 model to parse the detection semantics and marking the foot region detection results in the target detection image based on the parsing results of the detection semantics includes: The detected semantics are input into the pre-trained Florence-2 model for semantic understanding, and a semantic parsing vector is generated based on the semantic understanding result; Based on the semantic parsing vector, candidate bounding boxes for the foot region are located in the target detection image; The candidate bounding boxes for the foot region are filtered by confidence, and the candidate bounding boxes for the foot region that meet the preset conditions are taken as the detection results of the foot region.
3. The foot region segmentation method as described in claim 2, characterized in that, The step of performing confidence screening on the candidate bounding boxes of the foot region and selecting the candidate bounding boxes of the foot region that meet the preset conditions as the foot region detection result includes: Calculate the semantic matching score of each foot region candidate box, and remove the foot region candidate boxes whose semantic matching score is lower than a preset threshold; Non-maximum suppression is applied to the remaining candidate bounding boxes of the foot region to remove redundant detection boxes and obtain the detection result of the foot region.
4. The foot region segmentation method as described in claim 1, characterized in that, The step of performing a reverse NMS operation on the foot region detection results to obtain the target detection box includes: Iterate through the detection results of each foot region and calculate the overlapping area between the detection results of each foot region; The foot region detection result containing the most overlapping area is used as the target detection box. When there are multiple foot region detection results with the same overlapping area, the foot region detection result with the largest area is used as the target detection box.
5. The foot region segmentation method as described in claim 4, characterized in that, The step of using the detection result of the foot region containing the most overlapping area as the target detection box includes: Calculate the cross-union ratio (CUC) among the detection results of each foot region; The target detection box is determined from the foot region detection results based on the sum of the intersection-union ratios.
6. The foot region segmentation method as described in claim 1, characterized in that, The step of extracting the foreground image from the object detection image using the BiRefNet model, and capturing the target location as the foot region based on the position of the object detection box in the foreground image, includes: The target detection image is input into the BiRefNet model to obtain the multi-scale feature map output by the BiRefNet model; Foreground separation is performed based on the multi-scale feature map to generate a preliminary foreground mask; The initial foreground mask is optimized to obtain the foreground image.
7. The foot region segmentation method as described in claim 1, characterized in that, The step of extracting the foreground image from the object detection image using the BiRefNet model, and capturing the target location as the foot region based on the position of the object detection box in the foreground image, includes: The coordinates of the target detection box are mapped to the foreground image, and candidate foreground for the foot region are cropped based on the coordinate mapping; Edge optimization is performed on the candidate foreground of the foot region to obtain the foot region.
8. The foot region segmentation method as described in claim 7, characterized in that, The step of edge optimization of the candidate foreground region of the foot region as the foot region includes: Morphological closing operations are used to fill the internal holes of the candidate foreground of the foot region, and non-foot noise regions in the candidate foreground of the foot region are removed by edge detection algorithms; Output a smooth and continuous foot region segmentation mask to optimize the edges of the foot region candidate foreground.
9. A foot area segmentation device, characterized in that, The foot region segmentation device stores a computer program, which, when executed by a processor, implements the foot region segmentation method according to any one of claims 1-8.
10. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the foot region segmentation method according to any one of claims 1-8.