A Semantic Segmentation-Based Method for Guiding Group Photos at Self-Service Photo Booths

CN122679334APending Publication Date: 2026-09-01SHENZHEN EASY TOUCH TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610474493.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-11
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

然而对设备安装精度及多传感器标定提出了较高要求,在实际应用中存在维护复杂、稳定性差问题,难以适用于成本敏感且空间受限的自助照相亭场景

Benefits of technology

[0060] (1) This invention introduces a multi-resolution cascaded YOLOv11-Seg instance segmentation model and combines it with a high-resolution supplementary inference mechanism driven by occlusion risk. This enables high-precision pixel-level instance segmentation of multi-person high-density overlapping scenes under edge computing conditions. Candidate instance masks are quickly obtained through low-resolution inference. Then, occlusion risk scores are calculated based on the intersection-union ratio occlusion metric, boundary coverage metric, and single-unit anomaly penalty term. High-resolution fine inference is performed only on high-risk areas. This significantly improves the segmentation accuracy of complex occlusion boundaries while ensuring real-time performance. It can accurately restore the real contour boundaries of the human body and provide high signal-to-noise ratio input data for occlusion relationship determination. It maintains stable segmentation performance even when multiple people are in close contact, clothing colors are similar, and local occlusion is severe.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122679334A_ABST
    Figure CN122679334A_ABST
Patent Text Reader

Abstract

This invention discloses a self-service photo booth multi-person group photo guidance method based on semantic segmentation, belonging to the field of self-service photo booth technology. The method includes: Step 1, obtaining a standardized input image sequence; Step 2, inputting the standardized input image sequence into an improved YOLOv11-Seg instance segmentation model deployed on edge computing nodes, outputting an initial instance mask set and corresponding human keypoint set; Step 3, generating a refined instance mask set; Step 4, constructing a directed topology graph of occlusion relationships; Step 5, obtaining a hierarchically compressed directed topology graph of occlusion relationships; Step 6, generating a candidate pose adjustment strategy set; Step 7, mapping to spatially directional multimodal guidance instructions; Step 8, completing the multi-person group photo capture. This invention achieves highly consistent and interpretable multimodal guidance effects, significantly improving the stability and user experience of multi-person collaborative shooting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of self-service photo booth technology, and more particularly to a self-service photo booth method for guiding multiple people to take group photos based on semantic segmentation. Background Technology

[0002] With the widespread use of self-service photo booths in commercial entertainment, tourist attractions and social settings, the demand for automated group photos is gradually increasing. In situations where multiple people stand in high density, occlusion between people is inevitable.

[0003] In existing technologies, most self-service photo booth systems rely on face detection or body detection boxes for target recognition. These typically use rectangular boxes for coarse-grained human body localization, failing to accurately describe body boundaries and limb contours. When multiple people stand close together, the overlapping of detection boxes makes it difficult for the system to distinguish the true spatial relationships between different individuals, easily leading to recognition confusion when clothing colors are similar or limbs overlap.

[0004] To overcome the aforementioned problems, some existing technologies attempt to introduce binocular vision, structured light, or depth camera 3D perception devices to obtain spatial depth information of the human body. However, these technologies place high demands on the installation accuracy of the devices and the calibration of multiple sensors, resulting in complex maintenance and poor stability in practical applications, making them unsuitable for cost-sensitive and space-constrained self-service photo booth scenarios.

[0005] In the few systems with automatic guidance capabilities, a local judgment strategy based on a single target is typically employed. That is, when a user is detected to be occluded, a movement command is directly issued to that user. This decision-making approach based on isolated objects lacks the ability to model global spatial relationships, and is prone to command conflicts in densely populated environments. For example, the movement of one user may cause secondary occlusion of other users, leading to the system issuing conflicting guidance commands continuously, resulting in command oscillations or even cyclical adjustments, which seriously affects user experience and shooting efficiency. Summary of the Invention

[0006] One objective of this invention is to propose a self-service photo booth method for guiding multiple people to take photos based on semantic segmentation. This invention achieves highly consistent and interpretable multimodal guidance, significantly improving the stability and user experience of multi-person collaborative shooting.

[0007] A self-service photo booth group photo guidance method based on semantic segmentation according to an embodiment of the present invention includes:

[0008] Step 1: Use the imaging equipment inside the self-service photo booth to continuously acquire the original group photo video stream and perform preprocessing operations to obtain a standardized input image sequence;

[0009] Step 2: Input the standardized input image sequence into the improved YOLOv11-Seg instance segmentation model deployed on edge computing nodes, and output the initial instance mask set and the corresponding human key point set;

[0010] Step 3: Based on the set of human body key points, refine the initial instance mask set by key point prior auxiliary masking to generate a refined instance mask set;

[0011] Step 4: Perform pixel-level overlap relationship calculation and key region occlusion rate calculation on any two human body instances in the refined instance mask set to obtain the occlusion relationship pair set and construct the occlusion relationship directed topology graph;

[0012] Step 5: Perform a layer consistency determination on the occlusion relationship directed topology graph within multiple consecutive frames. If the determination result meets the merging condition, merge nodes at the same layer into block nodes to obtain a layer-compressed occlusion relationship directed topology graph.

[0013] Step 6: Input the directed topology graph of hierarchical compressed occlusion relationship into the graph neural network model to generate a set of candidate pose adjustment strategies;

[0014] Step 7: Perform a global occlusion weight change evaluation on the candidate attitude adjustment strategy set to obtain the optimal attitude adjustment strategy and map it into a spatially directional multimodal guidance command.

[0015] Step 8: Within the preset monitoring period after the multimodal guidance command is output, continuously collect the updated standardized input image sequence, repeat steps 2 to 7, and update the hierarchical compression occlusion relationship directed topology graph and optimal pose adjustment strategy in real time. When the global occlusion weight calculated based on the hierarchical compression occlusion relationship directed topology graph is lower than the preset occlusion convergence threshold and remains constant within the set stable time window, trigger the automatic group photo acquisition operation to complete the capture of multi-person group photos.

[0016] Optionally, the preprocessing operations include distortion correction, illumination equalization, and noise suppression.

[0017] The set of normalized input image sequences consists of several frames of normalized input images, with each frame corresponding to a time index.

[0018] Optionally, step two includes:

[0019] The normalized input image of any frame is input into the low-resolution inference branch of the improved YOLOv11-Seg instance segmentation model deployed on the edge computing node, and a fixed-size low-resolution resampling process is performed on the normalized input image to obtain the low-resolution input image.

[0020] The low-resolution input image is input into the low-resolution inference branch of the improved YOLOv11-Seg instance segmentation model to obtain low-resolution candidate boxes, low-resolution mask response sets, and low-resolution keypoint response sets.

[0021] Based on the low-resolution mask response set, a threshold determination is performed on the mask response corresponding to each low-resolution candidate box. Pixels with mask response values ​​greater than or equal to the preset mask threshold are marked as foreground pixels, and the remaining pixels are marked as background pixels, thus generating the corresponding low-resolution candidate instance mask.

[0022] Based on the low-resolution keypoint response set, the keypoint response corresponding to each low-resolution candidate box is decoded to obtain the corresponding low-resolution human keypoint set.

[0023] For each low-resolution candidate instance mask in the standardized input image, calculate the occlusion risk score, filter out the indices corresponding to the low-resolution candidate instance masks whose occlusion risk scores are greater than or equal to the preset occlusion risk judgment threshold, and construct an occlusion risk region index set.

[0024] For each index in the occlusion risk region index set, a cropping operation with extended boundaries is performed based on the position of the corresponding low-resolution candidate box in the original frame normalized input image to obtain an occlusion risk region image patch. The occlusion risk region image patch is then input into the high-resolution supplementary inference branch of the improved YOLOv11-Seg instance segmentation model. After high-resolution resampling processing, instance segmentation inference is performed to obtain a high-resolution mask response map and a high-resolution key point response map.

[0025] A high-resolution candidate instance mask is generated based on the high-resolution mask response map. A high-resolution human key point set is obtained by decoding the high-resolution key point response map. The high-resolution human key point set is then back-mapped to the image coordinate system of the original frame normalized input image through spatial coordinate mapping to obtain the back-mapped high-resolution instance mask and the back-mapped high-resolution human key point set.

[0026] For low-resolution candidate instance masks and corresponding human keypoint sets that are not included in the occlusion risk area index set, coordinate backmapping is performed. The back-mapped high-resolution instance mask and the back-mapped low-resolution instance mask are then fused together with the occlusion risk area index set. Cross-frame temporal association is then performed to obtain an initial instance mask set and human keypoint set carrying consistent identity labels.

[0027] Optionally, step three includes:

[0028] For each initial instance mask set and the human body key point set, a key point prior response field is established for each initial instance mask and its corresponding human body key point set.

[0029] Based on the preset connection relationships between key points in the human body key point set, a priori response field for skeleton connection is established.

[0030] Normalization is performed on the prior response fields of key points and skeleton connections respectively to obtain normalized prior response fields of key points and normalized prior response fields of skeleton connections.

[0031] The initial instance mask, the normalized keypoint prior response field, and the normalized skeleton connection prior response field are fused at the pixel level to obtain a refined mask response map. Threshold segmentation is then performed based on the refined mask response map to generate a refined instance mask.

[0032] Perform connected component partitioning on each refined instance mask to obtain the refined instance mask connected component. Determine whether the human key points in the human key point set fall inside the refined instance mask connected component to obtain the final refined instance mask.

[0033] All final refined instance masks are constructed into a set of refined instance masks, and the boundary contours, occupied areas, and image coordinate system positions are extracted.

[0034] Optionally, step four includes:

[0035] For any two different human body instances in the refined instance mask set, perform pixel-level overlap relationship calculation to obtain the full mask occlusion rate;

[0036] Based on the set of human body key points, construct the key region mask corresponding to each human body instance, and calculate the key region occlusion rate.

[0037] Based on the full mask occlusion rate and the key area occlusion rate, the directional occlusion judgment value is calculated to obtain the occlusion relationship pair;

[0038] Based on the full mask occlusion rate, the key area occlusion rate, and the proximity of the center position, calculate the corresponding edge weights of the occlusion relationship;

[0039] Using each human instance in the refined instance mask set as a node and the occlusion relationship pair set as directed edges, a directed topology graph of occlusion relationships is constructed. The edge weights in the directed topology graph of occlusion relationships are written to obtain the directed topology graph of occlusion relationships used for hierarchical consistency determination.

[0040] Optionally, step five includes:

[0041] A cross-frame node correspondence is established based on the occlusion relationship directed topology graph to form a node tracking sequence;

[0042] For each occlusion graph node in the directed topology graph of each frame, calculate the node level value;

[0043] Based on the node hierarchy values ​​within multiple consecutive frames, a hierarchy consistency determination is performed. For nodes at the same level that meet the merging conditions, block node merging processing is performed to obtain a block node set.

[0044] Step 54: Reconstruct the directed edges based on the block node set to obtain a directed topological graph of hierarchical compression and occlusion relationships.

[0045] Optionally, step six includes:

[0046] The hierarchical compressed occlusion relationship directed topology graph is input into the graph neural network model, and block node input features are constructed.

[0047] Based on the hierarchical compression occlusion relationship directed topology graph and block node input features, the graph neural network model is used for inference to obtain the action reward score corresponding to each block node.

[0048] The block node attitude adjustment judgment value is calculated by weighted summation based on the block node in-degree, the cumulative value of the block node in-edge weight, and the action gain score.

[0049] Based on the attitude adjustment judgment value of the block node, a set of candidate attitude adjustment strategies is generated.

[0050] Optionally, the criteria for determining the candidate pose adjustment strategy category are as follows:

[0051] When the block node attitude adjustment judgment value is less than the preset starting policy threshold, it is determined that the corresponding block node does not need to perform attitude adjustment at present, and no candidate attitude adjustment policy is generated for the block node.

[0052] When the block node attitude adjustment judgment value is greater than or equal to the preset starting strategy threshold and falls within the first preset strategy threshold range, the candidate attitude adjustment strategy category corresponding to the block node is determined as the probe candidate attitude adjustment strategy.

[0053] When the block node attitude adjustment judgment value falls into the second preset strategy threshold range, the candidate attitude adjustment strategy category corresponding to the block node is determined as a lateral offset candidate attitude adjustment strategy with left / right spatial pointing vectors, based on the spatial relative relationship between the block node centroid coordinates and the center axis of the screen.

[0054] When the block node attitude adjustment judgment value falls within the third preset policy threshold range, the candidate attitude adjustment policy category corresponding to the block node is determined as the longitudinal yielding candidate attitude adjustment policy.

[0055] Optionally, step seven includes:

[0056] For each candidate pose adjustment strategy, a global occlusion weight change evaluation is performed to obtain the global occlusion weight change corresponding to each candidate pose adjustment strategy;

[0057] Based on the change in global occlusion weight and the action gain score, the global gain corresponding to the candidate pose adjustment strategy is calculated, and the pose adjustment strategy with the largest global gain is selected as the optimal pose adjustment strategy.

[0058] Based on the optimal attitude adjustment strategy, a spatially directional multimodal guidance command is generated and output to the corresponding user in real time.

[0059] The beneficial effects of this invention are:

[0060] (1) This invention introduces a multi-resolution cascaded YOLOv11-Seg instance segmentation model and combines it with a high-resolution supplementary inference mechanism driven by occlusion risk. This enables high-precision pixel-level instance segmentation of multi-person high-density overlapping scenes under edge computing conditions. Candidate instance masks are quickly obtained through low-resolution inference. Then, occlusion risk scores are calculated based on the intersection-union ratio occlusion metric, boundary coverage metric, and single-unit anomaly penalty term. High-resolution fine inference is performed only on high-risk areas. This significantly improves the segmentation accuracy of complex occlusion boundaries while ensuring real-time performance. It can accurately restore the real contour boundaries of the human body and provide high signal-to-noise ratio input data for occlusion relationship determination. It maintains stable segmentation performance even when multiple people are in close contact, clothing colors are similar, and local occlusion is severe.

[0061] (2) This invention realizes the ability to infer the relative front-back relationship of multiple people under monocular two-dimensional imaging by constructing a directed occlusion relationship topology graph based on refined instance masks. By calculating the full mask occlusion rate and the key region occlusion rate, and introducing the category adaptive key region radius, the key parts of the head and face are given higher weights, which improves the semantic accuracy of occlusion judgment. By introducing the lowest effective boundary ordinate as a tie arbitration mechanism, a directed relationship is forcibly assigned when the occlusion judgment values ​​are equal, avoiding the problem of topological structure breakage. Without the need for additional hardware, a stable occlusion direction relationship can be derived based only on the topological coverage relationship of the two-dimensional mask, which significantly reduces the system cost and improves the robustness of the algorithm.

[0062] This invention introduces a joint modeling mechanism of hierarchical compressed occlusion relationship directed topology graph and graph neural network, and combines it with global occlusion weight change evaluation to optimize the guidance strategy from local decision-making to global optimality. It can complete global state inference before decision-making, effectively avoiding command conflicts and oscillation problems. At the same time, it generates accurate natural language voice commands and graphical prompts through spatial pointing label mapping, achieving highly consistent and interpretable multimodal guidance effect, significantly improving the stability and user experience of multi-person collaborative shooting. Attached Figure Description

[0063] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0064] Figure 1 This is a flowchart of the self-service photo booth group photo guidance method based on semantic segmentation proposed in this invention;

[0065] Figure 2 This diagram illustrates the structure of the multi-resolution cascaded YOLOv11-Seg instance segmentation model and the low-resolution inference and high-resolution supplementary inference processes in the self-service photo booth group photo guidance method based on semantic segmentation proposed in this invention. Detailed Implementation

[0066] Example 1: Reference Figures 1-2 A method for guiding multiple people to take self-service group photos at self-service photo booths based on semantic segmentation includes:

[0067] Step 1: Use the imaging equipment inside the self-service photo booth to continuously acquire the original group photo video stream and perform preprocessing operations to obtain a standardized input image sequence;

[0068] In this embodiment, the preprocessing operations include distortion correction, illumination equalization, and noise suppression.

[0069] The set of normalized input image sequences consists of several frames of normalized input images, with each frame corresponding to a time index.

[0070] Step 2: Input the standardized input image sequence into the improved YOLOv11-Seg instance segmentation model deployed on edge computing nodes, and output the initial instance mask set and the corresponding human key point set;

[0071] In this embodiment, step two includes:

[0072] The normalized input image of any frame is input into the low-resolution inference branch of the improved YOLOv11-Seg instance segmentation model deployed on the edge computing node, and a fixed-size low-resolution resampling process is performed on the normalized input image to obtain the low-resolution input image.

[0073] In Example 1, the low-resolution inference branch of the improved YOLOv11-Seg instance segmentation model performs spatial scaling on the standardized input image according to the preset low-resolution target size. By performing interpolation resampling on the pixel grid of the corresponding frame of the standardized input image, the width and height of the standardized input image are mapped to the preset low-resolution target width and preset low-resolution target height, respectively, to obtain the low-resolution input image.

[0074] The low-resolution input image is input into the low-resolution inference branch of the improved YOLOv11-Seg instance segmentation model to obtain low-resolution candidate boxes, low-resolution mask response sets, and low-resolution keypoint response sets.

[0075] In Example 1, the low-resolution inference branch is used to perform backbone feature extraction and multi-scale feature fusion on the low-resolution input image to obtain a low-resolution feature map. Based on the low-resolution feature map, the bounding box prediction head outputs the position offset corresponding to each candidate target and performs coordinate decoding on the position offset to obtain the low-resolution candidate box.

[0076] Based on the low-resolution feature map, the pixel-level foreground response results corresponding to each candidate target are output through the mask prediction head to obtain the low-resolution mask response set; based on the low-resolution feature map, the human key point position response results corresponding to each candidate target are output through the key point prediction head to obtain the low-resolution key point response set.

[0077] Based on the low-resolution mask response set, a threshold determination is performed on the mask response corresponding to each low-resolution candidate box. Pixels with mask response values ​​greater than or equal to the preset mask threshold are marked as foreground pixels, and the remaining pixels are marked as background pixels, thus generating the corresponding low-resolution candidate instance mask.

[0078] In Example 1, according to the position range of each low-resolution candidate box in the low-resolution candidate box set, the local mask response region corresponding to each low-resolution candidate box is extracted, and size alignment processing is performed on the local mask response region to make the spatial range of the local mask response region consistent with the corresponding low-resolution candidate box. Pixel-level threshold segmentation processing is performed on the size-aligned local mask response region, and pixels with mask response values ​​greater than or equal to a preset mask threshold are marked as foreground pixels, and pixels with mask response values ​​less than the preset mask threshold are marked as background pixels, thereby generating a low-resolution candidate instance mask that corresponds one-to-one with the low-resolution candidate box.

[0079] Based on the low-resolution keypoint response set, the keypoint response corresponding to each low-resolution candidate box is decoded to obtain the corresponding low-resolution human keypoint set.

[0080] In Example 1, according to the position range of each low-resolution candidate box in the low-resolution candidate box set, the local key point response region corresponding to each low-resolution candidate box is extracted. The response values ​​of multiple key point response channels representing different human key point categories in the local key point response region are searched to determine the maximum response position in each key point response channel. The maximum response position is used as the initial position of the corresponding human key point. Then, the initial position of each human key point is mapped according to the position range of the low-resolution candidate box to obtain the low-resolution human key point set corresponding to the low-resolution candidate box.

[0081] For each low-resolution candidate instance mask in the standardized input image, calculate the occlusion risk score, filter out the indices corresponding to the low-resolution candidate instance masks whose occlusion risk scores are greater than or equal to the preset occlusion risk judgment threshold, and construct an occlusion risk region index set.

[0082] In Example 1, for each low-resolution candidate instance mask, pairwise overlap analysis is performed with other low-resolution candidate instance masks. The cross-union ratio (CUNR) occlusion metric is obtained by calculating the ratio of the number of overlapping pixels between two low-resolution candidate instance masks to the number of pixels in their union. At the same time, the boundary coverage ratio between the boundary pixel set of the low-resolution candidate instance mask and other low-resolution candidate instance masks is calculated to obtain the boundary occlusion metric. Meanwhile, the aspect ratio features of the low-resolution candidate boxes and the pixel distribution density features of the low-resolution candidate instance masks are extracted. When the aspect ratio features or pixel distribution density features are determined to exceed the preset range of normal human body shape distribution of a single entity, a single entity abnormality penalty term representing an extremely crowded merging state is generated.

[0083] The intersection-union occlusion metric, boundary occlusion metric, and individual instance anomaly penalty are weighted and combined according to preset weights to obtain the occlusion risk score of the low-resolution candidate instance mask.

[0084] The preset occlusion risk judgment threshold is obtained through the statistical results of occlusion risk scores during the calibration phase. Specifically, the occlusion risk scores corresponding to all low-resolution candidate instance masks during the calibration phase are statistically analyzed. The mean and standard deviation of all occlusion risk scores are calculated. The mean and standard deviation of the occlusion risk scores are linearly combined according to the threshold adjustment coefficient to obtain the preset occlusion risk judgment threshold. Minimum and maximum allowable values ​​are then applied to the preset occlusion risk judgment threshold to obtain the final preset occlusion risk judgment threshold.

[0085] ;

[0086] in, Indicates the first The first frame of the normalized input image Occlusion risk score corresponding to each low-resolution candidate instance mask. Indicates the relationship with the first The index of another low-resolution candidate box is used for pairwise comparison of the low-resolution candidate instance mask. Indicates the first The first frame of the normalized input image A low-resolution candidate instance mask. Indicates the first The set of boundary pixels of a low-resolution candidate instance mask. This represents the weighting coefficients of the intersection-union ratio (IoU) of the low-resolution candidate instance mask. This represents the weighting coefficient for the coverage of the low-resolution candidate instance mask boundary. Indicates the first The mask of the first low-resolution candidate instance and the first The number of overlapping pixels in the low-resolution candidate instance mask. Indicates the first The mask of the first low-resolution candidate instance and the first The number of pixels in the union of the low-resolution candidate instance masks. Indicates the first The set of boundary pixels of the low-resolution candidate instance mask and the first The number of overlapping boundary pixels of a low-resolution candidate instance mask Indicates the first The total number of boundary pixels of a low-resolution candidate instance mask. This represents the individual anomaly penalty weighting coefficient. Indicates the first The first frame of the normalized input image The single-entity anomaly penalty term corresponding to the low-resolution candidate instance mask is implemented with the following logic: when the first... The aspect ratio of the first low-resolution candidate box exceeds the upper limit of the normal single-person aspect ratio threshold, or the first... When the pixel density distribution of a low-resolution candidate instance mask exhibits a multi-peak pattern, an extreme crowded merging state determination is triggered. Set to a preset high-risk constant; otherwise Set it to 0.

[0087] For each index in the occlusion risk region index set, a cropping operation with extended boundaries is performed based on the position of the corresponding low-resolution candidate box in the original frame normalized input image to obtain an occlusion risk region image patch. The occlusion risk region image patch is then input into the high-resolution supplementary inference branch of the improved YOLOv11-Seg instance segmentation model. After high-resolution resampling processing, instance segmentation inference is performed to obtain a high-resolution mask response map and a high-resolution key point response map.

[0088] In Example 1, for each index in the occlusion risk region index set, the corresponding cropping region is determined based on the position range of the corresponding low-resolution candidate box in the normalized input image of the original frame. When determining the cropping region, the low-resolution human key point set corresponding to the low-resolution candidate box is extracted simultaneously, and it is determined whether the coordinates of each key point in the low-resolution human key point set exceed the boundary of the cropping region. If there are far-end key point coordinates that exceed the boundary, the boundary of the cropping region is expanded outward to completely surround all far-end key point coordinates, generating the corrected cropping region.

[0089] Preset extended boundaries are superimposed on the horizontal and vertical boundaries of the corrected cropping region to obtain the extended cropping region. Based on the extended cropping region, a local image is extracted from the normalized input image of the original frame to obtain the occlusion risk area image block.

[0090] High-resolution resampling processing with the same input size as the high-resolution supplementary inference branch is performed on the image patch of the occlusion risk region to obtain the high-resolution supplementary inference input image. The high-resolution supplementary inference input image is then input into the high-resolution supplementary inference branch of the improved YOLOv11-Seg instance segmentation model. The high-resolution supplementary inference branch includes a high-resolution input adaptation unit, a backbone feature extraction unit, a multi-scale feature fusion unit, and a mask prediction unit and a key point prediction unit connected to the multi-scale feature fusion unit, respectively.

[0091] The high-resolution input adaptation unit performs size alignment and channel normalization on the high-resolution supplementary inference input image. The backbone feature extraction unit extracts local human texture features and boundary contour features from the occlusion risk area image patch. The multi-scale feature fusion unit fuses local human texture features and boundary contour features at different scales. The mask prediction unit outputs pixel-level foreground response results based on the fused high-resolution features to obtain a high-resolution mask response map. The key point prediction unit outputs human key point position response results based on the fused high-resolution features to obtain a high-resolution key point response map.

[0092] A high-resolution candidate instance mask is generated based on the high-resolution mask response map. A high-resolution human key point set is obtained by decoding the high-resolution key point response map. The high-resolution human key point set is then back-mapped to the image coordinate system of the original frame normalized input image through spatial coordinate mapping to obtain the back-mapped high-resolution instance mask and the back-mapped high-resolution human key point set.

[0093] In Example 1, based on the high-resolution mask response map, according to the response regions of each candidate target output by the high-resolution supplementary inference branch, the local high-resolution mask response region corresponding to each candidate target is extracted, and the size alignment processing is performed on the local high-resolution mask response region so that the local high-resolution mask response region is consistent with the high-resolution candidate box space range of the corresponding candidate target.

[0094] Pixel-level thresholding is performed on the local high-resolution mask response region after size alignment. Pixels with mask response values ​​greater than or equal to the preset high-resolution mask threshold are marked as foreground pixels, and pixels with mask response values ​​less than the preset high-resolution mask threshold are marked as background pixels, generating a high-resolution candidate instance mask that corresponds one-to-one with each candidate target.

[0095] Based on the high-resolution keypoint response map, the corresponding local high-resolution keypoint response region is extracted according to the high-resolution candidate box space range of each candidate target. Response value search is performed on multiple keypoint response channels representing different human keypoint categories in the local high-resolution keypoint response region to determine the maximum response position in each keypoint response channel. The maximum response position is then used as the keypoint position of the corresponding human keypoint within the high-resolution candidate box space range, resulting in a high-resolution human keypoint set that corresponds one-to-one with each candidate target.

[0096] Based on the cropping position of the occlusion risk region image patch relative to the original frame normalized input image, the scale transformation relationship corresponding to the high-resolution resampling process, and the local positional relationship of the high-resolution candidate box spatial range relative to the occlusion risk region image patch, inverse coordinate transformation is performed on the high-resolution candidate instance mask and the high-resolution human keypoint set, respectively. The foreground pixel positions in the high-resolution candidate instance mask are mapped back to the image coordinate system of the original frame normalized input image to obtain the back-mapped high-resolution instance mask. The positions of each human keypoint in the high-resolution human keypoint set are mapped back to the image coordinate system of the original frame normalized input image to obtain the back-mapped high-resolution human keypoint set.

[0097] For low-resolution candidate instance masks and corresponding human keypoint sets that are not included in the occlusion risk area index set, coordinate backmapping is performed. The back-mapped high-resolution instance mask and the back-mapped low-resolution instance mask are then fused together with the occlusion risk area index set. Cross-frame temporal association is then performed to obtain an initial instance mask set and human keypoint set carrying consistent identity labels.

[0098] In Example 1, for the low-resolution candidate instance mask and the corresponding human keypoint set that are not included in the occlusion risk region index set, the low-resolution candidate instance mask and the human keypoint set are back-mapped to the image coordinate system of the original frame normalized input image through spatial coordinate mapping to obtain the back-mapped low-resolution instance mask and the back-mapped low-resolution human keypoint set; based on the occlusion risk region index set, the back-mapped high-resolution instance mask and the back-mapped low-resolution instance mask are fused.

[0099] During the fusion process, if the boundary pixels of the back-mapping high-resolution instance mask and the back-mapping low-resolution instance mask have spatial position conflicts in the image coordinate system, a pixel-level arbitration mechanism is established: based on the high-frequency edge feature map output by the high-resolution supplementary inference branch, the pixels in the conflict area are assigned to the instance mask with higher edge gradient response value, and morphological smoothing filtering is performed on the fused boundary; the corresponding human key point set is fused synchronously.

[0100] Extract the historical initial instance mask set and historical human keypoint set with assigned temporal instance IDs from the previous frame's normalized input image. Based on the intersection-union ratio (IU) of each initial instance mask in the current frame's fused initial instance mask set and the historical initial instance mask set, and combined with the spatial Euclidean distance between each human keypoint in the current frame's human keypoint set and each historical human keypoint in the historical human keypoint set, construct a bipartite graph matching cost matrix. Use the Hungarian algorithm to solve for the optimal matching result of the bipartite graph matching cost matrix. Assign instance IDs with cross-frame temporal consistency to each initial instance mask and each human keypoint in the current frame's fused output, and output the initial instance mask set and human keypoint set carrying consistent identity labels.

[0101] Step 3: Based on the set of human body key points, refine the initial instance mask set by key point prior auxiliary masking to generate a refined instance mask set;

[0102] In this embodiment, step three includes:

[0103] For each initial instance mask set and the human body key point set, a key point prior response field is established for each initial instance mask and its corresponding human body key point set.

[0104] In Example 1, within the pixel region corresponding to the initial instance mask, a spatial distribution response is constructed in the pixel coordinate space of the standardized input image, with each human keypoint in the human keypoint set as the center position. The spatial distribution response is then weighted and superimposed to obtain the keypoint prior response field.

[0105] The spatial distribution response corresponding to each human body key point decays as the distance from the pixel coordinate to the human body key point increases.

[0106] Based on the preset connection relationships between key points in the human body key point set, a priori response field for skeleton connection is established.

[0107] In Example 1, the preset connection relationship represents the set of human skeleton lines. For each pair of human key points in the set of human key points that satisfy the preset connection relationship, a line segment connecting the two human key points is constructed according to the two-dimensional human key point coordinates in the image coordinate system, thus obtaining the set of human skeleton lines.

[0108] Within the pixel coordinate space of the standardized input image, for each pixel coordinate, the shortest distance from the pixel coordinate to each human skeleton connection in the set of human skeleton connections is calculated. Based on the shortest distance, a spatial distribution response is constructed for each pixel coordinate. For each pixel coordinate, the spatial distribution response with the largest response value is selected from the spatial distribution responses corresponding to all human skeleton connections as the skeleton connection prior response value of the corresponding pixel coordinate, thus obtaining the skeleton connection prior response field.

[0109] Normalization is performed on the prior response fields of key points and skeleton connections respectively to obtain normalized prior response fields of key points and normalized prior response fields of skeleton connections.

[0110] The initial instance mask, the normalized keypoint prior response field, and the normalized skeleton connection prior response field are fused at the pixel level to obtain a refined mask response map. Threshold segmentation is then performed based on the refined mask response map to generate a refined instance mask.

[0111] In Example 1, the values ​​of each pixel coordinate in the initial instance mask, the response values ​​in the normalized keypoint prior response field, and the response values ​​in the normalized skeleton connection prior response field are weighted and summed according to preset weight coefficients to obtain the mask refinement response map.

[0112] Threshold segmentation is performed based on the mask thinning response map. Pixels with mask thinning response values ​​greater than or equal to the mask thinning threshold are marked as foreground pixels, and pixels with mask thinning response values ​​less than the mask thinning threshold are marked as background pixels, thus generating a thinning instance mask.

[0113] Perform connected component partitioning on each refined instance mask to obtain the refined instance mask connected component. Determine whether the human key points in the human key point set fall inside the refined instance mask connected component to obtain the final refined instance mask.

[0114] In Example 1, for each refined instance mask connected region, it is determined whether the human key points in the human key point set fall inside the refined instance mask connected region:

[0115] When the refined instance mask connected region contains at least one human key point, the corresponding refined instance mask connected region is retained.

[0116] When the refined instance mask connected region does not contain any human key points, delete the corresponding refined instance mask connected region.

[0117] The connected regions of all retained refined instance masks are merged to obtain the final refined instance mask.

[0118] All final refined instance masks are constructed into a set of refined instance masks, and the boundary contours, occupied areas, and image coordinate system positions are extracted.

[0119] In Example 1, the boundary contour is extracted by filtering the pixel coordinates in the final refined instance mask that satisfy the condition that the current pixel is a foreground pixel and that there is at least one background pixel in its eight neighborhoods. The set of pixel coordinates constitutes the boundary contour.

[0120] The occupied area is calculated as follows: count all foreground pixels in the final refined instance mask to obtain the total number of foreground pixels, and use the total number of foreground pixels as the occupied area.

[0121] The image coordinate system position is calculated as follows: the horizontal position is obtained by weighted averaging of the horizontal coordinates of all foreground pixels in the final thinned instance mask, and the vertical position is obtained by weighted averaging of the vertical coordinates of all foreground pixels. The horizontal and vertical positions are then combined to form the image coordinate system position.

[0122] Step 4: Perform pixel-level overlap relationship calculation and key region occlusion rate calculation on any two human body instances in the refined instance mask set to obtain the occlusion relationship pair set and construct the occlusion relationship directed topology graph;

[0123] In this embodiment, step four includes:

[0124] For any two different human body instances in the refined instance mask set, perform pixel-level overlap relationship calculation to obtain the full mask occlusion rate;

[0125] In Example 1, for any two different human body instance masks, the overlapping pixel region of the two final refined instance masks in the pixel coordinate space is calculated. By counting the number of pixels in the overlapping pixel region and dividing the number of pixels by the area occupied by the final refined instance mask corresponding to the occluded human body instance, the full mask occlusion rate of the occluded human body instance is obtained. The full mask occlusion rate is used to represent the degree of occlusion of the entire area of ​​one human body instance by another human body instance.

[0126] Based on the set of human body key points, construct the key region mask corresponding to each human body instance, and calculate the key region occlusion rate.

[0127] In Example 1, the radius of the key region is constructed in the pixel coordinate space of the standardized input image, with each human key point in the human key point set as the center.

[0128] Different categories of human body key points correspond to different category adaptive key region radii. The category adaptive key region radius corresponding to human body key points on the head and face is larger than that corresponding to human body key points on the extremities.

[0129] For each pixel coordinate, determine whether the pixel coordinate falls within the radius of the category-adaptive key region corresponding to any human key point, and is also located inside the final refined instance mask corresponding to the human instance. When the above two conditions are met, mark the corresponding pixel as a key region pixel and obtain the key region mask.

[0130] Based on the overlapping pixel area between the final refined instance mask corresponding to the occluded human instance and the key region mask corresponding to the occluded human instance, the number of overlapping pixels in the key region is counted, and the number of overlapping pixels in the key region is divided by the total number of pixels in the key region mask to obtain the key region occlusion rate, which is used to represent the degree of occlusion of the key part area of ​​one human instance by another human instance.

[0131] Based on the full mask occlusion rate and the key area occlusion rate, the directional occlusion judgment value is calculated to obtain the occlusion relationship pair;

[0132] In Example 1, the full mask occlusion rate and the key area occlusion rate are weighted and combined according to preset weights to obtain the directional occlusion judgment value.

[0133] When the directional occlusion determination value of the occluding human instance on the occluded human instance is greater than or equal to the preset occlusion determination threshold, and the directional occlusion determination value of the occluding human instance on the occluded human instance is greater than the directional occlusion determination value of the occluded human instance on the occluding human instance, it is determined that the occluding human instance occludes the occluded human instance, and a corresponding directed occlusion relationship pair is generated.

[0134] When the directional occlusion judgment value of the occluded human instance to the occluded human instance is equal to the directional occlusion judgment value of the occluded human instance to the occluded human instance, and both are greater than or equal to the preset occlusion judgment threshold, the lowest effective boundary ordinate of the final refined instance mask corresponding to the two human instances in the image coordinate system is extracted.

[0135] Human instances with larger minimum effective boundary ordinates are identified as occluded human instances, and human instances with smaller minimum effective boundary ordinates are identified as occluded human instances, generating corresponding directed occlusion pairs.

[0136] Based on the full mask occlusion rate, the key area occlusion rate, and the proximity of the center position, calculate the corresponding edge weights of the occlusion relationship;

[0137] In Example 1, for any two human body instances, the image coordinate system positions corresponding to the two human body instances are extracted, the center position distance between the two human body instances is calculated, and the degree of closeness of the center positions is calculated based on the center position distance. The smaller the center position distance, the greater the degree of closeness of the center positions.

[0138] The occlusion rate of the full mask, the occlusion rate of the key area, and the proximity of the center position are weighted and combined to obtain the edge weights corresponding to the occlusion relationship pairs.

[0139] Using each human instance in the refined instance mask set as a node and the occlusion relationship pair set as directed edges, a directed topology graph of occlusion relationships is constructed. The edge weights in the directed topology graph of occlusion relationships are written to obtain the directed topology graph of occlusion relationships used for hierarchical consistency determination.

[0140] Step 5: Perform a layer consistency determination on the occlusion relationship directed topology graph within multiple consecutive frames. If the determination result meets the merging condition, merge nodes at the same layer into block nodes to obtain a layer-compressed occlusion relationship directed topology graph.

[0141] In this embodiment, step five includes:

[0142] A cross-frame node correspondence is established based on the occlusion relationship directed topology graph to form a node tracking sequence;

[0143] In Example 1, the occlusion relationship directed topology graphs corresponding to multiple consecutive standardized input images are arranged in chronological order to construct a temporal occlusion relationship directed topology graph sequence. For each standardized input image in the temporal occlusion relationship directed topology graph sequence, the corresponding human instance consistency identity label is extracted for each occlusion graph node. Based on the human instance consistency identity label, a cross-frame node correspondence relationship for the same human instance is established between multiple consecutive frames, so that the occlusion graph nodes of the same human instance in different frames form a node tracking sequence.

[0144] For each occlusion graph node in the directed topology graph of each frame, calculate the node level value;

[0145] In Example 1, for each occlusion graph node in any frame of the standardized input image, all directed occlusion relationship pairs pointing to the occlusion graph node are extracted, and the corresponding starting occlusion graph nodes are used to form a predecessor node set. When the predecessor node set is empty, the occlusion graph node is determined as a layer 0 node.

[0146] When the set of predecessor nodes is not empty, obtain the node level value corresponding to each predecessor node in the set of predecessor nodes, select the maximum value from all the node level values ​​corresponding to predecessor nodes, and add 1 to the maximum value to use as the node level value of the current occlusion graph node.

[0147] The node hierarchy values ​​corresponding to all occlusion map nodes in each frame of the standardized input image are obtained, which are used to represent the front and back hierarchy relationships of different human instances in the occlusion relationship.

[0148] Based on the node hierarchy values ​​within multiple consecutive frames, a hierarchy consistency determination is performed. For nodes at the same level that meet the merging conditions, block node merging processing is performed to obtain a block node set.

[0149] In Example 1, for each frame within a continuous frame window, the difference in node hierarchy values ​​between two nodes in that frame is calculated to obtain the hierarchy difference, which is used to represent the degree of hierarchy difference between the two nodes in the same frame.

[0150] For each frame within a consecutive frame window, determine whether there is a directed occlusion relationship between two nodes:

[0151] When there is a directional occlusion pair in any direction, the connection indicator value of the corresponding frame is marked as 1; otherwise, it is marked as 0, thus obtaining the connection indicator sequence.

[0152] When the level difference corresponding to each frame in a continuous frame window is 0, and the connection indicator value corresponding to each frame in a continuous frame window is 0, it is determined that the two nodes maintain the same level in multiple consecutive frames and do not have a direct occlusion relationship, and the two nodes meet the merging condition. When the condition is not met, it is determined that the two nodes do not meet the merging condition.

[0153] All occlusion map nodes that meet the merging conditions and have the same node level value are divided into same-layer node groups. Each same-layer node group corresponds to a set of human instances that can be merged. For each same-layer node group, a block node is generated. All block nodes in the normalized input image of frame t are constructed into a block node set. For occlusion map nodes that are not divided into any same-layer node group, their original node form is retained, and the original node is written into the block node set as a single-node block node.

[0154] Based on the block node set, the directed edges are reconstructed to obtain a directed topological graph of hierarchical compression occlusion relationships.

[0155] In Example 1, for any two different block nodes in the block node set, it is determined whether there is a directed occlusion relationship between any original node belonging to the previous block node and any original node belonging to the next block node. When there is a directed occlusion relationship, a compressed directed edge is established between the corresponding two block nodes.

[0156] For each compressed directed edge, extract the edge weights of all original directed occlusion pairs, select the maximum value from the edge weights as the compressed edge weight of the compressed directed edge, and construct a hierarchical compressed occlusion directed topology graph from all block nodes, compressed directed edges, and compressed edge weights.

[0157] Step 6: Input the directed topology graph of hierarchical compressed occlusion relationship into the graph neural network model to generate a set of candidate pose adjustment strategies;

[0158] In this embodiment, step six includes:

[0159] The hierarchical compressed occlusion relationship directed topology graph is input into the graph neural network model, and block node input features are constructed.

[0160] In Example 1, for each block node in the block node set, all compressed directed edges pointing to the block node are extracted, and the number of edges is counted to obtain the in-degree of the block node; at the same time, the weights of all compressed edges pointing to the block node are extracted and summed to obtain the cumulative value of the in-edge weights of the block node. The in-degree and cumulative value of the in-edge weights of the block node are normalized respectively, and the normalized centroid coordinates and relative size of the bounding box of the corresponding block node in the image coordinate system are extracted. The normalized in-degree, cumulative value of in-edge weights, and spatial geometric features are combined to form the input features of the block node.

[0161] Based on the hierarchical compression occlusion relationship directed topology graph and block node input features, the graph neural network model is used for inference to obtain the action reward score corresponding to each block node.

[0162] In Example 1, during the inference process of the graph neural network model, the set of block nodes and the set of compressed edges in the directed topological graph of hierarchical compression occlusion relationship are converted into an adjacency matrix.

[0163] The input feature vector of each block node is aggregated with the input feature vectors of its neighboring block nodes according to the connection relationship defined by the adjacency matrix, including the bidirectional feature interaction between the occluding and occluding parties, to obtain the hidden feature vector of the block node in the current network layer.

[0164] In the multi-layer structure of the graph neural network model, the hidden feature vectors of the block nodes in the last layer are fed into the output mapping layer. The action reward score of each block node is calculated through linear mapping and output activation function. The action reward score represents the expected reward and priority of the corresponding block node in the current frame when executing the pose adjustment strategy.

[0165] The block node attitude adjustment judgment value is calculated by weighted summation based on the block node in-degree, the cumulative value of the block node in-edge weight, and the action gain score.

[0166] Based on the attitude adjustment judgment values ​​of block nodes, a set of candidate attitude adjustment strategies is generated;

[0167] In Example 1, a correspondence rule is pre-established between the block node attitude adjustment judgment value and the candidate attitude adjustment strategy category. The candidate attitude adjustment strategy categories include probe candidate attitude adjustment strategy, lateral offset candidate attitude adjustment strategy and longitudinal clearance candidate attitude adjustment strategy.

[0168] The block node attitude adjustment judgment value is compared with the preset starting policy threshold and the threshold range of each preset policy in turn, and the candidate attitude adjustment policy category corresponding to the block node is determined according to the comparison result.

[0169] After completing the assignment of attitude adjustment strategy categories for block nodes, the block node, the candidate attitude adjustment strategy category corresponding to the block node, and the action gain score corresponding to the block node are combined to form a candidate attitude adjustment strategy for that block node. All generated candidate attitude adjustment strategies are summarized to obtain a set of candidate attitude adjustment strategies.

[0170] In this embodiment, the rule for determining the candidate attitude adjustment strategy category is as follows:

[0171] When the block node attitude adjustment judgment value is less than the preset starting policy threshold, it is determined that the corresponding block node does not need to perform attitude adjustment at present, and no candidate attitude adjustment policy is generated for the block node.

[0172] When the block node attitude adjustment judgment value is greater than or equal to the preset starting strategy threshold and falls within the first preset strategy threshold range, the candidate attitude adjustment strategy category corresponding to the block node is determined as the probe candidate attitude adjustment strategy.

[0173] When the block node attitude adjustment judgment value falls into the second preset strategy threshold range, the candidate attitude adjustment strategy category corresponding to the block node is determined as a lateral offset candidate attitude adjustment strategy with left / right spatial pointing vectors, based on the spatial relative relationship between the block node centroid coordinates and the center axis of the screen.

[0174] When the block node attitude adjustment judgment value falls within the third preset policy threshold range, the candidate attitude adjustment policy category corresponding to the block node is determined as the longitudinal yielding candidate attitude adjustment policy.

[0175] Step 7: Perform a global occlusion weight change evaluation on the candidate attitude adjustment strategy set to obtain the optimal attitude adjustment strategy and map it into a spatially directional multimodal guidance command.

[0176] In this embodiment, step seven includes:

[0177] For each candidate pose adjustment strategy, a global occlusion weight change evaluation is performed to obtain the global occlusion weight change corresponding to each candidate pose adjustment strategy;

[0178] In Example 1, based on the directed topological graph of the hierarchical compressed occlusion relationship corresponding to the standardized input image of frame t, the current global occlusion weight is calculated as the sum of all compressed edge weights.

[0179] The attitude adjustment attenuation coefficient is determined based on the candidate attitude adjustment strategy category, and a preset occlusion transfer penalty coefficient is set. (Empirical hyperparameters, pre-calibrated through physical crowding experiments in a real-world photo booth setting), are used to represent the degree of occlusion aggravation caused by the compression of surrounding non-target nodes when the target node moves, specifically for the first... Candidate pose adjustment strategies, with the target block node as the target. The target is all nodes that point to the target block. In the compressed directed edges, the source node corresponding to the directed edge with the largest compressed edge weight is extracted and defined as the main occlusion source node. Virtual edge weight updates, distinguishing between target and non-target nodes, are then performed on the compressed directed edges directly associated with the target block node to obtain the virtually adjusted compressed edge weights.

[0180] ;

[0181] in, As the primary source node for occlusion, Indicates the first In the frame-normalized input image, for the first The candidate pose adjustment strategy executes the virtual adjustment of the compressed edge weights. Indicates the first From the frame-normalized input image, from the first frame The block node points to the first The compressed directed edges of each block node have their corresponding compression edge weights, and g(c) represents the target block node index corresponding to the c-th candidate attitude adjustment strategy. Indicates the first The first frame of the normalized input image The attitude adjustment attenuation coefficient corresponding to each candidate attitude adjustment strategy. This represents the attitude adjustment attenuation coefficient corresponding to the candidate attitude adjustment strategies for the probe. This represents the attitude adjustment attenuation coefficient corresponding to the candidate attitude adjustment strategies for lateral offset. This represents the attitude adjustment attenuation coefficient corresponding to the longitudinal yielding candidate attitude adjustment strategy.

[0182] The calculation logic of the virtual adjusted compressed edge weight is to perform edge weight decay only on the associated edges between the target block node and the main occlusion source node, while applying occlusion transfer penalty to the associated edges between the target block node and other adjacent nodes, simulating the occlusion transfer phenomenon in physical space.

[0183] Based on all the virtually adjusted compressed edge weights, the corresponding virtually adjusted global occlusion weights are calculated. Based on the current global occlusion weights and the virtually adjusted global occlusion weights, the change in global occlusion weights for each candidate pose adjustment strategy is calculated. :

[0184] ;

[0185] in, : indicates the first The current global occlusion weights corresponding to the frame-normalized input image. This represents the set of compressed edges.

[0186] If the change in global occlusion weight is greater than 0, it means that the candidate pose adjustment strategy has achieved positive occlusion relief benefits from a global perspective; if the change in global occlusion weight is less than 0, it means that the cost of exacerbating surrounding occlusion caused by the strategy exceeds its benefit of alleviating the main occlusion, and it is an ineffective or deteriorating strategy.

[0187] Based on the change in global occlusion weight and the action gain score, the global gain corresponding to the candidate pose adjustment strategy is calculated, and the pose adjustment strategy with the largest global gain is selected as the optimal pose adjustment strategy.

[0188] In Example 1, the global occlusion weight change and action gain score of each candidate attitude adjustment strategy are weighted and combined according to a preset weight to obtain the corresponding global gain; the global gains of all candidate attitude adjustment strategies are compared to determine the candidate attitude adjustment strategy number with the largest global gain, and the candidate attitude adjustment strategy corresponding to the candidate attitude adjustment strategy number is determined as the optimal attitude adjustment strategy.

[0189] When multiple candidate pose adjustment strategies have their global benefits reaching their maximum values ​​simultaneously, the strategy with the largest change in global occlusion weight is selected first. When the change in global occlusion weight is still the same, the strategy with the largest action benefit score is selected first.

[0190] Based on the optimal attitude adjustment strategy, a spatially directional multimodal guidance command is generated and output to the corresponding user in real time.

[0191] In Example 1, the target block node corresponding to the optimal pose adjustment strategy and its candidate pose adjustment strategy category are extracted, the set of human instance numbers contained in the target block node is obtained, and the horizontal and vertical positions of the target block node in the image coordinate system are calculated.

[0192] Before generating guidance instructions, for candidate pose adjustment strategies with lateral displacement attributes, the lateral position of the target block node in the image coordinate system and the lateral position of the associated node that causes the maximum incoming edge weight are extracted. By calculating the difference between the lateral coordinates of the two, a safe escape direction vector (left or right) away from the maximum occlusion source is generated. Based on the lateral spatial pointing label, the vertical spatial pointing label, the candidate pose adjustment strategy category, and the safe escape direction vector, a natural language voice instruction is generated (e.g., the instruction: "Customer on the left in the front row, please shift to the left"). A synchronously displayed graphical prompt is generated based on the target block node and the candidate pose adjustment strategy category. The natural language voice instruction and the graphical prompt are combined to generate a multimodal guidance instruction.

[0193] Step 8: Within the preset monitoring period after the multimodal guidance command is output, continuously collect the updated standardized input image sequence, repeat steps 2 to 7, and update the hierarchical compression occlusion relationship directed topology graph and optimal pose adjustment strategy in real time. When the global occlusion weight calculated based on the hierarchical compression occlusion relationship directed topology graph is lower than the preset occlusion convergence threshold and remains constant within the set stable time window, trigger the automatic group photo acquisition operation to complete the capture of multi-person group photos.

[0194] Example 2: The implementer first organized the scene data of the self-service photo booth, collecting a total of 312 original video clips, all from the same type of monocular 2D imaging equipment, with a uniform resolution of 1920×1080 and a uniform frame rate of 30 frames per second. After screening, 271 valid video clips were retained, with a total of 28,640 frames, including 6,240 frames for 2-person scenes, 9,180 frames for 3-person scenes, 7,820 frames for 4-person scenes, and 5,400 frames for 5-person scenes. After completing the annotation at the instance level, a total of 104,376 human body instances were obtained, along with 1,774,392 human body keypoint annotations, 56,288 sets of occlusion direction relationship annotations, and 104,376 sets of key area visibility ratio annotations. Further, training samples, validation samples, and test samples were obtained: 19,840 frames for training samples, 4,160 frames for validation samples, and 4,640 frames for test samples. In the training samples, mildly occluded samples accounted for 29.8%, moderately occluded samples for 44.1%, and severely occluded samples for 26.1%. Severely occluded samples were defined as those where the key areas of the head and face were occluded by other human instances by more than 35%. During the training phase, the implementers recorded that traditional human detection bounding box methods achieved an effective target localization recall rate of 91.3% on the training set, but could not output instance-level human boundaries; traditional face detection methods had a false negative rate of 18.6% for faces of users in the back row that were occluded in the training samples; methods that only performed instance segmentation without occlusion relationship directed topological graph modeling had an occlusion direction determination accuracy of only 83.9% in severely occluded samples. The method of this invention, after model convergence within the same training cycle, achieved a human instance segmentation boundary F1 value of 0.891 on the validation samples, and a key area occlusion direction determination accuracy of 95.1%.

[0195] In a specific shooting task, four users entered a standard self-service photo booth. The standing area inside the booth was approximately 1.3 meters by 1.35 meters. After entering, the four users automatically formed two rows: the two in the front row were close to the camera, while the two in the back row, due to insufficient space, stood behind the two in the front row. Within the first second of acquiring the original group photo video stream, the system continuously read 30 frames of standardized input images. The implementer observed in the background log that the average brightness of the 12th frame of the standardized input image was 102.7, the standard deviation of brightness was 34.5, there was wide-angle stretching at the four corners of the image, and the average edge line offset was 4.8 pixels in the uncorrected state. After distortion correction, illumination equalization, and noise suppression, the average brightness of the 12th frame of the standardized input image became 118.2, the standard deviation of brightness became 28.4, and the average edge line offset decreased to 1.3 pixels, indicating that the input image met the requirements for subsequent inference. At this point, the four users were assigned consistent identity tags ID-01, ID-02, ID-03, and ID-04, respectively.

[0196] After inputting the normalized input image of frame 12 into the low-resolution inference branch of the improved YOLOv11-Seg instance segmentation model, the system outputs four low-resolution candidate boxes with confidence scores of 0.97, 0.95, 0.91, and 0.89, corresponding to four human instances. After decoding the low-resolution mask response set, the low-resolution candidate instance masks corresponding to ID-01 and ID-02 have complete boundaries, while the low-resolution candidate instance masks corresponding to ID-03 and ID-04 show obvious adhesion in the head-shoulder connection region. The system simultaneously decoded four sets of human keypoints: ID-01 (17 visible keypoints), ID-02 (17 visible keypoints), ID-03 (14 visible keypoints), and ID-04 (13 visible keypoints). The system calculated occlusion risk scores. Logs show that ID-01 had an occlusion risk score of 0.16, ID-02 0.21, ID-03 0.58, and ID-04 0.67. The system's current preset occlusion risk threshold is 0.40; therefore, ID-03 and ID-04 were selected for the occlusion risk region index set. The system then cropped local image patches from the original frame for these two high-risk regions and performed high-resolution supplementary inference. After supplementary inference, the upper boundary of the head in ID-03 changed from a jagged edge to a continuous edge, and its mask area was corrected from 81,246 pixels to 86,411 pixels; the right hairline of ID-04 was re-separated from the shoulder of the user in the front row, and its mask area was corrected from 76,803 pixels to 82,957 pixels. After the high-resolution human keypoint set was remapped back to the original image coordinate system, ID-03 added and recovered the left ear point, right ear point, and chin point, while ID-04 added and recovered the right eye point, nose tip point, and right shoulder point.

[0197] During the mask refinement stage, the system establishes a priori response fields for keypoints for each initial instance mask and its corresponding set of human keypoints. Backend records show that the priori response values ​​of keypoints in the head region of ID-03 are significantly higher than those in the waist region, with the radius of the high-response area around the human keypoints varying between 8 and 22 pixels. Simultaneously, the system establishes a priori response field for skeleton connections based on preset connectivity relationships, enhancing the responses of the corresponding regions along the shoulder-neck connection, the head midline connection, and the upper arm connection. After normalization of the two response fields, they are fused pixel-level with the initial instance mask. In the fusion result, an isolated connected component with an area of ​​436 pixels in the upper left shoulder region of ID-03 was removed during the connected region filtering stage because it contained no human keypoints. A missing boundary segment in the right outer side of ID-04, caused by occlusion, was filled in under the guidance of the skeleton connection priori response field, increasing the boundary continuity coefficient from 0.73 to 0.88. After refinement, the final refined instance masks for ID-01, ID-02, ID-03, and ID-04 occupy areas of 145,832 pixels, 138,614 pixels, 89,103 pixels, and 84,672 pixels, respectively. The system further extracts the boundary contour pixel set, recording 1,324 boundary contour pixels for ID-04; simultaneously, it calculates the image coordinate system positions: ID-01's horizontal position is 624.3, and its vertical position is 603.8; ID-02's horizontal position is 1128.7, and its vertical position is 618.4; ID-03's horizontal position is 601.5, and its vertical position is 481.2; and ID-04's horizontal position is 1137.6, and its vertical position is 468.9.

[0198] During the occlusion relationship determination phase, the system performs pixel-level overlap relationship calculation and key region occlusion rate calculation for any two different human body instances. In the valid results of frame 12, the full mask occlusion rate of ID-01 over ID-03 is 0.214, and the key region occlusion rate is 0.337; the full mask occlusion rate of ID-02 over ID-04 is 0.296, and the key region occlusion rate is 0.462; the full mask occlusion rate of ID-01 over ID-04 is only 0.051, and the key region occlusion rate is 0.083; the full mask occlusion rate of ID-02 over ID-03 is only 0.039, and the key region occlusion rate is 0.061. Because this invention uses category-adaptive key region radius for different human body key point categories, the category-adaptive key region radius of the head and face key points is set to 1.9 times that of the limb end key points. Therefore, the system assigns a higher weight to the case where the right side of the face of ID-04 is occluded. After calculating the directional occlusion determination values, the directional occlusion determination value from ID-01 to ID-03 is 0.288, and the directional occlusion determination value from ID-03 to ID-01 is 0.041; the directional occlusion determination value from ID-02 to ID-04 is 0.395, and the directional occlusion determination value from ID-04 to ID-02 is 0.053. The system's current preset occlusion determination threshold is 0.15, therefore, directional occlusion relationships ID-01→ID-03 and ID-02→ID-04 are generated. Simultaneously, ID-03 and ID-04 exhibit partial shoulder overlap in frame 13, and both have a directional occlusion determination value of 0.168, which is greater than the preset occlusion determination threshold. The system therefore triggers a tie-breaking mechanism, extracting the lowest valid boundary ordinates of the final refined instance masks of both instances in the image coordinate system. The lowest valid boundary ordinate for ID-03 is 871, and for ID-04 it is 858. The system classifies ID-03, with the larger lowest valid boundary ordinate, as an occluded human instance, and ID-04 as an occluded human instance, generating an additional directed occlusion pair ID-03→ID-04. This result would not occur in traditional methods because traditional human detection bounding box methods only obtain two overlapping detection boxes, failing to provide a stable orientation; traditional face detection methods directly lose half of ID-04's face in this frame, making it impossible to establish a correspondence.

[0199] During the multi-frame hierarchical consistency determination phase, the system reads the directed topological graph of occlusion relationships corresponding to the 7 frames of standardized input images from frame 12 to frame 18, establishes cross-frame node correspondences, and forms a node tracking sequence. The system calculates the node hierarchical value of each occlusion graph node in each frame. The results show that: ID-01 has a node hierarchical value of 0 in all 7 frames, ID-02 has a node hierarchical value of 0 in all 7 frames, ID-03 has a node hierarchical value of 1 in all 7 frames, and ID-04 has a node hierarchical value of 1 from frame 12 to frame 14. In frames 15 and 16, due to a local direct occlusion relationship with ID-03, the node hierarchical value briefly becomes 2, and returns to 1 in frames 17 and 18. The system further judges the layer difference and connection indicator value. ID-01 and ID-02 have a layer difference of 0 for 7 consecutive frames, and their connection indicator values ​​are also 0 for 7 consecutive frames. Therefore, they meet the merging condition and are assigned to the same layer node group and merged into block node B-01. Although ID-03 and ID-04 have the same layer in frames 12, 13, and 14, their connection indicator values ​​are not 0 due to the direct directed occlusion relationship in frames 13 and 15. Therefore, they do not meet the merging condition and are retained as block nodes B-02 and B-03 respectively. Finally, the system obtains the block node set {B-01, B-02, B-03} and reconstructs the compressed directed edges. After compression, the weight of the compressed edge from B-01 to B-02 is 0.337, the weight of the compressed edge from B-01 to B-03 is 0.462, and the weight of the compressed edge from B-02 to B-03 is 0.168. After hierarchical compression, the number of nodes in the graph is reduced from 4 to 3, and the number of edges is compressed from 3 stable edges plus 1 short-term edge to 3 effective compressed edges, which significantly reduces the processing complexity of the subsequent graph neural network model.

[0200] During the strategy generation phase, the system inputs the hierarchical compressed occlusion relation directed topology graph into the graph neural network model and constructs block node input features. For B-01, the number of compressed directed edges pointing to this block node is 0, the in-degree of the block node is 0, and the cumulative weight of the block node's incoming edges is 0; for B-02, the number of compressed directed edges pointing to this block node is 1, the in-degree of the block node is 1, and the cumulative weight of the block node's incoming edges is 0.337; for B-03, the number of compressed directed edges pointing to this block node is 2, the in-degree of the block node is 2, and the cumulative weight of the block node's incoming edges is 0.630. After normalization, the normalized in-degree of the block node of B-01 is 0, and the normalized cumulative weight of the block node's incoming edges is 0; the normalized in-degree of the block node of B-02 is 0.5, and the normalized cumulative weight of the block node's incoming edges is 0.535; the normalized in-degree of the block node of B-03 is 1.0, and the normalized cumulative weight of the block node's incoming edges is 1.0. The system inputs the features of these three block nodes along with the adjacency matrix into the graph neural network model. After the graph neural network model completes three layers of information propagation, it outputs action reward scores of 0.18, 0.72, and 0.61 for B-01, B-02, and B-03, respectively. The system then performs a weighted summation of the normalized block node in-degree, the normalized cumulative value of the block node's in-edge weights, and the action reward scores to obtain the block node attitude adjustment judgment values: 0.11 for B-01, 0.73 for B-02, and 0.68 for B-03. The currently used preset initial strategy threshold is 0.45, the first preset strategy threshold range is [0.45, 0.70), the second preset strategy threshold range is [0.70, 0.85), and the third preset strategy threshold range is [0.85, 1.00]. Therefore, no candidate attitude adjustment strategy is generated for B-01, B-02 is assigned as a lateral offset candidate attitude adjustment strategy, and B-03 is assigned as a probe candidate attitude adjustment strategy. The system combines the block node, the candidate attitude adjustment strategy category corresponding to the block node, and the action gain score corresponding to the block node to obtain two candidate attitude adjustment strategies, denoted as "B-02, lateral offset, 0.72" and "B-03, probe, 0.61" respectively.

[0201] During the optimal strategy selection phase, the system evaluates the global occlusion weight changes for each of the two candidate attitude adjustment strategies. The current global occlusion weight is the sum of all compressed edge weights, i.e., 0.337 plus 0.462 plus 0.168, which equals 0.967. For the candidate attitude adjustment strategy "B-02, lateral offset, 0.72", the system calls the attitude adjustment attenuation coefficient of 0.55 corresponding to the lateral offset candidate attitude adjustment strategy and performs virtual edge weight updates on all compressed directed edges directly associated with B-02. After the update, the edge weight of B-01→B-02 changes from 0.337 to 0.152, and the edge weight of B-02→B-03 changes from 0.168 to 0.076. Other edges remain unchanged, resulting in a virtual adjusted global occlusion weight of 0.690. The corresponding change in global occlusion weight is... For the candidate attitude adjustment strategy "B-03, probe, 0.61", the system calls the attitude adjustment attenuation coefficient of 0.35 corresponding to the probe candidate attitude adjustment strategy and performs virtual edge weight update on all compressed directed edges directly associated with B-03. After the update, the edge weight of B-01→B-03 changes from 0.462 to 0.300, and the edge weight of B-02→B-03 changes from 0.168 to 0.109, resulting in a virtual adjusted global occlusion weight of 0.746, with a corresponding change in global occlusion weight of 0.221. If only the change in global occlusion weight is considered, the lateral offset candidate attitude adjustment strategy of B-02 is better; however, the system also weights the change in global occlusion weight with the action gain score according to preset weights to obtain the global gain. The final calculation results show that the global benefit of "B-02, lateral offset, 0.72" is 0.547, and the global benefit of "B-03, probe, 0.61" is 0.586. Therefore, the system selects "B-03, probe, 0.61" as the optimal posture adjustment strategy, which is fundamentally different from traditional methods. Traditional human detection bounding box methods only provide a simple prompt such as "the person on the right, move a little to the right" in the same frame. Traditional instance segmentation without mapping prioritizes the user with the largest occluded area and prompts "back row, please move to the side," without evaluating the chain reaction of this instruction on other nodes. This invention, by introducing a global benefit comparison, ultimately did not choose the lateral offset strategy, which seems to improve occlusion more but has greater side effects, but instead chose the probe strategy, which has less disturbance to the overall formation.

[0202] During the instruction generation and output phase, the system extracts the target block node B-03 corresponding to the optimal posture adjustment strategy and its candidate posture adjustment strategy category "probe," and obtains the set of human instance numbers contained in the target block node. Here, target block node B-03 only contains ID-04. The system calculates the horizontal position of ID-04 in the image coordinate system as 1137.6, and the image width is 1920. Therefore, the horizontal position falls into the right area, and the horizontal spatial pointing label is determined to be "right side." The system also reads the block node level value corresponding to block node B-03 as 1, so the vertical spatial pointing label is determined to be "back row." Based on the horizontal spatial pointing label "right side," the vertical spatial pointing label "back row," and the candidate posture adjustment strategy category "probe," the system generates the natural language voice instruction "Please peek a little from the right side of the back row," and highlights the area corresponding to ID-04 on the screen, while superimposing a slightly upward-right adjusting arrow, generating a synchronously displayed graphical prompt. After the multimodal guidance instruction is output, ID-04 makes a slight upward-right head adjustment movement within 0.58 seconds. The system re-acquired standardized input images from frames 19 to 26. The recalculation results showed that the visible proportion of the key facial regions of ID-04 increased from 57.4% to 88.9%, the occlusion rate of the key regions corresponding to ID-04 decreased from 0.462 to 0.119, the compression weight from B-01 to B-03 decreased from 0.462 to 0.171, the compression weight from B-02 to B-03 decreased from 0.168 to 0.042, and the global occlusion weight decreased from 0.967 to 0.550, a reduction of 43.1%. Simultaneously, the visible proportion of the key facial regions of ID-03 remained above 84.6%, and was not obscured again due to the movement of the ID-04 probe. The system continuously monitored that the updated global occlusion weight remained below the preset occlusion convergence threshold of 0.60 for the next 41 frames, and no new hierarchical abrupt changes occurred for 1.5 seconds. Therefore, it automatically triggered shooting and finally output the final image. In this task, the total time from the entry of four users into the photo booth to the final automatic capture was 6.2 seconds. The first low-resolution inference took 18 milliseconds, the high-resolution supplementary inference for occlusion risk areas took 11 milliseconds, the refinement of the key point prior auxiliary mask took 6 milliseconds, the construction of the occlusion relationship directed topology graph took 4 milliseconds, the hierarchical consistency judgment and block node merging took 3 milliseconds, the graph neural network model inference took 4 milliseconds, the global occlusion weight change evaluation and optimal pose adjustment strategy selection took 3 milliseconds, and the instruction mapping and broadcasting took 2 milliseconds. The average processing latency of the entire algorithm per frame was stable at around 31 milliseconds.

[0203] Within a continuous verification period, the implementers selected 120 high-density group photo tasks involving four or five people to compare the method of this invention with three traditional methods. The comparison results showed that the first-pass success rate of the method of this invention in the 120 tasks was 92.5%, compared to 71.7% for the traditional human detection bounding box method, 66.7% for the traditional face detection method, and 83.3% for the method that only performs instance segmentation without occlusion relationship directed topological graph modeling. The average number of instructions for the method of this invention was 2.4, compared to 4.9 for the traditional human detection bounding box method, 5.3 for the traditional face detection method, and 3.8 for the method that only performs instance segmentation without occlusion relationship directed topological graph modeling. The instruction oscillation rate of the method of this invention was 5.0%, compared to 26.7% for the traditional human detection bounding box method, 31.7% for the traditional face detection method, and 18.3% for the method that only performs instance segmentation without occlusion relationship directed topological graph modeling. Regarding the improvement in the visibility ratio of key areas for users with heavy occlusion, the method of this invention achieves an average improvement of 29.6 percentage points, compared to 13.8 percentage points for traditional human body detection methods, 11.4 percentage points for traditional face detection methods, and 21.7 percentage points for methods that only perform instance segmentation without occlusion relationship directed topological graph modeling. In determining whether the final image achieves effective exposure for all users, the method of this invention achieves 94.2%, compared to 75.8% for traditional human body detection methods, 69.2% for traditional face detection methods, and 86.7% for methods that only perform instance segmentation without occlusion relationship directed topological graph modeling.

[0204] Within the same verification cycle, the implementers selected 20 failed samples for review. The main reason for the failure of traditional human detection bounding box methods is that the large overlap of the bounding boxes makes it impossible to accurately determine the visibility of the actual head, resulting in alternating instructions to "move towards the center" and "stand forward." The main reason for the failure of traditional face detection methods is that once the faces of users in the back row are more than half-obscured, the system directly loses the target and no longer generates effective guidance. The main reason for the failure of methods that only perform instance segmentation without occlusion relationship directed topology graph modeling is that although it knows who is blocked, it cannot make a global trade-off among multiple targets, often resulting in a local correction loop of "first move the left side, then move the right side back." In this invention, only one set of samples failed in these 20 review samples. The reason for the failure was that the simultaneous large movement of five people caused a short-term mismatch of the identity labels of human instances in three consecutive frames. As a result, the system was delayed by 0.8 seconds in entering the stable range, but still managed to complete the capture.

[0205] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for guiding group photos in self-service photo booths based on semantic segmentation, characterized in that, include: Step 1: Use the imaging equipment inside the self-service photo booth to continuously acquire the original group photo video stream and perform preprocessing operations to obtain a standardized input image sequence; Step 2: Input the standardized input image sequence into the improved YOLOv11-Seg instance segmentation model deployed on edge computing nodes, and output the initial instance mask set and the corresponding human key point set; Step 3: Based on the set of human body key points, refine the initial instance mask set by key point prior auxiliary masking to generate a refined instance mask set; Step 4: Perform pixel-level overlap relationship calculation and key region occlusion rate calculation on any two human body instances in the refined instance mask set to obtain the occlusion relationship pair set and construct the occlusion relationship directed topology graph; Step 5: Perform a layer consistency determination on the occlusion relationship directed topology graph within multiple consecutive frames. If the determination result meets the merging condition, merge nodes at the same layer into block nodes to obtain a layer-compressed occlusion relationship directed topology graph. Step 6: Input the directed topology graph of hierarchical compressed occlusion relationship into the graph neural network model to generate a set of candidate pose adjustment strategies; Step 7: Perform a global occlusion weight change evaluation on the candidate attitude adjustment strategy set to obtain the optimal attitude adjustment strategy and map it into a spatially directional multimodal guidance command. Step 8: Within the preset monitoring period after the multimodal guidance command is output, continuously collect the updated standardized input image sequence, repeat steps 2 to 7, and update the hierarchical compression occlusion relationship directed topology graph and optimal pose adjustment strategy in real time. When the global occlusion weight calculated based on the hierarchical compression occlusion relationship directed topology graph is lower than the preset occlusion convergence threshold and remains constant within the set stable time window, trigger the automatic group photo acquisition operation to complete the capture of multi-person group photos.

2. The self-service photo booth group photo guidance method based on semantic segmentation according to claim 1, characterized in that, The preprocessing operations include distortion correction, illumination equalization, and noise suppression. The set of normalized input image sequences consists of several frames of normalized input images, with each frame corresponding to a time index.

3. The self-service photo booth group photo guidance method based on semantic segmentation according to claim 1, characterized in that, Step two includes: The normalized input image of any frame is input into the low-resolution inference branch of the improved YOLOv11-Seg instance segmentation model deployed on the edge computing node, and a fixed-size low-resolution resampling process is performed on the normalized input image to obtain the low-resolution input image. The low-resolution input image is input into the low-resolution inference branch of the improved YOLOv11-Seg instance segmentation model to obtain low-resolution candidate boxes, low-resolution mask response sets, and low-resolution keypoint response sets. Based on the low-resolution mask response set, a threshold determination is performed on the mask response corresponding to each low-resolution candidate box. Pixels with mask response values ​​greater than or equal to the preset mask threshold are marked as foreground pixels, and the remaining pixels are marked as background pixels, thus generating the corresponding low-resolution candidate instance mask. Based on the low-resolution keypoint response set, the keypoint response corresponding to each low-resolution candidate box is decoded to obtain the corresponding low-resolution human keypoint set. For each low-resolution candidate instance mask in the standardized input image, calculate the occlusion risk score, filter out the indices corresponding to the low-resolution candidate instance masks whose occlusion risk scores are greater than or equal to the preset occlusion risk judgment threshold, and construct an occlusion risk region index set. For each index in the occlusion risk region index set, a cropping operation with extended boundaries is performed based on the position of the corresponding low-resolution candidate box in the original frame normalized input image to obtain an occlusion risk region image patch. The occlusion risk region image patch is then input into the high-resolution supplementary inference branch of the improved YOLOv11-Seg instance segmentation model. After high-resolution resampling processing, instance segmentation inference is performed to obtain a high-resolution mask response map and a high-resolution key point response map. A high-resolution candidate instance mask is generated based on the high-resolution mask response map. A high-resolution human key point set is obtained by decoding the high-resolution key point response map. The high-resolution human key point set is then back-mapped to the image coordinate system of the original frame normalized input image through spatial coordinate mapping to obtain the back-mapped high-resolution instance mask and the back-mapped high-resolution human key point set. For low-resolution candidate instance masks and corresponding human keypoint sets that are not included in the occlusion risk area index set, coordinate backmapping is performed. The back-mapped high-resolution instance mask and the back-mapped low-resolution instance mask are then fused together with the occlusion risk area index set. Cross-frame temporal association is then performed to obtain an initial instance mask set and human keypoint set carrying consistent identity labels.

4. The self-service photo booth group photo guidance method based on semantic segmentation according to claim 1, characterized in that, Step three includes: For each initial instance mask set and the human body key point set, a key point prior response field is established for each initial instance mask and its corresponding human body key point set. Based on the preset connection relationships between key points in the human body key point set, a priori response field for skeleton connection is established. Normalization is performed on the prior response fields of key points and skeleton connections respectively to obtain normalized prior response fields of key points and skeleton connections. The initial instance mask, the normalized keypoint prior response field, and the normalized skeleton connection prior response field are fused at the pixel level to obtain a refined mask response map. Threshold segmentation is then performed based on the refined mask response map to generate a refined instance mask. Perform connected component partitioning on each refined instance mask to obtain the refined instance mask connected component. Determine whether the human key points in the human key point set fall inside the refined instance mask connected component to obtain the final refined instance mask. All final refined instance masks are constructed into a set of refined instance masks, and the boundary contours, occupied areas, and image coordinate system positions are extracted.

5. The self-service photo booth group photo guidance method based on semantic segmentation according to claim 1, characterized in that, Step four includes: For any two different human body instances in the refined instance mask set, perform pixel-level overlap relationship calculation to obtain the full mask occlusion rate; Based on the set of human body key points, construct the key region mask corresponding to each human body instance, and calculate the key region occlusion rate. Based on the full mask occlusion rate and the key area occlusion rate, the directional occlusion judgment value is calculated to obtain the occlusion relationship pair; Based on the full mask occlusion rate, the key area occlusion rate, and the proximity of the center position, calculate the corresponding edge weights of the occlusion relationship; Using each human instance in the refined instance mask set as a node and the occlusion relationship pair set as directed edges, a directed topology graph of occlusion relationships is constructed. The edge weights in the directed topology graph of occlusion relationships are written to obtain the directed topology graph of occlusion relationships used for hierarchical consistency determination.

6. The self-service photo booth group photo guidance method based on semantic segmentation according to claim 1, characterized in that, Step five includes: A cross-frame node correspondence is established based on the occlusion relationship directed topology graph to form a node tracking sequence; For each occlusion graph node in the directed topology graph of each frame, calculate the node level value; Based on the node hierarchy values ​​within multiple consecutive frames, a hierarchy consistency determination is performed. For nodes at the same level that meet the merging conditions, block node merging processing is performed to obtain a block node set. Step 54: Reconstruct the directed edges based on the block node set to obtain a directed topological graph of hierarchical compression and occlusion relationships.

7. The self-service photo booth group photo guidance method based on semantic segmentation according to claim 1, characterized in that, Step six includes: The hierarchical compressed occlusion relationship directed topology graph is input into the graph neural network model, and block node input features are constructed. Based on the hierarchical compression occlusion relationship directed topology graph and block node input features, the graph neural network model is used for inference to obtain the action reward score corresponding to each block node. The block node attitude adjustment judgment value is calculated by weighted summation based on the block node in-degree, the cumulative value of the block node in-edge weight, and the action gain score. Based on the attitude adjustment judgment value of the block node, a set of candidate attitude adjustment strategies is generated.

8. The self-service photo booth group photo guidance method based on semantic segmentation according to claim 7, characterized in that, The criteria for determining the candidate pose adjustment strategy category are as follows: When the block node attitude adjustment judgment value is less than the preset starting policy threshold, it is determined that the corresponding block node does not need to perform attitude adjustment at present, and no candidate attitude adjustment policy is generated for the block node. When the block node attitude adjustment judgment value is greater than or equal to the preset starting strategy threshold and falls within the first preset strategy threshold range, the candidate attitude adjustment strategy category corresponding to the block node is determined as the probe candidate attitude adjustment strategy. When the block node attitude adjustment judgment value falls into the second preset strategy threshold range, the candidate attitude adjustment strategy category corresponding to the block node is determined as a lateral offset candidate attitude adjustment strategy with left / right spatial pointing vectors, based on the spatial relative relationship between the block node centroid coordinates and the center axis of the screen. When the block node attitude adjustment judgment value falls within the third preset policy threshold range, the candidate attitude adjustment policy category corresponding to the block node is determined as the longitudinal yielding candidate attitude adjustment policy.

9. The self-service photo booth group photo guidance method based on semantic segmentation according to claim 1, characterized in that, Step seven includes: For each candidate pose adjustment strategy, a global occlusion weight change evaluation is performed to obtain the global occlusion weight change corresponding to each candidate pose adjustment strategy; Based on the change in global occlusion weight and the action gain score, the global gain corresponding to the candidate pose adjustment strategy is calculated, and the pose adjustment strategy with the largest global gain is selected as the optimal pose adjustment strategy. Based on the optimal attitude adjustment strategy, a spatially directional multimodal guidance command is generated and output to the corresponding user in real time.