A whole body two-dimensional key point positioning method, device and medium
Patent Information
- Application Number
- CN202611319156.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-28
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本发明实施例提供了一种全身二维关键点定位方法、装置及介质,以至少解决现有全身二维关键点定位方法因缺乏自适应混淆识别与因果干预能力、难以兼顾多层级人体结构约束而导致的复杂场景下定位精度低与泛化能力差的技术问题
[0043]在本发明实施例中,通过构建基于混淆评估的自适应因果干预机制,能够显式量化每个全身二维关键点受遮挡、背景噪声和外观相似性干扰的程度,动态筛选高混淆关键点并执行自适应软替换校正,有效避免了传统硬替换策略造成的简单样本过干预、困难样本欠干预以及对罕见有效姿态过度正则化的问题,通过构建关键点层、肢体层和区域层的三层级分层图推理架构,同时建模了局部骨骼拓扑约束、肢体协同关系以及全身范围的高阶语义一致性,结合图推理与因果干预交替优化机制,能够在结构先验的指导下逐步消除局部预测偏差,抑制错误在骨架图中的传播,通过双路径联合优化机制与两阶段分层数据增强策略,引导模型学习因果相关而非统计相关的特征表示,在模拟真实复杂场景与保持特征分布稳定性之间取得了最优平衡,显著提升了模型在分布外样本上的泛化能力;第四,上述技术方案从特征学习、不确定性评估、结构约束到训练优化形成了完整的技术闭环,协同增效,有效解决了现有方法在严重遮挡、剧烈形变、复杂背景等挑战性场景下鲁棒性不足的问题,适用于智能安防、人机交互、动作分析、体育评估、自动驾驶及医学康复等多种应用场景。
Smart Images

Figure CN122820853A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of whole-body two-dimensional key point positioning technology, and more specifically, to a whole-body two-dimensional key point positioning method, device, and medium. Background Technology
[0002] With the rapid development of applications such as intelligent security, human-computer interaction, and autonomous driving, full-body 2D human pose estimation has become a research hotspot in the field of computer vision. Deep learning-based methods have achieved high-precision detection performance on standard datasets. However, existing full-body 2D keypoint localization methods still face severe challenges in complex real-world scenarios: First, existing methods essentially rely on observational correlations rather than causal relationships for learning. When the human body is severely occluded, undergoes drastic deformation, changes in viewpoint, or has complex background textures, the model is prone to incorrectly treating co-occurring features unrelated to the target keypoints as the basis for discrimination, resulting in anatomically unreasonable predictions and leading to localization errors, missed detections, or false detections. Second, existing methods lack dynamic recognition and adaptive intervention mechanisms for confused keypoints. They employ a uniform processing strategy for all keypoints, failing to explicitly assess the uncertainty of different keypoints and perform targeted feature correction. A few methods that introduce intervention mechanisms use a fixed number of interventions. Pre- or hard replacement strategies are prone to causing the problem of "over-intervention for simple samples and under-intervention for difficult samples." Third, there is a serious disconnect between prior knowledge of human anatomy and causal reasoning mechanisms. Most existing methods are limited to single-level or single-topological graph structure modeling, making it difficult to simultaneously consider local skeletal connections, limb coordination relationships, and high-order semantic consistency across the entire body. Furthermore, directly reasoning about potentially obfuscated initial features leads to errors propagating hierarchically throughout the skeleton graph. Fourth, conventional training strategies do not differentiate the enhancement intensity at different training stages, making it difficult to strike a balance between simulating real occlusion interference and maintaining feature distribution stability. The initialization method and joint optimization strategy for the canonical embedding table used for deobfuscation also lack targeted design. These shortcomings result in insufficient robustness and poor generalization ability of existing technologies in challenging scenarios such as occlusion, deformation, and complex backgrounds, which urgently need to be addressed.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This invention provides a method, apparatus, and medium for whole-body two-dimensional key point localization, which at least solves the technical problems of low localization accuracy and poor generalization ability in complex scenarios caused by the lack of adaptive confusion recognition and causal intervention capabilities and the difficulty in taking into account the constraints of multi-level human body structures in existing whole-body two-dimensional key point localization methods.
[0005] According to one aspect of the present invention, in order to achieve the above-mentioned objective, a method for locating two-dimensional key points throughout the body is provided, comprising the following steps:
[0006] The human image data to be processed is acquired, and feature extraction is performed on the human image data to obtain the whole body two-dimensional key point embedding features. Based on the whole body two-dimensional key point embedding features, the degree of confusion information is determined, whereby the degree of confusion information is used to characterize the degree of interference of each whole body two-dimensional key point.
[0007] Based on the degree of obfuscation information, an adjustment mode is determined. The adjustment mode is used to de-obfuscate and correct the embedded features of the whole body two-dimensional key points. The adjustment mode includes an adjustment mode and a fixed mode.
[0008] In response to the adjustment mode, a control instruction set is generated. This control instruction set is used to perform deobfuscation correction operations and human body structure modeling operations to output the coordinates of the two-dimensional key points of the whole body.
[0009] Furthermore, based on the embedding features of the whole-body two-dimensional keypoints, the level of confusion information is determined, including:
[0010] Based on the embedded features of the two-dimensional key points of the whole body, the discrete probability distribution of each key point in the horizontal and vertical directions is generated;
[0011] The confusion score for each two-dimensional key point in the whole body is calculated based on the discrete probability distribution.
[0012] The average level of confusion is calculated based on the confusion scores of all 2D key points throughout the body.
[0013] Furthermore, the confusion score for each two-dimensional keypoint in the whole body is calculated based on the discrete probability distribution, including:
[0014] The confusion score of each two-dimensional key point in the whole body is calculated by combining the maximum response value of the discrete probability distribution with the information entropy.
[0015] Based on the comparison between the average level of confusion and a preset first threshold, the number of key points to be intervened is determined. When the average level of confusion is lower than the first threshold, the number of key points to be intervened is zero.
[0016] Furthermore, based on the comparison between the average level of confusion and the first threshold, the number of key points to be intervened is determined, which also includes:
[0017] When the average level of confusion exceeds the preset second threshold, the number of key points to be intervened is the preset upper limit.
[0018] When the average level of confusion is between the first threshold and the second threshold, the number of key points to be intervened is determined by linear interpolation.
[0019] Key points to be intervened were selected in descending order of confusion scores.
[0020] Furthermore, feature extraction is performed on the human image data to obtain the whole-body two-dimensional keypoint embedding features, including:
[0021] Human image data is input into a multi-stage convolutional backbone structure for hierarchical feature extraction. Each convolutional stage sequentially performs convolution operations, normalization processing, and non-linear activation processing.
[0022] In the multi-stage convolutional backbone structure, a cross-stage partial connection mechanism is introduced to divide the input features into a side branch and a transformation branch. The side branch maintains the identity mapping, and the transformation branch is convolved and merged with the side branch after the transformation, and channel compression is performed through convolution.
[0023] The high-level semantic features output from the multi-stage convolutional backbone network are input into the gated attention enhancement module for channel response recalibration. The spatial features are then mapped to the two-dimensional keypoint dimension of the whole body through a learnable projection matrix to obtain the two-dimensional keypoint embedding features of the whole body.
[0024] Furthermore, the high-level semantic features are input into the gated attention enhancement module for channel response recalibration, including:
[0025] The global description vector for each channel is obtained through global average pooling;
[0026] Channel attention weights are generated by nonlinearly mapping the global description vector using a multilayer perceptron.
[0027] The channel attention weights are applied element-wise to the high-level semantic features to obtain the enhanced features.
[0028] Furthermore, based on the level of confusion information, the adjustment mode is determined, including:
[0029] When the confusion level information indicates that the average confusion level is greater than or equal to the first threshold, the adjustment mode is determined to be the regulation mode;
[0030] When the confusion level information indicates that the average confusion level is less than the first threshold, the adjustment mode is determined to be the fixed mode.
[0031] Furthermore, the de-obfuscation correction operation includes:
[0032] Construct a binary mask matrix, which is used to mark whether each key point has been intervened;
[0033] Obtain the preset specification embedding table, which contains specification embeddings that correspond one-to-one with each key point;
[0034] For each key point marked as requiring intervention, the fusion weight is determined based on its confusion score. The original embedded features are then weighted and fused with the corresponding canonical embedded features according to the fusion weight to obtain the key point features after intervention.
[0035] Human body structure modeling operations include:
[0036] The two-dimensional key point features of the whole body after intervention are input into the hierarchical graph inference network for human structure modeling. The hierarchical graph inference network includes at least three semantic levels: key point layer, limb layer and region layer. Through bottom-up inter-layer aggregation, graph structure update within each layer and top-down semantic feedback, the two-dimensional key point features of the whole body after structural correction are obtained.
[0037] The structurally corrected full-body 2D keypoint features are input into the full-body 2D keypoint prediction head, which outputs discrete probability distributions in the horizontal and vertical directions, and decodes the coordinates of the full-body 2D keypoints based on the discrete probability distributions.
[0038] According to one embodiment of the present invention, a whole-body two-dimensional key point positioning device is also provided, comprising:
[0039] The acquisition module is used to acquire human image data to be processed, extract features from the human image data to obtain the whole body two-dimensional key point embedding features, and determine the degree of confusion information based on the whole body two-dimensional key point embedding features. The degree of confusion information is used to characterize the degree of interference of each whole body two-dimensional key point.
[0040] The adjustment module is used to determine the adjustment mode based on the degree of obfuscation information. The adjustment mode is used to de-obfuscate and correct the embedded features of the whole body two-dimensional key points. The adjustment mode includes the adjustment mode and the fixed mode.
[0041] The generation module is used to generate a set of control instructions in response to the adjustment mode. The set of control instructions is used to perform deobfuscation correction operations and human body structure modeling operations to output the coordinates of the two-dimensional key points of the whole body.
[0042] According to another aspect of the present invention, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is executed, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of the present invention.
[0043] In this embodiment of the invention, by constructing an adaptive causal intervention mechanism based on confusion assessment, the degree of interference from occlusion, background noise, and appearance similarity of each full-body 2D keypoint can be explicitly quantified. Highly confused keypoints are dynamically screened and adaptive soft replacement correction is performed, effectively avoiding the problems of over-intervention on simple samples, under-intervention on difficult samples, and excessive regularization of rare effective poses caused by traditional hard replacement strategies. By constructing a three-level hierarchical graph inference architecture of keypoint layer, limb layer, and region layer, and simultaneously modeling local skeletal topological constraints, limb coordination relationships, and high-order semantic consistency across the entire body, combined with an alternating optimization mechanism of graph inference and causal intervention, local prediction bias can be gradually eliminated under the guidance of structural priors. To suppress the propagation of errors in the skeleton graph, a dual-path joint optimization mechanism and a two-stage hierarchical data augmentation strategy are used to guide the model to learn causal rather than statistically related feature representations. This achieves an optimal balance between simulating real complex scenarios and maintaining the stability of feature distribution, significantly improving the model's generalization ability on out-of-distribution samples. Fourth, the above technical solutions form a complete technical closed loop from feature learning, uncertainty assessment, structural constraints to training optimization, synergistically enhancing each other and effectively solving the problem of insufficient robustness of existing methods in challenging scenarios such as severe occlusion, drastic deformation, and complex backgrounds. It is applicable to various application scenarios such as intelligent security, human-computer interaction, motion analysis, sports assessment, autonomous driving, and medical rehabilitation. Attached Figure Description
[0044] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0045] Figure 1 This is a flowchart of a whole-body two-dimensional key point localization method according to one embodiment of the present invention;
[0046] Figure 2 This is a flowchart of a whole-body two-dimensional key point localization method according to one embodiment of the present invention;
[0047] Figure 3 This is a schematic diagram of the feature encoding network and gated attention enhancement module in a whole-body two-dimensional key point localization method according to one embodiment of the present invention;
[0048] Figure 4 This is a schematic diagram of the causal intervention module in a whole-body two-dimensional key point localization method according to one embodiment of the present invention;
[0049] Figure 5 This is a schematic diagram of the structure of a layered graph inference network in a whole-body two-dimensional key point localization method according to one embodiment of the present invention;
[0050] Figure 6 This is a schematic diagram of a dual-path joint optimization and two-stage training strategy in a whole-body two-dimensional keypoint localization method according to one embodiment of the present invention.
[0051] Figure 7 This is a structural block diagram of a two-dimensional whole-body key point positioning device according to one embodiment of the present invention. Detailed Implementation
[0052] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0053] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0054] According to an embodiment of the present invention, a method for locating two-dimensional key points of the whole body is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0055] This method embodiment can be executed in an electronic device or similar computing device that includes memory and a processor. Taking operation on a terminal as an example, the terminal may include one or more processors (processors may include, but are not limited to, central processing units (CPUs), graphics processing units (GPUs), digital signal processing (DSP) chips, microcontroller units (MCUs), field-programmable gate arrays (FPGAs), neural network processors (NPUs), tensor processors (TPUs), artificial intelligence (AI) type processors, etc.) and memory for storing data. Optionally, the terminal may also include transmission devices, input / output devices, and display devices for communication functions. Those skilled in the art will understand that the above structural description is merely illustrative and does not limit the structure of the terminal. For example, the terminal may include more or fewer components than described above, or have a different configuration than described above.
[0056] The memory can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the full-body two-dimensional keypoint localization method in this embodiment of the invention. The processor executes various functional applications and data processing by running the computer program stored in the memory, thereby realizing the aforementioned full-body two-dimensional keypoint localization method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the mobile terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0057] The transmission device is used to receive or send data via a network. Specific examples of the network mentioned above may include a wireless network provided by the mobile terminal's communication provider. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0058] Display devices can be, for example, touchscreen liquid crystal displays (LCDs) and touch displays (also referred to as "touchscreens" or "touch displays"). The LCD allows users to interact with the user interface of the mobile terminal. In some embodiments, the mobile terminal has a graphical user interface (GUI), which allows users to interact with the GUI through finger contact and / or gestures on a touch-sensitive surface. Optional human-computer interaction functions include: creating web pages, drawing, word processing, creating electronic documents, playing games, video conferencing, instant messaging, sending and receiving emails, call interfaces, playing digital video, playing digital music, and / or web browsing, etc. Executable instructions for performing the above human-computer interaction functions are configured / stored in one or more processor-executable computer program products or readable storage media.
[0059] Figure 1 This is a flowchart of a whole-body two-dimensional key point localization method according to one embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:
[0060] Step S110: Obtain the human image data to be processed, and extract features from the human image data to obtain the whole-body two-dimensional keypoint embedding features. Determine the degree of confusion based on the whole-body two-dimensional keypoint embedding features. The degree of confusion is used to characterize the degree of interference of each whole-body two-dimensional keypoint. The specific content is as follows:
[0061] In step S110, the human image data can be derived from any one or a combination of monocular RGB images, single human images, or single frames from videos. After acquiring the image data, the target region is cropped according to the human detection bounding box, extracting the region containing the human body from the original image to remove unnecessary background interference. The cropped human image is then uniformly scaled to a preset resolution, preferably 384×288 in this embodiment, to ensure consistency in the size of the input image. After size normalization, an affine transformation is performed on the image data to correct for posture and viewpoint deviations of the human body in the image. Based on this, pixel value standardization is performed on the image, using preset mean and variance to standardize the RGB three channels respectively, adjusting the pixel values of each channel to the standard range to reduce input distribution offsets caused by different acquisition devices, lighting conditions, and image quality. Through the above preprocessing steps, the differences in scale, posture, and color distribution of the input data can be effectively eliminated, providing standardized input for subsequent feature extraction.
[0062] After image data preprocessing, the preprocessed human image is input into a feature encoding network for hierarchical feature extraction. The feature encoding network comprises a multi-stage convolutional backbone, with each stage sequentially including convolution operations, normalization, and nonlinear activation. Convolution operations are used to extract local spatial context features from the image; normalization accelerates network convergence and improves training stability; and nonlinear activation functions introduce nonlinear transformation capabilities, enabling the network to fit complex feature distributions. In the multi-stage convolutional backbone, a cross-stage partial connection mechanism is introduced to optimize the feature propagation path. Specifically, the input features are divided into a bypass branch and a transformation branch. The bypass branch maintains an identity mapping, directly propagating the original features to subsequent layers to preserve fine-grained spatial location information. The transformation branch extracts higher-level semantic features through convolutional transformation. The features from the bypass branch and the transformation branch are then concatenated and fused, and channel compression is performed through convolution, effectively reducing redundant computation while maintaining feature expressive power.
[0063] The aforementioned feature encoding network ultimately outputs high-level semantic features. To enhance the discriminative power of the feature representation, the high-level semantic features are input into a gated attention enhancement module for channel response recalibration. Specifically, firstly, global average pooling is used to compress the spatial dimension features of each channel into a global description vector, allowing for the quantification of the global response of each channel. Subsequently, a multilayer perceptron is used to perform a nonlinear mapping on the global description vector to generate channel attention weights, which are used to characterize the contribution of different channels to the keypoint localization task. Finally, the generated channel attention weights are applied element-wise to the original high-level semantic features, strengthening channels with strong responses to key parts of the human body and suppressing irrelevant or interfering channel responses, thereby obtaining enhanced features. This gated attention enhancement mechanism effectively improves the model's sensitivity to keypoint-related features.
[0064] Based on this, the enhanced spatial features are mapped to the full-body 2D keypoint dimension using a learnable projection matrix. This involves converting the feature vector of each spatial location into an embedded representation corresponding to each keypoint, thus obtaining the full-body 2D keypoint embedding features. The dimension of the full-body 2D keypoint embedding features is determined by the batch size, the number of full-body 2D keypoints to be detected, and the number of feature channels. In this embodiment, the preferred number of full-body 2D keypoints to be detected is 133, covering the body, face, hands, and feet, and the preferred number of feature channels is 512, thus providing sufficient semantic expression for each keypoint. Through the above process, the original human image data is transformed into a structured keypoint-level embedding representation. The embedding features of each keypoint contain its corresponding local appearance information and global contextual semantics, providing a reliable feature foundation for subsequent confusion assessment and de-confussing correction.
[0065] After obtaining the embedded features of the full-body 2D keypoints, the confusion level information is further determined based on these features. This confusion level information characterizes the degree to which each full-body 2D keypoint is affected by occlusion, background noise, and appearance similarity interference. Specifically, a prediction modeling submodule is first constructed. This submodule generates the logistic values of each full-body 2D keypoint in the horizontal and vertical directions using linear mapping. A flexible maximum value function with a temperature parameter is then used to convert the logistic values into a discrete probability distribution. The temperature parameter adjusts the smoothness of the distribution; performing a maximum value shift operation before the flexible maximum value function significantly improves numerical stability and suppresses gradient explosion. Based on this, a confusion evaluation submodule is constructed, calculating the confusion score of each keypoint using its probability distribution. For each full-body 2D keypoint, a joint measurement is performed using the maximum response value and information entropy of its probability distribution. The maximum response value reflects the model's confidence in predicting the keypoint's location, while the information entropy characterizes the uncertainty of the prediction distribution. Lower maximum response values or higher information entropy indicate greater uncertainty in the keypoint's prediction and a higher likelihood of being affected by external interference factors. Finally, the average level of confusion is calculated based on the confusion scores of all full-body 2D keypoints. This average level of confusion serves as a global metric at the sample level, quantifying the overall predictive uncertainty of the current sample and providing a basis for subsequent adaptive intervention decisions. Through this method, accurate quantitative identification of the degree of interference at each full-body 2D keypoint is achieved.
[0066] Step S120: Based on the obfuscation level information, determine the adjustment mode. The adjustment mode is used to de-obfuscate and correct the embedded features of the whole-body two-dimensional keypoints. The adjustment mode includes an adjustment mode and a fixed mode, the details of which are as follows:
[0067] In step S120, based on the confusion level information determined in the preceding steps, an adjustment mode for the current sample is further determined. The confusion level information is quantified using the global metric of average confusion level, which comprehensively reflects the overall prediction uncertainty level of all full-body 2D keypoints in the current sample. The adjustment mode is used to de-confuse the embedded features of full-body 2D keypoints. Its purpose is to fuse the original embedded features of highly confused keypoints with their corresponding canonical embedded features to suppress the spurious correlation effects introduced by factors such as occlusion, background noise, and appearance similarity. The adjustment mode specifically includes two types: adjustment mode and fixed mode.
[0068] The core basis for determining the adjustment mode is the comparison between the average confusion level and a preset threshold. In this embodiment, a first threshold is preset, which is used to distinguish whether a sample is in a low-confusion or high-confusion state. Specifically, the average confusion level of all full-body 2D keypoints of the current sample is calculated, and the average confusion level is compared with the first threshold. When the average confusion level is lower than the first threshold, it indicates that the overall prediction uncertainty of the current sample is low, and the prediction results of each keypoint are relatively reliable. At this time, the adjustment mode is determined to be the fixed mode. In the fixed mode, no deconfusion correction operation is performed, that is, no intervention is made on the embedding features of any keypoint, thereby avoiding unnecessary interference to the already accurately predicted keypoints and preventing the problem of "over-intervention of simple samples". When the average confusion level is greater than or equal to the first threshold, it indicates that the current sample has high prediction uncertainty, and some keypoints may be significantly interfered with by occlusion, background noise, or similarity of parts. At this time, the adjustment mode is determined to be the adjustment mode. In the adjustment mode, deconfusion correction is required for the embedding features of high-confusion keypoints.
[0069] When the adjustment mode is determined to be the adjustment mode, the number of key points to be intervened is further determined based on the specific value of the average confusion level. In this embodiment, a second threshold is also set, which is higher than the first threshold. When the average confusion level is between the first threshold and the second threshold, the number of key points to be intervened is determined by linear interpolation, so that the number of interventions increases smoothly as the average confusion level increases. When the average confusion level is higher than the second threshold, it indicates that the sample as a whole is in a state of severe confusion. At this time, the number of key points to be intervened is set to a preset upper limit to ensure that a sufficient number of highly confused key points are corrected. Through the above adaptive mechanism, the number of interventions can be dynamically adjusted according to the actual confusion level of the sample, which avoids both the problem of "over-intervention for simple samples" caused by a fixed number of interventions and the problem of "under-intervention for difficult samples".
[0070] After determining the number of keypoints to be intervened in the adjustment mode, it is also necessary to determine which keypoints to intervene on. In this embodiment, keypoints are sorted based on their respective confusion scores, and a corresponding number of keypoints are selected as keypoints to be intervened on in descending order of confusion scores. Keypoints with higher confusion scores indicate more severe interference, and prioritizing intervention on them can more effectively improve the overall localization accuracy. For the keypoints selected for intervention, their original embedded features and corresponding canonical embeddings are weighted and fused according to the fusion weights. The fusion weights are determined by the confusion scores of the keypoints after Sigmoid mapping, so that the higher the confusion level of the keypoint, the greater the proportion of its original features being replaced, thus subjecting it to stronger canonical prior constraints. The canonical embedding is a learnable reference embedding set according to the keypoint index, used to provide a stable structural prior under high confusion conditions. Through the above soft replacement mechanism, de-confusion correction of highly confused keypoints is achieved.
[0071] By determining the adjustment mode and executing corresponding operations based on the aforementioned confusion level information, this embodiment can adaptively select a fixed mode or an adjustment mode for different sample confusion levels, and dynamically determine the number of interventions and the intervention targets in the adjustment mode, thus achieving sample-level adaptive causal intervention. This mechanism effectively avoids the "one-size-fits-all" approach in traditional methods, significantly improving the localization accuracy of highly confused samples while protecting the prediction accuracy of low-confusion samples.
[0072] Step S140: In response to the adjustment mode, a control instruction set is generated. This control instruction set is used to perform deobfuscation correction and human body structure modeling operations to output the coordinates of the full-body two-dimensional key points. The specific content is as follows:
[0073] In step S140, when performing the deobfuscation correction operation in adjustment mode, a binary mask matrix is first constructed based on the key points to be intervened determined in the preceding steps. The number of rows in this matrix corresponds to the batch size, and the number of columns corresponds to the total number of two-dimensional key points throughout the body. The values of the elements in the matrix are used to mark whether the corresponding key points are subject to intervention. Specifically, for key points selected for intervention after being sorted from high to low according to their obfuscation scores, their corresponding positions in the mask matrix are marked as 1, indicating that the key point needs feature intervention; for unselected key points, the corresponding positions are marked as 0, indicating that no intervention is performed and their original embedded features are preserved. Through this binary mask matrix, the scope of the intervention operation can be precisely controlled, avoiding indiscriminate processing of all key points.
[0074] After identifying the key points to be intervened, a pre-defined canonical embedding table is obtained. The canonical embedding table is a globally learnable parameter, its dimension determined by the number of 2D key points and feature channels throughout the body. In this embodiment, the canonical embedding table is initialized using a normal distribution and participates in end-to-end joint optimization along with the other network parameters during training. The canonical embedding table contains canonical embeddings corresponding to each key point, each representing the idealized baseline features of the corresponding key point under unobstructed, standard viewpoint conditions. For each key point marked for intervention, a fusion weight is determined based on its confusion score. Specifically, the fusion weight is obtained by mapping the confusion score to a sigmoid function, ensuring that key points with higher confusion levels have a greater proportion of their original features replaced, thus receiving stronger canonical prior constraints. Subsequently, the original embedding features and the corresponding canonical embeddings are weighted and fused according to the fusion weights to obtain the intervened key point features. Through this soft replacement mechanism, de-confusion correction of highly confused key points is achieved, avoiding information loss caused by hard replacement and ensuring that highly confused key points receive sufficient canonical prior constraints.
[0075] After de-obfuscation correction, the control instruction set continues to perform human structure modeling. The intervened full-body 2D keypoint features are input into a hierarchical graph inference network for human structure modeling. The hierarchical graph inference network includes at least three semantic layers: a keypoint layer, a limb layer, and a region layer. The keypoint layer preserves fine-grained spatial location information, the limb layer models the synergistic relationships of skeletal chains and functional joint combinations, and the region layer encodes high-order semantic information of the head, trunk, upper limbs, and lower limbs. Through this three-layer division, the model can simultaneously consider local skeletal connectivity, cross-limb synergistic relationships, and high-order semantic consistency across the entire body.
[0076] The hierarchical graph inference network first performs a bottom-up inter-layer aggregation operation. This process groups and aggregates low-level node features according to prior knowledge of human anatomy to generate high-level node representations: node features from the keypoint layer are aggregated according to anatomy to form node representations for the limb layer, and node features from the limb layer are further aggregated according to region to form node representations for the region layer. Through this aggregation process, the model can progressively integrate local observation information from multiple keypoints to form a more stable high-level representation of limb posture and overall body layout.
[0077] After inter-layer aggregation, graph structure updates within each layer are performed. Each layer's graph structure update propagates and corrects node features based on a normalized adjacency matrix or attention weight matrix. The adjacency matrix includes not only static topological priors based on human skeletal connections but also dynamic correction terms constructed based on the similarity of current sample node features, thereby enhancing the model's adaptability to unconventional poses and complex interaction scenarios. Residual connections are used during graph updates to enhance training stability. Through graph structure updates within each layer, the model can achieve information exchange and feature correction between nodes within the same semantic level.
[0078] After updating the graph structure within each layer, a top-down semantic feedback operation is performed. This process transmits the semantic representations of high-level regions and limbs to low-level keypoint nodes through a back-mapping matrix, enabling high-level global semantic information to correct local response biases in low-level nodes. Through this feedback mechanism, even if some keypoints lack reliable local observation features due to occlusion, reasonable predictions can still be obtained with the help of high-level semantic priors, thereby effectively enhancing the structural consistency of the overall pose.
[0079] It is worth noting that after at least one graph structure update phase, the control instruction set triggers the confusion assessment and feature intervention process again. Specifically, after each graph structure update phase executed by the hierarchical graph inference network, the confusion score of each full-body 2D keypoint is calculated again based on the current node features. The full-body 2D keypoints to be intervened are then redefined, and the original embedding features of the corresponding nodes are weighted and fused with the canonical embeddings. Through this mechanism, a closed-loop optimization process alternating between graph inference and causal intervention is formed, enabling the model to form a positive cycle between structural priors and causal deconfusion. This gradually weakens the spurious correlation effects introduced by occlusion, background noise, and part similarity, and suppresses the cumulative amplification of local prediction errors during graph propagation.
[0080] Finally, the structurally corrected 2D keypoint features obtained after the aforementioned alternating optimization are input into the 2D keypoint prediction head. The prediction head outputs the discrete probability distributions of each keypoint in the horizontal and vertical directions, with each discrete probability distribution representing the probability value of the keypoint potentially being located at each discrete position in the corresponding dimension. During the decoding stage, continuous coordinates are obtained using expectation calculation, i.e., the probability distributions in each dimension are weighted and summed, and the probability values are used as weights to perform a weighted average of each discrete position, resulting in sub-pixel precision continuous coordinate values. Through this method, the 2D coordinates of the 2D keypoints throughout the body are finally output.
[0081] Based on steps S110 to S140 above, in this embodiment of the invention, by constructing an adaptive causal intervention mechanism based on confusion assessment, the degree of interference from occlusion, background noise, and appearance similarity of each full-body 2D keypoint can be explicitly quantified. Highly confused keypoints are dynamically screened and adaptive soft replacement correction is performed, effectively avoiding the problems of over-intervention on simple samples, under-intervention on difficult samples, and over-regularization of rare effective poses caused by traditional hard replacement strategies. By constructing a three-level hierarchical graph reasoning architecture of keypoint layer, limb layer, and region layer, and simultaneously modeling local skeletal topological constraints, limb coordination relationships, and high-order semantic consistency across the entire body, combined with an alternating optimization mechanism of graph reasoning and causal intervention, it can progressively optimize under the guidance of structural priors. The first step eliminates local prediction bias and suppresses the propagation of errors in the skeleton graph. Through a dual-path joint optimization mechanism and a two-stage hierarchical data augmentation strategy, the model is guided to learn causal rather than statistically related feature representations. This achieves an optimal balance between simulating real complex scenarios and maintaining the stability of feature distribution, significantly improving the model's generalization ability on out-of-distribution samples. Fourth, the above technical solution forms a complete technical closed loop from feature learning, uncertainty assessment, structural constraints to training optimization, with synergistic effects. It effectively solves the problem of insufficient robustness of existing methods in challenging scenarios such as severe occlusion, drastic deformation, and complex backgrounds, and is applicable to various application scenarios such as intelligent security, human-computer interaction, motion analysis, sports assessment, autonomous driving, and medical rehabilitation.
[0082] The whole-body two-dimensional keypoint localization method of the present invention determines the degree of confusion information based on the embedding features of the whole-body two-dimensional keypoints, including: generating discrete probability distributions of each keypoint in the horizontal and vertical directions according to the embedding features of the whole-body two-dimensional keypoints; calculating the confusion score of each whole-body two-dimensional keypoint according to the discrete probability distribution; and calculating the average degree of confusion based on the confusion scores of all whole-body two-dimensional keypoints.
[0083] This embodiment quantifies the prediction uncertainty of each key point into a confusion score and further aggregates it into a sample-level average confusion degree, realizing a multi-level interference degree assessment from individual key points to the overall sample, providing a fine-grained and globally consistent quantitative basis for subsequent adaptive intervention decisions.
[0084] Furthermore, the confusion score of each two-dimensional key point in the whole body is calculated based on the discrete probability distribution, including: calculating the confusion score of each two-dimensional key point in the whole body by combining the maximum response value of the discrete probability distribution with the information entropy; and determining the number of key points to be intervened based on the comparison result of the average confusion degree and the preset first threshold, wherein when the average confusion degree is lower than the first threshold, the number of key points to be intervened is zero.
[0085] This embodiment combines the maximum response value and information entropy to measure the uncertainty of key point prediction in two ways, and sets that no intervention is performed when the average confusion level is lower than a first threshold, thereby protecting low-confusion samples and avoiding unnecessary feature correction from interfering with the accurate prediction of key points.
[0086] Furthermore, based on the comparison between the average level of confusion and the first threshold, the number of key points to be intervened is determined, which also includes: when the average level of confusion is higher than the preset second threshold, the number of key points to be intervened is the preset upper limit; when the average level of confusion is between the first threshold and the second threshold, the number of key points to be intervened is determined by linear interpolation; and key points to be intervened are selected in descending order of confusion score.
[0087] This embodiment dynamically determines the number of interventions by setting a first threshold and a second threshold and using linear interpolation. At the same time, it prioritizes the key points to be intervened according to the confusion score from high to low, thus achieving a precise match between the intervention intensity and the intervention object. This avoids over-intervention in low confusion scenarios and ensures that a sufficient number of key points can be effectively corrected in high confusion scenarios. Furthermore, it prioritizes the correction of the most severely disturbed key points, thereby improving the targeting of de-confusion correction and the overall positioning accuracy.
[0088] Furthermore, feature extraction is performed on the human image data to obtain the full-body two-dimensional keypoint embedding features. This includes: inputting the human image data into a multi-stage convolutional backbone structure for hierarchical feature extraction, wherein each convolutional stage sequentially performs convolution operations, normalization processing, and nonlinear activation processing; introducing a cross-stage partial connection mechanism in the multi-stage convolutional backbone structure to divide the input features into bypass branches and transformation branches. The bypass branches maintain an identity mapping, and the transformation branches are convolved and fused with the bypass branches after convolution transformation, and channel compression is performed through convolution; inputting the high-level semantic features output by the multi-stage convolutional backbone network into a gated attention enhancement module for channel response recalibration, and mapping the spatial features to the full-body two-dimensional keypoint dimension through a learnable projection matrix to obtain the full-body two-dimensional keypoint embedding features.
[0089] This embodiment uses a multi-stage convolutional backbone structure combined with a cross-stage partial connection mechanism to extract high-level semantic features while preserving fine-grained spatial location information. The channel responses are then recalibrated by a gated attention module, and finally, the spatial features are mapped to the keypoint dimension through a learnable projection matrix. This generates full-body two-dimensional keypoint embedding features that take into account both local details and global semantics, providing a high-quality feature foundation for subsequent confusion assessment and de-confusion correction.
[0090] Furthermore, the high-level semantic features are input into the gated attention enhancement module for channel response recalibration, including: obtaining the global description vector of each channel through global average pooling; performing nonlinear mapping on the global description vector through a multilayer perceptron to generate channel attention weights; and applying the channel attention weights element by element to the high-level semantic features to obtain the enhanced features.
[0091] This embodiment compresses the spatial information of each channel through global average pooling and generates attention weights by learning the nonlinear interaction between channels through a multilayer perceptron. The weights are then applied element by element to the original features, enabling the network to adaptively strengthen the channels that contribute significantly to key point localization and suppress irrelevant or interfering channel responses, thereby improving the discriminative power and robustness of feature representation.
[0092] Furthermore, based on the confusion level information, an adjustment mode is determined, including: when the confusion level information indicates that the average confusion level is greater than or equal to a first threshold, the adjustment mode is determined to be an adjustment mode; when the confusion level information indicates that the average confusion level is less than the first threshold, the adjustment mode is determined to be a fixed mode.
[0093] This embodiment adaptively switches adjustment modes by comparing the average level of confusion with a first threshold. When the level of confusion is low, it maintains a fixed mode to preserve the stability of the original features. When the level of confusion reaches or exceeds the threshold, it switches to adjustment mode to initiate de-confusion correction. This achieves precise intervention for highly confused samples and effective protection for low-confused samples, avoiding excessive interference with accurately predicted key points.
[0094] Further, the de-obfuscation correction operation includes: constructing a binary mask matrix, which is used to mark whether each keypoint has been intervened; obtaining a pre-defined canonical embedding table, which contains canonical embeddings corresponding to each keypoint; for each keypoint marked as having undergone intervention, determining the fusion weight based on its obfuscation score, and weighting and fusing its original embedding features with the corresponding canonical embeddings according to the fusion weight to obtain the keypoint features after intervention; the human body structure modeling operation includes: inputting the intervention-corrected full-body two-dimensional keypoint features into a hierarchical graph inference network for human body structure modeling. The hierarchical graph inference network includes at least three semantic layers: keypoint layer, limb layer, and region layer. Through bottom-up inter-layer aggregation, graph structure updates within each layer, and top-down semantic feedback, the structure-corrected full-body two-dimensional keypoint features are obtained; the structure-corrected full-body two-dimensional keypoint features are input into the full-body two-dimensional keypoint prediction head, which outputs discrete probability distributions in the horizontal and vertical directions, and decodes the coordinates of the full-body two-dimensional keypoints based on the discrete probability distributions.
[0095] This embodiment accurately locates key points to be intervened using a binary mask matrix, and uses a standardized embedding table to assign different fusion weights according to the confusion score to achieve adaptive soft replacement. Then, through bottom-up aggregation, intra-layer updates, and top-down semantic feedback of a three-level hierarchical graph inference network of key point layer, limb layer, and region layer, structural constraints and corrections are performed. Finally, the coordinates are decoded by the prediction head, forming a complete closed loop from individual feature correction to global structural optimization, which significantly improves the accuracy and structural consistency of whole-body two-dimensional key point localization.
[0096] Another method for locating two-dimensional key points of the whole body according to one embodiment of the present invention, such as Figures 2-6 As shown, the method includes the following steps:
[0097] Step 201: Obtain the human image data to be processed and perform preprocessing operations.
[0098] Human image data can be derived from static monocular RGB images or continuous frame results of video sequences. The target region is cropped and aligned based on the human detection bounding box. Subsequently, the input image is uniformly scaled to a resolution of 384×288 and affine transformation, pixel value standardization, and color perturbation processing are performed.
[0099] During the training phase, to enhance the model's adaptability to occlusion, viewpoint changes, and local missing data, this embodiment employs a hierarchical data augmentation strategy, including random horizontal flipping, random half-body transformation, random bounding box scaling, random rotation, HSV color jitter, and CoarseDropout occlusion simulation.
[0100] Step 202: Input the preprocessed human image into the feature encoding network to perform hierarchical feature extraction. The feature encoding network includes a multi-stage convolutional backbone structure, a cross-stage partial connection mechanism, and a gated attention enhancement module. Compared with the limitation of traditional convolutional neural networks that can only extract local features, this invention retains fine-grained positional information through the cross-stage partial connection mechanism and recalibrates the channel response through the gated attention enhancement module, thereby significantly improving the expressive ability of embedded features at the level of two-dimensional key points of the whole body.
[0101] like Figure 3 The diagram shows the structure of the feature encoding network and the gated attention enhancement module. The feature encoding network contains a multi-stage convolutional backbone structure. Each stage sequentially performs convolution, normalization, and non-linear activation operations. Its basic calculation formula is as follows:
[0102] (1) in For the first The layer's output features, where Conv represents the convolution operation. For the first The convolutional kernel weights of the layer, The convolution stride is... For fill size, It is a non-linear activation function. This is the normalization function.
[0103] In this embodiment, the convolutional kernel preferably adopts a combination of 3×3 convolution and 1×1 convolution, where the 3×3 convolution is used to extract local spatial context features, and the 1×1 convolution is used for channel transformation and compression calculation; the normalization layer can adopt batch normalization or layer normalization, and its calculation formula is as follows:
[0104] (2)
[0105] in The mean of features within a batch or layer. For characteristic variance, and For learnable scaling and translation parameters, For numerically stable terms, the preferred value range is: to .
[0106] To reduce redundant computation while maintaining feature representation capabilities, a cross-stage partial connection mechanism is introduced into the convolutional backbone structure. This mechanism divides the input features into a bypass branch and a transformation branch. The bypass branch directly retains the input features, while the transformation branch extracts higher-level semantic representations through several convolutional blocks. These representations are then concatenated according to the following formula:
[0107] (3)
[0108] in The identity mapping feature of the bypass branch. For the input features of the transform branch, This is a convolutional transformation operation; after fusion, channel compression is achieved through convolution.
[0109] The high-level semantic features output from the backbone network are further input into the gated attention enhancement module. Specifically, the channel description vector is first obtained through global average pooling, calculated using the following formula:
[0110] (4)
[0111] in For the first The spatial location of each channel The feature values at each location are then used; the channel weights are obtained by performing a nonlinear mapping through a multilayer perceptron, and the calculation formula is as follows:
[0112] (5)
[0113] in and For learnable weight matrix, Represents the ReLU or GELU activation function. This represents the Sigmoid function; then, the weights are applied element-wise to the original features to obtain the enhanced features, calculated using the following formula:
[0114] (6)
[0115] in The Hadamard product (element-wise multiplication) is represented; finally, the spatial features are converted into a full-body 2D keypoint-level embedding representation through flattening and linear mapping.
[0116] (7)
[0117] in For batch size, The number of two-dimensional key points on the whole body to be detected. This represents the number of feature channels.
[0118] Step 203: Construct an adaptive causal intervention module to de-obfuscate the embedded features of the whole-body two-dimensional keypoints. The adaptive causal intervention module includes a prediction modeling submodule, an obfuscation evaluation submodule, and a feature intervention submodule.
[0119] The predictive modeling submodule generates discrete predicted distributions of each full-body 2D keypoint in the horizontal and vertical directions. The confusion assessment submodule calculates the confusion score of each full-body 2D keypoint based on the maximum response value and information entropy of the predicted distribution, and determines the sample-level average confusion degree based on the confusion scores of all full-body 2D keypoints, thereby adaptively determining the number of full-body 2D keypoints to be intervened. Preferably, upper and lower limits are set for the number of full-body 2D keypoints to be intervened to avoid over-intervention or under-intervention.
[0120] The feature intervention submodule selects full-body 2D keypoints to be intervened based on their confusion scores from high to low. It then performs a weighted fusion of the original embedded features and the corresponding canonical embeddings according to fusion weights to obtain the intervened full-body 2D keypoint features. The fusion weights are obtained from the confusion scores through a monotonic mapping function, ensuring that full-body 2D keypoints with higher confusion levels are subject to stronger canonical prior constraints. The canonical embeddings are learnable reference embeddings set according to the full-body 2D keypoint indices, used to provide stable structural priors under high confusion conditions.
[0121] like Figure 4The diagram shows the structure of the causal intervention module: The causal intervention module includes a predictive modeling submodule, a confusion assessment submodule, and a feature intervention submodule. The predictive modeling submodule outputs the discrete logit values of the whole-body two-dimensional key points in the horizontal and vertical directions through the horizontal and vertical classification heads, respectively, denoted as... and .
[0122] After obtaining the logit values in the horizontal and vertical directions, a coordinate probability distribution is constructed using the Softmax function with a temperature parameter. The calculation formula is as follows:
[0123] (8)
[0124] (9)
[0125] in Used to adjust the smoothness of the distribution, the preferred value range is [value range missing]. Numerical stability can be significantly improved and gradient explosion can be suppressed by performing a maximum shift before Softmax.
[0126] The confusion assessment submodule uses the maximum response value of the coordinate distribution and information entropy to jointly model the uncertainty of the whole-body two-dimensional keypoints. For the first... The formula for calculating the confusion score of each full-body two-dimensional key point is as follows:
[0127] (10)
[0128] The formula for calculating information entropy is:
[0129] (11)
[0130] And the balance weights satisfy A higher confusion score indicates that the full-body 2D keypoint is more likely to be affected by occlusion, background noise, or similarity of parts.
[0131] An overall statistic is constructed based on the confusion scores of all full-body 2D keypoints, and the average confusion level is calculated using the following formula:
[0132] (12)
[0133] In this embodiment, the number of interventions N is adaptively determined based on the average confusion level: if the average confusion level is below a first threshold (e.g., 0.3), then N=0, and no intervention is performed to avoid over-intervention; if the average confusion level is above a second threshold (e.g., 0.7), then N=13 (the upper limit); otherwise, N is determined by linear interpolation. After determining N, N key points are selected from high to low confusion scores to perform soft replacement interventions. The intervention decision does not depend on the absolute confusion score threshold of a single key point, but is based on the sample-level average confusion level and relative ranking to achieve a balance between accuracy and computational cost.
[0134] The feature intervention submodule constructs a binary mask matrix based on the results of the full-body two-dimensional key point screening:
[0135] (13)
[0136] Furthermore, canonical embedding is used to perform soft replacement correction on highly confusing full-body 2D keypoint features. The formula for calculating the replacement features is as follows:
[0137] (14)
[0138] in Indicates the first Mask values for 2D key points of the whole body (1 indicates intervention, 0 indicates no intervention). This represents the fusion weights obtained by cropping the confusion scores after Sigmoid mapping. This represents the canonical embedding after adaptive adjustment based on the mean of global features.
[0139] The specification embedding table For globally learnable parameters, the optimal dimension is:
[0140] (15)
[0141] It is initialized using a normal distribution:
[0142] (16)
[0143] During training, the canonical embedding table, along with the rest of the network parameters, participates in end-to-end optimization to provide a more stable feature reference for highly confusing full-body 2D keypoints.
[0144] The above-mentioned specification embedding table The core learning objective is to learn an "idealized, context-invariant baseline feature" for each full-body 2D keypoint (K in total). This baseline feature aims to characterize the deep semantic representation that the keypoint should have under unoccluded, standard viewpoint, and typical pose conditions. In other words, it is a "prototype feature" aggregated from the training data, which removes individual appearance differences, occlusion artifacts, and background noise interference.
[0145] Specifically, specification embedding The canonical embedding of the k-th keypoint does not directly correspond to any specific training image. Instead, it implicitly extracts commonalities from all training samples through backpropagation. During training, when the k-th keypoint of a sample is judged to be highly confusing (e.g., occluded or motion-blurred), the network will remove its original embedding features. With specification embedding Soft fusion (weighted averaging) is performed. The final whole-body keypoint prediction loss (L) is minimized. keypoint ) and counterfactual consistency loss (L cf ), Standard Embedding Driven towards updating in a direction that enables the fused features to produce accurate predictions. In other words, the canonical embedding learns the feature center of the keypoint in its "normal" state across all training samples, thus serving as a stable and informative prior to repair the corrupted original features under high confusion conditions.
[0146] To prevent canonical embeddings from degenerating into ordinary learnable parameters during training (e.g., merely fitting the average features of the training set and losing robustness), this invention also introduces implicit constraints: canonical embeddings participate in gradient updates only through counterfactual branches (i.e., only when an intervention is performed), and the weights are fused. The obfuscation score is adaptively controlled to avoid interference from low-obfuscation samples by the canonical embedding. Furthermore, in a preferred embodiment, a slow update strategy (e.g., using exponential moving average) or a regularization term (e.g., constraining its KL divergence with the original feature distribution of the corresponding keypoint) can be applied to the canonical embedding to further maintain its stability as an "ideal baseline," but the invention is not limited thereto. Through the above mechanism, the canonical embedding can effectively encode the anatomical priors of human pose, becoming a "causal anchor" in obfuscation correction, thereby significantly improving the model's generalization ability and localization accuracy in complex scenes.
[0147] It should be noted that the reason this invention employs an alternating optimization mechanism of "graph reasoning-causal intervention," rather than performing a single global intervention before hierarchical graph reasoning, is that the initial causal intervention is based solely on the raw embedded features without structural constraints. At this point, the topological relationships between keypoints have not been utilized, leading to a biased assessment of confusion levels. Once the first round of graph structure update has been completed, keypoint features propagate information using prior knowledge of the human skeleton and dynamic adjacency matrices. Some keypoints that were initially predicted to be off-target will move closer to positions conforming to anatomical patterns, resulting in a sharper prediction distribution and reduced information entropy, thus significantly altering the confusion score of each keypoint. For example, a keypoint initially misclassified as "left elbow" may experience a significant decrease in prediction uncertainty after structural reasoning with adjacent nodes such as "left upper arm" and "left wrist." At this point, performing confusion assessment and causal intervention again allows for the re-screening of highly confused keypoints based on more accurate structural features and targeted deconfusion correction, achieving a progressive optimization of "structural prior-assisted deconfusion." Conversely, if intervention is only performed once before inference, subsequent structural inference cannot benefit from iterative obfuscation removal, and the initially incompletely corrected obfuscated features may accumulate and amplify during multi-layer graph propagation. Through a closed loop of "deobfuscation-structural inference-re-intervention," the model can form a positive cycle between structural priors and causal deobfuscation, gradually approaching the ideal human pose features, thereby significantly improving localization accuracy and robustness in challenging scenarios such as severe occlusion and similar appearances.
[0148] Step 204: Input the whole-body 2D keypoint features processed by adaptive causal intervention into a hierarchical graph inference network for human structure modeling. The hierarchical graph inference network includes three semantic layers: the whole-body 2D keypoint layer, the limb layer, and the region layer. The whole-body 2D keypoint layer is used to retain fine-grained positional information, the limb layer is used to model the collaborative relationship of skeletal chains and functional joint combinations, and the region layer is used to encode high-order semantic information of the head, trunk, upper limbs, and lower limbs.
[0149] The hierarchical graph inference network includes bottom-up inter-layer aggregation, graph structure updates within each layer, and top-down semantic feedback. Bottom-up inter-layer aggregation is used to aggregate low-level full-body 2D keypoint features into limb and region-level semantic representations; graph structure updates within each layer are used to propagate and correct features by combining human topological priors and inter-node relationships; top-down semantic feedback is used to backpropagate high-level structural semantics to low-level full-body 2D keypoint nodes to correct local prediction biases.
[0150] Preferably, after at least one graph structure update phase, confusion assessment and feature intervention are performed again to form an alternating optimization mechanism of "graph reasoning-causal intervention". Through this mechanism, the spurious correlation effects introduced by occlusion, background noise and part similarity can be gradually weakened under the guidance of structural priors, and the accumulation of local prediction errors during graph propagation can be suppressed.
[0151] like Figure 5 The diagram shows the structure of a hierarchical graph inference network: the network divides graph nodes into three semantic levels: a full-body 2D keypoint layer, a limb layer, and a region layer. The full-body 2D keypoint layer preserves fine-grained coordinate information, the limb layer models skeletal chains and functional joint combinations, and the region layer encodes higher-order semantic regions such as the head, torso, upper limbs, and lower limbs.
[0152] Mapping relationship definition from key point layer to limb layer: Based on human anatomy, this embodiment divides the two-dimensional key points of the whole body into the following limb nodes (each group is one limb node).
[0153] Table 1: Key Point Layer → Limb Layer (Mapping Relationship)
[0154] left upper arm left shoulder, left elbow left forearm Left elbow, left wrist left hand Key points of left wrist + left palm + 21 fingers of left hand right upper arm Right shoulder, right elbow right forearm Right elbow, right wrist right hand Key points of the right wrist, right palm, and 21 fingers of the right hand Left thigh Left hip, left knee left calf Left knee, left ankle left foot Six key points of the left ankle and left foot Right thigh Right hip, right knee Right calf Right knee, right ankle Right foot Six key points of the right ankle and right foot trunk Neck, left shoulder, right shoulder, left hip, right hip, spine (including but not limited to: nose, left eye, right eye, left ear, right ear, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, right ankle) head All 68 key points on the face (in the order marked on the COCOWholeBody face).
[0155] Table 2: Limb Layer → Region Layer (Mapping Relationship)
[0156] Left arm area Left upper arm, left forearm, left hand Right arm area Right upper arm, right forearm, right hand Left leg area Left thigh, left calf, left foot Right leg area Right thigh, right calf, right foot torso area trunk Head area head
[0157] Table 3: Key Point Layer → Region Layer (Direct mapping relationship, used for skip connections or fast semantic transfer)
[0158] Left arm area Key points of the left shoulder, left elbow, left wrist, left palm, and 21 fingers of the left hand Right arm area Key points of the right shoulder, right elbow, right wrist, right palm, and 21 fingers of the right hand Left leg area Six key foot points: left hip, left knee, left ankle, and left foot Right leg area Six key foot points: right hip, right knee, right ankle, and right foot. torso area Neck, left shoulder, right shoulder, left hip, right hip, nose, left eye, right eye, left ear, right ear, spine, etc. Head area All 68 key points of the face
[0159] Instructions for use:
[0160] Tables 1 and 2 are used for bottom-up aggregation: key point features → limb features → region features.
[0161] Table 3 is used for top-down feedback or cross-layer residual connections to directly pass regional semantics to key points, enhancing the global consistency of local features.
[0162] Mapping matrix Based on the table above, the matrix structure is as follows: if a lower-level node belongs to a higher-level node, the corresponding matrix element is set to 1; otherwise, it is set to 0. In a practical implementation, non-binary weights can also be assigned based on anatomical distance or learnable weights.
[0163] In the bottom-up aggregation process, the features of low-level nodes are mapped to the representations of high-level nodes according to the prior knowledge of human body structure. The calculation formula is as follows:
[0164] (17)
[0165] in Indicates the first Layer node characteristics, Indicates the first layer to the first Inter-layer mapping matrix of layers, This represents the learnable upward transformation parameters. Through this process, the model can progressively integrate local observation information from multiple 2D keypoints across the entire body, forming a more stable high-level representation of limb posture and overall body layout.
[0166] During the graph structure update process within each layer, residual graph convolution is used for node feature iteration, and the calculation formula is as follows:
[0167] (18)
[0168] in Indicates the first The normalized adjacency matrix of the layer, Indicates the first Graph transformation parameters of the layer, This represents a non-linear activation function. The adjacency matrix not only includes static skeletal topological priors but can also be dynamically adjusted based on the node feature similarity of the current sample, thereby enhancing the model's adaptability to unconventional poses and complex interaction scenarios.
[0169] In the top-down feedback process, the high-level region semantic representation and limb semantic representation are fed back to the low-level whole-body 2D keypoint nodes through a back-mapping matrix, and its update form is as follows:
[0170] (19)
[0171] in Indicates the first layer to the first The inverse mapping matrix of the layer, This represents the learnable down-transformation parameters. This process can utilize high-level global semantic information to correct local response biases and enhance the structural consistency of the overall attitude.
[0172] It is worth noting that after each layer of the graph structure is updated, the confusion assessment and causal intervention process can be performed again, thereby forming an alternating optimization mechanism of "graph reasoning-causal intervention" to gradually eliminate potential confusion factors in deep features and continuously strengthen the topological consistency and semantic coordination between the two-dimensional key points of the whole body.
[0173] Step 205: Input the final obtained features into the full-body 2D keypoint prediction head, output the probability distributions in the horizontal and vertical directions respectively, and obtain the 2D coordinates of the full-body 2D keypoints through expectation calculation.
[0174] The final obtained full-body 2D keypoint features are input into the full-body 2D keypoint prediction head, which outputs the probability distributions in the horizontal and vertical directions respectively. During the decoding stage, continuous coordinates are obtained using the expectation calculation method, with the following formula:
[0175] (20)
[0176] (twenty one)
[0177] Alternatively, the maximum response position can be used in conjunction with local offset estimation for sub-pixel level correction to further improve positioning accuracy.
[0178] Step 206: During model training, a dual-path joint optimization mechanism is introduced, which includes a counterfactual path with causal intervention and an observation path without intervention. A two-stage training strategy and a hierarchical data augmentation mechanism are adopted to improve the model's generalization ability and localization accuracy.
[0179] like Figure 6 The diagram illustrates a dual-path joint optimization and two-stage training strategy: During model training, a counterfactual path containing causal intervention and an observation path without intervention are introduced for dual-path joint optimization. The observation path stops gradient propagation and is used only to provide a control distribution under the uninterrupted condition; the counterfactual path is used to learn robust feature representations corrected for causal intervention.
[0180] Regarding the design of the loss function, this embodiment comprehensively adopts the whole-body 2D keypoint lateral classification loss, the whole-body 2D keypoint longitudinal classification loss, and the counterfactual consistency loss. The total loss function calculation formula is as follows:
[0181] (twenty two)
[0182] In this embodiment, the counterfactual consistency loss Lcf is used to measure the predictive consistency between the counterfactual branch (i.e., the branch with causal intervention) and the observed branch (i.e., the branch without intervention). Specifically, let the horizontal probability distribution of the observed branch output be... The probability distribution in the vertical direction is The corresponding probability distribution of the counterfactual branch output is: , ,but One of the following two forms or a combination thereof may be used:
[0183] KL divergence based on probability distribution:
[0184] (twenty three)
[0185] in This form directly constrains the output distribution of the two branches, which helps maintain the consistency of the probability distribution shape.
[0186] Euclidean distance based on decoded coordinates: First, the probability distribution is decoded into coordinate values:
[0187] (twenty four)
[0188] (25)
[0189] Then calculate the Euclidean distance between the coordinates of the two branches: (26)
[0190] This form more directly constrains the consistency of the final predicted coordinates and is more computationally efficient.
[0191] In this embodiment, a KL divergence-based approach is preferred because the probability distribution contains information about the uncertainty of the model's predictions, which can more finely guide the learning of canonical embeddings. Both approaches are within the scope of this invention. Furthermore, the observation branch stops gradient propagation during training (i.e., (This is only used as a reference for consistency constraints; the gradient of the counterfactual branch is obtained through...) Backpropagation is used to update network parameters and canonical embedding tables.
[0192] The weight hyperparameter of the counterfactual consistency loss. Preferred settings ;when When the size is too small, canonical embeddings struggle to learn effective deobfuscated representations; when... If the value is too large, the model will be subject to excessive constraints, affecting the convergence of the main task.
[0193] This embodiment further employs a two-stage training strategy:
[0194] Phase 1 (Rounds 1 to 270): Employing strong data augmentation techniques, including random horizontal flipping, random half-body transformations, Random scaling of the bounding box within the range Random rotation within the range, HSV color jitter, and application probability are The CoarseDropout occlusion simulation is used to enhance the model's robustness to occlusion, deformation, and complex backgrounds.
[0195] Phase 2 (Rounds 271 to 420): Reduce the intensity of data augmentation; bounding box transformations only perform scaling and rotation operations; translation factor is set to... The scaling range is adjusted to The rotation range is adjusted to At the same time, reduce the probability of applying CoarseDropout to This is to promote further convergence of the model on relatively clean data distributions and improve prediction accuracy.
[0196] To improve training stability, this embodiment may also introduce gradient clipping, exponential moving average parameter updates, and random deactivation strategies to reduce gradient oscillations and overfitting risks.
[0197] In the inference stage, to balance accuracy and speed, the number of whole-body 2D key points involved in the intervention can be reduced, and some dynamic adjustment items can be turned off, retaining only the high-yield confusion correction and hierarchical graph inference stages, thereby reducing computational overhead while maintaining the stability of whole-body 2D key point prediction results.
[0198] In a preferred embodiment, the training dataset may be the COCOWholeBody dataset or other datasets containing full-body 2D keypoint annotations, with a total number of full-body 2D keypoints. Preferred It covers the entire body, including the body (17 pieces), face (68 pieces), hands (21 x 2 pieces), and feet (6 x 2 pieces).
[0199] Based on steps S201 to S206 above, a multi-stage feature extraction architecture with fine-grained feature enhancement is constructed in this embodiment of the invention. This invention overcomes the shortcomings of traditional CNN models and benchmark schemes, which suffer from the loss of fine-grained positional information due to excessively large convolutional strides and single feature transmission paths. It adopts a multi-stage convolutional backbone structure combined with a gated attention enhancement module to construct a feature extraction and representation scheme that balances semantic expression and positional accuracy. This architecture, through a hierarchical feature fusion mechanism, effectively preserves the fine-grained spatial positional information of two-dimensional key points throughout the body while maintaining high-level semantic feature extraction capabilities. Simultaneously, the gated attention module enhances the response strength to features of key human body parts through dynamic adaptive allocation of channel weights, ensuring the effective transmission of features from small targets and easily occluded areas in the deep network from the source. This provides a reliable foundation for subsequent feature correction and inference optimization, which is one of the core reasons why this method achieves significant advantages in medium-sized target detection (AP(M) improvement of 12.26%) and high IoU threshold localization (AP.75 improvement of 9.92%).
[0200] Second, an adaptive causal intervention mechanism for pseudo-correlation feature correction is proposed. Addressing the prediction bias of full-body 2D keypoints caused by obfuscation factors such as background noise and occlusion artifacts in complex scenes, this invention introduces an adaptive causal intervention module, constructing an uncertainty-driven dynamic feature correction mechanism. This module quantifies the confidence and obfuscation level of each full-body 2D keypoint by jointly measuring the maximum response value and information entropy of the full-body 2D keypoint heatmap, achieving accurate screening of highly obfuscated full-body 2D keypoints. Subsequently, through targeted causal feature correction operations, the interference of pseudo-correlation features on the prediction results is effectively weakened, guiding the model to learn feature representations that are causally related to human posture rather than statistically correlated. This fundamentally improves the model's anti-interference ability and robustness under occlusion and complex background interference, resulting in a stable improvement in detection accuracy across various complex scenes.
[0201] Third, a three-level hierarchical graph reasoning structural constraint modeling scheme was designed. To address the problem of skeletal structure breakage caused by the loss of some full-body 2D keypoint features, this invention constructs a three-level hierarchical graph reasoning architecture consisting of a full-body 2D keypoint layer, a limb layer, and a region layer, achieving unified modeling of multi-scale human structure priors. This architecture models local skeletal topological constraints at the full-body 2D keypoint layer, models the cooperative motion relationships between adjacent limbs at the limb layer, and models high-order semantic consistency constraints across the entire body at the region layer. Simultaneously, combined with an alternating optimization mechanism of "graph reasoning-causal intervention," local prediction biases are gradually eliminated under the guidance of full-body structural priors, suppressing the propagation of erroneous predictions in the skeleton graph. Even when some full-body 2D keypoints are lost due to occlusion, reasonable inference can still be achieved through global structural constraints, effectively alleviating the problem of skeletal structure breakage caused by occlusion and further improving the overall detection accuracy and recall rate of the model.
[0202] Fourth, a dual-path joint optimization generalization training strategy is adopted. To further enhance the model's generalization ability to out-of-distribution scenes, this invention employs a dual-path joint optimization mechanism and a two-stage training strategy. On the one hand, by constraining the consistency between the counterfactual path and the observation path, the model is guided to learn causal feature representations unaffected by confounding variables, avoiding overfitting of the model to statistically relevant noise in the training set. On the other hand, the two-stage hierarchical data augmentation mechanism progressively simulates complex situations in real-world scenes such as occlusion, deformation, viewpoint changes, and illumination changes, achieving an optimal balance between fully simulating complex real-world scenes and maintaining the stability of the training feature distribution. This effectively improves the model's generalization ability on out-of-distribution samples, enabling this method to maintain stable high performance on test samples of different scales and difficulties.
[0203] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0204] This invention also provides a whole-body two-dimensional key point positioning device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as described herein. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0205] Figure 7 According to one embodiment of the present invention, a whole-body two-dimensional key point positioning device includes:
[0206] The acquisition module 301 is used to acquire human image data to be processed, extract features from the human image data to obtain the whole body two-dimensional key point embedding features, and determine the degree of confusion information based on the whole body two-dimensional key point embedding features. The degree of confusion information is used to characterize the degree of interference of each whole body two-dimensional key point.
[0207] The adjustment module 302 is used to determine the adjustment mode based on the degree of confusion information. The adjustment mode is used to perform de-confusion correction on the embedded features of the whole body two-dimensional key points. The adjustment mode includes an adjustment mode and a fixed mode.
[0208] The generation module 303 is used to generate a control instruction set in response to the adjustment mode. The control instruction set is used to perform deobfuscation correction operations and human body structure modeling operations to output the coordinates of two-dimensional key points of the whole body.
[0209] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.
[0210] According to one embodiment of the present invention, an electronic device is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the above-described whole-body two-dimensional key point localization method during runtime.
[0211] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0212] Step S1: Obtain human image data to be processed, and extract features from the human image data to obtain the whole body two-dimensional key point embedding features. Determine the degree of confusion based on the whole body two-dimensional key point embedding features, wherein the degree of confusion is used to characterize the degree of interference of each whole body two-dimensional key point.
[0213] Step S2: Based on the degree of confusion information, determine the adjustment mode. The adjustment mode is used to de-confusion correct the embedded features of the whole body two-dimensional key points. The adjustment mode includes the adjustment mode and the fixed mode.
[0214] Step S3: In response to the adjustment mode, a control instruction set is generated. The control instruction set is used to perform deobfuscation correction operations and human body structure modeling operations to output the coordinates of the two-dimensional key points of the whole body.
[0215] According to one embodiment of the present invention, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the storage medium is located to execute the above-described whole-body two-dimensional key point localization method.
[0216] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:
[0217] Step S1: Obtain human image data to be processed, and extract features from the human image data to obtain the whole body two-dimensional key point embedding features. Determine the degree of confusion based on the whole body two-dimensional key point embedding features, wherein the degree of confusion is used to characterize the degree of interference of each whole body two-dimensional key point.
[0218] Step S2: Based on the degree of confusion information, determine the adjustment mode. The adjustment mode is used to de-confusion correct the embedded features of the whole body two-dimensional key points. The adjustment mode includes the adjustment mode and the fixed mode.
[0219] Step S3: In response to the adjustment mode, a control instruction set is generated. The control instruction set is used to perform deobfuscation correction operations and human body structure modeling operations to output the coordinates of the two-dimensional key points of the whole body.
[0220] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0221] According to one embodiment of the present invention, a computer program product is also provided, including a computer program that, when executed by a processor, implements the above-described method for locating two-dimensional key points throughout the body.
[0222] Optionally, in this embodiment, the above-mentioned computer program product can be configured as a computer program that performs the following steps:
[0223] Step S1: Obtain human image data to be processed, and extract features from the human image data to obtain the whole body two-dimensional key point embedding features. Determine the degree of confusion based on the whole body two-dimensional key point embedding features, wherein the degree of confusion is used to characterize the degree of interference of each whole body two-dimensional key point.
[0224] Step S2: Based on the degree of confusion information, determine the adjustment mode. The adjustment mode is used to de-confusion correct the embedded features of the whole body two-dimensional key points. The adjustment mode includes the adjustment mode and the fixed mode.
[0225] Step S3: In response to the adjustment mode, a control instruction set is generated. The control instruction set is used to perform deobfuscation correction operations and human body structure modeling operations to output the coordinates of the two-dimensional key points of the whole body.
[0226] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.
[0227] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0228] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.
[0229] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0230] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0231] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0232] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for locating two-dimensional key points throughout the body, characterized in that, Includes the following steps: The human image data to be processed is acquired, and feature extraction is performed on the human image data to obtain the whole body two-dimensional key point embedding features. Based on the whole body two-dimensional key point embedding features, the degree of confusion information is determined, wherein the degree of confusion information is used to characterize the degree of interference of each whole body two-dimensional key point. Based on the obfuscation level information, an adjustment mode is determined. The adjustment mode is used to de-obfuscate and correct the embedded features of the whole body two-dimensional key points. The adjustment mode includes an adjustment mode and a fixed mode. In response to the adjustment mode, a set of control instructions is generated to perform deobfuscation correction and human body structure modeling operations to output the coordinates of two-dimensional key points of the whole body.
2. The whole-body two-dimensional key point localization method according to claim 1, characterized in that, Based on the embedded features of the whole-body two-dimensional key points, the degree of confusion information is determined, including: Based on the embedded features of the whole body two-dimensional key points, a discrete probability distribution of each key point in the horizontal and vertical directions is generated; The confusion score for each two-dimensional key point in the whole body is calculated based on the discrete probability distribution. The average level of confusion is calculated based on the confusion scores of all two-dimensional key points throughout the body.
3. The whole-body two-dimensional key point localization method according to claim 2, characterized in that, The confusion score for each two-dimensional keypoint in the whole body is calculated based on the discrete probability distribution, including: The confusion score of each two-dimensional key point in the whole body is calculated by combining the maximum response value of the discrete probability distribution with the information entropy. Based on the comparison between the average level of confusion and a preset first threshold, the number of key points to be intervened is determined, wherein when the average level of confusion is lower than the first threshold, the number of key points to be intervened is zero.
4. The whole-body two-dimensional key point localization method according to claim 3, characterized in that, Based on the comparison between the average level of confusion and the first threshold, the number of key points to be intervened is determined, which further includes: When the average level of confusion is higher than a preset second threshold, the number of key points to be intervened is a preset upper limit. When the average level of confusion is between the first threshold and the second threshold, the number of key points to be intervened is determined by linear interpolation. The key points to be intervened are selected in descending order of the confusion scores.
5. The whole-body two-dimensional key point localization method according to claim 1, characterized in that, Feature extraction is performed on the human image data to obtain the whole-body two-dimensional keypoint embedding features, including: The human image data is input into a multi-stage convolutional backbone structure for hierarchical feature extraction, wherein each convolutional stage sequentially performs convolution operations, normalization processing, and nonlinear activation processing. A cross-stage partial connection mechanism is introduced into the multi-stage convolutional backbone structure to divide the input features into a bypass branch and a transformation branch. The bypass branch maintains an identity mapping, and the transformation branch is convolved and fused with the bypass branch after convolution transformation. Channel compression is performed through convolution. The high-level semantic features output by the multi-stage convolutional backbone network are input into the gated attention enhancement module for channel response recalibration. The spatial features are then mapped to the two-dimensional key point dimension of the whole body through a learnable projection matrix to obtain the two-dimensional key point embedding features of the whole body.
6. The whole-body two-dimensional key point localization method according to claim 5, characterized in that, The high-level semantic features are input into the gated attention enhancement module for channel response recalibration, including: The global description vector for each channel is obtained through global average pooling; The global description vector is nonlinearly mapped using a multilayer perceptron to generate channel attention weights. The channel attention weights are applied element-wise to the high-level semantic features to obtain the enhanced features.
7. The whole-body two-dimensional key point localization method according to claim 1, characterized in that, Based on the aforementioned level of confusion information, an adjustment mode is determined, including: When the confusion level information indicates that the average confusion level is greater than or equal to the first threshold, the adjustment mode is determined to be the adjustment mode; When the confusion level information indicates that the average confusion level is less than the first threshold, the adjustment mode is determined to be the fixed mode.
8. The whole-body two-dimensional key point localization method according to claim 1, characterized in that, The deobfuscation correction operation includes: Construct a binary mask matrix, which is used to mark whether each key point has been intervened; Obtain a preset specification embedding table, which contains specification embeddings that correspond one-to-one with each key point; For each key point marked as requiring intervention, a fusion weight is determined based on its confusion score. The original embedded features are then weighted and fused with the corresponding canonical embedded features according to the fusion weight to obtain the key point features after intervention. The human body structure modeling operation includes: The two-dimensional key point features of the whole body after intervention are input into a hierarchical graph inference network for human structure modeling. The hierarchical graph inference network includes at least three semantic levels: key point layer, limb layer and region layer. Through bottom-up inter-layer aggregation, graph structure update within each layer and top-down semantic feedback, the two-dimensional key point features of the whole body after structural correction are obtained. The full-body two-dimensional keypoint features after structural correction are input into the full-body two-dimensional keypoint prediction head, which outputs discrete probability distributions in the horizontal and vertical directions respectively, and decodes the coordinates of the full-body two-dimensional keypoints based on the discrete probability distributions.
9. A two-dimensional whole-body key point positioning device, characterized in that, include: The acquisition module is used to acquire human image data to be processed, and to extract features from the human image data to obtain whole-body two-dimensional keypoint embedding features. Based on the whole-body two-dimensional keypoint embedding features, the degree of confusion information is determined, wherein the degree of confusion information is used to characterize the degree of interference of each whole-body two-dimensional keypoint. An adjustment module is used to determine an adjustment mode based on the obfuscation level information. The adjustment mode is used to perform deobfuscation correction on the embedded features of the whole body two-dimensional key points. The adjustment mode includes an adjustment mode and a fixed mode. A generation module is used to generate a set of control instructions in response to the adjustment mode. The set of control instructions is used to perform de-obfuscation correction operations and human body structure modeling operations to output the coordinates of two-dimensional key points of the whole body.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the storage medium is located to perform the whole-body two-dimensional key point localization method of any one of claims 1 to 8.