A method, medium and system for extracting a person's posture key point in a complex high-altitude environment
Patent Information
- Application Number
- CN202611052146.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]有鉴于此,本发明提供一种复杂高空环境下的吊篮人员姿态关键点提取方法、介质及系统,能够解决现有技术中存在高空吊篮作业场景下强逆光与剧烈抖动并存时人员姿态关键点提取精度严重下降的技术问题
[0029]本发明通过惯性测量单元预积分生成帧间运动补偿参数,并将运动补偿仿射变换与多曝光高动态范围融合策略前置于关键点提取流程,使输入图像在送入后续网络之前同时完成抖动抑制与光照均衡,从根本上消除了抖动与逆光对特征提取的双重干扰。在此基础上,飞行时间传感器深度图与运动掩膜联合驱动的前景分割网络将静止建筑构件从候选前景中剔除,避免了特征混淆;高空吊篮姿态解析模型以可变形卷积感受野由惯性测量单元数据驱动的方式对运动模糊区域进行定向补偿,动态骨架图构建模块将关节置信度与特征相似度实时编码进图结构,使信息传播路径随遮挡状态自适应调整;物理约束推理模块将解剖学先验以可微分形式嵌入网络,保证关键点预测在几何上的自洽性。上述各环节协同作用,使关键点提取在强逆光与剧烈抖动并存时仍能保持较高精度。
Smart Images

Figure CN122821633A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of construction safety technology, and specifically relates to a method, medium, and system for extracting key points of personnel posture in suspended platforms under complex high-altitude environments. Background Technology
[0002] Human posture key point extraction technology is widely used in the field of construction safety monitoring, especially in the analysis of personnel behavior in high-altitude suspended platform operations. Existing methods typically employ heatmap regression models based on convolutional neural networks, combined with optical flow tracing algorithms to achieve cross-frame localization of key points, and supplemented by graph convolutional networks to infer the skeleton topology. The above methods have achieved good results in indoor controlled lighting conditions or ground-based fixed camera scenarios, and have also been applied to some outdoor monitoring scenarios.
[0003] However, in the environment of high-altitude suspended platform operations, because the platform is suspended from the building facade, the camera experiences instantaneous acceleration due to wind loads as it moves along with the platform. The continuous jitter, coupled with extreme lighting conditions in the building exterior scene—including strong backlighting and large areas of shadow with a dynamic range exceeding 14 EV—caused the effective relative displacement between frames to exceed 30 pixels, far exceeding the linear assumption range of the optical flow algorithm, leading to frequent failures in cross-frame keypoint tracking. Simultaneously, the extreme lighting significantly reduced the distinguishability of work clothes and building components in the image feature space, making it impossible for the convolutional feature extraction module to differentiate between the foreground and background of the human body, thus causing a shift in the localization peak of the keypoint heatmap. The combined effect of these factors significantly reduced the keypoint extraction accuracy of existing methods.
[0004] In other words, existing technologies suffer from a severe decrease in the accuracy of extracting key points of personnel posture when strong backlighting and violent shaking coexist in high-altitude suspended platform operations. Summary of the Invention
[0005] In view of this, the present invention provides a method, medium and system for extracting key points of personnel posture in suspended baskets under complex high-altitude environments, which can solve the technical problem that the accuracy of personnel posture key point extraction is seriously reduced when strong backlight and violent shaking coexist in high-altitude suspended basket operation scenarios.
[0006] The present invention is implemented as follows: The first aspect of the present invention provides a method for extracting key points of personnel posture in a suspended platform under complex high-altitude environments, comprising the following steps:
[0007] Image sequences from the camera mounted on the suspended platform are acquired, and high-frequency data from the inertial measurement unit are acquired simultaneously. The inter-frame rotation matrix and translation vector are pre-integrated to generate inter-frame motion compensation parameters.
[0008] Based on the inter-frame motion compensation parameters, motion compensation affine transformation is applied to the image sequence frame by frame to compress the effective relative displacement between frames. Then, the hierarchical pyramid optical flow method is used to complete the cross-frame tracking of key points, and the random sampling consistency algorithm is used to eliminate outlier matching.
[0009] The motion-compensated image sequence is fused using a multi-exposure high dynamic range fusion strategy that aligns and fuses three frames (short exposure, medium exposure, and long exposure), and then compressed to 8 bits after camera response function calibration and local tone mapping.
[0010] By fusing the depth map from the time-of-flight sensor with the motion-compensated image, a joint red-green-blue depth segmentation network is constructed. A motion mask is generated by differentially superimposing three consecutive frames to remove static building components and retain the foreground region of moving human bodies.
[0011] The foreground region of the moving human body is input into the posture analysis model of the high-altitude hanging basket. The two-dimensional coordinates and occlusion confidence of 17 human body key points are extracted. When the number of occluded key points is greater than 0, the occlusion key point recovery algorithm based on compressed sensing and sparse signal reconstruction is started to reconstruct the coordinates of the occluded nodes.
[0012] Based on the confidence distribution of key points output by the aerial hoist attitude analysis model, the scene complexity assessment value is calculated. The dynamic edge weight update frequency and graph convolution iteration refinement number of the aerial hoist attitude analysis model are adjusted according to the scene complexity assessment value, and the final set of key point coordinates is output.
[0013] Specifically, the inertial measurement unit is a microelectromechanical system inertial device that integrates a gyroscope and a triaxial accelerometer. The sampling rate is set to a pre-integration sampling rate threshold so that the inter-frame integration error is at the sub-pixel level.
[0014] Specifically, the pre-integration involves accumulating the high-frequency angular velocity and acceleration measurements of the inertial measurement unit between the timestamps of two frames using median integration to obtain the inter-frame rotation matrix and translation vector.
[0015] Specifically, the motion-compensated affine transformation uses the rotation matrix and translation vector obtained from pre-integration as parameters to perform an inverse spatial transformation on the current frame image, compressing the effective relative displacement between frames to within the displacement compensation threshold.
[0016] Specifically, the layered pyramid optical flow method involves downsampling the image into multiple resolution levels, estimating the large displacement optical flow starting from the coarsest layer, and refining it layer by layer to the original resolution, with the number of layers set to a threshold.
[0017] In the multi-exposure high dynamic range fusion strategy, the camera response function calibration adopts the Debevec-Malik method, which captures multiple sets of image sequences with different exposure times for the same static scene, fits the nonlinear curve of pixel response, and maps the pixel values of the three frames of images to a linear radiosity map.
[0018] Specifically, the local tone mapping employs a bilateral filtering base layer separation method, decomposing the radiance map into a base layer and a detail layer, applying logarithmic linear compression to the base layer, and directly superimposing the detail layers.
[0019] Specifically, the red-green-blue depth joint segmentation network takes the red-green-blue three-channel image and the time-of-flight sensor depth map as input, extracts the joint features of space and appearance through a four-channel feature fusion encoder, and outputs a human foreground segmentation mask.
[0020] The high-altitude suspended basket attitude analysis model consists of three sub-modules connected in series: a visual feature extraction module, a dynamic skeleton graph construction module, and a physical constraint reasoning module. The visual feature extraction module uses an improved lightweight high-resolution network as its backbone and adds an adaptive receptive field branch based on deformable convolution. The sampling offset is generated by the acceleration and angular velocity output by the inertial measurement unit through a fully connected layer.
[0021] The dynamic skeleton graph construction module is based on a 17-node basic graph defined by human anatomy. The edge weights of the graph are jointly calculated by the peak intensity of the joint heatmap of each frame and the cosine similarity of the features of adjacent nodes, generating a dynamic weight matrix and updating it in real time. Message passing adopts anisotropic graph convolution.
[0022] The physical constraint inference module introduces a differentiable kinematic constraint layer at the end of the network, projects the predicted key point coordinates onto the skinned multi-person linear model space, calculates the bone length residual as an auxiliary loss term, and backpropagates the gradient to the entire network for end-to-end training.
[0023] Specifically, the occlusion key point recovery algorithm based on compressed sensing and sparse signal reconstruction uses the coordinates of observable key points to form a measurement vector. Under the constraint of the basis vector matrix obtained by principal component analysis decomposition of a large-scale attitude dataset, the sparse coefficients are solved by the L1 regularization minimization method. Then, the sparse coefficients are substituted into the complete dictionary matrix to reconstruct the coordinates of all key points. The solution adopts the alternating direction multiplier method.
[0024] Specifically, the scene complexity assessment value is obtained by taking the mean and variance of the peak confidence scores of the 17 key points in the heatmap output by the high-altitude suspended basket attitude analysis model, dividing the mean peak confidence score of the heatmap by the global mean benchmark, and dividing the variance by the global variance benchmark, and then taking a weighted sum to obtain the dimensionless scene complexity assessment value.
[0025] The training dataset of the high-altitude suspended basket attitude analysis model is composed of on-site collected samples and enhanced synthetic samples mixed according to a mixing ratio threshold. The enhanced synthetic samples use a multi-exposure high dynamic range fusion method to enhance the data of the outdoor building exterior wall scene and apply random affine perturbation to the image sequence.
[0026] Wherein, the pre-integration sampling rate threshold is 1000Hz; the displacement compensation threshold is 5 pixels; the layer threshold is 3 to 4 layers; the mixing ratio threshold is 6:4; the number of iterations of the alternating direction multiplier method is 50 times; and the scheduling segmentation threshold of the scene complexity evaluation value is 0.8 and 0.5.
[0027] A second aspect of the present invention provides a computer-readable storage medium storing program instructions, which, when executed in a computer, are used to perform the above-described method for extracting key points of personnel posture in a suspended basket under complex high-altitude conditions.
[0028] A third aspect of the present invention provides a system for extracting key points of personnel posture in a suspended basket under complex high-altitude environments, comprising the aforementioned computer-readable storage medium, wherein the system is a computer, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.
[0029] This invention generates inter-frame motion compensation parameters through pre-integration using an inertial measurement unit (IMU) and incorporates a motion compensation affine transformation and multi-exposure high dynamic range (HDR) fusion strategy into the keypoint extraction process. This allows the input image to simultaneously achieve jitter suppression and illumination equalization before being fed into subsequent networks, fundamentally eliminating the dual interference of jitter and backlighting on feature extraction. Furthermore, a foreground segmentation network jointly driven by a time-of-flight sensor depth map and a motion mask removes stationary building components from candidate foregrounds, avoiding feature confusion. A high-altitude gondola attitude analysis model uses deformable convolutional receptive fields driven by IMU data to perform directional compensation for motion-blurred regions. A dynamic skeleton graph construction module encodes joint confidence and feature similarity into the image structure in real time, enabling the information propagation path to adaptively adjust according to occlusion conditions. A physical constraint inference module embeds anatomical priors into the network in a differentiable form, ensuring geometric self-consistency in keypoint prediction. The synergistic effect of these components ensures high accuracy in keypoint extraction even under strong backlighting and severe jitter.
[0030] In summary, this invention solves the technical problem mentioned in the background art of severely reduced accuracy in extracting key points of personnel posture when strong backlighting and severe shaking coexist in high-altitude suspended platform operation scenarios. Attached Figure Description
[0031] Figure 1 This is a flowchart of the method of the present invention.
[0032] Figure 2 This is a comparison diagram of inter-frame displacement distribution before and after motion compensation.
[0033] Figure 3Box plots showing the error distribution of key point coordinates for different complexity evaluation value ranges in different scenarios. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below.
[0035] like Figure 1 The diagram shown is a flowchart of a method for extracting key points of personnel posture in a suspended platform under complex high-altitude environments, provided by the first aspect of this invention. This method includes the following steps:
[0036] S01. Acquire image sequences from the camera mounted on the suspended basket, and simultaneously acquire high-frequency data from the gyroscope and accelerometer of the inertial measurement unit. Perform pre-integration on the inter-frame rotation matrix and translation vector at a sampling rate of 1000Hz to generate inter-frame motion compensation parameters.
[0037] S02. Apply motion compensation affine transformation to the image sequence frame by frame according to the inter-frame motion compensation parameters to compress the effective relative displacement between frames to within 5 pixels. Then, use the 3-4 layered pyramid optical flow method to complete the cross-frame tracking of key points and use the random sampling consistency algorithm to eliminate outlier matching.
[0038] S03. The motion-compensated image sequence is fused using a multi-exposure high dynamic range fusion strategy with three frames aligned and fused (short exposure, medium exposure, and long exposure). After camera response function calibration and local tone mapping, the image is compressed to 8 bits and sent to subsequent processing.
[0039] S04. By fusing the depth map from the time-of-flight sensor with the motion-compensated image, a joint red-green-blue depth segmentation network is constructed. A motion mask is generated by differentially superimposing three consecutive frames to remove static building components and retain the foreground region of the moving human body.
[0040] S05. Input the foreground region of the moving human body into the posture analysis model of the high-altitude hanging basket, extract the two-dimensional coordinates and occlusion confidence of 17 human body key points, and when the number of occluded key points is greater than 0, start the occlusion key point recovery algorithm based on compressed sensing and sparse signal reconstruction to reconstruct the coordinates of the occluded nodes.
[0041] S06. Based on the confidence distribution of key points output by the aerial hoist attitude analysis model, calculate the scene complexity assessment value. When the scene complexity assessment value falls into different intervals, adjust the dynamic edge weight update frequency and the number of graph convolution iterations of the aerial hoist attitude analysis model, and output the final set of key point coordinates.
[0042] The inertial measurement unit is a microelectromechanical system inertial device that integrates a gyroscope and a triaxial accelerometer, with a sampling rate of 1000Hz, so that the accumulation of integral error within the frame interval (approximately 33ms) is at the sub-pixel level.
[0043] Pre-integration refers to the process of accumulating the high-frequency angular velocity and acceleration measurements taken by the inertial measurement unit between two image timestamps using median integration to obtain the inter-frame rotation matrix and translation vector, thereby acquiring the relative motion without relying on absolute pose estimation; this process can achieve instantaneous acceleration... The jittering motion is converted into geometric transformation parameters that can be applied to the image coordinates.
[0044] Among them, motion-compensated affine transformation refers to performing an inverse spatial transformation on the current frame image using the pre-integrated rotation matrix and translation vector as parameters, so that the relative motion of the camera between frames is canceled at the image level, thereby compressing the effective relative displacement between frames from more than 30 pixels to less than 5 pixels; the 5-pixel threshold is achieved by simulating wind speed. Repeated experiments were conducted on the suspended basket shaking platform to statistically analyze the inflection point range of the Lucas-Kanade optical flow tracking success rate curve under different compensation residuals. The upper limit of displacement corresponding to the change from a sharp drop to a gradual slope of the success rate curve was taken and determined through regression analysis of multiple rounds of experimental data.
[0045] Among them, the layered pyramid optical flow method refers to downsampling the image into multiple resolution levels, estimating the large displacement optical flow starting from the coarsest layer, and refining it layer by layer to the original resolution, thereby breaking through the limitation of the linear assumption of small displacement in single-layer optical flow; the number of layers is 3 to 4, which is determined by joint experiments on tracking accuracy and computation time under different jitter amplitudes.
[0046] Among them, the random sampling consensus algorithm refers to randomly selecting the smallest subset from the matched point pairs to fit the motion model, counting the number of interior points that conform to the model, and taking the model with the most interior points as the final result after several iterations, which is used to eliminate erroneous matching exterior points caused by the shaking of the suspended basket.
[0047] In the multi-exposure high dynamic range fusion strategy, the specific exposure time ratios of short, medium, and long exposures are determined by camera response function calibration experiments. The camera response function calibration adopts the Debevec-Malik method, which involves capturing at least 10 sets of image sequences with different exposure times for the same static scene, fitting the nonlinear curve of pixel response, and mapping the pixel values of the three frames of images to a linear radiosity map. The local tone mapping adopts the Durand bilateral filtering base layer separation method, which decomposes the radiosity map into a base layer and a detail layer. Logarithmic linear compression is applied to the base layer, and the detail layers are directly superimposed. Finally, scenes with a dynamic range exceeding 14 EV are compressed to 8 bits and fed into the subsequent network. The specific ratio of short, medium, and long exposure times is determined by simulating strong backlight and sidelight conditions in scenes of building exterior walls with different orientations. The ratio range corresponding to the highest signal-to-noise ratio in the human key point area of the fused image is taken.
[0048] Among them, the time-of-flight sensor is a sensor that actively emits modulated light pulses and receives reflected signals to calculate the depth of the scene. Its operating wavelength is usually in the range of 850 to 940 nm and is not affected by visible light illuminance, thus it is robust to bright backlight scenes.
[0049] Among them, the red-green-blue depth joint segmentation network refers to a convolutional neural network that takes red, green and blue three-channel images and time-of-flight sensor depth maps as input, extracts spatial and appearance joint features through a four-channel feature fusion encoder, and outputs a human foreground segmentation mask; the depth channel provides color-independent spatial position constraints to distinguish work clothes and building components with highly similar colors and textures.
[0050] Among them, the three consecutive frame differential superposition motion mask refers to subtracting the absolute values of three consecutive frames in time and superimposing them. The area where pixel changes continue to exist is marked as the moving foreground, and the area with stable pixel values is marked as the static background. Static building components have a response close to zero in the differential mask, so they are removed from the candidate foreground. The number of frames is 3, which is determined by the experimental trade-off between the mask missegmentation rate and the delay under different wind speed conditions.
[0051] The specific structure of the high-altitude suspended basket attitude analysis model is as follows: the high-altitude suspended basket attitude analysis model consists of three sub-modules connected in series: a visual feature extraction module, a dynamic skeleton graph construction module, and a physical constraint inference module. The visual feature extraction module uses an improved lightweight high-resolution network as its backbone. Based on standard multi-resolution parallel feature fusion, it adds an adaptive receptive field branch based on deformable convolution version 3. The sampling offset of this adaptive receptive field branch is generated by predicting the triaxial acceleration and triaxial angular velocity output from the inertial measurement unit through a fully connected layer, allowing the receptive field to adaptively extend in the jitter direction of the current frame, and to perform anisotropic feature acquisition on motion-blurred regions. The multi-resolution feature map is then fused through channel attention weighting to output a 256-dimensional semantic feature map. The dynamic skeleton graph construction module is based on a 17-node basic graph defined in human anatomy (i.e., COCO Keypoints). The dataset defines a skeletal topology. Each node embedding consists of three parts: two-dimensional coordinate components, a 256-dimensional semantic feature vector, and an occlusion confidence scalar. The edge weights of the graph are jointly calculated by the peak intensity of the joint heatmap in each frame and the cosine similarity of the features of adjacent nodes, generating a dynamic weight matrix that is updated in real time. Message passing uses anisotropic graph convolution to distinguish between two information flow directions: proximal to far-end and far-end to proximal, simulating the causal transmission relationship of the human motion chain. The inter-neuron weights of the anisotropic graph convolution are driven by a dynamic weight matrix. The number of feature transmission channels between layers is dynamically allocated between 128 and 512 based on the scene complexity evaluation value of the current frame. The CUDA stream schedules visible node subgraphs and occluded node subgraphs in parallel to different streams based on the occlusion confidence distribution of graph nodes to improve throughput. The memory allocation strategy dynamically adjusts the heatmap storage precision based on the proportion of occluded nodes in the batch. When the proportion of occluded nodes is less than 20%, 16-bit floating-point storage is used, and when it is more than 20%, it switches to 32-bit floating-point storage to ensure reconstruction accuracy. The physical constraint inference module introduces a differentiable kinematic constraint layer at the end of the network. Based on the prior that the length of human skeletal segments is approximately constant in the same video sequence, the predicted keypoint coordinates are projected onto the skinned multi-person linear model space. The skeletal length residual is calculated as an auxiliary loss term, and the gradient is backpropagated to the entire network for end-to-end training. An iterative refinement loop is started for occluded nodes. In each iteration, the updated features of visible nodes are propagated to the occluded nodes again. The iteration is terminated after a maximum of 3 iterations or when the change in the Euclidean distance between two adjacent output coordinates is lower than the convergence threshold. The convergence threshold is determined by experiments on a simulated dataset with an occlusion rate of 10% to 80% to statistically analyze the inflection point of the iterative stability curve. The network training incorporates a contrastive learning loss. The skeleton embeddings of adjacent frames of the same person attract each other, while the skeleton embeddings of different people repel each other, thereby improving the instance discrimination ability.The high-altitude suspended basket posture analysis model encodes joint confidence and feature similarity into the image structure in real time through dynamic edge weights, enabling skeleton inference to adaptively adjust the information propagation path according to scene lighting and occlusion. The deformable convolutional receptive field is driven by inertial measurement unit data, enabling visual feature extraction to be directionally compensated in the direction of jitter. The physical constraint layer embeds anatomical priors into the network in a differentiable form, ensuring that key point prediction maintains geometric self-consistency in human motion. The collaboration of the three modules enables the model to output a temporally stable and anatomically reasonable key point coordinate sequence even in high-altitude scenes with strong backlight, severe jitter, and severe occlusion.
[0052] The steps for establishing the training dataset for the high-altitude suspended platform attitude analysis model specifically include: collecting high-altitude suspended platform operation videos, labeling the two-dimensional coordinates and occlusion status of 17 human key points, and supplementing with time-of-flight sensor depth maps and synchronous inertial measurement unit data; addressing the scarcity of high-altitude backlight scenes, employing a multi-exposure high dynamic range fusion method to enhance the data of outdoor building facade scenes, simulating... Local oversaturation and Mixed lighting conditions with underexposure in shadow areas; to address basket jitter, random affine perturbations were applied to the image sequence, with the perturbation amplitude sampled based on the measured acceleration distribution of the inertial measurement unit; the dataset includes single-person and multi-person scenes, with occlusion rates covering 0% to 80%; the final dataset is composed of on-site collected samples and enhanced synthetic samples mixed in a 6:4 ratio, the ratio of which was determined by ablation experiments on the average accuracy index of key points on the validation set.
[0053] The training steps of the high-altitude suspended basket attitude analysis model specifically include: initializing the visual feature extraction module with pre-trained lightweight high-resolution network weights; initializing the dynamic skeleton graph construction module and the physical constraint inference module with uniformly distributed randomness; and using the AdamW optimizer with an initial learning rate of [value missing]. The cosine annealing strategy is used to decay the value to The total number of training rounds is determined by the early stopping strategy where the average accuracy of keypoints on the validation set no longer improves; the total loss is a weighted sum of three factors: the mean square error loss of the keypoint heatmap, the auxiliary loss of the bone length residual, and the contrastive learning loss, with the weight coefficients determined by grid search experiments on the validation set; during training, the CUDA flow allocation of the dynamic edge weight matrix is scheduled in real time according to the proportion of occluded nodes in the current batch, ensuring that the occluded node subgraph and the visible node subgraph are computed in parallel without causing flow synchronization blocking.
[0054] The occlusion keypoint recovery algorithm based on compressed sensing and sparse signal reconstruction works as follows: Human pose has sparse representation under the basis vector matrix obtained by principal component analysis decomposition of a large-scale pose dataset, meaning that most poses can be approximately reconstructed by a linear combination of only a few principal components. When some keypoints in a high-altitude scene are occluded, the known observable keypoint coordinates form the measurement vector, and the occluded keypoint coordinates are considered as missing components. The problem is transformed into solving an underdetermined linear equation system of sparse coefficient vectors under dictionary constraints, and the sparse coefficients are solved using the L1 regularization minimization method. The sparse coefficients are then substituted into the complete dictionary matrix to reconstruct the coordinates of all key points. The solution employs the alternating direction multiplier method, converging after approximately 50 iterations. The number of iterations is determined experimentally by statistically analyzing the convergence point of the reconstruction error reduction curve on test sets with different occlusion rates. This algorithm, without adding hardware sensors, can recover the coordinates of occluded nodes relying solely on observable nodes and an offline-built pose dictionary. It exhibits robust recovery capabilities against local occlusion caused by scaffolding, beams, and other components in high-altitude scenes, fundamentally reducing the problem of skipped or broken links in skeleton node trajectories and ensuring temporal continuity of the output pose. The key parameters of the algorithm are the L1 regularization coefficient and the constraint relaxation. The L1 regularization coefficient controls the sparsity of the sparse coefficient vector, while the constraint relaxation controls the tolerance range of the measurement fitting error. The optimal value range for both was determined experimentally using a grid search combined with cross-validation method on simulated test sets with different occlusion rates and noise levels.
[0055] The scene complexity assessment value is calculated as follows: The mean and variance of the peak confidence scores of the 17 key points in the heatmap output by the high-altitude suspended platform attitude analysis model are taken. The mean peak confidence score is divided by the global mean benchmark, and the variance is divided by the global variance benchmark. These are then weighted and summed to obtain the dimensionless scene complexity assessment value. The global mean benchmark and global variance benchmark are obtained by statistically analyzing the peak confidence score distribution of all samples in the training dataset. The weighting coefficients are determined by ablation experiments. The scheduling rule for the scene complexity assessment value is as follows: Let the scene complexity assessment value be... ,when At that time, the dynamic edge weight update frequency is once per frame, and the graph convolution iteration refinement number is 1; when At that time, the dynamic edge weight update frequency is once per frame, and the graph convolution iteration refinement number is 2; when At that time, the dynamic edge weight update frequency was set to once per frame, the graph convolution iteration refinement number was set to 3 times, the number of channels of the dynamic skeleton graph construction module was increased from 128 to 512, and the video memory storage precision was switched to 32-bit floating point; the above thresholds of 0.8 and 0.5 were determined by piecewise line fitting experiments on the validation set covering different lighting, occlusion rate and jitter amplitude scenes, by statistically analyzing the change of the average precision of key points with the scene complexity evaluation value.
[0056] The specific implementation of step S01 is as follows: While the camera is mounted on the suspended platform, an inertial measurement unit (IMU) integrating a gyroscope and a three-axis accelerometer is installed. The sampling rate of the IMU is set to 1000Hz. During the process of the camera acquiring image sequences at a frame rate of approximately 30Hz, the IMU continuously acquires approximately 33 high-frequency angular velocity and linear acceleration measurements between the timestamps of two adjacent frames. These measurements are gradually accumulated using median integration to obtain the inter-frame rotation matrix and translation vector, which is the pre-integration result. The core advantage of pre-integration is that it does not rely on absolute pose estimation; it only uses relative motion to describe the inter-frame camera motion, thereby enabling the instantaneous acceleration of the suspended platform to reach... The jitter motion is converted into geometric transformation parameters that can be applied to the image coordinates, providing input for subsequent affine transformations for motion compensation. Due to the sufficiently high sampling rate, the accumulation of inter-frame integration error is at the sub-pixel level, ensuring the accuracy of the compensation parameters.
[0057] The specific implementation of step S02 is as follows: Using the pre-integrated inter-frame rotation matrix and translation vector as parameters, an inverse affine transformation, i.e., motion-compensated affine transformation, is applied to the current frame image to cancel out the relative motion of the camera between frames at the image level. The effective relative displacement between frames is compressed from more than 30 pixels to less than 5 pixels. After completing motion compensation, a 3-4 layer hierarchical pyramid optical flow method is used to track key points across frames. The hierarchical pyramid optical flow method downsamples the image into multiple resolution levels, estimating the optical flow of larger displacements starting from the top layer with the lowest resolution, and refining it layer by layer to the original resolution, breaking through the linear assumption limitation of the single-layer Lucas-Kanade optical flow method, which is only applicable to small displacement scenes. After tracking is completed, the random sampling consensus algorithm is used to remove outliers from matching point pairs. This algorithm randomly selects the smallest subset from the matching point pairs to fit the motion model, counts the number of interior points that conform to the model, and after multiple iterations, takes the model with the most interior points as the final result, effectively removing erroneous matches caused by basket jitter.
[0058] The specific implementation of step S03 is as follows: For the motion-compensated image sequence, three frames of images—short exposure, medium exposure, and long exposure—are simultaneously acquired at each moment. The specific exposure time ratio of the three frames is determined by camera response function calibration experiments. Camera response function calibration uses the Debevec-Malik method, capturing at least 10 sets of image sequences with different exposure times for the same static scene, fitting a nonlinear curve of pixel response, and mapping the pixel values of the three frames to a linear radiance map. The linear radiance map preserves the true irradiance information of the scene and can cover scenes with both bright backlighting and deep shadows with a dynamic range exceeding 14 EV. Subsequently, the Durand bilateral filtering base layer separation method is used for local tone mapping, decomposing the radiance map into a base layer and a detail layer. Logarithmic domain linear compression is applied to the base layer to reduce the overall dynamic range, while the detail layers are directly superimposed to preserve local texture details. Finally, the result is compressed to 8 bits and sent to subsequent processing.
[0059] The specific implementation of step S04 is as follows: The depth map acquired by the time-of-flight sensor is pixel-level aligned with the motion-compensated image to construct a joint red-green-blue depth segmentation network. The time-of-flight sensor actively emits modulated light pulses, with a working wavelength typically in the range of 850–940 nm. It is unaffected by visible light illuminance, robust to strong backlighting scenes, and can provide scene depth information independent of color. The joint red-green-blue depth segmentation network takes four channels—red, green, and blue three-channel images and the depth map—as input. A four-channel feature fusion encoder extracts joint spatial and appearance features. The depth channel provides spatial position constraints to distinguish between work clothes and building components with highly similar colors and textures, outputting a human foreground segmentation mask. Based on this, the absolute values of the pairwise subtractions of three consecutive temporally connected images are superimposed to generate a motion mask. Regions with continuous pixel changes are marked as moving foregrounds, while regions with stable pixel values are marked as static backgrounds, thereby eliminating static building components and retaining the moving human foreground region.
[0060] The specific implementation of step S05 is as follows: The foreground region of the moving human body is fed into the posture analysis model of the high-altitude hanging basket. The visual feature extraction module uses an improved lightweight high-resolution network as its backbone, and adds an adaptive receptive field branch based on the third version of deformable convolution. The sampling offset is generated by the triaxial acceleration and triaxial angular velocity output by the inertial measurement unit through a fully connected layer, so that the receptive field extends adaptively in the jitter direction and anisotropic features are collected in the motion-blurred region. The multi-resolution feature map is fused by channel attention and outputs a 256-dimensional semantic feature map. The dynamic skeleton graph construction module is based on a 17-node basic graph defined by human anatomy. The edge weights are jointly calculated by the peak intensity of the joint heatmap and the cosine similarity of the features of adjacent nodes to generate a dynamic weight matrix. Message passing adopts anisotropic graph convolution to distinguish between two information flow directions: proximal to far end and far end to proximal. The physical constraint inference module introduces a differentiable kinematic constraint layer, projects the predicted key point coordinates onto the skinned multi-person linear model space, and calculates the bone length residual as an auxiliary loss term for end-to-end training. When the number of occluded keypoints is greater than 0, the occlusion keypoint recovery algorithm based on compressed sensing and sparse signal reconstruction is initiated: the measurement vector is constructed using the coordinates of the observable keypoints, and under the constraint of the basis vector matrix obtained by principal component analysis decomposition of a large-scale pose dataset, the sparse coefficient vector is solved using the L1 regularization minimization method. Then, the sparse coefficients are substituted into the complete dictionary matrix to reconstruct the coordinates of all keypoints. The solution process adopts the alternating direction multiplier method, and converges after about 50 iterations to recover the coordinates of the occluded nodes.
[0061] The specific implementation of step S06 is as follows: Take the mean and variance of the peak confidence scores of the 17 key point heatmaps output by the high-altitude suspended basket attitude analysis model, divide them respectively by the global mean benchmark and global variance benchmark obtained statistically from the training dataset, and then sum them according to the weighted coefficients to obtain the dimensionless scene complexity evaluation value. Let the scene complexity evaluation value be... ,when When, the graph convolution iteration refinement is performed once; when When, the graph convolution iteration refinement is performed twice; when At that time, the graph convolution iteration refinement was performed 3 times, while the number of channels in the dynamic skeleton graph construction module was increased from 128 to 512, and the video memory storage precision was switched to 32-bit floating point. The dynamic edge weights were updated once per frame in all levels. The above scheduling rules concentrate computing resources on complex scenes, improving the key point localization accuracy of difficult samples while ensuring real-time performance, and outputting the final set of key point coordinates.
[0062] It should be noted that the key technologies of this invention include: an inertial measurement unit-driven motion compensation mechanism, which transforms physical jitter into a reversible geometric transformation through high-frequency pre-integration, ensuring that the input of optical flow tracing always satisfies the linear displacement assumption, thereby restoring tracking accuracy without relying on additional hardware. Compared to pure vision methods that rely solely on the image itself to estimate motion, this method offers higher compensation accuracy and greater robustness to large jitter. A collaborative preprocessing mechanism combining multi-exposure high dynamic range fusion and time-of-flight sensor depth guidance is also included. The former recovers effective texture under extreme lighting conditions from the perspective of temporal multi-frame fusion, while the latter provides illumination-independent spatial information from the perspective of active perception. The two technologies complement each other at the feature level, enabling the segmentation and feature extraction networks to distinguish between human bodies and building components even in strong backlighting scenarios. The deformable convolutional receptive field is driven by data from the inertial measurement unit and coordinated in real time by the dynamic skeleton graph construction module. The former performs directional compensation for motion blur in the feature extraction stage, while the latter encodes the joint confidence of the current frame into the graph structure in real time in the skeleton inference stage, so that the information propagation path is dynamically adjusted according to occlusion and lighting conditions. The synergy of the three key technologies enables the entire process from original image acquisition to final key point output to have the ability to adapt to jitter and backlighting, forming a complete robust pose extraction closed loop.
[0063] It's important to note that in high-altitude suspended platform operations, when workers experience significant limb occlusion within the platform—such as bending over, squatting, or being partially obscured by the platform frame—some key points remain invisible for multiple consecutive frames. This leads to technical issues like trajectories of skeleton nodes jumping and chain breaks. The reason for this is that existing heatmap regression models show near-zero peak confidence when key points are occluded. The network lacks effective constraints to estimate the coordinates of invisible nodes, resulting in significant jumps in output coordinates between occluded and unoccluded frames. While graph convolutional networks can infer occluded nodes using neighbor node information, the confidence of neighbor nodes themselves decreases with high occlusion rates, leading to accumulated inference errors and ultimately causing chain breaks in the skeleton sequence, affecting the coherence of subsequent behavior analysis. A common solution to this problem is to add temporal smoothing filters, such as Kalman filtering or exponential moving average, to interpolate and estimate the coordinates of occluded frames. However, the core assumption of temporal smoothing filtering is that the occlusion duration is short. When the occlusion lasts for multiple frames, the coordinates estimated by the filter will systematically deviate from the true coordinates. Furthermore, the filtering parameters lack adaptability to workers with different movement speeds, and the estimation error is further amplified in scenarios where the shaking of the suspended platform and human movement are superimposed. This invention effectively solves this technical problem. Specifically, it uses an occlusion keypoint recovery algorithm based on compressed sensing and sparse signal reconstruction. It leverages the prior knowledge that human pose has sparse representation under the basis vector matrix obtained by principal component analysis decomposition of a large-scale pose dataset. The observable keypoint coordinates are used to construct the measurement vector, and the occluded keypoint coordinates are treated as missing components. This is transformed into a sparse solution problem of an underdetermined linear equation system. The sparse coefficients are solved using the L1 regularization minimization method through the alternating direction multiplier method, and then substituted into the complete dictionary matrix to reconstruct the coordinates of all nodes. This method does not rely on the assumption of temporal continuity. It can recover the coordinates of occluded nodes by relying only on the observable nodes of the current frame and the offline constructed pose dictionary. It is also effective for depth occlusion that lasts for multiple frames. The skeleton length residual auxiliary loss of the physical constraint inference module further ensures the geometric rationality of the recovered coordinates, so that the skeleton sequence maintains temporal continuity before and after occlusion, thereby eliminating trajectory jumps and chain breaks.
[0064] A second aspect of the present invention provides a computer-readable storage medium storing program instructions, which, when executed in a computer, are used to perform the above-described method for extracting key points of personnel posture in a suspended basket under complex high-altitude conditions.
[0065] A third aspect of the present invention provides a system for extracting key points of personnel posture in a suspended basket under complex high-altitude environments, comprising the aforementioned computer-readable storage medium. The system can be any one of a computer, a server, or a microcontroller. The computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.
[0066] Specifically, the principle of this invention is:
[0067] The reason why this invention can solve the above-mentioned technical problems lies in the following technical logic.
[0068] First, the jitter of the suspended platform causes the relative displacement between frames to far exceed the linear assumption range of the optical flow algorithm, which is the direct cause of keypoint tracking failure. This invention, before sending the image to the tracking module, uses an inertial measurement unit to sample angular velocity and acceleration at a high frequency of 1000Hz. Through median integral pre-integration, the inter-frame rotation matrix and translation vector are obtained. Then, these motion parameters are used to apply an inverse affine transformation to the image, compressing the effective relative displacement between frames to within 5 pixels. This operation removes the physical rigid body motion from the image coordinate space, ensuring that the input for subsequent optical flow tracking satisfies the linear assumption, thereby restoring the effectiveness of tracking. The high sampling rate of the inertial measurement unit ensures that the inter-frame integration error is at the sub-pixel level, bringing the compensation accuracy within the acceptable range for the optical flow algorithm.
[0069] Secondly, the dynamic range of images in strong backlight scenes exceeds the linear response range of the camera's 8-bit sensor, which is the root cause of feature extraction failure. This invention employs a multi-exposure high dynamic range strategy that fused short, medium, and long exposure frames. After camera response function calibration, pixel values are mapped to a linear radiosity map. Then, local tone mapping separated by a bilateral filtering base layer compresses the dynamic range exceeding 14 EV to 8 bits, restoring effective texture details in both oversaturated and underexposed areas, thereby providing balanced input for subsequent networks.
[0070] Furthermore, work clothes and building components are highly similar in color and texture features, making it difficult to distinguish foreground from background using only visible light images. This invention introduces a time-of-flight sensor depth map, whose operating wavelength is unaffected by visible light illuminance, and can stably provide scene depth information even under strong backlight. The red-green-blue depth joint segmentation network uses a four-channel fusion encoder to extract joint spatial and appearance features. The depth channel provides color-independent spatial position constraints, ensuring that foreground segmentation remains distinguishable even in scenes with similar appearances. A three-frame differentially superimposed motion mask further excludes static building components from the candidate foreground, ensuring that the region fed into the pose analysis model is a real human foreground.
[0071] Finally, occlusion and jitter in high-altitude scenes cause a decrease in the peak confidence of some key point heatmaps, making it impossible to recover the coordinates of occluded nodes in single-frame inference. The high-altitude gondola posture analysis model of this invention achieves directional feature compensation through deformable convolutional receptive fields driven by inertial measurement unit data. The dynamic skeleton graph construction module encodes joint confidence and feature similarity into the edge weights of the graph in real time, driving anisotropic graph convolution to propagate information along the human motion chain, enabling visible node features to effectively infer the coordinates of occluded nodes. The physical constraint inference module embeds the anatomical prior of constant bone segment length into end-to-end training in the form of a differentiable bone length residual loss, ensuring the geometric self-consistency of the predicted coordinates. The scene complexity evaluation value dynamically schedules the number of graph convolution iterations and channels, concentrating computational resources on difficult scenes to further improve accuracy. These modules eliminate interference sources sequentially from four levels: physical compensation, illumination equalization, foreground segmentation, and skeleton inference, forming a complete closed-loop solution.
[0072] The following provides a specific embodiment 1 of the present invention, and the specific implementation of each step in this embodiment 1 is described in detail below.
[0073] The specific implementation of step S01 is as follows: A camera mounted on a suspended platform acquires image sequences at a frame rate of approximately 30Hz, while a microelectromechanical system (MEMS) inertial measurement unit integrating a gyroscope and a three-axis accelerometer simultaneously acquires angular velocity and acceleration data at a sampling rate of 1000Hz. Timestamps are generated for two consecutive frames of images. and The inertial measurement unit measurements between frames are pre-integrated, and the angular velocities are accumulated using the median integration method to obtain the inter-frame rotation matrix. Accumulated acceleration yields the inter-frame translation vector. The pre-integral formula is expressed as follows:
[0074] ;
[0075] ;
[0076] In the formula, For the first The triaxial angular velocity vector at each sampling time, in units of For the first The triaxial acceleration vector at each sampling time, in units of The sampling interval between adjacent inertial measurement units is taken as... s This refers to the total number of inter-frame sampling points, i.e., the number of inertial measurement unit sampling points between the timestamps of two image frames. This is the gravitational acceleration vector, in units of . For the first Sampling time relative to the first The velocity integral at each frame time, in units of It is obtained by iteratively deriving from the prior acceleration integral. For from the first Frame time to the The rotation matrix at the sampling time is obtained by iteratively deriving the integral of the preceding angular velocity. Using the Lie group exponential mapping, the 3D rotation vector is mapped to the rotation matrix space. This outputs the inter-frame motion compensation parameters. .
[0077] The specific implementation of step S02 is as follows: using the pre-integrated result and Constructing the affine transformation matrix The inverse spatial transformation is applied to the current frame image, compressing the effective relative displacement between frames from more than 30 pixels to less than 5 pixels. After compensation, the image uses a 3-4 layer hierarchical pyramid optical flow method to complete cross-frame tracking of key points, estimating large displacements from the coarsest resolution layer and refining layer by layer. The random sampling consensus algorithm randomly selects the smallest subset from the matched point pairs to fit the homography matrix, counts the number of interior points and iterates, and selects the model with the most interior points to eliminate outliers.
[0078] The specific implementation of step S03 is as follows: A three-frame alignment and fusion strategy of short exposure, medium exposure, and long exposure is adopted, and the exposure time ratio of the three frames is determined by camera response function calibration experiments. The camera response function calibration uses the Derbervik-Malik method, capturing at least 10 sets of image sequences with different exposure times for the same static scene, fitting pixel response nonlinear curves, and mapping the pixel values of the three frames to a linear radiosity map. The unit is The radiance map was decomposed into basic layers using the Durant bilateral filtering method. With detail layer Logarithmic linear compression is applied to the base layer, and the detail layers are directly stacked, ultimately compressing scenes with a dynamic range exceeding 14 EV to 8-bit output.
[0079] The specific implementation of step S04 is as follows: The time-of-flight sensor depth map... Compared with motion-compensated red-green-blue images The input is concatenated into a four-channel array and fed into a red-green-blue depth joint segmentation network. A four-channel feature fusion encoder extracts joint spatial and appearance features and outputs a human foreground segmentation mask. A motion mask is generated by superimposing the pairwise differences of three consecutive temporally linked frames. The calculation method is described as follows:
[0080] ;
[0081] In the formula, , , These are the image pixel value matrices for the current frame, the previous frame, and the two previous frames, respectively. As the reference for pixel value normalization, we set it to 255, so that... It is a dimensionless matrix; regions with difference responses approaching zero are identified as static backgrounds and removed, while foreground regions of moving human figures are retained.
[0082] The specific implementation of step S05 is as follows: The foreground region of the moving human body is fed into the posture analysis model of the high-altitude suspended basket, and the two-dimensional coordinates and occlusion confidence of 17 key points are output. When the number of occluded key points is greater than 0, the occlusion key point recovery algorithm based on compressed sensing and sparse signal reconstruction is activated. The algorithm principle is that the human posture is derived from the basis vector matrix obtained by principal component analysis decomposition of a large-scale dataset. The following has sparse representation, where The total number of key points is set to 17. Let the number of basis vectors be and the complete keypoint coordinate vector be . The sparse coefficient vector is Then we have:
[0083] ;
[0084] Let the observation matrix corresponding to the observable key points be... ,in The number of observable keypoints is given by the observation vector. The problem of restoring key points that are obscured can be transformed into the following: Regularization minimization problem:
[0085] ;
[0086] In the formula, For optimal sparse coefficient estimation To observe the standard deviation of noise, the units are... Consistent, in pixels, used for normalizing data fitting terms. for Regularization coefficient, dimensions and The two are the same, and their ratio is a dimensionless regulating factor. The regularization coefficient is a reference value, determined by grid search combined with cross-validation experiments on a simulation test set with different occlusion rates and noise levels. for norm square Norm. The alternating direction multiplier method is used to solve this problem, and it converges after approximately 50 iterations. Substitution Reconstruct the coordinates of all key points. This is the reconstructed complete keypoint coordinate vector.
[0087] The specific implementation of step S06 is as follows: Based on the peak confidence of the heatmap of 17 key points output by the attitude analysis model of the high-altitude suspended basket, calculate the scene complexity evaluation value. Let the first... The confidence level of the peak value of the heatmap for each key point is: , The global mean benchmark is The global variance benchmark is The calculation formula is expressed as follows:
[0088] ;
[0089] In the formula, and Let be the weighting coefficient, satisfying Determined by ablation experiments The global mean of the peak confidence scores of heatmaps for all samples in the training dataset. The corresponding global variances are both calculated statistically from the training dataset and used for normalization. This is a dimensionless evaluation value. When... When, the graph convolution iteration refinement is performed once; when When, the graph convolution iteration refinement is performed twice; when At that time, the number of graph convolution iterations for refinement is 3, the number of channels is increased from 128 to 512, the video memory storage precision is switched to 32-bit floating point, and the final set of key point coordinates is output.
[0090] The specific implementation of the high-altitude suspended basket attitude analysis model is as follows: the model consists of three sub-modules connected in series: a visual feature extraction module, a dynamic skeleton graph construction module, and a physical constraint inference module. The visual feature extraction module uses an improved lightweight high-resolution network as its backbone, and adds an adaptive receptive field branch based on deformable convolution version 3, building upon multi-resolution parallel feature fusion. Layer feature map at location The sampling offset at that point is The triaxial acceleration output by the inertial measurement unit With triaxial angular velocity After splicing, prediction is performed through a fully connected layer, as expressed in the following formula:
[0091] ;
[0092] In the formula, For the first The fully connected layer mapping function corresponding to the layer takes a six-dimensional feature vector of the inertial measurement unit as input and outputs a two-dimensional offset corresponding to the spatial location of the feature map, in pixels. For vector concatenation operations The unit is , The unit is The two images are concatenated and then linearly transformed by a fully connected layer to output the pixel-level offset. The weights of the fully connected layer learn the dimensional transformation relationship during training. The multi-resolution feature maps are then fused using channel attention weighting to output a 256-dimensional semantic feature map. , and This represents the feature map space size, in pixels.
[0093] The dynamic skeleton diagram construction module is based on a 17-node basic diagram defined by human anatomy. The embedding vector of each node is ,in These are the two-dimensional image coordinate components of the key point, in pixels. This is the semantic feature vector corresponding to this node. The occlusion confidence scalar has a range of values. .node With nodes Edge weights between The calculation formula is expressed as follows:
[0094] ;
[0095] In the formula, For nodes Peak intensity of heat map The ratio of the peak intensity of the heatmaps of all nodes in the training dataset to the global mean is a dimensionless normalization term. For nodes With nodes dot product of eigenvectors and For the corresponding feature vector norm The cosine similarity normalization benchmark is set to 1, making the second term a dimensionless quantity. The weighting adjustment coefficient satisfies Determined by validation set experiments, making The overall edge weights are dimensionless. Anisotropic graph convolution distinguishes the near-to-far direction. From distal to proximal direction Two types of information flow, the first Layer nodes The update formula is expressed as follows:
[0096] ;
[0097] In the formula, and The first Learnable weight matrices in the proximal-to-distal and distal-to-proximal directions of the layer and They are nodes The set of neighbor nodes in both directions For the first Layer bias vector Activation function For the first Layer nodes Embedded vector For the first Layer nodes The updated embedding vector.
[0098] The physics constraint inference module projects the predicted keypoint coordinates onto the skinned multi-person linear model space and calculates the bone length residual as an auxiliary loss term. Let the... The predicted coordinates of the key points at both ends of the segment bone are as follows: and The unit is pixels, the first The reference length of a segmental bone is Units are pixels, bone length residual auxiliary loss. The statement is as follows:
[0099] ;
[0100] In the formula, Total number of skeletal segments To normalize the denominator, make Dimensionless loss value The total loss is obtained by statistically averaging the predicted coordinates of the skeleton endpoints of visible frames in the same video sequence. Loss due to mean square error of key point heatmap , bone length residual auxiliary loss Comparative learning loss The weighted sum of the three terms is expressed as follows:
[0101] ;
[0102] In the formula, , , These are weighting coefficients, all of which are dimensionless scalars, determined by grid search experiments on the validation set. The mean squared error loss of the key point heatmap is a dimensionless value. To compare the learning loss, we use a dimensionless value where the skeleton embeddings of adjacent frames of the same person attract each other, while the skeleton embeddings of different people repel each other, thus improving the ability to distinguish instances.
[0103] To better understand and implement this invention, the following is a specific application scenario of the invention, Example 2: To illustrate the effects of the invention, technicians constructed a test environment for extracting key points of the high-altitude suspended platform's posture. Using the exterior wall cleaning operation of a high-rise building as the test scenario, the suspended platform was suspended at a height of approximately 40m, and the measured wind speed range was... The suspended platform experienced continuous shaking under wind load, with the average measured effective relative displacement between camera frames being approximately 34 pixels. The test period covered the morning period with strong backlighting, and the illuminance in some areas of the building's exterior walls exceeded [a certain level]. The illuminance in the shaded area was approximately 25 lux, and the scene dynamic range was approximately 13.8 EV. There were two workers on site, both wearing gray work clothes that matched the color of the exterior walls, with an occlusion coverage of 20%–65%.
[0104] The test hardware configuration includes: an industrial camera with a frame rate of 30Hz, a microelectromechanical system (MEMS) inertial measurement unit (INS) with a sampling rate of 1000Hz, a time-of-flight sensor with a working wavelength of 940nm, and an edge computing unit equipped with a graphics processor. The INS and camera are synchronized with timestamps via hardware trigger signals, while the time-of-flight sensor and camera are aligned at the frame level via external triggers.
[0105] In step S01, the inertial measurement unit continuously acquires triaxial angular velocity and triaxial acceleration at 1000Hz, and performs pre-integration within approximately 33ms between the timestamps of two adjacent frames to obtain the inter-frame rotation matrix and translation vector. The pre-integration process uses median integration, and the inter-frame integration error is approximately 0.3 pixels, which is in the sub-pixel range. In step S02, a motion-compensated affine transformation is applied to each frame using the pre-integration parameters. After compensation, the average effective relative displacement between frames is reduced to approximately 3.8 pixels, meeting the threshold requirement of less than 5 pixels. Figure 2 The image shows a comparison of inter-frame displacement distribution before and after motion compensation. Subsequently, a four-layer pyramid optical flow method was used for cross-frame tracking of key points, and outliers were eliminated using a random sampling consistency algorithm, with an outlier elimination rate of approximately 18%.
[0106] In step S03, for strong backlight scenes, the exposure times for short, medium, and long exposures were set to 0.5ms, 4ms, and 16ms, respectively, with a ratio of approximately 1:8:32. The camera response function was calibrated using the Debevec-Malik method, fitting 12 different exposure sequences captured from an exterior wall scene. After multi-exposure high dynamic range fusion, local tone mapping employed the Durand bilateral filter base layer separation method, compressing the 13.8EV scene to 8 bits, resulting in a significant improvement in the signal-to-noise ratio of the human keypoint region compared to a single-frame image. The pixel response consistency of the fusion results is shown in Table 1.
[0107] Table 1. Comparison of signal-to-noise ratio of key human body regions before and after multi-exposure fusion.
[0108]
[0109] In step S04, the time-of-flight sensor provides a 640×480 resolution depth map, which is then pixel-aligned with the image and input into the red-green-blue depth joint segmentation network. The network uses a four-channel feature fusion encoder to extract joint spatial and appearance features, outputting a human foreground segmentation mask. A motion mask is generated by differentially stacking three consecutive frames of images. Static building components have response values close to zero in the differential mask and are therefore removed from the candidate foreground.
[0110] In step S05, the foreground region of the moving human body is fed into the attitude analysis model of the high-altitude suspended basket. The adaptive receptive field branch of the visual feature extraction module, based on the real-time acceleration and angular velocity output by the inertial measurement unit, predicts the sampling offset of the deformable convolution through a fully connected layer, extending the receptive field in the jitter direction and collecting anisotropic features from the motion-blurred region. The dynamic skeleton graph construction module updates the edge weights of the 17-node base graph in real-time per frame. During the test period, the number of occluded keypoints reached 6 in some frames, corresponding to an occlusion rate of approximately 35%. At this point, the occlusion keypoint recovery algorithm based on compressed sensing and sparse signal reconstruction is activated, using the alternating direction multiplier method iteratively 50 times to solve for the sparse coefficients and reconstruct the coordinates of the occluded nodes. The average accuracy of keypoints at various occlusion rates is shown in Table 2.
[0111] Table 2. Statistical table of average accuracy of key points under different occlusion rates.
[0112]
[0113] In step S06, during periods of strong backlighting, the scene complexity evaluation value is... The mean is approximately 0.42, falling into In the specified interval, the number of graph convolution iterations for refinement is automatically adjusted to 3, the number of channels in the dynamic skeleton graph construction module is increased to 512, and the video memory storage precision is switched to 32-bit floating-point. During periods of cloud cover, The mean is approximately 0.74, falling into The interval and the number of iterations are adjusted to 2. For example... Figure 3 The figure shows a box plot of the key point coordinate error distribution corresponding to different complexity evaluation value ranges. It can be seen that the adaptive scheduling mechanism effectively controls the coordinate error in each complexity range.
[0114] Compared with traditional methods, this invention achieves the following technical advancements: Traditional methods rely on pure visual optical flow tracking, which fails under significant camera shake. In contrast, this invention uses inertial measurement unit pre-integration to transform physical motion into reversible geometric compensation, ensuring that tracking always operates within an effective range. Traditional single-exposure images cannot simultaneously preserve texture details in both highlight and shadow areas in extreme dynamic range scenarios. This invention, through multi-exposure high dynamic range fusion, physically expands the effective perception range. Traditional pure visual segmentation networks struggle to distinguish foregrounds when work clothes and background colors are similar. This invention introduces a time-of-flight sensor depth map as an independent spatial constraint channel, eliminating the need for color features in segmentation. Traditional occlusion recovery methods rely on the assumption of temporal continuity, leading to accumulated estimation errors over multiple frames of occlusion. In contrast, this invention's compressed sensing sparse reconstruction method relies solely on observable nodes and the pose dictionary for the current frame, making it equally effective for depth occlusion. These technological advancements stem from targeted modeling of the physical causes of various interference factors, rather than simply stacking network structures.
[0115] It should be noted that the variables involved in this invention are explained in detail in Tables 3 and 4.
[0116] Table 3. Variable Explanation Table (Part 1)
[0117]
[0118] Table 4. Variable Explanation Table (Part Two)
[0119]
[0120] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for extracting key points of personnel posture in a suspended platform under complex high-altitude environments, characterized in that, Includes the following steps: Image sequences from the camera mounted on the suspended platform are acquired, and high-frequency data from the inertial measurement unit are acquired simultaneously. The inter-frame rotation matrix and translation vector are pre-integrated to generate inter-frame motion compensation parameters. Based on the inter-frame motion compensation parameters, motion compensation affine transformation is applied to the image sequence frame by frame to compress the effective relative displacement between frames. Then, the hierarchical pyramid optical flow method is used to complete the cross-frame tracking of key points, and the random sampling consistency algorithm is used to eliminate outlier matching. The motion-compensated image sequence is fused using a multi-exposure high dynamic range fusion strategy that aligns and fuses three frames (short exposure, medium exposure, and long exposure), and then compressed to 8 bits after camera response function calibration and local tone mapping. By fusing the depth map from the time-of-flight sensor with the motion-compensated image, a joint red-green-blue depth segmentation network is constructed. A motion mask is generated by differentially superimposing three consecutive frames to remove static building components and retain the foreground region of moving human bodies. The foreground region of the moving human body is input into the posture analysis model of the high-altitude hanging basket. The two-dimensional coordinates and occlusion confidence of 17 human body key points are extracted. When the number of occluded key points is greater than 0, the occlusion key point recovery algorithm based on compressed sensing and sparse signal reconstruction is started to reconstruct the coordinates of the occluded nodes. Based on the confidence distribution of key points output by the aerial hoist attitude analysis model, the scene complexity assessment value is calculated. The dynamic edge weight update frequency and graph convolution iteration refinement number of the aerial hoist attitude analysis model are adjusted according to the scene complexity assessment value, and the final set of key point coordinates is output.
2. The method for extracting key points of personnel posture in a suspended platform under complex high-altitude environments according to claim 1, characterized in that, The inertial measurement unit is specifically a microelectromechanical system inertial device that integrates a gyroscope and a triaxial accelerometer. The sampling rate is set to a pre-integration sampling rate threshold so that the inter-frame integration error is at the sub-pixel level.
3. The method for extracting key points of personnel posture in a suspended platform under complex high-altitude environments according to claim 2, characterized in that, The pre-integration specifically involves accumulating the high-frequency angular velocity and acceleration measurements of the inertial measurement unit between the timestamps of two frames using median integration to obtain the inter-frame rotation matrix and translation vector.
4. The method for extracting key points of personnel posture in a suspended platform under complex high-altitude environments according to claim 3, characterized in that, The motion-compensated affine transformation specifically uses the rotation matrix and translation vector obtained from pre-integration as parameters to perform an inverse spatial transformation on the current frame image, compressing the effective relative displacement between frames to within the displacement compensation threshold.
5. The method for extracting key points of personnel posture in a suspended platform under complex high-altitude environments according to claim 4, characterized in that, The layered pyramid optical flow method specifically involves downsampling the image into multiple resolution levels, estimating the large displacement optical flow starting from the coarsest layer, and refining it layer by layer to the original resolution, with the number of layers set to a threshold.
6. The method for extracting key points of personnel posture in a suspended platform under complex high-altitude environments according to claim 5, characterized in that, In the multi-exposure high dynamic range fusion strategy, the camera response function calibration adopts the Debevec-Malik method, which captures multiple sets of image sequences with different exposure times for the same static scene, fits the nonlinear curve of pixel response, and maps the pixel values of the three frames of images to a linear radiosity map.
7. The method for extracting key points of personnel posture in a suspended platform under complex high-altitude environments according to claim 6, characterized in that, The local tone mapping specifically employs a bilateral filtering base layer separation method, which decomposes the radiance map into a base layer and a detail layer. Logarithmic linear compression is applied to the base layer, and the detail layers are directly superimposed.
8. The method for extracting key points of personnel posture in a suspended platform under complex high-altitude environments according to claim 7, characterized in that, The red-green-blue depth joint segmentation network specifically uses red-green-blue three-channel images and time-of-flight sensor depth maps as inputs, extracts spatial and appearance joint features through a four-channel feature fusion encoder, and outputs a human foreground segmentation mask.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, which, when executed in a computer, are used to perform the method for extracting key points of personnel posture in a suspended basket under complex high-altitude environments as described in any one of claims 1-8.
10. A system for extracting key points of personnel posture in a suspended platform under complex high-altitude environments, characterized in that, The system comprises the computer-readable storage medium of claim 9, wherein the system is a computer, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes program instructions stored in the computer-readable storage medium.