A matchstick figure posture intelligent monitoring method based on orthogonal amplification

CN122598077APending Publication Date: 2026-08-18NANCHANG WOMENS HOME CARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611003466.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

这些方法能增加样本数量,但对于跌倒和床边滑落等动作,随意扰动会破坏骨长比例、支撑接触点和动作阶段顺序

Benefits of technology

1、全程以去除人脸、服饰纹理的火柴人骨架图像作为输入与推理载体,无需存储、传输清晰人体画面,从数据源规避养老监护场景隐私泄露风险;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122598077A_ABST
    Figure CN122598077A_ABST
Patent Text Reader

Abstract

The application discloses a matchstick person posture intelligent monitoring method based on orthogonal amplification, collects a monitoring scene video, extracts human key points, generates a non-private matchstick person skeleton sample, extracts four types of features of a skeleton structure, motion, support contact and data distribution, adaptively matches a plurality of Hadamard, DCT and Wavelet orthogonal sampling matrices for different skeleton states, realizes sample amplification through point-by-point multiplication of the matrices, and adds a skeleton semantic check to remove invalid samples; and a BiGRU-Attention, TCN or Transformer time sequence model is used to complete fall and long-time lying risk identification. The application solves the problems of traditional monitoring privacy leakage, high-risk posture sample scarcity and random amplification to generate invalid skeleton samples, and is applicable to home, nursing ward and old-age care institution scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent monitoring, human posture recognition, sample amplification, and elderly safety monitoring, specifically to a stick figure posture intelligent monitoring method based on orthogonal amplification. Background Technology

[0002] In home-based elderly care, nursing home care, and hospital ward nursing scenarios, incidents such as falls, slips off the bedside, and prolonged lying down after a fall need to be identified as early as possible. Directly using raw video footage raises privacy concerns. Stick figure surveillance images simplify the human body to key points and skeletal lines, downplaying facial features, clothing, and interior details, while still preserving posture, movement, and support relationships.

[0003] The shortcomings of stick figure samples differ from those of blurred images. The main problems with blurred images often stem from brightness, texture, occlusion, compression distortion, and privacy blur intensity; the main problems with stick figure samples, however, arise from missed keypoint detections, broken skeletons, low joint confidence, confusion between left and right limbs, bone length ratio drift, and a limited number of samples across different action phases. If amplification schemes are still selected based on blur intensity or texture degradation, the true risk factors in stick figure samples may be overlooked.

[0004] Existing posture augmentation methods often employ joint coordinate perturbation, mirroring, scaling, or temporal cropping. While these methods increase the number of samples, for actions such as falls and slips off the bedside, arbitrary perturbations can disrupt bone length ratios, support contact points, and the sequence of action phases. A specialized orthogonal augmentation selection algorithm for stick figure surveillance images is needed, which selects orthogonal bases and auxiliary augmentation methods based on skeletal structure and posture motion. Summary of the Invention

[0005] To address the existing problems in stick figure datasets for elderly care monitoring, such as class imbalance, invalid amplified samples, and insufficient accuracy in recognizing high-risk actions, this invention aims to provide an intelligent stick figure posture monitoring method based on orthogonal amplification.

[0006] To achieve the above objectives, the present invention provides a method for intelligent monitoring of stick figure postures based on orthogonal amplification, comprising the following steps: S1. Collect video streams of elderly monitoring scenarios, perform human detection and human pose estimation on the video streams, and obtain the coordinates of key points of the monitored object, key point confidence, human detection box and time series; S2. Generate a stick figure monitoring image or a stick figure posture matrix based on the coordinates of the human body key points. The stick figure monitoring image includes key point nodes, skeleton lines and gray values ​​used to represent the confidence of key points or the visibility of the skeleton. S3. Calculate the stick figure augmentation selection features of the stick figure monitoring image or stick figure posture matrix. The stick figure augmentation selection features include at least two of the following: skeleton integrity, key point missing location, left and right limb symmetry, bone length ratio stability, joint angular velocity, center of gravity descent amplitude, trunk tilt angle, human body orientation, support contact point, action stage, category risk level, and corresponding category sample quantity. S4. Establish a candidate amplification scheme library, which includes a variety of orthogonal sampling matrix sets and optional auxiliary attitude amplification schemes; wherein the orthogonal sampling matrix sets include, but are not limited to, CC-Hadamard, Walsh-Hadamard, Natural-Hadamard, ROCC-Hadamard, discrete cosine orthogonal basis, Wavelet orthogonal basis, and sampling matrix sets corresponding to learnable orthogonal bases that satisfy orthogonal constraints; the auxiliary attitude amplification schemes include at least one of joint coordinate perturbation, bone length constraint perturbation, temporal resampling, keypoint confidence masking, left and right mirroring, or attitude scale perturbation. S5. Calculate the matching score of each candidate amplification scheme based on the stick figure amplification selection features, and select one or more orthogonal sampling matrix sets and auxiliary posture amplification schemes to determine the amplification factor, number of sampling matrices, sampling matrix index and auxiliary posture amplification intensity for each category or each sample. S6. Select a sampling matrix from the selected set of orthogonal sampling matrices. The stick figure pose samples to be amplified respectively with the sampling matrix By multiplying each dot individually, we obtain the orthogonal dot product samples of the stick figures. , where k is the sampling matrix index; S7. Perform skeleton semantic consistency verification on the stick figure orthogonal dot product samples. After the verification is passed, add them to the stick figure training dataset. S8. Use the amplified stick figure training dataset to train the elderly monitoring and recognition model, and output the risk recognition results of falls, non-falls, or prolonged lying down during the online monitoring stage.

[0007] Preferably, the matching score in step S5 is obtained by weighting the skeleton integrity, key point reliability, bone length stability, joint motion, support contact, motion risk, and number of category samples; the matching score does not take image brightness, image texture, or blur intensity as the dominant factor.

[0008] When the head, shoulder, hip, knee, and ankle main chain key points are continuous and the bone length ratio is stable, the CC-Hadamard or Natural-Hadamard sampling matrix set is preferred; when the joint angular velocity, center of gravity descent amplitude, or trunk tilt angle change exceeds the threshold, the Walsh-Hadamard sampling matrix set is preferred; when the proportion of missing key points at the elbow, knee, wrist, or ankle exceeds the threshold, the ROCC-Hadamard sampling matrix set is preferred; when the postural scale changes or the body orientation changes significantly, the DCT sampling matrix set is preferred; when the support contact point switches between the wrist, knee, ankle, and the ground or bedside, the Wavelet sampling matrix set is preferred.

[0009] Preferably, the stick figure orthogonal dot product sample satisfies = ⊙ ,in This represents the amplified stick figure pose sample. represents the k-th orthogonal sampling matrix, and ⊙ represents point-by-point multiplication; multiple amplified samples corresponding to the same amplified sample are all subjected to the point-by-point multiplication process.

[0010] Preferably, when the auxiliary posture augmentation scheme is used in conjunction with the orthogonal sampling matrix set, the stick figure posture sample is first multiplied point by point with the orthogonal sampling matrix, and then joint coordinate perturbation, temporal resampling or key point confidence masking is performed within the range of bone length ratio constraints and joint angle constraints.

[0011] Preferably, the skeleton semantic consistency verification in step S7 includes at least one of the following: bone length ratio deviation verification, left and right limb topology verification, joint angle range verification, center of gravity change direction verification, support contact point verification, and action phase sequence verification; stick figure augmentation samples that fail the verification are not added to the training set.

[0012] When the number of samples in a category is lower than the preset threshold or the risk level of the monitored event is higher than the preset level, the number of samples in the corresponding category is increased; among them, the amplification factor of the categories of pre-fall instability, falling, landing, slipping off the bedside, recovery from a partial fall, or prolonged lying down is higher than that of the normal walking category.

[0013] Preferably, the elderly monitoring and recognition model includes a stick figure posture feature encoding part and a temporal classification part; the temporal classification part adopts one or more of the following: bidirectional gated recurrent unit, attention mechanism, temporal convolutional network or Transformer temporal encoder.

[0014] Preferably, the stick figure monitoring image is obtained by extracting human key points from privacy-blurred images, depth images, infrared images, or ordinary video frames. During the online monitoring stage, only the stick figure posture sequence or stick figure monitoring image is input, and it is not necessary to retain clear human face or clothing texture.

[0015] An orthogonal augmentation selection system for stick figure surveillance images includes: a video acquisition module, a human pose estimation module, a stick figure image generation module, a stick figure augmentation selection feature calculation module, a candidate augmentation scheme library, an augmentation scheme selection module, an orthogonal sampling matrix generation module, a skeleton semantic verification module, a model training module, and an online monitoring and recognition module; wherein the augmentation scheme selection module is used to perform the augmentation scheme selection step, and the orthogonal sampling matrix generation module is used to generate a sampling matrix and multiply it point-by-point with the stick figure pose samples.

[0016] The present invention has the following beneficial effects: 1. The entire process uses stick figure skeleton images with faces and clothing textures removed as the input and reasoning medium, eliminating the need to store and transmit clear human images, thus avoiding the risk of privacy leaks in elderly care monitoring scenarios from the data source. 2. Abandoning the traditional amplification selection logic based on image texture and blur, it uses skeleton integrity, joint movement, and limb support relationship as the core judgment criteria, and matches exclusive orthogonal sampling matrices for different human postures and occlusion scenarios, which greatly improves the diversity and effectiveness of amplified samples; 3. The sample is expanded by multiplying multiple orthogonal bases point by point, and all auxiliary transformations are subject to human physiological constraints to avoid generating invalid samples that violate human structure. 4. Added multi-dimensional skeleton logic verification to remove amplified samples with topological errors and contradictory motion logic, and eliminated the interference of invalid data on the time series recognition model; 5. Increase the amplification factor for rare high-risk categories such as falls, bed falls, and prolonged lying down to solve the problem of dataset category imbalance. The accuracy of high-risk identification is significantly better than no amplification and traditional random amplification schemes. 6. It is compatible with ordinary visible light, infrared, depth, and privacy-blurred cameras. The backend temporal model supports multiple mainstream temporal networks such as BiGRU, TCN, and Transformer, and is suitable for various elderly care monitoring scenarios such as home, institutions, and hospital wards. Attached Figure Description

[0017] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments; Figure 1 This is a system block diagram of the present invention; Figure 2 This is a visual illustration of the point-by-point multiplication and amplification of the original stick figure sample with multiple orthogonal sampling matrices in this invention. Figure 3This is a comparison chart of the recognition accuracy of different amplification schemes and no amplification scheme in the embodiments of the present invention. Detailed Implementation

[0018] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.

[0019] Reference Figure 1-3 The specific implementation adopts the following technical solution: 1. Implementation process for video acquisition and stick figure monitoring image generation Data acquisition devices were deployed at the head of bedrooms, beside nursing beds, and in common living areas of elderly care facilities to continuously output real-time monitoring video streams. Human detection and pose estimation were performed on the video stream frame by frame or in fixed time windows to calculate the two-dimensional coordinates and single keypoint confidence values ​​of key points on the entire body, including the head, shoulders, elbows, wrists, hips, knees, and ankles. All keypoint coordinates are normalized to the human detection bounding box or a standardized canvas of uniform size. Following preset human skeleton topology connection rules, a human skeleton is generated by connecting keypoints with their coordinates. The grayscale values ​​are then mapped based on the confidence scores of each keypoint, ultimately outputting a standardized stick figure surveillance image as the original sample for subsequent augmentation processing. .

[0020] 2. Implementation Logic of Stick Figure Augmentation Feature Classification and Matching Scoring This embodiment divides the features used to screen amplification schemes into four categories, all based on stick figure skeleton data extraction, without relying on image texture or brightness information: (1) Skeletal structural features: skeleton integrity, main chain connectivity of head, shoulder, hip, knee and ankle, left and right limb symmetry, stability of the proportion of bone length in the whole body, location and proportion of missing key points in local areas; (2) Motion timing characteristics: single joint angular velocity, change in center of gravity height in consecutive frames, change in torso tilt angle, overall orientation shift of the human body, and division of motion timing stages; (3) Support contact characteristics: the spatial positional relationship and contact switching status of the wrist, knee, and ankle limb endpoints relative to the ground, the edge of the nursing bed, and the handrail; (4) Data set distribution characteristics: the number of samples in each posture category and the risk level of the corresponding monitoring event.

[0021] Based on the above feature weighting, the matching score of each orthogonal sampling matrix scheme is calculated, and the dominant terms of the scores of different orthogonal bases are distinguished as follows: CC-Hadamard and Natural-Hadamard sampling matrices: The matching score is dominated by a weighted average of two features: main chain integrity and whole-body bone length stability. Walsh-Hadamard sampling matrix: The matching score is dominated by joint angular velocity, center of gravity height descent, and temporal variation characteristics of the movement phase; ROCC-Hadamard sampling matrix: The matching score is dominated by the proportion of missing local keypoints and the degree of confidence decay of joint keypoints; DCT orthogonal sampling matrix: The matching score is dominated by the overall scale of human posture, human orientation offset, and global skeleton displacement features; Wavelet orthogonal sampling matrix: The matching score is dominated by the frequency of contact point switching supported by limb endpoints, local limb deformation, and multi-scale skeletal structure features.

[0022] The system selects one or more sets of orthogonal sampling matrices with the highest matching scores, and then determines the number of orthogonal matrices used in a single amplification and the total amplification factor of a single sample based on the scarcity of samples for the corresponding posture category.

[0023] Example 1: High-risk action scenario of falling and dropping The number of fall-related samples in the original dataset is relatively low. The system can generate two or more augmented samples for a single stick figure posture sequence. When the time sequence detects a rapid decrease in the height of the center of gravity and a trunk tilt angle greater than 70° (close to horizontal), the matching weights of the Walsh-Hadamard and Wavelet matrices are automatically increased, and these two types of orthogonal bases are preferentially selected for amplification. If there are short-term missed detections and the confidence level drops to zero at key points of the wrist and knee in the sequence, the matching weight of the ROCC-Hadamard matrix is ​​increased to simulate a real occlusion scenario of local skeletal fracture. If there are no missing key points in the human main chain and the bone length ratio fluctuates very little, and only the sample coverage needs to be expanded, the CC-Hadamard or Natural-Hadamard matrix can be used to complete the balanced sampling amplification.

[0024] By performing point-by-point multiplication of orthogonal matrices, we obtain = ⊙ Subsequently, auxiliary amplifications such as joint perturbation and temporal resampling are superimposed, and skeleton semantic consistency verification is performed: the verification dimensions include bone length ratio deviation, joint physiological angle range, center of gravity descent direction, support contact logic, and action phase temporal sequence. Samples that fail the verification are directly removed, and only valid amplified samples are incorporated into the training set.

[0025] Example 2: For monitoring scenarios involving bedridden or sedentary conditions, the original stick figure pose time sequence is directly fed into a recognition model constructed using BiGRU-Attention, Temporal Convolutional Network (TCN), or Transformer Temporal Encoder; the model training dataset consists of the original stick figure samples and orthogonally amplified samples filtered by skeleton semantic verification.

[0026] During the online monitoring and inference phase, the system only inputs the stick figure posture sequence into the recognition model. There is no need to store, transmit, or input clear facial images, clothing textures, or other private image information throughout the process, thus achieving a balance between privacy and posture recognition.

[0027] Compared to orthogonal augmentation selection algorithms for blurred surveillance images, the input features, matching scores, and verification conditions in this embodiment are all different. The blurred image version focuses on image degradation and scene texture; the stick figure version focuses on skeleton topology, joint motion, support contact, and keypoint reliability. Both can use orthogonal sampling matrices such as Hadamard, DCT, and Wavelet, but the reasons for selection and the combination methods are different.

[0028] Experimental Example: To verify the recognition effect of the amplification scheme of the present invention, a uniform control experiment was set up: The original training set contains 70 independent stick figure pose time series samples; all amplification schemes uniformly adopt a 3x amplification scale, that is, each original sample adds 2 amplified samples, and the total number of training sets after expansion is 210; no sample amplification operations are performed on the validation set and test set.

[0029] All experimental groups used identical training / validation / test set partitioning rules and a unified BiGRU-Attention temporal recognition model for training and testing. For a detailed comparison of the recognition accuracy and F1 score of each scheme, please refer to the appendix. Figure 3 The actual recognition accuracy can vary slightly depending on the total number of samples in the dataset and the combination of orthogonal sampling matrices.

[0030] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A method for intelligent monitoring of a matchstick man's posture based on orthogonal amplification, characterized in that, Includes the following steps: (S1) Collect video streams of elderly monitoring scenarios, perform human detection and human pose estimation on the video streams, and obtain the coordinates of human key points, key point confidence, human detection box and time series of the monitored object. (S2) Generate a stick figure monitoring image or a stick figure posture matrix based on the coordinates of the human body key points. The stick figure monitoring image includes key point nodes, skeleton lines and gray values ​​used to represent the confidence of key points or the visibility of skeletons. (S3) Calculate the stick figure augmentation selection features of the stick figure monitoring image or stick figure posture matrix. The stick figure augmentation selection features include at least two of the following: skeleton integrity, key point missing location, left and right limb symmetry, bone length ratio stability, joint angular velocity, center of gravity descent amplitude, trunk tilt angle, human body orientation, support contact point, action stage, category risk level, and corresponding category sample quantity. (S4) Establish a candidate amplification scheme library, which includes a variety of orthogonal sampling matrix sets and optional auxiliary attitude amplification schemes; wherein the orthogonal sampling matrix sets include, but are not limited to, CC-Hadamard, Walsh-Hadamard, Natural-Hadamard, ROCC-Hadamard, discrete cosine orthogonal basis, Wavelet orthogonal basis, and sampling matrix sets corresponding to learnable orthogonal bases that satisfy orthogonal constraints; the auxiliary attitude amplification schemes include at least one of joint coordinate perturbation, bone length constraint perturbation, temporal resampling, key point confidence masking, left and right mirroring, or attitude scale perturbation. (S5) Calculate the matching score of each candidate amplification scheme based on the stick figure amplification selection features, and select one or more orthogonal sampling matrix sets and auxiliary posture amplification schemes to determine the amplification factor, number of sampling matrices, sampling matrix index and auxiliary posture amplification intensity for each category or each sample. (S6) selecting a sampling matrix in the selected set of orthogonal sampling matrixes the matchstick posture sample to be amplified respectively with the sampling matrix point-by-point multiplication, obtaining a matchstick orthogonal point multiplication sample wherein k is the sampling matrix serial number (S7) Perform skeleton semantic consistency verification on the stick figure orthogonal dot product samples. After the verification is passed, add them to the stick figure training dataset. (S8) Use the amplified stick figure training dataset to train the elderly monitoring and recognition model, and output the risk recognition results of falls, non-falls or prolonged lying down during the online monitoring stage.

2. The method according to claim 1, wherein, The matching score in step (S5) is obtained by weighting the skeleton integrity, key point reliability, bone length stability, joint motion, support contact, motion risk, and number of category samples; the matching score does not take image brightness, image texture, or blur intensity as the dominant factor.

3. The method of claim 1, wherein the method is based on a quadratically amplified matchstick figure posture intelligent monitoring method. When the head, shoulder, hip, knee, and ankle main chain key points are continuous and the bone length ratio is stable, the CC-Hadamard or Natural-Hadamard sampling matrix set is preferred; when the joint angular velocity, center of gravity descent amplitude, or trunk tilt angle change exceeds the threshold, the Walsh-Hadamard sampling matrix set is preferred; when the proportion of missing key points at the elbow, knee, wrist, or ankle exceeds the threshold, the ROCC-Hadamard sampling matrix set is preferred; when the postural scale changes or the body orientation changes significantly, the DCT sampling matrix set is preferred; when the support contact point switches between the wrist, knee, ankle, and the ground or bedside, the Wavelet sampling matrix set is preferred.

4. The method for intelligent monitoring of stick figure posture based on orthogonal amplification according to claim 1, characterized in that, The stick figure orthogonal dot product sample satisfies = ⊙ ,in This represents the amplified stick figure pose sample. represents the k-th orthogonal sampling matrix, and ⊙ represents point-by-point multiplication; multiple amplified samples corresponding to the same amplified sample are all subjected to the point-by-point multiplication process.

5. The stick figure posture intelligent monitoring method based on orthogonal amplification according to claim 1, wherein when the auxiliary posture amplification scheme is used in conjunction with the orthogonal sampling matrix set, the stick figure posture sample is first multiplied point by point with the orthogonal sampling matrix, and then joint coordinate perturbation, temporal resampling or key point confidence masking is performed within the range of bone length ratio constraint and joint angle constraint.

6. The stick figure posture intelligent monitoring method based on orthogonal amplification according to claim 1, wherein the skeleton semantic consistency verification in step (S7) includes at least one of the following: bone length ratio deviation verification, left and right limb topology verification, joint angle range verification, center of gravity change direction verification, support contact point verification, and action stage sequence verification; stick figure amplification samples that fail the verification are not added to the training set.

7. The stick figure posture intelligent monitoring method based on orthogonal amplification according to claim 1, when the number of category samples is lower than a preset threshold or the risk level of the monitored event is higher than a preset level, the number of sampling matrices for the corresponding category is increased; wherein the amplification factor of the categories of pre-fall instability, falling, landing, slipping off the bedside, half-fall recovery or lying down for a long time is higher than that of the normal walking category.

8. The stick figure posture intelligent monitoring method based on orthogonal amplification according to claim 1, wherein the elderly monitoring and recognition model includes a stick figure posture feature encoding part and a temporal classification part; wherein the temporal classification part adopts one or more of bidirectional gated recurrent units, attention mechanisms, temporal convolutional networks or Transformer temporal encoders.

9. The stick figure posture intelligent monitoring method based on orthogonal amplification according to claim 1, wherein the stick figure monitoring image is obtained by extracting human key points from privacy blurred images, depth images, infrared images or ordinary video frames, and only the stick figure posture sequence or stick figure monitoring image is input during the online monitoring stage, without needing to retain clear human face or clothing texture.

10. An orthogonal augmentation selection system for stick figure surveillance images, characterized in that, include: The system includes a video acquisition module, a human pose estimation module, a stick figure image generation module, a stick figure amplification selection feature calculation module, a candidate amplification scheme library, an amplification scheme selection module, an orthogonal sampling matrix generation module, a skeleton semantic verification module, a model training module, and an online monitoring and recognition module. The amplification scheme selection module is used to perform the amplification scheme selection step according to any one of claims 1 to 9, and the orthogonal sampling matrix generation module is used to generate a sampling matrix and multiply it point by point with the stick figure posture sample.