Human body posture abnormity monitoring method based on multi-sensor fusion
By using a multi-sensor fusion method, visual, radar, and audio data are adaptively acquired and processed. Multimodal fusion and individual baseline comparison are performed to solve the problems of false alarms and missed alarms and privacy protection in human posture monitoring, and achieve efficient and accurate posture anomaly monitoring.
Patent Information
- Application Number
- CN202511466538.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-03-06
AI Technical Summary
Existing human posture monitoring technologies are susceptible to environmental interference, have high false alarm and false negative rates, and lack privacy protection. Furthermore, multi-sensor fusion methods fail to fully utilize the spatiotemporal complementarity of multi-source data, making it difficult to achieve continuous and reliable monitoring.
A multi-sensor fusion method is adopted to adaptively acquire hierarchical blurred images, hierarchical radar data, and spatial audio data. Visual, radar, and audio pose data are extracted separately and multimodal fusion is performed. Based on the fused pose data, multi-indicator joint judgment is performed and compared with individual baselines. The radar resolution and privacy protection strength are dynamically adjusted to reduce the false alarm rate and improve the recognition accuracy.
It achieves accurate monitoring of human posture under different environmental conditions, reduces false alarm and false alarm rates, meets privacy protection requirements, adapts to individual differences, and provides hierarchical alarms with rapid response.
Smart Images

Figure CN121606284A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human posture technology, and in particular to a method for monitoring abnormal human posture using multi-sensor fusion. Background Technology
[0002] With the increasing demand for home-based elderly care, rehabilitation care, protection of people living alone, and hospital and institutional care, the monitoring of abnormal human postures (such as falls, abnormal lying down, prolonged stillness, and convulsions) in adults is becoming increasingly important. Existing monitoring solutions mostly use single sensors such as cameras (visible light / infrared), radar, or audio sensors to identify and monitor human behavior and states. However, single sensors are susceptible to environmental interference (lighting, obstruction, noise, changes in viewing angle, metal reflection, etc.) or are insensitive to specific postures, leading to high false alarm and false negative rates; they also struggle to achieve continuous and reliable monitoring when data quality is unstable.
[0003] To compensate for the limitations of a single signal source, some studies have attempted to connect multiple sensors (such as cameras and radar) in parallel. However, the common approach is to interpret each signal independently or to simply vote on / weight the outputs of each sensor. This method fails to fully utilize the spatiotemporal complementarity of multi-source data, making it difficult to establish dynamic correlations between behavioral, physical, and acoustic cues. When conflicts arise between different modalities or when single-factor false alarms occur, it still lacks end-to-end fusion and discrimination capabilities. In addition, existing technologies often use static thresholds for comparison. Static thresholds are difficult to cover differences in age, weight, physical fitness, and lifestyle habits, as well as contextual differences such as day / night / room conditions, which can also easily lead to false positives.
[0004] Meanwhile, privacy compliance and data security are particularly critical in home and healthcare settings. Common high-resolution camera solutions collect sensitive information about users and family members, and existing methods often rely on backend algorithms to blur faces or identify individuals. If the device is attacked or the algorithm is bypassed, the original image may still be leaked, failing to meet the privacy protection requirement of "unidentifiable at the source." However, excessive blurring of images can also lead to the loss of human posture information, resulting in more serious false alarms.
[0005] The purpose of this invention is to design a multi-sensor fusion method for detecting abnormal human posture in order to address the problems existing in the prior art. Summary of the Invention
[0006] In view of this, the purpose of this invention is to propose a multi-sensor fusion method for detecting abnormal human posture, which can solve the above-mentioned problems.
[0007] This invention provides a multi-sensor fusion method for detecting abnormal human posture, comprising: Adaptively acquire graded blurred images, graded radar data, and spatial audio data, and extract visual pose data, radar pose data, and audio pose data respectively; Multimodal fusion is performed on visual pose data, radar pose data, and audio pose data to obtain corresponding fused pose data; Based on the fused posture data, multiple indicators are jointly judged and compared with individual baselines to determine whether there are posture anomalies and to issue graded alarms.
[0008] The beneficial effects of this invention are: First, the radar resolution and sampling are dynamically adjusted based on distance, speed, and number of targets. This automatically improves spatiotemporal resolution during high-speed, close-range, and rapidly changing scenarios to avoid missing crucial instantaneous actions. In long-range, low-speed, and multi-target scenarios, bandwidth and computing power are balanced to maintain stable global tracking. Privacy is protected through active optical interference blurring, which is linked to radar resolution levels. Interference intensity is reduced when risk increases to preserve necessary contour / motion information, while interference is increased when risk decreases to maximize privacy. Noise reduction, sound source separation, and anomalous sound classification are used to obtain sound source direction and confidence levels for semantic and spatial verification with vision / radar.
[0009] Secondly, human behavior type verification is performed through trajectory continuity to avoid mistaking noise / false targets for real behavior and reduce false alarms. False alarms are effectively suppressed when posture is normal but noise is abnormal, and missed alarms are prevented when posture is abnormal and weak sound evidence is present, thus improving overall recognition accuracy. Dynamic weighting based on correlation / confidence and knowledge graphs addresses the inconsistency in recognition across different modalities.
[0010] Third, anomaly candidates and initial risk scores are output through a unified posture health model; then, individual baselines are constructed based on age, weight, lifestyle habits, and historical trajectories for dynamic threshold calibration, taking into account both recall and individual differences, significantly reducing false positives in atypical populations. Deviations between core indicators and dynamic thresholds are calculated to form an interpretable anomaly intensity representation, avoiding the problem of a one-size-fits-all approach using static thresholds. The initial risk score and deviation are jointly used to generate a tiered strategy, achieving accurate and rapid response. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings required in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a system module diagram of Embodiment 1.
[0013] Figure 2 This is a flowchart of the method in Example 2. Detailed Implementation
[0014] To facilitate understanding by those skilled in the art, the structure of the present invention will now be described in further detail with reference to the accompanying drawings. It should be understood that, unless otherwise specified, the order of the steps mentioned in this embodiment can be adjusted according to actual needs, and they can even be executed simultaneously or partially simultaneously.
[0015] Example 1 like Figure 1 As shown, Embodiment 1 provides a multi-sensor fusion-based human posture anomaly monitoring system, including: Radar is used to acquire graded radar data and detect abnormal human postures through graded radar data. Privacy-protecting cameras are installed in the same monitoring area as radar to acquire hierarchically blurred spatial images, and to detect abnormal human postures through hierarchically blurred spatial images. Specifically, the privacy-protecting camera includes a camera and an active optical interference blurring structure. The active optical interference blurring structure is set on the shooting path of the camera and is used to perform dynamic physical blurring processing on the image acquired by the camera. The active optical interference blurring structure includes: a switch, a light-emitting device, and a brightness adjustment device. The light-emitting device faces the camera or the camera's shooting direction. By turning the light-emitting device on or off, and cooperating with the brightness adjustment device to adjust the brightness of the light-emitting device, different degrees of camera exposure interference can be achieved. In this embodiment, an active optical interference blurring structure is positioned along the camera's shooting path. This structure emits light of appropriate intensity to actively interfere with the camera's imaging. The purpose of the light-emitting device facing the camera and brightening it is to induce overexposure, glare, and curtain light in the imaging path, thereby erasing details while preserving outlines and creating an unrecognizable "privacy-blurred" image. By controlling the on / off state or brightness level of the light-emitting device through a switch and brightness adjustment device, the degree of overexposure of the image can be flexibly adjusted, i.e., the blur range can be controlled to adapt to privacy protection requirements. For example, when strong privacy protection is required, the light-emitting device can operate continuously with strong light, so that the image captured by the camera consists only of bright areas and blurred images, making it difficult to restore facial and body features as well as environmental details. If it is necessary to reduce interference, the light-emitting device can be appropriately weakened to balance image usability and privacy protection.
[0016] This active optical interference blurring structure requires no additional physical light-transmitting material at the front of the camera, nor does it rely on changes in the physical state or mechanical movement of optical components. It achieves the unrecognizable nature of the source image solely through optical intervention of the light-emitting device. Even if the camera's underlying system is attacked or its algorithm is maliciously tampered with, the captured original image will only be an overexposed and blurred image, further enhancing the system's privacy and security level and meeting data compliance requirements in scenarios such as home, medical, and elderly care.
[0017] Example 2 like Figure 2 As shown, Embodiment 2 provides a multi-sensor fusion method for detecting abnormal human posture, including: S1 adaptively acquires hierarchical blurred images, hierarchical radar data, and spatial audio data, and extracts visual pose data, radar pose data, and audio pose data respectively. S101 acquires the radar characteristics of the detected target through radar, the target radar characteristics include: distance, speed, and number of radar targets, and adjusts the radar resolution according to the detected target characteristics to obtain graded radar data; S1011 If the distance is greater than the long-range threshold or the speed is less than or equal to the low-speed threshold, adjust the radar resolution to a low level. S1012 If the near-range threshold < distance < far-range threshold or the low-speed threshold < speed < high-speed threshold, adjust the radar resolution to the medium level. If the distance threshold or speed is greater than the high speed threshold, the radar resolution will be adjusted to a high level. S1014 If the number of radar targets is greater than or equal to the target threshold, adjust the radar resolution to the medium level.
[0018] In this step, the closer the target is, the stronger the echo signal, and the higher the resolution requirement. At close range, the target has richer details and changes faster, and more fine-grained features can be utilized (such as breathing / micro-movements). Close-range targets are prone to behavioral abnormalities (such as falling or struggling) and have less occlusion, requiring high resolution to detect movement details; at far range, the requirement is lower, and lower resolution can be used for coverage.
[0019] High-speed targets (such as running or sudden movement) indicate rapid scene changes, and ordinary frame rates or resolutions may miss key points, requiring higher time and spatial resolutions to capture the motion. Low-speed targets (walking or stationary) generally have low risk and high information redundancy, and can be captured at low resolutions; high-speed or violent motion requires high sampling and frequency to prevent missed detections.
[0020] In multi-target scenarios, radar beams and resources are scarce. To ensure overall tracking accuracy and separation under mutual obstruction, a balance needs to be struck between total bandwidth and the number of samples. For single targets, high-resolution coverage can be maximized, while for multi-target scenarios, medium resolution combined with a partitioning strategy can be used to avoid exceeding computing power limits and facilitate resource allocation for each target.
[0021] S102 adjusts the interference intensity of the active optical interference blurring structure of the blurring camera according to the radar resolution level corresponding to the detected target features, and obtains a graded blurring image. In this step, the radar resolution setting is essentially an adjustment signal driven by target characteristics (range, speed, behavior, number of targets). A high setting usually means "close range, high speed, or special behavior," requiring more temporal and structural cues from the visual end (reducing blurring within the privacy limit, i.e., lowering the intensity of active optical interference). A low setting means "long range, low speed, or low risk," allowing the visual end to maintain stronger interference intensity to prioritize privacy.
[0022] S103 obtains human behavior types and estimates human poses from hierarchical blurred images; S1031 inputs a hierarchical blurred image of the whole human body region into a pre-trained visual behavior recognition model to obtain the human behavior type; S1032 extracts skeletal key points from a hierarchical blurred image of the entire human body, adapts and fits the joints through spatial positional relationships, and outputs the estimated human posture.
[0023] In this step, the hierarchically blurred image preserves contour and motion shape information through low-passing while suppressing identity features (facial texture, fine-grained texture). This allows the pre-trained model to perform behavior recognition and pose estimation without revealing the identity. Human action semantics can be directly obtained through behavior classification, and pose can be obtained through skeletal keypoints to achieve structured geometric representations, which can be used independently for rule judgment, as well as for verification and enhancement.
[0024] S104 performs temporal correlation and tracking of human targets in three-dimensional point cloud data formed by radar data, extracts micro-motion indicators from temporal features to form human micro-motion trajectory, performs skeleton fitting on three-dimensional point cloud data, and outputs human radar estimated attitude. In this step, the human body's micro-motion trajectory and estimated posture are extracted by radar. The estimated posture by radar can obtain the three-dimensional posture features of the human body, and the human body's micro-motion trajectory can be used for subsequent accuracy verification with other data.
[0025] S105 acquires and classifies abnormal human sounds using spatial audio data.
[0026] S1051 performs environmental noise suppression and sound source separation on spatial audio data to obtain human audio data; S1052 uses a pre-trained acoustic recognition model to detect and classify abnormal sounds in human audio data, thus obtaining the abnormal sound type.
[0027] In this step, environmental noise suppression and sound source separation (array beamforming, DOA, blind source separation) are performed first to extract the target human voice / collision sound from the environmental noise, improving its discernibility. Then, a pre-trained acoustic model is used for anomaly detection and classification to avoid high false alarms caused by contaminated data. Spatial audio provides sound source direction and confidence levels, which are used for joint judgment with subsequent visual and radar analysis.
[0028] S2 obtains corresponding fused attitude data by performing multimodal fusion of visual attitude data, radar attitude data, and audio attitude data respectively; S201 determines the continuity of the current human micro-movement trajectory within the current sliding window. If the current human micro-movement trajectory is continuous, the human behavior type is output; otherwise, the human behavior type is marked as unreliable. S2011 counts the actual number of trajectory frames obtained by the current sliding window and compares it with the theoretical number of frames to be collected. If the frame loss rate is lower than the frame loss threshold, it is determined that the time sampling is basically continuous and a time stability test is performed. Otherwise, the human micro-movement trajectory is marked as discontinuous. S2012 calculates the time interval between adjacent trajectory points. If the maximum time interval is less than the interval threshold, it is determined that the time interval is continuous and stable, and spatial continuity is judged. Otherwise, it is marked as discontinuous human micro-movement trajectory. S2013 calculates the displacement and velocity of the trajectory points. If the displacement and velocity are within the physiological velocity threshold range, the micro-motion trajectory is judged to be continuous; otherwise, the human micro-motion trajectory is marked as discontinuous.
[0029] In this step, if the trajectory itself is discontinuous or distorted, sensor noise will be mistaken for real behavior, leading to incorrect judgment. Therefore, determining the continuity of the trajectory before outputting the behavior type ensures the accuracy of the behavior type.
[0030] The sliding window can be set to 3 seconds or 5 seconds. By comparing the frame drop rate, unreliable windows caused by severe frame drops can be quickly filtered out, avoiding invalid inferences made on low-quality data in subsequent calculations. Assuming the frame drop rate is acceptable, further verification is made to ensure that the temporal sampling is uniform and stable, avoiding pseudo-continuities caused by intermittent jitter or clock drift. After confirming temporal continuity, further elimination of instantaneous jumps / drifts in space (a common problem in multi-source fusion or radar / visual positioning) ensures that the trajectory conforms to physical and physiological constraints in spatial variations.
[0031] S202 cross-validates human image pose estimation, human radar pose estimation, and abnormal sound type to obtain pose and sound anomaly cross-validation data. S2021 determines whether the human body radar estimated attitude and the human body image estimated attitude are abnormal. If either human body estimated attitude is abnormal, the abnormal sound type is marked as reliable. If both the human body radar attitude estimation and the human body image attitude estimation are normal, then the abnormal sound type is marked as to be evaluated.
[0032] In this step, in real-world environments, audio events such as roaring, coughing, and groaning might be misjudged as high-risk abnormal sounds by the algorithm. However, if the human posture is normal (e.g., sitting, standing, walking steadily), the actual risk and response requirements are far lower than when the posture is abnormal (e.g., falling, lying down for a long time, convulsions). If any posture is judged as abnormal (e.g., falling, abnormal lying down, convulsions), even if the confidence level of the abnormal sound detection is average, as long as an abnormal sound type (e.g., impact, groaning, calling for help) is detected, the system will automatically increase the risk level and quickly issue an alarm, minimizing the false negative rate. If both postures are normal, the abnormal sound type is marked as "to be evaluated" and downweighted during subsequent multimodal fusion, effectively suppressing false alarms caused by environmental noise or occasional sounds.
[0033] S203 aligns motion-related data and posture-sound anomaly cross-validation data in terms of spatiotemporal features, and performs adaptive weighting based on posture correlation to obtain fused posture data.
[0034] In this step, posture relevance can be obtained through a knowledge graph constructed from expert experience. An algorithm aligns multimodal data to a unified time axis and spatial coordinates. Based on the anomaly relevance and confidence level of each posture in the current time period, weights are adaptively adjusted to construct fused posture data that best represents the real event. For example, if an elderly person living alone gets up at night, the visual posture judgment signal is weak due to poor lighting and viewing angle. However, if radar tracks continuously detect the person moving quickly and then suddenly stopping, the radar track anomaly is heavily weighted, while the posture fusion weight is low. Sound without anomalies can also be encoded with low weight. The final output shows posture dominating the movement trajectory, promptly increasing the anomaly risk index.
[0035] S3 uses fused posture data to perform multi-indicator joint judgment and compares it with individual baselines to determine whether there are posture anomalies and issue graded alarms.
[0036] S301 uses a posture health model to perform a global interpretation of the fused posture data and outputs abnormal candidate postures and their initial risk scores. If an anomaly candidate exists, S302 will call the individual baseline for comparison and calibration according to the attitude type to determine the final anomaly and alarm level.
[0037] S3021 constructs individual historical baseline parameters based on current human age and weight, and sets dynamic thresholds for individual baselines in combination with time scenarios; S3022 calculates the deviation of the current value of an abnormal candidate pose from the individual baseline dynamic threshold. If the deviation amount does not reach the classification threshold, it is marked as continuously monitored. If the deviation amount reaches the classification threshold, a classification alarm is issued based on the deviation amount and the initial risk score, and the corresponding abnormal posture is output.
[0038] In this step, a unified posture health model is used to globally interpret multimodal fusion postures, quickly identifying anomalous candidates. Different ages, weights, physical abilities, and lifestyles lead to significant differences in posture and micro-motion feature distributions. Introducing an individual baseline allows for re-evaluation of candidates, significantly reducing false alarms caused by individual differences. The risks of being still in bed at night are completely different from being still on the ground during the day; a dynamic threshold is set using time-scene conditions to avoid a one-size-fits-all approach. The degree of anomaly is quantified by the deviation from the individual baseline, which is more adaptable to cross-population and cross-time period differences than a fixed threshold. The initial risk score is combined with the deviation amount for tiered alerts, suppressing noise and enabling rapid response to high-risk scenarios.
[0039] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0040] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0041] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0042] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0043] It should be noted that any reference signs placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The words first, second, and third, etc., do not indicate any order. These words can be interpreted as names.
[0044] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention.
[0045] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
[0046] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0047] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms should not be construed as necessarily referring to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
Claims
1. A multi-sensor fusion based human posture anomaly monitoring method, characterized in that, The application relates to a human posture anomaly monitoring system based on multi-sensor fusion, which comprises the following steps: a radar is used to acquire hierarchical radar data, and the hierarchical radar data are used to detect abnormal human postures; a privacy protection camera is arranged in the same monitoring area as the radar and is used to acquire hierarchical fuzzy space images, and the hierarchical fuzzy space images are used to detect abnormal human postures; Specifically, the privacy protection camera comprises a camera and an active optical interference fuzzy structure, the active optical interference fuzzy structure is arranged on a shooting path of the camera, and is used to dynamically physically blur the images acquired by the camera; the active optical interference fuzzy structure comprises a switch, a light emitting device and a brightness adjusting device, the light emitting device faces the camera or a shooting direction of the camera, the light emitting device is turned on or off, the brightness adjusting device is used to adjust the brightness of the light emitting device, and different degrees of camera exposure interference are realized; the method comprises the following steps: adaptive acquisition of hierarchical fuzzy images, hierarchical radar data and spatial audio data, extraction of visual posture data, radar posture data and audio posture data; multi-modal fusion of the visual posture data, the radar posture data and the audio posture data to obtain corresponding fusion posture data; multi-index joint determination and individual baseline comparison based on the fusion posture data to determine whether there is a posture anomaly and to perform hierarchical alarm.
2. The multi-sensor fusion based human posture anomaly monitoring method according to claim 1, wherein, the adaptive acquisition of hierarchical fuzzy images, hierarchical radar data and spatial audio data, and the extraction of visual posture data, radar posture data and audio posture data comprise the following steps: acquisition of target radar features by the radar, wherein the target radar features comprise distance, speed and radar target quantity, adjustment of the radar resolution according to the target features to obtain hierarchical radar data; adjustment of the interference intensity of the active optical interference fuzzy structure of the fuzzy camera according to the radar resolution adjustment gear corresponding to the target features to obtain hierarchical fuzzy images; acquisition of human behavior types and human image estimated postures by the hierarchical fuzzy images; time sequence correlation and tracking of human targets in three-dimensional point cloud data formed by the radar data, extraction of micro-motion indexes from time sequence features to form human micro-motion trajectories, skeleton fitting of the three-dimensional point cloud data, and output of human radar estimated postures; acquisition of abnormal sounds of the human body and classification thereof by the spatial audio data.
3. The multi-sensor fusion based human posture anomaly monitoring method according to claim 2, wherein, the acquisition of target radar features by the radar, wherein the target radar features comprise distance, speed and radar target quantity, and the adjustment of the radar resolution according to the target features to obtain hierarchical radar data comprise the following steps: if the distance is greater than a long-distance threshold value or the speed is less than or equal to a low-speed threshold value, the radar resolution is adjusted to a low gear; if the distance is between a near-distance threshold value and a long-distance threshold value or the speed is between a low-speed threshold value and a high-speed threshold value, the radar resolution is adjusted to a middle gear; if the distance is less than or equal to the near-distance threshold value or the speed is greater than a high-speed threshold value, the radar resolution is adjusted to a high gear; if the number of radar targets is greater than or equal to a target threshold value, the radar resolution is adjusted to the middle gear.
4. The multi-sensor fusion based human posture anomaly monitoring method according to claim 2, wherein, the acquisition of human behavior types and human image estimated postures by the hierarchical fuzzy images comprises the following steps: the hierarchical fuzzy images of the whole human body are input into a pre-trained visual behavior recognition model to obtain human behavior types; The skeletal key point extraction is performed on the hierarchical fuzzy image of the whole body region of the human body, the joint points are adapted and fitted through the spatial position relationship, and the estimated posture of the human body is output.
5. The multi-sensor fusion based human posture anomaly monitoring method according to claim 2, wherein, The abnormal sound of the human body and the classification thereof through the spatial audio data acquisition include: The environmental noise suppression and sound source separation are performed on the spatial audio data to obtain human body audio data; The abnormal sound detection and type classification are performed on the human body audio data through the pre-trained acoustic recognition model to obtain the abnormal sound type.
6. The multi-sensor fusion based human posture anomaly monitoring method according to claim 1, wherein, The multi-modal fusion of the visual posture data, the radar posture data and the audio posture data respectively to obtain the corresponding fusion posture data includes: The continuity of the current human body micro-motion trajectory in the current sliding window is judged, if the current human body micro-motion trajectory is continuous, the human body behavior type is output, otherwise the human body behavior type is marked as untrustworthy; The posture sound abnormal cross-validation data is obtained by cross-verification of the human body image estimated posture, the human body radar estimated posture and the abnormal sound type; The motion correlation data and the posture sound abnormal cross-validation data are time-space feature aligned, and the fusion posture data is obtained through adaptive weighting according to the posture correlation.
7. The multi-sensor fusion based human posture anomaly monitoring method according to claim 6, wherein, The continuity of the current human body micro-motion trajectory in the current sliding window is judged, if the current human body micro-motion trajectory is continuous, the human body behavior type is output, otherwise the human body behavior type is marked as untrustworthy includes: The actual number of trajectory frames obtained in the current sliding window is counted and compared with the theoretical number of frames to be collected, if the frame loss rate is lower than the frame loss threshold, it is determined that the time sampling is basically continuous, the time stability test is performed, otherwise the human body micro-motion trajectory is marked as discontinuous; The time interval of adjacent trajectory points is calculated, if the maximum time interval is less than the interval threshold, it is determined that the time interval is continuous and stable, the spatial continuity is judged, otherwise the human body micro-motion trajectory is marked as discontinuous; The trajectory point displacement and its speed are calculated, if the displacement speed is within the physiological speed threshold range, it is determined that the micro-motion trajectory is continuous, otherwise the human body micro-motion trajectory is marked as discontinuous.
8. The multi-sensor fusion based human posture anomaly monitoring method according to claim 6, wherein, The posture sound abnormal cross-validation data is obtained by cross-verification of the human body image estimated posture, the human body radar estimated posture and the abnormal sound type includes: It is respectively judged whether the human body radar estimated posture and the human body image estimated posture are abnormal, if any of the human body estimated postures is abnormal, the abnormal sound type is marked as trustworthy; If the human body radar estimated posture and the human body image estimated posture are both normal, the abnormal sound type is marked as to be evaluated.
9. The multi-sensor fusion based human posture anomaly monitoring method according to claim 1, wherein, The multi-index joint determination and individual baseline comparison based on the fusion posture data are performed to determine whether there is a posture abnormality and to perform hierarchical alarm includes: The fusion posture data is globally interpreted through the posture health model, and the abnormal candidate posture and its initial risk score are output; If there is an abnormal candidate, the individual baseline is called for comparison and calibration according to the posture type, and the final abnormality and alarm level are determined.
10. The multi-sensor fusion based human posture anomaly monitoring method according to claim 9, wherein, If there is an abnormal candidate, the individual baseline is called for comparison and calibration according to the posture type, and the final abnormality and alarm level are determined includes: The individual historical baseline parameters are constructed through the current human body age and weight, and the individual baseline dynamic threshold is set in combination with the time scene; For the abnormal candidate posture, the deviation amount of its current value from the individual baseline dynamic threshold is calculated; If the deviation amount does not reach the grading threshold, it is marked as a sustained concern, and if the deviation amount reaches the grading threshold, graded alarm is performed according to the deviation amount and the initial risk score, and the corresponding abnormal posture is output.