Anti-falling identification early warning method based on image identification and related equipment thereof

Through the anti-fall recognition early warning method based on image recognition, the multimodal fusion of RGB-D data flow and inertial monitoring data is solved, and the problem of insufficient accuracy of traditional methods in complex environments is achieved, and an efficient anti-fall warning is achieved.

CN120279658AInactive Publication Date: 2025-07-08FOSHAN CHANCHENG DISTRICT GLOBAL ELECTRICAL PORCELAIN ELECTRICAL MATERIALS CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510512569.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional fall warning methods rely on a single sensor or simple threshold detection, making it difficult to accurately capture changes in human movement in complex environments, resulting in false alarms or missed alarms, and the deep fusion of multi-dimensional information cannot be achieved, and the response speed and accuracy are insufficient.

Method used

The anti-fall recognition warning method based on image recognition is adopted, and multi-modal pre-processing is performed by acquiring real-time image acquisition data sets, target detection and three-dimensional attitude modeling of RGB-D data streams are carried out, and multi-modal fusion is carried out in combination with inertial monitoring data to achieve risk quantification and hierarchical early warning.

Benefits of technology

It improves the identification accuracy and robustness in complex environments, can predict the free fall movement trajectory of personnel in real time, issue early warnings in a timely manner, and reduce the risk of accident injury.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279658A_ABST
    Figure CN120279658A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, and provides an anti-falling recognition early warning method based on image recognition and related equipment thereof. The method comprises the following steps: carrying out multi-modal preprocessing on a real-time image acquisition data set to obtain an RGB-D data stream, carrying out target detection and three-dimensional attitude modeling on the RGB-D data stream to obtain a target personnel label set and a personnel attitude parameter set, and carrying out trajectory prediction on the personnel attitude parameter set through a physical kinematics model to obtain predicted motion trajectory data. Acquiring an inertial monitoring data set in real time according to the target person label set, performing multi-modal fusion in combination with the predicted motion trajectory data to obtain a confidence evaluation value, and performing risk quantification on the confidence evaluation value and the predicted motion trajectory data to obtain a graded early warning instruction. And performing protocol coding and signal conversion on the graded early warning instruction to obtain a control signal. According to the invention, through image identification, motion prediction, multi-modal data fusion and fine processing, the accuracy of anti-falling monitoring in a complex scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image recognition, and in particular, to a fall prevention recognition and warning method based on image recognition and related devices thereof. Background Art

[0002] With the increasing complexity of personnel activities in working environments and life scenarios, timely and accurate warnings can provide effective intervention before accidents occur, saving lives and reducing economic losses. Therefore, fall prevention warnings are of extremely important significance for ensuring personnel safety and reducing the risk of accident injuries.

[0003] Traditional fall prevention warning methods usually rely on single sensors or simple threshold detection, such as using an accelerometer or a pressure sensor alone for monitoring. Their data processing means are relatively rough, often difficult to accurately capture the complex changes in human movement, and prone to false alarms or missed alarms. At the same time, these methods have obvious deficiencies in data collection, preprocessing, and motion state analysis, and are unable to deeply fuse multi-dimensional information, resulting in the response speed and accuracy in complex environments being difficult to meet actual needs. Summary of the Invention

[0004] In view of this, the present application provides a fall prevention recognition and warning method based on image recognition and related devices thereof to solve the problem of low accuracy of fall prevention warnings in complex scenarios.

[0005] The first aspect of the present application provides a fall prevention recognition and warning method based on image recognition, the method comprising: Obtaining a real-time image acquisition data set of a target area, and performing multi-modal preprocessing on the real-time image acquisition data set to obtain a standardized RGB-D data stream; Performing target detection and three-dimensional pose modeling processing on the RGB-D data stream to obtain a target personnel label set and a personnel pose parameter set; Performing trajectory prediction processing on the personnel pose parameter set through a preset physical kinematic model to obtain predicted motion trajectory data; According to the target personnel label set, obtaining a corresponding inertial monitoring data set in real time, and performing multi-modal fusion processing on the predicted motion trajectory data and the inertial monitoring data set to obtain a confidence evaluation value corresponding to each target personnel; Performing risk quantification processing on the confidence evaluation value and the predicted motion trajectory data to obtain a hierarchical warning instruction corresponding to each target personnel; Performing protocol encoding and signal conversion processing on the hierarchical warning instruction according to a preset conversion method to obtain a control signal for driving a fall prevention device.

[0006] In an alternative embodiment, the real-time image acquisition dataset includes an original RGB image and an original depth map, and the multi-modal preprocessing of the real-time image acquisition dataset to obtain a standardized RGB-D data stream includes: Perform denoising processing on the original RGB image and the original depth map to obtain a noise-optimized RGB image and a noise-optimized depth map; Perform white balance correction processing on the noise-optimized RGB image to obtain a color-balanced optimized RGB image; Perform depth compensation processing on the noise-optimized depth map to obtain a compensated optimized depth map; Perform pixel alignment processing on the color-balanced optimized RGB image and the compensated optimized depth map to obtain initial RGB-D data; Perform histogram equalization and scale normalization processing on the initial RGB-D data to obtain scale-optimized RGB-D data; Perform format conversion processing on the scale-optimized RGB-D data to obtain a standardized RGB-D data stream.

[0007] In an alternative embodiment, the object detection and three-dimensional pose modeling processing of the RGB-D data stream to obtain a target person label set and a person pose parameter set includes: Perform foreground segmentation processing on the RGB-D data stream to obtain foreground region data containing people, and perform object detection processing on the foreground region data to obtain a set of object detection boxes and a set of person key point coordinates; Perform feature extraction processing on the RGB-D data stream according to the set of object detection boxes to obtain a set of target feature vectors; Perform identity association processing according to the set of target feature vectors to obtain the target person label set; Perform skeleton fitting processing according to the set of person key point coordinates to obtain a three-dimensional human skeleton model, and perform joint point pose calculation processing on the three-dimensional human skeleton model to obtain a set of joint point Euler angle parameters; Perform pose stability analysis processing according to the set of joint point Euler angle parameters to obtain the person pose parameter set.

[0008] In an alternative embodiment, the trajectory prediction processing of the person pose parameter set through a preset physical kinematic model to obtain predicted motion trajectory data includes: Perform temporal feature extraction processing on the person pose parameter set to obtain a sequence of hidden state vectors; Perform behavior classification processing according to the sequence of hidden state vectors to obtain a behavior category; When the behavior category is a preset target behavior, perform a fall motion prediction process on the implicit state vector sequence through the physical kinematics model to obtain a free fall motion trajectory function corresponding to each target person; Perform a time domain analysis process on the free fall motion trajectory function to obtain collision time data; Perform a combination process according to the free fall motion trajectory function and the collision time data to obtain the predicted motion trajectory data.

[0009] In an optional implementation manner, the obtaining the corresponding inertial monitoring data set in real time according to the target person tag set, and performing a multimodal fusion process on the predicted motion trajectory data and the inertial monitoring data set to obtain a confidence evaluation value corresponding to each target person includes: Perform a look-up table process on the target person tag set to obtain an inertial measurement unit tag set, and obtain the corresponding inertial monitoring data set in real time according to the inertial measurement unit tag set; Perform a state vector construction process on the predicted motion trajectory data and the inertial monitoring data set to obtain a visual state vector set and an inertial state vector set; Perform a Kalman filtering process on the visual state vector set and the inertial state vector set to obtain a predicted state covariance matrix; Perform a trace value calculation process on the predicted state covariance matrix to obtain a state covariance trace value; Perform a ratio calculation process on the state covariance trace value according to a preset maximum state covariance trace value to obtain a normalized confidence value, and perform a data combination process on the normalized confidence value to obtain a confidence evaluation value corresponding to each target person.

[0010] In an optional implementation manner, the predicted motion trajectory data includes a behavior category and collision time data, and the performing a risk quantification process on the confidence evaluation value and the predicted motion trajectory data to obtain a hierarchical warning instruction corresponding to each target person includes: Perform a feature vector component process on the behavior category, the collision time data, and the confidence evaluation value to obtain an input feature vector; Perform an exponential function calculation process on the input feature vector to obtain a risk index, and perform a threshold comparison process on the risk index according to a preset risk threshold to obtain a warning level; Perform a conditional judgment process on the collision time data and the confidence evaluation value to obtain an adjustment flag; Perform a conditional correction process on the warning level according to the adjustment flag to obtain a final warning level; Perform control signal mapping processing on the final warning level to obtain a hierarchical warning instruction corresponding to each target person.

[0011] In an optional implementation manner, the protocol encoding and signal conversion processing of the hierarchical warning instruction according to a preset conversion method to obtain a control signal for driving a fall prevention device includes: Perform device type identification processing on the hierarchical warning instruction to obtain a target device type identifier, and perform format mapping processing on the hierarchical warning instruction according to the target device type identifier to obtain device-specific instruction data; Perform electrical signal conversion processing on the device-specific instruction data to obtain an electrical signal set; Perform protocol classification and encoding processing on the electrical signal set according to the conversion method to obtain a communication control signal set; Perform signal comprehensive editing processing on the communication control signal set to obtain a control signal for driving a fall prevention device.

[0012] The second aspect of the present application provides a fall prevention recognition and warning device based on image recognition, and the device includes: An image acquisition module, configured to acquire a real-time image acquisition data set of a target area, and perform multi-modal preprocessing on the real-time image acquisition data set to obtain a standardized RGB-D data stream; A target detection module, configured to perform target detection and three-dimensional pose modeling processing on the RGB-D data stream to obtain a target person label set and a person pose parameter set; A trajectory prediction module, configured to perform trajectory prediction processing on the person pose parameter set through a preset physical kinematic model to obtain predicted motion trajectory data; A confidence evaluation module, configured to obtain a corresponding inertial monitoring data set according to the target person label set in real time, and perform multi-modal fusion processing on the predicted motion trajectory data and the inertial monitoring data set to obtain a confidence evaluation value corresponding to each target person; A hierarchical warning module, configured to perform risk quantification processing on the confidence evaluation value and the predicted motion trajectory data to obtain a hierarchical warning instruction corresponding to each target person; A fall prevention control module, configured to perform protocol encoding and signal conversion processing on the hierarchical warning instruction according to a preset conversion method to obtain a control signal for driving a fall prevention device.

[0013] In a third aspect of the present application, an electronic device is provided. The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned anti-fall recognition and warning method based on image recognition are implemented.

[0014] In a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned anti-fall recognition and warning method based on image recognition are implemented.

[0015] In summary, the present application includes at least the following beneficial technical effects: 1. Through preprocessing RGB images and depth maps and then combining inertial monitoring data, the complementarity of visual information and sensor data is achieved, enhancing the robustness and recognition accuracy in complex environments.

[0016] 2. By extracting temporal features, behavior classification, and physical motion prediction, the free-fall motion trajectory of the target person can be predicted in real time, and the collision time can be analyzed, providing a forward-looking judgment basis for anti-fall warning.

[0017] 3. Real-time operations are realized from data acquisition, processing, prediction to warning signal generation, enabling a warning to be issued in time before a person falls, thereby reducing the risk of accident injury. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 is a flowchart of an anti-fall recognition and warning method based on image recognition provided by an embodiment of the present application; Figure 2 is a functional module diagram of an anti-fall recognition and warning device based on image recognition provided by an embodiment of the present application; Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] Next, in combination with the accompanying drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0021] As Figure 1 shown, it is a flowchart of a fall prevention recognition and warning method based on image recognition provided by an embodiment of the present application. The fall prevention recognition and warning method based on image recognition provided by an embodiment of the present application includes the following steps.

[0022] Step S1: Obtain a real-time image acquisition data set of a target area, and perform multi-modal preprocessing on the real-time image acquisition data set to obtain a standardized RGB-D data stream.

[0023] It should be understood that the real-time image acquisition data set is environmental data obtained by detecting the target area through an industrial camera and a depth camera. The real-time image acquisition data set includes, but is not limited to, the original RGB image and the original depth map. Due to the inconsistent space of the camera settings, the real-time image acquisition data set is prone to data noise (for example, environmental noise, uneven brightness, sensor error, etc.). Therefore, after obtaining the real-time image acquisition data set of the target area, it is necessary to first perform preprocessing such as noise reduction, light compensation, alignment, and format standardization on the collected original RGB image and the original depth map respectively to ensure the spatial, brightness, and scale consistency between the data, and ensure the unity of the data in the subsequent processing stage. The specific operations are as follows: First, the industrial camera and the depth camera set near the target area respectively collect the original RGB image and the original depth map, where the RGB image records the color image information, and the depth map records the distance between each pixel and the camera. For the real-time image acquisition data set, noise reduction processing needs to be performed first. The non-local means filtering (for example, NL-Means) is used to denoise the collected original RGB image. By comparing the similarity of small areas around the pixels, the weighted average value is calculated to remove the random noise points in the image. The same corresponding denoising algorithm (for example, median filtering for the depth map or a variant of NL-Means) is also used for the collected original depth map, so that the noise in the depth information is effectively suppressed.

[0024] Furthermore, due to the phenomenon of local overbrightness or overdarkness in environmental light, it is necessary to perform light compensation on the noise-optimized RGB image after noise suppression. In the embodiments of the present application, adaptive histogram equalization is used to perform local brightness correction on the image. Specifically, the noise-optimized RGB image is divided into several small regions of a fixed size (for example, 16x16 pixel blocks), and the local average value and local standard deviation of the pixels are calculated within each region, denoted as μ local and σ local . The global average value and global standard deviation μ global and σ global are used to correct the local image. The histogram equalization process can be represented by the following formula: where I(x,y) represents the pixel value of the noise-optimized RGB image at the coordinate (x,y). μ local represents the average brightness of the local region (for example, 16×16 pixel block). σ local represents the standard deviation of the brightness of the local region. μ global represents the average brightness of the entire image. σ global represents the standard deviation of the brightness of the entire image. By combining the local brightness information with the global brightness standard, the purpose of balancing the image brightness distribution and compensating for local overdark or overbright regions is achieved. After the white balance correction process, a color-balanced optimized RGB image is obtained, and this image will be used for further feature extraction and target detection in the subsequent steps.

[0025] At the same time, since there may be depth value deviations in the depth map due to sensor errors or environmental interference, it is necessary to perform depth compensation processing on the noise-optimized depth map. In the embodiments of the present application, a statistical analysis-based method is used to smooth and correct the local depth data, so that the depth value of each pixel point is closer to the true value. After the depth compensation processing, a compensated optimized depth map is obtained, and the depth value of each pixel in this image represents an accurate physical distance.

[0026] After obtaining the color-balanced optimized RGB image and the compensated optimized depth map, since there may be a position deviation between the RGB image and the depth map due to different camera installation angles, it is necessary to ensure the one-to-one correspondence of their pixels through pixel alignment processing. Using the relative position parameters between the cameras (i.e., the rotation matrix R and the translation vector T), an affine transformation is performed on the depth map, so that each pixel in the compensated optimized depth map matches the corresponding pixel position in the color-balanced optimized RGB image. The pixel alignment process is based on geometric principles and is completed by calculating the transformation matrix, and the RGB information and depth information at specific positions in the obtained initial RGB-D data are completely corresponding.

[0027] For the initial RGB-D data after alignment processing, further histogram equalization and scale normalization processing are required to provide a consistent data basis for the next step of object detection and pose modeling. Histogram equalization processing adjusts the distribution of each gray level in the image, increases the contrast of the image, and makes the details clearer. The principle of the histogram equalization process of the initial RGB-D data is the same as that of the histogram equalization process of the noise-optimized RGB image, and will not be elaborated here. For details, please refer to the calculation process of the histogram equalization process of the noise-optimized RGB image. Further, scale normalization processing is performed on the initial RGB-D data after histogram equalization processing. Scale normalization processing adjusts the image data to a preset standard scale range, ensuring that each data set has a unified size and resolution in subsequent processing, so as to obtain scale-optimized RGB-D data.

[0028] Finally, format conversion is performed on the scale-optimized RGB-D data to meet the requirements of the preset data structure, thereby generating standardized RGB-D data. Further, the standardized RGB-D data is sorted according to the time series, thereby forming time series data (i.e., RGB-D data stream). The RGB-D data stream includes, but is not limited to, the RGB image and the corresponding depth map at each moment, and is accompanied by timestamp information for subsequent time series processing and fusion.

[0029] Through layer-by-layer multimodal preprocessing, not only the quality of the image and depth data is improved, but also the spatial, brightness, and scale consistency between the data is ensured. Among them, noise reduction processing eliminates environmental noise; light compensation solves the problem of uneven brightness; depth compensation corrects sensor errors; alignment processing ensures the precise matching of multimodal data; histogram equalization and normalization ensure the unity of the data in the subsequent processing stage. It lays a solid data foundation for subsequent object detection, pose modeling, and time series analysis, ensuring the high quality and high consistency of the data.

[0030] Step S2: Perform object detection and three-dimensional pose modeling processing on the RGB-D data stream to obtain a target person label set and a person pose parameter set.

[0031] After obtaining the RGB-D data, it is necessary to extract the personnel information existing in the target area from the RGB-D data stream, and perform three-dimensional reconstruction on the personnel actions and postures, so as to generate a detailed pose parameter set. The processing process of this step sequentially includes foreground segmentation, object detection, key point detection, skeleton fitting, and pose parameter calculation. The specific operations are as follows: First, the embodiment of the present application processes images using a background subtraction algorithm. A background model is constructed using the image data of consecutive frames, and the value of each pixel in the background model is a long-term stable value. For each pixel in the current frame image, it is compared with the value of the corresponding pixel in the background model. When the difference between the two exceeds a predetermined threshold, the pixel is marked as a foreground pixel. Here, depth map information can also be used for auxiliary judgment: if there is an obvious difference in the depth value within a certain area compared to the depth value in the background model, then this area is more likely to be the foreground. In this way, after foreground segmentation processing, the output result is the foreground area data that only contains the personnel targets within the target area at the corresponding moment. Reducing the interference information in the background through foreground segmentation processing helps improve the computational efficiency and accuracy of subsequent target detection and key point detection algorithms.

[0032] Furthermore, the foreground area data is processed using a target detection algorithm. The target detection algorithm uses a trained deep neural network model (e.g., the improved YOLOv7 algorithm), which can automatically identify personnel targets in the foreground area and give the bounding box (i.e., the target detection box) for each target. After foreground segmentation and target detection processing, the output data includes a set of target detection boxes and a set of personnel key point coordinates. Among them, the set of target detection boxes represents the position and size of each detected person in the image, usually represented by the upper left and lower right coordinates of a rectangular box. For the set of personnel key point coordinates, a dedicated key point detection model (e.g., HRNet-W48) is used to process the detected target area, and the two-dimensional or three-dimensional coordinate information of each key part of the person (e.g., head, shoulders, elbows, wrists, hips, knees, ankles, etc.) is output. Through foreground segmentation processing, only the image areas related to the person are retained to eliminate background noise. At the same time, the target detection box provides a clear position localization, and the key point coordinates provide spatial data for constructing a three-dimensional skeleton model.

[0033] Furthermore, detailed feature extraction is performed on the regions within each target detection box, thereby converting the visual information of the human targets in the image into high-dimensional numerical vectors (i.e., target feature vectors). The target feature vectors include, but are not limited to, information such as the texture, color, and shape of the person. First, for the image regions within each target detection box in the RGB image data, cropping is performed to extract the image segments containing the target persons. A pre-trained deep neural network (e.g., convolutional neural network) is used to perform convolutional operations on the image segments to extract local features, and then pooling operations are used to reduce the dimension and suppress noise, ultimately converting the original image data into a set of dense feature vectors. At the same time, for the depth map data, similar convolutional networks can also be used to process the depth map regions, and then fuse with the features of the RGB image. This fusion process can adopt simple concatenation, weighted average, or more complex fusion strategies to obtain the fused set of target feature vectors. The set of feature vectors contains the appearance features of the target persons, and each feature vector has a fixed dimension, facilitating subsequent distance measurement and similarity comparison.

[0034] After obtaining the set of target feature vectors, the currently detected target persons are compared with the historical detection records to generate unique target person labels, thereby achieving continuous tracking of the same target persons. The identity association process uses a feature matching method to determine the target identity by calculating the distance or similarity between the current feature vector and the historical feature vectors. Specifically, each vector in the set of target feature vectors is compared with the pre-stored historical feature library, usually using Euclidean distance, cosine similarity, or other measurement methods. In the embodiments of this application, Euclidean distance is used for identity association processing. Exemplarily, assume there are two feature vectors F1 and F2. Among them, F1 comes from the set of target feature vectors, and F2 comes from the historical feature library. The formula for calculating its Euclidean distance is as follows: where d(F1,F2) represents the Euclidean distance between vector F1 and F2. F 1,i represents the i-th component activity in vector F1. n represents the dimension of the feature vector. According to the calculation results, the distances between the current feature vector and all vectors in the historical feature library are sorted. If a certain distance is lower than the predetermined threshold, it is considered that the currently detected target is the same as the target in the historical record, and thus the same identity label is assigned; otherwise, a new identity label is assigned to the current target.

[0035] Meanwhile, perform line skeleton fitting processing on the obtained set of human key point coordinates to construct a three-dimensional human skeleton model. Subsequently, calculate the pose of each joint in the skeleton model to obtain the Euler angle parameters of each joint. The Euler angle parameters are used to describe the rotation state of the joint and are an important basis for judging the human motion pose. First, use geometric modeling methods to connect the key point coordinates into a skeleton structure, and connect the key points in sequence according to the human anatomical structure. For example, connect the head to the shoulders, the shoulders to the elbows, and the elbows to the wrists, etc. to form a series of bone segments. Accurately estimate the spatial positions of each key point to form a three-dimensional skeleton model. This model not only reflects the general shape of the person in the image but also provides a spatial reference for joint angle calculation. When calculating the pose of joint points of the skeleton model, the Euler angle is usually used to represent the rotation angle of the joint. Exemplarily, taking two adjacent bone segments as an example, the vector represents the upper arm direction, and the vector represents the forearm direction, then the angle between the two vectors can be calculated by the dot product formula. The calculation formula for the Euler angle parameters of the joint point is as follows: where θ represents the calculated angle between joint points. and respectively represent the vectors of adjacent bone segments. For each joint in the three-dimensional skeleton model, calculate its rotation angle in sequence through a similar method to form a complete set of Euler angle parameters of joint points. Each Euler angle in this parameter set corresponds to a human joint and reflects its rotation state in space, providing a quantitative basis for subsequent pose stability analysis.

[0036] After obtaining the set of Euler angle parameters of joint points, perform further pose stability analysis on the set of Euler angle parameters of joint points to quantify the balance and stability of the person's movement, so as to provide a more detailed judgment basis for risk warning. The pose stability analysis judges whether there are abnormal rapid changes or uncoordinated movements of the person by comparing the changes of each joint Euler angle in consecutive frames. The specific operation is as follows: First, construct a time series curve of the Euler angle parameters of each joint changing with time, and obtain the stability index of joint movement through smoothing and statistical analysis of the curve. For example, the mean and standard deviation of the change rate of each joint Euler angle in consecutive frames can be calculated. Exemplarily, assume that the Euler angle of the i-th joint at the t-th frame is , then the angle change between adjacent frames can be calculated as . Further, perform statistical analysis on the change amounts of all joints to obtain the overall pose change rate index . The mean of the change rates of all joint Euler angles is used as the pose stability index. The calculation formula for the change rate of joint Euler angle is as follows: Among them, represents the overall attitude change rate index. represents the Euler angle change rate of the i-th joint. N represents the total number of joints (for example, 17). By calculating the average value of the overall joint angle change, the smoothness of the personnel's movement is reflected, and the smaller the value, the more stable the attitude.

[0037] Finally, the attitude change index is compared with a preset stability threshold to determine whether the personnel is in a balanced state. This analysis result is integrated into a complete set of personnel attitude parameters, which includes joint Euler angles, attitude change rates, and other statistical indicators, comprehensively describing the attitude stability of the personnel.

[0038] Step S3: Perform trajectory prediction processing on the set of personnel attitude parameters through a preset physical kinematic model to obtain predicted motion trajectory data.

[0039] It should be understood that the set of personnel attitude parameters contains data on key indicators of the human body in consecutive frames (such as joint angles, center of mass height, motion speed, etc.), and these data record the motion state of the personnel in space in the form of discrete time points. In order to capture the dynamic change law of continuous actions, it is necessary to convert the discrete attitude data into a sequence of hidden state vectors that reflect the temporal dependence relationship between the front and back frames. For this purpose, a bidirectional long short-term memory network is used to process the input attitude parameter sequence. This network structure contains two processing units, forward and backward, which generate a hidden state vector that comprehensively considers the temporal dependence relationship before and after by capturing the motion information in the past and future respectively. The specific operations are as follows: First, normalize the attitude parameter vector S t (where S t includes all joint angles, center of mass height, motion speed, etc. of this frame) to ensure that each parameter is compared under the same dimension. Then, sequentially input the normalized vectors into the bidirectional long short-term memory network. The network internally calculates the forward output and the backward output at each time step, and finally splices the two parts of the output to form the hidden state vector at time t. Convert the discrete attitude parameter sequence into a sequence of hidden state vectors containing temporal information before and after, so as to provide sufficient time-dependent features for subsequent behavior classification and motion trajectory prediction.

[0040] Furthermore, by classifying the hidden state vectors, it is possible to promptly distinguish whether a person is in an abnormal state, thereby determining whether to enter the subsequent motion trajectory prediction stage. Using the sequence of hidden state vectors, classify the continuous motion states to determine the current behavior type of the person (e.g., normal, imbalance, or fall). Map the continuous actions to discrete behavior categories through a pre-trained classification model (e.g., support vector machine or multi-layer perceptron), which facilitates determining whether free fall motion trajectory prediction should be performed in subsequent steps. Specifically, input the hidden state vector corresponding to each time step into the pre-trained classifier. The classifier performs a linear or non-linear mapping on the vector based on the weight parameters obtained during training and outputs the behavior category. The behavior category is usually represented by digital encoding. For example, 0 represents "normal"; 1 represents "imbalance"; 2 represents "fall". During this process, the motion information carried by each hidden state vector enables the classifier to determine the current behavior state of the person. By classifying the hidden state vectors, it is possible to promptly distinguish whether a person is in an abnormal state, thereby determining whether to enter the subsequent motion trajectory prediction stage. After being fully trained, the classification model can accurately distinguish various behavior states, ensuring that the system can respond promptly when detecting a target behavior (such as a fall).

[0041] Furthermore, use a preset physical kinematic model to process the sequence of hidden state vectors, thereby predicting the motion trajectory of a person in a falling state. Only when the behavior classification determines a preset target behavior (i.e., fall), use this prediction model for processing to obtain a trajectory function describing free fall motion. This process first extracts the key physical quantities in the hidden state vector, such as the current center-of-mass height and the initial vertical velocity. These two quantities are the key parameters for free fall motion prediction. According to the law of free fall motion in physical kinematics, the physical kinematic model maps the key parameters in the hidden state vector to generate the following free fall motion trajectory function.

[0042] where Z(t) represents the predicted height of the person at time t. H COM represents the height of the center of mass of the person in the current frame. v Z0 represents the initial vertical velocity, which refers to the velocity of the person in the vertical direction when starting to fall. g represents the acceleration due to gravity. t represents the time variable. Based on the current center-of-mass height and initial vertical velocity of the person, combined with the acceleration due to gravity, predict the height change of the person in the future time to form a free fall motion trajectory function.

[0043] After obtaining the free-fall motion trajectory function, perform time-domain analysis processing on the free-fall motion trajectory function to calculate the collision time data of the person descending from the current state to the preset safe height (or the ground). The collision time data is an important parameter for judging whether a person will have a falling accident in a short time and is also the key basis for risk warning decision-making. Specifically, first determine the safe height threshold, which represents the minimum height considered safe during the movement of the person (for example, the height of the guardrail or the ground height). When the predicted height is less than or equal to the safe height threshold, it is considered that the person touches the safe area or the ground. The following formula is used to calculate the collision time data: where t impact represents the predicted collision time, that is, the time required for the person to reach the safe height H safe . v Z0 represents the initial vertical velocity. g represents the acceleration due to gravity. H COM represents the current height of the person's center of mass. H safe represents the preset safe height threshold. By calculating the time required for the person's free-fall motion trajectory to drop to the safe height, the collision time data is determined, providing a time basis for risk assessment.

[0044] Finally, comprehensively integrate the free-fall motion trajectory function and the collision time data to form complete predicted motion trajectory data. This data not only contains the functional relationship describing the change of the person's height with time in the free-fall state but also contains the key moment information of the collision time, which can provide a comprehensive motion prediction basis for subsequent risk quantification and warning decision-making. Specifically, combine the free-fall motion trajectory function with the calculated collision time to form a data structure containing the motion function and the key time node. This combination process can be implemented through data structures in programming languages (such as dictionaries, structures, etc.), and the two parts of information are saved and transmitted in a unified format.

[0045] Step S4: Real-time obtain the corresponding inertial monitoring data set according to the target person tag set, and perform multi-modal fusion processing on the predicted motion trajectory data and the inertial monitoring data set to obtain the confidence evaluation value corresponding to each target person.

[0046] It should be understood that the people in the target area are all wearing inertial measurement unit (IMU) devices, which are used to collect the acceleration and angular velocity of the target person in real time. According to the label information of each target person, determine the corresponding IMU device identifier, and use this identifier to obtain the corresponding inertial monitoring data in real time. Among them, the target person tag set can uniquely identify each detected person in the image. The specific operation is as follows: First, construct a predefined lookup table that pre-stores the correspondence between target person tags and inertial measurement unit tags. This lookup table may exist in the form of a database, a configuration file, or a hard-coded array, where each record contains two fields: the target person tag and the corresponding IMU device identifier. After the system receives the target person tag, it obtains the corresponding IMU tag through table lookup processing. Through table lookup processing, it is ensured that for each target person, the source device of their inertial monitoring data can be accurately identified. Next, based on the obtained IMU device identifier, the system collects data from the corresponding IMU device in real time through a communication interface (such as RS-485, CAN bus, or wireless transmission protocol). The IMU device usually outputs acceleration, angular velocity, and other motion-related parameters periodically, forming an inertial monitoring data set. This data set contains inertial measurement values within a continuous time and marks each data frame with a timestamp to synchronize with the motion prediction information in the visual data.

[0047] Furthermore, convert the predicted motion trajectory data obtained from the visual system and the inertial monitoring data obtained from the IMU sensor into state vectors respectively. The construction process of the state vector aims to convert the raw data from different sources into a numerical representation in a unified format, so that subsequent data fusion processing such as Kalman filtering can be performed under the same data structure. The visual state vector usually includes parameters such as position and velocity predicted from the motion trajectory; while the inertial state vector is composed of information such as acceleration and angular velocity in the IMU data.

[0048] First, perform parsing processing on the predicted motion trajectory data, which usually consists of a motion trajectory function and key time nodes (such as the collision time). By sampling the motion trajectory function, corresponding position, velocity, and other information are obtained at different time points to construct a visual state vector of a continuous time series. For example, if the motion trajectory function is z(t), it can be sampled at a fixed time interval to obtain the position z(t i ) and the velocity v(t i ) obtained by differential calculation. This information constitutes the visual state vector X v .

[0049] At the same time, process the inertial monitoring data set. Each data frame in the inertial data set usually contains acceleration a(t) and angular velocity w(t) data. After performing necessary filtering and correction (such as low-pass filtering) on this data to eliminate noise interference, each frame of data is represented as an inertial state vector X imu , whose dimensions include acceleration components and angular velocity components in each axis. The visual state vector can be represented in a discrete sampling manner as follows: Among them, X v (t i ) represents the visual state vector formed at time t i . Z(t i ) represents the predicted position at time t i . Z(t i-1 ) represents the predicted position at the previous time point t i-1 . represents the time interval between two consecutive sampling time points. represents the velocity calculated by difference. Converting the predicted motion trajectory data into a visual state vector represents position and velocity information, facilitating unified data processing with the inertial state vector.

[0050] Similarly, the inertial state vector can be expressed as follows: Among them, each component represents the acceleration and angular velocity of each axis collected by the IMU within the corresponding time t i .

[0051] By converting the two data sources into state vectors respectively, the unified data representation form can be obtained, providing a structured input for subsequent data fusion algorithms (such as Kalman filtering). The visual state vector reflects the motion information predicted by the camera, while the inertial state vector provides the motion data directly collected by the IMU. The two are complementary, and more accurate motion state estimation can be obtained after fusion. Among them, the visual state vector set and the inertial state vector set are both arranged in chronological order, and the two are time-aligned.

[0052] To eliminate the errors of the two sensor data and fuse the two data, the Kalman filtering algorithm is used to perform multimodal fusion processing on the visual state vector set and the inertial state vector set. During the fusion process, the errors of the two sensor data are dynamically corrected, so as to obtain a more accurate motion state estimation. Among them, Kalman filtering is a mathematical method based on linear system state estimation, which can perform recursive optimal estimation on data disturbed by noise and output the state estimation and its covariance matrix. The specific operations are as follows: First, set the state transition model and the observation model. The state vector x tRepresents the fused motion state at time t, which may contain information such as position and velocity. The state transition model describes the law of state change over time, and the observation model describes how to observe the state vector from the visual state vector and the inertial state vector. The core steps of Kalman filtering include prediction (i.e., using the state transition model to predict the current state and covariance) and update (i.e., using the actual observation data to correct the prediction result). In the prediction stage, prediction is carried out based on the state estimate and the state transition matrix at the previous moment, and the predicted state covariance matrix is calculated simultaneously. The update of the state estimate in the prediction stage of Kalman filtering can be expressed by the following formula: Where, Represents the predicted state estimate at time t. Represents the state estimate at the previous moment. F represents the state transition matrix, which describes the linear change of the state over time. P t|t-1 Represents the predicted state covariance matrix, which reflects the uncertainty of the predicted state. P t-1 Represents the state covariance matrix at the previous moment. F T Represents the transpose of matrix F. Q represents the process noise covariance matrix, which describes the influence of the internal noise of the system. Predicting the state and its covariance through the state transition model lays the foundation for subsequent correction using observation data.

[0053] In the update stage, the fused observation data (combining the visual state vector and the inertial state vector) is used to update the state estimate and the covariance matrix, calculate the Kalman gain and perform state correction. The correction of the predicted state in the update stage of Kalman filtering can be expressed by the following formula: Where, K t Represents the Kalman gain calculated at time t. H represents the observation matrix, which maps the state vector to the observation space. R represents the observation noise covariance matrix, which describes the noise from the sensor observation data. Z t Represents the observation data at time t, which combines visual and inertial data. I represents the identity matrix. Represents the updated state estimate at time t. P t Represents the updated state covariance matrix. (.) -1 Represents matrix inversion. H T Represents the transpose of the observation matrix. Correcting the predicted state according to the actual observation data and calculating the updated state covariance matrix reflect the uncertainty of the fused state estimate.

[0054] After Kalman filtering, the output result is the predicted state covariance matrix, which details the uncertainty of the fused state estimate and is an important intermediate variable for subsequent calculation of the confidence evaluation value. The reason for Kalman filtering is that by fusing visual and inertial data, their respective advantages can be utilized and compensated for each other, thereby obtaining a more accurate and stable motion state estimate. Using the Kalman filter algorithm can update the state estimate in real time and automatically adjust the tolerance to noise, providing a rigorous mathematical basis for multimodal data fusion.

[0055] After the Kalman filtering is completed, the predicted state covariance matrix, as a key indicator representing the uncertainty of the state estimate, has diagonal elements that reflect the estimation errors of each state component. To convert this matrix into a single value for subsequent confidence calculation, this step uses the trace value calculation method, which is to sum all the diagonal elements of the matrix. The larger the trace value, the greater the state estimation error and the higher the uncertainty; conversely, the uncertainty is lower. The specific operation is as follows: Sum all the diagonal elements in the predicted state covariance matrix to obtain the state covariance trace value. The trace value calculation formula is as follows: where, tr(P t ) represents the trace value of matrix P t . P t (i,i) represents the element in the i-th row and i-th column of matrix P t , that is, the variance of the -th state component. n represents the dimension of the state vector. Converting the uncertainty information in matrix form into a scalar facilitates subsequent comparison with the preset maximum trace value and further calculation of the normalized confidence value.

[0056] The single value obtained by summation can intuitively reflect the uncertainty of the entire state estimate and provide a quantitative basis for risk assessment. The state covariance trace value, as a comprehensive indicator of the Kalman filter output result, plays an important role in the multimodal data fusion process. The lower its value, the more accurate the fused state estimate and the higher the system confidence.

[0057] Finally, calculate the ratio of the state covariance trace value to the preset maximum state covariance trace value to obtain a normalized confidence value, which ranges between 0 and 1 and is used to quantitatively describe the reliability of the fused state estimate. Among them, the maximum state covariance trace value is the maximum tolerance error value determined based on historical data or experimental calibration. When the actual trace value is small, the normalized confidence value will approach 1; conversely, it will approach 0. The ratio calculation formula is as follows: where, ρ trepresents the normalized confidence value. tr(P t ) represents the trace value of the current predicted state covariance matrix. tr(P max ) represents the preset maximum state covariance trace value. represents the ratio of the state covariance trace value. "1 -" means converting the ratio value to a confidence level. The smaller the ratio value, the higher the confidence level. Converting the state covariance trace value to a normalized confidence value enables the reliability of the fused state estimate to be represented by a value within a fixed range, facilitating subsequent decision-making and judgment by the system.

[0058] Subsequently, the normalized confidence value is combined with other state estimate information through data combination processing to form the final confidence evaluation value for each target person. The process of data combination processing can adopt a simple data packaging method, associating the confidence value with the target person label to form a complete evaluation result data structure. Each target person in this structure corresponds to a confidence evaluation value, indicating the reliability of the state estimate of this person and providing a basis for risk early warning.

[0059] Step S5: Perform risk quantification processing on the confidence evaluation value and the predicted motion trajectory data to obtain a hierarchical early warning instruction corresponding to each target person.

[0060] It should be understood that the predicted motion trajectory data includes but is not limited to behavior categories and collision time data. The behavior category directly reflects the current action state of the person, and the collision time reveals the time distance between the person and the safety boundary, while the confidence evaluation value represents the reliability of the multi-modal data fusion result. Combining these three pieces of data into a vector can integrate various risk indicators in a quantitative manner, ensuring that the system can use unified input data for mathematical modeling and calculation in subsequent steps. First, for each target person, the above three pieces of data are arranged and combined in a predetermined order to form a unified input feature vector X = [C, t impact , ρ], where C represents the behavior category, usually a discrete value (e.g., 0, 1, 2) representing different states. t impact represents the collision time data, reflecting the key time nodes of the person in free fall motion. ρ represents the confidence evaluation value.

[0061] Further, the unified input feature vector is mapped into a risk index through a mathematical model, and the early warning level is determined based on the comparison result between the risk index and the preset risk threshold. The main purpose of using the exponential function is to perform a non-linear mapping on the input features, enabling the system to more sensitively reflect the comprehensive impact of each input index on the risk index, thereby realizing the quantification of potential risks. To this end, the exponential function in the logistic regression model is introduced. This model linearly combines the input features and, through the mapping of the exponential function, outputs a risk index between 0 and 1. The calculation formula of the exponential function is as follows: where R represents the calculated risk index, and its value range is from 0 to 1. β0, β1, β2, β3 represent the weight coefficients determined through training with historical data, and each coefficient reflects the influence degree of the corresponding input variable on the risk index. C represents the behavior category. t impact represents the collision time data. ρ represents the confidence evaluation value. exp(.) represents the natural exponential function, with its base being the mathematical constant e. The input features are linearly combined and then mapped to a non-linear risk index, reflecting the comprehensive impact of each input factor on the overall risk. The non-linear characteristic of the exponential function helps the risk index to quickly respond to changes when there are large changes in the input data, thereby realizing the sensitive detection of dangerous situations.

[0062] Subsequently, the calculated risk index is compared with the preset risk threshold to determine the preliminary early warning level. Usually, the following rules are set: If R ≥ 0.8, the early warning level is set to 1 (indicating the highest risk); if 0.6 ≤ R < 0.8, the early warning level is set to 2 (indicating medium risk); if R < 0.6, the early warning level is set to 3 (indicating low risk). By presetting the risk threshold, the continuous risk index can be converted into a discrete early warning level, facilitating the system to issue corresponding control instructions according to the early warning level in subsequent steps.

[0063] Further, through conditional judgment on the collision time data and the confidence evaluation value, an adjustment flag is generated to correct the early warning level. This processing step is based on an important premise: when it is detected that a person will reach the collision threshold in an extremely short time and the data fusion result has a high confidence, the highest-level early warning measures should be taken preferentially to shorten the response time and strengthen the safety intervention. In the specific implementation process, the following judgment rules are set: If t impact <T crit and ρ > ρ crit , then the adjustment flag is generated as 1; otherwise, the adjustment flag is 0. Where t impact represents the collision time data. ρ represents the confidence evaluation value. T critis a preset critical value of the collision time, for example, set to 2 seconds. ρ crit is a preset confidence threshold, for example, set to 0.7. By judging the key parameters, the risk level can be corrected in an emergency to ensure that the system has the highest response priority for serious risks (such as a person about to touch the ground and the sensing data is very reliable).

[0064] It should be understood that the early warning system is required to be able to respond quickly in an emergency and raise the early warning level to the highest level in order to trigger the strongest safety intervention measures. Combine the generated adjustment identifier with the early warning level and correct the early warning level through conditional correction to obtain the final early warning level. The core idea of the conditional correction process is: when an emergency condition is detected (that is, the adjustment identifier is 1), regardless of the early warning level, the early warning level is corrected to the highest level (that is, early warning level 1); when there is no emergency condition, the early warning level remains unchanged.

[0065] Finally, converting the abstract risk assessment result into a specific electrical signal is the key link to realize the transition from risk perception to physical intervention. By presetting the mapping rules, the system can output control signals of different levels according to different risk levels, so as to drive the corresponding devices to make safety responses. The control signal mapping process converts the abstract early warning level into an electrical signal or communication message that can be recognized by the actuator, ensuring that the system can quickly trigger physical intervention measures when a risk is detected (for example, triggering an electromagnetic lock, starting air cushion inflation, sending an alarm message, etc.).

[0066] First, preset a set of mapping rules to correspond different early warning levels with the corresponding control signal parameters one by one. For example, early warning level 1 corresponds to the highest-level control signal, which may include high-voltage PWM pulses, emergency CAN instructions, and Modbus-TCP messages containing detailed alarm information; early warning level 2 corresponds to the medium-level control signal with secondary parameters; while early warning level 3 corresponds to the low-risk state, only logging is required and no physical intervention needs to be activated.

[0067] After obtaining the final early warning level, look up the preset control signal mapping table according to the final early warning level. The mapping table clearly stipulates the data format and parameters of the control signal corresponding to each early warning level. Then, convert the data format corresponding to the early warning level into a specific physical control signal. For example, for early warning level 1, generate a combined signal containing the PWM pulse parameters required for triggering the electromagnetic lock (such as pulse width, frequency, current voltage), CAN bus data frame (including air cushion inflation duration and target pressure), and Modbus-TCP message (including device ID, timestamp, early warning level information). Finally, through comprehensive editing and processing of the above various signals, integrate them into a final control signal output that can drive the relevant components in the fall prevention device to work simultaneously.

[0068] Step S6: Perform protocol encoding and signal conversion processing on the hierarchical warning instruction according to a preset conversion method to obtain a control signal for driving the fall prevention device.

[0069] Perform a preliminary analysis on the generated hierarchical warning instruction, classify it according to the preset device type, and map the general warning instruction to device-specific instruction data according to the specific communication requirements of each device. The hierarchical warning instruction contains key information such as risk level, collision time, confidence level, etc. However, different fall prevention devices (such as electromagnetic locks, air cushion devices, central monitoring units, etc.) have different requirements for the format of the instruction data. Therefore, it is necessary to identify the device type and convert the format of the hierarchical warning instruction. First, by looking up a table or using a preset mapping relationship, analyze the identification field in the input hierarchical warning instruction to extract the target device type identifier. This device type identifier is a sequence of numbers or characters used to uniquely identify the device type to be controlled by the warning instruction. When looking up the table, the system compares the parameters carried in the hierarchical warning instruction according to the preset mapping table to determine the output device type identifier. This process ensures that instructions for different devices can be correctly distinguished, avoiding incorrect instruction transmission or device response confusion.

[0070] After the device type identification is completed, the format mapping process is carried out next. This processing step converts the general hierarchical warning instruction data into the data format required by each target device according to preset rules. The format mapping process includes operations such as data structure adjustment, data domain rearrangement, and necessary parameter unit conversion. For example, for a device controlling an electromagnetic lock, its instruction may need to include parameters such as voltage, current, pulse width, and pulse frequency, while for an air cushion device, it may need to include parameters such as inflation duration and target pressure. To achieve this conversion, the system has pre-designed a set of format mapping rules to map each parameter in the general instruction to the device-specific parameters one by one. For example, assuming that the risk level L, collision time t0, and confidence level ρ in the warning instruction are input parameters, for the electromagnetic lock device, the mapping rule maps the risk level L to the pulse width W in the control signal, the collision time t0 to the pulse frequency f, and the confidence level ρ to the voltage V; this mapping process is achieved by looking up a table or function mapping, and the output is the device-specific instruction data.

[0071] Furthermore, after obtaining the dedicated instruction data required by each target device, these device-specific instruction data are converted into electrical signals at the physical level, that is, the digitized control parameters are converted into electrical signals that can drive the actual hardware actuators. The electrical signal conversion process includes operations such as digital-to-analog conversion, pulse signal generation, and signal modulation. The purpose is to convert the warning information into specific physical quantities such as voltage, current, pulse width, and frequency to meet the hardware interface requirements of each fall prevention device. First, the device-specific instruction data are used as input, and the digital signal is converted into an analog electrical signal through a dedicated digital-to-analog conversion module. During the digital-to-analog conversion process, the system converts the input digital control parameters into corresponding voltage values or current values according to the preset scale factor and bias. This process is strictly carried out in accordance with industrial standards to ensure that the output signal has a stable amplitude and low-noise characteristics.

[0072] In addition to the basic digital-to-analog conversion, for different devices, it is also necessary to generate pulse signals or perform specific waveform modulation on the converted signals. For example, for electromagnetic lock control, it is usually required to output a PWM (pulse width modulation) signal, which controls the triggering time and intensity of the device by changing the pulse width; for the air cushion device, it may be required to output an electrical signal in the form of a data frame that conforms to the CAN bus protocol. For these different requirements, the electrical signal conversion process will perform different modulation methods on the signal according to the different device-specific instruction data. For example, by setting fixed sampling frequency and pulse width parameters, the continuous signal after digital-to-analog conversion is discretized into a PWM signal; or a dedicated modulation circuit is used to encode the digital signal into a CAN bus data frame.

[0073] After obtaining the electrical signal set, it is classified and encoded into a communication control signal set that meets the requirements of various industrial communication protocols through a preset conversion method. Different anti-falling devices have different requirements for communication protocols. For example, an electromagnetic lock may directly control using PWM signals, while an air cushion device may be controlled through a CAN bus, and the central monitoring system reports alarm information using the Modbus-TCP protocol. Therefore, classifying and encoding the electrical signal set according to the protocol can make the generated control signals meet the requirements of multiple device interfaces and achieve unified management of multiple communication protocols. First, according to the device type identification information, the electrical signal set is classified according to the protocol. The preset conversion method includes the communication protocol parameters corresponding to various devices, and the system groups the electrical signals according to these parameters. For example, the PWM signals for the electromagnetic lock are grouped separately, the CAN bus data frames for the air cushion device are grouped into another group, and the signals for monitoring information reporting are grouped into the Modbus-TCP group. After classification, each group of signals is processed using the corresponding encoding method. The encoding process includes operations such as data encapsulation, checksum addition, and head and tail identification addition to ensure that the generated communication control signals not only meet the protocol specifications but also have anti-interference capabilities. Similarly, for PWM signals, Modbus-TCP messages, etc., the system processes them using the corresponding encoding functions respectively.

[0074] Finally, the communication control signal set (such as PWM signals, CAN bus data frames, Modbus-TCP messages, etc.) is comprehensively edited and integrated to form a complete final control signal that can directly drive the anti-falling device. The signal comprehensive editing process aims to coordinate and schedule multiple protocol signals so that each part of the signal can work together to achieve synchronous control of the anti-falling device. The comprehensive editing process not only includes signal merging but also unified scheduling and correction of aspects such as timing, priority, and data consistency. First, the various signals in the communication control signal set are synchronized, and each signal is aligned according to a preset time window to ensure that the control signals received by each device at the same moment have consistent timestamps. The timing synchronization process uses a timer and buffer mechanism to ensure that all signals are accurately time-corrected before transmission. Then, the various signals are sorted according to priority. Usually, when an extreme risk situation is detected, the priority of the electromagnetic lock trigger signal and the air cushion inflation command should be higher than the monitoring data reporting signal. According to the preset priority strategy, the signal data is sorted and merged. For this purpose, methods such as weighted average or priority override can be used to merge the key parameters of each signal into a comprehensive control data packet. In the signal comprehensive editing process, it is also necessary to perform a final standardization process on the formats of each signal to ensure that the finally generated control signal can be directly transmitted to the interface of the driving device.

[0075] This application is applied to the field of image recognition technology. By performing multi-modal preprocessing on the real-time image acquisition data set, an RGB-D data stream is obtained. Target detection and three-dimensional pose modeling are performed on the RGB-D data stream to obtain a target person label set and a person pose parameter set. A trajectory prediction is performed on the person pose parameter set through a physical kinematic model to obtain predicted motion trajectory data. An inertial monitoring data set is obtained in real time according to the target person label set, and multi-modal fusion is performed in combination with the predicted motion trajectory data to obtain a confidence evaluation value. The confidence evaluation value and the predicted motion trajectory data are used for risk quantification to obtain a graded early warning instruction, and protocol encoding and signal conversion are performed on the graded early warning instruction to obtain a control signal. Through image recognition, motion prediction, multi-modal data fusion, and refined processing, this application improves the accuracy of anti-fall monitoring in complex scenarios, and at the same time provides a powerful control signal for anti-fall devices, which is an important technical means to achieve intelligent safety protection.

[0076] As Figure 2 shown, it is a functional module diagram of an anti-fall recognition and early warning device based on image recognition provided by an embodiment of this application.

[0077] In some embodiments, the anti-fall recognition and early warning device 2 based on image recognition may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the anti-fall recognition and early warning device 2 based on image recognition may be stored in the memory of the server and executed by at least one processor to execute (see details in Figure 1 the description) the functions of the anti-fall recognition and early warning method based on image recognition.

[0078] In this embodiment, the anti-fall recognition and early warning device 2 based on image recognition can be divided into multiple functional modules according to the functions it performs. The functional modules may include: an image acquisition module 21, a target detection module 22, a trajectory prediction module 23, a confidence evaluation module 24, a graded early warning module 25, and an anti-fall control module 26. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0079] The image acquisition module 21 is used to acquire a real-time image acquisition data set of a target area and perform multi-modal preprocessing on the real-time image acquisition data set to obtain a standardized RGB-D data stream.

[0080] In an optional implementation manner, the image acquisition module 21 is specifically used for: Performing denoising processing on the original RGB image and the original depth map to obtain a noise-optimized RGB image and a noise-optimized depth map; Perform white balance correction processing on the noise-optimized RGB image to obtain a color-balanced optimized RGB image; Perform depth compensation processing on the noise-optimized depth map to obtain a compensated optimized depth map; Perform pixel alignment processing on the color-balanced optimized RGB image and the compensated optimized depth map to obtain initial RGB-D data; Perform histogram equalization and scale normalization processing on the initial RGB-D data to obtain scale-optimized RGB-D data; Perform format conversion processing on the scale-optimized RGB-D data to obtain a standardized RGB-D data stream.

[0081] The target detection module 22 is used to perform target detection and three-dimensional pose modeling processing on the RGB-D data stream to obtain a target person label set and a person pose parameter set.

[0082] In an optional embodiment, the target detection module 22 is specifically used for: Perform foreground segmentation processing on the RGB-D data stream to obtain foreground region data containing people, and perform target detection processing on the foreground region data to obtain a target detection box set and a set of person key point coordinates; Perform feature extraction processing on the RGB-D data stream according to the target detection box set to obtain a target feature vector set; Perform identity association processing according to the target feature vector set to obtain the target person label set; Perform skeleton fitting processing according to the set of person key point coordinates to obtain a three-dimensional human skeleton model, and perform joint point pose calculation processing on the three-dimensional human skeleton model to obtain a set of joint point Euler angle parameters; Perform pose stability analysis processing according to the set of joint point Euler angle parameters to obtain the person pose parameter set.

[0083] The trajectory prediction module 23 is used to perform trajectory prediction processing on the person pose parameter set through a preset physical kinematic model to obtain predicted motion trajectory data.

[0084] In an optional embodiment, the trajectory prediction module 23 is specifically used for: Perform time series feature extraction processing on the person pose parameter set to obtain a sequence of hidden state vectors; Perform behavior classification processing according to the sequence of hidden state vectors to obtain a behavior category; When the behavior category is a preset target behavior, perform a fall motion prediction process on the implicit state vector sequence through the physical kinematics model to obtain a free-fall motion trajectory function corresponding to each target person; Perform a time-domain analysis process on the free-fall motion trajectory function to obtain collision time data; Perform a combination process according to the free-fall motion trajectory function and the collision time data to obtain the predicted motion trajectory data.

[0085] The confidence evaluation module 24 is used to obtain the corresponding inertial monitoring data set in real time according to the target person label set, and perform a multi-modal fusion process on the predicted motion trajectory data and the inertial monitoring data set to obtain a confidence evaluation value corresponding to each target person.

[0086] In an optional implementation manner, the confidence evaluation module 24 is specifically used for: Perform a look-up table process on the target person label set to obtain an inertial measurement unit label set, and obtain the corresponding inertial monitoring data set in real time according to the inertial measurement unit label set; Perform a state vector construction process on the predicted motion trajectory data and the inertial monitoring data set to obtain a visual state vector set and an inertial state vector set; Perform a Kalman filtering process on the visual state vector set and the inertial state vector set to obtain a predicted state covariance matrix; Perform a trace value calculation process on the predicted state covariance matrix to obtain a state covariance trace value; Perform a ratio calculation process on the state covariance trace value according to a preset maximum state covariance trace value to obtain a normalized confidence value, and perform a data combination process on the normalized confidence value to obtain a confidence evaluation value corresponding to each target person.

[0087] The hierarchical early warning module 25 is used to perform a risk quantification process on the confidence evaluation value and the predicted motion trajectory data to obtain a hierarchical early warning instruction corresponding to each target person.

[0088] In an optional implementation manner, the hierarchical early warning module 25 is specifically used for: Perform a feature vector component process on the behavior category, the collision time data, and the confidence evaluation value to obtain an input feature vector; Perform an exponential function calculation process on the input feature vector to obtain a risk index, and perform a threshold comparison process on the risk index according to a preset risk threshold to obtain an early warning level; Perform conditional judgment processing on the collision time data and the confidence evaluation value to obtain an adjustment identifier; Perform conditional correction processing on the warning level according to the adjustment identifier to obtain the final warning level; Perform control signal mapping processing on the final warning level to obtain a hierarchical warning instruction corresponding to each target person.

[0089] The anti-falling control module 26 is configured to perform protocol encoding and signal conversion processing on the hierarchical warning instruction according to a preset conversion method to obtain a control signal for driving the anti-falling device.

[0090] In an alternative embodiment, the anti-falling control module 26 is specifically configured to: Perform device type identification processing on the hierarchical warning instruction to obtain a target device type identifier, and perform format mapping processing on the hierarchical warning instruction according to the target device type identifier to obtain device-specific instruction data; Perform electrical signal conversion processing on the device-specific instruction data to obtain an electrical signal set; Perform protocol classification and encoding processing on the electrical signal set according to the conversion method to obtain a communication control signal set; Perform signal comprehensive editing processing on the communication control signal set to obtain a control signal for driving the anti-falling device.

[0091] It should be understood that the various change methods and specific embodiments in the methods provided in the above embodiments are equally applicable to the anti-falling identification and warning device based on image recognition in this embodiment. Through the foregoing detailed description of the anti-falling identification and warning method based on image recognition, those skilled in the art can clearly know the implementation method of the anti-falling identification and warning device based on image recognition in this embodiment. For the sake of simplicity of the specification, it will not be elaborated herein.

[0092] As Figure 3 shown, it is a schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0093] In a preferred embodiment of the present invention, the electronic device 3 may include, but is not limited to: a memory 31, at least one processor 32, and at least one communication bus 33.

[0094] Those skilled in the art should understand that Figure 3 the structure of the electronic device 3 shown does not constitute a limitation on the embodiments of the present invention. The electronic device 3 may further include more or fewer other hardware or software than shown, or different component arrangements.

[0095] In some embodiments, the electronic device 3 is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, a microprocessor, an application specific integrated circuit, a programmable gate array, a digital signal processor, and an embedded device, etc.

[0096] It should be noted that the electronic device 3 is only an example, and other existing or future electronic products that can be adapted to this application should also be included within the protection scope of this application and are incorporated herein by reference.

[0097] In some embodiments, a computer program is stored in the memory 31, and when the computer program is executed by the at least one processor 32, all or part of the steps in the above-described anti-falling recognition and warning method based on image recognition are implemented. The memory 31 includes a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, magnetic tape memories, or any other computer-readable medium capable of carrying or storing data. Further, the computer-readable storage medium mainly includes a storage program area and a storage data area. Among them, the storage program area can store an operating system, application programs required for at least one function, etc.

[0098] In some embodiments, the at least one processor 32 is the control core (Control Unit) of the electronic device 3, connecting various components of the entire electronic device 3 through various interfaces and circuits. By running or executing programs or modules stored in the memory 31, and by invoking data stored in the memory 31, it performs various functions of the electronic device 3 and processes data. For example, when the at least one processor 32 executes the computer program stored in the memory 31, it implements all or part of the steps of the anti-fall recognition and warning method based on image recognition described in the embodiments of the present application; or implements all or part of the functions of the anti-fall recognition and warning device based on image recognition. The at least one processor 32 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions packaged together, including a combination of one or more central processing units (Central Processing Unit, CPU), microprocessors, digital processing chips, graphics processors, and various control chips, etc.

[0099] In some embodiments, the at least one communication bus 33 is configured to enable connection communication between the memory 31 and the at least one processor 32, etc. Although not shown, the electronic device 3 may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to the at least one processor 32 through a power management device, so as to implement functions such as management of charging, discharging, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device 3 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0100] The above-mentioned integrated unit implemented in the form of software function modules can be stored in a computer-readable storage medium. The above-mentioned software function modules are stored in a storage medium, including several instructions for causing an electronic device (which may be a personal computer, an electronic device, or a network device, etc.) or a processor to execute part of the methods described in the various embodiments of the present application.

[0101] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0102] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical unit, and it may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0103] The above are all preferred embodiments of this application. The protection scope of this application is not limited thereby. Therefore, any equivalent changes made according to the structure, shape, and principle of this application shall be covered within the protection scope of this application.

Claims

1. A fall prevention recognition and warning method based on image recognition, characterized in that, The method includes: Obtaining a real-time image acquisition dataset of a target area, and performing multi-modal preprocessing on the real-time image acquisition dataset to obtain a standardized RGB-D data stream; Performing target detection and three-dimensional pose modeling processing on the RGB-D data stream to obtain a target personnel label set and a personnel pose parameter set; Performing trajectory prediction processing on the personnel pose parameter set through a preset physical kinematic model to obtain predicted motion trajectory data; According to the target personnel label set, obtaining a corresponding inertial monitoring dataset in real time, and performing multi-modal fusion processing on the predicted motion trajectory data and the inertial monitoring dataset to obtain a confidence evaluation value corresponding to each target personnel; Performing risk quantification processing on the confidence evaluation value and the predicted motion trajectory data to obtain a hierarchical early warning instruction corresponding to each target personnel; Performing protocol encoding and signal conversion processing on the hierarchical early warning instruction according to a preset conversion method to obtain a control signal for driving a fall prevention device.

2. The anti-falling recognition and warning method based on image recognition according to claim 1, characterized in that, The real-time image acquisition dataset includes an original RGB image and an original depth map, and the performing multi-modal preprocessing on the real-time image acquisition dataset to obtain a standardized RGB-D data stream includes: Performing denoising processing on the original RGB image and the original depth map to obtain a noise-optimized RGB image and a noise-optimized depth map; Performing white balance correction processing on the noise-optimized RGB image to obtain a color-balanced optimized RGB image; Performing depth compensation processing on the noise-optimized depth map to obtain a compensated optimized depth map; Performing pixel alignment processing on the color-balanced optimized RGB image and the compensated optimized depth map to obtain initial RGB-D data; Performing histogram equalization and scale normalization processing on the initial RGB-D data to obtain scale-optimized RGB-D data; Performing format conversion processing on the scale-optimized RGB-D data to obtain a standardized RGB-D data stream.

3. The anti-falling recognition and warning method based on image recognition according to claim 1, characterized in that, The performing target detection and three-dimensional pose modeling processing on the RGB-D data stream to obtain a target personnel label set and a personnel pose parameter set includes: Performing foreground segmentation processing on the RGB-D data stream to obtain foreground area data containing personnel, and performing target detection processing on the foreground area data to obtain a target detection box set and a personnel key point coordinate set; Performing feature extraction processing on the RGB-D data stream according to the target detection box set to obtain a target feature vector set; Performing identity association processing according to the target feature vector set to obtain the target personnel label set; Performing skeleton fitting processing according to the personnel key point coordinate set to obtain a three-dimensional human skeleton model, and performing joint point pose calculation processing on the three-dimensional human skeleton model to obtain a joint point Euler angle parameter set; Performing pose stability analysis processing according to the joint point Euler angle parameter set to obtain the personnel pose parameter set.

4. The anti-fall recognition and warning method based on image recognition according to claim 1, wherein Performing trajectory prediction processing on the set of human posture parameters through a preset physical kinematic model to obtain predicted motion trajectory data includes: Performing temporal feature extraction processing on the set of human posture parameters to obtain a sequence of hidden state vectors; Performing behavior classification processing according to the sequence of hidden state vectors to obtain a behavior category; When the behavior category is a preset target behavior, performing free fall motion prediction processing on the sequence of hidden state vectors through the physical kinematic model to obtain a free fall motion trajectory function corresponding to each target person; Performing time-domain analysis processing on the free fall motion trajectory function to obtain collision time data; Performing combination processing according to the free fall motion trajectory function and the collision time data to obtain the predicted motion trajectory data.

5. The anti-falling recognition and warning method based on image recognition according to claim 1, characterized in that The step of obtaining the confidence evaluation value corresponding to each target person by obtaining the corresponding inertial monitoring data set according to the target person label set in real time and performing multi-modal fusion processing on the predicted motion trajectory data and the inertial monitoring data set includes: Performing a look-up table process on the target person label set to obtain an inertial measurement unit label set, and obtaining the corresponding inertial monitoring data set in real time according to the inertial measurement unit label set; Performing state vector construction processing on the predicted motion trajectory data and the inertial monitoring data set to obtain a visual state vector set and an inertial state vector set; Performing Kalman filtering processing on the visual state vector set and the inertial state vector set to obtain a predicted state covariance matrix; Performing a trace value calculation process on the predicted state covariance matrix to obtain a state covariance trace value; Performing a ratio calculation process on the state covariance trace value according to a preset maximum state covariance trace value to obtain a normalized confidence value, and performing data combination processing on the normalized confidence value to obtain the confidence evaluation value corresponding to each target person.

6. The anti-falling recognition and warning method based on image recognition according to claim 4, characterized in that, The predicted motion trajectory data includes a behavior category and collision time data. The step of performing risk quantification processing on the confidence evaluation value and the predicted motion trajectory data to obtain a hierarchical warning instruction corresponding to each target person includes: Performing feature vector component processing on the behavior category, the collision time data, and the confidence evaluation value to obtain an input feature vector; Performing an exponential function calculation process on the input feature vector to obtain a risk index, and performing a threshold comparison process on the risk index according to a preset risk threshold to obtain a warning level; Performing a conditional judgment process on the collision time data and the confidence evaluation value to obtain an adjustment flag; Performing conditional correction processing on the warning level according to the adjustment flag to obtain a final warning level; Performing a control signal mapping process on the final warning level to obtain a hierarchical warning instruction corresponding to each target person.

7. The anti-falling recognition and warning method based on image recognition according to claim 1, characterized in that The step of performing protocol encoding and signal conversion processing on the hierarchical warning instruction according to a preset conversion method to obtain a control signal for driving a fall prevention device includes: Perform device type identification processing on the hierarchical warning instruction to obtain a target device type identifier, and perform format mapping processing on the hierarchical warning instruction according to the target device type identifier to obtain device-specific instruction data; Perform electrical signal conversion processing on the device-specific instruction data to obtain an electrical signal set; Perform protocol classification and encoding processing on the electrical signal set according to the conversion method to obtain a communication control signal set; Perform signal comprehensive editing processing on the communication control signal set to obtain a control signal for driving the anti-falling device.

8. An anti-falling recognition and warning device based on image recognition, characterized in that, The device includes: An image acquisition module, configured to acquire a real-time image acquisition data set of a target area, and perform multi-modal preprocessing on the real-time image acquisition data set to obtain a standardized RGB-D data stream; A target detection module, configured to perform target detection and three-dimensional pose modeling processing on the RGB-D data stream to obtain a target person label set and a person pose parameter set; A trajectory prediction module, configured to perform trajectory prediction processing on the person pose parameter set through a preset physical kinematic model to obtain predicted motion trajectory data; A confidence evaluation module, configured to obtain a corresponding inertial monitoring data set in real time according to the target person label set, and perform multi-modal fusion processing on the predicted motion trajectory data and the inertial monitoring data set to obtain a confidence evaluation value corresponding to each target person; A hierarchical warning module, configured to perform risk quantification processing on the confidence evaluation value and the predicted motion trajectory data to obtain a hierarchical warning instruction corresponding to each target person; An anti-fall control module, configured to perform protocol encoding and signal conversion processing on the hierarchical warning instruction according to a preset conversion method to obtain a control signal for driving the anti-falling device.

9. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the anti-falling identification and warning method based on image recognition according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the anti-falling identification and warning method based on image recognition according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • High-altitude operation anti-falling visual monitoring and early warning system and method

    CN120472642A

  • Electric power engineering construction safety early warning method and system

    CN120932363A

  • A power engineering construction safety early warning method and system

    CN120932363B

  • Power lithium battery multi-parameter automatic sorting and conveying system based on visual identification

    CN121266850A

  • A vision-based automated sorting and conveying system for multi-parameter lithium-ion batteries.

    CN121266850B