Personnel status recognition method and system based on target recognition detection
By parameterizing the original image data, intelligent BIM data is generated, combined with multi-disciplinary data exchange and fuzzy logic evaluation, the accuracy and real-time problems of personnel state recognition in the existing technology are solved, and security management and early warning are realized throughout the life cycle, improving the accuracy and real-timeness of security management.
Patent Information
- Application Number
- CN202411470135.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-21
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2044-10-21
AI Technical Summary
The existing personnel status recognition technology has unstable identification accuracy in complex environments, making it difficult to make full use of multi-source heterogeneous data, lacks real-time risk assessment and intelligent decision-making support, and cannot achieve security management throughout the life cycle.
By parameterizing the original image data, intelligent BIM data is generated, multi-disciplinary data exchange, personnel location and dangerous area marking are carried out, and intelligent attitude analysis and fuzzy logic evaluation are combined to realize behavior recognition and real-time early warning, providing full life cycle management.
The conversion from 2D images to 3D models is realized, the degree of data visualization and understanding is improved, the movements and postures of staff are accurately captured, safety hazards are discovered in a timely manner, and real-time warning and long-term safety management strategies are provided, which significantly improves the accuracy and real-time nature of safety management.
Smart Images

Figure CN119516161B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of target detection, and in particular to a method and system for identifying personnel status based on target recognition detection. Background Art
[0002] Existing human presence recognition technologies primarily rely on traditional computer vision algorithms and simple machine learning models. These methods typically employ fixed feature extraction methods, such as HOG (Histogram of Oriented Gradients) or SIFT (Scale-Invariant Feature Transform), to identify human posture and behavior. Meanwhile, existing technologies are also beginning to incorporate deep learning models, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), to improve recognition accuracy. In fields such as engineering construction and industrial production, some systems are beginning to integrate human presence recognition with BIM (Building Information Modeling) technology to achieve precise positioning and risk assessment of the work environment.
[0003] However, existing technologies still have many shortcomings in practical applications. First, traditional algorithms have poor adaptability to complex environments and are easily affected by factors such as occlusion and changes in lighting, resulting in unstable recognition accuracy. Second, although deep learning models have been introduced, they are often single models or simple model combinations, making it difficult to fully utilize the advantages of multi-source heterogeneous data. Furthermore, the existing BIM and personnel status recognition are not closely integrated enough, making it difficult to achieve real-time, dynamic risk assessment and early warning. In addition, most systems lack intelligent decision-making support and full-life cycle safety management capabilities, and are unable to provide managers with timely and effective intervention recommendations and long-term safety optimization strategies. Summary of the Invention
[0004] The present application provides a method and system for personnel status recognition based on target recognition detection, which are used to improve the efficiency of personnel status recognition based on target recognition detection.
[0005] In the first aspect, the present application provides a personnel status recognition method based on target recognition detection, and the personnel status recognition method based on target recognition detection includes: performing parameter processing on the collected original image data to obtain intelligent BIM data; performing multidisciplinary data exchange processing on the intelligent BIM data to obtain target detection data, wherein the target detection data includes personnel location information, safety equipment recognition results and danger zone markings; performing intelligent posture analysis processing on the human body information in the target detection data to obtain optimized posture data; performing behavior recognition processing on the optimized posture data to obtain behavior analysis data, wherein the behavior analysis data includes the type of violation behavior, climbing status judgment and risk level assessment; performing fuzzy logic evaluation processing on the behavior analysis data to obtain safety compliance data; performing intelligent early warning processing on the safety compliance data and real-time monitoring data to obtain full life cycle management data, wherein the full life cycle management data includes real-time early warning information, intervention measure recommendations and safety incident reports.
[0006] In a second aspect, the present application provides a device for identifying a person's status based on target recognition detection, the device comprising:
[0007] The acquisition module is used to perform parameterized processing on the collected original image data to obtain intelligent BIM data;
[0008] an exchange module for performing multidisciplinary data exchange processing on the intelligent BIM data to obtain target detection data, wherein the target detection data includes personnel location information, safety equipment identification results, and danger zone markings;
[0009] An analysis module, configured to perform intelligent posture analysis on the human body information in the target detection data to obtain optimized posture data;
[0010] an identification module for performing behavior recognition processing on the optimized posture data to obtain behavior analysis data, wherein the behavior analysis data includes the type of violation behavior, the determination of the climbing state, and the risk level assessment;
[0011] An evaluation module, configured to perform fuzzy logic evaluation processing on the behavior analysis data to obtain safety compliance data;
[0012] The early warning module is used to perform intelligent early warning processing on the safety compliance data and real-time monitoring data to obtain full life cycle management data, wherein the full life cycle management data includes real-time early warning information, intervention measure recommendations and safety incident reports.
[0013] In the technical solution provided by this application, raw image data is parametrically processed to generate intelligent BIM data. This step enables the conversion from 2D images to 3D models, laying the foundation for subsequent spatial analysis and target positioning, while also improving data visualization and comprehension. Next, the intelligent BIM data undergoes multidisciplinary data exchange processing to generate target detection data containing personnel location information, safety equipment identification results, and hazardous area markings. This step enables comprehensive perception and analysis of the work scene, providing detailed environmental information for safety management. The human body information in the target detection data is then subjected to intelligent posture analysis to generate optimized posture data. This step accurately captures the movements and postures of workers, helping to identify potential unsafe behaviors. Subsequently, the optimized posture data undergoes behavior recognition processing to generate behavior analysis data, including violation type, climbing status determination, and risk level assessment. This step enables intelligent identification and risk assessment of worker behavior, enabling timely identification of safety hazards. The behavior analysis data is then subjected to fuzzy logic evaluation to generate safety compliance data. This step uses a fuzzy inference system to comprehensively evaluate worker behavior, providing a more objective and comprehensive assessment of safety compliance. Finally, intelligent early warning processing is performed on safety compliance data and real-time monitoring data, resulting in full lifecycle management data, including real-time warning information, intervention recommendations, and safety incident reports. This step completes the closed loop from data analysis to safety management, enabling not only timely warnings but also specific intervention recommendations and supporting the formulation of long-term safety management strategies. This entire approach, driven by data and using intelligent algorithms, shifts the safety management model from passive response to proactive prevention, significantly improving the accuracy and real-time nature of safety management. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0015] Figure 1 Schematic diagram of an embodiment of a method for identifying a person's status based on target recognition and detection in an embodiment of the present application;
[0016] Figure 2 This is a schematic diagram of an embodiment of a personnel status recognition device based on target recognition detection in an embodiment of the present application. DETAILED DESCRIPTION
[0017] The embodiments of the present application provide a method and system for identifying the status of a person based on target recognition detection. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0018] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiments of the present application, an embodiment of a method for identifying a person's status based on target recognition detection includes:
[0019] Step S101: performing parameter processing on the collected original image data to obtain intelligent BIM data;
[0020] Step S102: Perform multidisciplinary data exchange processing on the intelligent BIM data to obtain target detection data, wherein the target detection data includes personnel location information, safety equipment identification results, and danger zone markings;
[0021] Step S103: performing intelligent posture analysis on the human body information in the target detection data to obtain optimized posture data;
[0022] Step S104: performing behavior recognition processing on the optimized posture data to obtain behavior analysis data, wherein the behavior analysis data includes the type of violation behavior, the judgment of the climbing state, and the risk level assessment;
[0023] Step S105: Perform fuzzy logic evaluation on the behavior analysis data to obtain safety compliance data;
[0024] Step S106: Perform intelligent early warning processing on the safety compliance data and real-time monitoring data to obtain full life cycle management data, wherein the full life cycle management data includes real-time early warning information, intervention measure recommendations and safety incident reports.
[0025] It is understandable that the execution subject of this application can be a personnel status recognition device based on target recognition detection, or a terminal or a server, which is not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.
[0026] Specifically, the collected raw image data undergoes parametric processing to generate intelligent BIM data. This process includes image denoising, histogram equalization, geometric correction, illumination compensation, feature extraction, feature mapping, spatial information fusion, visualization, semantic segmentation, and parametric modeling. For example, image denoising is performed using a median filter, histogram equalization is performed using the cumulative distribution function, geometric correction is performed using an affine transformation, and illumination compensation is performed using the Retinex algorithm. Next, a convolutional neural network is used to extract image features, which are then associated with the BIM model through feature mapping. Spatial information fusion uses the ICP algorithm to align the 2D image with the 3D BIM model, and visualization uses WebGL technology to render the 3D scene. Semantic segmentation is performed using a U-Net network, and finally, intelligent BIM data is generated through parametric modeling. Multidisciplinary data exchange processing is performed on the intelligent BIM data to generate object detection data. This step includes data format conversion, data classification, feature extraction, object detection, non-maximum suppression, coordinate mapping, object classification, safety equipment identification, and hazard zone delineation. Data format conversion utilizes the IFC standard, and data classification is based on a hierarchical clustering algorithm. Feature extraction uses the ResNet50 network, and object detection employs the YOLOv5 algorithm. Non-maximum suppression filters objects by calculating the intersection over union (IoU), while coordinate mapping uses a homography matrix to convert 2D coordinates into 3D spatial coordinates. Object classification utilizes the SVM algorithm, safety device identification uses the Faster R-CNN network, and hazardous area demarcation is based on a spatial clustering algorithm.
[0027] Intelligent posture analysis is performed on human body information in target detection data to generate optimized posture data. This process includes body region extraction, multi-scale feature extraction, keypoint localization, temporal information fusion, occlusion inference, 3D reconstruction, posture angle calculation, motion trajectory analysis, and confidence assessment. Body region extraction uses the Mask R-CNN algorithm, and multi-scale feature extraction uses the FPN network. Keypoint localization uses the improved HRNet algorithm, and temporal information fusion utilizes a Kalman filter. Occlusion inference is based on a graph convolutional network, and 3D reconstruction uses a monocular depth estimation algorithm. Pose angle calculation uses the law of cosines, motion trajectory analysis uses the DTW algorithm, and confidence assessment is based on Monte Carlo dropout. The optimized posture data is then processed for behavior recognition to generate behavior analysis data. This step includes spatiotemporal graph construction, feature extraction, multi-level convolution, attention mechanism, behavior classification, climbing state determination, height estimation, and safety risk assessment. The spatiotemporal graph construction uses the ST-GCN model, and feature extraction uses the I3D network. Multi-level convolution uses a 3D convolutional neural network, and the attention mechanism uses a self-attention model. Behavior classification is based on an LSTM network, and climbing status determination utilizes a decision tree algorithm. Height estimation is achieved through triangulation, and safety risk assessment utilizes the Fuzzy AHP method. Fuzzy logic evaluation is performed on behavioral analysis data to generate safety compliance data. This process involves constructing a fuzzy rule base, fuzzifying input variables, fuzzy reasoning, dynamic weight adjustment, fuzzy output generation, defuzzification, multidimensional evaluation, comprehensive analysis, and safety compliance calculation. The fuzzy rule base is constructed based on expert knowledge, and input variable fuzzification utilizes triangular membership functions. Fuzzy reasoning utilizes a modified Mamdani fuzzy inference system, and dynamic weight adjustment is based on the Adaptive Neuro-Fuzzy Inference System (ANFIS). Fuzzy output generation utilizes the center of gravity method, while defuzzification utilizes the maximum membership method. The multidimensional evaluation encompasses safety, efficiency, and compliance. The comprehensive analysis utilizes the Analytic Hierarchy Process (AHP), and safety compliance calculation is based on the weighted average method.
[0028] Finally, intelligent early warning processing is performed on safety compliance data and real-time monitoring data to generate full lifecycle management data. This step includes data fusion, risk level assessment, multi-threshold early warning mechanism setup, early warning trigger determination, early warning information personalization, early warning information generation, intelligent intervention decision-making, remote expert support, closed-loop feedback, and event analysis. Data fusion utilizes DS evidence theory, and risk level assessment uses the random forest algorithm. The multi-threshold early warning mechanism is based on the dynamic threshold method, and early warning trigger determination utilizes pattern recognition technology. Early warning information personalization is achieved through a collaborative filtering algorithm, and early warning information generation utilizes a multimodal output algorithm. Intelligent intervention decision-making is based on reinforcement learning, and the remote expert support system utilizes knowledge graph technology. The closed-loop feedback mechanism utilizes PID control principles, and event analysis employs an association rule mining algorithm.
[0029] For example, at a large construction site, this method first processed images captured by surveillance cameras. Image denoising reduced noise by 30%, and histogram equalization increased image contrast by 40%. Geometric correction corrected the 5° tilt caused by the camera's mounting angle, and illumination compensation improved visibility in shadowed areas. A convolutional neural network extracted 128-dimensional feature vectors, which were mapped to 3,000 components in the BIM model. Spatial information fusion kept the registration error between the 2D image and the 3D BIM model within 5 cm. Semantic segmentation achieved 95% accuracy, and the resulting intelligent BIM data contained detailed information on 50 work areas. In the object detection phase, the YOLOv5 algorithm identified 100 workers, 50 pieces of safety equipment, and 10 hazardous areas on the construction site with 98% accuracy. In the human pose analysis phase, the improved HRNet algorithm located 17 key points on the human body with an average error of less than 2 cm. Temporal information fusion reduced jitter by 30%, keeping the depth error of the 3D reconstruction within 10 cm. Behavior recognition processing identified five common violations, such as not wearing a hard hat and entering an unauthorized area, with an accuracy rate of 92%. The accuracy of altitude status assessment was 95%, with a height estimation error within 20 cm. A fuzzy logic evaluation system evaluated each worker's behavior based on 100 fuzzy rules, resulting in a safety compliance score ranging from 0 to 100. Finally, the intelligent early warning system successfully warned of three potential safety incidents based on real-time data, provided 10 targeted intervention recommendations, and generated a safety incident report containing 30 days of data.
[0030] In the present embodiment, raw image data is parametrically processed to generate intelligent BIM data. This step enables the conversion from 2D images to 3D models, laying the foundation for subsequent spatial analysis and target positioning, while also improving data visualization and comprehension. Next, the intelligent BIM data undergoes multidisciplinary data exchange processing to generate target detection data containing personnel location information, safety equipment identification results, and hazardous area markings. This step enables comprehensive perception and analysis of the work scene, providing detailed environmental information for safety management. The human body information in the target detection data is then subjected to intelligent posture analysis to generate optimized posture data. This step accurately captures the movements and postures of workers, helping to identify potential unsafe behaviors. Subsequently, the optimized posture data is subjected to behavior recognition processing to generate behavior analysis data, including violation types, height status determination, and risk level assessment. This step enables intelligent identification and risk assessment of worker behavior, enabling timely identification of safety hazards. The behavior analysis data is then subjected to fuzzy logic evaluation to generate safety compliance data. This step uses a fuzzy inference system to comprehensively evaluate worker behavior, providing a more objective and comprehensive assessment of safety compliance. Finally, intelligent early warning processing is performed on safety compliance data and real-time monitoring data, resulting in full lifecycle management data, including real-time warning information, intervention recommendations, and safety incident reports. This step completes the closed loop from data analysis to safety management, enabling not only timely warnings but also specific intervention recommendations and supporting the formulation of long-term safety management strategies. This entire approach, driven by data and using intelligent algorithms, shifts the safety management model from passive response to proactive prevention, significantly improving the accuracy and real-time nature of safety management.
[0031] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0032] (1) Performing image denoising on the original image data to obtain denoised image data, and performing histogram equalization on the denoised image data to obtain enhanced image data;
[0033] (2) performing geometric correction processing on the enhanced image data to obtain corrected image data, and performing illumination compensation processing on the corrected image data to obtain compensated image data;
[0034] (3) Perform feature extraction processing on the compensated image data through a convolutional neural network to obtain image feature data, and perform feature mapping processing on the image feature data to obtain BIM feature data;
[0035] (4) Perform spatial information fusion processing on BIM feature data to obtain fused feature data, and visualize the fused feature data through graphical processing to obtain visualized BIM data;
[0036] (5) Perform semantic segmentation processing on the visual BIM data to obtain segmented BIM data, and perform parametric modeling processing on the segmented BIM data through attribute association processing to obtain intelligent BIM data.
[0037] Specifically, the original image data is subjected to image denoising to obtain denoised image data. Image denoising is a technique for removing or reducing unnecessary noise in an image. Common methods include median filtering, Gaussian filtering, and wavelet transform. The non-local means (NLM) denoising algorithm is employed. This algorithm estimates the true value of each pixel by searching for similar pixel blocks throughout the image, effectively preserving image detail while reducing noise. The denoised image data is then subjected to histogram equalization to obtain enhanced image data. Histogram equalization is a commonly used image enhancement technique that redistributes the grayscale value distribution of an image to increase image contrast. This step uses the cumulative distribution function (CDF) to implement histogram equalization, thereby enhancing the overall visual quality and detail visibility of the image. The enhanced image data is then subjected to geometric correction to obtain corrected image data. Geometric correction corrects geometric distortion caused by factors such as camera angle and lens distortion. This method employs an affine transformation based on control points for geometric correction. The transformation matrix is calculated using the least squares method to correct the image geometry.
[0038] Illumination compensation is performed on the corrected image data to obtain compensated image data. Illumination compensation aims to eliminate the effects of uneven illumination on image quality. The Retinex algorithm is used for illumination compensation. This algorithm, based on the working principles of the human visual system, can effectively improve the local contrast and color reproduction of the image. Subsequently, feature extraction is performed on the compensated image data using a convolutional neural network to obtain image feature data. A convolutional neural network (CNN) is a deep learning model particularly suitable for processing data with grid structures, such as images. In this method, ResNet-50 is used as the feature extraction network. This network solves the vanishing gradient problem of deep networks through residual learning and is capable of extracting rich multi-scale features.
[0039] Feature mapping is performed on the image feature data to obtain BIM feature data. Feature mapping is the process of associating extracted image features with BIM (Building Information Modeling) data. This step uses fully connected layers to convert the features extracted by the CNN into feature representations compatible with the BIM model, mapping 2D image features to 3D BIM features. Next, spatial information fusion is performed on the BIM feature data to obtain fused feature data. Spatial information fusion aims to integrate 2D image features with the spatial information of the 3D BIM model. Here, the Iterative Closest Point (ICP) algorithm is used for 2D-3D registration to accurately locate image features at corresponding locations on the BIM model.
[0040] The fused feature data is visualized through graphical processing to obtain visualized BIM data. This visualization uses WebGL technology to render the fused feature data into an intuitive 3D scene, facilitating subsequent analysis and interaction. Next, semantic segmentation is performed on the visualized BIM data to obtain segmented BIM data. Semantic segmentation is a pixel-level classification task that aims to assign a semantic label to each pixel in an image. This method uses a U-Net network for semantic segmentation. This network has an encoder-decoder structure that effectively captures multi-scale contextual information and achieves accurate pixel-level segmentation.
[0041] Finally, the segmented BIM data is parametrically modeled through attribute association, generating intelligent BIM data. Parametric modeling involves associating the segmented BIM data with a predefined parametric model. This step uses a rule-based reasoning system to associate the segmentation results with the geometric and semantic attributes of the BIM model, generating intelligent BIM data with rich semantic information.
[0042] For example, in a task involving personnel status recognition at a large construction site, a 4000x3000 pixel original image was collected. Using the NLM denoising algorithm, the image's signal-to-noise ratio (SNR) was increased from 20dB to 28dB. Histogram equalization increased the image's contrast by 40%, making the previously blurry worker outlines clearly discernible. Geometric correction corrected the 5-degree tilt caused by the camera's mounting angle, ensuring vertical lines in the image were truly vertical. Illumination compensation effectively brightened shadow areas, increasing visibility from only 20% to 80%. The ResNet-50 network extracted 2048-dimensional feature vectors, which were then mapped to a 500-dimensional BIM feature space through fully connected layers. The ICP algorithm registered the 2D image features with the 3D BIM model, achieving an average registration error of less than 5cm. WebGL rendering generated a 3D scene model containing 100,000 vertices. The U-Net network performed semantic segmentation on the scene, accurately identifying 95% of workers, equipment, and building structures. Finally, parametric modeling linked the segmentation results with the BIM model, generating intelligent BIM data containing 500 parametric objects, each with an average of 20 attributes, such as location, size, material, and safety status.
[0043] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0044] (1) Perform data format conversion on intelligent BIM data to obtain standardized BIM data, and perform data classification on the standardized BIM data to obtain classified BIM data;
[0045] (2) Feature extraction is performed on the classified BIM data to obtain BIM feature vectors, and target detection is performed on the BIM feature vectors using the YOLOv5 algorithm to obtain candidate target data;
[0046] (3) Perform non-maximum suppression processing on the candidate target data to obtain the filtered target data, and perform coordinate mapping processing on the filtered target data to obtain spatial positioning data;
[0047] (4) Perform target classification processing on the spatial positioning data to obtain classified target data, and perform security device identification processing on the classified target data through a deep learning algorithm to obtain device identification data;
[0048] (5) The equipment identification data is processed into dangerous areas to obtain area marking data, and the area marking data is processed into data integration to obtain target detection data.
[0049] Specifically, intelligent BIM data undergoes data format conversion to obtain standardized BIM data. Using the Industry Foundation Classes (IFC) standard, BIM data from various sources and formats is uniformly converted into a standardized IFC format. IFC is an international standard for building information modeling that ensures data interoperability and consistency. During the conversion process, the geometric, attribute, and relationship information in the original BIM data is mapped to the corresponding entities and attributes in the IFC schema.
[0050] The standardized BIM data is classified to obtain classified BIM data. A hierarchical clustering algorithm is used for data classification to divide BIM objects into different categories, such as structural elements, equipment, personnel, etc., according to their attributes and characteristics. The hierarchical clustering algorithm calculates the similarity between objects and gradually merges similar objects to form a tree-like hierarchical structure. This classification method can effectively organize and manage complex BIM data, laying the foundation for subsequent feature extraction and target detection. Feature extraction is performed on the classified BIM data to obtain a BIM feature vector. Feature extraction uses the ResNet50 deep convolutional neural network, which solves the gradient vanishing problem in deep network training through residual learning and can extract rich multi-scale features. ResNet50 contains 50 convolutional layers. By extracting and combining features layer by layer, it finally outputs a 2048-dimensional feature vector that contains high-level semantic information of the BIM object.
[0051] Next, the YOLOv5 algorithm is used to perform object detection on the BIM feature vector to obtain candidate object data. YOLOv5 is an advanced single-stage object detection algorithm that transforms the object detection problem into a regression problem, capable of simultaneously predicting the object's bounding box and category. The YOLOv5 algorithm first divides the input image into a grid, then predicts multiple bounding boxes and confidence scores for each grid, and finally selects candidate objects based on a preset threshold. Non-maximum suppression (NMS) is performed on the candidate object data to obtain filtered target data. NMS is a post-processing technique used to eliminate overlapping detection boxes and retain the best detection results. NMS first sorts all candidate boxes by confidence score and then calculates the intersection over union (IoU) between overlapping boxes. If the IoU exceeds a preset threshold, the box with the highest confidence score is retained and the remaining overlapping boxes are removed. This step effectively reduces redundant detections and improves object detection accuracy.
[0052] Subsequently, the filtered target data undergoes coordinate mapping to obtain spatial positioning data. Coordinate mapping converts target coordinates in the 2D image space into coordinates in the 3D BIM model space. This process uses the homography matrix, which describes the projective relationship between two planes. By solving the homography equation, the 2D image coordinates can be mapped to the 3D BIM coordinate system, achieving precise spatial positioning of the target. The spatial positioning data undergoes target classification processing to obtain classified target data. Target classification utilizes the support vector machine (SVM) algorithm. SVM is an effective binary classifier that separates data points of different categories by finding an optimal hyperplane. For multi-class classification problems, a one-vs-all strategy is employed, training an SVM classifier for each category. During the classification process, each SVM calculates the probability that the target belongs to that class, ultimately selecting the class with the highest probability as the classification result.
[0053] Next, a deep learning algorithm is used to perform safety equipment identification on the classified target data, generating equipment identification data. The Faster R-CNN (Faster Region-based Convolutional Neural Network) algorithm is employed here. This algorithm is a two-stage object detection model consisting of a region proposal network (RPN) and an object classification network. The RPN first generates candidate regions that may contain objects, and then the object classification network performs accurate classification and bounding box regression on these regions. Faster R-CNN is particularly suitable for identifying safety equipment such as helmets and seat belts because it can handle objects of varying scales and shapes. The equipment identification data is then processed for hazardous area segmentation, generating area labeling data. This is done using the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) spatial clustering algorithm. Based on the concept of density, DBSCAN can discover clusters of arbitrary shapes and is robust to noisy data. By setting an appropriate density threshold and neighborhood radius, DBSCAN can effectively identify potential hazardous areas, such as high-altitude work areas and machinery operating areas.
[0054] Finally, the area marking data is consolidated to generate target detection data. This data consolidation process combines personnel location information, safety equipment identification results, and hazardous area markings into a unified data structure. This data structure uses the JSON (JavaScript Object Notation) format, which is highly readable and extensible. The consolidated target detection data includes each detected target's category, location coordinates, confidence level, associated safety equipment information, and the area type it is located in.
[0055] For example, in a task involving personnel status recognition at a large construction site, intelligent BIM data containing 10,000 BIM objects was first converted to IFC format, resulting in a standardized data size of 500MB. Using a hierarchical clustering algorithm, these objects were classified into 15 main categories: four structural elements, six equipment types, and five personnel categories. A ResNet50 network extracted 2048-dimensional feature vectors from each BIM object, generating a total of 10,000 feature vectors. The YOLOv5 algorithm processed these feature vectors with a confidence threshold of 0.5 and initially detected 1,500 candidate objects. After NMS processing with an IoU threshold of 0.5, 1,200 valid objects were retained. Coordinate mapping then mapped these 1,200 objects from 2D image space (resolution 3840 × 2160) to 3D BIM space (dimensions 100m × 50m × 30m). The SVM classifier classified these 1,200 objects with an accuracy rate of 95%, correctly identifying the categories of 1,140 objects. The Faster R-CNN algorithm identified safety equipment, successfully identifying 480 of 500 safety-related objects with an accuracy rate of 96%. The DBSCAN algorithm, with a neighborhood radius of 5 meters and a minimum sample size of 3, identified 10 potentially hazardous areas within the construction site. The final integrated object detection data contained detailed information on 1,140 objects, with an average of 20 attribute fields per object, for a total data volume of approximately 5 MB.
[0056] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0057] (1) Perform human region extraction processing on the target detection data to obtain human ROI data, and perform feature extraction processing on the human ROI data through a multi-scale feature pyramid network to obtain multi-scale feature data;
[0058] (2) The improved HRNet algorithm is used to locate the key points of the multi-scale feature data to obtain two-dimensional key point data, and the two-dimensional key point data is fused with time series information to obtain continuous posture data;
[0059] (3) Perform occlusion inference processing on the continuous posture data to obtain the completed posture data, and perform three-dimensional reconstruction processing on the completed posture data through a monocular depth estimation algorithm to obtain three-dimensional posture data;
[0060] (4) Calculate the posture angle of the three-dimensional posture data to obtain angle feature data, and analyze the motion trajectory of the angle feature data to obtain motion feature data;
[0061] (5) Perform confidence evaluation on the motion feature data to obtain credibility data, and perform data fusion on the credibility data to obtain optimized posture data.
[0062] Specifically, the target detection data is processed for human region extraction to obtain human ROI (Region of Interest) data. The Mask R-CNN algorithm is used, which adds a branch network on the basis of Faster R-CNN to generate an accurate segmentation mask for the target. Mask R-CNN can not only detect the bounding box of the human body, but also generate pixel-level human contours, thereby accurately extracting the human body area. The human ROI data is processed for feature extraction through a multi-scale feature pyramid network (FPN) to obtain multi-scale feature data. FPN is a top-down architecture that can use the multi-layer feature map of the convolutional neural network to generate a feature pyramid with strong semantic information. The processing process of FPN can be expressed by the following formula:
[0063] ;
[0064] in, Indicates the The feature map of the layer, U represents the upsampling operation, represents the convolutional feature map of the corresponding layer, Represents a convolution operation. This structure effectively captures human features at different scales, providing rich information for subsequent pose estimation. The improved HRNet (High-Resolution Network) algorithm performs key point localization processing on multi-scale feature data, generating two-dimensional key point data. HRNet's improvement lies in the introduction of an attention mechanism, which enhances the network's ability to perceive key points. The key point localization process can be expressed as:
[0065] ;
[0066] Among them, K represents the key point coordinates, A represents the attention weight, and F represents the feature map. represents element-wise multiplication, is a differentiable operation used to transform the heatmap into precise coordinates.
[0067] The two-dimensional key point data is subjected to time series information fusion processing to obtain continuous posture data. This step uses the Kalman filter, which can optimally estimate the posture based on historical observations and current measurements. The state update equation of the Kalman filter is:
[0068] ;
[0069] in, represents the current state, F is the state transfer matrix, B is the control input matrix, is the control vector, is process noise. Occlusion inference is performed on the continuous pose data to obtain the completed pose data. Occlusion inference uses a graph convolutional network (GCN), which can effectively infer the positions of occluded key points by leveraging the topological structure of the human skeleton. A forward pass of the GCN can be expressed as:
[0070] ;
[0071] in, Represents the node features of the lth layer, A is the adjacency matrix, D is the degree matrix, is the weight matrix, is the activation function. Then, the monocular depth estimation algorithm is used to perform 3D reconstruction on the completed pose data to obtain 3D pose data. Monocular depth estimation uses a DenseDepth network, which estimates depth information from a single image through an encoder-decoder structure and skip connections. The loss function for depth estimation includes:
[0072] ;
[0073] in, is the depth reconstruction error, is the gradient consistency error, is the structural similarity error, 、 、 Is the weight coefficient. The three-dimensional posture data is processed for posture angle calculation to obtain angle feature data. Angle calculation is based on the cosine theorem, for example, to calculate the elbow angle :
[0074] ;
[0075] in, and are the vectors of the upper arm and forearm respectively.
[0076] The angle feature data is processed by motion trajectory analysis to obtain motion feature data. The motion trajectory analysis uses the Dynamic Time Warping (DTW) algorithm, which can calculate the similarity between two time series and is not affected by changes in motion speed. The DTW distance is defined as:
[0077] ;
[0078] Where X and Y are two time series, w is the alignment path, d is the distance function, and K is the length of the alignment path. represents the element in the X sequence at the kth step of the alignment path. Represents the element in the Y sequence at step k of the alignment path. This means selecting the path that minimizes the total distance among all possible alignment paths. Finally, the motion feature data is confidence evaluated to obtain credibility data, which is then fused to produce optimized pose data. Confidence evaluation uses Monte Carlo dropout techniques, estimating model uncertainty through multiple forward propagations without dropout. Data fusion uses a weighted average method to combine the confidence scores of each component.
[0079] For example, in a construction site personnel status recognition task, 100 human ROIs with an average size of 256×256 pixels were first extracted from a 4K resolution (3840×2160) image. The FPN network generated feature maps at five scales: 32×32, 64×64, 128×128, 256×256, and 512×512, each containing 256 channels. The improved HRNet algorithm localized 17 human keypoints on these feature maps, with an average localization error of less than 3 pixels. A Kalman filter performed temporal fusion on a 30-frame (1-second) sequence of keypoints, reducing pose jitter by 40%. The GCN network successfully inferred 15% of occluded keypoints with an average error of 5 cm. Monocular depth estimation converted 2D poses to 3D poses, achieving an average relative error of 10%. Pose angle calculation determined 20 joint angles with an accuracy of 1 degree. The DTW algorithm analyzed a 100-frame (approximately 3-second) motion sequence and identified five common construction actions, such as carrying, bending, and climbing, with an accuracy rate of 92%. Monte Carlo dropout was used to perform 50 samplings to obtain confidence intervals for each posture parameter, with an average confidence level of 0.85. The resulting fused optimized posture data includes each worker's 3D joint position, joint angle, motion type, and corresponding confidence level, generating approximately 1KB of data per worker per second. This highly accurate and reliable posture data provides a solid foundation for subsequent safety monitoring and efficiency analysis.
[0080] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0081] (1) Perform spatiotemporal graph construction on the optimized posture data to obtain posture graph structure data, and perform feature extraction on the posture graph structure data through the spatiotemporal graph convolutional network to obtain spatiotemporal feature data;
[0082] (2) Perform multi-level convolution processing on the spatiotemporal feature data to obtain high-level action feature data, and then perform weight distribution processing on the high-level action feature data through the attention mechanism to obtain weighted feature data;
[0083] (3) Performing behavior classification processing on the weighted feature data to obtain violation behavior type data, and performing climbing state judgment processing on the violation behavior type data to obtain climbing state data;
[0084] (4) Perform altitude estimation processing on the climbing status data to obtain altitude information data, and perform safety risk assessment processing on the altitude information data through the risk assessment model to obtain risk level data;
[0085] (5) Perform data integration processing on the risk level data to obtain behavioral analysis data.
[0086] Specifically, the optimized pose data is processed through spatiotemporal graph construction to obtain pose graph structure data. Spatiotemporal graph construction represents human pose sequences as a graph structure, where nodes represent key points and edges represent connections between key points. This representation effectively captures both spatial and temporal features of human pose. The spatiotemporal graph construction uses a sliding window approach to combine pose data from T consecutive frames into a graph structure, where each node contains the spatial coordinates and temporal information of a key point. Feature extraction is performed on the pose graph structure data using a spatiotemporal graph convolutional network (ST-GCN), generating spatiotemporal feature data. ST-GCN is a deep learning model specifically designed for processing spatiotemporal graph data, which simultaneously considers information in both spatial and temporal dimensions. The core of ST-GCN is the graph convolution operation, which defines convolution and pooling operations on the graph structure to effectively extract spatiotemporal features from pose sequences. The ST-GCN network consists of multiple ST-GCN layers, each of which transforms and aggregates features of the input graph structure, ultimately outputting a high-dimensional spatiotemporal feature vector.
[0087] The spatiotemporal feature data is then subjected to multi-layer convolution to generate high-level motion feature data. This multi-layer convolution utilizes a 3D convolutional neural network (3D CNN), which performs convolution operations simultaneously in both spatial and temporal dimensions, thereby capturing more complex motion patterns. The 3D CNN consists of multiple 3D convolutional layers, 3D pooling layers, and fully connected layers, each layer extracting progressively higher-level motion features. Through this hierarchical feature extraction process, the network is able to learn features ranging from low-level body movements to high-level behavioral patterns. An attention mechanism is used to assign weights to the high-level motion feature data, generating weighted feature data. The introduction of the attention mechanism enables the model to automatically learn the importance of different spatiotemporal regions, thereby better focusing on key motion details. The self-attention mechanism is employed here, which allows the model to consider information from all locations when computing the representation of a given location and dynamically assigns weights based on relevance. The self-attention mechanism assigns weights by calculating the similarity between the query, key, and value, resulting in a weighted feature representation.
[0088] Subsequently, the weighted feature data is processed for behavioral classification to obtain violation type data. This classification utilizes a long short-term memory (LSTM) network, a specialized recurrent neural network that effectively processes long sequences of data and captures long-term dependencies. LSTM uses a gating mechanism, including input, forget, and output gates, to control the flow of information, enabling it to learn complex temporal patterns. In this method, the LSTM network receives the weighted feature sequence as input and outputs a probability distribution for the violation category. The violation type data is then processed for height determination to obtain height determination data. This determination utilizes a decision tree algorithm, which classifies data based on a series of conditional judgments. Each internal node in the decision tree represents a test on a feature or attribute, each branch represents a test output, and each leaf node represents a category or decision outcome. In this method, the decision tree utilizes the violation type, posture characteristics, and environmental information to determine whether a worker is in a height determination.
[0089] The ascent status data is then processed for altitude estimation to generate altitude information. This altitude estimation utilizes triangulation, a method that calculates the target's actual height based on known camera parameters and reference objects in the image. Specifically, by comparing the target's size in the image with the size of reference objects of known height (such as buildings or standard equipment), combined with the camera's internal and external parameters, the target's actual height can be inferred.
[0090] Finally, the risk assessment model performs a safety risk assessment on the altitude information data to obtain risk level data. This risk level data is then integrated to generate behavioral analysis data. The risk assessment model utilizes a fuzzy comprehensive evaluation method, which combines fuzzy mathematics and comprehensive evaluation concepts to address uncertainty and ambiguity in the evaluation process. The fuzzy comprehensive evaluation method first establishes a set of evaluation factors and a set of comments. The factor set is then mapped to the comment set using a fuzzy relationship matrix, ultimately resulting in a membership vector for the risk level. The data integration process then combines the risk level with the previously generated behavioral classification results, climbing status, and altitude information to form the final behavioral analysis data.
[0091] For example, in a task involving personnel status recognition at a large construction site, optimized posture data of 100 workers was first processed. A spatiotemporal graph was constructed using a sliding window of 30 frames (1 second), generating a graph structure consisting of 17 nodes (corresponding to 17 key points of the human body). The ST-GCN network consists of eight ST-GCN layers, each outputting a 256-dimensional feature vector. The 3D CNN network comprises five 3D convolutional layers, ultimately outputting 1024-dimensional high-level motion features. The self-attention mechanism computes 16 attention heads, each with a dimension of 64, to produce weighted feature data. The LSTM network consists of two layers, each with 128 hidden units, ultimately outputting probability distributions for 10 common traffic violations. A decision tree algorithm uses five key features (such as height, posture angle, and tool usage) to determine the height status, achieving an accuracy of 95%. The average error in height estimation is kept within 10 cm. The fuzzy comprehensive evaluation model considers five risk factors (height, postural stability, protective measures, environmental conditions, and operational complexity) and categorizes risk levels into five levels (low, lower, medium, higher, and high). The resulting behavioral analysis data includes each worker's violation type, climbing status, estimated height, and risk level, generating approximately 2KB of data per worker per second. Through this comprehensive behavioral analysis, the method successfully identified 15 potentially high-risk behaviors, including eight instances of working at height without a safety belt and seven instances of improper operation of heavy equipment. This provided site managers with timely and accurate safety warnings, effectively reducing the incidence of safety incidents.
[0092] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0093] (1) Build a fuzzy rule base on the behavior analysis data to obtain fuzzy rule set data, and perform fuzzification processing on the input variables of the fuzzy rule set data to obtain input fuzzy set data;
[0094] (2) The improved Mamdani fuzzy inference system is used to perform fuzzy inference processing on the input fuzzy set data to obtain inference result data, and the weight of the inference result data is dynamically adjusted to obtain weighted result data;
[0095] (3) Perform fuzzy output generation processing on the weighted result data to obtain fuzzy output set data, and defuzzify the fuzzy output set data using the centroid method to obtain accurate scoring data;
[0096] (4) Perform multi-dimensional evaluation processing on the precise scoring data to obtain multi-dimensional evaluation data, and perform comprehensive analysis and processing on the multi-dimensional evaluation data to obtain comprehensive evaluation data;
[0097] (5) Calculate and process the comprehensive assessment data to obtain safety compliance data.
[0098] Specifically, a fuzzy rule base is constructed on the behavioral analysis data to generate fuzzy rule set data. A fuzzy rule base is a set of IF-THEN statements that describe the relationship between input variables and output variables. These rules are based on the knowledge and experience of domain experts and cover a wide range of possible scenarios and corresponding evaluation results. For example, a rule might be "IF High risk AND Inadequate protective measures THEN Safety compliance is low." The rule base construction process includes determining the input and output variables, defining fuzzy sets for the linguistic variables, and designing the rule set. The fuzzy rule set data is then fuzzified to generate input fuzzy set data. Fuzzification is the process of converting precise input values into fuzzy sets. This step uses membership functions to describe the membership of input variables to different fuzzy sets. Common membership functions include triangular, trapezoidal, and Gaussian functions. For example, for the input variable "High risk," three fuzzy sets can be defined: "Low," "Medium," and "High." A triangular membership function can be used to describe the membership of a specific height value to these three fuzzy sets.
[0099] Then, the improved Mamdani fuzzy inference system performs fuzzy inference on the input fuzzy set data to obtain the inference result data. The Mamdani fuzzy inference system is a commonly used fuzzy inference method that reaches conclusions through four steps: fuzzification of input variables, rule evaluation, aggregation of rule outputs, and defuzzification. The improved Mamdani system introduces a dynamic weight adjustment mechanism based on traditional methods, dynamically adjusting the importance of rules based on the characteristics of the input data. This improvement improves the system's adaptability to different situations and inference accuracy. Dynamic weight adjustment is performed on the inference result data to obtain weighted result data. This dynamic weight adjustment is based on the concept of the adaptive neural-fuzzy inference system (ANFIS), which optimizes rule weights through feedback learning. ANFIS combines the learning capabilities of neural networks with the reasoning capabilities of fuzzy systems, enabling continuous adjustment and optimization of fuzzy rule weights based on historical data and real-time feedback. This step enables the system to adapt to changing work environments and safety requirements.
[0100] Subsequently, the weighted result data is processed for fuzzy output generation to obtain fuzzy output set data. Fuzzy output generation is the process of aggregating the outputs of each rule. Common aggregation methods include the maximum method, weighted average method, etc. In this method, an improved weighted average method is adopted, which takes into account the weight and activation strength of each rule to obtain a more accurate fuzzy output set. The fuzzy output set data is defuzzified by the centroid method to obtain accurate scoring data. Defuzzification is the process of converting a fuzzy output set into an accurate numerical value. The centroid method is a commonly used defuzzification method, which obtains the final accurate value by calculating the centroid of the fuzzy output set. The advantage of the centroid method is that it can comprehensively consider the entire output distribution and obtain a relatively smooth and stable result.
[0101] The precise scoring data is processed through a multidimensional evaluation process to produce multidimensional assessment data. This multidimensional assessment considers multiple aspects, such as safety, efficiency, and compliance, with each dimension having its own specific evaluation criteria and weighting. This step expands the single precise score into a multidimensional score, providing a more comprehensive assessment result. Next, the multidimensional assessment data is comprehensively analyzed to produce comprehensive assessment data. This comprehensive analysis utilizes the Analytic Hierarchy Process (AHP), which decomposes complex problems into a hierarchical structure and performs pairwise comparisons to determine the relative importance of each factor. AHP effectively processes both qualitative and quantitative factors, resulting in an objective, comprehensive assessment result.
[0102] Finally, the comprehensive assessment data is processed for safety compliance, resulting in compliance data. This compliance calculation uses a weighted average method, averaging the scores of each dimension based on their importance to produce a final safety compliance score. This score reflects the degree to which workers' behavior complies with safety regulations, providing a quantitative basis for subsequent safety management and intervention.
[0103] For example, in a task involving personnel status identification at a large construction site, a fuzzy rule base containing 100 rules was first constructed, covering aspects such as work at height, equipment operation, and personal protection. Input variables included five variables: height risk (low, medium, high), protective measures (inadequate, general, sufficient), and operational compliance (low, medium, high). Three fuzzy sets were defined for each variable using triangular membership functions. For example, for a worker working at a height of 10 meters, the height risk input values were 0.8 (high), 0.2 (medium), and 0 (low). The improved Mamdani fuzzy inference system processed these inputs and initially determined a safety compliance score of 0.6. Dynamic weight adjustment increased the weight of the rules related to work at height by 10% based on the safety record of the previous week, resulting in a recalculated weighted result of 0.55. The fuzzy output generation stage aggregated the outputs of multiple rules to generate a composite fuzzy set. Defuzzification was performed using the centroid method, resulting in a final accuracy score of 75 (out of 100). A multi-dimensional assessment expanded this score into three dimensions: safety (80 points), efficiency (70 points), and compliance (75 points). The AHP comprehensive analysis assigned weights of 0.5, 0.3, and 0.2, respectively, resulting in a calculated overall assessment score of 76. Finally, based on industry standards, this 76 score was converted to a safety compliance score of 85%. This result indicates that the worker's behavior generally complied with safety regulations, but there was room for improvement, particularly in terms of work efficiency. This comprehensive safety compliance assessment not only provided quantitative safety assessment results but also provided a basis for site managers to develop targeted safety training and improvement measures, effectively improving overall safety management.
[0104] In a specific embodiment, the process of executing step S106 may specifically include the following steps:
[0105] (1) Perform data fusion processing on safety compliance data and real-time monitoring data to obtain fusion status data, and perform risk level assessment processing on the fusion status data using the random forest algorithm to obtain dynamic risk data;
[0106] (2) Perform multi-threshold warning mechanism setting processing on dynamic risk data to obtain warning threshold data, and perform warning trigger judgment processing on the warning threshold data to obtain warning trigger data;
[0107] (3) Performing personalized warning information processing on the warning trigger data to obtain personalized warning data, and performing warning information generation processing on the personalized warning data through a multimodal output algorithm to obtain multimodal warning data;
[0108] (4) Using the decision tree algorithm to process the multimodal warning data for intelligent intervention decision-making, obtain intervention recommendation data, and perform remote expert support processing on the intervention recommendation data to obtain expert guidance data;
[0109] (5) Perform closed-loop feedback processing on the expert guidance data to obtain optimization strategy data, and perform event analysis on the optimization strategy data through data mining algorithms to obtain full life cycle management data.
[0110] Specifically, safety compliance data and real-time monitoring data are fused to produce fused status data. Data fusion utilizes the DS evidence theory, which effectively handles the uncertainty and conflict of multi-source information. DS evidence theory integrates information from different sources by calculating basic probability distribution functions, trust functions, and likelihood functions, ultimately yielding a more reliable and comprehensive status description. The fused status data is then processed using the random forest algorithm to assess risk levels and generate dynamic risk data. A random forest is an ensemble learning method composed of multiple decision trees. Each decision tree independently performs classification or regression on the input data, ultimately generating a result through voting or averaging. The advantages of random forests lie in their ability to handle high-dimensional data, robustness to noise, and the ability to assess feature importance. In this method, the random forest model uses the fused status data as input and outputs a risk level assessment result.
[0111] The dynamic risk data is then processed using a multi-threshold early warning mechanism to obtain early warning threshold data. This multi-threshold early warning mechanism is based on a dynamic threshold method, which dynamically adjusts the early warning threshold based on historical data and current trends. The calculation of dynamic thresholds takes into account the time series characteristics of the data, such as seasonality, periodicity, and trends, making early warnings more sensitive and accurate. The early warning threshold data contains thresholds corresponding to different risk levels, providing a basis for subsequent early warning triggering. The early warning threshold data is processed for early warning trigger judgment to obtain early warning trigger data. Early warning trigger judgment utilizes pattern recognition technology, which can identify data patterns that match predefined risk patterns. The pattern recognition process includes three steps: feature extraction, pattern matching, and decision-making. When the identified risk pattern exceeds the preset threshold, an early warning of the corresponding level is triggered.
[0112] Subsequently, the warning trigger data is personalized to generate personalized warning data. This personalized warning information utilizes a collaborative filtering algorithm, which recommends the most appropriate warning information format and content based on user historical behavior and preferences. Collaborative filtering calculates similarities between users to identify similar user groups, then generates personalized recommendations for target users based on their behavioral characteristics. The personalized warning data is then processed using a multimodal output algorithm to generate multimodal warning data. The multimodal output algorithm can generate warning information in various formats, including text, images, and audio, tailored to different scenarios and user needs. This approach comprehensively considers the urgency of the information, the user's ability to receive it, and environmental factors to select the most effective warning method.
[0113] Then, the multimodal warning data is processed for intelligent intervention decisions using a decision tree algorithm to obtain intervention recommendation data. The decision tree algorithm constructs a decision model through a series of if-then rules and can automatically generate intervention recommendations based on the characteristics of the warning data. The advantage of the decision tree is that it has a clear structure and is easy to understand and implement. In this method, the decision tree model takes into account multiple dimensions such as warning level, risk type, and environmental factors, and outputs specific intervention measures. The intervention recommendation data is processed with remote expert support to obtain expert guidance data. The remote expert support system uses knowledge graph technology, which can effectively organize and represent the knowledge and experience of domain experts. The knowledge graph describes complex knowledge structures through entities, relationships, and attributes, supporting intelligent reasoning and querying. Remote experts can quickly access relevant knowledge through this system and review and supplement the automatically generated intervention recommendations.
[0114] Expert guidance data is processed through closed-loop feedback to generate optimized strategy data. This closed-loop feedback mechanism, based on the PID (proportional-integral-derivative) control principle, continuously adjusts control parameters to optimize system response. In security management, this mechanism can continuously optimize early warning and intervention strategies based on the actual effectiveness of intervention measures. Event analysis and processing of optimized strategy data using data mining algorithms generates full lifecycle management data. Data mining utilizes association rule mining algorithms, such as the Apriori algorithm, which can discover implicit and valuable association rules from large amounts of data. By analyzing the effectiveness of optimization strategy implementation and historical event data, deep-seated patterns and trends in security management can be uncovered, providing data support for full lifecycle management.
[0115] For example, in a large-scale construction site personnel status recognition task, safety compliance data (average 85%) and real-time monitoring data (including location, posture, and environmental parameters) for 100 workers were first fused. DS evidence theory set the credibility of these two types of data to 0.7 and 0.8, respectively, resulting in a more comprehensive description of the status. A random forest algorithm, using 1,000 decision trees, performed risk assessment on the fused data, outputting a five-level risk probability distribution. A multi-threshold warning mechanism, based on data from the past three months, set dynamic thresholds for each risk level. For example, the high-risk threshold was adjusted from a fixed 0.8 to a time-varying 0.75-0.85. Warning triggering identified 15 potential risk patterns, three of which exceeded the thresholds, triggering warnings. Warning information was personalized, tailoring the format and content of each worker's warning based on their position, experience, and preferences. A multimodal output algorithm generated warnings in three formats: text messages, voice alerts, and visual charts. A decision tree algorithm generated 50 specific intervention recommendations based on 20 key features. The remote expert support system connected 10 security experts and, through rapid retrieval and reasoning using the knowledge graph, optimized automatically generated intervention recommendations within 5 minutes. A closed-loop feedback mechanism, leveraging 30 consecutive days of data, increased early warning accuracy from an initial 80% to 95%. A data mining algorithm analyzed 500 security incidents over the past year and identified 10 key risk factor combinations. These findings were integrated into the full lifecycle management data, providing a basis for long-term security strategy formulation.
[0116] The above describes the personnel status recognition method based on target recognition detection in the embodiment of the present application. The following describes the personnel status recognition device based on target recognition detection in the embodiment of the present application. Figure 2 In the embodiment of the present application, an embodiment of a personnel status recognition device based on target recognition detection includes:
[0117] The acquisition module 201 is used to perform parameterization processing on the acquired original image data to obtain intelligent BIM data;
[0118] An exchange module 202 is configured to perform multidisciplinary data exchange processing on the intelligent BIM data to obtain target detection data, wherein the target detection data includes personnel location information, safety equipment identification results, and dangerous area markings;
[0119] An analysis module 203 is configured to perform intelligent posture analysis on the human body information in the target detection data to obtain optimized posture data;
[0120] Identification module 204, configured to perform behavior recognition processing on the optimized posture data to obtain behavior analysis data, wherein the behavior analysis data includes the type of violation behavior, the determination of the climbing state, and the risk level assessment;
[0121] Evaluation module 205, configured to perform fuzzy logic evaluation processing on the behavior analysis data to obtain safety compliance data;
[0122] The early warning module 206 is used to perform intelligent early warning processing on the safety compliance data and real-time monitoring data to obtain full life cycle management data, wherein the full life cycle management data includes real-time early warning information, intervention measure recommendations and safety incident reports.
[0123] Through the collaborative efforts of these components, the raw image data is parameterized to generate intelligent BIM data. This step transforms 2D images into 3D models, laying the foundation for subsequent spatial analysis and target location, while also improving data visualization and comprehension. Next, the intelligent BIM data undergoes multidisciplinary data exchange processing to generate target detection data, including personnel location information, safety equipment identification results, and hazardous area markings. This step enables comprehensive perception and analysis of the work scene, providing detailed environmental information for safety management. The human body information in the target detection data is then subjected to intelligent posture analysis to generate optimized posture data. This step accurately captures worker movements and postures, helping to identify potentially unsafe behaviors. Subsequently, the optimized posture data undergoes behavior recognition processing to generate behavioral analysis data, including violation type, height determination, and risk level assessment. This step enables intelligent identification and risk assessment of worker behavior, enabling timely identification of safety hazards. Fuzzy logic evaluation is then performed on the behavioral analysis data to generate safety compliance data. This step uses a fuzzy inference system to comprehensively evaluate worker behavior, providing a more objective and comprehensive assessment of safety compliance. Finally, intelligent early warning processing is performed on safety compliance data and real-time monitoring data, resulting in full lifecycle management data, including real-time warning information, intervention recommendations, and safety incident reports. This step completes the closed loop from data analysis to safety management, enabling not only timely warnings but also specific intervention recommendations and supporting the formulation of long-term safety management strategies. This entire approach, driven by data and using intelligent algorithms, shifts the safety management model from passive response to proactive prevention, significantly improving the accuracy and real-time nature of safety management.
[0124] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying personnel status based on target recognition detection, characterized in that: The personnel status recognition method based on target recognition detection includes: Perform parameterized processing on the collected original image data to obtain intelligent BIM data; The step of performing parameterized processing on the collected original image data to obtain intelligent BIM data includes: Performing image denoising on the original image data to obtain denoised image data, and performing histogram equalization on the denoised image data to obtain enhanced image data; performing geometric correction on the enhanced image data to obtain corrected image data, and performing illumination compensation on the corrected image data to obtain compensated image data; performing feature extraction on the compensated image data through a convolutional neural network to obtain image feature data, and performing feature mapping on the image feature data to obtain BIM feature data; performing spatial information fusion on the BIM feature data to obtain fused feature data, and visualizing the fused feature data through graphical processing to obtain visual BIM data; performing semantic segmentation on the visual BIM data to obtain segmented BIM data, and performing parametric modeling on the segmented BIM data through attribute association processing to obtain intelligent BIM data; Performing multidisciplinary data exchange processing on the intelligent BIM data to obtain target detection data, wherein the target detection data includes personnel location information, safety equipment identification results, and danger zone markings; The step of performing multidisciplinary data exchange processing on the intelligent BIM data to obtain target detection data includes: Perform data format conversion processing on the intelligent BIM data to obtain standardized BIM data, and perform data classification processing on the standardized BIM data to obtain classified BIM data; perform feature extraction processing on the classified BIM data to obtain BIM feature vectors, and perform target detection processing on the BIM feature vectors through the YOLOv5 algorithm to obtain candidate target data; perform non-maximum suppression processing on the candidate target data to obtain screened target data, and perform coordinate mapping processing on the screened target data to obtain spatial positioning data; perform target classification processing on the spatial positioning data to obtain classified target data, and perform safety equipment identification processing on the classified target data through a deep learning algorithm to obtain equipment identification data; perform dangerous area division processing on the equipment identification data to obtain area marking data, and perform data integration processing on the area marking data to obtain target detection data; Performing intelligent posture analysis on human body information in the target detection data to obtain optimized posture data; Performing behavior recognition processing on the optimized posture data to obtain behavior analysis data, wherein the behavior analysis data includes the type of illegal behavior, the determination of the climbing state, and the risk level assessment; Performing fuzzy logic evaluation processing on the behavior analysis data to obtain safety compliance data; Intelligent early warning processing is performed on the safety compliance data and real-time monitoring data to obtain full life cycle management data, wherein the full life cycle management data includes real-time early warning information, intervention measure recommendations and safety incident reports.
2. The method for identifying a person's status based on target recognition detection according to claim 1, characterized in that: The intelligent posture analysis and processing of the human body information in the target detection data to obtain optimized posture data includes: Performing human region extraction processing on the target detection data to obtain human ROI data, and performing feature extraction processing on the human ROI data through a multi-scale feature pyramid network to obtain multi-scale feature data; The multi-scale feature data is processed by performing key point positioning processing by introducing the HRNet algorithm with an attention mechanism to obtain two-dimensional key point data, and the two-dimensional key point data is processed by performing time series information fusion processing to obtain continuous posture data; The expression for key point positioning processing of the multi-scale feature data by the HRNet algorithm using the attention mechanism is: , Among them, K represents the key point coordinates, A represents the attention weight, and F represents the feature map. represents element-wise multiplication, is a differentiable operation; Performing occlusion inference processing on the continuous posture data to obtain completed posture data, and performing three-dimensional reconstruction processing on the completed posture data using a monocular depth estimation algorithm to obtain three-dimensional posture data; Performing posture angle calculation processing on the three-dimensional posture data to obtain angle feature data, and performing motion trajectory analysis processing on the angle feature data to obtain motion feature data; Confidence evaluation processing is performed on the motion feature data to obtain credibility data, and data fusion processing is performed on the credibility data to obtain optimized posture data.
3. The method for identifying a person's status based on target recognition detection according to claim 1, characterized in that: The optimized posture data is subjected to behavior recognition processing to obtain behavior analysis data, wherein the behavior analysis data includes the type of violation behavior, the determination of the climbing state, and the risk level assessment, including: Performing a spatiotemporal graph construction process on the optimized posture data to obtain posture graph structure data, and performing feature extraction process on the posture graph structure data through a spatiotemporal graph convolutional network to obtain spatiotemporal feature data; Performing multi-level convolution processing on the spatiotemporal feature data to obtain high-level motion feature data, and performing weight distribution processing on the high-level motion feature data through an attention mechanism to obtain weighted feature data; Performing behavior classification processing on the weighted feature data to obtain violation behavior type data, and performing climbing state judgment processing on the violation behavior type data to obtain climbing state data; Performing height estimation processing on the climbing status data to obtain height information data, and performing safety risk assessment processing on the height information data using a risk assessment model to obtain risk level data; Perform data integration processing on the risk level data to obtain behavior analysis data.
4. The method for identifying a person's status based on target recognition detection according to claim 1, characterized in that: The performing of fuzzy logic evaluation processing on the behavior analysis data to obtain safety compliance data includes: Performing fuzzy rule base construction processing on the behavior analysis data to obtain fuzzy rule set data, and performing input variable fuzzification processing on the fuzzy rule set data to obtain input fuzzy set data; The Mamdani fuzzy inference system with a dynamic weight adjustment mechanism is used to perform fuzzy inference processing on the input fuzzy set data to obtain inference result data, and the inference result data is dynamically adjusted in weight to obtain weighted result data; wherein the dynamic weight adjustment mechanism is used to dynamically adjust the weight value of the rule according to the characteristics of the input data; Performing fuzzy output generation processing on the weighted result data to obtain fuzzy output set data, and performing defuzzification processing on the fuzzy output set data by a centroid method to obtain accurate scoring data; Performing multi-dimensional evaluation processing on the precise scoring data to obtain multi-dimensional evaluation data, and performing comprehensive analysis processing on the multi-dimensional evaluation data to obtain comprehensive evaluation data; Safety compliance calculation is performed on the comprehensive evaluation data to obtain safety compliance data.
5. The method for identifying a person's status based on target recognition detection according to claim 1, characterized in that: The intelligent early warning processing of the safety compliance data and real-time monitoring data is performed to obtain full life cycle management data, wherein the full life cycle management data includes real-time early warning information, intervention measure recommendations and safety incident reports, including: Performing data fusion processing on the safety compliance data and the real-time monitoring data to obtain fusion status data, and performing risk level assessment processing on the fusion status data using a random forest algorithm to obtain dynamic risk data; Performing a multi-threshold warning mechanism setting process on the dynamic risk data to obtain warning threshold data, and performing a warning trigger judgment process on the warning threshold data to obtain warning trigger data; Performing personalized warning information processing on the warning trigger data to obtain personalized warning data, and performing warning information generation processing on the personalized warning data through a multimodal output algorithm to obtain multimodal warning data; Performing intelligent intervention decision processing on the multimodal warning data through a decision tree algorithm to obtain intervention suggestion data, and performing remote expert support processing on the intervention suggestion data to obtain expert guidance data; The expert guidance data is subjected to closed-loop feedback processing to obtain optimization strategy data, and the optimization strategy data is subjected to event analysis processing through a data mining algorithm to obtain full life cycle management data.
6. A device for identifying a person's state based on target recognition detection, for implementing the method for identifying a person's state based on target recognition detection as claimed in any one of claims 1 to 5, characterized in that: The personnel status recognition device based on target recognition detection includes: The acquisition module is used to perform parameterized processing on the collected original image data to obtain intelligent BIM data; The acquisition module is specifically used to: perform image denoising on the original image data to obtain denoised image data, and perform histogram equalization on the denoised image data to obtain enhanced image data; perform geometric correction on the enhanced image data to obtain corrected image data, and perform illumination compensation on the corrected image data to obtain compensated image data; perform feature extraction on the compensated image data through a convolutional neural network to obtain image feature data, and perform feature mapping on the image feature data to obtain BIM feature data; perform spatial information fusion on the BIM feature data to obtain fused feature data, and perform visualization on the fused feature data through graphical processing to obtain visualized BIM data; perform semantic segmentation on the visualized BIM data to obtain segmented BIM data, and perform parametric modeling on the segmented BIM data through attribute association processing to obtain intelligent BIM data; an exchange module for performing multidisciplinary data exchange processing on the intelligent BIM data to obtain target detection data, wherein the target detection data includes personnel location information, safety equipment identification results, and danger zone markings; The exchange module is specifically used to: perform data format conversion processing on the intelligent BIM data to obtain standardized BIM data, and perform data classification processing on the standardized BIM data to obtain classified BIM data; perform feature extraction processing on the classified BIM data to obtain BIM feature vectors, and perform target detection processing on the BIM feature vectors through the YOLOv5 algorithm to obtain candidate target data; perform non-maximum suppression processing on the candidate target data to obtain screened target data, and perform coordinate mapping processing on the screened target data to obtain spatial positioning data; perform target classification processing on the spatial positioning data to obtain classified target data, and perform safety equipment identification processing on the classified target data through a deep learning algorithm to obtain equipment identification data; perform dangerous area division processing on the equipment identification data to obtain area marking data, and perform data integration processing on the area marking data to obtain target detection data; An analysis module, configured to perform intelligent posture analysis on the human body information in the target detection data to obtain optimized posture data; an identification module for performing behavior recognition processing on the optimized posture data to obtain behavior analysis data, wherein the behavior analysis data includes the type of violation behavior, the determination of the climbing state, and the risk level assessment; An evaluation module, configured to perform fuzzy logic evaluation processing on the behavior analysis data to obtain safety compliance data; The early warning module is used to perform intelligent early warning processing on the safety compliance data and real-time monitoring data to obtain full life cycle management data, wherein the full life cycle management data includes real-time early warning information, intervention measure recommendations and safety incident reports.
Citation Information
Patent Citations
Construction site risk intelligent identification method and system based on BIM model
CN117474321A
Construction personnel warning protection system based on attitude risk degree analysis
CN117711131A