Construction personnel dangerous area border crossing detection method and system based on neural network
Through the neural network-based detection method for dangerous areas of construction personnel, the limitations of traditional manual supervision methods at the construction site are solved, and intelligent identification, evaluation and early warning of dangerous areas of construction personnel are realized, and the safety management efficiency and reliability of the construction site are improved.
Patent Information
- Application Number
- CN202510093889.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The traditional manual supervision method has problems such as visual fatigue, high cost, strong subjectivity and lack of quantitative evaluation and data analysis in the cross-border monitoring of dangerous areas at the construction site, resulting in low safety management efficiency and reliability.
The neural network-based construction personnel's dangerous area out-of-bounds detection method is adopted, and the risk level is divided and visualized by keyframe extraction of video streams, target detection of YOLO11 neural network, target tracking of ByteTrack algorithm, spline interpolation is used to build a boundary model of dangerous area, ray projection is used to conduct out-of-bounds determination and behavioral analysis. Finally, risk level division and visualization are performed through a multi-level early warning mechanism.
It realizes intelligent identification, evaluation and early warning of dangerous areas of construction personnel, reduces the cost of manual supervision, improves the accuracy and real-timeness of inspection, and improves the safety management efficiency and reliability of construction sites.
Smart Images

Figure CN120014549A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular to a method and system for detecting construction personnel crossing a dangerous area based on a neural network. Background Art
[0002] With the rapid development of the power industry, the safety maintenance of power facilities has become increasingly important. Traditional monitoring of construction workers crossing the dangerous area mainly relies on manual supervision. Supervisors use their naked eyes to judge whether construction workers have entered the dangerous area by watching the monitoring video in real time. This manual supervision method has been widely used in industrial practice. By arranging special safety supervisors, the safety status of the construction site is monitored and managed in real time.
[0003] However, traditional manual supervision methods have obvious limitations. First, the human eye is prone to visual fatigue when staring at the monitoring screen for a long time, which leads to reduced supervision efficiency and the risk of missed detection and false detection. Secondly, special supervisors are required to conduct 24-hour uninterrupted monitoring, which has high human resource costs. In addition, manual judgment is highly subjective, and it is impossible to quantitatively evaluate and classify cross-border behaviors, and there is a lack of systematic recording and data analysis capabilities for cross-border incidents. These problems seriously restrict the efficiency and reliability of safety management at power construction sites. Summary of the invention
[0004] The present invention provides a method and system for detecting construction personnel crossing dangerous areas based on a neural network. While reducing the cost of manpower supervision, the present invention provides an accurate and reliable detection function for construction personnel crossing dangerous areas, thereby realizing intelligent identification, evaluation and early warning of crossing-border behaviors.
[0005] According to a first aspect of the present disclosure, a method for detecting construction personnel crossing a dangerous area based on a neural network is provided, comprising: extracting a key frame sequence from a collected construction site video stream to obtain a standardized image frame sequence; performing multi-scale feature extraction and target detection processing on the standardized image frame sequence through a YOLO11 neural network to obtain a detection result including target position information and a confidence score; performing target tracking processing on the detection result through a ByteTrack algorithm to obtain tracking data including a target ID and a motion trajectory; constructing a dangerous area boundary model through a spline interpolation algorithm based on a boundary point sequence collected through human-computer interaction to obtain dangerous area data; performing cross-border judgment and behavior analysis through a ray projection algorithm based on the tracking data and the dangerous area data to obtain a cross-border behavior assessment result; and performing risk level classification and visualization processing through a multi-level warning mechanism based on the cross-border behavior assessment result to obtain an enhanced video stream and a warning data packet.
[0006] According to a second aspect of the present disclosure, a construction worker dangerous area crossing detection system based on a neural network is provided, comprising:
[0007] An extraction module is used to extract key frame sequences from the collected construction site video stream to obtain a standardized image frame sequence;
[0008] A detection module, used to perform multi-scale feature extraction and target detection processing on the standardized image frame sequence through a YOLO11 neural network to obtain a detection result including target position information and a confidence score;
[0009] A tracking module, used to perform target tracking processing on the detection results through the ByteTrack algorithm to obtain tracking data including target ID and motion trajectory;
[0010] A construction module is used to construct a dangerous area boundary model through a spline interpolation algorithm according to a boundary point sequence collected by human-computer interaction to obtain dangerous area data;
[0011] An analysis module, used to perform cross-border judgment and behavior analysis through a ray casting algorithm according to the tracking data and the dangerous area data, and obtain a cross-border behavior assessment result;
[0012] The processing module is used to classify and visualize the risk levels according to the cross-border behavior assessment results through a multi-level warning mechanism to obtain enhanced video streams and warning data packets.
[0013] The present invention realizes efficient preprocessing and standardization of video data by extracting key frame sequences from the collected construction site video stream, and provides reliable input data for subsequent target detection. The YOLO11 neural network is used for multi-scale feature extraction and target detection, which fully utilizes the advantages of deep learning models in target recognition and accurately obtains the location information and confidence score of construction personnel. The ByteTrack algorithm is introduced for target tracking processing, which can not only assign a unique ID to each detected target, but also generate a continuous and stable motion trajectory, effectively solving the tracking break problem of traditional tracking algorithms under target occlusion and detection quality fluctuations. Based on human-computer interaction, the boundary point sequence is collected and the dangerous area boundary model is constructed through the spline interpolation algorithm, which provides a flexible and accurate definition method of dangerous areas to meet the safety management needs of different construction scenarios. Through the ray projection algorithm for cross-border judgment and behavior analysis, an objective and accurate cross-border behavior evaluation mechanism is established, which converts subjective safety judgment into quantifiable evaluation results. Finally, the risk level division and visualization processing are carried out through a multi-level warning mechanism, which not only realizes the timely warning of cross-border behavior, but also provides intuitive visual feedback and complete data records, greatly improving the safety management efficiency and reliability of the construction site. The entire solution realizes the full process automation from video acquisition, target detection, trajectory tracking to cross-border judgment and early warning processing, effectively reducing the cost of manual supervision and improving the accuracy and real-time performance of cross-border detection in dangerous areas.
[0014] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:
[0016] Figure 1 A flowchart of a method for detecting construction personnel crossing a dangerous area based on a neural network according to an embodiment of the present disclosure is shown;
[0017] Figure 2 A network structure diagram of a YOLO11 neural network according to an embodiment of the present disclosure is shown;
[0018] Figure 3 A flow chart of cross-border detection according to an embodiment of the present disclosure is shown;
[0019] Figure 4A block diagram of a construction worker dangerous area crossing detection system based on a neural network according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0021] In addition, the term "and / or" in this article is only a description of the association relationship between the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0022] Figure 1 FIG. 1 is a flow chart of a method 100 for detecting construction personnel crossing a dangerous area based on a neural network in an embodiment of the present disclosure. Figure 1 As shown, the method 100 includes:
[0023] S110: extracting a key frame sequence from the collected construction site video stream to obtain a standardized image frame sequence;
[0024] Optionally, the construction site video stream is segmented by a frame rate controller to obtain an initial image frame set; the initial image frame set is resolution standardized to adjust the image size to a standard size to obtain image data of uniform size; the image data of uniform size is subjected to illumination compensation by a brightness equalization algorithm to obtain image data with balanced illumination; the image data with balanced illumination is subjected to noise suppression by an adaptive filtering algorithm to obtain image data after noise reduction; the image data after noise reduction is scored and calculated by an image quality assessment algorithm to screen out candidate images whose quality scores exceed a preset threshold; the candidate images are arranged and organized in time sequence to obtain the standardized image frame sequence.
[0025] The video stream is processed by the frame rate controller, and the video stream is segmented and sampled at fixed time intervals. The frame rate controller extracts video frames at a sampling frequency of 25 frames / second according to the timestamp information of the video to form an initial image frame set. In practical applications, for a 30-minute construction site video, a total of 45,000 initial image frames are generated at a sampling rate of 25 frames / second. The obtained initial image frame set is subjected to resolution standardization, and the image size is uniformly adjusted to 1920×1080 pixels. Resolution standardization uses a bilinear interpolation algorithm to determine the new pixel value by calculating the weighted average of the four source pixels around the target pixel position. For each target pixel point (x, y), its pixel value P(x, y) is calculated by the linear combination of the four surrounding source pixels, and the weight coefficient is determined by the relative position of the target point to the source pixel point. For example, the original image is 1280×720 pixels, and after bilinear interpolation, a uniform size image data of 1920×1080 pixels is obtained.
[0026] Then, the uniform-sized image data is illuminated by brightness equalization algorithm. The brightness equalization adopts the histogram equalization method. First, the grayscale histogram of the image is calculated, and the number of pixels at each grayscale level is counted. Then, the grayscale value is remapped by the cumulative distribution function to make the grayscale distribution of the image more uniform. For construction site images with uneven illumination, such as when the local area is too dark or too bright, the brightness equalization process is used to obtain illumination-balanced image data with more reasonable contrast. In view of the noise problem in the image, the adaptive filtering algorithm is used for noise reduction. The adaptive filter dynamically adjusts the filtering parameters according to the statistical characteristics of the local image area, and maintains the edge and detail information of the image while removing noise. The filter calculates the local mean and variance of the neighborhood area of each pixel, and adaptively determines the filtering strength in combination with the estimated variance of the noise. For different degrees of noise pollution, the adaptive filter can generate appropriate filtering parameters to obtain image data with good noise reduction effect.
[0027] The image quality assessment algorithm calculates multiple indicators such as image clarity, signal-to-noise ratio and brightness distribution, and comprehensively scores the processed images. The evaluation indicators include gradient amplitude statistics, spectrum analysis and brightness variance. The scoring results are compared with the preset quality threshold to screen out candidate images with higher quality scores. In practical applications, by setting a reasonable quality threshold, image frames with blur, severe noise or abnormal lighting can be effectively eliminated. The screened candidate images are rearranged and organized in time sequence to construct a standardized image frame sequence. The image frames are sorted by timestamp information to ensure the temporal continuity of the image sequence, which is convenient for subsequent target detection and tracking processing. For video surveillance at construction sites, the standardized image sequence has uniform resolution, appropriate brightness contrast and low noise level, providing a reliable data basis for subsequent dangerous area boundary crossing detection.
[0028] For example, during the maintenance of power equipment at a construction site, the original surveillance video may have problems such as screen shaking, lighting changes, and noise interference. Through the above standardized processing flow, the original video with a resolution of 1280×720 and 30 frames / second for 30 minutes is first sampled at 25 frames / second to obtain 45,000 initial image frames. After resolution standardization, it is uniformly adjusted to 1920×1080 pixels. For areas with insufficient local light, brightness equalization processing makes the outline of the staff more clearly visible. The adaptive filtering algorithm processes the noise characteristics of different areas, effectively suppressing noise while maintaining edge details. Image quality assessment selects 42,000 image frames of qualified quality, and reorganizes them in time sequence to obtain a standardized image frame sequence.
[0029] S120: performing multi-scale feature extraction and target detection processing on the standardized image frame sequence through a YOLO11 neural network to obtain a detection result including target position information and a confidence score;
[0030] Optionally, the standardized image frame sequence is subjected to convolution operation and feature aggregation processing through the YOLO11 backbone network to obtain a multi-level feature map; the multi-level feature map is subjected to feature fusion processing through the C3k2 module, wherein a standard bottleneck structure or a C3 module is selected to perform feature extraction according to the C3k parameter to obtain fused feature data; the fused feature data is subjected to point-wise spatial attention calculation through the C2PSA module, and the feature map is multiplied by the attention weight to obtain enhanced feature data; the enhanced feature data is subjected to upsampling and feature splicing processing through the Neck network to obtain multi-scale target feature data; the multi-scale target feature data is subjected to target detection and coordinate regression calculation through the Head detector to obtain target frame coordinates and category probability; the target frame coordinates and category probability are subjected to non-maximum suppression processing and a confidence score is calculated to obtain the detection result containing target position information and a confidence score.
[0031] Among them, the YOLO11 neural network consists of three main parts: the backbone network (Backbone), the neck network (Neck) and the detection head (Head). Figure 2 As shown, it is a network structure diagram of the YOLO11 neural network in the embodiment of the present application. Accurate positioning of construction workers is achieved through layer-by-layer feature extraction and fusion. The backbone network first extracts features from the input standardized image frame sequence, and is composed of multiple convolutional layers and feature aggregation modules. The convolutional layer scans the image through a sliding window to extract local feature information, and the window size is 3×3 or 1×1. The feature aggregation module includes a C3k2 module, an SPPF module, and a C2PSA module, wherein the C3k2 module is an improved feature fusion structure, and its working mode is controlled by the c3k parameter: when c3k is False, a standard bottleneck structure is used for feature extraction; when c3k is True, the C3 module is enabled for feature processing. The bottleneck structure reduces the amount of computation by 1×1 convolution dimensionality reduction, 3×3 convolution feature extraction, and 1×1 convolution dimensionality increase, while the C3 module uses a more complex branch structure to enhance feature expression capabilities.
[0032] The C2PSA module introduces a point-based spatial attention mechanism to calculate the attention weight for each spatial position of the feature map. First, the input feature map is divided into multiple channel groups, and self-attention calculation is performed in each group to obtain a weight matrix that reflects the importance of space. The weight matrix is then multiplied by the original feature map to highlight the feature response of important areas. The attention mechanism enables the network to adaptively focus on key areas in the image and improve the detection accuracy of construction workers. The neck network adopts a feature pyramid structure to fuse the multi-scale features extracted by the backbone network. The resolution of high-level feature maps is enlarged to the same level as low-level features through upsampling operations. Figure 1This multi-scale feature fusion strategy can simultaneously utilize high-level semantic information and low-level detail information to enhance the detection capability of objects of different scales.
[0033] The detection head network performs target detection and coordinate regression on the fused feature map. The target category probability and bounding box coordinate offset of each position are predicted through a multi-layer convolutional network. The prediction results are decoded to obtain the actual coordinates of the target box. Then, the non-maximum suppression algorithm is used to eliminate overlapping boxes and retain the detection box with the highest confidence. The final output detection result contains the category, location coordinates and confidence score of each target.
[0034] For example, for a standardized image of 1920×1080 pixels at a construction site, the backbone network first converts it into multiple feature maps through convolution operations, with the resolution decreasing and the number of channels increasing. The C3k2 module selects the appropriate feature extraction method according to the setting of the c3k parameter, such as enabling the C3 module to obtain a more detailed feature representation when detecting small targets. The C2PSA module calculates point-based spatial attention, and the area where the construction workers are located will receive a higher attention weight. The neck network combines feature maps of different scales to generate a feature representation containing multi-scale information. The detection head analyzes the feature map and outputs the bounding box coordinates and confidence of the construction workers. For multiple detection boxes in the image, the non-maximum suppression algorithm is used to screen out the optimal detection results, and finally generate clear target positioning information. The whole process ensures both the accuracy of the detection and the real-time requirements.
[0035] S130: performing target tracking processing on the detection result by using the ByteTrack algorithm to obtain tracking data including the target ID and motion trajectory;
[0036] Optionally, the detection results are divided into a high-score detection set and a low-score detection set by a score analyzer; the motion state of the target position in the high-score detection set is predicted by a Kalman predictor to obtain predicted position data; the correlation between the predicted position data and the high-score detection set in the new frame is calculated by a similarity calculation module to obtain a target matching pair; the detection results in the low-score detection set and the unmatched trajectories are secondary matched by a trajectory association processor to obtain a supplementary matching pair; the trajectory state of the target matching pair and the supplementary matching pair is updated by a trajectory manager, including assigning a target ID to the trajectories that have been successfully matched continuously and keeping count of the unmatched trajectories; the target ID is combined with the corresponding detection result position sequence to obtain the tracking data including the target ID and the motion trajectory.
[0037] Among them, the ByteTrack algorithm is responsible for the target tracking task in the detection of construction workers crossing the boundary in the dangerous area, and converts the detection results into continuous motion trajectories. First, the detection results are divided by the confidence threshold through the score analyzer, and the threshold is set to 0.7, and the detection results are divided into a high-score detection set and a low-score detection set. The confidence of the target detection box in the high-score detection set is greater than 0.7, indicating that the detection accuracy is high; the confidence of the target detection box in the low-score detection set is between 0.1 and 0.7. Although the detection confidence of these targets is not high, they still contain valuable tracking information. The Kalman predictor predicts the motion state of the target position in the high-score detection set, using a constant speed motion model. The state vector contains the center point coordinates (x, y) and velocity components (vx, vy) of the target, and estimates the possible position of the target in the next frame through the state of the previous frame. The prediction process takes into account the influence of measurement noise and system noise, and dynamically adjusts the Kalman gain to balance the weight of the predicted value and the observed value, so as to obtain more accurate predicted position data.
[0038] The similarity calculation module is responsible for associating and matching the predicted position data with the high-scoring detection set in the new frame. The calculation uses two indicators: one is the intersection over union (IOU) of the detection frame, which reflects the spatial overlap of the target frame; the other is the similarity of the appearance features of the target, which measures the proximity of the feature vectors through the cosine distance. The two similarity indicators are weighted and combined to obtain a comprehensive correlation score. For each predicted position, the detection frame with the highest correlation score that exceeds the matching threshold is selected as the matching pair.
[0039] For tracks that fail to find a match in the high-score detection set, the track association processor attempts to perform a secondary match with the detection box in the low-score detection set. This step is particularly important for dealing with situations where the target is temporarily occluded or the detection quality is reduced. The secondary match is also based on the similarity calculation of spatial position and appearance features, but a lower matching threshold is used to obtain a complementary matching pair. The track manager is responsible for maintaining and updating the status of all tracks. For newly detected targets, a new track is created and a unique target ID is assigned. For existing tracks, the status is updated according to the continuous matching situation: the status of the track that has been successfully matched continuously is set to active, and its position and motion information are updated; the unmatched track enters the hold state, records the number of unmatched frames, and deletes the track when the preset number of frames (usually 30 frames) is exceeded.
[0040] Finally, the target ID is combined with the corresponding position sequence to form complete tracking data. The tracking data not only contains the spatial location information of the target, but also records the movement trajectory in the time dimension, providing basic data support for subsequent cross-border behavior analysis.
[0041] For example, in a construction site video surveillance, a worker moves from a safe area to a dangerous area. The YOLO11 network first detects the person and outputs a detection frame containing the position coordinates and confidence. The score analyzer divides the detection results into a high-score detection set (confidence>0.7) and a low-score detection set (0.1<confidence<0.7). The Kalman predictor predicts the possible position in the next frame based on the worker's historical position and speed. The similarity calculation module matches the predicted position with the high-score detection frame in the new frame and finds the best matching pair by calculating the spatial overlap and feature similarity. When the worker passes through certain obstructions (such as equipment or structures) and the detection quality decreases, the confidence of the detection frame may drop to 0.5. At this time, the secondary matching mechanism of the trajectory association processor is used to maintain the continuity of the trajectory using the detection frame in the low-score detection set. The trajectory manager assigns a unique target ID to the worker and continuously updates his motion trajectory. The entire process achieves stable tracking of construction workers, and the continuity of the trajectory can be maintained even in the case of partial occlusion or fluctuations in detection quality.
[0042] S140: constructing a dangerous area boundary model through a spline interpolation algorithm according to the boundary point sequence collected by human-computer interaction to obtain dangerous area data;
[0043] Optionally, the click positions generated by human-computer interaction are sampled and recorded by a mouse click sequence recorder to obtain the boundary point sequence; the boundary point sequence is mapped from pixel coordinates to actual coordinates by a coordinate converter to obtain an actual space coordinate sequence; the actual space coordinate sequence is curve-fitted by a cubic spline calculator to obtain a smooth boundary curve equation; the boundary curve equation is tested for regional closure by a connected domain analyzer, and non-closed areas are automatically closed to obtain a closed area curve; the closed area curve is discretized and sampled by a grid projector to evenly generate boundary reference points on the curve to obtain regional boundary sampling data; the regional boundary sampling data is spatially indexed by a spatial index builder to obtain the dangerous area data.
[0044] The boundary points of the dangerous area are collected by the mouse click sequence recorder. The operator marks the boundary outline of the dangerous area by clicking the mouse on the monitoring screen, and the coordinate position of each click is recorded to form a boundary point sequence. The recorder responds to the click event, captures the pixel coordinate value (x, y) of the mouse click, and organizes these coordinate points into an ordered sequence according to the time sequence of the click.
[0045] The coordinate converter converts the recorded pixel coordinates into real space coordinates. This conversion is based on the calibration parameters of the camera, including the intrinsic matrix and the extrinsic matrix. The intrinsic matrix contains the focal length and principal point coordinate information, and the extrinsic matrix describes the position and posture of the camera in the world coordinate system. Through these parameters, the two-dimensional pixel coordinates are mapped to the three-dimensional world coordinate system to obtain the position coordinates of the boundary points in the real space. The cubic spline calculator smoothly interpolates the real space coordinate sequence to generate a continuous boundary curve. Spline interpolation uses a piecewise cubic polynomial function to construct smooth curve segments between adjacent control points. Each curve segment satisfies the continuity conditions of the position, first-order derivative, and second-order derivative at the endpoint to ensure the smoothness of the entire boundary curve. The interpolation process takes into account the spatial distribution of the control points and adaptively adjusts the tension coefficient of the curve to avoid overfitting.
[0046] The connected domain analyzer performs a closure check on the boundary curve. By calculating the distance between the first and last points of the curve, it is determined whether the boundary forms a closed area. For non-closed boundary curves, a curve segment connecting the first and last points is automatically added. Bezier curves are used for closure processing to ensure a smooth transition with the original boundary curve at the connection point. The closure check also includes self-intersection detection to avoid self-intersection of the boundary curve. The grid projector discretizes and samples the closed area curve. According to the length and complexity of the curve, the sampling interval is determined, and boundary reference points are evenly generated on the curve. The sampling process takes into account the curvature change of the curve and increases the sampling density in areas with larger curvature. At the same time, the sampling points are screened and redundant sampling points are removed to obtain boundary sampling data that can accurately express the boundary shape and is easy to calculate.
[0047] The spatial index builder builds an efficient spatial retrieval structure based on the boundary sampling data. The boundary points are organized into a hierarchical tree structure using spatial indexing methods such as quadtree or R-tree. The index structure supports fast point relationship queries, which facilitates the subsequent determination of whether the target point is located in the danger zone. The construction process includes determining the tree depth, node splitting conditions and merging strategies, balancing retrieval efficiency and storage overhead.
[0048] For example, at a power facility maintenance site, it is necessary to define a dangerous area around high-voltage equipment. The operator uses the mouse to click on 12 boundary points on the monitoring screen to depict an irregular polygonal area. These click positions are first recorded as pixel coordinates and then converted to actual space coordinates based on the camera calibration parameters. The cubic spline interpolator connects the 12 control points into a smooth boundary curve. The connected domain analysis finds that there is a 0.5-meter gap between the first and last points, and a smooth closed segment is automatically added. The closed boundary curve is uniformly sampled, and a boundary reference point is generated every 0.3 meters, forming a total of 80 sampling points. Finally, a quadtree spatial index is constructed to organize the sampling points in layers to support subsequent fast area queries.
[0049] S150: performing cross-border determination and behavior analysis through a ray casting algorithm according to the tracking data and the dangerous area data to obtain a cross-border behavior assessment result;
[0050] Optionally, the motion trajectory in the tracking data is analyzed for the current frame target position by a position extractor to obtain target center point coordinate data; a horizontal ray is constructed for the target center point coordinate data by a ray generator, and the boundary point connection line of the dangerous area data is used as a judgment basis to obtain intersection data of the ray and the boundary; parity statistics are performed on the intersection data of the ray and the boundary by an intersection counter, and the inside and outside of the target position is determined by the parity of the number of intersections to obtain a position judgment result; the target ID in the tracking data is time-series associated by a trajectory analyzer, and the continuous crossing time of each target is counted to obtain crossing duration data; the position judgment result and the crossing duration data are comprehensively analyzed by a behavior evaluator, the severity of the crossing behavior is calculated, and the behavior score data is obtained; the risk analyzer performs quantitative calculation of the degree of danger on the behavior score data to obtain the crossing behavior assessment result.
[0051] Among them, the position extractor extracts the position information of the target in the current frame from the tracking data. Figure 3 As shown, it is a flow chart of out-of-bounds detection in an embodiment of the present application. In the present application, the tracking data includes a target ID and a corresponding motion trajectory. The real-time spatial position of the target center point is obtained by parsing the latest coordinate point in the trajectory data. The position information is expressed in three-dimensional coordinates (x, y, z), reflecting the exact position of the target in the monitoring scene. The ray generator constructs horizontal rays based on the target center point to realize the determination of the point relationship. The specific method is to construct a ray from the target center point along the horizontal direction (usually the positive direction of the x-axis is selected), and the ray is mathematically expressed in the form of a parametric equation. The boundary of the dangerous area is formed by discrete boundary points connected by line segments to form a closed polygon, and each boundary line segment can be expressed as a straight line equation between two endpoints.
[0052] The intersection counter performs statistical analysis on the intersections of rays and boundaries. By solving the intersections of the ray equation and each boundary segment equation, a series of intersection coordinates are obtained. Determine whether each intersection is valid (i.e., whether it falls on the line segment rather than the extension of the straight line), and record the number of valid intersections. According to the ray law, if the number of intersections between the rays emitted from a point and the boundary of the closed area is an odd number, the point is inside the area; if the number of intersections is an even number, it is outside the area. The trajectory analyzer performs timing correlation analysis on the target ID and calculates the duration of the out-of-bounds state. By checking the position determination results of the target in consecutive frames, the start time and the number of frames of the out-of-bounds state are counted. For each target ID, its historical out-of-bounds events are recorded, including the start time, end time, and duration. These timing data reflect the temporal characteristics of the out-of-bounds behavior.
[0053] The behavior evaluator comprehensively analyzes the location determination results and the cross-border duration data to evaluate the severity of the cross-border behavior. The evaluation indicators include the cross-border distance (the shortest distance between the target and the boundary), the cross-border duration (the duration of continuous cross-border) and the cross-border frequency (the number of cross-border times per unit time). These indicators are weighted and combined to calculate the scoring data reflecting the severity of the cross-border behavior. The risk analyzer performs quantitative analysis on the behavior scoring data to generate the cross-border behavior assessment results. According to the distribution characteristics of the scoring data, multiple risk level thresholds are set to divide the cross-border behavior into different danger levels. The risk analysis takes into account multiple dimensions of cross-border behavior, including the spatial dimension (cross-border distance), the temporal dimension (duration) and the behavioral dimension (cross-border frequency), so as to obtain a comprehensive risk assessment result.
[0054] For example, at a power facility maintenance site, the motion trajectory data of a worker is obtained through target tracking. The position extractor parses the center point coordinates of the worker's current position from the trajectory data. The ray generator constructs a horizontal ray from this point to the right, generating multiple intersections with the predefined dangerous area boundary. The intersection counter statistics find that the ray has three intersections with the boundary, and determines that the worker is inside the dangerous area based on the parity. The trajectory analyzer associates the worker's target ID and finds that this has been a continuous out-of-bounds state for 30 seconds. The behavior evaluator combines the 2-meter out-of-bounds distance and the 30-second duration to calculate a higher behavior score. The risk analyzer determines that this is a high-risk out-of-bounds behavior based on the score data, and requires immediate warning processing.
[0055] S160: Based on the cross-border behavior assessment results, risk level classification and visualization processing are performed through a multi-level warning mechanism to obtain enhanced video streams and warning data packets;
[0056] Optionally, the risk level analyzer performs threshold segmentation processing on the cross-border behavior assessment result, divides the risk degree into multiple levels, and obtains graded warning data; the spatiotemporal association processor marks the target position and timestamp of the graded warning data, establishes the spatiotemporal mapping relationship of the cross-border event, and obtains associated warning data; the visual mark generator performs graphical processing on the associated warning data, generates corresponding visual mark symbols for different risk levels, and obtains mark rendering data; the layer overlay processor performs synthesis processing on the mark rendering data and the original video stream, overlays the visual mark and the boundary of the dangerous area for display, and obtains the enhanced video stream; the data structuring processor extracts and organizes the associated warning data, organizes the cross-border event records in a unified format, and obtains a structured warning record; the structured warning record and the enhanced video stream are packaged and integrated by a data packager to obtain the warning data packet.
[0057] Among them, the risk level analyzer performs graded processing on the assessment results of the cross-border behavior. According to the risk score value of the cross-border behavior, multiple threshold intervals are set to divide the risk level into three levels: low risk (score 0-30), medium risk (score 31-70) and high risk (score 71-100). Through threshold segmentation processing, the continuous score value is mapped to discrete risk levels to form structured graded warning data. The spatiotemporal association processor adds spatiotemporal tag information to the graded warning data. Each warning data is associated with the target ID, occurrence timestamp and spatial position coordinates. The timestamp is recorded in a unified time format, accurate to milliseconds, which is convenient for subsequent time series analysis. The spatial position uses the actual coordinate system to record the three-dimensional coordinate value of the target, reflecting the precise location of the cross-border event. The spatiotemporal tag establishes the spatiotemporal mapping relationship of the cross-border event to form associated warning data with complete context information.
[0058] The visual marker generator is responsible for converting the associated warning data into visual graphic markers. For out-of-bounds events of different risk levels, visual marker symbols of different colors and shapes are generated: low risk uses yellow circular markers, medium risk uses orange triangle markers, and high risk uses red square markers. The size of the marker symbol increases with the increase of risk level to enhance visual prominence. At the same time, text annotations of the target ID and risk level are added next to the marker to generate complete marker rendering data. The layer overlay processor synthesizes the marker rendering data with the original video stream. First, the original video screen is displayed on the bottom layer, and then the boundary outline of the dangerous area is superimposed, and the scope of the dangerous area is highlighted by a translucent fill effect. The visual marker of the out-of-bounds warning is superimposed on the top layer, and the position of the marker is updated in real time as the target moves. The overlay of multiple layers uses Alpha blending to ensure the coordination of visual effects and generate an enhanced video stream with augmented reality effects.
[0059] The data structured processor organizes and organizes the associated warning data in a unified format. The warning information is encapsulated according to a predefined data structure, including fields such as event ID, target ID, timestamp, spatial coordinates, risk level, cross-border distance, and duration. Each cross-border event record follows the same structured format to facilitate data storage, query, and analysis. The processed structured warning record contains complete information about the cross-border event. The data packager integrates and encapsulates the structured warning record and the enhanced video stream. Using a unified data packet format, the warning record is saved as text data in JSON format, and the enhanced video stream is saved as H.264 encoded video data. An index relationship between the warning record and the video frame is established in the data packet to support rapid positioning of the corresponding video clip by timestamp. Finally, a warning data packet containing complete warning information is formed.
[0060] For example, at the maintenance site of power facilities, the boundary crossing detection system detected a worker entering a dangerous area. The risk level analyzer calculated the risk score of 85 points based on the distance (2 meters) and duration (30 seconds) of the worker's boundary crossing, and determined it to be a high risk level. The spatiotemporal correlation processor recorded the time (2024-01-16 14:30:25.345) and location coordinates (X=10.5, Y=8.2, Z=1.6) of the boundary crossing. The visual marker generator generates a red square marker for the high-risk event with a marker size of 40×40 pixels. The layer overlay processor overlays the red marker on top of the target position in the video screen and displays the red boundary of the dangerous area. The data structured processor generates a complete warning record, including the event number, target information, and risk assessment data. The data packager packages the warning record with the corresponding video clip to form a complete warning data packet.
[0061] Figure 4 FIG. 2 is a block diagram of a construction worker dangerous zone crossing detection system 200 based on a neural network according to an embodiment of the present disclosure. Figure 4 As shown, the device 200 includes:
[0062] An extraction module 210 is used to extract a key frame sequence from the collected construction site video stream to obtain a standardized image frame sequence;
[0063] A detection module 220 is used to perform multi-scale feature extraction and target detection processing on the standardized image frame sequence through a YOLO11 neural network to obtain a detection result including target position information and a confidence score;
[0064] A tracking module 230 is used to perform target tracking processing on the detection result through a ByteTrack algorithm to obtain tracking data including a target ID and a motion trajectory;
[0065] A construction module 240 is used to construct a dangerous area boundary model through a spline interpolation algorithm according to the boundary point sequence collected by human-computer interaction to obtain dangerous area data;
[0066] An analysis module 250 is used to perform cross-border determination and behavior analysis based on the tracking data and the dangerous area data through a ray casting algorithm to obtain a cross-border behavior assessment result;
[0067] The processing module 260 is used to perform risk level classification and visualization processing through a multi-level warning mechanism according to the cross-border behavior assessment result, and obtain an enhanced video stream and a warning data packet.
[0068] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0069] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for detecting construction workers crossing dangerous areas based on neural networks, characterized in that: include: Extract key frame sequences from the collected construction site video stream to obtain standardized image frame sequences; Performing multi-scale feature extraction and target detection processing on the standardized image frame sequence through the YOLO11 neural network to obtain a detection result including target position information and a confidence score; Performing target tracking processing on the detection results through the ByteTrack algorithm to obtain tracking data including target ID and motion trajectory; According to the boundary point sequence collected by human-computer interaction, the dangerous area boundary model is constructed through the spline interpolation algorithm to obtain the dangerous area data; According to the tracking data and the dangerous area data, cross-border determination and behavior analysis are performed through a ray casting algorithm to obtain a cross-border behavior assessment result; According to the assessment results of the cross-border behavior, risk level classification and visualization are carried out through a multi-level warning mechanism to obtain enhanced video streams and warning data packets.
2. The method for detecting construction personnel crossing a dangerous area based on a neural network according to claim 1, characterized in that: The key frame sequence is extracted from the collected construction site video stream to obtain a standardized image frame sequence, including: Segmenting the construction site video stream by a frame rate controller to obtain an initial image frame set; Performing resolution standardization processing on the initial image frame set, adjusting the image size to a standard size, and obtaining image data of uniform size; Performing illumination compensation processing on the image data of uniform size by using a brightness equalization algorithm to obtain image data with balanced illumination; Performing noise suppression processing on the illumination-balanced image data by using an adaptive filtering algorithm to obtain noise-reduced image data; The image data after noise reduction is scored and calculated by an image quality assessment algorithm, and candidate images having quality scores exceeding a preset threshold are screened out; The candidate images are arranged and organized in time sequence to obtain the standardized image frame sequence.
3. The method for detecting construction personnel crossing a dangerous area based on a neural network according to claim 1, characterized in that: The YOLO11 neural network is used to perform multi-scale feature extraction and target detection processing on the standardized image frame sequence to obtain a detection result including target position information and confidence score, including: Performing convolution operation and feature aggregation processing on the standardized image frame sequence through the YOLO11 backbone network to obtain a multi-level feature map; Performing feature fusion processing on the multi-level feature graphs through the C3k2 module, wherein the standard bottleneck structure or the C3 module is selected to perform feature extraction according to the C3k parameter to obtain fused feature data; Performing point-wise spatial attention calculation on the fused feature data through the C2PSA module, multiplying the feature map by the attention weight to obtain enhanced feature data; Upsampling and feature concatenation are performed on the enhanced feature data through a Neck network to obtain multi-scale target feature data; Performing target detection and coordinate regression calculation on the multi-scale target feature data through a Head detector to obtain target frame coordinates and category probability; Non-maximum suppression processing is performed on the target frame coordinates and category probability and a confidence score is calculated to obtain the detection result including the target position information and the confidence score.
4. The method for detecting construction personnel crossing a dangerous area based on a neural network according to claim 1, characterized in that: The target tracking process is performed on the detection result by the ByteTrack algorithm to obtain tracking data including the target ID and the motion trajectory, including: The detection results are divided into a high-score detection set and a low-score detection set by using a score analyzer according to a confidence threshold. Predicting the motion state of the target position in the high-score detection set by using a Kalman predictor to obtain predicted position data; The similarity calculation module calculates the correlation between the predicted position data and the high-score detection set in the new frame to obtain a target matching pair; Performing secondary matching calculation on the detection results in the low-score detection set and the unmatched trajectories through a trajectory association processor to obtain a supplementary matching pair; Update the trajectory status of the target matching pair and the supplementary matching pair through the trajectory manager, including assigning target IDs to the trajectories that have been matched successfully, and keeping counts of the unmatched trajectories; The target ID is combined with the corresponding detection result position sequence to obtain the tracking data including the target ID and the motion trajectory.
5. The method for detecting construction personnel crossing dangerous areas based on a neural network according to claim 1 is characterized in that: The method of constructing a dangerous area boundary model by using a spline interpolation algorithm based on a sequence of boundary points collected by human-computer interaction to obtain dangerous area data includes: The click positions generated by human-computer interaction are sampled and recorded by a mouse click sequence recorder to obtain the boundary point sequence; A coordinate converter is used to map pixel coordinates of the boundary point sequence to actual coordinates to obtain an actual space coordinate sequence. Performing curve fitting processing on the actual space coordinate sequence by a cubic spline calculator to obtain a smooth boundary curve equation; Performing a regional closure test on the boundary curve equation by a connected domain analyzer, automatically closing the non-closed area, and obtaining a closed area curve; Discretely sampling the closed area curve through a grid projector, evenly generating boundary reference points on the curve, and obtaining area boundary sampling data; The spatial index builder constructs a spatial index structure for the area boundary sampling data to obtain the dangerous area data.
6. The method for detecting construction personnel crossing dangerous areas based on a neural network according to claim 1, characterized in that: The step of performing cross-border determination and behavior analysis based on the tracking data and the dangerous area data by using a ray casting algorithm to obtain a cross-border behavior assessment result includes: Performing a current frame target position analysis on the motion trajectory in the tracking data by a position extractor to obtain target center point coordinate data; A ray generator is used to construct a horizontal ray for the target center point coordinate data, and a line connecting the boundary points of the dangerous area data is used as a determination reference to obtain intersection data of the ray and the boundary; The parity statistics of the intersection data of the ray and the boundary are performed by using an intersection counter, and the inside and outside of the target position are determined by the parity of the number of intersections to obtain a position determination result; The target IDs in the tracking data are temporally correlated by a trajectory analyzer, and the continuous crossing time of each target is counted to obtain the crossing duration data; Comprehensively analyzing the position determination result and the boundary crossing duration data through a behavior evaluator, calculating the severity of the boundary crossing behavior, and obtaining behavior score data; The risk analyzer performs a quantitative calculation of the degree of danger on the behavior scoring data to obtain the cross-border behavior assessment result.
7. The method for detecting construction personnel crossing dangerous areas based on a neural network according to claim 1, characterized in that: According to the cross-border behavior assessment results, risk level classification and visualization processing are performed through a multi-level warning mechanism to obtain enhanced video streams and warning data packets, including: The risk level analyzer performs threshold segmentation processing on the cross-border behavior assessment result, divides the risk degree into multiple levels, and obtains graded warning data; The hierarchical warning data is marked with target positions and timestamps by a spatiotemporal correlation processor, a spatiotemporal mapping relationship of the boundary crossing event is established, and associated warning data is obtained; Graphically processing the associated warning data through a visual marker generator, generating corresponding visual marker symbols for different risk levels, and obtaining marker rendering data; The mark rendering data and the original video stream are synthesized by a layer overlay processor, and the visual mark and the boundary of the dangerous area are overlaid and displayed to obtain the enhanced video stream; Extracting and organizing the associated warning data through a data structuring processor, arranging the boundary crossing event records in a unified format, and obtaining structured warning records; The structured warning record and the enhanced video stream are packaged and integrated by a data packager to obtain the warning data packet.
8. A construction worker danger zone crossing detection system based on a neural network, used to implement the construction worker danger zone crossing detection method based on a neural network as described in any one of claims 1 to 7, characterized in that: The neural network-based construction worker dangerous area crossing detection system includes: An extraction module is used to extract key frame sequences from the collected construction site video stream to obtain a standardized image frame sequence; A detection module, used to perform multi-scale feature extraction and target detection processing on the standardized image frame sequence through a YOLO11 neural network to obtain a detection result including target position information and a confidence score; A tracking module, used to perform target tracking processing on the detection results through the ByteTrack algorithm to obtain tracking data including target ID and motion trajectory; A construction module is used to construct a dangerous area boundary model through a spline interpolation algorithm according to a boundary point sequence collected by human-computer interaction to obtain dangerous area data; An analysis module, used to perform cross-border judgment and behavior analysis through a ray casting algorithm according to the tracking data and the dangerous area data, and obtain a cross-border behavior assessment result; The processing module is used to classify and visualize the risk levels according to the cross-border behavior assessment results through a multi-level warning mechanism to obtain enhanced video streams and warning data packets.
Citation Information
Cited By
Real-time monitoring and abnormity early warning shooting training safety system and method
CN120298979A
Image analysis early warning method and system based on visual large model
CN120655982A
Image analysis and early warning method and system based on visual large model
CN120655982B
Method and system for detecting personnel intrusion in operation dangerous area of bridge crane
CN120783437A
Neural network image recognition method and multistage feedback construction monitoring system
CN120808097A