Construction site safety risk real-time early warning method and system based on artificial intelligence
By generating separate driving identifiers and performing gradient segmentation, independent motion segments are extracted, and the attitude offset angle and center of gravity swing amplitude are calculated to construct independently identifiable dangerous behavior units. This solves the problem of difficulty in identifying individual actions when multiple people are densely packed on construction sites, and enables accurate early warning of safety risks at construction sites.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SICHUAN LVSHU CONSTR ENG CO LTD
- Filing Date
- 2026-06-18
- Publication Date
- 2026-07-24
AI Technical Summary
In real-time safety risk early warning at construction sites, when multiple people gather in dense crowds, the overlapping of individual outlines in the video perception results forms an adhesive boundary, making it difficult to identify and extract abnormal actions individually, thus resulting in dangerous behaviors not being captured and warned of in a timely manner.
By acquiring pixel displacement amplitude, contour edge transition frequency, and local occlusion overlap ratio, a separation-driven identifier is generated. Gradient segmentation and directional difference constraint processing are performed to extract independent motion segments. The attitude offset angle and center of gravity swing amplitude are calculated. Local rearrangement and dynamic contraction are implemented to define the dangerous behavior unit that can be independently identified. The warning information is then output through reverse mapping processing.
Effectively separating continuous adhesion areas in densely populated environments improves the completeness of abnormal movement identification and the timeliness of early warning response, ensures continuous perception of individual movement status, and enhances the accuracy and response speed of risk assessment at construction sites.
Smart Images

Figure CN122454512A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of safety risk early warning technology, specifically to a method and system for real-time early warning of safety risks at construction sites based on artificial intelligence. Background Technology
[0002] Real-time early warning of safety risks at construction sites refers to the continuous acquisition of personnel behavior, machinery operation status, and environmental change characteristics within the construction area by relying on sensing devices, video acquisition devices, and environmental monitoring units. Artificial intelligence technology is introduced to fuse and analyze multi-source data, constructing a risk evolution trajectory in the time dimension. By identifying abnormal patterns, deviation characteristics, and potential dangerous trends, it enables early judgment and dynamic alerts for risks such as falls from heights, machinery collisions, and violations of regulations. This allows for the output of early warning information and the triggering of intervention measures before the risk actually occurs, transforming construction site safety management from post-event response to pre-event control and dynamic process suppression.
[0003] The existing technology has the following shortcomings: During real-time safety risk early warning at construction sites, when multiple people gather densely in the construction area, the individual silhouettes in the video perception results will overlap and gradually form contiguous boundaries. This causes previously clearly separated personnel targets to be aggregated into a single continuous area during the identification process. As this continuous area continues to expand, the differences in actions between individuals will be masked by the overall motion characteristics, making it difficult to extract and accurately determine locally abnormal actions. Furthermore, the triggering characteristics of dangerous behaviors will be diluted by the overall behavioral trend, leading to a deviation in the identification results of high-risk actions. Ultimately, this results in individual dangerous behaviors not being captured in a timely manner and early warning information not being output.
[0004] For example, when multiple workers are concentrated around a hoisted object in a hoisting operation area, the outlines of their bodies gradually overlap in the video footage, forming a continuous, contiguous area. At this point, one worker is bending down to retrieve an object from below the hoisted object. This action should have been classified as high-risk; however, due to the overlapping outlines, the bending action is obscured by the overall standing posture of the group and categorized as a uniform movement, thus failing to identify the dangerous action separately. Furthermore, as the hoisted object sways slightly, the worker is actually within the potential strike range, but because the dangerous action was not identified and a warning was not triggered, the worker is ultimately exposed to a high-risk area without receiving timely alerts.
[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this application is to provide a method and system for real-time early warning of safety risks at construction sites based on artificial intelligence, so as to solve the problems in the background art mentioned above.
[0007] To achieve the above objectives, this application provides the following technical solution: a real-time early warning method for safety risks at construction sites based on artificial intelligence, comprising the following steps: The pixel displacement amplitude, contour edge jump frequency, and local occlusion overlap ratio in the densely populated construction area are obtained. Based on the pixel displacement amplitude, contour edge jump frequency, and local occlusion overlap ratio, a separation driving identifier is generated to characterize the degree of boundary adhesion in the state of personnel gathering. Based on the separation-driven identifier, gradient segmentation is performed on the continuous and connected regions within the image. Orientation difference constraints are introduced at each connection boundary to extract independent motion segments from the segmented regions and restore the compressed individual contour information. Calculate the posture offset angle and center of gravity swing amplitude corresponding to each independent motion segment, and perform local rearrangement based on the posture offset angle and center of gravity swing amplitude. During the rearrangement process, highlight non-consistent motion blocks and reveal abnormal motions that are covered by the group. Based on the non-uniform action blocks, dynamic contraction is performed at the corresponding positions, and spatial compression ratio constraints are superimposed during the contraction process to separate local abnormal actions from the continuous adhesion area and construct independently identifiable dangerous behavior units. Based on the dangerous behavior units, the original dense area is reverse-mapped, and the stripped action features are backfilled into the corresponding positions. The warning trigger threshold range is adjusted simultaneously, and the individual dangerous actions are accurately captured and warning information is output in the state of dense gathering.
[0008] Preferably, the separation driving identifier is generated based on pixel displacement amplitude, contour edge transition frequency, and local occlusion overlap ratio. The specific steps are as follows: The adjacent frames of the construction area are collected and numbered in time sequence. The search area is set based on the pixel position to complete the brightness comparison. The coordinate difference of the corresponding point is extracted to obtain the pixel displacement amplitude. At the same time, the brightness difference is scanned to extract the edge points and the number of position changes in the continuous frame is recorded to obtain the contour edge jump frequency. By combining pixel displacement amplitude and contour edge jump frequency, the level intervals are divided to generate motion change markers and boundary change markers. Pixel statistics are performed on the overlapping areas of personnel contours to obtain the local occlusion overlap ratio and divide the coverage level to generate spatial coverage markers. The motion change markers, boundary change markers, and spatial coverage markers are summarized, recorded, and grouped into the same category area to complete the generation of separation-driven identifiers. At the same time, the range of changes in the combined values of the analysis unit is defined, and the degree of boundary adhesion is marked.
[0009] Preferably, in the process of completing the brightness comparison by setting the search area based on the pixel position, a limited search area is set around the corresponding position in the next frame for the pixel coordinates in the previous frame. The position with the smallest difference is selected as the corresponding point by comparing the brightness value difference point by point. The pixel displacement amplitude is obtained by the difference between the coordinates of the previous and next frames. At the same time, the number of edge point position changes is recorded in the continuous frame to obtain the contour edge jump frequency.
[0010] Preferably, independent motion segments are extracted from the divided regions, and the specific steps are as follows: Obtain the spatial distribution of the separation driving identifier, delineate the set of pixels with consistent combined markers and spatially adjacent pixels as the base region, and filter the region containing spatial covering markers and whose corresponding pixels are continuously distributed as the continuous adhesion region. The difference between adjacent pixel combination numbers is compared point by point along the continuous adhesion region. The difference is recorded as the gradient change and the spatial direction is marked. The gradient change and direction are summarized to obtain the gradient change distribution. The intersection of gradient change paths is marked as the potential segmentation position. The initial segmentation boundary is delineated and the directional difference constraint is introduced to correct the boundary. Based on the corrected region division results, pixels with displacement changes in continuous frames are collected as candidate sets according to the pixel spatial adjacency relationship. The candidate sets are determined to be independent motion segments, and boundary expansion is performed to restore individual contour information, thus completing the extraction of independent motion segments from continuous adhered regions.
[0011] Preferably, the difference between adjacent pixel combination numbers is compared point by point along the continuous adhesion region, the difference is recorded as the gradient change amount and the spatial direction is marked, the gradient change amount and direction are summarized to obtain the gradient change distribution, the intersection of gradient change paths is marked as potential segmentation position, and directional difference constraints are introduced to correct the boundary to realize the division of the continuous adhesion region.
[0012] Preferably, non-consistent action blocks are highlighted during the rearrangement process, and the specific steps are as follows: The pixel distribution of independent motion segments is analyzed frame by frame. The average coordinate position is calculated by accumulating the horizontal and vertical coordinates of the pixels to determine the center position. Boundary points are obtained by extending along the vertical direction to construct the attitude reference line. The attitude offset angle is determined by combining the angle changes in consecutive frames. At the same time, the maximum offset distance is calculated based on the change in the center position to determine the amplitude of the center of gravity swing. Spatial mapping is performed based on the attitude offset angle and the center of gravity swing amplitude to establish motion feature identifiers. The differences between adjacent independent motion segments are analyzed by comparing the combined identifiers of angle intervals and amplitude intervals, the number of differences in the boundary area is recorded, and the areas that reach the preset number standard are selected as non-consistent motion candidate areas. Local rearrangement is performed on candidate regions of inconsistent actions. The pixels are sorted according to the range of changes in the pose offset angle and the amplitude of the center of gravity swing. The pixel affiliation is adjusted to expand the spatial range of the difference fragments and compress the spatial range of other fragments. The inconsistent action blocks are identified by comparing the combined identifiers before and after the rearrangement, thereby revealing the abnormal actions.
[0013] Preferably, the process of obtaining candidate regions and candidate blocks for inconsistent actions is as follows: the number of differences in the boundary regions is accumulated and regions that reach the preset number standard are selected as candidate regions for inconsistent actions, and candidate blocks for inconsistent actions are obtained through spatial continuous distribution merging.
[0014] Preferably, the construction of independently identifiable dangerous behavior units involves the following steps: Extract the spatial coordinates corresponding to the pixel distribution range of the non-consistent action block, filter the pixels that have no adjacent pixels of the same type in four directions as the boundary pixel set, connect the boundary pixel set to obtain the closed boundary curve and determine the initial boundary range. Based on the distribution of the attitude offset angle and center of gravity swing amplitude range within the initial defined range, the range combination with the most pixels is selected to determine the dominant range. Based on the dominant range, a dynamic shrinkage definition is performed, advancing inward pixel by pixel, eliminating non-dominant range pixels until a dominant range pixel appears and the advancement stops. By combining the remaining pixel count ratio within the shrinkage boundary with a spatial compression ratio constraint to limit the shrinkage degree, the final boundary pixel set is removed one by one from the continuous adhesion area and assigned an independent label to construct an independently identifiable dangerous behavior unit.
[0015] Preferably, the stripped motion features are backfilled into the corresponding positions, and the warning trigger threshold range is adjusted simultaneously. The specific steps are as follows: Read the original spatial coordinate information corresponding to the dangerous behavior unit, find the same coordinate pixels in the original dense area one by one, replace the corresponding position pixel identifier with the dangerous behavior unit pixel, and write the attitude offset angle and center of gravity swing amplitude to complete the reverse mapping process. Around the pixel set after the reverse mapping process is completed, organize the motion feature information of all pixels in the core area, expand a layer of pixels outward along the four neighboring directions, and assign the attitude offset angle and the center of gravity swing amplitude to the expanded pixels to realize motion feature backfilling. For the original dense area after the action feature backfilling is completed, the warning trigger threshold range is adjusted. The trigger thresholds for posture offset angle and center of gravity swing amplitude in the area containing dangerous behavior units are reduced, while the thresholds for areas not containing dangerous behavior units remain unchanged, so as to achieve accurate capture of individual dangerous actions and output of warning information.
[0016] The AI-based real-time early warning system for safety risks at construction sites includes a multi-source feature extraction module, an adhesion area segmentation module, an action difference recognition module, an abnormal action stripping module, and a feature backfilling early warning module. The multi-source feature extraction module obtains the pixel displacement amplitude, contour edge jump frequency, and local occlusion overlap ratio in the densely populated images of the construction area, and generates separation driving identifiers based on the pixel displacement amplitude, contour edge jump frequency, and local occlusion overlap ratio. The adhesion region segmentation module performs gradient segmentation on continuous adhesion regions within the image based on the separation driving identifier, introduces directional difference constraints at each adhesion boundary position, and extracts independent motion segments from the segmented regions. The motion difference recognition module calculates the posture offset angle and center of gravity swing amplitude corresponding to each independent motion segment, and performs local rearrangement processing based on the posture offset angle and center of gravity swing amplitude, highlighting non-consistent motion blocks during the rearrangement process; The abnormal action stripping module performs dynamic contraction and delimitation at the corresponding position based on the non-consistent action block, and superimposes spatial compression ratio constraints during the contraction process to strip local abnormal actions from the continuous adhesion area and construct independently identifiable dangerous behavior units. The feature backfilling and early warning module performs reverse mapping processing on the original dense area based on the dangerous behavior unit, backfills the stripped action features to the corresponding position, and simultaneously adjusts the early warning trigger threshold range. It completes the accurate capture of individual dangerous actions and outputs early warning information in the state of dense aggregation.
[0017] The technical effects and advantages provided by this application in the above technical solution are as follows: This application constructs a separation-driven identifier and guides gradient segmentation and directional difference constraint processing to achieve effective separation of continuous and adhered regions in a densely gathered state of multiple people. This allows individual action information that was originally covered by the overall motion characteristics to be extracted from the adhered regions, thereby avoiding the masking effect of group behavior on local actions. It can still maintain the ability to continuously perceive the motion state of individuals under complex occlusion conditions, improve the completeness of abnormal action identification, and enable the construction site to still have a stable risk assessment basis in high-density operation scenarios.
[0018] This application uses a combination of dynamic shrinkage definition and spatial compression ratio constraints to separate abnormal action areas from the contiguous environment and construct them into independent dangerous behavior units. At the same time, through reverse mapping and action feature backfilling mechanisms, the dangerous behavior units can be accurately located in the original dense area. With the synchronous adjustment of the early warning trigger threshold range, local abnormal actions can be captured and early warnings can be triggered at an earlier stage, thereby improving the timeliness and accuracy of risk early warning response at the construction site. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings.
[0020] Figure 1 This is a flowchart illustrating the overall method of this application.
[0021] Figure 2 This is a schematic diagram of the system modules of this application. Detailed Implementation
[0022] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.
[0023] like Figure 1 As shown, this application provides a real-time early warning method for construction site safety risks based on artificial intelligence, including the following steps: Step 1: Obtain the pixel displacement amplitude, contour edge jump frequency, and local occlusion overlap ratio in the densely populated construction area. Generate a separation driving identifier based on the pixel displacement amplitude, contour edge jump frequency, and local occlusion overlap ratio to characterize the degree of boundary adhesion in the state of personnel gathering. The separation driving identifier is generated based on pixel displacement amplitude, contour edge transition frequency, and local occlusion overlap ratio. The specific steps are as follows: Image acquisition devices with fixed viewing angles are deployed within the construction area to cover the area where people are active. During continuous acquisition, adjacent frames are acquired at fixed time intervals, and each frame is numbered sequentially by time. Between two adjacent frames, for each pixel position, a search area is set around the coordinate position in the next frame, using the pixel's coordinates in the previous frame as a reference. Within the search area, the brightness value changes are compared point by point, and the pixel position with the smallest difference in brightness change is selected as the corresponding point. The displacement distance of the pixel is determined by the difference in coordinates between the two frames, and this displacement distance is recorded as the pixel displacement amplitude at that position.
[0024] After calculating the pixel displacement of the entire image, all pixel displacement amplitudes are stored according to spatial coordinates to form a displacement distribution covering the entire image. Subsequently, contour extraction is performed on the personnel area in the same image. By scanning the brightness difference between adjacent pixels, when the brightness difference between adjacent pixels exceeds a preset range, the position is marked as an edge point. The spatial position change of the edge point is recorded in multiple consecutive frames. When the same edge point moves or disappears and reappears between consecutive frames, the change is counted as a jump. The number of jumps of the edge point is accumulated within a set time interval, and the position offset distance corresponding to each jump is recorded. Thus, the contour edge jump frequency of the position is obtained.
[0025] Based on the above, the area of people in the image is divided into regions. By detecting the overlapping areas between different people's contours, the areas where two or more contour boundaries overlap are marked as occlusion areas. By statistically analyzing the proportion of the number of pixels contained in the occlusion area to the total number of pixels in the entire image, the local occlusion overlap ratio is obtained. This ratio value is marked on the corresponding spatial position, so that the pixel displacement amplitude, contour edge jump frequency and local occlusion overlap ratio are established in the same spatial coordinate system.
[0026] After obtaining the pixel displacement amplitude, contour edge transition frequency, and local occlusion overlap ratio, for each spatial location, the corresponding pixel displacement amplitude value is divided into multiple level intervals according to a preset interval, and the level interval number is used as the motion change mark for that location; at the same time, the contour edge transition frequency corresponding to that location is divided into multiple frequency intervals according to the cumulative number within the time interval, and the frequency interval number is used as the boundary change mark for that location.
[0027] Furthermore, the proportion of local occlusion overlap is divided into several coverage levels according to the proportion value, and the coverage level number is used as the spatial coverage mark of the location; at the same spatial coordinate position, the motion change mark, boundary change mark and spatial coverage mark are combined and recorded so that each pixel position corresponds to a combined value containing the three types of marks.
[0028] Within the entire image, the combination values of all pixel positions are uniformly organized, and pixels with the same combination value are grouped into the same category area and formed a continuous distribution in space, thereby obtaining the separation drive identifier. This separation drive identifier presents a distribution form in space that corresponds to different combination states in different areas. The different combination states reflect the comprehensive characteristics of motion changes, boundary changes, and occlusion coverage in the area.
[0029] After the separation driving identifier is constructed, the image is divided into regions based on the spatial distribution of the separation driving identifier. Continuous pixel regions with the same or similar combination values are defined as the same analysis unit. The separation driving identifiers in each analysis unit are statistically analyzed, and the range of change of combination values and spatial distribution within the analysis unit are recorded.
[0030] During the analysis, the pixel displacement amplitude within the analysis unit is compared with the contour edge transition frequency. When multiple different combinations of values exist within the same analysis unit and these combinations are spatially staggered, the analysis unit is marked as a boundary adhesion region. At the same time, combined with the local occlusion overlap ratio, when there is an occlusion region within the analysis unit and the occlusion region is distributed throughout the entire analysis unit, the analysis unit is further marked as a high adhesion region.
[0031] After marking all analysis units, the marking results of each analysis unit are integrated according to their spatial position in the image to form a complete distribution result of the degree of boundary adhesion. This result can reflect the boundary adhesion situation under the state of personnel gathering in the construction area on a regional basis, and provide a clear spatial positioning basis for subsequent processing of continuous adhesion areas.
[0032] Step 2: Based on the separation-driven identifier, perform gradient segmentation on the continuous and connected regions within the image, introduce directional difference constraints at each connection boundary, extract independent motion segments from the segmented regions, and restore the compressed individual contour information. The specific steps for extracting independent motion segments from the divided regions are as follows: Based on the spatial distribution of the separation drive markers in the image, continuous adhesion regions are identified and defined. In the process of defining the regions, the set of pixels with completely identical combined markers and spatially adjacent pixels in the separation drive markers is defined as the base region. Then, the regions containing spatial overlay markers and whose corresponding pixel numbers are continuously distributed are further selected from the base regions as continuous adhesion regions.
[0033] After the continuous adhesion region is determined, take any pixel position in the region as the starting point, and compare the difference of the separation driving identifier combination number of adjacent pixels point by point in its four neighborhoods. When there is a numerical difference between the combination numbers of adjacent pixels, the difference is recorded as the gradient change amount at that position, and the gradient change amount is marked according to the spatial direction to form the gradient change direction from the current pixel to the adjacent pixel.
[0034] Subsequently, all pixels within the entire continuous and adhered region are traversed in the same manner, and the gradient changes between all pixels and their corresponding directions are summarized to obtain the gradient change distribution covering the entire continuous and adhered region. In this distribution, when the gradient change directions between multiple adjacent pixels show a continuous extension relationship, the extension path is recorded as a gradient change path, and the locations where the gradient change paths intersect or separate are marked as potential segmentation locations. Based on the connectivity of these potential segmentation locations in space, the preliminary segmentation boundary is delineated, so that the continuous and adhered region is divided into several sub-regions with relatively consistent internal gradient changes.
[0035] After obtaining the initial sub-region division results, directional difference constraints are introduced around the boundary positions between each sub-region. During the processing, the direction of the spatial displacement corresponding to the pixel displacement amplitude within each sub-region is extracted. Specifically, by recording the coordinate changes of each pixel between adjacent frames, the coordinate difference of the pixel in the previous frame and the next frame is converted into a spatial displacement vector, and the displacement direction of the pixel is determined based on the displacement vector.
[0036] After extracting the displacement direction of all pixels in the sub-region, these displacement directions are classified and statistically analyzed according to their angle range. The number of pixels falling into the same angle range is accumulated, and the angle range with the most pixels is selected as the dominant motion direction of the sub-region. Subsequently, at the boundary of adjacent sub-regions, the dominant motion directions of the two sub-regions are differentially processed, and the angle difference between the two dominant motion directions is calculated.
[0037] When the angle difference exceeds the preset directional change range, the boundary position is marked as the directional difference boundary, and the initial segmentation boundary is corrected along the boundary, so that the segmentation boundary moves towards the area with greater directional difference, thereby making the corrected boundary more in line with the actual motion distribution. After the above processing is completed on the boundary positions of all sub-regions, the region division result after directional difference constraint correction is obtained, so that adjacent sub-regions have a distinguishable relationship in the direction of motion.
[0038] Based on the region division after completing the orientation difference constraint correction, independent motion segments are extracted from the divided regions. During the extraction process, each sub-region is used as the basic processing unit. Within the sub-region, based on the spatial adjacency relationship between pixels, pixels that are adjacent to each other and maintain displacement changes in consecutive frames are grouped into the same candidate set, and the continuity of the candidate set is judged in the time dimension.
[0039] When a set exists in multiple consecutive frames and the positions of its internal pixels change continuously over time, the candidate set is identified as an independent motion segment. After the extraction of the independent motion segments, contour restoration is performed on the boundary position of each independent motion segment. During the restoration process, starting from the pixel at the boundary of the independent motion segment, the process gradually expands along its outer adjacent pixels. When the expansion encounters a position where the separation drive identifier combination number changes, this position is identified as the boundary expansion termination position, and all pixels on the expansion path are included in the current independent motion segment range, so that the segment boundary extends outward from its original position to the area where the separation drive identifier changes, thereby restoring the individual contour information that was compressed in the continuous adhesion state.
[0040] After extracting and restoring the boundaries of all independent motion segments, each independent motion segment is mapped onto the original image according to its spatial location, so that each independent motion segment corresponds to a person area with independent motion characteristics, providing a clear input basis for subsequent posture analysis and dangerous behavior identification.
[0041] Step 3: Calculate the posture offset angle and center of gravity swing amplitude corresponding to each independent motion segment, and perform local rearrangement based on the posture offset angle and center of gravity swing amplitude. During the rearrangement process, highlight non-consistent motion blocks and reveal abnormal motions that are covered by the group. The steps to highlight inconsistent action blocks during the rearrangement process are as follows: The pixel distribution of each independent motion segment in consecutive frames is analyzed frame by frame. In each frame, the spatial coordinates of all pixels within the independent motion segment are counted. The horizontal and vertical coordinates of all pixels are summed and the average coordinate position is calculated based on the total number of pixels. This average coordinate position is determined as the center position of the independent motion segment in the current frame.
[0042] After obtaining the center position, the vertical direction of the screen is used as a unified reference direction. The center position is extended upward and downward along the vertical direction, and the two boundary points that intersect with the boundary of the independent motion segment are recorded. The line connecting the center position to the upper boundary point is used as the attitude reference line of the current independent motion segment. The angle change value between the attitude reference line and the vertical direction of the screen is recorded between consecutive frames. The range of the angle change value in consecutive frames is determined as the attitude offset angle corresponding to the independent motion segment.
[0043] Based on the calculation of the attitude offset angle, the change of the center position of the independent motion segment in consecutive frames is recorded. By accumulating the coordinate difference of the center position in adjacent frames frame by frame, and calculating the maximum offset distance of the center position relative to the initial position over the entire time interval, the maximum offset distance is determined as the center of gravity swing amplitude of the independent motion segment, so that both the attitude offset angle and the center of gravity swing amplitude have a clear calculation basis.
[0044] After obtaining the attitude offset angle and center of gravity swing amplitude corresponding to each independent motion segment, based on the spatial distribution of all independent motion segments within the same frame, the attitude offset angle and center of gravity swing amplitude of each independent motion segment are spatially mapped according to its center position. A corresponding motion feature identifier is established for each independent motion segment in the frame coordinate system. Multiple angle intervals are divided according to the numerical range of the attitude offset angle, and multiple amplitude intervals are divided according to the numerical range of the center of gravity swing amplitude, so that each independent motion segment corresponds to a combination identifier of angle interval number and amplitude interval number.
[0045] After assigning combined identifiers to all independent motion segments, the combined identifiers of spatially adjacent independent motion segments are compared one by one. When there is a difference between the angle interval number or amplitude interval number of two adjacent independent motion segments, the difference is recorded in the boundary area, and the number of differences is accumulated. When the number of accumulated differences in the same area reaches the preset number standard, the area is determined as a non-consistent action candidate area, and these candidate areas are merged in space to form a complete non-consistent action candidate block by making adjacent and continuously distributed candidate areas.
[0046] After the candidate block of non-consistent motion is determined, local rearrangement is performed on the independent motion segments involved in the block. During the process, the candidate block of non-consistent motion is used as the processing range. All independent motion segments within the range are sorted according to the combination of their attitude offset angle and center of gravity swing amplitude. Independent motion segments with larger attitude offset angle variation range are arranged in the front area of the spatial position, and independent motion segments with larger center of gravity swing amplitude are arranged in the center area of the spatial position. At the same time, independent motion segments with attitude offset angle and center of gravity swing amplitude in the same range are arranged in the direction away from the center area.
[0047] In practical implementation, independent motion segments with a large range of attitude deviation angle changes refer to: Within the same time interval, the attitude offset angle change range of each independent motion segment is calculated, and the attitude offset angle change ranges of each independent motion segment are uniformly sorted. The independent motion segments that fall within a pre-defined proportion range in the sorting results are identified as independent motion segments with larger attitude offset angle change ranges. The pre-defined proportion range is the top 30% of independent motion segments after sorting by attitude offset angle change range from largest to smallest, or independent motion segments whose attitude offset angle change range exceeds the average value of all independent motion segments.
[0048] Independent motion segments with large amplitude of center of gravity oscillation refer to: Within the same time interval, the center of gravity swing amplitude is calculated for each independent motion segment, and the center of gravity swing amplitudes corresponding to each independent motion segment are uniformly sorted. The independent motion segments that fall within a pre-defined proportion interval in the sorting results are identified as independent motion segments with larger center of gravity swing amplitudes. The pre-defined proportion interval is the top 30% of independent motion segments after sorting by center of gravity swing amplitude from largest to smallest, or independent motion segments whose center of gravity swing amplitude exceeds the average value of all independent motion segments. Thus, the larger value is clearly defined by the sorting position or statistical distribution, avoiding uncertain interpretations.
[0049] After sorting is completed, the affiliation of all pixels in the region is redistributed, so that pixels that originally belonged to different independent motion segments are reassigned to new spatial positions according to the sorting results. By expanding the pixels corresponding to the independent motion segments that are ranked higher outward and compressing the pixel spatial range corresponding to the independent motion segments that are ranked lower, the independent motion segments with large differences obtain a larger display area in spatial distribution, thereby enhancing their visual differentiation from surrounding segments.
[0050] After the local rearrangement is completed, the non-consistent motion blocks are finally determined based on the spatial distribution after the rearrangement. During the determination process, the combination identifiers of independent motion segments at the same spatial position before and after the rearrangement are compared. When a certain spatial region corresponds to multiple independent motion segments before the rearrangement and to a single independent motion segment after the rearrangement, the region is marked as a motion separation region. At the same time, the corresponding posture offset angle and center of gravity swing amplitude in the region are used as motion feature descriptions of the region.
[0051] Furthermore, in all action separation areas, regions with different posture offset angles from the surrounding areas and whose center of gravity swing amplitude changes beyond a preset range are selected. These regions are identified as non-consistent action blocks, and their spatial positions in the original image are marked so that they can be identified individually in subsequent processing. This enables the exposure of abnormal actions that are obscured by the crowd, providing a clear and distinguishable input basis for subsequent dangerous behavior identification and early warning information output.
[0052] Step 4: Based on the non-consistent action blocks, perform dynamic contraction and delimitation at the corresponding positions, and superimpose spatial compression ratio constraints during the contraction process to separate local abnormal actions from the continuous adhesion area and construct independently identifiable dangerous behavior units. The specific steps for constructing independently identifiable dangerous behavior units are as follows: Based on the pixel distribution range of the non-uniform action block in the original image, boundary extraction processing is performed. During the extraction process, the spatial coordinates of all pixels in the non-uniform action block are counted, and pixels that do not have adjacent pixels of the same type in any of the four directions are selected. This set of pixels is determined as the boundary pixel set, and they are connected in order of spatial position to form a closed boundary curve. The area enclosed by the closed boundary curve is determined as the initial defined range.
[0053] Within this initial defined range, the attitude offset angle and center of gravity swing amplitude corresponding to each pixel are statistically analyzed, and the ranges are classified according to the intervals of attitude offset angle and center of gravity swing amplitude. The interval combination with the most pixels is determined as the dominant interval. The dominant interval is obtained by comparing the number of pixels corresponding to each interval combination. The interval combination with the largest number of pixels is the unique dominant interval, thus providing a clear basis for subsequent dynamic shrinkage definition.
[0054] After the dominant region is obtained, dynamic shrinkage boundary processing is performed around the initial boundary range. During the processing, the boundary pixel set is used as the starting position, and the process moves inward pixel by pixel along the boundary normal direction in space. Each move removes the outermost pixel from the current boundary pixel set and updates the remaining pixel set within the boundary range.
[0055] After each removal operation, the attitude offset angle and center of gravity swing amplitude range corresponding to the removed pixels are statistically analyzed. When all removed pixels fall into the non-dominant range, the next inward push operation is continued. When a pixel belonging to the dominant range appears for the first time among the removed pixels, the push along that direction is stopped immediately, thereby forming independent contraction termination boundaries in different directions, so that the final defined range only retains the pixel set region with the dominant range as the core.
[0056] During the dynamic shrinkage and demarcation process, a spatial compression ratio constraint is introduced to control the degree of shrinkage of the demarcation range. In specific implementation, the total number of pixels within the initial demarcation range is used as the baseline number. After each boundary advancement, the number of remaining pixels within the current demarcation range is counted, and the ratio between the current number of remaining pixels and the baseline number is calculated. When this ratio drops to a preset ratio threshold range, the shrinkage operation in all directions is stopped. The preset ratio threshold range is limited to 30% to 60% of the number of pixels in the initial demarcation range. When the current percentage of remaining pixels is less than 30%, the shrinkage is stopped to avoid over-compression of the effective action area. When the current percentage of remaining pixels is greater than 60%, the shrinkage operation continues to eliminate peripheral interference areas. This ensures that the final demarcation range can both eliminate interference from non-dominant areas and retain the complete abnormal action feature area in space.
[0057] After completing the dynamic shrinkage definition and spatial compression ratio constraint processing, local abnormal action stripping processing is performed around the final defined range. During the processing, the pixel set within the final defined range is removed one by one from the pixel set of the original continuous and adhered region, and an independent identifier is assigned to the removed pixel set. This makes the pixel set no longer belong to the original continuous and adhered region in the data record, but exists as a new independent pixel set. At the same time, the spatial position of the pixel set in the original image coordinates remains unchanged, so that it is still in its original position in space but independent in terms of belonging relationship. This constructs a dangerous behavior unit corresponding to the pixel set, and retains the corresponding posture offset angle and center of gravity swing amplitude as behavioral feature description in the dangerous behavior unit. This allows the dangerous behavior unit to participate as an independent object in subsequent dangerous behavior identification and early warning processing, thereby achieving the purpose of stripping local abnormal actions from the continuous and adhered region and constructing independently identifiable dangerous behavior units.
[0058] Step 5: Perform reverse mapping processing on the original dense area based on the dangerous behavior unit, backfill the stripped action features to the corresponding position, and simultaneously adjust the warning trigger threshold range to accurately capture individual dangerous actions and output warning information in the dense cluster state; The stripped motion features are then backfilled into the corresponding positions, and the warning trigger threshold range is adjusted synchronously. The specific steps are as follows: After the dangerous behavior unit is constructed, based on the original spatial coordinate information recorded before the dangerous behavior unit is stripped, a reverse mapping process is performed on its corresponding pixel set. During the process, the original coordinate position of each pixel in the dangerous behavior unit is read one by one, and a pixel point with the same coordinate position is found in the pixel set of the original dense area. The pixel in the dangerous behavior unit replaces the pixel identifier of the corresponding position in the original dense area one by one, so that the pixel corresponding to the position changes from the original continuous and adhered area to the dangerous behavior unit. At the same time, during the replacement process, the posture offset angle and center of gravity swing amplitude of the dangerous behavior unit are written as action feature information into the corresponding pixel position, so that each pixel that has completed the mapping has both spatial position and action feature information. Thus, the accurate spatial positioning of the dangerous behavior unit is restored in the original dense area, realizing the complete execution of the reverse mapping process.
[0059] Subsequently, around the dangerous behavior units that have been mapped to the original dense area, motion feature backfilling processing is performed on their corresponding positions. During the processing, the pixel set corresponding to the dangerous behavior unit is taken as the core area. First, the pose offset angle and center of gravity swing amplitude of all pixels in the core area are uniformly sorted out, and the motion feature information is assigned to all pixels in the area.
[0060] Subsequently, starting from the boundary pixel of the core region, a layer of pixels is extended outward along the four neighboring regions. Pixels adjacent to the boundary pixel that do not yet contain motion feature information are included in the extension range, and the attitude offset angle and center of gravity swing amplitude of the dangerous behavior unit are assigned to the extended pixel.
[0061] After completing one layer of expansion, the expansion is stopped, thus strictly limiting the action feature backfilling range to the set of dangerous behavior unit pixels and one layer of adjacent pixels. This avoids excessive spatial diffusion of action features while maintaining the spatial structure of the original dense area. Action feature identifiers are added only at the pixel attribute level, thereby enabling the stripped action features to be backfilled to the corresponding positions.
[0062] After the motion feature backfilling is completed, the warning trigger threshold range is synchronously adjusted based on the spatial distribution of dangerous behavior units in the original dense area. During the adjustment process, the warning trigger threshold range is used as the benchmark, and the warning trigger threshold in the area containing dangerous behavior units is uniformly adjusted to 70% of the original threshold. That is, the attitude offset angle trigger threshold and the center of gravity swing amplitude trigger threshold are reduced to 70% of the original threshold, making it easier for motion changes in the area to trigger warning judgment. The original warning trigger threshold range remains unchanged in the area that does not contain dangerous behavior units, thus forming differentiated warning trigger standards in the same image.
[0063] After the threshold adjustment is completed, when the attitude offset angle of any pixel exceeds the attitude offset angle trigger threshold corresponding to its region, or when the center of gravity swing amplitude of the pixel exceeds the center of gravity swing amplitude trigger threshold corresponding to its region, it is determined that dangerous behavior has occurred at the pixel position, and an early warning message is immediately output. This enables the accurate capture of individual dangerous actions and the output of early warning information in a densely clustered state.
[0064] like Figure 2 As shown, this application provides an artificial intelligence-based real-time early warning system for construction site safety risks, including a multi-source feature extraction module, an adhesion area segmentation module, an action difference recognition module, an abnormal action stripping module, and a feature backfilling early warning module; The multi-source feature extraction module obtains the pixel displacement amplitude, contour edge jump frequency, and local occlusion overlap ratio in the densely populated images of the construction area, and generates separation driving identifiers based on the pixel displacement amplitude, contour edge jump frequency, and local occlusion overlap ratio. The adhesion region segmentation module performs gradient segmentation on continuous adhesion regions within the image based on the separation driving identifier, introduces directional difference constraints at each adhesion boundary position, and extracts independent motion segments from the segmented regions. The motion difference recognition module calculates the posture offset angle and center of gravity swing amplitude corresponding to each independent motion segment, and performs local rearrangement processing based on the posture offset angle and center of gravity swing amplitude, highlighting non-consistent motion blocks during the rearrangement process; The abnormal action stripping module performs dynamic contraction and delimitation at the corresponding position based on the non-consistent action block, and superimposes spatial compression ratio constraints during the contraction process to strip local abnormal actions from the continuous adhesion area and construct independently identifiable dangerous behavior units. The feature backfilling and early warning module performs reverse mapping processing on the original dense area based on the dangerous behavior unit, backfills the stripped action features to the corresponding position, and simultaneously adjusts the early warning trigger threshold range. It completes the accurate capture of individual dangerous actions and outputs early warning information in the state of dense aggregation.
[0065] The artificial intelligence-based real-time early warning method for construction site safety risks provided in this embodiment of the invention is implemented through the aforementioned artificial intelligence-based real-time early warning system for construction site safety risks. For details of the specific methods and processes of the artificial intelligence-based real-time early warning system for construction site safety risks, please refer to the embodiments of the aforementioned artificial intelligence-based real-time early warning method for construction site safety risks, which will not be repeated here.
[0066] The foregoing has only described certain exemplary embodiments of this application by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of this application. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of this application.
Claims
1. A real-time early warning method for safety risks at construction sites based on artificial intelligence, characterized in that, Includes the following steps: The pixel displacement amplitude, contour edge jump frequency, and local occlusion overlap ratio in the densely populated construction area are obtained, and a separation driving identifier is generated based on the pixel displacement amplitude, contour edge jump frequency, and local occlusion overlap ratio. Based on the separation-driven identifier, gradient segmentation is performed on the continuous and connected regions within the image. Orientation difference constraints are introduced at each connection boundary to extract independent motion segments from the segmented regions. Calculate the posture offset angle and center of gravity swing amplitude corresponding to each independent motion segment, and perform local rearrangement based on the posture offset angle and center of gravity swing amplitude, highlighting non-consistent motion blocks during the rearrangement process; Based on the non-uniform action blocks, dynamic contraction is performed at the corresponding positions, and spatial compression ratio constraints are superimposed during the contraction process to separate local abnormal actions from the continuous adhesion area and construct independently identifiable dangerous behavior units. Based on the dangerous behavior units, the original dense area is reverse-mapped, and the stripped action features are backfilled into the corresponding positions. The warning trigger threshold range is adjusted simultaneously, and the individual dangerous actions are accurately captured and warning information is output in the state of dense gathering.
2. The method for real-time early warning of construction site safety risks based on artificial intelligence according to claim 1, characterized in that, The separation driving identifier is generated based on pixel displacement amplitude, contour edge transition frequency, and local occlusion overlap ratio. The specific steps are as follows: The adjacent frames of the construction area are collected and numbered in time sequence. The search area is set based on the pixel position to complete the brightness comparison. The coordinate difference of the corresponding point is extracted to obtain the pixel displacement amplitude. At the same time, the brightness difference is scanned to extract the edge points and the number of position changes in the continuous frame is recorded to obtain the contour edge jump frequency. By combining pixel displacement amplitude and contour edge jump frequency, the level intervals are divided to generate motion change markers and boundary change markers. Pixel statistics are performed on the overlapping areas of personnel contours to obtain the local occlusion overlap ratio and divide the coverage level to generate spatial coverage markers. The motion change markers, boundary change markers, and spatial coverage markers are summarized, recorded, and grouped into the same category area to complete the generation of separation-driven identifiers.
3. The method for real-time early warning of construction site safety risks based on artificial intelligence according to claim 2, characterized in that, In the process of brightness comparison based on pixel position setting search area, a limited search area is set around the corresponding position in the next frame for the pixel coordinates in the previous frame. The position with the smallest difference is selected as the corresponding point by comparing the brightness value difference point by point. The pixel displacement amplitude is obtained by the difference between the coordinates of the previous and next frames. At the same time, the number of edge point position changes in the continuous frame is recorded to obtain the contour edge jump frequency.
4. The method for real-time early warning of construction site safety risks based on artificial intelligence according to claim 2, characterized in that, The specific steps for extracting independent motion segments from the divided regions are as follows: Obtain the spatial distribution of the separation driving identifier, delineate the set of pixels with consistent combined markers and spatially adjacent pixels as the base region, and filter the region containing spatial covering markers and whose corresponding pixels are continuously distributed as the continuous adhesion region. The difference between adjacent pixel combination numbers is compared point by point along the continuous adhesion region. The difference is recorded as the gradient change and the spatial direction is marked. The gradient change and direction are summarized to obtain the gradient change distribution. The intersection of gradient change paths is marked as the potential segmentation position. The initial segmentation boundary is delineated and the directional difference constraint is introduced to correct the boundary. Based on the corrected region division results, pixels with displacement changes in continuous frames are collected as candidate sets according to the pixel spatial adjacency relationship. The candidate sets are determined to be independent motion segments, and boundary expansion is performed to restore individual contour information, thus completing the extraction of independent motion segments from continuous adhered regions.
5. The method for real-time early warning of construction site safety risks based on artificial intelligence according to claim 4, characterized in that, The difference between adjacent pixel combination numbers is compared point by point along the continuous and adhered region. The difference is recorded as the gradient change and the spatial direction is marked. The gradient change and direction are summarized to obtain the gradient change distribution. The intersection of gradient change paths is marked as the potential segmentation position. The directional difference constraint is introduced to correct the boundary and realize the division of the continuous and adhered region.
6. The method for real-time early warning of construction site safety risks based on artificial intelligence according to claim 4, characterized in that, The steps to highlight inconsistent action blocks during the rearrangement process are as follows: The pixel distribution of independent motion segments is analyzed frame by frame. The average coordinate position is calculated by accumulating the horizontal and vertical coordinates of the pixels to determine the center position. Boundary points are obtained by extending along the vertical direction to construct the attitude reference line. The attitude offset angle is determined by combining the angle changes in consecutive frames. At the same time, the maximum offset distance is calculated based on the change in the center position to determine the amplitude of the center of gravity swing. Spatial mapping is performed based on the attitude offset angle and the center of gravity swing amplitude to establish motion feature identifiers. The differences between adjacent independent motion segments are analyzed by comparing the combined identifiers of angle intervals and amplitude intervals, the number of differences in the boundary area is recorded, and the areas that reach the preset number standard are selected as non-consistent motion candidate areas. For candidate regions of inconsistent actions, local rearrangement is performed. By sorting according to the range of changes in attitude offset angle and the amplitude of center of gravity swing, the pixel affiliation is adjusted to expand the spatial range of the difference segment and compress the spatial range of other segments. The inconsistent action blocks are determined by comparing the combined identifiers before and after rearrangement.
7. The method for real-time early warning of construction site safety risks based on artificial intelligence according to claim 6, characterized in that, The process of obtaining candidate regions and candidate blocks for inconsistent actions is as follows: the number of differences in the boundary regions is accumulated and regions that reach the preset number standard are selected as candidate regions for inconsistent actions. Candidate blocks for inconsistent actions are obtained through spatial continuous distribution merging.
8. The method for real-time early warning of construction site safety risks based on artificial intelligence according to claim 6, characterized in that, The specific steps for constructing independently identifiable dangerous behavior units are as follows: Extract the spatial coordinates corresponding to the pixel distribution range of the non-consistent action block, filter the pixels that have no adjacent pixels of the same type in four directions as the boundary pixel set, connect the boundary pixel set to obtain the closed boundary curve and determine the initial boundary range. Based on the distribution of the attitude offset angle and center of gravity swing amplitude range within the initial defined range, the range combination with the most pixels is selected to determine the dominant range. Based on the dominant range, a dynamic shrinkage definition is performed, advancing inward pixel by pixel, eliminating non-dominant range pixels until a dominant range pixel appears and the advancement stops. By combining the remaining pixel count ratio within the shrinkage boundary with a spatial compression ratio constraint to limit the shrinkage degree, the final boundary pixel set is removed one by one from the continuous adhesion area and assigned an independent label to construct an independently identifiable dangerous behavior unit.
9. The method for real-time early warning of construction site safety risks based on artificial intelligence according to claim 8, characterized in that, The stripped motion features are then backfilled into the corresponding positions, and the warning trigger threshold range is adjusted synchronously. The specific steps are as follows: Read the original spatial coordinate information corresponding to the dangerous behavior unit, find the same coordinate pixels in the original dense area one by one, replace the corresponding position pixel identifier with the dangerous behavior unit pixel, and write the attitude offset angle and center of gravity swing amplitude to complete the reverse mapping process. Around the pixel set after the reverse mapping process is completed, organize the motion feature information of all pixels in the core area, expand a layer of pixels outward along the four neighboring directions, and assign the attitude offset angle and the center of gravity swing amplitude to the expanded pixels to realize motion feature backfilling. For the original dense area after the action feature backfilling is completed, the warning trigger threshold range is adjusted. The trigger thresholds for posture offset angle and center of gravity swing amplitude in the area containing dangerous behavior units are reduced, while the thresholds for areas not containing dangerous behavior units remain unchanged, so as to achieve accurate capture of individual dangerous actions and output of warning information.
10. A real-time early warning system for construction site safety risks based on artificial intelligence, used to implement the real-time early warning method for construction site safety risks based on artificial intelligence as described in any one of claims 1-9, characterized in that, It includes a multi-source feature extraction module, an adhesion region segmentation module, an action difference recognition module, an abnormal action stripping module, and a feature backfilling early warning module; The multi-source feature extraction module obtains the pixel displacement amplitude, contour edge jump frequency, and local occlusion overlap ratio in the densely populated images of the construction area, and generates separation driving identifiers based on the pixel displacement amplitude, contour edge jump frequency, and local occlusion overlap ratio. The adhesion region segmentation module performs gradient segmentation on continuous adhesion regions within the image based on the separation-driven identifier, introduces directional difference constraints at each adhesion boundary, and extracts independent motion segments from the segmented regions. The motion difference recognition module calculates the posture offset angle and center of gravity swing amplitude corresponding to each independent motion segment, and performs local rearrangement processing based on the posture offset angle and center of gravity swing amplitude, highlighting non-consistent motion blocks during the rearrangement process; The abnormal action stripping module performs dynamic contraction and delimitation at the corresponding position based on the non-consistent action block, and superimposes spatial compression ratio constraints during the contraction process to strip local abnormal actions from the continuous adhesion area and construct independently identifiable dangerous behavior units. The feature backfilling and early warning module performs reverse mapping processing on the original dense area based on the dangerous behavior unit, backfills the stripped action features to the corresponding position, and simultaneously adjusts the early warning trigger threshold range. It completes the accurate capture of individual dangerous actions and outputs early warning information in the state of dense aggregation.