Construction site management method and system based on image recognition

Through image recognition technology, the personnel density classification and monitoring of the construction area is solved, and the real-time supervision of the wear of protective equipment for construction personnel is improved and the efficiency of safety management of construction sites is improved.

CN120544231APending Publication Date: 2025-08-26POWERCHINA JIANGXI ELECTRIC POWER ENGINEERING CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510598245.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the prior art, it is impossible to monitor whether construction personnel have wear protective equipment in real time, resulting in safety hazards.

Method used

The target image of the construction area is obtained through image recognition technology, the crowd image is screened, and the area is divided according to the personnel density, the personnel density that cannot be identified in all areas is identified, and whether it meets the density threshold is determined. A list of construction personnel is generated for safety warnings.

Benefits of technology

Real-time monitoring of the construction area is realized, the supervision efficiency of the wear of protective equipment for construction personnel is improved, and the safety management of construction sites is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544231A_ABST
    Figure CN120544231A_ABST
Patent Text Reader

Abstract

The invention provides a construction site management method and system based on image recognition, and relates to the technical field of construction site management. According to the image density of the personnel density, carrying out region division to obtain a first-level region, a second-level region and a third-level region, screening identification results to obtain a crowd image of which the identification features cannot be unidirectionally identified, and obtaining the personnel density in the crowd image of which the identification features cannot be unidirectionally identified; judging whether the personnel density in the crowd image of which the identification features cannot be unidirectionally recognized accords with a density threshold value or not; if yes, a corresponding constructor list is determined according to the crowd image of which the identification features cannot be recognized in a global mode, and safety early warning is conducted on constructors according to the constructor list, so that construction site safety management is achieved. The technical problem that whether a constructor wears protective equipment cannot be supervised in real time is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of construction site management, and in particular to a construction site management method and system based on image recognition. Background Art

[0002] Safety is a difficult problem and key to engineering construction. It is particularly important to make full use of advanced technical means to improve the safety supervision level of engineering projects.

[0003] During the construction process, safety risks are everywhere due to workers' non-standard operations and a high-risk working environment. Therefore, it is necessary to conduct safety monitoring of high-risk targets such as workers' protective equipment (such as safety helmets), scaffolding clips, and manhole covers at the construction site to reduce the occurrence of safety accidents. However, since the monitoring targets are relatively small, this poses a huge challenge to management personnel.

[0004] In the existing technology, the management of whether construction workers at the construction site are wearing protective equipment is relatively lax, and even if there are some management units, they often conduct manual supervision at the entrance of the construction site. However, during the operation, construction workers are generally prone to remove the protective equipment they have worn when entering the site due to the work content or weather factors, such as removing protective equipment when it is hot; in the current situation of special operations, protective equipment needs to be removed but people forget to continue wearing it afterwards. Therefore, the management method of manual supervision at the entrance of the construction site to check whether protective equipment is worn in the existing technology cannot provide real-time supervision of construction workers, and thus poses a safety hazard to the lives of construction workers. Summary of the Invention

[0005] Based on this, the purpose of the present invention is to provide a construction site management method and system based on image recognition, which is used to solve the technical problem in the existing technology that it is impossible to monitor in real time whether construction workers are wearing protective equipment.

[0006] In one aspect, the present invention provides a construction site management method based on image recognition, the method comprising: Acquire a target image of a construction area, the target image including a crowd image, an equipment image, and a site image, and filter the target image to obtain a crowd image, the crowd image including a plurality of construction worker images; Obtaining a density of people based on a crowd image, performing regional division based on the image density of the people density to obtain a regional division result, wherein the regional division result includes, in descending order of image density, a primary region, a secondary region, and a tertiary region; sequentially identifying the primary region, the secondary region, and the tertiary region and screening the identification results to obtain a crowd image with identification features that cannot be fully identified, and obtaining a density of people in the crowd image with identification features that cannot be fully identified, wherein the identification features include a helmet; Determine whether the density of people in a crowd image whose identification features cannot be fully identified meets the density threshold; If not, return to the step of dividing the area according to the image density of the population density until the population density in the crowd image where the identification features cannot be fully identified meets the density threshold; If so, a corresponding list of construction personnel is determined based on the crowd image whose identification features cannot be fully recognized, and a safety warning is issued to the construction personnel based on the list of construction personnel to achieve construction site safety management.

[0007] The above-mentioned construction site management method based on image recognition realizes real-time monitoring of the construction area through the real-time collected images of the crowd in the construction area, which solves the technical problem in the existing technology that it is impossible to supervise in real time whether the construction workers are wearing protective equipment. Specifically, in order to improve the recognition efficiency of real-time monitoring, the personnel density is obtained according to the crowd image, and the area is divided according to the image density of the personnel density to obtain the first-level area, the second-level area and the third-level area, and the recognition results are filtered to obtain the crowd image with the identification feature that cannot be fully identified, and the personnel density in the crowd image with the identification feature that cannot be fully identified is obtained. In order to ensure the recognition and matching efficiency of the system for the subsequent construction personnel list, it is judged whether the personnel density in the crowd image with the identification feature that cannot be fully identified meets the density threshold. If so, the corresponding construction personnel list is determined according to the crowd image with the identification feature that cannot be fully identified, and the construction personnel are given a safety warning based on the construction personnel list to achieve construction site safety management.

[0008] In addition, the above-mentioned construction site management method based on image recognition according to the present invention may also have the following additional technical features: Furthermore, the steps of sequentially identifying the primary area, the secondary area, and the tertiary area and screening the identification results to obtain a crowd image with identification features that cannot be fully identified include: Identify the first-level area and record the first-level area identification time; when the first-level area identification time is greater than the first time threshold, identify the second-level area and record the second-level area identification time; when the second-level area identification time is greater than the second time threshold, identify the third-level area and record the third-level area identification time, wherein the first time threshold is greater than the second time threshold.

[0009] Furthermore, the steps of sequentially identifying the primary area, the secondary area, and the tertiary area and screening the identification results to obtain a crowd image with identification features that cannot be fully identified also include: When the third-level area recognition time is longer than the third time threshold, the system returns to continue recognizing the first-level area and records the first-level area recognition time again; When the first-level area recognition time is greater than the first time threshold again, the second-level area is recognized again and the second-level area recognition time is recorded. When the second-level area recognition time is greater than the second time threshold again, the process returns to the step of continuing to recognize the first-level area and recording the first-level area recognition time again, until the first-level area recognition is completed, and then the third-level area is recognized and the third-level area recognition time is recorded, wherein the second time threshold is greater than the third time threshold.

[0010] Furthermore, the step of obtaining a target image of the construction area includes: Obtaining an initial image of the construction area and extracting image features through multiple convolutional layers. The initial image includes an image of a crowd, an image of equipment, and an image of the site. Determining anchor points by locating corresponding locations in the original image based on the extracted image features, including safety helmets. A rectangular frame is drawn according to the anchor point to obtain a candidate region frame when the target confidence in the area corresponding to the rectangular frame exceeds a confidence threshold, a candidate region is obtained according to the candidate region frame, the candidate region is output to a classification regression network, and a pooling operation is performed through a region of interest pooling layer to achieve size normalization to obtain a feature vector; The feature vector is input into the classification regression network, and the target category in the rectangular box is determined by the classification module. At the same time, the optimal confidence rectangular box and the rectangular box coordinate information corresponding to the optimal confidence rectangular box are determined. The real target box is obtained according to the optimal confidence rectangular box and the rectangular box coordinate information, and the target image is obtained according to the real target box.

[0011] Furthermore, in the step of delineating a rectangular frame based on the anchor point to obtain a candidate region frame when the target confidence within the region corresponding to the rectangular frame exceeds a confidence threshold, the candidate region determination criterion is: Criterion 1: For each image, the image with the largest overlap ratio with the pre-calibrated frame is considered a positive sample; Criterion 2: Preset the overlap ratio IoU threshold range. Assume that the upper limit of the threshold is A and the lower limit of the threshold is B. If IoU>A, it is a positive sample, and if IoU<B, it is a negative sample. Criterion 3: For candidate region bounding boxes that do not meet criteria 1 and 2, no sample processing is performed.

[0012] Furthermore, the classification loss function and the regression loss function are associated to form a multi-task loss function for model training, wherein the expression of the multi-task loss function is: ; Where, i Anchor frame index object; N cls is the number of anchor boxes in the classification loss; Nreg Feature map size; P i represents the target type probability; Represents an identification tag, where When it indicates that there is no target to be identified in the index area, otherwise When , we have; λ is the weight coefficient; t i ( t x , t y , t w , t h ) is a four-point vector prediction coordinate, is the real coordinate of a four-point vector, where the specific expression is: ; ; Where: represents the classification loss; ( x,y ) is the center coordinate of the border, and the horizontal axis coordinate x also represents the coordinate of the border; w、h is the size information of the prediction box.

[0013] Furthermore, the classification loss function is expressed as: ; Where, Represents the bounding box regression loss function, which is expressed as: ; Where, smooth L1 Represents the robust loss function, which is expressed as: .

[0014] Another aspect of the present invention provides a construction site management system based on image recognition, the system comprising: an acquisition module, configured to acquire a target image of a construction area, the target image including a crowd image, an equipment image, and a site image, and to filter the target image to obtain a crowd image, the crowd image including a plurality of construction worker images; a screening module for obtaining a density of people based on a crowd image, performing regional division based on the density of the crowd image to obtain a regional division result, wherein the regional division result includes, in descending order of image density, a primary region, a secondary region, and a tertiary region; sequentially identifying the primary region, the secondary region, and the tertiary region and screening the identification results to obtain a crowd image with identification features that cannot be fully identified, and obtaining the density of people in the crowd image with identification features that cannot be fully identified, wherein the identification features include helmets; A judgment module is used to judge whether the density of people in a crowd image whose identification features cannot be fully identified meets a density threshold; A first execution module is configured to, when the density of people in the crowd image whose identification features cannot be fully identified does not meet the density threshold, return to the step of dividing the area according to the image density of people until the density of people in the crowd image whose identification features cannot be fully identified meets the density threshold; The second execution module is used to determine the corresponding list of construction personnel based on the crowd image whose identification features cannot be fully identified when the density of people in the crowd image meets the density threshold, and to issue safety warnings to the construction personnel based on the construction personnel list to achieve construction site safety management.

[0015] Another aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned construction site management method based on image recognition.

[0016] On the other hand, the present invention also provides a data processing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the construction site management method based on image recognition as described above is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Flowchart of a construction site management method based on image recognition in an embodiment of the present invention; Figure 2 Schematic diagram of the VGG-16 network structure in an embodiment of the present invention; Figure 3 Schematic diagram of a feature fusion and multi-scale detection method in an embodiment of the present invention; Figure 4 Schematic diagram of the structure of the classification regression network in an embodiment of the present invention; Figure 5 Schematic diagram of the RoIPooling layer in an embodiment of the present invention; Figure 6 Schematic diagram of frame regression in an embodiment of the present invention; The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION

[0018] To facilitate understanding of the present invention, the present invention will be described more fully below with reference to the accompanying drawings. The drawings illustrate several embodiments of the present invention. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the present invention.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0020] In order to solve the technical problem that the existing technology cannot monitor in real time whether construction workers are wearing protective equipment, this application provides a construction site management method and system based on image recognition, which realizes real-time monitoring of the construction area through real-time collected images of people in the construction area, solving the technical problem that the existing technology cannot monitor in real time whether construction workers are wearing protective equipment.

[0021] Specifically, in order to improve the recognition efficiency of real-time monitoring, a drone vision system or a vision pod at a fixed high point is used to collect crowd images, and the population density is obtained based on the crowd image. The area is divided according to the image density of the population density to obtain the first-level area, the second-level area and the third-level area, and the recognition results are filtered to obtain the crowd image with the identification feature that cannot be fully identified, and the population density in the crowd image with the identification feature that cannot be fully identified is obtained; in order to ensure the recognition and matching efficiency of the system for the subsequent construction personnel list, it is judged whether the population density in the crowd image with the identification feature that cannot be fully identified meets the density threshold; if so, the corresponding construction personnel list is determined based on the crowd image with the identification feature that cannot be fully identified, and safety warnings are issued to the construction personnel based on the construction personnel list to achieve site safety management.

[0022] To facilitate understanding of the present invention, several embodiments of the present invention are provided below. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive disclosure of the present invention.

[0023] Example 1 See also Figure 1 , which shows a construction site management method based on image recognition in a first embodiment of the present invention, the method includes steps S101 to S104: S101 . Acquire a target image of a construction area, the target image including a crowd image, an equipment image, and a site image. Filter the target image to obtain a crowd image, the crowd image including multiple construction worker images.

[0024] S102. Obtain the density of people based on the crowd image, perform regional division according to the image density of the people density to obtain a regional division result, the regional division result includes the primary region, the secondary region and the tertiary region in descending order according to the image density; identify the primary region, the secondary region and the tertiary region in turn and filter the recognition results to obtain a crowd image with identification features that cannot be fully recognized, and obtain the density of people in the crowd image with identification features that cannot be fully recognized.

[0025] In this embodiment, the identification feature includes a hard hat. As a specific example, construction areas are generally large in area. In order to improve the system's recognition efficiency for target objects and quickly obtain construction workers who are not wearing hard hats, in this embodiment, when identifying the target object, the construction area is first divided into multiple units. Furthermore, since the number and distribution of construction workers in the construction area are random in actual work scenarios, multiple different construction areas often appear depending on the work content. For example, in construction, the work content includes unloading raw materials, mixing sand and gravel, and pouring mortar. Therefore, multiple different construction areas are likely to appear. In the process of identifying the construction area, the initial recognition is likely to default to forming different areas based on the type of work or the work content, and then divide the areas into primary, secondary, and tertiary areas based on the density of each area. This is a single division. If, after the first division, there are still a large number of people in the crowd images of each area, a second division is required to further improve recognition efficiency. At this point, the construction workers in the second division generally belong to the same type of work, but their spatial distribution may differ due to different work content. For example, in mortar pouring, the required personnel include material transfer personnel and material control personnel. Based on this, the spatial layout varies due to differences in work content. At some large-scale construction sites, each type of work is often accompanied by a large number of people. Therefore, the second division provides conditions for recognition efficiency. In actual operation projects, multi-level division can be performed according to actual needs, including three-level division, four-level division, or five-level division.

[0026] Furthermore, in the actual recognition process, due to work requirements, it is often difficult to identify each construction worker in one go. For example, when the system is performing image recognition, a construction worker may be bent over. At this time, the system will not recognize the worker's identification features, that is, it will not recognize the safety helmet. If the construction worker is then labeled as a non-standard person and punished, it is obviously unreasonable. Therefore, it is necessary to mark the construction workers whose identification features are not recognized for secondary recognition, and summarize all construction workers whose identification features are not recognized in the same area to reconstruct the recognition pool. Secondary division is performed based on the density of people in the recognition pool to improve recognition efficiency. In actual operation projects, multiple recognitions can be performed according to actual needs.

[0027] To prevent the system from freezing due to incomplete recognition of all construction workers within the same area, as a specific example, a preset recognition time for each area is set. Specifically, a time threshold is set for each area's recognition time. When the recognition time exceeds this threshold, the system automatically advances to the next area for recognition, and so on, thus avoiding system freezes. Furthermore, since the image density of the primary, secondary, and tertiary areas decreases sequentially, to improve recognition efficiency and avoid situations where the recognition time for the primary area is insufficient and the recognition time for the tertiary area is excessive, in this embodiment, the first time threshold for the primary area is greater than the second time threshold for the secondary area, and greater than the third time threshold for the tertiary area. When the recognition time for the primary area exceeds the first time threshold, the secondary area is recognized and its recognition time is recorded. When the recognition time for the secondary area exceeds the second time threshold, the tertiary area is recognized and its recognition time is recorded. When the recognition time for the tertiary area exceeds the third time threshold, the system records the recognition result and exits.

[0028] In actual operation, in order to fully identify the identification features within each level of area and collect all construction workers not wearing helmets, multiple rounds of recognition are required after the first round of recognition. Specifically, when the recognition time for the third-level area exceeds the third time threshold, the process returns to continue recognizing the first-level area and recording the recognition time for the first-level area again. When the recognition time for the first-level area again exceeds the first time threshold, the process returns to continue recognizing the second-level area and recording the recognition time for the second-level area. When the recognition time for the second-level area again exceeds the second time threshold, the process returns to continue recognizing the first-level area and recording the recognition time for the first-level area again. After the recognition of the first-level area is completed, the third-level area is recognized and the recognition time for the third-level area is recorded, where the second time threshold is greater than the third time threshold. It should be further explained that since the density of the first-level area is higher and the number of construction workers is larger, the density of the third-level area is the lowest, and the probability of successful recognition in one go is higher. Even if successful, the remaining number of recognitions will be lower than that of other areas. Therefore, in the second round of recognition, the focus should be on the first-level area, followed by the second and third-level areas. It should be further explained that, since construction workers are often in an active state, they can be fully identified when the time threshold is reasonable. Since construction workers are often working continuously, when a second round of identification is performed on the area after one round of identification, the number of construction workers that need to be identified in each area has been greatly reduced. Therefore, it is necessary to perform refined identification on each area. Therefore, in this embodiment, when performing the second round of identification, after all the identification features in the first-level area are identified, the second-level area and the third-level area are identified in turn, and the cycle is repeated. In addition, during the multi-round identification process, as the number of identification rounds increases, the number of construction workers to be identified in each area is definitely decreasing. Therefore, in order to further improve the identification efficiency, in this embodiment, the first time threshold, the second time threshold, and the third time threshold are also reduced in equal proportion step by step. For example, in the first round of recognition, the first time threshold is 100s, the second time threshold is 80s, and the third time threshold is 60s; in the second round of recognition, the first time threshold is 50s, the second time threshold is 40s, and the third time threshold is 30s. Compared with the first round of recognition, the recognition time threshold is proportionally reduced by 1 / 2. To avoid the situation where multiple rounds of recognition freeze, as a specific example, based on actual experience, the system generally completes all recognition after 4 rounds. Therefore, in this embodiment, the upper limit of the number of cycles is set to 5 rounds. That is, after 5 rounds of recognition, if there are still areas where the identification features have not been fully recognized, the system will no longer recognize, collect the recognition results as the final recognition results, and exit.

[0029] S103: Determine whether the density of people in the crowd image whose identification features cannot be fully identified meets the density threshold.

[0030] Since there are often many construction workers in the construction area, in order to avoid the system crash or low recognition efficiency caused by the large number of people in the crowd image with non-fully identifiable identification features in the same area, in this embodiment, the system first judges the density of people in the crowd image with non-fully identifiable identification features. When the density of people in the crowd image with non-fully identifiable identification features meets the conditions, the corresponding list of construction workers is determined to improve the efficiency of subsequent image recognition.

[0031] If not, the process returns to step S102 until the density of people in the crowd image whose identification features cannot be fully recognized meets the density threshold. If so, the process proceeds to step S104.

[0032] S104: Determine a corresponding list of construction personnel based on the crowd image whose identification features cannot be fully recognized, and issue safety warnings to the construction personnel based on the list of construction personnel to achieve construction site safety management.

[0033] Specifically, safety warning methods include loudspeaker warnings, telephone warnings, etc.

[0034] In this embodiment, the step of acquiring the target image of the construction area includes steps S1011 to S1013: S1011. Obtain an initial image of the construction area to extract image features through multiple convolutional layers. The initial image includes a crowd image, an equipment image, and a site image. Based on the extracted image features, obtain the corresponding position in the original image to determine the anchor point. The image features include a safety helmet.

[0035] S1012. A rectangular frame is drawn according to the anchor point to obtain a candidate region border when the target confidence in the area corresponding to the rectangular frame exceeds the confidence threshold. A candidate region is obtained according to the candidate region border to output the candidate region to a classification regression network, and a pooling operation is performed through a region of interest pooling layer to achieve size normalization and thereby obtain a feature vector.

[0036] S1013. Input the feature vector into the classification regression network, determine the target category in the rectangular box through the classification module, and simultaneously determine the optimal confidence rectangular box and the rectangular box coordinate information corresponding to the optimal confidence rectangular box, obtain the true target box based on the optimal confidence rectangular box and the rectangular box coordinate information, and obtain the target image based on the true target box.

[0037] As a specific example, due to the outstanding performance of the Faster R-CNN object detection model, in this embodiment, a Faster R-CNN model was constructed for object detection. All stages of this network are completed within a convolutional neural network, resulting in an end-to-end, two-stage detection model. To address the large discrepancies in the scales of security risk targets in captured images, feature fusion and multi-scale detection methods based on the Faster R-CNN model were introduced during the security risk feature extraction process, significantly improving the accuracy of security risk target detection.

[0038] Specifically, the Faster R-CNN model consists of three main components: a VGG-16-based feature extraction backbone network, a Region Proposal Network (RPN), and a classification and regression network. First, the target detection image is input, scaled to a fixed size, and features are extracted using multiple convolutional layers. The output feature map is then fed into the RPN network, which locates corresponding locations in the original image, known as anchor points, and adds rectangular boxes of varying sizes and aspect ratios. When the target confidence within the rectangular box exceeds a threshold, the box is considered a precise candidate region. The candidate region is then fed into the classification and regression network, where it is pooled and resized using the Region Proposal Pooling (RoIPooling) layer to generate a feature vector. This feature vector is then fed into the classification and regression network, where the classification module determines the target category within the box and determines the optimal confidence box and its coordinates.

[0039] First, the backbone network features are extracted and fused. The backbone network of the Faster R-CNN model uses the VGG-16 network for feature extraction. The VGG series network structure is simple and very practical, such as Figure 2 As shown, the 13 convolutional layers and non-weighted pooling layers are divided into five groups. The first two groups each contain two convolutional layers and one pooling layer, while the last three groups each contain three convolutional layers and one pooling layer. Each convolutional layer group undergoes a max-pooling operation after convolution, reducing computational complexity and improving model scalability and classification decision-making capabilities. The VGG-16 network, consisting of five network stages consisting of 13 convolutional and pooling layers, is connected to three fully connected layers and a softmax layer for classification. This network performs the backbone image feature extraction and plays a critical role in the performance of the entire object detection model. The VGG-16 network uses the same 3×3 convolution kernel size and max-pooling size. The Feature Pyramid Network (FPN) utilizes feature information of different scales applied to different convolutional layers to predict objects of different scales.

[0040] Small objects contain relatively concentrated and less useful information due to their small pixels. Using a fixed 3×3 convolution kernel for convolution, image features are not fully extracted. In order to obtain more information during prediction, the RPN network uses feature fusion to extract features, such as Figure 3 As shown in the figure, a variety of different size feature maps are constructed by integrating the five network stages of VGG-16, and the feature pyramid is selectively fused multiple times to complete the detection target feature fusion and multi-scale detection operations.

[0041] In the feature fusion operation, the convolutional layer C4 in the fourth network stage undergoes a 1×1×256 convolution operation to reduce its number of channels to 256. Deconvolution is then used to restore C4 to the same size as C2. Simultaneously, C2 undergoes the same 1×1×256 convolution operation to reduce its number of channels to 256. C2 and C4 are then fused using an additive fusion function. To remove overlap caused by deconvolution, the pixel-superimposed feature map undergoes a 3×3×256 convolution operation to reduce its number of channels to 256, forming a new feature layer P2. The same method is used to fuse the feature layers C5 in the fifth network stage and C3 in the third network stage to form a new feature layer P3. C4 undergoes a 1×1×256 convolution and a 3×3×256 convolution to obtain P4, and C5 undergoes the same operation to obtain P5.

[0042] At this point, four pyramid feature layers (P5, P4, P3, and P2) of fixed dimensions are obtained. Their shallow features are integrated with deep features, fully extracting the features of small objects while ensuring the recognition of medium-to-large objects. These new features are then integrated into the RPN network to obtain the model's predicted regions of interest (RoIs), enabling multi-scale detection. RPN-P4 and RPN-P5 have larger receptive fields, with anchor window sizes of 64×64, 128×128, and 256×256, respectively. RPN-P2 and RPN-P3 have smaller receptive fields of 24×24, 32×32, and 64×64, respectively, with aspect ratios of 1:1, 2:1, and 1:2, respectively. The candidate regions of the four parts are input into the RoIPooling layer and mapped back to RPN-P2, RPNP3, RPN-P4, and RPN-P5 respectively to obtain 7×7 size normalized feature maps. The cascade fusion function is introduced to vertically splice the feature maps from P2 to P4 for joint classification and regression operations.

[0043] After the above feature fusion operation, four feature pyramids, P5, P4, P3, and P2, are obtained. P5 and P4 fully exploit the semantic information of high-level features, while P3 and P2 fully exploit the positional information of low-level feature maps. Multi-scale detection is then achieved through the RPN network, making the RPN more accurate in judging foreground objects. This method is used to extract features of construction safety risk targets, fully mining information at the feature layer and significantly improving the identification and classification of construction safety risk targets.

[0044] The RPN network then generates candidate regions. Specifically, the main function of the RPN is to batch-generate candidate regions within Faster R-CNN. As a fully convolutional neural network, the RPN solves the problem of the R-CNN series's selective search taking too long. Faster R-CNN extracts and generates feature maps and directly inputs them into the RPN network. By setting anchor points to divide the windows into a fixed number, it outputs rectangular candidate regions of various sizes, thereby generating several candidate regions as initial detection boxes. Each initial detection box is then passed through the Softmax classifier in the classification and regression network to select foreground candidate regions, and then adjusted using bounding box regression to obtain a feature sub-graph.

[0045] Assume that the feature extraction network generates a feature map of size N×M and feeds it into the RPN. If the feature extraction network generates a fixed feature map of size 256×N×M, the VGG16 network generates a feature map of dimension 512×N×M. For the RPN network, the input feature map first passes through a sliding window. The sliding window takes the n×n sized region in the grid as input, processes it through the intermediate layers, and outputs a uniform number of low-dimensional features, typically n=3. At each position in the feature map where the sliding window passes, a multi-scale approach is used to set k anchor points to predict the target location. The aspect ratios of the anchor points are typically width:height∈{1:1, 1:2, 2:1}. Table 1 shows the candidate region sizes of different anchor scales during feature extraction.

[0046] Table 1: RPN different anchor point scales and ratios

[0047] As shown in Table 1, for each feature point, the anchor point generates 9 candidate regions according to 3 scales of 128, 256, and 512 and three different aspect ratios of 1:1, 1:2, and 2:1, namely, the feature Figure 1 A total of N×M×9 anchor boxes are generated.

[0048] To train the RPN network, N×M×k anchor boxes need to be classified and marked after 1×1 convolution processing, and positive and negative sample information is added to obtain N×M×2k classification feature vectors. The Softmax classifier is used to determine whether the candidate box contains the target. If so, the positive sample parameter is set to 1 and the negative sample is set to 0. If not, the negative sample parameter is set to 1 and the positive sample is set to 0, thereby extracting the positive sample containing the target and obtaining the candidate region. The judgment criteria are: (1) For each image, the image with the largest overlap ratio with the pre-calibrated frame is determined to be a positive sample; (2) By setting the Intersection over Union (IoU) threshold range, assuming that the upper limit of the threshold is A and the lower limit of the threshold is B, if IoU>A is a positive sample, if IoU<B is a negative sample; (3) For candidate region borders that do not meet the two criteria 1 and 2, no sample processing is performed.

[0049] Furthermore, regarding the classification and regression network, specifically: after the model obtains candidate regions through the RPN network, it needs to perform detailed classification and regression. The initial size of the image input model is fixed. The RPN network regresses and adjusts the anchor boxes with positive scores to generate region candidate boxes of varying sizes and shapes. The features extracted by the backbone network and the candidate regions generated by the RPN network are connected to the classification and regression network. Using detailed regression and classification processing, the correction parameters and confidence levels of each category feature are obtained. Figure 4 For the classification and regression network structure, first connect a RoIPooling layer to pool the candidate areas into feature maps of the same size. To complete the mapping of the feature maps, two fully connected layers fc6 and fc7 are connected after the RoIPooling layer. Then, parallel fully connected layers fc / cls and fc / bbox_reg are connected. fc / bbox_reg is used to obtain the correction parameters of the features, and fc / cls is used to obtain the scores of the features. Finally, a Softmax layer is added to convert the feature scores into confidence levels.

[0050] The RoIPooling pooling method is introduced for standardization and the size of the anchor box is fixed, such as Figure 5As shown in the figure, an 800×800 image is fed into a convolutional neural network. The convolutional layers use VGG16 with a stride of 32. After the convolutional layers, the image size is reduced to 1 / 32 of its original size, and a 25×25 feature map is output. The original image contains a 665×665 region proposal. The size of the proposed region mapped to the image becomes 665 / 32 = 20.78, resulting in a 20.78×20.78 feature map. The network then rounds the region down to 20×20, completing the first quantization step of the RoIPooling layer. The 20×20 feature map needs to be pooled to a size of 7×7, so the 20×20 feature map is divided into 49 equal-sized areas. The size of each small area is 20 / 7=2.86, and the network is quantized for the second time to make the area size 2×2. Finally, in each 2×2 small area, the maximum value of the four pixels is taken out as the "representative" value of the area. The 49 small areas output a total of 49 pixel values, which together constitute a 7×7 region of interest, achieving fixed-length output, which is conducive to output to the next layer of network for further processing.

[0051] For confidence setting, Faster R-CNN uses Softmax in the final stage of candidate region classification to obtain the confidence of detecting targets of each category. When setting a fixed threshold, a threshold that is too high will mix in many false targets, while a threshold that is too low will miss some true targets. The traditional approach is to perform multiple tests on the training set and take the average value. The applicability of the threshold obtained in this way has certain limitations. Therefore, based on the fixed threshold acquisition method, an adaptive threshold is introduced to improve the model's decision-making ability for different types of security risk targets. The adaptive threshold introduces a second-order difference method, so that the discrete confidence arrays are correlated with each other and jointly reflect the confidence change trend, so as to obtain the corresponding adaptive threshold. An image is searched for targets in the RPN network, and the candidate region generates a recognition value of 300. Each candidate region obtains a confidence representing a category, and a total of 300 a×1 arrays are obtained. Take the maximum value of each array and sort it in descending order, discard the value less than 0.1, and obtain an n×1 array C. The confidence level decreases, so let f (﹒) is the function of the dynamic trend of the associated confidence, f (C k ) takes its maximum value, C k The confidence threshold for this test target is as follows: .

[0052] Finally, a loss function is constructed. Loss is the gap between the predicted results and the actual results described in the neural network training process. This gap can be optimized by the loss function to reduce the loss by optimizing the network parameters. In the early R-CNN target detection network model, network training is generally carried out in a multi-link pipeline form. This method requires training modules separately, such as CNN network, classifier, and regional regression. This increases the burden on the computer GPU and consumes huge hard disk space. The Faster R-CNN model associates the classification loss function and the regression loss function to form a multi-task loss function for model training, thereby improving the efficiency of storage space utilization. The expression of multi-task loss is: ; Where, i Anchor frame index object; N cls is the number of anchor boxes in the classification loss; N reg Feature map size; P i represents the target type probability; Represents an identification tag, where When it indicates that there is no target to be identified in the index area, otherwise When , we have; λ is the weight coefficient; t i ( t x , t y , t w , t h ) is a four-point vector prediction coordinate, is the real coordinate of a four-point vector, where the specific expression is: ; ; Where: represents the classification loss; ( x,y ) is the center coordinate of the border, and the horizontal axis coordinate x also represents the coordinate of the border; w、h is the size information of the prediction box; taking the helmet as an example, Figure 6 This is the bounding box regression of the helmet, where the manually annotated green bounding box is the true target box, the RPN outputs a blue bounding box, and the helmet predicted target box is a red bounding box.

[0053] Furthermore, the classification loss function is expressed as: ; Where, Represents the bounding box regression loss function, which is expressed as: ; Where, smooth L1 Represents the robust loss function, which is expressed as: .

[0054] In summary, the construction site management method based on image recognition in the above-mentioned embodiment of the present invention realizes real-time monitoring of the construction area through the real-time collected crowd images of the construction area, which solves the technical problem in the prior art that it is impossible to supervise in real time whether the construction workers are wearing protective equipment; specifically, in order to improve the recognition efficiency of real-time monitoring, the personnel density is obtained according to the crowd image, and the area is divided according to the image density of the personnel density to obtain the first-level area, the second-level area and the third-level area, and the recognition results are filtered to obtain the crowd image with the identification feature that cannot be recognized globally, and the personnel density in the crowd image with the identification feature that cannot be recognized globally is obtained; in order to ensure the recognition and matching efficiency of the system for the later construction personnel list, it is judged whether the personnel density in the crowd image with the identification feature that cannot be recognized globally meets the density threshold; if so, the corresponding construction personnel list is determined according to the crowd image with the identification feature that cannot be recognized globally, and the construction personnel are given a safety warning based on the construction personnel list to realize construction site safety management.

[0055] Example 2 A second embodiment of the present invention provides a construction site management system based on image recognition, comprising: an acquisition module, configured to acquire a target image of a construction area, the target image including a crowd image, an equipment image, and a site image, and to filter the target image to obtain a crowd image, the crowd image including a plurality of construction worker images; a screening module for obtaining a density of people based on a crowd image, performing regional division based on the density of the crowd image to obtain a regional division result, wherein the regional division result includes, in descending order of image density, a primary region, a secondary region, and a tertiary region; sequentially identifying the primary region, the secondary region, and the tertiary region and screening the identification results to obtain a crowd image with identification features that cannot be fully identified, and obtaining the density of people in the crowd image with identification features that cannot be fully identified, wherein the identification features include helmets; A judgment module is used to judge whether the density of people in a crowd image whose identification features cannot be fully identified meets a density threshold; A first execution module is configured to, when the density of people in the crowd image whose identification features cannot be fully identified does not meet the density threshold, return to the step of dividing the area according to the image density of people until the density of people in the crowd image whose identification features cannot be fully identified meets the density threshold; The second execution module is used to determine the corresponding list of construction personnel based on the crowd image whose identification features cannot be fully identified when the density of people in the crowd image meets the density threshold, and to issue safety warnings to the construction personnel based on the construction personnel list to achieve construction site safety management.

[0056] In summary, the construction site management system based on image recognition in the above-mentioned embodiment of the present invention realizes real-time monitoring of the construction area through the real-time collected crowd images of the construction area, which solves the technical problem in the prior art that it is impossible to supervise in real time whether the construction workers are wearing protective equipment; specifically, in order to improve the recognition efficiency of real-time monitoring, the personnel density is obtained according to the crowd image, and the area is divided according to the image density of the personnel density to obtain the first-level area, the second-level area and the third-level area, and the recognition results are filtered to obtain the crowd image with the identification feature that cannot be recognized globally, and the personnel density in the crowd image with the identification feature that cannot be recognized globally is obtained; in order to ensure the recognition and matching efficiency of the system for the later construction personnel list, it is judged whether the personnel density in the crowd image with the identification feature that cannot be recognized globally meets the density threshold; if so, the corresponding construction personnel list is determined according to the crowd image with the identification feature that cannot be recognized globally, and the construction personnel are given a safety warning based on the construction personnel list to realize construction site safety management.

[0057] In addition, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the method in the above embodiment when the program is executed by a processor.

[0058] In addition, an embodiment of the present invention further provides a data processing device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method in the above embodiment when executing the program.

[0059] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0060] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting, or processing it in another suitable manner as necessary, and then storing it in a computer memory.

[0061] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.

[0062] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0063] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

Claims

1. A construction site management method based on image recognition, characterized in that: Methods include: Acquire a target image of a construction area, the target image including a crowd image, an equipment image, and a site image, and filter the target image to obtain a crowd image, the crowd image including a plurality of construction worker images; Obtaining a density of people based on a crowd image, performing regional division based on the image density of the people density to obtain a regional division result, wherein the regional division result includes, in descending order of image density, a primary region, a secondary region, and a tertiary region; sequentially identifying the primary region, the secondary region, and the tertiary region and screening the identification results to obtain a crowd image with identification features that cannot be fully identified, and obtaining a density of people in the crowd image with identification features that cannot be fully identified, wherein the identification features include a helmet; Determine whether the density of people in a crowd image whose identification features cannot be fully identified meets the density threshold; If not, return to the step of dividing the area according to the image density of the population density until the population density in the crowd image where the identification features cannot be fully identified meets the density threshold; If so, the corresponding list of construction personnel is determined based on the crowd image whose identification features cannot be fully recognized, and safety warnings are issued to the construction personnel based on the list of construction personnel to achieve construction site safety management.

2. The construction site management method based on image recognition according to claim 1, characterized in that: The steps of sequentially identifying the primary area, the secondary area, and the tertiary area and screening the identification results to obtain a crowd image with identification features that cannot be fully identified include: Identify the first-level area and record the time it takes to identify the first-level area; When the first-level area recognition time is longer than the first time threshold, the second-level area is recognized and the second-level area recognition time is recorded; When the second-level area recognition time is greater than the second time threshold, the third-level area is recognized and the third-level area recognition time is recorded, wherein the first time threshold is greater than the second time threshold.

3. The construction site management method based on image recognition according to claim 2, characterized in that: The steps of sequentially identifying the primary area, the secondary area, and the tertiary area and screening the identification results to obtain a crowd image with identification features that cannot be fully identified also include: When the third-level area recognition time is longer than the third time threshold, the system returns to continue recognizing the first-level area and records the first-level area recognition time again; When the first-level area recognition time is greater than the first time threshold again, the second-level area is recognized again and the second-level area recognition time is recorded. When the second-level area recognition time is greater than the second time threshold again, the process returns to the step of continuing to recognize the first-level area and recording the first-level area recognition time again, until the first-level area recognition is completed, and then the third-level area is recognized and the third-level area recognition time is recorded, wherein the second time threshold is greater than the third time threshold.

4. The construction site management method based on image recognition according to claim 1, characterized in that: The steps to obtain the target image of the construction area include: Obtaining an initial image of the construction area and extracting image features through multiple convolutional layers. The initial image includes an image of a crowd, an image of equipment, and an image of the site. Determining anchor points by locating corresponding locations in the original image based on the extracted image features, including safety helmets. A rectangular frame is drawn according to the anchor point to obtain a candidate region frame when the target confidence in the area corresponding to the rectangular frame exceeds a confidence threshold, a candidate region is obtained according to the candidate region frame, the candidate region is output to a classification regression network, and a pooling operation is performed through a region of interest pooling layer to achieve size normalization to obtain a feature vector; The feature vector is input into the classification regression network, and the target category in the rectangular box is determined by the classification module. At the same time, the optimal confidence rectangular box and the rectangular box coordinate information corresponding to the optimal confidence rectangular box are determined. The real target box is obtained according to the optimal confidence rectangular box and the rectangular box coordinate information, and the target image is obtained according to the real target box.

5. The construction site management method based on image recognition according to claim 4, characterized in that: In the step of delineating a rectangular frame based on the anchor points to obtain a candidate region frame when the target confidence within the region corresponding to the rectangular frame exceeds a confidence threshold, the candidate region determination criteria are: Criterion 1: For each image, the image with the largest overlap ratio with the pre-calibrated frame is considered a positive sample; Criterion 2: Preset the overlap ratio IoU threshold range. Assume that the upper limit of the threshold is A and the lower limit of the threshold is B. If IoU>A, it is a positive sample, and if IoU<B, it is a negative sample. Criterion 3: For candidate region bounding boxes that do not meet criteria 1 and 2, no sample processing is performed.

6. The construction site management method based on image recognition according to claim 4, characterized in that: The classification loss function and the regression loss function are associated to form a multi-task loss function for model training. The expression of the multi-task loss function is: ; Where, i Anchor frame index object; N cls is the number of anchor boxes in the classification loss; N reg Feature map size; P i represents the target type probability; Represents an identification tag, where When it indicates that there is no target to be identified in the index area, otherwise When , we have; λ is the weight coefficient; t i ( t x , t y , t w , t h ) is a four-point vector prediction coordinate, is the real coordinate of a four-point vector, where the specific expression is: ; ; Where: represents the classification loss; ( x,y ) is the center coordinate of the border, and the horizontal axis coordinate x also represents the coordinate of the border; w、h is the size information of the prediction box.

7. The construction site management method based on image recognition according to claim 6, characterized in that: The expression of the classification loss function is: ; Where, Represents the bounding box regression loss function, which is expressed as: ; Where, smooth L1 Represents the robust loss function, which is expressed as: 。 8. A construction site management system based on image recognition, characterized in that: The system comprises: an acquisition module, configured to acquire a target image of a construction area, the target image including a crowd image, an equipment image, and a site image, and to filter the target image to obtain a crowd image, the crowd image including a plurality of construction worker images; a screening module for obtaining a density of people based on a crowd image, performing regional division based on the density of the crowd image to obtain a regional division result, wherein the regional division result includes, in descending order of image density, a primary region, a secondary region, and a tertiary region; sequentially identifying the primary region, the secondary region, and the tertiary region and screening the identification results to obtain a crowd image with identification features that cannot be fully identified, and obtaining the density of people in the crowd image with identification features that cannot be fully identified, wherein the identification features include helmets; A judgment module is used to judge whether the density of people in a crowd image whose identification features cannot be fully identified meets a density threshold; A first execution module is configured to, when the density of people in the crowd image whose identification features cannot be fully identified does not meet the density threshold, return to the step of dividing the area according to the image density of people until the density of people in the crowd image whose identification features cannot be fully identified meets the density threshold; The second execution module is used to determine the corresponding list of construction personnel based on the crowd image whose identification features cannot be fully identified when the density of people in the crowd image meets the density threshold, and to issue safety warnings to the construction personnel based on the construction personnel list to achieve construction site safety management.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the construction site management method based on image recognition as described in any one of claims 1 to 7 is implemented.

10. A data processing device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the construction site management method based on image recognition as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Safety helmet detection and identity recognition method and device for electric power operation site

    CN114067268A

  • Weather consultation video conference roll call method and system

    CN116052260A

  • Dense area target detection method based on density cascade clustering

    CN116188757A