A seat belt violation detection method combining object detection and semantic segmentation
By combining object detection and semantic segmentation technology, a network model is constructed to identify the wearing status of the seat belt, which solves the problem of not being able to judge the standardization and misidentification of seat belt wearing in the prior art, and achieves efficient and accurate detection of illegal wearing of seat belts.
Patent Information
- Application Number
- CN202210766081.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-01
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-07-01
AI Technical Summary
The existing high-altitude seat belt detection methods cannot effectively judge the wearing normativeness of seat belts, and it is easy to misidentify unrelated personnel when used in mobile devices.
Combining object detection and semantic segmentation technology, the object detection network model and semantic segmentation network model are constructed, which are used to predict personnel key point positioning information, full-body lacing detection information and scaffolding area positioning information, and identify the position of the lanyard to determine whether there are illegal wearing behaviors of the seat belt.
It realizes intelligent detection of illegal wearing of seat belts, improves the real-time and accuracy of detection, avoids misidentification of irrelevant personnel, and is suitable for image detection collected by mobile devices.
Smart Images

Figure CN115131732B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of seat belt wearing detection, and particularly to a method for detecting seat belt violation wearing combining object detection and semantic segmentation. Background Art
[0002] During the process of urban construction, a scaffolding is a personnel working platform erected for the orderly progress of relevant construction activities. An aerial safety belt is the safety guarantee for scaffolding constructors, which is divided into two parts: a full-body harness and a lanyard. When working, the aerial safety belt should be worn in a standardized manner, that is, the lanyard part should be higher than the position above the operator's waist to play a certain buffering role for personnel in case of accidental fall. However, due to the lack of supervision, some constructors do not wear the safety belt in a standardized manner or even do not wear it at all when working in the scaffolding area, which finally leads to accidents of personnel falling from height and causes serious negative social impacts. In order to prevent such accidents from occurring, the detection of seat belt violation wearing of relevant construction personnel can be completed through deep learning, which can accelerate the construction process of a smart construction site.
[0003] At present, some scholars have studied the wearing of aerial safety belts by personnel. Generally, they directly use a convolutional neural network to extract features from relevant images and finally complete a classification task, that is, two results of wearing and not wearing, without further studying the wearing standardization. Since the above research does not limit the area of personnel, when traditional detection methods are used on mobile devices, it is possible to misjudge irrelevant personnel on the ground, resulting in waste of unnecessary computing resources.
[0004] To sum up, the existing methods for detecting aerial safety belts mainly have the following technical problems:
[0005] (1) The detection target is single and cannot further determine the wearing standardization of the seat belt;
[0006] (2) There is no area restriction, and when the algorithm is applied to mobile devices, it will cause misidentification of irrelevant personnel. Summary of the Invention
[0007] In view of the above deficiencies in the prior art, the present invention provides a method for detecting seat belt violation wearing combining object detection and semantic segmentation.
[0008] In order to achieve the above invention purpose, the technical solution adopted by the present invention is:
[0009] A method for detecting seat belt violation wearing combining object detection and semantic segmentation, comprising the following steps:
[0010] S1. Collect image data of the construction site;
[0011] S2. Build a target detection network model, and use the collected construction site image data to predict the personnel key point positioning information, the whole body lanyard detection information, and the scaffolding area positioning information;
[0012] S3. Judge whether to conduct the safety belt hanging rope detection according to the relative positions of the personnel key point positioning information, the whole body lanyard detection information, and the scaffolding area positioning information; if so, execute step S4; otherwise, end the process;
[0013] S4. Extract the rectangular frame image of the hanging rope area according to the personnel key point positioning information;
[0014] S5. Build a semantic segmentation network model to identify the hanging rope position from the extracted rectangular frame image of the hanging rope area;
[0015] S6. Determine whether there is any violation in wearing the safety belt according to the hanging rope position.
[0016] Optionally, step S2 specifically includes the following sub-steps:
[0017] S2-1. Use a feature fusion convolutional neural network to extract a feature image from the construction site image data;
[0018] S2-2. Upsample the extracted feature image to generate a branch feature map;
[0019] S2-3. Use a multi-branch prediction neural network to predict the branch feature map to obtain the personnel key point positioning information, the whole body lanyard detection information, and the scaffolding area positioning information.
[0020] Optionally, the calculation process of the feature fusion convolutional neural network in step S2-1 is as follows:
[0021]
[0022]
[0023]
[0024] Among them, I represents the iterative output of the depth feature of the nth layer network; x n represents the output of the nth layer network; Node represents the aggregation function; represents the output of the network from the mth layer to the nth layer by the intermediate calculation module; T m (x) represents the output of the mth layer feature fusion convolutional neural network; T n represents the output of the nth layer feature fusion convolutional neural network. represents the output of the network from the first layer to the nth layer by the intermediate calculation module; B represents the convolutional function; represents the output of the subsequent calculation module from the first layer to the nth layer network.
[0025] Optionally, the calculation method for upsampling the extracted feature image in step S2-2 is as follows:
[0026] o = s(i - 1)+2p - k + 2
[0027] Where i represents the input size of the feature image; s represents the size of the stride; p represents the padding value of the feature image boundary; k represents the size of the convolution kernel; o represents the size of the output branch feature map.
[0028] Optionally, in step S3, determining whether to perform the safety belt hanging rope detection according to the relative positions of the human key point positioning information, the whole body lacing detection information, and the scaffolding area positioning information specifically includes:
[0029] Calculating the intersection relationship between the human target box and the scaffolding rotation target box according to the human key point positioning information and the scaffolding area positioning information;
[0030] Judging whether to perform the safety belt hanging rope detection according to the intersection relationship between the human target box and the scaffolding rotation target box in combination with the whole body lacing detection information.
[0031] Optionally, the specific method for calculating the intersection relationship between the human target box and the scaffolding rotation target box according to the human key point positioning information and the scaffolding area positioning information is as follows:
[0032] Select one side of the human target box as the first line segment, select one side of the scaffolding rotation target box as the second line segment, and calculate the intersection result of the first line segment and the second line segment using the following formula:
[0033] result1 = sin(θ1)×sin(θ2)
[0034] result2 = sin(θ3)×sin(θ4)
[0035] Where result1 represents the first intersection result; θ1 represents the angle between one endpoint of the first line segment and the two endpoints of the second line segment; θ2 represents the angle between the other endpoint of the first line segment and the two endpoints of the second line segment; result2 represents the second intersection result; θ3 represents the angle between one endpoint of the second line segment and the two endpoints of the first line segment; θ4 represents the angle between the other endpoint of the second line segment and the two endpoints of the first line segment;
[0036] When both the first intersection result and the second intersection result are less than or equal to 0, the first line segment and the second line segment intersect; otherwise, the first line segment and the second line segment do not intersect.
[0037] Optionally, step S4 specifically includes:
[0038] Select the leg corresponding to the larger ordinate among the left ankle key point and the right ankle key point in the key point positioning information of the person as the discrimination target, and use the distance from the knee key point to the ankle key point in this leg as the height of the rectangular frame;
[0039] Use a set multiple of the larger Euclidean distance from the left shoulder key point to the left hip key point and from the right shoulder key point to the right hip key point as the width of the rectangular frame;
[0040] Use the knee key point and the ankle key point as the midpoints of the upper and lower sides of the rectangular frame respectively.
[0041] Optionally, step S5 specifically includes the following sub-steps:
[0042] S5-1. Perform unified size processing on the rectangular frame image of the extracted lanyard area;
[0043] S5-2. Use a convolutional neural network to extract the convolutional features of the rectangular frame image processed in step S5-1;
[0044] S5-3. Use an ASPP pyramid structure to perform feature fusion on the convolutional features extracted in step S5-2 to obtain an encoded feature map;
[0045] S5-4. Use the bilinear interpolation method to decode the obtained encoded feature map to identify the lanyard position.
[0046] Optionally, in step S5-3, the ASPP pyramid structure uses 6, 12, and 18 dilated convolutions for feature extraction of different receptive fields of the image, and then performs feature fusion processing through a pooling layer.
[0047] Optionally, the calculation method of the dilated convolution is:
[0048]
[0049] l = w+(w - 1)*(u - 1)
[0050] where w represents the size of the ordinary convolution kernel; l represents the size of the dilated convolution kernel; t represents the stride of the dilated convolution; q represents the image padding value; u represents the number of pixels with value 0 in the dilated convolution; v represents the size of the feature map output after the dilated convolution.
[0051] The present invention has the following beneficial effects:
[0052] The present invention predicts the positioning information of personnel key points, the detection information of full-body lanyards, and the positioning information of the scaffolding area by constructing a target detection network model for the collected image data of the construction site, and determines whether to perform the detection of the safety belt hanging rope according to the relative positions of the positioning information of personnel key points and the detection information of full-body lanyards and the positioning information of the scaffolding area; then, by extracting the rectangular frame image of the hanging rope area and using a semantic segmentation network model to identify the hanging rope position, the intelligent detection of the illegal wearing of the safety belt is realized, the real-time performance and accuracy of the detection are improved, and it is more suitable for the image detection collected by mobile devices. Description of the Drawings
[0053] Figure 1 It is a schematic flowchart of a method for detecting illegal wearing of a safety belt combining target detection and semantic segmentation in an embodiment of the present invention. Detailed Embodiments
[0054] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0055] As Figure 1 shown, an embodiment of the present invention provides a method for detecting illegal wearing of a safety belt combining target detection and semantic segmentation, including the following steps S1 to S6:
[0056] S1. Collect image data of the construction site;
[0057] In an alternative embodiment of the present invention, step S1 of the present invention can use an intelligent safety helmet to collect image data of the construction site of scaffolding construction workers, and then perform subsequent detection of illegal wearing of safety belts. The present invention can effectively utilize the existing intelligent safety helmets, so that the construction site data is no longer limited to fixed cameras, and the real-time performance and accuracy of the detection of illegal wearing of safety belts are improved.
[0058] S2. Construct a target detection network model, and use the collected image data of the construction site to predict the positioning information of personnel key points, the detection information of full-body lanyards, and the positioning information of the scaffolding area;
[0059] In an alternative embodiment of the present invention, by analyzing the characteristics of the on-site image data of scaffolding construction workers, when the workers wear safety belts in a standard manner, since the hanging position of the hook on the hanging rope is relatively high, there is no safety belt hanging rope below the knees of the workers; on the contrary, when not wearing it in a standard manner, the safety belt hanging rope will appear below the knees of the workers. Therefore, the present invention first constructs a target detection network model to obtain the positioning information of human key points, the detection information of full-body lacing, and the positioning information of the scaffolding area.
[0060] Step S2 of the present invention specifically includes the following sub-steps S2-1 to S2-3:
[0061] S2-1. Use a feature fusion convolutional neural network to extract a feature image from the on-site image data of the construction site;
[0062] In an alternative embodiment of the present invention, in step S2-1 of the present invention, the input on-site image data of the construction site is first uniformly scaled to a size of 3*512*512, and if the short side is insufficient, it is padded with zeros; then a feature fusion convolutional neural network is used to extract a feature image from the on-site image data of the construction site. The feature fusion convolutional neural network constructed by the present invention is different from the traditional convolutional neural network. By adding a convolutional network with a feature fusion structure, the shallow features can be continuously iterated and then used as the cross-layer input of the deep features, so that a single node can obtain more feature information of different levels. The calculation process of the feature fusion convolutional neural network is as follows:
[0063]
[0064]
[0065]
[0066] Among them, I represents the iterative output of the deep features of the nth layer of the network; x n represents the output of the nth layer of the network; Node represents the aggregation function; represents the output of the network from the mth layer to the nth layer of the intermediate calculation module; T m (x) represents the output of the feature fusion convolutional neural network of the mth layer; T n represents the output of the feature fusion convolutional neural network of the nth layer; represents the output of the network from the first layer to the nth layer of the intermediate calculation module; B represents the convolution function; represents the output of the subsequent calculation module from the first layer to the nth layer of the network. Among them, the intermediate calculation module is built by a convolutional neural network with two or more layers; the subsequent calculation module is after the intermediate calculation module and is built by a convolutional neural network with three or more layers.
[0067] The feature fusion convolutional neural network includes 33 convolutional layers, 5 pooling layers and 1 fully connected layer. The size of the convolutional kernel is 7*7, the stride is 2, the edge padding is 3, and the size of the finally generated feature map is 2048*16*16.
[0068] S2-2. Upsample the extracted feature image to generate a branch feature map;
[0069] In an alternative embodiment of the present invention, in order to ensure that the image is displayed at a higher resolution, step S2-2 of the present invention performs an upsampling operation on the feature image. Here, the transposed convolution method is used for upsampling, and the calculation method is:
[0070] o=s(i-1)+2p-k+2
[0071] Where, i represents the input size of the feature image; s represents the size of the stride; p represents the padding value of the feature image boundary; k represents the size of the convolutional kernel; o represents the size of the output branch feature map.
[0072] The transposed convolution layer includes 3 layers, the size of the transposed convolution kernel is 3*3, the stride is 1, and the edge padding is 0. The size of the generated branch feature map after upsampling is 64*128*128.
[0073] S2-3. Use a multi-branch prediction neural network to predict the branch feature map to obtain personnel key point location information, full-body lanyard detection information and scaffolding area location information.
[0074] In an alternative embodiment of the present invention, step S2-3 of the present invention plans a center point for the target according to the target annotation and trains around the center point. Considering that the personnel wearing the intelligent safety helmet may be in various postures, which may cause a certain degree of horizontal offset of the scaffolding in the image. If the traditional method based on the vertical rectangular frame is used for target bounding, more irrelevant areas will be selected, and it is possible to determine that irrelevant personnel have violated regulations. Therefore, the present invention uses a rotated target box method for the scaffolding area to locate the scaffolding area, so as to avoid detecting violations of irrelevant personnel through area restrictions, making the present invention more suitable for data detection collected by mobile devices.
[0075] The present invention uses a multi-branch prediction neural network to predict the branch feature map. The multi-branch prediction neural network includes a first branch prediction neural network, a second branch prediction neural network and a third branch prediction neural network, which are respectively used for predicting the center point offset value, predicting the length and width of the target box centered on the center point, and predicting the angle value. Among them, the third branch prediction neural network is set for the scaffolding area location, mainly for predicting the angle offset value (the angle range is [0°, 180°]) between the scaffolding area and the horizontal direction.
[0076] The first branch prediction neural network, the second branch prediction neural network, and the third branch prediction neural network constructed in the present invention all include 1 convolutional layer, 1 pooling layer, and 1 convolutional layer, where the convolutional kernel size is 3, the stride is 1, and the edge padding is 1. The first branch prediction neural network and the second branch prediction neural network output prediction maps with a size of 2*128*128, and respectively obtain the human key points, the detection and recognition results of the full-body lanyard, the center point offset value corresponding to the scaffolding area, and the length and width of the target box. Among them, there are 8 human key points, namely the left shoulder, the right shoulder, the left hip, the right hip, the left knee, the right knee, the left ankle, and the right ankle, and the recognition result of the full-body lanyard is used as the human target box. The prediction branch 3 outputs a prediction map with a size of 1*128*128, indicating the offset angle between the rotated target box of the scaffolding area and the horizontal direction.
[0077] S3. Determine whether to perform the safety belt hanging rope detection according to the relative positions of the human key point positioning information, the full-body lanyard detection information, and the scaffolding area positioning information; if so, execute step S4; otherwise, end the process;
[0078] In an optional embodiment of the present invention, after obtaining the human target box and the scaffolding rotated target box, it is necessary to judge the positional relationship between the two. That is, when the human target box is within the scaffolding rotated target box or the two intersect, it is determined that the person is within the scaffolding area. When the person is not within the scaffolding area, no subsequent safety belt hanging rope detection operation is required; when the person is within the scaffolding area and has a full-body lanyard, subsequent safety belt hanging rope detection operations can be performed; when the person is within the scaffolding area and does not have a full-body lanyard, the person is in a violation state at this time, and no subsequent safety belt hanging rope detection operation is required.
[0079] Since the construction site image data are all two-dimensional images, the two belong to the positional relationship judgment of two arbitrary quadrilaterals on a two-dimensional plane. The present invention determines whether to perform the safety belt hanging rope detection according to the relative positions of the human key point positioning information, the full-body lanyard detection information, and the scaffolding area positioning information, specifically including:
[0080] Calculate the intersection relationship between the human target box and the scaffolding rotated target box according to the human key point positioning information and the scaffolding area positioning information;
[0081] Judge whether to perform the safety belt hanging rope detection according to the intersection relationship between the human target box and the scaffolding rotated target box in combination with the full-body lanyard detection information.
[0082] The solution adopted by the present invention is to select one side of the human body target frame as the first line segment, and then select one side of the scaffolding rotation target frame as the second line segment. The image coordinate system regards the upper left corner of the image as the origin, the horizontal direction is the x-axis, and the vertical direction is the y-axis. The intersection result of the first line segment and the second line segment is calculated using the following formula:
[0083] result1=sin(θ1)×sin(θ2)
[0084] result2=sin(θ3)×sin(θ4)
[0085] Wherein, result1 represents the first intersection result; θ1 represents the angle between one endpoint of the first line segment and the two endpoints of the second line segment; θ2 represents the angle between the other endpoint of the first line segment and the two endpoints of the second line segment; result2 represents the second intersection result; θ3 represents the angle between one endpoint of the second line segment and the two endpoints of the first line segment; θ4 represents the angle between the other endpoint of the second line segment and the two endpoints of the first line segment;
[0086] When the first intersection result and the second intersection result are both less than or equal to 0, the first line segment and the second line segment intersect; otherwise, the first line segment and the second line segment do not intersect.
[0087] According to the intersection relationship between the human body target frame and the scaffolding rotating target frame, the present invention performs subsequent safety belt hanging rope detection when it is determined that the human body target frame intersects with the scaffolding rotating target frame and the whole body lacing detection information determines that the person has a whole body lacing.
[0088] S4, extracting a rectangular frame image of the lanyard area according to the key point positioning information of the personnel;
[0089] In an optional embodiment of the present invention, since the basis for construction workers to wear safety belts in violation of regulations is that there will be no safety belt hanging rope below the knee, it is necessary to first locate the hanging rope area and obtain a rectangular frame of the hanging rope area.
[0090] Step S4 of the present invention specifically includes:
[0091] The corresponding leg with a larger ordinate between the left ankle key point and the right ankle key point in the key point positioning information of the personnel is selected as the discrimination target, and the distance between the ordinate of the knee key point and the ankle key point in the leg is used as the height of the rectangular frame;
[0092] The width of the rectangular frame is set as a multiple of the larger Euclidean distance between the left shoulder key point and the left hip key point and the right shoulder key point and the right hip key point;
[0093] Use the knee keypoint and the ankle keypoint as the midpoints of the upper and lower sides of the rectangle respectively.
[0094] S5. Construct a semantic segmentation network model to identify the position of the lanyard from the rectangular frame image of the extracted lanyard area;
[0095] In an alternative embodiment of the present invention, the present invention first performs image encoding processing on the rectangular frame image of the lanyard area using the encoder of the semantic segmentation network model, and then performs a decoding operation on the encoded image using the bilinear interpolation method to obtain the identified position of the lanyard.
[0096] Step S5 of the present invention specifically includes the following sub-steps:
[0097] S5-1. Perform unified size processing on the rectangular frame image of the extracted lanyard area;
[0098] In an alternative embodiment of the present invention, in order to improve the model processing efficiency, the present invention uniformly adjusts the size of the rectangular frame image to 1000*400.
[0099] S5-2. Use a convolutional neural network to extract the convolutional features of the rectangular frame image processed in step S5-1;
[0100] In an alternative embodiment of the present invention, the present invention uses a convolutional neural network to extract the global features of the rectangular frame image; the convolutional neural network used here includes 13 convolutional layers, 4 pooling layers, and 14 activation layers, where the convolutional layers are before the activation layers, the pooling layers are after the activation layers, the convolutional kernel size is 3*3, the stride is 2, and the image padding is 0.
[0101] S5-3. Use the ASPP pyramid structure to perform feature fusion on the convolutional features extracted in step S5-2 to obtain an encoded feature map;
[0102] In an alternative embodiment of the present invention, the ASPP pyramid structure used in the present invention uses 6, 12, and 18 dilated convolutions for feature extraction of different receptive fields of the image, and then performs feature fusion processing through one pooling layer. The calculation method of the dilated convolution is as follows:
[0103]
[0104] l = w+(w - 1)*(u - 1)
[0105] where, w represents the size of the ordinary convolutional kernel; l represents the size of the dilated convolutional kernel; t represents the stride of the dilated convolution; q represents the image padding value; u represents the number of pixels with 0 in the dilated convolution; v represents the size of the feature map output after the dilated convolution.
[0106] S5-4. Use the bilinear interpolation method to decode the obtained encoded feature map to identify the position of the lanyard.
[0107] In an alternative embodiment of the present invention, the calculation method of the bilinear interpolation method adopted by the present invention is as follows:
[0108]
[0109]
[0110]
[0111] where Q 11 (x1, y1), Q 12 (x1, y2), Q 21 (x2, y1), Q 22 (x2, y2) respectively correspond to the endpoints of the lower left corner, upper left corner, upper right corner, and lower right corner of the rectangular area of the pixel to be inserted; R1(x, y1) corresponds to the point inserted in the x-axis direction between the upper left corner and the upper right corner, and R2(x, y2) corresponds to the point inserted in the x-axis direction between the lower left corner and the lower right corner; P(x, y) represents the pixel point inserted after using bilinear interpolation; f represents the output function.
[0112] S6. Determine whether there is any violation in the wearing of the safety belt according to the position of the lanyard.
[0113] In an alternative embodiment of the present invention, according to the judgment principle of safety belt violation wearing, that is, when the construction worker wears the safety belt properly, there is no safety belt lanyard below the knees of the construction worker; on the contrary, when wearing improperly, there will be a safety belt lanyard below the knees of the construction worker; thus, it is possible to determine whether there is any violation in the wearing of the safety belt according to the detected position of the lanyard.
[0114] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks the device for the specified function.
[0115] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements in the processFigure 1 one process or multiple processes and / or blocks Figure 1 the functions specified in one block or multiple blocks.
[0116] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple processes and / or the functions specified in one block or multiple blocks.
[0117] Specific embodiments are used in the present invention to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
[0118] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A seat belt violation wearing detection method combining object detection and semantic segmentation, characterized in that, It includes the following steps: S1. Collect image data of the construction site; S2. Build a target detection network model, and use the collected image data of the construction site to predict the positioning information of personnel key points, the detection information of the full-body lanyard, and the positioning information of the scaffolding area; S3. Judge whether to conduct the safety belt hanging rope detection according to the relative positions of the personnel key point positioning information, the full-body lanyard detection information, and the scaffolding area positioning information; specifically including: Calculate the intersection relationship between the human target box and the scaffolding rotation target box according to the personnel key point positioning information and the scaffolding area positioning information, specifically: Select one side of the human target box as the first line segment, select one side of the scaffolding rotation target box as the second line segment, and use the following formula to calculate the intersection result of the first line segment and the second line segment: Among them, represents the first intersection result; represents the angle between one endpoint of the first line segment and the two endpoints of the second line segment; represents the angle between the other endpoint of the first line segment and the two endpoints of the second line segment; represents the second intersection result; represents the angle between one endpoint of the second line segment and the two endpoints of the first line segment; represents the angle between the other endpoint of the second line segment and the two endpoints of the first line segment; When both the first intersection result and the second intersection result are less than or equal to 0, the first line segment and the second line segment intersect; otherwise, the first line segment and the second line segment do not intersect; Judge whether to conduct the safety belt hanging rope detection according to the intersection relationship between the human target box and the scaffolding rotation target box combined with the full-body lanyard detection information; If so, execute step S4; otherwise, end the process; S4. Extract the rectangular frame image of the hanging rope area according to the personnel key point positioning information; specifically including: Select the leg with the larger ordinate among the left ankle key point and the right ankle key point in the personnel key point positioning information as the discrimination target, and use the distance from the knee key point to the ankle key point in this leg as the height of the rectangular frame; Use a set multiple of the larger Euclidean distance between the left shoulder key point and the left hip key point and the right shoulder key point and the right hip key point as the width of the rectangular frame; Use the knee key point and the ankle key point as the midpoints of the upper and lower sides of the rectangular frame respectively; S5. Build a semantic segmentation network model to identify the hanging rope position from the extracted rectangular frame image of the hanging rope area; S6. Determine whether there is any illegal wearing behavior of the safety belt according to the hanging rope position.
2. The seat belt violation wearing detection method combining object detection and semantic segmentation according to claim 1, characterized in that, Step S2 specifically includes the following sub-steps: S2-1. Use a feature fusion convolutional neural network to extract a feature image from the image data of the construction site; S2-2. Upsample the extracted feature image to generate a branch feature map; S2-3. Use a multi-branch prediction neural network to predict the branch feature map to obtain the positioning information of personnel key points, the detection information of the full-body lanyard, and the positioning information of the scaffolding area.
3. The seat belt violation wearing detection method combining object detection and semantic segmentation according to claim 2, characterized in that, The calculation process of the feature fusion convolutional neural network in step S2-1 is: Among them, I represents the iterative output of the depth features of the n -th layer network; represents the output of the n -th layer network; Node represents the aggregation function; represents the output of the intermediate calculation module from the m -th layer to the n -th layer network; represents the output of the m -th layer feature fusion convolutional neural network; represents the output of the n -th layer feature fusion convolutional neural network; represents the output of the intermediate calculation module from the 1st layer to the n -th layer network; B represents the convolution function; represents the output of the subsequent calculation module from the 1st layer to the n -th layer network.
4. The seat belt violation wearing detection method combining object detection and semantic segmentation according to claim 2, characterized in that, The calculation method of upsampling the extracted feature image in step S2-2 is: Among them, i represents the input size of the feature image; s represents the size of the stride; p represents the padding value of the feature image boundary; k represents the size of the convolutional kernel; o represents the size of the output branch feature map.
5. The seat belt violation wearing detection method combining object detection and semantic segmentation according to claim 1, characterized in that, Step S5 specifically includes the following sub-steps: S5-1. Process the extracted rectangular frame image of the hanging rope area to a unified size; S5-2. Use a convolutional neural network to extract the convolutional features of the rectangular frame image processed in step S5-1; S5-3. Use an ASPP pyramid structure to fuse the convolutional features extracted in step S5-2 to obtain an encoded feature map; S5-4. Use the bilinear interpolation method to decode the obtained encoded feature map to identify the hanging rope position.
6. The seat belt violation wearing detection method combining object detection and semantic segmentation according to claim 5, wherein, In step S5-3, the ASPP pyramid structure uses dilated convolutions with dilation rates of 6, 12, and 18 to extract features of different receptive fields of the image, and then performs feature fusion processing through a pooling layer.
7. The seat belt violation wearing detection method combining object detection and semantic segmentation according to claim 6, wherein, The calculation method of the dilated convolution is as follows: Among them, w represents the size of the ordinary convolution kernel; l represents the size of the dilated convolution kernel; t represents the stride of the dilated convolution; q represents the image padding value; u represents the number of pixels with value 0 in the dilated convolution; v represents the size of the feature map output after the dilated convolution.
Citation Information
Patent Citations
Deep learning-based violation detection method for low hanging and high use of safety belt
CN112215138A
Substation pointer instrument detection method based on deep learning
CN114463558A