A worker positioning method based on double-precision feature enhancement

Through double-precision features, the network and progressive anchoring loss function are enhanced, and the problems of high-cost and complex hardware in the existing technology are solved, low-cost worker positioning is achieved, and the safety of the construction site is improved.

CN115205754BActive Publication Date: 2025-07-18FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210868674.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-22
Publication Date
2025-07-18
Estimated Expiration
2042-07-22

AI Technical Summary

Technical Problem

In the construction site safety accident, binocular visual ranging method that relies on high-cost and complex hardware is difficult to effectively locate the worker's position, resulting in expensive equipment and complex structure, which cannot be widely used at the construction site.

Method used

The worker positioning method based on double-precision feature enhancement is adopted, and data is collected through electronic cameras, a double-precision worker detection model is built, and the workers' relative position is calculated using the similar triangle principle, and the feature enhancement module FEM and the progressive anchor loss function PAL are combined to realize worker positioning.

Benefits of technology

It reduces hardware requirements, simplifies structure, reduces costs, realizes accurate positioning of workers' locations, and provides safety guarantees at the construction site.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115205754B_ABST
    Figure CN115205754B_ABST
Patent Text Reader

Abstract

The present invention provides a worker positioning method based on double-precision feature enhancement, including the following steps: Step S1: Collect construction data of construction site workers through an electronic camera, randomly sample video frames into images, label the workers in the images through an image annotation tool labelimg, and generate a worker detection data set; Step S2: Build a double-precision worker detection model, determine model parameters and the loss function of the neural network, and optimize the model performance to the best; Step S3: Use the principle of similar triangles to calculate the relative position of the worker from the camera to achieve worker positioning. Applying this technical solution can calculate the position where the worker is located through the length and width of the predicted box of the worker detected by the double-precision feature enhancement network in the image and the camera parameter matrix. Compared with binocular vision ranging, the monocular vision ranging method has the characteristics of low hardware requirements, simple structure, and low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology for unsupervised machine learning, and particularly to a worker positioning method based on double-precision feature enhancement. Background Art

[0002] The main safety accidents that may occur during the construction process include cave-ins, water and mud gushing, gas explosions, and landslides. These accidents have caused great losses to people's lives and property and severely restricted the progress of transportation construction. Estimating depth information from images has been widely used in object detection. Relying on the depth information obtained by cameras can timely and accurately determine the orientation and distance of workers from the cameras, providing a golden time for rescuing trapped workers. In the method of combining two-dimensional images with depth information sources, the camera obtains the two-dimensional coordinate values of the imaging plane, and at the same time uses information sources such as structured light ranging, laser ranging, ultrasonic ranging, and position sensors to obtain depth information. However, the equipment is expensive, and many construction sites cannot be equipped with facilities due to various factors. In the method of calculating three-dimensional information only using two-dimensional images based on the principle of optical imaging, binocular ranging has higher requirements for hardware, is complex in structure, and has a higher cost. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a worker positioning method based on double-precision feature enhancement, which calculates the position of the worker by the length and width of the predicted box of the worker detected by the double-precision feature enhancement network in the image and the camera parameter matrix. Compared with binocular vision ranging, the monocular vision ranging method has the characteristics of low hardware requirements, simple structure, and low cost.

[0004] To achieve the above purpose, the present invention adopts the following technical solutions: A worker positioning method based on double-precision feature enhancement, including the following steps:

[0005] Step S1: Collect construction data of workers at the construction site through an electronic camera, randomly sample video frames into images, and labelimg annotate the workers in the images through an image annotation tool to generate a worker detection data set;

[0006] Step S2: Build a double-precision worker detection model, determine the model parameters and the loss function of the neural network, and optimize the model performance to the best;

[0007] Step S3: Use the principle of similar triangles to calculate the relative position of the worker from the camera to achieve worker positioning.

[0008] In a preferred embodiment, the specific implementation of the step S1 is as follows:

[0009] Step S11: Collect construction data of workers at the construction site through an electronic camera, randomly sample video frames to generate 30,000 construction site images as the data set;

[0010] Step S12: Use the image annotation tool labelimg to label the workers in the image, obtain the position of the worker detection frame in each image, export a txt file storing the worker labels, and divide the data set into three data sets with different difficulty levels;

[0011] Step S13: randomly extract the data set into training set and test set in proportion, generate training set and test set txt files consisting of image addresses respectively, and count all worker labels of each image into a label file named after the image address according to the image address in txt.

[0012] In a preferred embodiment, step S2 is specifically implemented as follows:

[0013] Step S21: Use VGG16 as the backbone network structure, select conv3_3, conv4_3, conv5_3, conv_fc7, conv6_2 and conv7_2 as the first shot detection layer, generate 6 original feature maps, and name them of1, of2, of3, of4, of5, of6 respectively;

[0014] Step S22: using a feature enhancement module FEM, the original feature maps are converted into six enhanced feature maps, which are named ef1, ef2, ef3, ef4, ef5 and ef6 respectively;

[0015] FEM uses the original pixel unit oc of the upper layer (i,j,l) and the non-local pixel unit nc of the current layer (i-ε,j-ε,l) 、nc (i-ε,j,l) 、nc (i-ε,j+ε,l) 、nc (i,j-ε,l) 、nc (i,j+ε,l) 、nc (i+ε,j-ε,l) 、nc (i+ε,j,l) and nc (i+ε,j+δ,l) These two different dimensional information are used to enhance the original pixel unit; that is, in the feature enhancement module FEM, the feature map unit of the current layer interacts with the neighboring units of the current feature map and the neighboring units in the upper feature map;

[0016] Step S23: design a progressive anchor loss function PAL;

[0017] A set of smaller anchor sizes is given in the first shot layer to assist supervision, and a larger anchor size is used in the second shot; progressive anchor sizes are designed in different layers and shots, thereby obtaining the multi-task loss function of the first shot respectively. and the multi-task loss function for the second lens

[0018]

[0019]

[0020] Among them, sa i represents the i-th smaller anchor point in the first shot layer, and N conf represents the number of positive and negative anchors, and N loc represents the number of positive anchors; L conf is the softmax loss on the two classes of face and background; L loc is the smooth L1 loss between the predicted box t i and the ground truth box g i with the anchor point being a i ; takes values between 0 and 1. When is 1, a i is a positive anchor, and at the same time the localization loss is activated; β is the balancing weight, and p i is the weight;

[0021] The two losses are weighted into a complete progressive anchor loss function L PAL : The anchor size of the first shot is half of that of the second shot, and λ is the weight;

[0022] L PAL = L FSL (sa)+λL SSL (a).

[0023] In a preferred embodiment, in step S22, first, a 1×1 convolutional kernel is used to normalize the feature map; then, the upper-layer feature map is upsampled to generate the current feature map element-wise; next, the feature map is divided into three sub-networks containing different numbers of dilated convolutional layers; finally, they are connected and integrated into an enhanced feature map; thus, the enhanced pixel unit is defined as follows:

[0024] ec (i,j,l) = f concat (f dilation (nc (i,j,l) ))

[0025] nc (i,j,l) = f prod (oc (i,j,l) , f up (oc (i,j,l+1) ))

[0026] Among them, c (i,j,l)Denote the pixel unit located at the l-th layer feature map with coordinates (i, j), f represents a series of basic dilated convolutions, element-wise products, upsamplings, or concatenation operations, ec (i,j,l) Denote the enhanced l-th layer feature map, nc (i,j,l) Denote the non-local l-th layer feature, oc (i,j,l) Denote the original pixel unit feature of the upper layer.

[0027] In a preferred embodiment, the specific implementation of step S3 is as follows:

[0028] Step S31: Offline part;

[0029] First, calibrate the camera using the Zhang-Zhengyou calibration algorithm or the Tsai two-step method. First, use the auxiliary photographing program and a 140-degree wide-angle camera to photograph a relatively dense checkerboard printed on A4 paper, and then use the improved TPS algorithm for distortion correction to obtain the parameter matrix; Camera calibration is the link between camera measurement and real three-dimensional world measurement, and is a necessary process for obtaining three-dimensional stereo information from a planar image;

[0030] Secondly, perform inverse perspective projection according to the optical imaging principle; Take a pedestrian image in the tunnel environment, perform distortion correction through the interpolation map obtained by camera calibration, set the area, and then calculate the perspective transformation matrix based on the relationship between the detected inner corner points of the checkerboard and the expected corner points;

[0031] Step S32: Online part;

[0032] First, perform image acquisition; Set the size of the acquired image, delay the automatic exposure of the camera, and detect whether the image acquisition is successful;

[0033] Secondly, perform distortion correction; Load the interpolation map generated in the offline part, and detect whether the loading is successful. Use the bilinear interpolation algorithm to perform interpolation correction on the distorted image;

[0034] Then, perform perspective transformation; Use the transformation matrix in the offline part to convert the undistorted image into a top view; Perform the conversion between the three-dimensional space points in the world and the corresponding points of the pixel coordinates in the phase plane;

[0035] Then, image preprocessing and object detection are carried out; for any input tunnel image, the illumination compensation image is obtained by using the block-based local homomorphic filtering (BLHF). The prior boxes output by the low-illumination feature enhancement module and the person object detection module are matched with the ground truth to find the prior box with the largest overlap degree. The color of the safety helmet is recognized by using the Lab color feature, the centroid method is used to locate the center of the safety helmet, the contour is recognized by comparing the use of the K-means clustering algorithm and the Canny edge detection algorithm, and the center and radius of its target area are determined by the Hough gradient circle transformation detection algorithm;

[0036] Finally, according to the geometric similarity method of a single camera and the principle of perspective projection, the distance of the target object is measured, the angles of the detected corner pairs are calculated, sorted according to the angle values, 3-6 middle point pairs are selected for random sample consensus (RANSAC) detection, and the point with the smallest error is selected to calculate the rotation and translation matrix; the large map is rotated and translated, and then the current frame is fused into the map by using the direct fusion method; the object detection box is output to the system in real time, and the distance information is marked in real time.

[0037] Compared with the prior art, the present invention has the following beneficial effects: by using the feature enhancement module (FEM), the single-lens detector is extended to a double-lens detector, and the progressive anchor loss function (PAL) is used. Through the improved anchor matching (IAM), the features are effectively enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a flowchart of a preferred embodiment of the present invention.

[0039] Figure 2 is a schematic diagram of the double-precision feature enhancement network structure used in the preferred embodiment of the present invention.

[0040] Figure 3 is a schematic diagram of the feature enhancement module (FEM) structure used in the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0042] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0043] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0044] A worker positioning algorithm based on double-precision feature enhancement, referring to Figures 1 to 3 , includes the following steps:

[0045] Step S1: Collect construction data of workers on the construction site through an electronic camera, randomly sample video frames into images, and use an image annotation tool labelimg to annotate the workers in the images to generate a worker detection data set;

[0046] Step S2: Build a double-precision worker detection model, determine model parameters and the loss function of the neural network, and optimize the model performance to the best;

[0047] Step S3: Use the principle of similar triangles to calculate the relative position of the worker from the camera to achieve worker positioning.

[0048] The specific implementation of step S1 is as follows:

[0049] Step S11: Collect construction data of workers on the construction site through an electronic camera, and randomly sample video frames to generate 30,000 construction site images as a data set;

[0050] Step S12: Manually annotate the workers in the images through the image annotation tool labelimg to obtain the positions of the worker detection frames in each image, and export a txt file storing worker labels. The worker labels include: the upper left coordinates (x, y) of the prior box annotated in the image where the worker is located, the width w of the prior box, the height h, and the worker type label. The worker type label consists of three 01 binary labels, respectively representing workers in three cases: workers with low detection difficulty, workers with medium detection difficulty, and workers with the highest detection difficulty;

[0051] Step S13: Divide the 30,000 datasets into three categories according to the detection difficulty level in the images. Among them, the Easy dataset represents that most workers in the dataset are relatively easy to detect, and the corresponding worker type label for the dataset is [1, 0, 0]; the Medium dataset represents that most workers in the dataset are of medium detection difficulty, and the corresponding worker type label for the dataset is [0, 1, 0]; the Hard dataset represents that most workers in the dataset are of great detection difficulty, and the corresponding worker type label for the dataset is [0, 0, 1]. The three datasets are randomly sampled as the training set and the test set at a ratio of 3:1 respectively, and training sets and test sets are generated as txt files composed of image addresses. According to the image addresses in the txt, the labels of all workers in each image are counted into a label file named after the image address.

[0052] The specific implementation of step S2 is as follows:

[0053] Step S21: Use VGG16 as the backbone network structure, select conv3_3, conv4_3, conv5_3, conv_fc7, conv6_2, and conv7_2 as the first shot detection layer, generate 6 original feature maps, and name them of1, of2, of3, of4, of5, and of6 respectively.

[0054] Step S22: Use the Feature EnhanceModule (FEM) to convert these original feature maps into 6 enhanced feature maps, and name them ef1, ef2, ef3, ef4, ef5, and ef6 respectively.

[0055] The feature enhancement module can enhance the original feature maps, making them more discriminable and robust, extend the single-shot detector to a double-shot detector, and combine multi-layer extended convolutional layers to enhance the semantics of the features. It does not ignore the context relationship between anchor points, simply referred to as FEM. FEM uses the original pixel unit oc of the upper layer (i,j,l) and the non-local pixel unit nc of the current layer (i-ε,j-ε,l) 、nc (i-ε,j,l) 、nc (i-ε,j+ε,l) 、nc (i,j-ε,l) 、nc (i,j+ε,l) 、nc (i+ε,j-ε,l) 、nc (i+ε,j,l) 、nc (i+ε,j+ε,l) These two different-dimensional information to enhance the original pixel unit. That is, in the feature enhancement module, the feature map units of the current layer interact with the neighbor units of the current feature map and the neighbor units in the upper layer feature map.

[0056] Here, a 1×1 convolutional kernel is first used to normalize the feature map. Then, the upper-layer feature map is upsampled to generate the current feature map element-wise. Next, the feature map is divided into three sub-networks containing different numbers of dilated convolutional layers. Finally, they are connected and integrated into an enhanced feature map. Thus, the enhanced pixel unit is defined as follows:

[0057] ec (i,j,l) = f concat (f dilation (nc (i,j,l) ))

[0058] nc (i,j,l) = f prod (oc (i,j,l) , f up (oc (i,j,l+1) ))

[0059] Among them, c (i,j,l) represents the pixel unit located at the coordinate (i, j) of the feature map at the l-th layer, f represents a series of basic dilation, element-wise product (element generation), upsampling, or concatenation operations, ec (i,j,l) represents the enhanced feature map at the l-th layer, nc (i,j,l) represents the non-local feature at the l-th layer, and oc (i,j,l) represents the original pixel unit feature of the upper layer.

[0060] Step S23: Design a progressive anchor loss function PAL.

[0061] Considering the progressive learning ability of the feature map at different levels and different shots, compared with the enhanced feature map at the same level, the original feature map has less semantic information for classification but more high-resolution position information for detection. That is, the original low-level feature map is more suitable for detecting and classifying smaller human face features. Therefore, a set of smaller anchor sizes are given in the first shot layer to assist supervision, and larger anchor sizes are used in the second shot. Designing progressive anchor sizes in different layers and different shots can effectively enhance the features, and thus obtain the multi-task loss function of the first shot and the multi-task loss function of the second shot

[0062]

[0063]

[0064] Among them, sa i represents the i-th smaller anchor in the first shot layer, Nconf Represents the number of positive and negative anchors, N loc Represents the number of positive anchors; L conf Is the softmax loss on two classes of face and background; L loc Is the predicted box t i And the ground truth box g i The smooth L1 loss between the parameterizations of, with the anchor point being a i ; Takes values between 0 and 1. When Is 1, a i Is a positive anchor, and at the same time the localization loss is activated; β is the balancing weight, p i Is the weight;

[0065] The two losses are weighted into a complete progressive anchor loss function L PAL : The anchor size of the first shot is half of that of the second shot, and λ is the weight;

[0066] L PAL = L FSL (sa)+λL SSL (a).

[0067] The specific implementation of step S3 is as follows:

[0068] Step S31: Offline part.

[0069] First, calibrate the camera using Zhang Zhengyou calibration algorithm or Tsai two-step method. First, use the auxiliary photographing program and a 140-degree wide-angle camera to photograph a relatively dense checkerboard printed on A4 paper, and then use the improved TPS algorithm for distortion correction to obtain the parameter matrix. Camera calibration is the link between camera measurement and real three-dimensional world measurement, and is a necessary process to obtain three-dimensional stereo information from plane images.

[0070] Secondly, perform inverse perspective projection according to the optical imaging principle. Take a pedestrian image in a tunnel environment, perform distortion correction through the interpolation map obtained by camera calibration, set the physical pixel ratio to 25cm / 56.74775 pixels (modifiable), set the area, and then calculate the perspective transformation matrix based on the relationship between the detected inner corner points of the checkerboard and the expected corner points.

[0071] Step S32, Online part.

[0072] First, perform image acquisition. Set the size of the acquired image, delay the automatic exposure of the camera, and detect whether the image acquisition is successful;

[0073] Secondly, perform distortion correction. Load the interpolation map generated in the offline part, and detect whether the loading is successful. Use the bilinear interpolation algorithm to perform interpolation correction on the distorted image.

[0074] Then perform a perspective transformation. Using the transformation matrix of the offline part, convert the undistorted image into a top view. Perform the conversion between the three-dimensional spatial points in the world and the corresponding pixel coordinates in the phase plane.

[0075] Then perform image preprocessing and target detection. For any input tunnel image, use block-based local homomorphic filtering (BLHF) to obtain the illumination compensation image. Match the prior boxes output by the low-light feature enhancement module and the person target detection module in Section 3.1.1 with the ground truth to find the prior box with the largest overlap. Use the Lab color feature to identify the color of the safety helmet, use the centroid method to locate the center of the safety helmet, compare the use of the K-means clustering algorithm and the Canny edge detection algorithm to identify the contour, and determine the center and radius of its target area through the Hough gradient circle transformation detection algorithm.

[0076] Finally, according to the geometric similarity method of a single camera and the principle of perspective projection, measure the distance of the target object, calculate the angles for the detected corner pairs, sort according to the angle values, select 3-6 middle point pairs, perform random sample consensus (RANSAC) detection, select the point with the smallest error, and calculate the rotation and translation matrix. Rotate and translate the large map, and then use the direct fusion method to fuse the current frame into the map. Output the target detection box to the system in real time and mark the distance information in real time.

[0077] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention in terms of the functions and effects produced belong to the protection scope of the present invention.

Claims

1. A worker positioning method based on double-precision feature enhancement, characterized in that It includes the following steps: Step S1: Collect construction data of workers at the construction site through an electronic camera, randomly sample video frames into images, label the workers in the images using the image annotation tool labelimg, and generate a worker detection dataset; Step S2: Build a double-precision worker detection model, determine model parameters and the loss function of the neural network, and optimize the model performance to the best; Step S3: Use the principle of similar triangles to calculate the relative position of the worker from the camera to achieve worker positioning; The specific implementation of step S2 is as follows: Step S21: Use VGG16 as the backbone network structure, select conv3_3, conv4_3, conv5_3, conv_fc7, conv6_2, and conv7_2 as the first lens detection layer to generate 6 original feature maps, and name them of1, of2, of3, of4, of5, and of6 respectively; Step S22: Use the Feature Enhance Module (FEM) to convert these original feature maps into 6 enhanced feature maps, named ef1, ef2, ef3, ef4, ef5, and ef6 respectively; The FEM uses the original pixel units oc in the upper layer (i,j,l) and the non-local pixel units nc in the current layer (i-ε,j-ε,l) 、nc (i-ε,j,l) 、nc (i-ε,j+ε,l) 、nc (i,j-ε,l) 、nc (i,j+ε,l) 、nc (i+ε,j-ε,l) 、nc (i+ε,j,l) and nc (i+ε,j+ε,l) to enhance the original pixel units with these two different-dimensional information; that is, in the feature enhancement module FEM, the feature map units in the current layer interact with the neighbor units in the current feature map and the neighbor units in the upper feature map; Step S23: Design a Progressive Anchor Loss function (PAL); A set of small anchor sizes is given in the first shot layer to assist in supervision, and large anchor sizes are used in the second shot; progressive anchor sizes are designed in different layers and different shots, and the multi-task loss functions of the first shot and the second shot are obtained respectively and the multi-task loss function of the second shot Among them, sa i represents the i-th small anchor point in the first shooting layer, N conf represents the number of positive and negative anchors, N loc represents the number of positive anchors; L conf is the softmax loss on the two classes of face and background; L loc is the smooth L1 loss between the predicted box t i and the parameterization of the ground truth box g i , with the anchor point being a i ; takes values between 0 and 1. When is 1, a i is a positive anchor, and at the same time the localization loss is activated; β is the balance weight, p i is the weight; Two losses are weighted to form a complete progressive anchor loss function $L$ PAL : The anchor size of the first shot is half of that of the second shot, and $\lambda$ is the weight; L PAL = L FSL (sa) + λL SSL (a).

2. The worker positioning method based on double-precision feature enhancement according to claim 1, wherein The specific implementation of step S1 is as follows: Step S11: Collect construction data of workers at the construction site through an electronic camera, randomly sample video frames to generate 30,000 construction site images as the dataset; Step S12: Use the image annotation tool labelimg to label the workers in the images, obtain the positions of the worker detection boxes in each image, export the txt file storing the worker labels, and divide the dataset into three datasets with different levels of difficulty; Step S13: Randomly extract the dataset into a training set and a test set according to a ratio, generate txt files of the training set and the test set composed of image addresses respectively, and count all the worker labels of each image into a label file named after the image address according to the image addresses in the txt; 3. The worker positioning method based on double-precision feature enhancement according to claim 1, characterized in that, In step S22, first use a 1×1 convolution kernel to normalize the feature map; then, upsample the upper-layer feature map to generate the current feature map element-wise; then, divide the feature map into three sub-networks containing different numbers of dilated convolution layers; finally, connect and integrate them into an enhanced feature map; thus, the enhanced pixel unit is defined as follows: ec (i,j,l) = f concat (f dilation (nc (u,j,l) )) nc (i,j,l) = f prod (oc (i,j,l) , f up (oc (i,j,l+1) ) Among them, c (i,j,l) represents a pixel unit located in the l-th layer feature map with coordinates (i, j). f represents a series of basic dilated convolutions, element-wise products, upsamplings, or concatenation operations. ec (i,j,l) represents the enhanced l-th layer feature map. nc (i,j,l) represents the non-local l-th layer feature. oc (i,j,k) represents the original pixel unit feature of the upper layer.

4. A worker positioning method based on double-precision feature enhancement according to claim 1, characterized in that The specific implementation of step S3 is as follows: Step S31: Offline part; First, calibrate the camera using the Zhang-Zhengyou calibration algorithm or the Tsai two-step method. First, use the auxiliary photographing program and a 140-degree wide-angle camera to photograph a relatively dense checkerboard printed on A4 paper, and then use the improved TPS algorithm for distortion correction to obtain the parameter matrix; Camera calibration is the link between camera measurement and real three-dimensional world measurement, and is a necessary process to obtain three-dimensional stereo information from a planar image; Secondly, perform inverse perspective projection according to the optical imaging principle; capture a pedestrian image in the tunnel environment, perform distortion correction through the interpolation map obtained by camera calibration, set the region, and then calculate the perspective transformation matrix based on the relationship between the detected inner corner points of the checkerboard and the expected corner points. Step S32: Online part; First, perform image acquisition; set the size of the captured image, delay the automatic exposure of the camera, and detect whether the image is successfully acquired. Secondly, perform distortion correction; load the interpolation map generated in the offline part and detect whether the loading is successful, and use the bilinear interpolation algorithm to perform interpolation correction on the distorted image. Then, perform perspective transformation; use the transformation matrix in the offline part to convert the undistorted image into a top view; perform the conversion between the three-dimensional space points in the world and the corresponding pixel coordinates in the phase plane. Then, perform image preprocessing and target detection; for any input tunnel image, use the block-based local homomorphic filtering (BLHF) to obtain the illumination compensation image, match the prior boxes output by the low-illumination feature enhancement module and the person target detection module with the ground truth, and find the prior box with the largest overlap; use the Lab color feature to identify the color of the safety helmet, use the centroid method to locate the center of the safety helmet, compare the use of the K-means clustering algorithm and the Canny edge detection algorithm to identify the contour, and determine the center and radius of its target area through the Hough gradient circle transformation detection algorithm; finally, according to the geometric similarity method of a single camera and the perspective projection principle, measure the distance of the target object, calculate the angle for the detected corner point pairs, sort according to the angle values, select 3-6 middle point pairs, perform the random sample consensus (RANSAC) detection, select the point with the smallest error to calculate the rotation and translation matrix; rotate and translate the large map, and then use the direct fusion method to fuse the current frame into the map. Output the target detection box to the system in real time and mark the distance information in real time.