Working site object measurement system and method based on three-dimensional point cloud distance measurement
By installing high-precision lidar and deep learning models at the work site to process point cloud data, the precise measurement of the distance between objects is achieved, and the problem of inability to effectively prevent safety accidents caused by too close distances in the prior art is solved, and the safety of the work site is improved.
Patent Information
- Application Number
- CN202510158501.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art video surveillance system is difficult to accurately measure the distance between objects at the operation site, and cannot effectively prevent safety accidents caused by too close distances.
The operation site object measurement system based on three-dimensional point cloud distance measurement is adopted. The operation site is scanned by installing high-precision lidar, point cloud data is acquired and deep learning model processing is performed, the three-dimensional position information of the object is extracted and the relative distance between the objects is calculated.
It realizes accurate measurement of the distance between objects on the work site, can effectively prevent safety accidents caused by too close distance, and improves the safety of the work site.
Smart Images

Figure CN120195685A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of on-site safety supervision, and particularly to an on-site object measurement system and method based on three-dimensional point cloud ranging. Background Art
[0002] In on-site operations such as power construction, due to the complex environment and various safety risks, such as the distance between personnel and mechanical equipment being too close, and personnel entering dangerous areas, traditional safety supervision means are difficult to achieve comprehensive and real-time monitoring and early warning.
[0003] The prior art CN112989110A includes a camera group, an intelligent computer processing cluster, and a database warehouse; the camera group is used to collect images in the spatial scene and transmit them to the intelligent computer processing cluster; the intelligent computer processing cluster is used to annotate the images collected by the camera group and transmit the annotated images to the database warehouse for storage; the database warehouse is used to store the data sent by the intelligent computer processing cluster and classify the data.
[0004] Although the video monitoring system of the prior art can achieve a certain degree of visual supervision, it lacks the ability to accurately measure the distance between objects and cannot effectively prevent safety accidents caused by too close distances. Summary of the Invention
[0005] The purpose of the present invention is to provide an on-site object measurement system and method based on three-dimensional point cloud ranging, which solves the problem that although the video monitoring system in the prior art can achieve a certain degree of visual supervision, it lacks the ability to accurately measure the distance between objects and cannot effectively prevent safety accidents caused by too close distances.
[0006] To achieve the above purpose, the present invention provides an on-site object measurement system and method based on three-dimensional point cloud ranging, including the following steps:
[0007] Install a high-precision lidar at the on-site operation to scan the on-site operation, obtain point cloud data containing the three-dimensional coordinate information of objects, and preprocess the collected point cloud data;
[0008] Input the preprocessed point cloud data into a deep learning model. The model extracts features from the point cloud data through a convolutional layer to generate feature vectors, and uses a maximum saliency pooling layer to process the feature map to retain key feature information;
[0009] Use the deep learning model to identify and classify objects in the point cloud data, and determine the position information of each object in three-dimensional space, including coordinates and bounding boxes;
[0010] Calculate the relative distance between objects based on the position information of the objects. For multiple objects, calculate the distance from them to the key area;
[0011] Transmit the calculated ranging result to the safety supervision terminal in real time through the wireless communication module. The safety supervision terminal displays the distance information between objects in real time. When the distance between objects is lower than the preset safety threshold, the terminal issues an alarm to remind the on-site personnel to take safety measures.
[0012] Among them, in "preprocessing the collected point cloud data", the following steps are included:
[0013] Use voxel filtering to downsample the point cloud, divide the point cloud into three-dimensional voxel grids, and retain the mean point within each voxel in each voxel;
[0014] Use the moving least squares method to perform weighted least squares fitting on each point and its neighborhood points through local polynomial fitting;
[0015] Simulate the scene from different perspectives by randomly rotating the point cloud data, simulate the point cloud data at different distances by random scaling, and simulate the noise interference in the actual scene by injecting Gaussian noise;
[0016] Normalize the coordinate values of the point cloud to a unified scale range, and achieve coordinate normalization by calculating the relative position of each point's coordinate to the center of the point cloud.
[0017] Among them, in "the model extracts features from the point cloud data through convolutional layers to generate feature vectors, and uses the saliency maximum pooling layer to process the feature maps to retain key feature information", the method further includes the following steps:
[0018] Perform a convolution operation on the point cloud data through a convolution kernel to extract local features in the data. Each convolution kernel slides on the point cloud data to calculate the feature response of the local area and generate multiple feature maps;
[0019] Introduce the saliency maximum pooling layer. The saliency maximum pooling layer selects the maximum response value in the local area of each feature map, retains the most significant feature information, and removes redundant details at the same time;
[0020] The fully connected layer flattens the pooled feature maps into one-dimensional vectors, and performs weighted combination of the features by learning weights to generate the final feature vector. The global pooling operation directly generates a fixed-length feature vector by aggregating all elements of the feature map.
[0021] Among them, in "using a deep learning model to identify and classify objects in the point cloud data, and determine the position information of each object in the three-dimensional space, including coordinates and bounding boxes", the method includes:
[0022] Input the pre - processed and feature - extracted point cloud data into a deep - learning model. Use an instance segmentation algorithm to process the point cloud data, accurately label each object instance in the point cloud, and generate an independent segmentation mask for each instance;
[0023] Through a convolutional neural network and a point - cloud - specific network, the model learns the geometric features and semantic information of the objects, and classifies each object instance into predefined categories;
[0024] By combining the instance segmentation results and the three - dimensional coordinate information in the point cloud data, calculate the center coordinates, size, and orientation of each object;
[0025] Generate its bounding box in three - dimensional space according to the center coordinates and size information of each object.
[0026] Among them, in "By combining the instance segmentation results and the three - dimensional coordinate information in the point cloud data, calculate the center coordinates, size, and orientation of each object", the method includes the following steps:
[0027] Extract the point cloud data corresponding to each object instance from the segmentation mask generated by the instance segmentation algorithm, and perform clustering analysis of the extracted point cloud data using a density - based clustering algorithm;
[0028] For each clustering result, calculate the centroid by taking the average of the coordinates of all points in the cluster, which is used as the center coordinates of the object;
[0029] Determine the bounding box of the object according to the distribution range of the point cloud data. Calculate the maximum and minimum values of the point cloud in the x, y, and z directions to determine the length, width, and height of the bounding box. At the same time, calculate the direction vector of the bounding box according to the distribution direction of the point cloud.
[0030] Among them, in "For each clustering result, calculate the centroid by taking the average of the coordinates of all points in the cluster, which is used as the center coordinates of the object", the method includes:
[0031] The calculation method of the centroid is to take the average of the coordinates of all points in the cluster, and the centroid reflects the geometric center position of the object in three - dimensional space and is a key parameter of the object's position information.
[0032] A job - site object measurement system based on three - dimensional point - cloud ranging, used to execute the job - site object measurement method based on three - dimensional point - cloud ranging, includes a lidar scanning module, a deep - learning processing module, a data - transmission module, and a safety supervision terminal. The deep - learning processing module receives the signal from the lidar scanning module, the data - transmission module receives the signal from the deep - learning processing module, and the safety supervision terminal receives the signal from the data - transmission module;
[0033] The lidar scanning module is used to perform three-dimensional scanning on the operation site to obtain point cloud data containing the three-dimensional coordinate information of objects;
[0034] The deep learning processing module is used to extract features and match the point cloud data, and calculate the exact distance between objects;
[0035] The data transmission module is used to transmit the ranging result to the safety supervision terminal in real time;
[0036] The safety supervision terminal is used to receive and display the ranging result, and issue an alarm when the distance does not meet the safety requirements.
[0037] An object measurement system and method for an operation site based on three-dimensional point cloud ranging according to the present invention includes installing a high-precision lidar at the operation site, scanning the operation site to obtain point cloud data containing the three-dimensional coordinate information of objects, and preprocessing the collected point cloud data; inputting the preprocessed point cloud data into a deep learning model, the model extracts features from the point cloud data through a convolutional layer to generate feature vectors, uses a saliency maximum pooling layer to process the feature map, and retains key feature information; uses the deep learning model to identify and classify objects in the point cloud data, and determine the position information of each object in three-dimensional space, including coordinates and bounding boxes; calculates the relative distance between objects according to the position information of the objects, and for multiple objects, calculates the distance between them and the key area; transmits the calculated ranging result to the safety supervision terminal in real time through a wireless communication module, the safety supervision terminal displays the distance information between objects in real time, and when the distance between objects is lower than a preset safety threshold, the terminal issues an alarm to remind the on-site personnel to take safety measures. In this way, by determining the position information of each object in three-dimensional space and obtaining the distance data between objects, safety accidents caused by too close distances can be effectively prevented. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.
[0039] Figure 1 is a flowchart of the object measurement method for the operation site based on three-dimensional point cloud ranging according to the present invention.
[0040] Figure 2 is a flowchart of preprocessing the collected point cloud data according to the present invention.
[0041] Figure 3 is a flowchart of the model extracting features from the point cloud data through a convolutional layer to generate feature vectors, using a saliency maximum pooling layer to process the feature map, and retaining key feature information according to the present invention.
[0042] Figure 4 This is a flowchart of the present invention for identifying and classifying objects in point cloud data using a deep learning model, and determining the position information of each object in three-dimensional space, including coordinates and bounding boxes.
[0043] Figure 5 This is a flowchart of the present invention for calculating the center coordinates, dimensions, and orientation of each object by combining the instance segmentation results and the three-dimensional coordinate information in the point cloud data.
[0044] Figure 6 This is a system diagram of the object measurement system at the operation site based on three-dimensional point cloud ranging of the present invention.
[0045] In the figure: lidar scanning module 101, deep learning processing module 102, data transmission module 103, and safety supervision terminal 104. Specific implementation manners
[0046] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.
[0047] Please refer to Figures 1 to 5 , a method for measuring objects at an operation site based on three-dimensional point cloud ranging, including the following steps:
[0048] S1: Install a high-precision lidar at the operation site, scan the operation site to obtain point cloud data containing the three-dimensional coordinate information of the objects, and preprocess the collected point cloud data;
[0049] S11: Downsample the point cloud using voxel filtering, divide the point cloud into three-dimensional voxel grids, and retain the mean point within each voxel in each voxel;
[0050] S12: Use the moving least squares method to perform weighted least squares fitting on each point and its neighboring points through local polynomial fitting;
[0051] S13: Simulate the scenes under different perspectives by randomly rotating the point cloud data, simulate the point cloud data at different distances by random scaling, and simulate the noise interference in the actual scene by injecting Gaussian noise;
[0052] S14: Normalize the coordinate values of the point cloud to a unified scale range, and achieve coordinate normalization by calculating the relative position of each point's coordinate to the center of the point cloud;
[0053] In this embodiment, the purpose of data filtering is to remove abnormal points and noise in the point cloud data, retain effective point cloud information, and improve the purity of the data. Method: Voxel Grid Filter: Principle: The point cloud data is divided into voxel grids (3D grids) in space, and the points within each voxel are downsampled to a representative point (usually the mean point or median point within the voxel). Function: Reduce the density of the point cloud data, reduce the data volume, and at the same time remove some noise points. Parameter: Voxel size (select an appropriate voxel size according to the application scenario and computing resources). Statistical Outlier Removal: Principle: Calculate the average distance of the neighboring points of each point, and regard the points with a distance much greater than the average value as noise points and remove them. Function: Effectively remove isolated noise points and is suitable for dealing with sparse noise. Parameters: Number of neighboring points (k) and standard deviation threshold (σ).
[0054] The purpose of data denoising is to further remove random noise in the point cloud, smooth the point cloud surface, and enhance the geometric features of the point cloud. Method: Moving Least Squares (MLS): Principle: Smooth the point cloud data by locally fitting a polynomial surface, and perform weighted least squares fitting on each point and its neighboring points. Function: Smooth the point cloud surface while retaining geometric features and is suitable for dealing with high-density point clouds. Parameters: Order of the fitting polynomial, neighborhood search radius, etc. Gaussian Filter: Principle: Smooth the point cloud based on the Gaussian kernel function. The new position of each point is obtained by weighted averaging of its neighboring points, and the weights are determined by the Gaussian kernel function. Function: Smooth the point cloud data, reduce noise, and at the same time avoid the loss of geometric features caused by excessive smoothing. Parameters: Standard deviation (σ) of the Gaussian kernel and number of neighboring points.
[0055] Purpose of data augmentation: Through data augmentation techniques, increase the diversity and richness of point cloud data and improve the generalization ability of subsequent deep learning models. Method: Rotation augmentation: Principle: Randomly rotate the point cloud data (rotate around the x, y, and z axes) to generate new point cloud samples. Function: Simulate point cloud data from different perspectives and enhance the model's learning ability for rotational invariance. Parameter: Rotation angle range (such as ±30°). Scaling augmentation: Principle: Randomly scale the point cloud data to generate point cloud samples of different scales. Function: Simulate point cloud data at different distances and enhance the model's adaptability to scale changes. Parameter: Scaling ratio range (such as 0.8 to 1.2). Noise injection: Principle: Randomly add Gaussian noise to the point cloud data to simulate noise interference in the actual scenario. Function: Enhance the model's robustness to noise and improve the reliability of the model in actual applications. Parameter: Standard deviation (σ) of the noise. Cropping augmentation: Principle: Randomly crop a part of the point cloud data to simulate partially occluded or incomplete point cloud data. Function: Enhance the model's processing ability for incomplete data and improve the model's robustness. Parameter: Cropping ratio (such as cropping 10% to 30% of the point cloud). 4. Data normalization Purpose: Normalize the coordinate range of point cloud data to a unified scale for subsequent processing and model training. Method: Coordinate normalization: Principle: Normalize the coordinate values of the point cloud to the range of [-1, 1] or [0, 1] by calculating the relative position of each point's coordinate to the center of the point cloud. Function: Eliminate the scale differences between different point cloud data and improve the model's generalization ability. Parameter: Normalization range (such as [-1, 1] or [0, 1]). Intensity normalization (if the point cloud contains intensity information): Principle: Normalize the intensity values of the point cloud to the range of [0, 1] through linear transformation. Function: Enhance the intensity information of the point cloud data for subsequent processing. Parameter: Minimum and maximum values of the intensity.
[0056] S2: Input the preprocessed point cloud data into the deep learning model. The model extracts features from the point cloud data through the convolutional layer to generate feature vectors, and uses the saliency maximum pooling layer to process the feature map to retain key feature information;
[0057] S21: Perform a convolution operation on the point cloud data through a convolution kernel to extract local features in the data. Each convolution kernel slides on the point cloud data to calculate the feature response of the local area and generate multiple feature maps;
[0058] S22: Introduce the saliency maximum pooling layer. The saliency maximum pooling layer selects the maximum response value in the local area of each feature map to retain the most significant feature information and remove redundant details at the same time;
[0059] S23: The fully connected layer flattens the pooled feature map into a one-dimensional vector and performs a weighted combination of the features by learning weights to generate the final feature vector. The global pooling operation directly generates a feature vector of a fixed length by aggregating all the elements of the feature map.
[0060] In this embodiment, during the point cloud data processing, the preprocessed point cloud data is first input into a convolutional neural network (CNN). As the core part of the network, the convolutional layer performs convolutional operations on the point cloud data through a series of convolutional kernels to extract local features in the data. Each convolutional kernel slides over the point cloud data to calculate the feature responses of local regions and generates multiple feature maps. These feature maps capture the geometric structure, shape, and spatial distribution information in the point cloud data, thus converting the original point cloud data into a high-dimensional feature representation, providing a basis for subsequent feature analysis and processing.
[0061] After the convolutional layer extracts the feature maps, in order to further retain key feature information and reduce the feature dimension, a saliency-max pooling layer is introduced. The role of the saliency-max pooling layer is to select the maximum response value in the local region of each feature map, thereby retaining the most significant feature information while removing redundant details. In this way, the pooling layer can enhance the robustness of the features, making the model more invariant to small-scale changes (such as noise or minor geometric deformations) in the input data. In addition, saliency-max pooling can also reduce the spatial resolution of the feature map, reduce the computational amount, and retain the most important features, providing a more compact feature representation for subsequent classification or distance calculation tasks.
[0062] After being processed by the saliency-max pooling, the obtained feature maps are further integrated into feature vectors. This process is usually completed through a fully connected layer or a global pooling operation. The fully connected layer flattens the pooled feature map into a one-dimensional vector and performs a weighted combination of the features by learning weights to generate the final feature vector. The global pooling operation directly generates a feature vector of a fixed length by aggregating all the elements of the feature map (such as global max pooling or global average pooling). The generated feature vector contains the high-level semantic information and key geometric features of the point cloud data, and can be used for subsequent object recognition, classification, or distance calculation tasks, providing a compact and informative feature representation for the model's decision-making.
[0063] S3: Use a deep learning model to identify and classify the objects in the point cloud data, and determine the position information of each object in the three-dimensional space, including coordinates and bounding boxes.
[0064] S31: Input the preprocessed and feature-extracted point cloud data into a deep learning model, and use an instance segmentation algorithm to process the point cloud data, accurately label each object instance in the point cloud, and generate an independent segmentation mask for each instance.
[0065] S32: Through the convolutional neural network and the dedicated point cloud network, the model learns the geometric features and semantic information of the objects, and classifies each object instance into predefined categories;
[0066] S33: By combining the instance segmentation results and the three-dimensional coordinate information in the point cloud data, calculate the center coordinates, dimensions, and orientation of each object;
[0067] S331: Extract the point cloud data corresponding to each object instance from the segmentation mask generated by the instance segmentation algorithm, and perform clustering analysis of the extracted point cloud data using a density-based clustering algorithm;
[0068] S332: Calculate the centroid for each clustering result by taking the average of the coordinates of all points in the cluster as the center coordinates of the object. The centroid is calculated by taking the average of the coordinates of all points in the cluster, and the centroid reflects the geometric center position of the object in three-dimensional space and is a key parameter of the object's position information;
[0069] S333: Determine the bounding box of the object according to the distribution range of the point cloud data, calculate the maximum and minimum values of the point cloud in the x, y, and z directions, determine the length, width, and height of the bounding box. At the same time, according to the distribution direction of the point cloud, calculate the direction vector of the bounding box
[0070] In this embodiment: First, extract the point cloud data corresponding to each object instance from the segmentation mask generated by the instance segmentation algorithm. The segmentation mask assigns each point in the point cloud to a specific object instance, so all points belonging to the same object can be filtered out through the mask information. This step ensures that the subsequent analysis is only for the point cloud data of a specific object, providing a basis for accurately calculating the geometric features of the object.
[0071] Next, perform clustering analysis on the extracted point cloud data to further optimize the point cloud representation of the object. The purpose of clustering analysis is to divide the points in the point cloud into multiple subsets according to the spatial distribution, and each subset represents a local area of the object. Common clustering methods include Euclidean clustering or density-based clustering algorithms (such as DBSCAN). Through clustering, noise points and outliers can be removed, and at the same time, the geometric consistency of the object can be enhanced, providing more accurate data for subsequent centroid calculation and bounding box determination.
[0072] After completing the clustering analysis, calculate the centroid for each clustering result as the center coordinates of the object. The centroid is calculated by taking the average of the coordinates of all points in the cluster. The centroid reflects the geometric center position of the object in three-dimensional space and is a key parameter of the object's position information. By calculating the centroid, an accurate reference point can be provided for subsequent distance calculation and object positioning.
[0073] Finally, determine the bounding box of the object according to the distribution range of the point cloud data. The bounding box is a minimum bounding box used to describe the position and size of the object in three-dimensional space. By calculating the maximum and minimum values of the point cloud in the x, y, and z directions, the length, width, and height of the bounding box are determined. At the same time, according to the distribution direction of the point cloud, the direction vector of the bounding box is calculated to ensure that the bounding box can tightly enclose the object. The bounding box not only provides the exact range of the object but also provides important geometric information for subsequent distance calculations and safety monitoring.
[0074] S34: Generate the bounding box of each object in three-dimensional space according to the center coordinates and size information of each object
[0075] In this embodiment: First, input the preprocessed and feature-extracted point cloud data into a deep learning model, and use an instance segmentation algorithm (such as Mask R-CNN or PointRCNN) to process the point cloud data. The instance segmentation algorithm precisely labels each object instance in the point cloud through pixel-level classification and segmentation, and generates an independent segmentation mask for each instance. This process not only identifies the object categories in the point cloud but also distinguishes different instances under the same category, providing a basis for subsequent object localization and distance calculation.
[0076] Based on instance segmentation, the deep learning model further performs category recognition and classification on each object instance in the point cloud. Through a convolutional neural network (CNN) or a point cloud-specific network (such as PointNet++), the model learns the geometric features and semantic information of the object and classifies each object instance into predefined categories (such as personnel, mechanical equipment, tools, etc.). This classification process utilizes the powerful feature extraction ability of the deep learning model to accurately identify different objects in the point cloud and assign corresponding category labels to them.
[0077] After completing the category recognition of the object, the model further determines the position information of each object in three-dimensional space. By combining the instance segmentation results and the three-dimensional coordinate information in the point cloud data, the center coordinates, size, and direction of each object are calculated. The specific method is to perform clustering analysis on the point cloud data within the segmentation mask of each object instance, calculate its centroid as the center coordinate of the object, and determine the bounding box of the object (including length, width, height, and direction) according to the distribution range of the point cloud. This process provides accurate object position information for subsequent distance calculations.
[0078] Finally, based on the center coordinates and size information of each object, a bounding box in three-dimensional space is generated. The bounding box is the minimum bounding box of the object and can accurately describe the position and range of the object in three-dimensional space. By calculating the vertex coordinates and directions of the bounding box, the model can integrate the geometric information and semantic information of the object, providing a complete object description for subsequent distance calculation and safety monitoring. This process ensures that the position and range of each object are accurately represented, thus supporting the high-precision ranging and safety warning functions of the system.
[0079] S4: Calculate the relative distance between objects based on the position information of the objects. For multiple objects, calculate the distance between them and the key area.
[0080] S5: Transmit the calculated ranging results to the safety supervision terminal in real time through the wireless communication module. The safety supervision terminal displays the distance information between objects in real time. When the distance between objects is lower than the preset safety threshold, the terminal issues an alarm to remind the on-site personnel to take safety measures.
[0081] Please refer to Figure 6 , a job site object measurement system based on three-dimensional point cloud ranging, used to execute the job site object measurement method based on three-dimensional point cloud ranging, including a lidar scanning module 101, a deep learning processing module 102, a data transmission module 103, and a safety supervision terminal 104. The deep learning processing module 102 receives the signal from the lidar scanning module 101, the data transmission module 103 receives the signal from the deep learning processing module 102, and the safety supervision terminal 104 receives the signal from the data transmission module 103;
[0082] The lidar scanning module 101 is used to perform three-dimensional scanning on the job site to obtain point cloud data containing the three-dimensional coordinate information of the objects;
[0083] The deep learning processing module 102 is used to extract features and match the point cloud data to calculate the accurate distance between objects;
[0084] The data transmission module 103 is used to transmit the ranging results to the safety supervision terminal 104 in real time;
[0085] The safety supervision terminal 104 is used to receive and display the ranging results and issue an alarm when the distance does not meet the safety requirements.
[0086] In this embodiment, the convolutional layer is used to extract features from the input power operation scene image to obtain multiple feature maps; the saliency maximum pooling layer is used to perform maximum pooling operations on the multiple feature maps to obtain a pooling result; the saliency component localization learning layer is used to perform weighted summation on the pooling result to obtain a summation result; and the classification layer is used to output the recognition result of the violation behavior in the power operation scene image according to the summation result.
[0087] The convolutional layer, the saliency maximum pooling layer, and the classification layer constitute the prototype learning network structure. The convolutional layer is used to perform a more in-depth analysis of the input image to obtain features with a higher level of abstraction. The specific number of layers of the convolutional layer can be set according to the actual situation. The saliency maximum pooling layer is connected after the last layer of the convolutional layer. Using the saliency maximum pooling layer connected after the convolutional layer can retain the position information of the convolutional layer, making the network structure interpretable. At the same time, using saliency maximum pooling instead of average pooling is because the salient objects in the power operation scene are relatively clear. Using the saliency maximum pooling layer in the model can capture the objects in the power operation scene more accurately, thus making the final recognition result more accurate. The saliency component localization learning layer specifically overlays the feature maps with weights to obtain the saliency component localization map corresponding to a specific category. Through the saliency component localization learning layer, it can be shown which feature maps among the multiple feature maps the basis for the model's classification decision mainly comes from, that is, making the final recognition result traceable. The classification layer is essentially a normalization network. The purpose of the classification layer is to map multiple scalars into a probability distribution. Each value output by the classification layer ranges from 0 to 1. Thus, through the classification layer, the probability distribution of different types in the power operation scene image can be obtained.
[0088] The operation of the saliency maximum pooling layer is represented by the following formula:
[0089] F k =Max(f k (x,y))
[0090] In the formula, fk(x,y) represents the value at the position (x, y) on the k-th feature map.
[0091] The operation of the saliency component localization learning layer is represented by the following formula:
[0092]
[0093] In the formula, c represents the category, and w represents the weight of the k-th feature map.
[0094] Based on the above formula, the saliency map belonging to a specific category c is defined as:
[0095]
[0096] As can be seen from this formula, Mc(x, y) represents the weighted sum of different saliency component feature maps for identifying a certain category c. The generated Mc(x, y) is enlarged to the size of the original image to obtain the saliency component localization map corresponding to a certain category c. Among them, it should be noted that the weight w, as a parameter of the model, can be adjusted and determined during the model training process. The network structure including a convolutional layer, a saliency max pooling layer, a saliency component localization learning layer, and a classification layer is trained using the labeled power operation scenario images, and the weight w is adjusted and determined during the training process.
[0097] The operation of the classification layer is represented by the following formula:
[0098]
[0099] In the formula, Pc represents the classification value corresponding to category c. The result Sc after weighted summation by the saliency component localization learning layer is input to the classification layer, and the classification layer uses this formula to calculate the specific score belonging to category c.
[0100] Power operation scenario violation behavior recognition method: Obtain the power operation scenario image to be recognized through video monitoring; input the power operation scenario image to be recognized into the power operation scenario violation behavior recognition model to obtain the recognition result; the recognition result includes the category of the violation behavior.
[0101] The above-disclosed are only one or more preferred embodiments of the present application, and the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A method for measuring objects at a work site based on three-dimensional point cloud ranging, characterized in that: The steps include: Install high-precision laser radar at the work site to scan the work site, obtain point cloud data containing the three-dimensional coordinate information of the object, and pre-process the collected point cloud data; The preprocessed point cloud data is input into the deep learning model. The model extracts features from the point cloud data through the convolution layer, generates feature vectors, and processes the feature map using the saliency maximum pooling layer to retain key feature information. Use deep learning models to identify and classify objects in point cloud data and determine the location information of each object in three-dimensional space, including coordinates and bounding boxes; According to the position information of the objects, the relative distance between the objects is calculated. For multiple objects, the distance between them and the key area is calculated. The calculated distance measurement results are transmitted to the safety monitoring terminal in real time through the wireless communication module. The safety monitoring terminal displays the distance information between objects in real time. When the distance between objects is lower than the preset safety threshold, the terminal issues an alarm to remind on-site personnel to take safety measures.
2. A method for measuring objects on a work site based on three-dimensional point cloud ranging as claimed in claim 1, characterized in that: In "Preprocessing the collected point cloud data", the following steps are included: The point cloud is downsampled using voxel filtering, which divides the point cloud into a three-dimensional voxel grid and retains a mean point within each voxel; The moving least squares method is used to perform weighted least squares fitting on each point and its neighboring points through local polynomial fitting; By randomly rotating the point cloud data, we simulate scenes at different perspectives; by randomly scaling, we simulate point cloud data at different distances; by injecting Gaussian noise, we simulate noise interference in actual scenes; Normalize the coordinate values of the point cloud to a uniform scale range, and achieve coordinate normalization by calculating the relative position of the coordinates of each point with the center of the point cloud.
3. The method for measuring objects on a work site based on three-dimensional point cloud ranging according to claim 1, characterized in that: In "the model extracts features from point cloud data through a convolutional layer, generates feature vectors, and processes feature maps using a saliency maximum pooling layer to retain key feature information", the method further includes the following steps: The convolution kernel is used to perform convolution operation on the point cloud data to extract local features in the data. Each convolution kernel slides on the point cloud data to calculate the feature response of the local area and generate multiple feature maps. The saliency max pooling layer is introduced. The saliency max pooling layer selects the maximum response value in the local area of each feature map, retains the most significant feature information, and removes redundant details. The fully connected layer flattens the pooled feature map into a one-dimensional vector, and performs weighted combination of the features through learning weights to generate the final feature vector. The global pooling operation directly generates a fixed-length feature vector by aggregating all elements of the feature map.
4. The method for measuring objects on a work site based on three-dimensional point cloud ranging according to claim 1, characterized in that: In “Using a deep learning model to identify and classify objects in point cloud data and determine the location information of each object in three-dimensional space, including coordinates and bounding boxes”, the method includes: The preprocessed and feature-extracted point cloud data is input into the deep learning model, and the instance segmentation algorithm is used to process the point cloud data, accurately mark each object instance in the point cloud, and generate an independent segmentation mask for each instance; Through convolutional neural networks and point cloud-specific networks, the model learns the geometric features and semantic information of objects and classifies each object instance into predefined categories; By combining the instance segmentation results and the 3D coordinate information in the point cloud data, the center coordinates, size and orientation of each object are calculated; Based on the center coordinates and size information of each object, its bounding box in three-dimensional space is generated.
5. The method for measuring objects on a work site based on three-dimensional point cloud ranging according to claim 4, characterized in that: In “calculating the center coordinates, size and orientation of each object by combining the instance segmentation results and the three-dimensional coordinate information in the point cloud data”, the method includes the following steps: Extract the point cloud data corresponding to each object instance from the segmentation mask generated by the instance segmentation algorithm, and perform clustering analysis on the extracted point cloud data using a density-based clustering algorithm; For each clustering result, the centroid is calculated by taking the average of the coordinates of all points in the cluster as the center coordinate of the object; Determine the bounding box of the object according to the distribution range of the point cloud data, calculate the maximum and minimum values of the point cloud in the x, y, and z directions, determine the length, width, and height of the bounding box, and at the same time, calculate the direction vector of the bounding box according to the distribution direction of the point cloud.
6. The method for measuring objects on a work site based on three-dimensional point cloud ranging according to claim 5, characterized in that: In “calculating the centroid for each clustering result by averaging the coordinates of all points in the cluster as the center coordinates of the object”, the method includes: The centroid is calculated by averaging the coordinates of all points in the cluster. The centroid reflects the geometric center position of the object in three-dimensional space and is a key parameter of the object's position information.
7. A system for measuring objects at a work site based on three-dimensional point cloud ranging, used to execute the method for measuring objects at a work site based on three-dimensional point cloud ranging as claimed in any one of claims 1 to 6, characterized in that: It includes a laser radar scanning module, a deep learning processing module, a data transmission module and a safety supervision terminal, wherein the deep learning processing module receives the signal of the laser radar scanning module, the data transmission module receives the signal of the deep learning processing module, and the safety supervision terminal receives the signal of the data transmission module; The laser radar scanning module is used to perform three-dimensional scanning on the work site to obtain point cloud data containing three-dimensional coordinate information of the object; The deep learning processing module is used to extract and match features of point cloud data and calculate the precise distance between objects; The data transmission module is used to transmit the ranging result to the safety supervision terminal in real time; The safety monitoring terminal is used to receive and display the distance measurement result, and to issue an alarm when the distance does not meet the safety requirements.
Citation Information
Patent Citations
Intelligent video monitoring system for power operation
CN112989110A