A Substation Safety Control Method and System Based on a 3D Reconstruction Model
By adopting a safety control method based on three-dimensional reconstruction model in the substation, using laser point clouds and image data to identify dynamic flight targets, the problems of low efficiency of traditional monitoring methods and high false alarm rate are solved, and high-precision safety monitoring and early warning are achieved.
Patent Information
- Application Number
- CN202510175153.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Traditional substation monitoring and early warning methods are inefficient and have a small coverage, making it difficult to meet the needs of modern power systems for real-time monitoring and efficient safety warning. Especially when dealing with diverse scenarios in complex environments, false alarms and underreporting are more common.
The substation safety management and control method based on the three-dimensional reconstruction model is adopted, and the dynamic area is extracted by obtaining the real-time data of the laser point cloud sensor and camera, and the background difference algorithm is used to extract dynamic areas, and the point cloud and image data are combined to match to identify and locate dynamic flight targets.
It significantly improves the safety control accuracy of the substation, reduces the false alarm and missed alarm rates, and realizes efficient monitoring and positioning of small and fast targets such as drones, ensuring comprehensive, accurate and efficient safety control of the substation environment.
Smart Images

Figure CN119672233B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of security control, and particularly to a substation security control method and system based on a three-dimensional reconstruction model. Background Art
[0002] In the power system, as the core hub for power transmission and conversion, the safety and stability of substations are directly related to the normal operation of the entire power grid. However, traditional substation monitoring and early warning methods mainly rely on manual inspections and a small number of fixed monitoring devices, which not only limit the monitoring efficiency and coverage but also result in high operating costs. More critically, these traditional methods are relatively slow in responding to emergencies and are difficult to meet the stringent requirements of modern power systems for real-time monitoring and efficient safety early warning of substations. Therefore, there is an urgent need for an efficient and intelligent monitoring system to improve the safety and response capabilities of substations and ensure the stable operation of the power grid.
[0003] Currently, with the development of three-dimensional laser scanning technology and computer vision technology, the application of three-dimensional reconstruction technology in the power system has gradually received attention. Using three-dimensional point cloud data for substation equipment monitoring and early warning can obtain the three-dimensional spatial information of equipment more comprehensively and accurately. However, in current abnormal behavior recognition algorithms, false alarms and missed detections are relatively common when dealing with diverse scenarios in complex environments. For example, for drones or small target objects, the accuracy of recognition is affected by the size and speed of the objects, and missed detections often occur. Summary of the Invention
[0004] This application provides a substation security control method and system based on a three-dimensional reconstruction model, which can improve the accuracy of substation security control.
[0005] In the first aspect of this application, a substation security control method based on a three-dimensional reconstruction model is provided. The method includes:
[0006] Obtain the real-time point cloud data input by the laser point cloud sensor and the real-time images collected by the camera;
[0007] Compare the real-time point cloud data with the background point cloud model point by point, and use the background difference algorithm to extract the changed dynamic area;
[0008] Based on the dynamic reconstruction module of the real-time point cloud data, identify the first position of the first dynamic target in the dynamic area;
[0009] Calculate the second position of the second dynamic target included in the real-time image;
[0010] If the matching between the first position and the second position is successful, and the matching between the sparse structure of the real-time point cloud data and the dense pixels of the real-time image is successful, it is determined that both the first dynamic target and the second dynamic target are dynamic flying targets;
[0011] Determine the position of the dynamic flying target according to the first position and the second position, and mark it in the three-dimensional model.
[0012] Optionally, the step of comparing the real-time point cloud data with the background point cloud model point by point and using the background difference algorithm to extract the changed dynamic area specifically includes:
[0013] Generate a real-time point cloud model according to the real-time point cloud data;
[0014] Perform clustering segmentation on the real-time point cloud model to obtain multiple real-time voxel units, and perform clustering segmentation on the background point cloud model to obtain multiple background voxel units;
[0015] Compare each of the multiple real-time voxel units with each of the multiple background voxel units one by one, and determine the first voxel unit corresponding to the background voxel unit among the multiple real-time voxel units;
[0016] Remove the first voxel unit from the multiple real-time voxel units to obtain a second voxel unit, and the second voxel unit is the dynamic area.
[0017] Optionally, the step of the dynamic reconstruction module based on the real-time point cloud data to identify the first position of the first dynamic target in the dynamic area specifically includes:
[0018] Determine the relative position between the second voxel unit and the first voxel unit according to the real-time point cloud model;
[0019] Align the real-time point cloud model with the background point cloud model through a point cloud registration algorithm;
[0020] Obtain the actual coordinate position of the background point cloud model;
[0021] Determine the first position of the first dynamic target according to the relative position and the actual coordinate position.
[0022] Optionally, before the matching between the sparse structure of the real-time point cloud data and the dense pixels of the real-time image is successful, the method further includes:
[0023] Align the laser point cloud sensor and the camera spatially to establish a unified coordinate system;
[0024] Calculate a feature descriptor for the first key points in the real-time image to obtain a first descriptor, and calculate a feature descriptor for the second key points in the real-time point cloud data to obtain a second descriptor;
[0025] Compare the first descriptor with the second descriptor, and match the first descriptor and the second descriptor whose similarity exceeds a preset threshold;
[0026] If it is determined that the matching rates of multiple first descriptors reach a preset condition, or the matching rates of multiple second descriptors reach a preset condition, then map the coordinates of the first key points to the dynamic reconstruction model to calculate the position of the dynamic flying target.
[0027] Optionally, the step of comparing the real-time point cloud data with the background point cloud model point by point and using the background difference algorithm to extract the changed dynamic region specifically further includes:
[0028] According to a predefined side length dimension, divide the real-time point cloud model into multiple first voxel points, and divide the background point cloud model into multiple second voxel points;
[0029] According to the geometric data of the first voxel points included in the real-time voxel unit, calculate the attribute features of each real-time voxel unit, and the attribute features include point density, average coordinates, and geometric center;
[0030] According to the geometric data of the second voxel points included in the background voxel unit, calculate the attribute features of each background voxel unit, and the attribute features include point density, average coordinates, and geometric center;
[0031] Compare the real-time voxel units with the background voxel units one by one, determine the dynamic voxel units in which the change value of the attribute features exceeds a preset standard among the multiple real-time voxel units, and determine the dynamic voxel units as the dynamic region.
[0032] Optionally, calculating a feature descriptor for the first key points in the real-time image to obtain a first descriptor specifically includes:
[0033] Downsample the real-time image multiple times to generate a series of image hierarchical structures with different resolutions. For each resolution of the image hierarchical structure, apply Gaussian filtering to the image to smooth the image and generate Gaussian blurred images at different scales. The Gaussian blurred image is represented as follows:
[0034]
[0035] where I(x, y) is the Gaussian blurred image, L(x, y, σ) is the Gaussian blurred image, and G(x, y, σ) is the Gaussian kernel, Denotes a convolution operation;
[0036] Calculate the difference between Gaussian blurred images of adjacent scales to generate a Difference of Gaussian (DoG) image, which is expressed as follows:
[0037]
[0038] where \(D(x,y,\sigma)\) is the DoG image, \(L(x,y,\sigma)\) is the Gaussian blurred image, \(L(x,y,k\sigma)\) is the adjacent Gaussian blurred image of the Gaussian blurred image, and \(k\) is a constant multiple;
[0039] Compare any pixel point in the DoG image with multiple adjacent pixel points. If the value of the any pixel point is greater than or less than the values of multiple adjacent pixel points spatially, and the value of the any pixel point is greater than or less than the values of multiple adjacent pixel points in terms of scale, then the any pixel point is marked as a key point to obtain the first key point;
[0040] Calculate the amplitude and direction of pixel points in the neighborhood of the first key point, specifically calculated through the following formula:
[0041]
[0042]
[0043] where \(m(x,y)\) is the amplitude of pixel points in the neighborhood of the first key point, \(\theta\) is the direction of pixel points in the neighborhood of the first key point, \(L(x + 1,y)\) and \(L(x - 1,y)\) respectively represent the pixel values of the left and right adjacent pixel points of the first key point in the horizontal direction of the DoG image, and \(L(x,y + 1)\) and \(L(x,y - 1)\) respectively represent the pixel values of the upper and lower adjacent pixel points of the first key point in the vertical direction of the DoG image;
[0044] Divide the neighborhood of the first key point into multiple sub - regions of the same size, and calculate the histogram of each sub - region according to the amplitude and direction of pixel points in the neighborhood of the key point;
[0045] Merge multiple histograms to form a multi - dimensional vector to obtain the first descriptor.
[0046] Optionally, calculating the feature descriptor for the second key point in the real - time point cloud data to obtain the second descriptor specifically includes:
[0047] Calculate the point curvature of any point cloud point corresponding to the real - time point cloud data, and the point curvature is calculated through the following formula:
[0048]
[0049] Where C is the point curvature, n i is the normal vector of the adjacent point of any one point cloud point, n is the normal vector of any one point cloud point, and k is the number of adjacent points of any one point cloud point;
[0050] If it is determined that the curvature of any one point cloud point is greater than the preset curvature, then it is determined that any one point cloud point is a key point, and the second key point is obtained;
[0051] Taking the second key point as the center, according to the direction of the normal vector of the second key point, the neighborhood is divided into a plurality of spherical sub-regions;
[0052] In each of the spherical sub-regions, a normal vector direction histogram is calculated, and the features of the spherical sub-regions are accumulated to generate an overall direction histogram, and the data of the direction histogram is the second descriptor.
[0053] In the second aspect of the present application, a substation safety control system based on a three-dimensional reconstruction model is provided. The system includes an acquisition module, an extraction module, an identification module, a processing module, and a judgment module, where:
[0054] The acquisition module is configured to acquire real-time point cloud data input by a laser point cloud sensor and real-time images collected by a camera;
[0055] The extraction module is configured to compare the real-time point cloud data with the background point cloud model point by point, and use the background difference algorithm to extract the changed dynamic region;
[0056] The identification module is configured to identify the first position of the first dynamic target in the dynamic region based on the dynamic reconstruction module performed on the real-time point cloud data;
[0057] The processing module is configured to calculate the second position of the second dynamic target included in the real-time image;
[0058] The judgment module is configured to determine that both the first dynamic target and the second dynamic target are dynamic flying targets if the first position and the second position are successfully matched, and the sparse structure of the real-time point cloud data and the dense pixels of the real-time image are successfully matched;
[0059] The processing module is configured to determine the position of the dynamic flying target according to the first position and the second position, and mark it in the three-dimensional model.
[0060] Optionally, the processing module is configured to generate a real-time point cloud model according to the real-time point cloud data;
[0061] The processing module is configured to perform clustering segmentation on the real-time point cloud model to obtain a plurality of real-time voxel units, and perform clustering segmentation on the background point cloud model to obtain a plurality of background voxel units;
[0062] The judgment module is configured to compare the plurality of real-time voxel units with the plurality of background voxel units one by one, and determine, among the plurality of real-time voxel units, the first voxel units corresponding to the background voxel units;
[0063] The processing module is configured to remove the first voxel units from the plurality of real-time voxel units to obtain second voxel units, and the second voxel units are the dynamic regions.
[0064] Optionally, the processing module is configured to determine the relative positions of the second voxel units and the first voxel units according to the real-time point cloud model;
[0065] The processing module is configured to align the real-time point cloud model with the background point cloud model through a point cloud registration algorithm;
[0066] The acquisition module is configured to acquire the actual coordinate positions of the background point cloud model;
[0067] The judgment module is configured to determine the first positions of the first dynamic targets according to the relative positions and the actual coordinate positions.
[0068] Optionally, the processing module is configured to perform spatial alignment on the lidar sensor and the camera to establish a unified coordinate system;
[0069] The processing module is configured to calculate feature descriptors for the first key points in the real-time image to obtain first descriptors, and calculate feature descriptors for the second key points in the real-time point cloud data to obtain second descriptors;
[0070] The judgment module is configured to compare the first descriptors with the second descriptors, and match the first descriptors and the second descriptors whose similarity exceeds a preset threshold;
[0071] The recognition module is configured to, if it is determined that the matching rates of the plurality of first descriptors reach a preset condition, or the matching rates of the plurality of second descriptors reach a preset condition, map the coordinates of the first key points to the dynamic reconstruction model to calculate the positions of the dynamic flying targets.
[0072] Optionally, the processing module is configured to subdivide the real-time point cloud model into a plurality of first voxel points according to a predefined side length dimension, and subdivide the background point cloud model into a plurality of second voxel points;
[0073] The processing module is configured to calculate the attribute features of each real-time voxel unit according to the geometric data of the first voxel points included in the real-time voxel unit, where the attribute features include point density, average coordinates, and geometric center;
[0074] The processing module is configured to calculate the attribute features of each background voxel unit according to the geometric data of the second voxel points included in the background voxel unit, where the attribute features include point density, average coordinates, and geometric center;
[0075] The judgment module is configured to compare the real-time voxel units with the background voxel units one by one, determine the dynamic voxel units among the multiple real-time voxel units where the change value of the attribute features exceeds a preset standard, and determine the dynamic voxel units as the dynamic region.
[0076] Optionally, the processing module is configured to perform multiple downsamplings on the real-time image to generate a series of image hierarchical structures with different resolutions. For each image hierarchical structure with a resolution, apply Gaussian filtering to the image to smooth the image and generate Gaussian blurred images at different scales. The Gaussian blurred image is expressed as follows:
[0077]
[0078] where I(x, y) is the Gaussian blurred image, L(x, y, σ) is the Gaussian blurred image, G(x, y, σ) is the Gaussian kernel, represents the convolution operation;
[0079] The processing module is configured to calculate the difference between Gaussian blurred images at adjacent scales to generate a difference Gaussian image. The difference Gaussian image is expressed as follows:
[0080]
[0081] where D(x, y, σ) is the difference Gaussian image, L(x, y, σ) is the Gaussian blurred image, L(x, y, kσ) is the adjacent Gaussian blurred image of the Gaussian blurred image, and k is a constant multiple;
[0082] The processing module is configured to compare any pixel point in the difference Gaussian image with multiple adjacent pixel points. If the value of any pixel point in space is greater than or less than the values of multiple adjacent pixel points in space, and the value of any pixel point in scale is greater than or less than the values of multiple adjacent pixel points in space, then the any pixel point is marked as a key point to obtain the first key point;
[0083] The processing module is used to calculate the amplitude and direction of the pixel points in the neighborhood of the first key point, and is specifically calculated by the following formula:
[0084]
[0085]
[0086] Where m(x, y) is the amplitude of the pixel points in the neighborhood of the first key point, θ is the direction of the pixel points in the neighborhood of the first key point, L(x + 1, y) and L(x - 1, y) respectively represent the pixel values of the pixel points adjacent to the first key point on the left and right in the horizontal direction of the difference Gaussian image, and L(x, y + 1) and L(x, y - 1) respectively represent the pixel values of the pixel points adjacent to the first key point above and below in the vertical direction of the difference Gaussian image;
[0087] The processing module is used to divide the neighborhood of the first key point into multiple sub-regions of the same size, and calculate the histogram of each sub-region according to the amplitude and direction of the pixel points in the neighborhood of the key point;
[0088] The processing module is used to merge multiple histograms to form a multi-dimensional vector to obtain the first descriptor.
[0089] Optionally, the processing module is used to calculate the point curvature of any point cloud point corresponding to the real-time point cloud data, and the point curvature is calculated by the following formula:
[0090]
[0091] Where C is the point curvature, n i is the normal vector of the adjacent points of any point cloud point, n is the normal vector of any point cloud point, and k is the number of adjacent points of any point cloud point;
[0092] The processing module is used to determine that any point cloud point is a key point if it is determined that the curvature of any point cloud point is greater than the preset curvature, and obtain the second key point;
[0093] The processing module is used to divide the neighborhood into multiple spherical sub-regions with the second key point as the center according to the direction of the normal vector of the second key point;
[0094] The processing module is used to calculate the normal vector direction histogram in each spherical sub-region, accumulate the features of the spherical sub-region, and generate an overall direction histogram, and the data of the direction histogram is the second descriptor.
[0095] In the third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. Both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory, so that the electronic device executes the method described in any one of the above.
[0096] In the fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, and when the instructions are executed, the method described in any one of the above is executed.
[0097] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0098] 1. In the present application, real-time data is collected through a laser point cloud sensor and a camera. The background difference algorithm is used to detect the dynamic area, and the matching results of the real-time point cloud data and the image data are combined, so as to achieve the accurate recognition and positioning of dynamic flight targets. By integrating the three-dimensional spatial information of the point cloud and the dense visual details of the image, after identifying the dynamic target, its type and position are confirmed through multi-data source matching, significantly reducing the false alarm and missed alarm rates, and ensuring the efficient monitoring and positioning of small and fast targets such as unmanned aerial vehicles. Finally, by real-time annotating the position of the dynamic target in the three-dimensional model, the comprehensive, accurate and efficient safety control of the substation environment is realized, thus greatly improving the safety control accuracy of the substation.
[0099] 2. By comparing the real-time point cloud data with the background point cloud model point by point, and combining voxel segmentation and the background difference algorithm, the dynamic area is extracted. Specifically, the point cloud model is subdivided into voxel units, and the geometric attributes of the real-time voxel units and the background voxel units are compared one by one to identify the dynamically changing voxel areas with obvious changes. This process can effectively filter out the static background and only retain the dynamically changing areas, enabling the system to accurately identify abnormal activities around the equipment. In a complex environment, the false alarm and missed alarm probabilities are significantly reduced, thus greatly improving the accuracy and response efficiency of substation safety monitoring.
[0100] 3. The real-time image is processed at multiple scales, and Gaussian blur and differential Gaussian calculations are applied to accurately extract key points and generate their feature descriptors. This process effectively improves the detection ability of key points at different scales, ensuring the robustness of key points to image rotation and illumination changes. By calculating the amplitude and direction of pixels in the neighborhood and constructing a direction histogram, the finally generated first descriptor not only comprehensively describes the local features of the key points, but also can effectively distinguish different objects in the image. The efficient extraction and accurate description of key points in a dynamic environment provide a solid foundation for subsequent target recognition and matching, thus significantly improving the overall performance and reliability of the substation safety monitoring system. Brief Description of the Drawings
[0101] Figure 1 is a schematic flowchart of a substation safety control method based on a three-dimensional reconstruction model disclosed in an embodiment of the present application;
[0102] Figure 2 is a schematic diagram of modules of a substation safety control system based on a three-dimensional reconstruction model disclosed in an embodiment of the present application;
[0103] Figure 3 is a schematic structural diagram of an electronic device.
[0104] Description of Reference Numerals: 201, acquisition module; 202, extraction module; 203, recognition module; 204, processing module; 205, judgment module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Embodiments
[0105] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0106] In the description of the embodiments of the present application, words such as "for example" or "for illustration" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "for example" or "for illustration" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "for example" or "for illustration" is intended to present relevant concepts in a specific manner.
[0107] In the description of the embodiments of the present application, the meaning of the term "a plurality" refers to two or more. For example, a plurality of systems refers to two or more systems, and a plurality of screen terminals refers to two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0108] As the core hub of the power system, the traditional monitoring method of the substation relies on manual inspections and fixed equipment. Not only is the monitoring efficiency low, the operating cost high, but also the response is slow in emergencies, making it difficult to meet the real-time monitoring requirements. With the development of 3D laser scanning and computer vision technologies, the monitoring method based on 3D point clouds has gradually been applied to substations, enabling accurate acquisition of the spatial information of equipment and enhancing safety. However, the current abnormal behavior recognition algorithms are prone to false alarms and missed detections in complex scenarios, and the detection effect for small and fast targets (such as drones) is poor. Therefore, a more efficient safety control method is urgently needed to ensure the stable operation of the power grid.
[0109] This embodiment discloses a substation safety control method based on a 3D reconstruction model, referring to Figure 1 , including the following steps S110 - S160:
[0110] S110, obtain the real-time point cloud data input by the laser point cloud sensor and the real-time images collected by the camera.
[0111] The substation safety control method based on a 3D reconstruction model disclosed in the embodiments of this application is applied to a server. The server includes, but is not limited to, electronic devices such as mobile phones, tablets, wearable devices, PCs (Personal Computers), etc., and can also be a background server running a substation safety control method based on a 3D reconstruction model. The server can be implemented by an independent server or a server cluster composed of multiple servers.
[0112] Deploy a laser point cloud sensor in the substation to collect the 3D point cloud data of the substation area. The point cloud data can accurately describe the shape and position of the equipment in space, generating the basic data required for the 3D reconstruction model. Install a camera in the monitoring area to collect real-time image data for capturing the 2D information of the equipment, personnel, and other moving objects. These image data can provide texture and color information to make up for the deficiencies of the point cloud data.
[0113] Transmit the data collected by the laser point cloud sensor and the camera to the server through the network. To ensure that the point cloud and image data are collected within the same time period, it is necessary to synchronize the timestamps of the laser sensor and the camera for accurate registration in subsequent processing.
[0114] Process the collected point cloud data, such as denoising, downsampling, and filtering, to remove redundant data and noise. Under the condition that it is manually determined that there are no dynamic interference objects (such as drones), use multiple frames of point cloud data to construct a static background model of the substation, representing the static layout and spatial position of the equipment, providing a benchmark for subsequent dynamic area detection. Generate a high-precision 3D reconstruction model of the substation on the server using the background point cloud model and use it for display and monitoring.
[0115] S120 compares the real-time point cloud data with the background point cloud model point by point, and uses the background difference algorithm to extract the dynamic area where changes occur.
[0116] In a possible implementation, comparing the real-time point cloud data with the background point cloud model point by point and using the background difference algorithm to extract the dynamic area where changes occur specifically includes: generating a real-time point cloud model according to the real-time point cloud data; performing clustering segmentation on the real-time point cloud model to obtain multiple real-time voxel units, and performing clustering segmentation on the background point cloud model to obtain multiple background voxel units; comparing the multiple real-time voxel units with the multiple background voxel units one by one to determine the first voxel units corresponding to the background voxel units among the multiple real-time voxel units; removing the first voxel units from the multiple real-time voxel units to obtain second voxel units, and the second voxel units are the dynamic areas.
[0117] In a possible implementation, comparing the real-time point cloud data with the background point cloud model point by point and using the background difference algorithm to extract the dynamic area where changes occur specifically further includes: subdividing the real-time point cloud model into multiple first voxel points according to a predefined side length dimension, and subdividing the background point cloud model into multiple second voxel points; calculating the attribute features of each real-time voxel unit according to the geometric data of the first voxel points included in the real-time voxel unit, and the attribute features include point density, average coordinates, and geometric center; calculating the attribute features of each background voxel unit according to the geometric data of the second voxel points included in the background voxel unit, and the attribute features include point density, average coordinates, and geometric center; comparing the real-time voxel units with the background voxel units one by one to determine the dynamic voxel units among the multiple real-time voxel units where the change value of the attribute features exceeds a preset standard, and determining the dynamic voxel units as the dynamic areas.
[0118] Specifically, according to the point cloud data of the current substation collected in real time from the laser point cloud sensor, a three-dimensional point cloud model containing the real-time state of the substation is generated. Determine a suitable voxel grid size (such as a side length of 0.1 meter or 0.2 meter) to divide the spatial range of the point cloud model, and the voxel size depends on the size of the target object and the detection accuracy requirements.
[0119] Perform voxel division on the point data of the real-time point cloud model, refine the point cloud points within the spatial range into multiple "first voxel points", and perform the same voxel division on the point data in the background point cloud model to divide the background point cloud data into multiple "second voxel points". For each voxel point, calculate its voxel index (V x ,V y ,V z ) according to its spatial coordinates (x, y, z), and allocate the point cloud points to the corresponding voxel units. The voxel index is specifically calculated by the following formula:
[0120]
[0121] Among them, (x, y, z) are the spatial coordinates of the voxel point, and (V x , V y , V z ) is the voxel index of the voxel point, and d is the side length of the voxel unit.
[0122] The voxel index divides the three-dimensional space into individual voxel units (i.e., cubic small blocks). Through the index, the voxel unit to which each point belongs can be quickly found, facilitating the division of the entire space into multiple small regions. This can simplify the management and processing of large-scale point cloud data, making the data within each voxel unit easier to access and process. For example, when it is necessary to find the neighboring points of a certain point in the point cloud, the voxel unit where the point is located and the adjacent voxel units can be directly found through the voxel index, thereby accelerating the speed of neighboring point query.
[0123] Calculate the attribute features of each real-time voxel unit according to the geometric data of the first voxel point contained in the real-time voxel unit. The attribute features include point density, average coordinates, and geometric center. Calculate the attribute features of each background voxel unit according to the geometric data of the second voxel point contained in the background voxel unit. The attribute features include point density, average coordinates, and geometric center.
[0124] Specifically, calculate the total number of points in the voxel unit, that is, the point density. The point density can reflect the density of the voxel unit and is often used to detect whether a new object appears. Calculate the average coordinates of all points in the voxel unit, which represents the average position of the point cloud in space. The calculation formula for the average coordinates is:
[0125]
[0126] Among them, (X, Y, Z) are the average coordinates of the voxel unit, N is the number of voxel points contained in the voxel unit, and (x i , y i , z i ) are the coordinates of the i-th voxel point, that is, the geometric data of each voxel point. Use the average coordinates of the voxel unit as the geometric center, which represents the central position of the voxel unit in the three-dimensional space.
[0127] Among them, the calculation methods of the attribute features of the real-time voxel unit and the background voxel unit both adopt the above methods. However, these characteristic values of the background voxel unit are fixed in most cases and reflect the static structure information of the substation.
[0128] Pair the real-time voxel units with the background voxel units one by one, and compare whether there are significant differences in their attribute characteristics. For each voxel unit, compare the differences in attribute characteristics (point density, average coordinates, and geometric center) between the real-time voxel unit and the corresponding background voxel unit, and adopt the following comparison methods: Calculate the difference in point density between the real-time voxel unit and the background voxel unit. If the difference exceeds the preset threshold, it indicates that a new object or dynamic change may have occurred in this area. Calculate the difference in average coordinates between the real-time voxel unit and the background voxel unit. If the difference exceeds the threshold, it means that the position of this area has changed. Judge the change of the central position of the voxel unit through the difference in geometric center. If the change is large, it may indicate the movement of an object or the addition of a new target.
[0129] Set a predefined change threshold. For voxel units whose attribute characteristic differences exceed this threshold, they are determined as "dynamic voxel units", that is, areas with significant changes. Mark all voxel units determined as dynamic voxel units as "second voxel units", that is, areas with significant changes. This dynamic voxel unit represents a possible abnormal dynamic area in the substation, such as the entry of a drone or a person. Perform clustering analysis on the marked dynamic voxel units, and merge spatially adjacent voxel units to form a complete dynamic area. The DBSCAN or connected component labeling algorithm can be used to aggregate nearby dynamic voxel units into a coherent dynamic area.
[0130] Or identify and remove static voxel units to obtain the dynamic area. Specifically, mark the voxel units whose comparison results conform to the static area as "first voxel units", that is, the static area. Remove all first voxel units (static voxel units) from the set of voxel units in the real-time point cloud model, so as to obtain the remaining "second voxel units", that is, the voxel units containing dynamic changes. Cluster the extracted second voxel units and merge spatially adjacent dynamic voxel units. A distance-based clustering algorithm (such as DBSCAN) can be used to aggregate nearby voxel units into a complete dynamic area. The clustered dynamic area is the area where changes have occurred, which may indicate the activities of drones or other objects. Mark the voxel units of the dynamic area in the 3D model for subsequent monitoring and early warning.
[0131] S130, a dynamic reconstruction module based on real-time point cloud data, identifies the first position of the first dynamic target in the dynamic area.
[0132] In a possible implementation, the dynamic reconstruction module based on real-time point cloud data identifies the first position of the first dynamic target in the dynamic region, which specifically includes: determining the relative position between the second voxel unit and the first voxel unit according to the real-time point cloud model; aligning the real-time point cloud model with the background point cloud model through a point cloud registration algorithm; obtaining the actual coordinate position of the background point cloud model; and determining the first position of the first dynamic target according to the relative position and the actual coordinate position.
[0133] Specifically, according to the previously extracted dynamic voxel units (second voxel units) and static voxel units (first voxel units), record their three-dimensional positions and voxel indices respectively. According to the voxel indices of the second voxel unit and the first voxel unit, calculate their relative position differences. For example: Δx = V x2 -V x1 , Δy = V y2 -V y1 , Δz = V z2 -V z1 , where (V x1 , V y1 , V z1 ) is the index position of the first voxel unit (static unit), and (V x2 , V y2 , V z2 ) is the index position of the second voxel unit (dynamic unit).
[0134] Select an appropriate point cloud registration algorithm (such as the ICP algorithm) to align the real-time point cloud model with the background point cloud model. The ICP (Iterative Closest Point) algorithm is commonly used for rigid body registration and can align the models by minimizing the Euclidean distance between two point clouds. Specifically, select an initial alignment method to estimate the initial poses of the real-time point cloud model and the background point cloud model. For the points in the real-time point cloud model, find the points in the background point cloud model that are closest in distance as corresponding points. Iteratively update the rotation and translation parameters to minimize the squared error between the matching point pairs until convergence. Obtain a transformation matrix, which is used to align the real-time point cloud data and the background point cloud data in the same coordinate system.
[0135] According to the preset benchmark of the background point cloud model, obtain the actual coordinate position of each background voxel unit. For example, the background point cloud data can include the three-dimensional spatial coordinates (X, Y, Z) of each static voxel unit for positioning. In the registered unified coordinate system, the coordinate position (X, Y, Z) of each background voxel unit (i.e., the first voxel unit) can be used as a reference benchmark to provide the actual spatial position.
[0136] Using the calculated relative position differences (Δx, Δy, Δz), the actual coordinate positions of the background model are combined with the relative position transformation to determine the actual positions of the dynamic voxel units. Specifically, the actual position (X b , Y b , Z b ) of a background voxel unit is selected as the reference position. According to the relative position differences, the positions of the dynamic voxel units are mapped into the actual coordinate system: X d = X b + Δx·d, Y d = Y b + Δy·d, Z d = Z b + Δz·d where d is the side length of the voxel unit, and (X d , Y d , Z d ) are the actual spatial coordinates of the dynamic target. If the dynamic region contains multiple voxel units, the coordinates of all dynamic voxel units can be averaged to obtain the center position of the first dynamic target.
[0137] S140, calculate the second position of the second dynamic target included in the real-time image.
[0138] Use feature point detection algorithms (such as SIFT, ORB, or FAST) to detect key feature points in the real-time image collected by the camera. Use the SIFT algorithm to detect feature points and calculate descriptors. Through the selected feature points, important positions that may represent dynamic targets (such as drones) can be found in the image. The specific real-time method will be described in the subsequent steps.
[0139] Use object detection algorithms (such as YOLO, Faster R-CNN, etc.) to identify dynamic targets in the real-time image. Extract the bounding box positions of the dynamic targets from the object detection results, which usually include the coordinates of the upper left corner and the lower right corner (or the center coordinates), and then calculate the second dynamic position of the second dynamic target in the real-time image according to the bounding box positions. In this application, calculating the second dynamic position of the second dynamic target in the real-time image is only a conventional technical means in the relevant technical field and will not be further elaborated here.
[0140] S150, if the first position and the second position are successfully matched, and the sparse structure of the real-time point cloud data and the dense pixels of the real-time image are successfully matched, then it is determined that both the first dynamic target and the second dynamic target are dynamic flying targets.
[0141] Ensure that the first position and the second position are in the same coordinate system. If one is in the image coordinate system and the other is in the point cloud coordinate system, they need to be transformed into a unified three-dimensional coordinate system through corresponding transformations (such as the extrinsic matrix). Extract descriptors of feature points (such as SIFT, ORB, FPFH, etc.) near the first position and the second position to obtain the local feature information of the target. For the transformed first position (X 1 , Y 1 , Z 1 ) and the second position (X 2 , Y 2 , Z 2 ), calculate the Euclidean distance between them, and set a suitable distance threshold (such as 0.1 meter) to determine whether these two positions can be considered the same target. If the distance d is less than the threshold, it is considered a successful match.
[0142] In a possible implementation, before successfully matching the sparse structure of the real-time point cloud data with the dense pixels of the real-time image, the method further includes: spatially aligning the lidar sensor and the camera to establish a unified coordinate system; calculating feature descriptors for the first key points in the real-time image to obtain the first descriptors, and calculating feature descriptors for the second key points in the real-time point cloud data to obtain the second descriptors; comparing the first descriptors with the second descriptors to match the first descriptors and the second descriptors whose similarity exceeds a preset threshold; if it is determined that the matching rate of multiple first descriptors reaches a preset condition, or the matching rate of multiple second descriptors reaches a preset condition, then map the coordinates of the first key points to the dynamic reconstruction model to calculate the position of the dynamic flying target.
[0143] Specifically, calibrate the lidar sensor and the camera to determine their relative positions and orientations in three-dimensional space. Usually, some known calibration boards or specific geometric shapes are used to record the acquisition data at different angles and positions. Obtain the internal parameters (focal length, principal point, distortion parameters, etc.) and external parameters (position and attitude relative to the camera) of the lidar sensor to establish the spatial relationship between the two. Use the calibration results to map the image data captured by the camera and the lidar data into a common three-dimensional coordinate system. This can ensure that subsequent data processing (such as feature point matching) is carried out in the same coordinate system, facilitating mutual comparison. Use feature point detection algorithms (such as SIFT, ORB, etc.) in the real-time image to extract the first key points and record their positions in the image. For each extracted first key point, calculate its feature descriptor. Taking SIFT as an example, generate a 128-dimensional descriptor vector.
[0144] In a possible implementation, calculating the feature descriptor for the first key point in the real-time image to obtain the first descriptor specifically includes:
[0145] Downsample the real-time image multiple times and downsample the original image multiple times (for example, halve it each time) to generate a hierarchical structure of images with different resolutions, which can capture features at different scales. For each hierarchical structure of images at a specific resolution, apply Gaussian filtering to the image to smooth the image and generate Gaussian blurred images at different scales. The Gaussian blurred image is represented as follows:
[0146]
[0147] where \(I(x,y)\) is the Gaussian blurred image, \(L(x,y,\sigma)\) is the Gaussian blurred image, and \(G(x,y,\sigma)\) is the Gaussian kernel, represents the convolution operation, and the definition of the Gaussian kernel is:
[0148]
[0149] Calculate the difference between Gaussian blurred images at adjacent scales to generate a Difference of Gaussian (DoG) image. The DoG image is represented as follows:
[0150]
[0151] where \(D(x,y,\sigma)\) is the DoG image, \(L(x,y,\sigma)\) is the Gaussian blurred image, and \(L(x,y,k\sigma)\) is the adjacent Gaussian blurred image, and \(k\) is a constant multiple. By calculating the difference between Gaussian blurred images at adjacent scales, the DoG image is obtained. This processing method can effectively highlight the edges and significant features in the image. The DoG image can help the algorithm identify the parts of the image with obvious changes, thus being more accurate in feature point detection.
[0152] For each pixel in the DoG image, check its local extrema in space and scale. Specifically, compare a pixel point with its 26 adjacent points. If the point is a local extremum (i.e., the value of the point is larger or smaller than the values of its surrounding points), then the point is marked as a candidate key point.
[0153] Compare any pixel point in the DoG image with multiple adjacent pixel points (including the four directions of up, down, left, and right and the eight diagonal directions). If the value of any pixel point in space is greater than or less than the values of multiple adjacent pixel points in space, and the value of any pixel point in scale is greater than or less than the values of multiple adjacent pixel points in space, then any pixel point is marked as a key point, obtaining the first key point. By comparing the value of the current pixel point with the values of adjacent pixel points, local extremum points (key points) can be identified. Local extremum points correspond to significant structures or edges in the image and can represent important features in the image. This step ensures that only important key points are extracted, thereby improving the efficiency and accuracy of subsequent processing.
[0154] Calculate the amplitude and direction of the pixel points in the neighborhood of the first key point, which are specifically calculated through the following formulas:
[0155]
[0156]
[0157] Among them, m(x, y) is the amplitude of the pixel points in the neighborhood of the first key point. The magnitude of the gradient amplitude reflects the degree of gray-scale change of the image at this point, that is, the edge intensity. Generally speaking, points with a larger gradient amplitude appear at the edges or texture-significant regions of the image. θ is the direction of the pixel points in the neighborhood of the first key point. The value range of the gradient direction is [-π, π], indicating the direction angle of the gray-scale change. The direction is usually used to construct a direction histogram in the generation of the feature descriptor, so as to assign a main direction to the key point and make the SIFT feature rotation-invariant. L(x + 1, y) and L(x - 1, y) respectively represent the pixel values of the adjacent pixel points on the left and right of the first key point in the horizontal direction of the difference-of-Gaussians image, and L(x, y + 1) and L(x, y - 1) respectively represent the pixel values of the adjacent pixel points above and below the first key point in the vertical direction of the difference-of-Gaussians image;
[0158] Divide the neighborhood of the first key point into multiple sub-regions of the same size. According to the amplitude and direction of the pixel points in the neighborhood of the key point, calculate the histogram of each sub-region. Usually, a 16×16 neighborhood is divided into 4×4 sub-regions (that is, each sub-region contains 4×4 pixels, a total of 16 sub-regions) for local calculation and construction of the direction histogram. Calculate the direction histogram for each 4×4 sub-region respectively, which can refine the gradient information around the key point and enhance the discrimination of the descriptor.
[0159] Each sub-region generates a histogram with 8 directions (it can also be other numbers of directions, but 8 directions are a commonly used choice) to represent the main direction distribution of the gradient. Classify the gradient direction of each pixel in the neighborhood into the corresponding direction interval (usually divided into 8 directions, each direction corresponding to a 45° angle interval), and add the gradient amplitude m(x, y) as a weight to the histogram count of this direction. Divide the direction into 8 equal intervals, and each interval represents a 45° direction range (the range from 0° to 360° is evenly divided into 8 intervals). Combine multiple histograms to form a multi-dimensional vector to obtain the first descriptor.
[0160] In a possible implementation manner, calculate the feature descriptor for the second key point in the real-time point cloud data to obtain the second descriptor, specifically including:
[0161] Calculate the point curvature of any point in the real-time point cloud data. The point curvature is used to measure the surface curvature of each point in the point cloud. The larger the point curvature, the more significant the geometric structure of the area where the point is located. Therefore, the point curvature can be used to detect key points. The point curvature is calculated by the following formula:
[0162]
[0163] where C is the point curvature, n i is the normal vector of the adjacent points of any point in the point cloud, n is the normal vector of any point in the point cloud, k is the number of adjacent points of any point in the point cloud (define a radius or a fixed number of neighbor points). For any point in the point cloud, calculate its normal vector n. The normal vector is usually estimated through the neighborhood points of the point, for example, by PCA (Principal Component Analysis) to find the normal vector of the local plane.
[0164] Compare the point curvature of each point. If the point curvature value of the point is greater than the preset curvature threshold, mark the point as a key point to obtain the second key points. Selecting an appropriate curvature threshold can effectively reduce the redundant information in the point cloud and only retain the points with higher curvature and rich information as key points.
[0165] Collect all the points that meet the curvature conditions to generate a set containing the second key points. These key points are usually located at the edges of the object or in areas with significant geometric changes and can better describe the geometric features of the point cloud data. Taking the second key point as the center, divide its neighborhood into multiple spherical sub-regions along the direction of its normal vector. The spherical sub-regions are a series of spherical shells (layers) centered on the key point, with each spherical shell having the same thickness, the center of the sphere being the key point, and the direction of the normal vector being used to determine the direction of sub-region division. Further divide the spherical sub-regions in the direction of the normal vector. Usually, the neighborhood of the key point is divided into several sub-regions (such as eight directions) to ensure the rotational invariance of the descriptor and improve the regional resolution at the same time.
[0166] In each spherical sub-region, calculate the direction of the normal vector of the points in it, and statistically analyze the distribution of these directions to generate a direction histogram. Each direction histogram represents the direction distribution of the normal vectors of the points in the spherical sub-region. The direction histogram is usually constructed by dividing the normal vector direction into several direction intervals, and each interval represents a fixed direction range. For example, when divided into 8 intervals, each interval corresponds to an angular range of 45°. In each spherical sub-region, accumulate the normal vector directions of all points to generate a histogram representing the direction distribution of the spherical sub-region. This histogram can effectively characterize the local geometric shape of the region. Accumulate the direction histograms of all spherical sub-regions to form an overall direction histogram. This overall direction histogram contains the characteristics of each spherical sub-region and forms a high-dimensional vector, that is, the second descriptor.
[0167] To reduce the amplitude difference of data, the generated orientation histogram is usually normalized so that the values of the descriptor are within a relatively consistent range. The integrated and normalized orientation histogram forms the feature descriptor of the key point (i.e., the second descriptor). This descriptor not only contains the spatial position and orientation information of the key point, but also reflects the local geometric structure features of the point, making it highly robust in the matching process.
[0168] Further, the first descriptor of the image is compared with the second descriptor of the point cloud, and the Euclidean distance or Hamming distance is used as the similarity metric. Calculate the distance between the first descriptor and the second descriptor. If the distance is less than the preset similarity threshold, it is considered that this pair of descriptors is successfully matched.
[0169] In the first descriptor set and the second descriptor set, count the number of matching pairs that meet the similarity condition (i.e., the distance is less than the threshold). Calculate the matching success rate between the two descriptor sets. The matching rate refers to the proportion of the number of successfully matched feature point pairs to the total number of feature points, which is used to judge the reliability of the matching. If the matching rate of the first descriptor or the matching rate of the second descriptor reaches the preset condition (e.g., 80%), it is determined that the image feature points and the point cloud feature points in this area are successfully matched. At this time, it can be determined that the first dynamic target in the dynamic reconstruction model and the second dynamic target in the real-time image are both dynamic flying targets (e.g., drones).
[0170] S160, Determine the position of the dynamic flying target according to the first position and the second position, and mark it in the three-dimensional model.
[0171] Jointly calibrate the camera and the laser point cloud sensor to obtain their relative position relationship in space, that is, the external parameters (rotation matrix R and translation vector T). During the calibration process, multiple shots need to be taken through a known three-dimensional calibration board or feature marker to accurately calculate the external parameter values. Use the calibration parameters to align the image coordinate system of the camera and the three-dimensional coordinate system of the laser point cloud to ensure that the first position (in the image) and the second position (in the point cloud) can be mapped to the same coordinate system for position matching.
[0172] Set the first position coordinates in the image as (x img , y img ). Using the internal and external parameters of the camera, map the image plane coordinates to the three-dimensional coordinate system to form the three-dimensional position (X img , Y img , Z img ) in the image. The conversion formula is as follows (assuming the depth Z img can be estimated from the point cloud data or the known model):
[0173]
[0174] Among them, R and T are the rotation matrix and translation vector respectively, which are used to convert the image coordinate system to the 3D point cloud coordinate system. The second position (X pc , Y pc , Z pc ) has been obtained in the point cloud, so it is already in the 3D coordinate system. The first position (X img , Y img , Z img ) and the second position (X pc , Y pc , Z pc ) are weighted averaged to reduce the uncertainty of a single data source. The following weighted formula can be used:
[0175] X target = αX img + (1 - α)X pc
[0176] Y target = αY img + (1 - α)Y pc
[0177] Z target = αZ img + (1 - α)Z pc
[0178] Among them, α is the weight coefficient, which can be set according to the accuracy of the image data and the point cloud data. Usually, 0.5 (average fusion) can be taken or adjusted according to the reliability of the actual data source.
[0179] In the 3D reconstruction model, the calculated position of the dynamic flying target is used as a marker point and displayed on the 3D monitoring interface for the operation and maintenance personnel to monitor in real time. When marking, special symbols (such as dots, arrows or drone icons) can be used to highlight the dynamic flying target in the 3D model for easy identification and tracking. If the target moves within the monitoring area, its marked position is updated in real time so that the 3D model dynamically reflects the current position of the target.
[0180] This embodiment also discloses a substation safety control system based on the 3D reconstruction model. Referring to Figure 2 , the system includes an acquisition module 201, an extraction module 202, an identification module 203, a processing module 204 and a judgment module 205, where:
[0181] The acquisition module 201 is used to acquire the real-time point cloud data input by the laser point cloud sensor and the real-time images collected by the camera.
[0182] The extraction module 202 is used to compare the real-time point cloud data with the background point cloud model point by point, and extract the changed dynamic area by using the background difference algorithm.
[0183] The recognition module 203 is used to recognize the first position of the first dynamic target in the dynamic area based on the dynamic reconstruction module performed on the real-time point cloud data.
[0184] The processing module 204 is used to calculate the second position of the second dynamic target included in the real-time image.
[0185] The judgment module 205 is used to determine that both the first dynamic target and the second dynamic target are dynamic flying targets if the first position and the second position are successfully matched, and the sparse structure of the real-time point cloud data and the dense pixels of the real-time image are successfully matched.
[0186] The processing module 204 is used to determine the position of the dynamic flying target according to the first position and the second position, and mark it in the three-dimensional model.
[0187] In a possible implementation manner, the processing module 204 is used to generate a real-time point cloud model according to the real-time point cloud data.
[0188] The processing module 204 is used to perform clustering segmentation on the real-time point cloud model to obtain multiple real-time voxel units, and perform clustering segmentation on the background point cloud model to obtain multiple background voxel units.
[0189] The judgment module 205 is used to compare the multiple real-time voxel units with the multiple background voxel units one by one, and determine the first voxel unit corresponding to the background voxel unit among the multiple real-time voxel units.
[0190] The processing module 204 is used to remove the first voxel unit from the multiple real-time voxel units to obtain the second voxel unit, and the second voxel unit is the dynamic area.
[0191] In a possible implementation manner, the processing module 204 is used to determine the relative position between the second voxel unit and the first voxel unit according to the real-time point cloud model.
[0192] The processing module 204 is used to align the real-time point cloud model with the background point cloud model through the point cloud registration algorithm.
[0193] The acquisition module 201 is used to acquire the actual coordinate position of the background point cloud model.
[0194] The judgment module 205 is used to determine the first position of the first dynamic target according to the relative position and the actual coordinate position.
[0195] In a possible implementation, the processing module 204 is configured to spatially align the laser point cloud sensor and the camera to establish a unified coordinate system.
[0196] The processing module 204 is configured to calculate feature descriptors for the first key points in the real-time image to obtain first descriptors, and calculate feature descriptors for the second key points in the real-time point cloud data to obtain second descriptors.
[0197] The judgment module 205 is configured to compare the first descriptors with the second descriptors and match the first descriptors and the second descriptors whose similarity exceeds a preset threshold.
[0198] The recognition module 203 is configured to, if it is determined that the matching rate of multiple first descriptors reaches a preset condition, or the matching rate of multiple second descriptors reaches a preset condition, map the coordinates of the first key points to the dynamic reconstruction model to calculate the position of the dynamic flying target.
[0199] In a possible implementation, the processing module 204 is configured to subdivide the real-time point cloud model into multiple first voxel points according to a predefined side length dimension, and subdivide the background point cloud model into multiple second voxel points.
[0200] The processing module 204 is configured to calculate the attribute features of each real-time voxel unit according to the geometric data of the first voxel points included in the real-time voxel unit. The attribute features include point density, average coordinates, and geometric center.
[0201] The processing module 204 is configured to calculate the attribute features of each background voxel unit according to the geometric data of the second voxel points included in the background voxel unit. The attribute features include point density, average coordinates, and geometric center.
[0202] The judgment module 205 is configured to compare each real-time voxel unit with the background voxel units one by one, determine the dynamic voxel units among the multiple real-time voxel units whose change values of the attribute features exceed a preset standard, and determine the dynamic voxel units as dynamic regions.
[0203] In a possible implementation, the processing module 204 is configured to downsample the real-time image multiple times to generate a series of image hierarchical structures with different resolutions. For each image hierarchical structure with a resolution, apply Gaussian filtering to the image to smooth the image and generate Gaussian blurred images at different scales. The Gaussian blurred images are represented as follows:
[0204]
[0205] where I(x, y) is the Gaussian blurred image, L(x, y, σ) is the Gaussian blurred image, G(x, y, σ) is the Gaussian kernel, represents the convolution operation.
[0206] The processing module 204 is configured to calculate the difference between Gaussian blurred images of adjacent scales to generate a difference Gaussian image, and the difference Gaussian image is represented as follows:
[0207]
[0208] Where D(x, y, σ) is the difference Gaussian image, L(x, y, σ) is the Gaussian blurred image, L(x, y, kσ) is the adjacent Gaussian blurred image of the Gaussian blurred image, and k is a constant multiple.
[0209] The processing module 204 is configured to compare any pixel point in the difference Gaussian image with multiple adjacent pixel points. If the value of any pixel point in space is greater than or less than the values of multiple adjacent pixel points in space, and the value of any pixel point in scale is greater than or less than the values of multiple adjacent pixel points in space, then any pixel point is marked as a key point to obtain the first key point.
[0210] The processing module 204 is configured to calculate the amplitude and direction of the pixel points in the neighborhood of the first key point, specifically calculated by the following formula:
[0211]
[0212]
[0213] Where m(x, y) is the amplitude of the pixel points in the neighborhood of the first key point, θ is the direction of the pixel points in the neighborhood of the first key point, L(x + 1, y) and L(x - 1, y) respectively represent the pixel values of the adjacent pixel points on the left and right of the first key point in the horizontal direction of the difference Gaussian image, and L(x, y + 1) and L(x, y - 1) respectively represent the pixel values of the adjacent pixel points above and below the first key point in the vertical direction of the difference Gaussian image.
[0214] The processing module 204 is configured to divide the neighborhood of the first key point into multiple sub-regions of the same size, and calculate the histogram of each sub-region according to the amplitude and direction of the pixel points in the neighborhood of the key point.
[0215] The processing module 204 is configured to merge multiple histograms to form a multi-dimensional vector to obtain the first descriptor.
[0216] In a possible implementation manner, the processing module 204 is configured to calculate the point curvature of any point cloud point corresponding to the real-time point cloud data, and the point curvature is calculated by the following formula:
[0217]
[0218] Where C is the point curvature, n iis the normal vector of the adjacent points of any point cloud point, n is the normal vector of any point cloud point, and k is the number of adjacent points of any point cloud point.
[0219] The processing module 204 is configured to determine any point cloud point as a key point if it is determined that the curvature of any point cloud point is greater than a preset curvature, and obtain a second key point.
[0220] The processing module 204 is configured to divide the neighborhood into multiple spherical sub-regions with the second key point as the center according to the direction of the normal vector of the second key point.
[0221] The processing module 204 is configured to calculate a normal vector direction histogram within each spherical sub-region, accumulate the features of the spherical sub-region, generate an overall direction histogram, and the data of the direction histogram is a second descriptor.
[0222] It should be noted that when the device provided in the above embodiment realizes its functions, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0223] This embodiment also discloses an electronic device. Referring to Figure 3 , the electronic device may include: at least one processor 301, at least one communication bus 302, a user interface 303, a network interface 304, and at least one memory 305.
[0224] Among them, the communication bus 302 is used to realize the connection and communication between these components.
[0225] Among them, the user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.
[0226] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0227] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server through various interfaces and lines, and executes various functions of the server and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling the data stored in the memory 305. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately through a single chip.
[0228] Among them, the memory 305 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store the data involved in the above-mentioned method embodiments. Optionally, the memory 305 may further be at least one storage device located far from the aforementioned processor 301. The memory 305, as a computer storage medium, may include an operating system, a network communication module, a user interface 303 module, and an application program for a substation safety control method based on a three-dimensional reconstruction model.
[0229] In Figure 3In the electronic device shown, the user interface 303 is mainly used to provide an interface for the user to input and obtain the data input by the user; and the processor 301 can be used to call the application program stored in the memory 305 for a substation safety control method based on a three-dimensional reconstruction model. When executed by one or more processors 301, the electronic device is caused to execute the method of one or more of the above embodiments.
[0230] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0231] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0232] In several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0233] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0234] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0235] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a memory 305 and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned memory 305 includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.
[0236] The present application also discloses a computer-readable storage medium storing instructions. When executed by one or more processors 301, it causes the electronic device to execute the method as described in one or more of the above embodiments.
[0237] The above are only exemplary embodiments of the present disclosure, and the scope of the present disclosure cannot be limited thereby. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the specification and the disclosure of the practical truth, those skilled in the art will easily think of other implementation schemes of the present disclosure. The present application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A substation safety management and control method based on a three-dimensional reconstruction model, characterized in that: The method comprises: Obtain real-time point cloud data input by the laser point cloud sensor and real-time images collected by the camera; Comparing the real-time point cloud data with the background point cloud model point by point, and extracting the changed dynamic area using a background difference algorithm; A dynamic reconstruction module based on the real-time point cloud data identifies a first position of a first dynamic target in the dynamic area; Calculating a second position of a second dynamic target included in the real-time image; If the first position is successfully matched with the second position, and the sparse structure of the real-time point cloud data is successfully matched with the dense pixels of the real-time image, it is determined that the first dynamic target and the second dynamic target are both dynamic flying targets; Determine the position of the dynamic flying target according to the first position and the second position, and mark it in the three-dimensional model; Before successfully matching the sparse structure of the real-time point cloud data with the dense pixels of the real-time image, the method further includes: Spatially align the laser point cloud sensor and the camera to establish a unified coordinate system; Calculating a feature descriptor for a first key point in the real-time image to obtain a first descriptor, and calculating a feature descriptor for a second key point in the real-time point cloud data to obtain a second descriptor; comparing the first descriptor with the second descriptor, and matching the first descriptor with the second descriptor whose similarity exceeds a preset threshold; If it is determined that the matching rate of the plurality of first descriptors reaches a preset condition, or the matching rate of the plurality of second descriptors reaches a preset condition, the coordinates of the first key point are mapped to the dynamic reconstruction model to calculate the position of the dynamic flying target.
2. A substation safety management and control method based on a three-dimensional reconstruction model according to claim 1, characterized in that: The real-time point cloud data is compared point by point with the background point cloud model, and the changed dynamic area is extracted using a background difference algorithm, specifically including: Generate a real-time point cloud model according to the real-time point cloud data; Clustering and segmenting the real-time point cloud model to obtain a plurality of real-time voxel units, and clustering and segmenting the background point cloud model to obtain a plurality of background voxel units; Compare the plurality of real-time voxel units with the plurality of background voxel units one by one, and determine a first voxel unit corresponding to the background voxel unit among the plurality of real-time voxel units; A first voxel unit among the plurality of real-time voxel units is removed to obtain a second voxel unit, where the second voxel unit is the dynamic area.
3. A substation safety management and control method based on a three-dimensional reconstruction model according to claim 2, characterized in that: The dynamic reconstruction module based on the real-time point cloud data identifies a first position of a first dynamic target in the dynamic area, specifically comprising: Determining a relative position of the second voxel unit and the first voxel unit according to the real-time point cloud model; Aligning the real-time point cloud model with the background point cloud model through a point cloud registration algorithm; Obtaining the actual coordinate position of the background point cloud model; The first position of the first dynamic target is determined according to the relative position and the actual coordinate position.
4. A substation safety management and control method based on a three-dimensional reconstruction model according to claim 2, characterized in that: The step of comparing the real-time point cloud data with the background point cloud model point by point and extracting the changed dynamic area using a background difference algorithm specifically includes: According to a predefined side length size, subdividing the real-time point cloud model into a plurality of first voxel points and subdividing the background point cloud model into a plurality of second voxel points; Calculate the attribute characteristics of each real-time voxel unit according to the geometric data of the first voxel point contained in the real-time voxel unit, wherein the attribute characteristics include point density, average coordinates and geometric center; Calculate the attribute characteristics of each background voxel unit according to the geometric data of the second voxel point contained in the background voxel unit, wherein the attribute characteristics include point density, average coordinates and geometric center; The real-time voxel units are compared with the background voxel units one by one, and dynamic voxel units whose attribute feature change values exceed a preset standard among the multiple real-time voxel units are determined, and the dynamic voxel units are determined to be the dynamic areas.
5. A substation safety management and control method based on a three-dimensional reconstruction model according to claim 1, characterized in that: Calculating a feature descriptor for a first key point in the real-time image to obtain a first descriptor specifically includes: The real-time image is downsampled multiple times to generate a series of image hierarchical structures with different resolutions. For each image hierarchical structure with different resolutions, Gaussian filtering is applied to the image to smooth the image and generate Gaussian blurred images of different scales. The Gaussian blurred images are represented as follows: ; Wherein, I(x,y) is the Gaussian blurred image, L(x,y,σ) is the Gaussian blurred image, G(x,y,σ) is the Gaussian kernel, Represents the convolution operation; The difference between Gaussian blurred images of adjacent scales is calculated to generate a differential Gaussian image, which is represented as follows: ; Wherein, D(x, y, σ) is a differential Gaussian image, L(x, y, σ) is the Gaussian blurred image, L(x, y, kσ) is a Gaussian blurred image adjacent to the Gaussian blurred image, and k is a constant multiple; Compare any pixel point in the differential Gaussian image with multiple adjacent pixel points. If the spatial value of any pixel point is greater than or less than the spatial values of multiple adjacent pixel points, and the scale value of any pixel point is greater than or less than the spatial values of multiple adjacent pixel points, then the any pixel point is marked as a key point, and the first key point is obtained. The amplitude and direction of the pixel points in the neighborhood of the first key point are calculated, specifically by the following formula: ; ; Wherein, m(x, y) is the amplitude of the pixel points in the neighborhood of the first key point, θ is the direction of the pixel points in the neighborhood of the first key point, L(x+1, y) and L(x-1, y) respectively represent the pixel values of the left and right adjacent pixels of the first key point in the horizontal direction of the differential Gaussian image, and L(x, y+1) and L(x, y-1) respectively represent the pixel values of the upper and lower adjacent pixels of the first key point in the vertical direction of the differential Gaussian image; Divide the area of the first key point into a plurality of sub-areas of the same size, and calculate a histogram of each of the sub-areas according to the amplitude and direction of the pixel points in the neighborhood of the key point; The multiple histograms are combined to form a multi-dimensional vector to obtain the first descriptor.
6. A substation safety management and control method based on a three-dimensional reconstruction model according to claim 1, characterized in that: The calculating a feature descriptor for the second key point in the real-time point cloud data to obtain a second descriptor specifically includes: Calculate the point curvature of any point cloud point corresponding to the real-time point cloud data, and the point curvature is calculated by the following formula: ; Where C is the curvature of the point, n i is the normal vector of the neighboring points of any one of the point cloud points, n is the normal vector of any one of the point cloud points, and k is the number of neighboring points of any one of the point cloud points; If it is determined that the curvature of any one of the point cloud points is greater than a preset curvature, the any one of the point cloud points is determined as a key point, and the second key point is obtained; Taking the second key point as the center, dividing the neighborhood into a plurality of spherical cells according to the direction of the normal vector of the second key point; A normal vector direction histogram is calculated in each of the spherical cells, and features of the spherical cells are accumulated to generate an overall direction histogram, wherein data of the direction histogram is the second descriptor.
7. A substation safety management and control system based on a three-dimensional reconstruction model, characterized in that: The system comprises an acquisition module (201), an extraction module (202), an identification module (203), a processing module (204) and a judgment module (205), wherein: The acquisition module (201) is used to acquire real-time point cloud data input by the laser point cloud sensor and real-time images collected by the camera; The extraction module (202) is used to compare the real-time point cloud data with the background point cloud model point by point, and extract the dynamic area that has changed using a background difference algorithm; The identification module (203) is used for identifying the first position of the first dynamic target in the dynamic area based on the dynamic reconstruction module of the real-time point cloud data; The processing module (204) is used to calculate a second position of a second dynamic target included in the real-time image; The judgment module (205) is used to determine that the first dynamic target and the second dynamic target are both dynamic flying targets if the first position is successfully matched with the second position and the sparse structure of the real-time point cloud data is successfully matched with the dense pixels of the real-time image; The processing module (204) is used to determine the position of the dynamic flying target according to the first position and the second position, and mark it in the three-dimensional model; The processing module (204) is used to spatially align the laser point cloud sensor and the camera to establish a unified coordinate system; The processing module (204) is used to calculate a feature descriptor for a first key point in the real-time image to obtain a first descriptor, and calculate a feature descriptor for a second key point in the real-time point cloud data to obtain a second descriptor; The processing module (204) is used to compare the first descriptor with the second descriptor, and match the first descriptor and the second descriptor whose similarity exceeds a preset threshold; The processing module (204) is used to map the coordinates of the first key point to the dynamic reconstruction model to calculate the position of the dynamic flying target if it is determined that the matching rate of the plurality of the first descriptors reaches a preset condition, or the matching rate of the plurality of the second descriptors reaches a preset condition.
8. An electronic device, characterized in that: The electronic device comprises a processor (301), a communication bus (302), a user interface (303), a network interface (304) and a memory (305), wherein the memory (305) is used to store instructions, the user interface (303) and the network interface (304) are both used to communicate with other devices, and the processor (301) is used to execute the instructions stored in the memory (305) so that the electronic device executes the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 6 is performed.
Citation Information
Patent Citations
Three-dimensional digital modeling system and method based on real scene video
CN118628663A