Alarm method and device based on intrusion distance detection
By combining two-dimensional images and three-dimensional point clouds, an intrusion detection method is used to calculate the three-dimensional distance between the intruder and the power equipment. This solves the problems of insufficient accuracy and high cost in traditional methods, and realizes efficient and low-cost three-dimensional distance alarm.
Patent Information
- Application Number
- CN202511241280.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-11-28
AI Technical Summary
Traditional intrusion detection methods cannot accurately calculate the three-dimensional physical distance between the intrusion target and critical facilities, resulting in a high false alarm rate, high cost, and insufficient real-time performance and adaptability.
Based on the intrusion detection model and a pre-established pixel point cloud map, the three-dimensional distance between the intruder and power-related equipment is calculated by combining two-dimensional images with three-dimensional point clouds, and alarm information is generated.
It achieves 3D distance calculation with centimeter-level accuracy, reduces hardware costs, improves real-time performance and adaptability, reduces computational load, and overcomes the shortcomings of traditional methods.
Smart Images

Figure CN121033992A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of power warning, in particular to an alarm method and device based on intrusion distance detection. BACKGROUND
[0002] Intrusion detection is widely used in security and industry fields, but pure intrusion alarm cannot meet the fine needs. Key scenes (such as substations and dangerous operation areas) need to accurately perceive the three-dimensional physical distance (such as "distance from equipment <1 meter") between the intrusion target and the key facility, so as to realize risk grading warning and improve the safety protection efficiency.
[0003] Traditional pure visual two-dimensional intrusion detection method mainly relies on ordinary cameras and image analysis algorithms. The fundamental defect is the lack of depth perception ability, which can only be judged in the two-dimensional image plane. Since the real three-dimensional information cannot be obtained, the system cannot calculate the actual position of the target in the physical world and the accurate distance (meter / centimeter level) to the key reference point. The alarm is greatly affected by perspective distortion and is sensitive to environmental light changes, with high false alarm rate.
[0004] Real-time radar detection (millimeter wave / laser radar) method, due to the dependence on high-cost special sensors, results in high system cost. In addition, the deployment is complex (requires precise calibration), and the performance of laser radar decreases in rainy and foggy weather, and the real-time point cloud processing calculation burden is heavy.
[0005] Real-time depth camera (RGB-D) detection method, due to the dependence on active light source, results in the failure of depth measurement in outdoor strong light. Moreover, the effective ranging range and field of view angle are limited, and the cost and power consumption are higher than ordinary cameras.
[0006] Real-time vision + real-time point cloud fusion detection method, due to the need for dual sensor cooperation, results in the highest system cost and complexity. Furthermore, the development of multi-source data space-time synchronization and fusion algorithm is difficult, and the data processing amount is large and the real-time performance will be affected. SUMMARY
[0007] The present disclosure at least provides an alarm method and device based on intrusion distance detection to solve the problems of poor ranging effect of traditional pure visual scheme on intrusion objects, high cost and accuracy decline in rainy weather of real-time radar detection method, small ranging range of real-time depth camera detection method, low real-time performance and large calculation amount of real-time vision + real-time point cloud fusion detection method.
[0008] According to an aspect of the present disclosure, an alarm method based on intrusion distance detection is provided, comprising:
[0009] Obtaining a pixel point cloud map of a target area which is established in advance; the target area includes power-related equipment;
[0010] acquire a to-be-detected image in the target region, and input the to-be-detected image into the trained intruder detection model, so that the intruder detection model processes the to-be-detected image and outputs information of an intruder detection frame;
[0011] determine a to-be-matched point according to the information of the intruder detection frame; the to-be-matched point includes a left lower corner pixel point and a right lower corner pixel point of the intruder detection frame;
[0012] screen a matched target three-dimensional point in the pixel point cloud map based on the two-dimensional coordinates of the to-be-matched point;
[0013] determine a three-dimensional estimated height of the intruder based on the height of the intruder detection frame and the internal parameters of the camera used to shoot the to-be-detected image;
[0014] draw a three-dimensional point enclosing line from the target three-dimensional point in a vertical direction according to a preset step size and the three-dimensional estimated height, draw a three-dimensional point enclosing line between last extension points of the target three-dimensional point according to a preset step size and the three-dimensional estimated height, form an intruder three-dimensional frame by using the drawn three-dimensional point enclosing line, and wherein the last extension point is the last point drawn from the target three-dimensional point in a vertical direction according to a preset step size and the three-dimensional estimated height, and the height of the last point is equal to or higher than the three-dimensional estimated height;
[0015] determine the distance between the intruder and the power-related equipment by using the intruder three-dimensional frame, and form an alarm information by using the distance;
[0016] wherein, the pixel point cloud map is generated according to the following steps:
[0017] transform a historical point cloud in a historical three-dimensional point cloud coordinate system into a camera coordinate system;
[0018] screen a target historical point cloud located in a camera field of view according to a camera horizontal field of view, a camera vertical field of view, and coordinates of the historical point cloud in the camera coordinate system;
[0019] perform normalization processing and distortion correction processing on the coordinates of each target historical point cloud in the camera coordinate system, perform pixel coordinate mapping by using the corrected coordinates and the internal parameters of the camera, and obtain the coordinates of each target historical point cloud in a pixel point cloud coordinate system; and calculate the depth of each target historical point cloud by using the coordinates of each target historical point cloud in the camera coordinate system.
[0020] map each target historical point cloud into the pixel point cloud map according to the coordinates of each target historical point cloud in the pixel point cloud coordinate system in ascending order of depth.
[0021] In a possible implementation, the transforming the historical point cloud located in the historical three-dimensional point cloud coordinate system into the camera coordinate system comprises:
[0022] determining a rigid transformation matrix based on the camera position and the camera pose; and transforming the historical point cloud located in the historical three-dimensional point cloud coordinate system into the camera coordinate system by using the rigid transformation matrix and the camera pose.
[0023] In a possible implementation, the screening the target historical point cloud located in the camera field of view according to the camera horizontal field of view angle, the camera vertical field of view angle and the coordinates of the historical point cloud in the camera coordinate system comprises:
[0024] the historical three-dimensional point cloud satisfying the following three conditions is regarded as the target historical point cloud: ; ; ;
[0025] wherein, represents the coordinates of the historical three-dimensional point cloud in the camera coordinate system; represents the camera horizontal field of view angle; represents the camera vertical field of view angle.
[0026] In a possible implementation, the target historical point cloud is subjected to a distortion correction process, comprising:
[0027] calculating the radial distance and distortion of the target historical point cloud;
[0028] calculating the radial distortion by using the radial distance and distortion and the camera intrinsic parameter;
[0029] calculating the tangential distortion by using the camera intrinsic parameter and the coordinates of the target historical point cloud after normalization processing;
[0030] determining the coordinates of the target historical point cloud after correction by using the coordinates of the target historical point cloud after normalization processing, the radial distortion and the tangential distortion.
[0031] In a possible implementation, the radial distance and distortion of the target historical point cloud are calculated by using the following formula: ;
[0032] wherein, r represents the radial distance and distortion, (x n ,y n ) represents the coordinates of the target historical point cloud after normalization.
[0033] In a possible implementation, the radial distortion is calculated by using the following formula:
[0034]
[0035] wherein, denotes radial distortion, k1, k2 denote camera intrinsic parameters, and r denotes radial distance and distortion.
[0036] In one possible implementation, tangential distortion is calculated using the following formula: ;
[0037] wherein, r denotes radial distance and distortion, (x n ,y n ) denotes normalized coordinates of the target historical point cloud, and p1, p2 denote camera intrinsic parameters. 、 denotes tangential distortion.
[0038] In one possible implementation, the corrected coordinates of the target historical point cloud are determined using the following formula: ;
[0039] wherein, (x d ,y d ) denotes the corrected coordinates of the target historical point cloud, (x n ,y n ) denotes normalized coordinates of the target historical point cloud, denotes radial distortion, 、 denotes tangential distortion.
[0040] In one possible implementation, pixel coordinate mapping is performed using the following formula: ; ; ; ;
[0041] wherein, denotes coordinates of the target historical point cloud in the pixel point cloud coordinate system; f x , f y , c x , c y denote camera intrinsic parameters, (x d ,y d ) denotes the corrected coordinates of the target historical point cloud, and W, H denote the width and height of a set image size.
[0042] According to another aspect of the present disclosure, there is provided an alarm device based on intrusion distance detection, comprising:
[0043] A pixel point cloud map determination module is configured to acquire a pixel point cloud map of a target region, which is established in advance; the target region includes power-related equipment.
[0044] A two-dimensional detection module is configured to acquire a to-be-detected image in the target region, and input the to-be-detected image into a trained intruder detection model, so that the intruder detection model processes the to-be-detected image and outputs information of an intruder detection frame.
[0045] A matching point processing module is configured to determine a to-be-matched point according to the information of the intruder detection frame; the to-be-matched point includes a left lower corner pixel point and a right lower corner pixel point of the intruder detection frame; and based on the two-dimensional coordinates of the to-be-matched point, a matching target three-dimensional point is screened in the pixel point cloud map.
[0046] A height estimation module is configured to determine a three-dimensional estimated height of the intruder based on a height of the intruder detection frame and an intrinsic parameter of a camera for shooting the to-be-detected image.
[0047] A three-dimensional frame generation module is configured to draw a three-dimensional point enclosing line from the target three-dimensional point along a vertical direction according to a preset step size and the three-dimensional estimated height; draw a three-dimensional point enclosing line between last extension points of the target three-dimensional point according to a preset step size and the three-dimensional estimated height; form an intruder three-dimensional frame by using the drawn three-dimensional point enclosing line; wherein the last extension point is a last point for dotting along the vertical direction according to the preset step size and the three-dimensional estimated height from the target three-dimensional point, and the height of the last point is equal to or higher than the three-dimensional estimated height.
[0048] An alarm module is configured to determine a distance between the intruder and the power-related equipment by using the intruder three-dimensional frame, and form alarm information by using the distance.
[0049] The pixel point cloud map determination module is further configured to perform the following steps to generate the pixel point cloud map:
[0050] transform a historical point cloud located in a historical three-dimensional point cloud coordinate system into a camera coordinate system;
[0051] screen a target historical point cloud located within a camera field of view according to a camera horizontal field of view, a camera vertical field of view, and coordinates of the historical point cloud in the camera coordinate system;
[0052] perform normalization processing and distortion correction processing on the coordinates of each target historical point cloud in the camera coordinate system, and perform pixel coordinate mapping by using the corrected coordinates and an intrinsic parameter of the camera, to respectively obtain coordinates of each target historical point cloud in a pixel point cloud coordinate system; and calculate the depth of each target historical point cloud by using the coordinates of each target historical point cloud in the camera coordinate system.
[0053] In ascending order of depth, each target's historical point cloud is mapped onto the pixel point cloud map based on its coordinates in the pixel point cloud coordinate system.
[0054] This disclosure discloses an alarm method and apparatus based on intrusion distance detection. It filters historical point clouds of targets within the camera's field of view based on camera pose and field of view, and performs operations such as normalization, distortion correction, and pixel coordinate mapping on the historical point clouds to generate a pixel point cloud map. This map only needs to be generated once and can be used for subsequent distance detection in the same area without regeneration. Combining this map with the detection bounding boxes obtained from 2D image detection, the distance between the intruder and power-related equipment can be determined. This not only achieves distance-based intrusion alarms based on real-time images and static historical 3D point clouds, but also utilizes the offline constructed pixel point cloud map as a spatial reference to achieve centimeter-level distance calculations, overcoming the lack of depth in pure vision and providing accurate 3D distance. Secondly, this disclosed solution does not require real-time radar or depth cameras, only a regular camera + static historical point clouds, with hardware costs approaching those of pure vision solutions, significantly reducing costs. Furthermore, regular cameras are unaffected by lighting conditions, avoiding the outdoor failure problem of RGB-D cameras and exhibiting strong adaptability. Additionally, this disclosed solution eliminates real-time point cloud acquisition / processing and complex fusion steps, providing high real-time performance and effectively reducing computational load.
[0055] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0056] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0057] Figure 1 This is one of the flowcharts of the alarm method based on intrusion distance detection in the embodiments of this disclosure;
[0058] Figure 2 This is the second flowchart of the alarm method based on intrusion distance detection in this embodiment of the present disclosure;
[0059] Figure 3 This is a schematic diagram of the three-dimensional bounding box of the intruder in an embodiment of this disclosure;
[0060] Figure 4 This is a schematic diagram of the alarm device based on intrusion distance detection in an embodiment of this disclosure. Detailed Implementation
[0061] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0062] This disclosure addresses the problems in intrusion detection, such as the poor ranging performance of traditional pure vision solutions, the high cost and decreased accuracy of real-time radar detection methods in rainy weather, the small ranging range of real-time depth camera detection methods, and the low real-time performance of real-time vision + real-time point cloud fusion detection methods. It proposes an alarm method and device based on intrusion distance detection. The technical solution of this disclosure not only achieves distance-based intrusion alarms based on real-time images and static historical 3D point clouds, but also utilizes offline constructed pixel point cloud maps as spatial references to achieve centimeter-level distance calculations, overcoming the lack of depth in pure vision and providing accurate 3D distance. Furthermore, it eliminates the need for real-time radar or depth cameras, requiring only a regular camera and static historical point clouds, resulting in hardware costs close to those of pure vision solutions, significantly reducing costs. Moreover, regular cameras are unaffected by lighting conditions, avoiding the outdoor failure problem of RGB-D cameras and exhibiting strong adaptability. Additionally, the solution of this disclosure eliminates the real-time point cloud acquisition / processing and complex fusion steps, providing high real-time performance and effectively reducing computational load.
[0063] The technical solution of this disclosure will be described below through specific embodiments.
[0064] like Figure 1 The diagram shown is a flowchart of the alarm method based on intrusion distance detection in this embodiment. The execution subject of this embodiment is a computing device or component with data processing capabilities. Specifically, the method of this embodiment may include the following steps:
[0065] S110. Obtain a pre-built pixel cloud map of the target area; the target area includes power-related equipment.
[0066] Pixel cloud maps can be generated before intrusion detection is performed, or they can be generated during the first intrusion detection.
[0067] like Figure 2 As shown, it can be generated using the following steps:
[0068] Historical point clouds of targets within the camera's field of view are filtered based on camera pose and FOV, and pixel point cloud maps are generated based on camera intrinsic parameters.
[0069] S120. Obtain the image to be detected within the target area, and input the image to be detected into the trained intrusion detection model so that the intrusion detection model processes the image to be detected and outputs the information of the intrusion detection box.
[0070] The intrusion detection model uses a pre-trained YOLOv13 network. The information of the intrusion detection bounding box can be composed of x, y, w, h, etc., where x and y represent the pixel coordinates of the top-left corner of the target box on the image, and w and h represent the width and height of the target box, respectively.
[0071] S130. Based on the information of the intruder detection box, determine the matching point; the matching point includes the lower left and lower right pixel points of the intruder detection box; based on the two-dimensional coordinates of the matching point, filter the matching target three-dimensional points in the pixel cloud map.
[0072] If there are multiple target 3D points, then the first 3D point or the 3D point with the smallest depth d is taken as the target 3D point.
[0073] The map(v',u') function stores a set of 3D points; see the construction process of a pixel point cloud map. Therefore, the matched 3D point map(x,y+h) will yield a set of 3D points (one or more), and map(x+w,y+h) will also yield a set of 3D points (one or more).
[0074] S140. Based on the height of the intruder detection frame and the intrinsic parameters of the camera that captured the image to be detected, determine the estimated three-dimensional height of the intruder.
[0075] The estimated three-dimensional height of the intruder is determined using the following formula. : ;
[0076] S150. Starting from the target three-dimensional point, draw a three-dimensional point enclosing line along the vertical direction according to a preset step length and a three-dimensional estimated height; draw a three-dimensional point enclosing line between the last extension points of the target three-dimensional point according to a preset step length and a three-dimensional estimated height; use the drawn three-dimensional point enclosing line to form an intruder three-dimensional frame.
[0077] The final extension point is the last point marked along the vertical direction from the target three-dimensional point, according to a preset step size and a three-dimensional estimated height. The height of this last point is equal to or higher than the three-dimensional estimated height.
[0078] Specifically, P is the target 3D point corresponding to the pixel point at the lower left corner of the intruder. objL The target 3D point P corresponding to the bottom right pixel objRBegin by drawing a 3D point enclosing line in the upward direction from the ground, with a step size of 0.1m and a height of [height]. Then, [P]... objL and P objR Between the final extension points, bounding lines are drawn around the three-dimensional points with a step size of 0.1m, ultimately forming the three-dimensional bounding box of the intruder, as shown below. Figure 3 As shown. One of the final extension points is from P. objL The starting point is marked with points in the direction of "ground upwards, step size 0.1m, length height". The height of this last point is equal to or higher than the specified height. One of the final extension points is from point P. objR The starting point is marked with the following steps: "ground upwards, step size 0.1m, length height". The height of the last point is equal to or higher than the height.
[0079] S160. Using the three-dimensional bounding box of the intruder, determine the distance between the intruder and the power-related equipment, and use the distance to generate alarm information.
[0080] The distance from the point set corresponding to the 3D bounding box of the intruder to the point cloud of power-related equipment (such as the point cloud of energized targets / power lines / electronic fences) is calculated as the alarm distance.
[0081] The above pixel point cloud map was generated according to the following steps:
[0082] Step 1: Transform the historical point cloud located in the historical 3D point cloud coordinate system to the camera coordinate system.
[0083] The camera coordinate system is as follows:
[0084] Origin: Position of the camera's optical center ;
[0085] Z-axis: Camera optical axis direction (pointing towards the shooting direction);
[0086] X-axis: Right side of the camera;
[0087] Y-axis: Downward direction of the camera (right-hand coordinate system);
[0088] Step one may specifically include:
[0089] Based on the camera position and camera orientation, a rigid body transformation matrix is determined; using the rigid body transformation matrix and camera orientation, the historical point cloud located in the historical 3D point cloud coordinate system is transformed to the camera coordinate system.
[0090] Camera location: Camera pose: rotation matrix (Quaternions or Euler angles), the rigid body transformation matrix is: ;
[0091] Where the translation vector is: ;
[0092] The historical point cloud, located in the historical 3D point cloud coordinate system, is transformed to the camera coordinate system using the following formula: ;
[0093] The coordinates of the point in the historical 3D point cloud coordinate system. : The coordinates of the point in the camera coordinate system.
[0094] Step 2: Based on the camera's horizontal field of view, camera's vertical field of view, and the coordinates of the historical point cloud in the camera coordinate system, filter the target historical point cloud located within the camera's field of view.
[0095] The historical 3D point cloud that meets the following three conditions will be used as the target historical point cloud: ; ; ;
[0096] in, Represents the coordinates of the historical 3D point cloud in the camera coordinate system; Indicates the camera's horizontal field of view; This indicates the camera's vertical field of view (in radians).
[0097] That is, the target historical point cloud : ;
[0098] Steps one and two above extracted the historical point cloud within the camera's field of view.
[0099] In the following steps, the camera intrinsic parameter is: f x , f y , c x , c y , k1, k2, p1, p2. W and H are the width and height of the image size set manually, and the camera intrinsic parameters are also at this scale.
[0100] Step 3: Normalize and correct the coordinates of each target historical point cloud in the camera coordinate system, and use the corrected coordinates and camera intrinsic parameters to perform pixel coordinate mapping to obtain the coordinates of each target historical point cloud in the pixel point cloud coordinate system; and use the coordinates of each target historical point cloud in the camera coordinate system to calculate the depth of each target historical point cloud.
[0101] The normalization formula is as follows:
[0102] For point : ; ;
[0103] Distortion correction includes the following steps:
[0104] 1) Calculate the radial distance and distortion of the target's historical point cloud; use the following formula to calculate the radial distance and distortion of the target's historical point cloud: ;
[0105] In the formula, r represents the radial distance and distortion, (x n ,y n () represents the normalized coordinates of the target historical point cloud.
[0106] 2) Calculate the radial distortion using the radial distance, distortion, and camera intrinsic parameters;
[0107] Radial distortion is calculated using the following formula: ;
[0108] In the formula, The value represents radial distortion, k1 and k2 represent camera intrinsic parameters, and r represents radial distance and distortion.
[0109] 3) Calculate tangential distortion using camera intrinsic parameters and the coordinates of the target's historical point cloud after normalization.
[0110] Tangential distortion is calculated using the following formula: ;
[0111] In the formula, r represents the radial distance and distortion, (x n ,y n () represents the normalized coordinates of the target's historical point cloud, and p1 and p2 represent the camera's intrinsic parameters. , This indicates tangential distortion.
[0112] 4) Determine the corrected coordinates of the target's historical point cloud using the normalized coordinates, radial distortion, and tangential distortion.
[0113] The corrected coordinates of the target's historical point cloud are determined using the following formula: ;
[0114] In the formula, (x d ,y d(x) represents the corrected coordinates of the target's historical point cloud. n ,y n () represents the coordinates of the target historical point cloud after normalization. Indicates radial distortion. , This indicates tangential distortion.
[0115] After distortion correction, pixel coordinate mapping is performed using the following formula: ; ; ; ;
[0116] in,
[0117] In the formula, This represents the coordinates of the target historical point cloud in the pixel point cloud coordinate system; f x f y c x c y Indicates camera intrinsic parameters, (x d ,y d The coordinates of the target historical point cloud after correction are represented by ), and W and H represent the width and height of the set image size.
[0118] The depth of the target historical point cloud is determined according to the following formula: ;
[0119] Step 4: In ascending order of depth, map each target's historical point cloud onto the pixel point cloud map based on its coordinates in the pixel point cloud coordinate system.
[0120] Pixel cloud map ;
[0121] Mapping rules:
[0122] Insert in ascending order by d.
[0123] Compared to pure image ranging techniques, the embodiments of this disclosure offer more controllable accuracy and greater adaptability to complex scenes: ranging relies on scale estimation using monocular vision or disparity calculation using binocular vision, which is easily affected by shooting angle, lens distortion, and missing scene textures, resulting in large error fluctuations (e.g., errors can reach several meters for distant targets). Furthermore, in complex scenes (e.g., drastic changes in lighting, numerous occlusions, or areas lacking texture), image feature blurring can easily lead to matching failures. This disclosure, relying on absolute scale references provided by known historical 3D points, locks the error within the range of camera intrinsic calibration accuracy and 3D point measurement accuracy. Multiple measurements of the same object can be controlled at the centimeter level. Even if the target portion in the image is occluded or affected by lighting interference, as long as the target bounding box can define the key area (e.g., the visible parts at the bottom and top), the conversion calculation can be completed using a small number of effective 3D points and camera intrinsic parameters. In complex scenes such as foggy weather, backlighting at night, and multiple objects intersecting and occluding, the measurement error can still be stably controlled within a preset range, significantly improving the anti-interference capability compared to pure image techniques.
[0124] Compared to technologies that rely on dedicated sensors, the technical solution disclosed herein has a lower hardware threshold: some high-precision measurement technologies rely on dedicated equipment such as LiDAR and structured light cameras, with the cost of a single device reaching tens or even hundreds of thousands of yuan. This disclosure can achieve the same measurement accuracy based on images acquired by ordinary RGB cameras (costing around a thousand yuan) combined with a small number of key 3D points obtained by low-cost 3D scanning equipment (such as consumer-grade LiDAR), making it more suitable for widespread adoption in small and medium-sized enterprises or civilian scenarios.
[0125] Compared to LiDAR detection technology, this method offers superior real-time response speed: LiDAR real-time detection requires complete preprocessing (denoising, registration), target segmentation, and 3D measurement of massive point clouds (typically hundreds of thousands to millions of points per second) for each frame, with single-frame processing often taking hundreds of milliseconds. In dynamic scenes, it is prone to detection delays (such as lag in calculating the position of rapidly moving targets). This method only needs to focus on a small number of 3D points corresponding to the target bounding box in the image, eliminating the need to process the entire scene's point cloud. It can track changes in target height in real time, making it particularly suitable for continuous monitoring scenarios of dynamic targets (such as real-time monitoring of material height on production lines and dynamic detection of vehicle height on roads).
[0126] The method disclosed in this embodiment is a three-dimensional spatial distance intrusion alarm method based on images and historical three-dimensional point clouds. It solves the problem of accurate three-dimensional distance alarm and is suitable for indoor and outdoor security scenarios with stable backgrounds.
[0127] Based on the same inventive concept, this disclosure provides an alarm and device based on intrusion distance detection. The steps performed by the components of this device are the same as or similar to those of the method described above, therefore, similar details will not be repeated. Figure 4 As shown, the alarm device based on intrusion distance detection in this embodiment includes:
[0128] The pixel point cloud map determination module 410 is used to acquire a pre-established pixel point cloud map of the target area; the target area includes power-related equipment.
[0129] The two-dimensional detection module 420 is used to acquire the image to be detected in the target area and input the image to be detected into the trained intrusion detection model so that the intrusion detection model processes the image to be detected and outputs the information of the intrusion detection box.
[0130] The matching point processing module 430 is used to determine the matching point based on the information of the intruder detection box; the matching point includes the lower left and lower right pixel points of the intruder detection box; and to filter the matching target three-dimensional points in the pixel cloud map based on the two-dimensional coordinates of the matching point.
[0131] The height estimation module 440 is used to determine the estimated three-dimensional height of the intruder based on the height of the intruder detection frame and the intrinsic parameters of the camera that captured the image to be detected.
[0132] The 3D bounding box generation module 450 is used to draw a 3D bounding line of points along the vertical direction, starting from the target 3D point, according to a preset step size and a 3D estimated height; and to draw another 3D bounding line of points between the last extension points of the target 3D point, according to a preset step size and a 3D estimated height; and to form a 3D bounding box of the intruder using the drawn 3D bounding lines. The last extension point is the last point marked along the vertical direction, starting from the target 3D point, according to a preset step size and a 3D estimated height, and the height of this last point is equal to or higher than the 3D estimated height.
[0133] The alarm module 460 is used to determine the distance between the intruder and the power-related equipment using the three-dimensional bounding box of the intruder, and to generate alarm information using the distance.
[0134] The pixel point cloud map determination module 410 is further configured to perform the following steps to generate a pixel point cloud map:
[0135] Transform the historical point cloud located in the historical 3D point cloud coordinate system to the camera coordinate system;
[0136] Based on the camera's horizontal field of view, camera's vertical field of view, and the coordinates of the historical point cloud in the camera coordinate system, filter the target historical point cloud located within the camera's field of view.
[0137] The coordinates of each target historical point cloud in the camera coordinate system are normalized and distortion corrected. Then, the corrected coordinates and camera intrinsic parameters are used to perform pixel coordinate mapping to obtain the coordinates of each target historical point cloud in the pixel point cloud coordinate system. In addition, the depth of each target historical point cloud is calculated using the coordinates of each target historical point cloud in the camera coordinate system.
[0138] In ascending order of depth, each target's historical point cloud is mapped onto the pixel point cloud map based on its coordinates in the pixel point cloud coordinate system.
[0139] The various embodiments of the techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0140] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0141] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0142] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0143] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0144] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0145] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0146] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An alarm method based on intrusion distance detection, characterized in that, include: Obtain a pre-built pixel point cloud map of the target area; The target area includes power-related equipment; The image to be detected within the target area is acquired, and the image to be detected is input into the trained intrusion detection model so that the intrusion detection model processes the image to be detected and outputs the information of the intrusion detection box. Based on the information of the intruder detection box, the matching point is determined; the matching point includes the lower left and lower right corner pixels of the intruder detection box. Based on the two-dimensional coordinates of the point to be matched, target three-dimensional points are selected for matching in the pixel point cloud map; The estimated three-dimensional height of the intruder is determined based on the height of the intruder detection frame and the intrinsic parameters of the camera that captured the image to be detected. Starting from the target 3D point, draw a 3D point enclosing line along the vertical direction according to a preset step size and a 3D estimated height; between the last extension point of the target 3D point, draw another 3D point enclosing line according to a preset step size and a 3D estimated height; use the drawn 3D point enclosing line to form an intruder 3D bounding box; wherein, the last extension point is the last point marked along the vertical direction starting from the target 3D point according to a preset step size and a 3D estimated height, and the height of the last point is equal to or higher than the 3D estimated height; use the intruder 3D bounding box to determine the distance between the intruder and the power-related equipment, and use the distance to generate an alarm message; The pixel point cloud map is generated according to the following steps: Transform the historical point cloud located in the historical 3D point cloud coordinate system to the camera coordinate system; Based on the camera's horizontal field of view, camera's vertical field of view, and the coordinates of the historical point cloud in the camera coordinate system, filter the target historical point cloud located within the camera's field of view. The coordinates of each target historical point cloud in the camera coordinate system are normalized and distortion corrected. Then, the corrected coordinates and camera intrinsic parameters are used to perform pixel coordinate mapping to obtain the coordinates of each target historical point cloud in the pixel point cloud coordinate system. In addition, the depth of each target historical point cloud is calculated using the coordinates of each target historical point cloud in the camera coordinate system. In ascending order of depth, each target's historical point cloud is mapped onto the pixel point cloud map based on its coordinates in the pixel point cloud coordinate system.
2. The alarm method based on intrusion distance detection according to claim 1, characterized in that, The transformation of the historical point cloud located in the historical 3D point cloud coordinate system to the camera coordinate system includes: Based on the camera position and camera orientation, a rigid body transformation matrix is determined; using the rigid body transformation matrix and camera orientation, the historical point cloud located in the historical 3D point cloud coordinate system is transformed to the camera coordinate system.
3. The alarm method based on intrusion distance detection according to claim 1, characterized in that, The step of filtering target historical point clouds located within the camera's field of view based on the camera's horizontal field of view, camera's vertical field of view, and the coordinates of the historical point cloud in the camera coordinate system includes: The historical 3D point cloud that meets the following three conditions will be used as the target historical point cloud: ; ; ; in, Represents the coordinates of the historical 3D point cloud in the camera coordinate system; Indicates the camera's horizontal field of view; This indicates the camera's vertical field of view.
4. The alarm method based on intrusion distance detection according to claim 1, characterized in that, Distortion correction processing is performed on the target historical point cloud, including: Calculate the radial distance and distortion of the target's historical point cloud; The radial distortion is calculated using the radial distance, distortion, and camera intrinsic parameters. Tangential distortion is calculated using camera intrinsic parameters and the coordinates of the target's historical point cloud after normalization. The corrected coordinates of the target's historical point cloud are determined using the normalized coordinates, radial distortion, and tangential distortion of the target's historical point cloud.
5. The alarm method based on intrusion distance detection according to claim 4, characterized in that, The radial distance and distortion of the target's historical point cloud are calculated using the following formula: ; In the formula, r represents the radial distance and distortion, (x n ,y n () represents the normalized coordinates of the target historical point cloud.
6. The alarm method based on intrusion distance detection according to claim 4, characterized in that, Radial distortion is calculated using the following formula: ; In the formula, The value represents radial distortion, k1 and k2 represent camera intrinsic parameters, and r represents radial distance and distortion.
7. The alarm method based on intrusion distance detection according to claim 4, characterized in that, Tangential distortion is calculated using the following formula: ; In the formula, r represents the radial distance and distortion, (x n ,y n () represents the normalized coordinates of the target's historical point cloud, and p1 and p2 represent the camera's intrinsic parameters. , This indicates tangential distortion.
8. The alarm method based on intrusion distance detection according to claim 4, characterized in that, The corrected coordinates of the target's historical point cloud are determined using the following formula: ; In the formula, (x d ,y d (x) represents the corrected coordinates of the target's historical point cloud. n ,y n () represents the coordinates of the target historical point cloud after normalization. Indicates radial distortion. , This indicates tangential distortion.
9. The alarm method based on intrusion distance detection according to claim 1, characterized in that, Pixel coordinate mapping is performed using the following formula: ; ; ; ; In the formula, This represents the coordinates of the target historical point cloud in the pixel point cloud coordinate system; f x f y c x c y Indicates camera intrinsic parameters, (x d ,y d The coordinates of the target historical point cloud after correction are represented by ), and W and H represent the width and height of the set image size.
10. An alarm device based on intrusion distance detection, characterized in that, include: The pixel point cloud map determination module is used to obtain a pre-built pixel point cloud map of the target area; The target area includes power-related equipment; A two-dimensional detection module is used to acquire the image to be detected within the target area and input the image to be detected into a trained intrusion detection model so that the intrusion detection model processes the image to be detected and outputs information of the intrusion detection box. The matching point processing module is used to determine the matching point based on the information of the intruder detection box; the matching point includes the lower left and lower right corner pixels of the intruder detection box; Based on the two-dimensional coordinates of the point to be matched, target three-dimensional points are selected for matching in the pixel point cloud map; The height estimation module is used to determine the three-dimensional estimated height of the intruder based on the height of the intruder detection frame and the intrinsic parameters of the camera that captured the image to be detected. A 3D bounding box generation module is used to draw a 3D bounding line of points along the vertical direction starting from the target 3D point, according to a preset step size and a 3D estimated height; draw a 3D bounding line of points between the last extension points of the target 3D point, according to a preset step size and a 3D estimated height; and form an intrusion 3D bounding box using the drawn 3D bounding lines; wherein, the last extension point is the last point marked along the vertical direction starting from the target 3D point, according to a preset step size and a 3D estimated height, and the height of the last point is equal to or higher than the 3D estimated height; The alarm module is used to determine the distance between the intruder and the power-related equipment using the three-dimensional bounding box of the intruder, and to generate alarm information using the distance; The pixel point cloud map determination module is further configured to perform the following steps to generate a pixel point cloud map: Transform the historical point cloud located in the historical 3D point cloud coordinate system to the camera coordinate system; Based on the camera's horizontal field of view, camera's vertical field of view, and the coordinates of the historical point cloud in the camera coordinate system, filter the target historical point cloud located within the camera's field of view. The coordinates of each target historical point cloud in the camera coordinate system are normalized and distortion corrected. Then, the corrected coordinates and camera intrinsic parameters are used to perform pixel coordinate mapping to obtain the coordinates of each target historical point cloud in the pixel point cloud coordinate system. In addition, the depth of each target historical point cloud is calculated using the coordinates of each target historical point cloud in the camera coordinate system. In ascending order of depth, each target's historical point cloud is mapped onto the pixel point cloud map based on its coordinates in the pixel point cloud coordinate system.