Natural resource element monitoring method and system based on scene event camera
Through a natural resource element monitoring method based on scene event cameras, combined with computer vision and artificial intelligence technologies, the problems of high cost, resolution and imaging frequency limitations in natural resource monitoring in existing technologies are solved, and high-precision monitoring and intelligent response to typical natural resource elements are achieved.
Patent Information
- Application Number
- CN202510768312.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies in natural resource monitoring have problems such as high manual patrol costs, limited satellite remote sensing resolution and imaging frequency, high false alarm rates of traditional cameras, and lack of spatial positioning capabilities, making it difficult to achieve real-time tracking and accurate monitoring of dynamic targets such as engineering vehicles.
A natural resource element monitoring method based on scene event cameras is adopted, combined with computer vision and artificial intelligence technology. Video images are collected through scene event cameras to obtain camera parameters. The observer viewpoint model and geographic registration model are used to spatialize the video. Combined with the engineering vehicle recognition model and event judgment model, high-precision video target perception and positioning are achieved, and changes in typical natural resource elements are monitored.
It achieves high-precision monitoring of typical elements of natural resources, reduces manual monitoring costs, and can respond to damage incidents in a timely manner. It is suitable for intelligent monitoring of geographic spatial scenarios such as plains, hills and plateaus, and has been expanded to areas such as smart fishing ports and urban management.
Smart Images

Figure CN120707628A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a natural resource element monitoring method and system based on a scene event camera. Background Art
[0002] Scene cameras are a type of imaging device capable of dynamically adjusting viewing angle, focal length, and orientation. They are widely used in intelligent surveillance, target tracking, geographic information collection, and other fields. With the acceleration of urbanization and the continuous advancement of infrastructure construction, the supervision and protection of typical natural resource elements such as arable land has become increasingly important. Commonly used technical methods include manual inspections and satellite remote sensing monitoring.
[0003] However, these traditional monitoring methods still have many shortcomings in practical application. Although manual inspections are intuitive, they are limited by high labor costs, long inspection cycles, and limited coverage, making them difficult to meet the needs of large-scale natural resource supervision. Although satellite remote sensing technology has the ability to monitor large areas, it is limited by weather, imaging frequency, and resolution, making it difficult to achieve real-time tracking of dynamic targets such as construction vehicles. Traditional fixed surveillance cameras mainly rely on manual or simple motion detection algorithms to identify anomalies. They are easily interfered with by factors such as lighting and weather, and have a high false alarm rate. At the same time, ordinary video cameras can only provide two-dimensional information at the pixel level and lack spatial positioning capabilities. As a result, it is impossible to effectively distinguish whether construction vehicles have entered sensitive areas, such as the red line area of arable land, in complex terrain. Summary of the Invention
[0004] The present invention provides a natural resource element monitoring method and system based on scene event cameras, which is used to address the defects of traditional camera video targets in the existing technology, such as the lack of spatial information. By utilizing the panoramic video provided by the scene event camera, combined with computer vision and artificial intelligence technologies, high-precision video target perception and positioning, and monitoring changes in typical natural resource elements can be achieved.
[0005] In a first aspect, the present invention provides a natural resource element monitoring method based on a scene event camera, comprising: Capturing a target video image through a scene event camera and obtaining camera parameters of the scene event camera; Based on the preset observer viewpoint model and geo-reference model, obtain the video spatialization model construction parameters; Based on the trained engineering vehicle target recognition model, obtaining the first coordinate system coordinates of the engineering vehicle in the target video image, and converting the first coordinate system coordinates of the engineering vehicle into the fourth coordinate system coordinates using the parameters constructed by the video spatialization model; Obtain the fourth coordinate system data of typical elements of natural resources, combine it with the event judgment model, and monitor the status changes of typical elements of natural resources.
[0006] According to a natural resource element monitoring method based on a scene event camera provided by the present invention, the camera parameters include external parameters and internal parameters, the external parameters include azimuth and pitch angle, and the internal parameters include focal length, field of view angle and image distortion; Correspondingly, obtain the external parameters of the scene event camera, including: Adjusting the scene event camera to a preset horizontal state; Determine the position coordinates and geographic orientation of the control point relative to the scene event camera, and calculate the spherical distance between the control point and the scene event camera; Obtaining the angle between the line of sight of the scene event camera and the geographical north direction when the horizontal azimuth angle is 0°, and calculating the azimuth angle based on the geometric angle output by the scene event camera; Calculating the straight-line distance and the actual pitch angle between the scene event camera and the control point according to the position and elevation of the scene event camera and the position and elevation of the control point, and obtaining a correction value of the pitch angle according to the pitch angle returned by the scene event camera and the correction value of the actual pitch angle; Correspondingly, obtain the internal parameters of the scene event camera, including: Obtaining the intrinsic parameter matrix and distortion parameters of the scene event camera; Based on the optimization criteria, relevant parameters are optimized and adjusted under different conditions, and the optimal parameter configuration is obtained through an iterative method.
[0007] According to a natural resource element monitoring method based on a scene event camera provided by the present invention, based on a preset observer viewpoint model and a geo-referenced model, video spatialization model construction parameters are obtained, including: Calculating a rotation matrix between the fifth coordinate system and the third coordinate system according to the external parameters and the internal parameters of the scene event camera; Establishing a projection transformation matrix based on the calibrated vertical field of view angle parameters and the calibrated horizontal field of view angle parameters, and obtaining a transformation matrix according to the rotation matrix and the projection transformation matrix; Based on the transformation matrix, an inverse transformation matrix of the transformation matrix is obtained, and the coordinates of the target point in the target video image in a third coordinate system are calculated according to the inverse transformation matrix; Based on the inverse transformation matrix of the rotation matrix and the negative translation vector, the coordinates of the target point in the third coordinate system are converted into coordinates in the fifth coordinate system, and the initial fourth coordinate system coordinates of the target point are obtained according to the coordinates in the fifth coordinate system; Based on a preset line-of-sight analysis method, the initial fourth coordinate system coordinates of the target point are positioned and corrected in line of sight in combination with the external parameters, internal parameters and digital elevation model (DEM) data of the scene event camera to obtain the true fourth coordinate system coordinates of the target point.
[0008] According to a natural resource element monitoring method based on a scene event camera provided by the present invention, based on a preset line-of-sight analysis method, the initial fourth coordinate system coordinates of the target point are positioned and corrected in line-of-sight in combination with the external parameters, internal parameters and DEM data of the scene event camera to obtain the true fourth coordinate system coordinates of the target point, including: Determine the sight line direction of the scene event camera according to the external parameters and internal parameters of the scene event camera and the angle between the viewpoint of the scene event camera and the center of the scene event camera, and determine the first intersection point of the sight line direction of the scene event camera and the DEM surface; Calculating the height of the scene event camera on the ground based on the fourth coordinate system coordinates of the scene event camera and the DEM data, calculating the fourth coordinate system coordinates of the scene event camera based on the height of the scene event camera, and converting the fourth coordinate system coordinates of the target point and the scene event camera into sixth coordinate system coordinates; The real fourth coordinate system coordinates of the target point are obtained according to the first intersection point of the sight line direction of the scene event camera and the DEM surface, the first coordinate system coordinates of the target point and the sixth coordinate system coordinates.
[0009] According to a natural resource element monitoring method based on a scene event camera provided by the present invention, based on a trained engineering vehicle target recognition model, the first coordinate system coordinates of the engineering vehicle in the target video image are obtained, and the first coordinate system coordinates of the engineering vehicle are converted into fourth coordinate system coordinates using parameters constructed by the video spatialization model, including: Manually annotate the target video image to construct an engineering vehicle recognition dataset, and expand the engineering vehicle recognition dataset by combining it with a public vehicle detection dataset; Based on a single-stage object detection deep neural network, we set preset hyperparameters and trained it for a preset number of rounds on an expanded engineering vehicle recognition dataset to obtain an engineering vehicle recognition model for detection in natural resource scenarios. Using the engineering vehicle recognition model, a preset confidence threshold is determined, and a first coordinate system coordinate set of a vehicle center point of the engineering vehicle target in the target video image is obtained through network reasoning; Based on the coordinate set of the first coordinate system of the vehicle center point, the DeepSort algorithm is combined to perform multi-target tracking, and output the identification ID of the engineering vehicle target and the coordinate set of the first coordinate system of the trajectory point; Based on the video spatialization model construction parameters, the first coordinate system coordinate set of the vehicle center point and the first coordinate system coordinate set of the trajectory point are combined to convert into corresponding fourth coordinate system coordinates.
[0010] According to a natural resource element monitoring method based on a scene event camera provided by the present invention, based on the coordinate set of the first coordinate system of the vehicle center point, combined with the DeepSort algorithm, multi-target tracking is performed to output the identification ID of the engineering vehicle target and the coordinate set of the first coordinate system of the trajectory point, including: A pre-trained convolutional neural network is used to extract features from the image within each detection frame to obtain a deep feature vector and the appearance information of the engineering vehicle target; The state of the engineering vehicle target is predicted based on the Kalman filter to obtain the estimated position of the vehicle in the current frame. A composite distance metric of motion similarity based on Mahalanobis distance and appearance similarity based on cosine distance is used to perform data association between the current detection result and the predicted trajectory according to the Hungarian algorithm to ensure that each target is correctly matched; For successfully matched targets, the corresponding Kalman filter state and appearance feature records are updated. For unsuccessfully matched targets, new trajectories are initialized, and unmatched trajectories are marked and deleted according to the preset loss threshold.
[0011] According to the present invention, a natural resource element monitoring method based on a scene event camera is provided, which obtains fourth coordinate system data of typical natural resource elements and monitors the state changes of the typical natural resource elements in combination with an event judgment model, including: Obtain the geometric information and attribute information of typical natural resource elements in the area within the camera's field of view based on field measurements, and output them in GeoJson format. The geometric information includes the fourth coordinate system coordinates of the typical natural resource elements, and the attribute information includes the type and category of the typical natural resource elements. Combined with the identification ID of the engineering vehicle target and the coordinate set of the fourth coordinate system of the trajectory point, the trajectory lines of engineering vehicles with different identification IDs are obtained. The intersection of the trajectory line and the element polygon is calculated through the line-polygon intersection algorithm to determine whether the engineering vehicle target has entered the cultivated land area. If the intersection is not empty, it is determined that the engineering vehicle target has entered the interior of the feature at least at some point in time.
[0012] According to a natural resource element monitoring method based on a scene event camera provided by the present invention, the event determination model includes: Determine the preset time threshold and the preset number threshold based on the requirements and experience of actual protection of typical natural resource elements; If the engineering vehicle target stays in the element area continuously for more than the preset time threshold, or repeatedly enters the cultivated land more than the preset number threshold within a certain period of time, it is determined that the engineering vehicle target has the risk of damaging the element.
[0013] In a second aspect, the present invention further provides a natural resource element monitoring system based on a scene event camera, comprising: An acquisition module, configured to capture a target video image through a scene event camera and obtain camera parameters of the scene event camera; A construction module, used for obtaining video spatialization model construction parameters based on a preset observer viewpoint model and a geo-referenced model; a conversion module, configured to obtain the first coordinate system coordinates of the engineering vehicle in the target video image based on the trained engineering vehicle target recognition model, and convert the first coordinate system coordinates of the engineering vehicle into the fourth coordinate system coordinates using the parameters constructed by the video spatialization model; The monitoring module is used to obtain the fourth coordinate system data of typical elements of natural resources, and monitor the status changes of typical elements of natural resources in combination with the event judgment model.
[0014] In a third aspect, the present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for monitoring natural resource elements based on a scene event camera as described above is implemented.
[0015] The natural resource element monitoring method and system based on scene event cameras provided by the present invention realize the monitoring of typical natural resource elements based on scene event cameras. It can be applied to the intelligent monitoring of changes in natural resource elements in geographical spatial scenes with different elevations such as plains, hills and plateaus, and realize rapid and accurate perception of changes in typical elements such as cultivated land, so that staff can respond and handle events such as destruction of cultivated land and illegal construction in a timely manner, reducing the time and economic costs generated by manual visual monitoring. It can also be expanded to other intelligent monitoring and emergency command fields such as smart fishing ports and urban management, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 Schematic diagram of the process of natural resource element monitoring method based on scene event camera provided by the present invention; Figure 2 This is a schematic diagram of monitoring and positioning of typical elements of natural resources provided by the present invention; Figure 3 Schematic diagram of the structure of the natural resource element monitoring system based on scene event camera provided by the present invention; Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0018] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0019] Figure 1 FIG. 1 is a flow chart of a natural resource element monitoring method based on a scene event camera according to an embodiment of the present invention. Figure 1 Shown, including: Step 100: Capturing a target video image through a scene event camera and obtaining camera parameters of the scene event camera; The scene event camera in the embodiment of the present invention is a frame acquisition-based camera that captures a complete image of the entire scene through a fixed exposure time, such as a panoramic PTZ camera, which can achieve pan, tilt and optical zoom through remote control.
[0020] The external parameters of the scene event camera include azimuth and pitch angles, and the internal parameters of the scene event camera include focal length, field of view angle, and image distortion; The azimuth angle is the angle formed by rotating clockwise from east to south to west with due north as the starting point at 0°; the pitch angle is the angle between the Z axis of the scene event third coordinate system and the horizontal plane. The external parameters of the scene event camera are obtained through the scene event camera external parameter calibration and correction. The purpose of obtaining the external parameters of the scene event camera is to standardize the installation of the scene event camera and ensure that the external parameters of the scene event camera are within a certain error range; the internal parameters of the scene event camera are obtained through the scene event camera internal parameter calibration and correction. The purpose of obtaining the internal parameters of the scene event camera is to determine the internal parameters of the scene event camera and their correction parameters through a small number of measured control points.
[0021] The scene event camera can obtain a larger field of view by rotating the lens in the horizontal and vertical directions. When the scene event camera rotates, the horizontal and vertical rotation angles will affect the imaging. Therefore, the embodiment of the present invention needs to obtain the horizontal and vertical rotation angles of the scene event camera. In the horizontal direction, the rotation angle of the scene event camera is usually determined by the azimuth angle. The azimuth angle is the angle formed by rotating clockwise from east to south to west with due north as the starting point at 0°. In the vertical direction, the rotation angle of the scene event camera is usually determined by the pitch angle. The pitch angle is the angle between the Z axis of the scene event third coordinate system and the horizontal plane.
[0022] Optionally, in some embodiments, obtaining external parameters of the scene event camera includes: adjusting the scene event camera to a preset horizontal state; determining the position coordinates and geographic orientation of the control point relative to the scene event camera, and calculating the spherical distance between the control point and the scene event camera; obtaining the angle between the line of sight of the scene event camera and the geographic north direction when the horizontal azimuth angle is 0°, and calculating the azimuth angle based on the geometric angle output by the scene event camera; calculating the straight-line distance and actual pitch angle between the scene event camera and the control point based on the position and elevation of the scene event camera and the position and elevation of the control point, and obtaining the pitch angle correction value based on the pitch angle returned by the scene event camera and the correction value of the actual pitch angle.
[0023] Among them, the control point is a certain position in the area to be measured. The camera is calibrated by setting the control point as a reference standard; the elevation refers to the vertical distance from the position of the scene event camera to the ground.
[0024] Specifically, when installing a scene event camera, you can use a spirit level to adjust the inclination of the bracket or use the scene event camera's built-in level calibration function to adjust the scene event camera to a preset horizontal state. After determining the position coordinates and geographic orientation of the control point using a total station or real-time kinematic (RTK) technology, calculate the control point according to the spherical distance formula. Obtain the angle between the line of sight and the geographic north direction when the scene event camera's azimuth in the horizontal direction is 0°, obtain the geometric angle output by the scene event camera, and calculate the azimuth of the scene event camera using the formula. If the geometric angle is greater than 90°, add 360° to the calculated result. Obtain the position and elevation of the scene event camera and the position and elevation of the control point. Calculate the straight-line distance and actual pitch angle between the scene event camera and the control point according to the formula. Correct the pitch angle based on the difference between the pitch angle returned by the scene event camera itself and the actual pitch angle to obtain the corrected pitch angle value. Similarly, calculate the correction value of the azimuth angle at different magnifications.
[0025] Spherical distance The calculation formulas are shown in (1) and (2): (1) (2) in, is the latitude of point A, is the longitude of point A, is the latitude of point B, is the longitude of point B, is the radius of the Earth.
[0026] Optionally, in some embodiments, obtaining the internal parameters of the scene event camera includes: obtaining the internal parameter matrix and distortion parameters of the scene event camera; optimizing and adjusting the relevant parameters under different conditions based on optimization criteria, and obtaining a better parameter configuration through an iterative method.
[0027] Specifically, embodiments of the present invention utilize the Zhang Zhengyou calibration method to obtain the camera's intrinsic parameter matrix and distortion parameters. This involves capturing an image of a checkerboard calibration plate with a scene event camera. A corner detection algorithm is then used to obtain the first coordinate coordinates of each black and white grid corner point, and the image's distortion parameters are calculated. Based on optimization criteria, the relevant parameters are optimized and adjusted under different conditions, and an iterative method is used to achieve the optimal parameter configuration.
[0028] For example, using a checkerboard calibration plate, after obtaining an image of the calibration plate, use the corner detection algorithm to obtain the first coordinate system coordinates of each black and white grid corner point Since the tangential distortion has a small effect, it is not considered here. Only the radial distortion is considered. The calculation formula is as follows (3): (3) in, is the corrected coordinate of the first coordinate system, is the distortion parameter, Represents the distance from a pixel to the center of the image. When an image is distorted, the above formula can be used to calculate the correct coordinates of each pixel in the image, restoring the image to its pre-distortion state. Based on the optimization criteria, the relevant parameters are optimized and adjusted under different conditions. The optimal parameter configuration is obtained through an iterative method, which can achieve automatic parameter calibration.
[0029] Step 200: Obtaining video spatialization model construction parameters based on a preset observer viewpoint model and a geo-referenced model; The georeference of the video frame is defined using seven specified parameters, including latitude, longitude, altitude, azimuth, field of view, pitch, and roll, which are applied to the video georeferencing model of the scene event camera.
[0030] Optionally, in some embodiments, based on a preset observer viewpoint model, external parameters of the scene event camera, and internal parameters of the scene event camera, the first coordinate system coordinates of the target point in the video image are converted into fourth coordinate system coordinates, including: calculating the rotation matrix between the fifth coordinate system and the third coordinate system based on the external parameters of the scene event camera and the internal parameters of the scene event camera; establishing a projection transformation matrix based on the calibrated vertical field of view angle parameters and the calibrated horizontal field of view angle parameters, and obtaining a transformation matrix based on the rotation matrix and the projection transformation matrix; obtaining an inverse transformation matrix of the transformation matrix based on the transformation matrix, and calculating the coordinates of the target point in the video image in the third coordinate system based on the inverse transformation matrix; converting the coordinates of the target point in the video image in the third coordinate system into fifth coordinate system coordinates based on the inverse transformation matrix of the rotation matrix and the negative translation vector, and obtaining the fourth coordinate system coordinates of the target point in the video image based on the fifth coordinate system coordinates.
[0031] Specifically, the embodiment of the present invention can obtain the transformation matrix of converting the coordinates of the first coordinate system into the coordinates of the fourth coordinate system by obtaining the inverse transformation matrix of each step of converting the coordinates of the fourth coordinate system into the coordinates of the first coordinate system. The inverse transformation matrix of the projection transformation matrix can be used to calculate the coordinates of the image pixels converted into the third coordinate system, and then the inverse matrix of the rotation matrix and the negative translation vector are calculated to convert the coordinates of the third coordinate system into the coordinates of the fifth coordinate system.
[0032] The fourth coordinate system coordinates are converted to the first coordinate system coordinates. First, the fourth coordinate system coordinates of the control point are converted to the third coordinate system coordinates. The calculation formulas are shown in (4) and (5): (4) (5) in, is the first eccentricity of the Earth, is the radius of curvature of the Maoyou circle, , is the Earth's semi-major axis, Indicates altitude.
[0033] Scene event camera parameter, represents the azimuth, Represents the pitch angle and calculates the rotation matrix between the two coordinate systems , use the rotation matrix to describe the conversion relationship between the coordinates, and then substitute the elevation data of the control point to calculate the translation component between the coordinate systems , the calculation formula is shown in (6): (6) Based on the calibrated vertical field of view angle parameters and horizontal field of view angle parameters, a projection transformation matrix is established to describe the transformation relationship between the coordinates of the second coordinate system and the coordinates of the third coordinate system. The calculation formulas are shown in (7) and (8): (7) (8) in, is the camera’s intrinsic parameter matrix, is the coordinate in the second coordinate system, is the coordinate in the first coordinate system; Finally, the dimensions of the above transformation relationships are unified, and the calculation process is merged into one formula. In order to facilitate algorithm implementation and optimize efficiency, homogeneous coordinates and matrix operations are introduced to obtain the transformation matrix of the fourth coordinate system coordinates converted into the first coordinate system coordinates corresponding to the video. The coordinates of any point in the fourth coordinate system can be substituted into it to calculate its first coordinate system coordinates on the video screen.
[0034] Optionally, in some embodiments, based on a preset line-of-sight analysis method, combined with the external parameters of the scene event camera, the internal parameters of the scene event camera and the DEM data, the initial fourth coordinate system coordinates of the video image target point are positioned and corrected to obtain the true fourth coordinate system coordinates of the video image target point.
[0035] An embodiment of the present invention determines whether a scene event camera and a target point in a video image are visible to each other by determining whether there is a corner point between the sight line direction of the scene event camera and the DEM surface based on a line of sight analysis method. When the scene event camera and the target point in the video image are visible to each other, the fourth coordinate system coordinates of the scene event camera are calculated, and the fourth coordinate system coordinates of the target point in the video image and the scene event camera are converted into sixth coordinate system coordinates. Based on a preset search algorithm, the true fourth coordinate system coordinates of the target point in the video image are obtained according to the first intersection point between the sight line direction of the scene event camera and the DEM surface, the first coordinate system coordinates of the target point in the video image, and the sixth coordinate system coordinates, so as to determine the actual position of the target point in the video image.
[0036] Specifically, the preset line-of-sight analysis method combines the external parameters of the scene event camera, the internal parameters of the scene event camera and the DEM data to perform line-of-sight positioning and coordinate correction on the initial fourth coordinate system coordinates of the video image target point to obtain the true fourth coordinate system coordinates of the video image target point, including: determining the line of sight direction of the scene event camera according to the external parameters of the scene event camera, the internal parameters of the scene event camera, and the angle between the viewpoint of the scene event camera and the center of the scene event camera, and determining the first intersection point of the line of sight direction of the scene event camera and the DEM surface; calculating the height of the scene event camera at the ground according to the fourth coordinate system coordinates of the scene event camera and the DEM data, and then calculating the fourth coordinate system coordinates of the scene event camera according to the hanging height of the scene event camera, and converting the fourth coordinate system coordinates of the video image target point and the scene event camera into sixth coordinate system coordinates; obtaining the true fourth coordinate system coordinates of the video image target point according to the first intersection point of the line of sight direction of the scene event camera and the DEM surface, the first coordinate system coordinates of the video image target point and the sixth coordinate system coordinates.
[0037] Step 300: Based on the trained engineering vehicle target recognition model, obtain the first coordinate system coordinates of the engineering vehicle in the target video image, and use the video spatialization model construction parameters to convert the first coordinate system coordinates of the engineering vehicle into fourth coordinate system coordinates.
[0038] Specifically, it includes: obtaining a sufficient number of video frame images output by scene event cameras, manually annotating them, producing an engineering vehicle recognition dataset, and combining them with public vehicle detection datasets for data expansion; based on a single-stage target detection deep neural network, setting appropriate hyperparameters, and fully training enough rounds on the engineering vehicle dataset to obtain an engineering vehicle recognition model with excellent detection effect in natural resource scenarios; based on the obtained engineering vehicle recognition model, setting an appropriate confidence threshold, and obtaining the first coordinate system coordinate set of the vehicle center point of the engineering vehicle target that may exist in the video image through network reasoning; based on the obtained engineering vehicle target center image pixel set, combining the DeepSort algorithm to realize multi-target tracking, and output the identification ID of the engineering vehicle target and the first coordinate system coordinate set of the trajectory point; based on the constructed video spatialization model, combined with the first coordinate system coordinates of the vehicle center point of the engineering vehicle target and the first coordinate system coordinate set of the trajectory point output by the engineering vehicle recognition model, convert them into corresponding fourth coordinate system coordinates.
[0039] Furthermore, a trained engineering vehicle target recognition model is obtained, including: obtaining a sufficient number of video frame images output by scene event cameras, manually annotating them, creating an engineering vehicle recognition dataset, and combining them with a public vehicle detection dataset for data expansion; based on a single-stage target detection deep neural network, setting appropriate hyperparameters, and fully training for sufficient rounds on the engineering vehicle dataset to obtain an engineering vehicle recognition model with excellent detection effect in natural resource scenarios.
[0040] The real-world scene video within the camera's field of view is captured through the scene event camera's video stream address interface, and frames are extracted and saved at regular intervals. A sufficient number of frames refers to at least 500 video images containing valid engineering vehicle targets after extraction. Manual annotation is performed using the Labelme annotation tool, selecting the locations of valid engineering vehicle targets in the video images and recording the corresponding category labels. The public vehicle detection dataset includes UA-DETRAC and some real-world engineering vehicle images acquired using web crawlers. The single-stage object detection algorithm is the popular real-time object detection algorithm YOLO11. Hyperparameters include learning rate, batch size, loss function, and optimizer.
[0041] Improvements were made to the YOLO11 baseline model, including optimizing the backbone network, optimizing the neck structure, and refining the loss function. Although the underlying target network possesses certain feature extraction capabilities and achieves high accuracy, it can still encounter issues such as missed and false detections in actual detection tasks involving finely classified vehicle models in complex scenarios. The backbone network was optimized by introducing an attention module that embeds the two-dimensional spatial position information from any feature map into the channel features. This module aggregates more features horizontally and vertically, ultimately outputting feature weights for the height and width dimensions, which are then combined with the original feature map. This allows the long-range dependencies between different features to be discovered while preserving spatial position information.
[0042] Factors such as the scene event camera's shooting angle and the high speed of vehicles cause vehicles in the video to deform significantly as they drive. Furthermore, some vehicle models have similar appearances, making it difficult to distinguish between different types of vehicles when scaled to a certain size, leading to false detections. Furthermore, some vehicles are located farther away from the scene event camera, resulting in their size being significantly reduced in the surveillance video, leading to missed detections. The neck network plays a connecting role in feature fusion within the model, integrating vehicle feature information at different scales. The combination of these multi-level features directly impacts the final detection results. The neck network structure is optimized to capture more vehicle feature information by increasing its depth and modifying the feature fusion method.
[0043] Loss functions are often used to measure the difference between a model's prediction and the ground-truth object, and are an important metric for evaluating model performance. YOLO11's loss function primarily consists of three components: bounding box localization loss, object classification loss, and confidence loss. The bounding box localization loss uses the CIoU loss function. The penalty is calculated based on the geometric relationship between the predicted and ground-truth boxes, without considering their directional alignment. This results in slow model convergence and potential oscillation around the optimal prediction accuracy. By redesigning the penalty metric to account for the angular difference between the two vectors, we shorten the model training cycle and improve model detection performance.
[0044] Optionally, in some embodiments, obtaining the first coordinate system coordinates of the engineering vehicle in the test video includes: extracting frames from the test video at a certain time interval, selecting a continuous set of extracted frame images, inputting them into a trained engineering vehicle target recognition model, setting an appropriate confidence threshold, and obtaining the first coordinate system coordinate set of the vehicle center point of the engineering vehicle target that may exist in the video image through network reasoning.
[0045] Based on the obtained coordinate set of the first coordinate system of the center of the engineering vehicle target, the DeepSort algorithm is combined to realize multi-target tracking, and the identification ID of the engineering vehicle target and the coordinate set of the first coordinate system of the trajectory point are output.
[0046] The input of the DeepSort algorithm is a series of detection boxes output by the recognition model, each of which contains location information and category information. , then use a deep neural network to extract the appearance features of the target and express them as feature vectors, as shown in formula (9): (9) in, is the image after the detection frame is cropped, is the deep neural network mapping function, is a high-dimensional feature vector.
[0047] Furthermore, the target state is estimated based on the Kalman filter, and the target state vector is defined as , the state transfer equation is , the observation model is , based on the state of the previous frame Calculate the state prediction value of the current frame , , calculate the covariance forecast ; Calculate Kalman gain ; Update status ; Update covariance , combining Mahalanobis distance and cosine similarity of appearance features to perform target matching. If the detection box does not match the existing target, a new trajectory is initialized. If the detection box successfully matches a target, the state of the target is updated and its appearance features are stored. If the target fails to match the new detection box within a certain number of frames, the target trajectory is deleted. The algorithm finally outputs the unique ID, position and size of each target to form trajectory information .
[0048] Based on the constructed video spatialization model, combined with the first coordinate system coordinates of the vehicle center point of the engineering vehicle target and the first coordinate system coordinates of the trajectory point output by the engineering vehicle recognition model, they are converted into corresponding fourth coordinate system coordinates. Specifically, based on the preset observer viewpoint model, the external parameters of the scene event camera and the internal parameters of the scene event camera, the first coordinate system coordinates of the target trajectory point are converted into the fourth coordinate system, including: calculating the rotation matrix between the fifth coordinate system and the third coordinate system based on the external parameters of the scene event camera and the internal parameters of the scene event camera; establishing a projection transformation matrix based on the calibrated vertical field of view angle parameters and the calibrated horizontal field of view angle parameters, and obtaining a transformation matrix based on the rotation matrix and the projection transformation matrix; obtaining an inverse transformation matrix of the transformation matrix based on the transformation matrix, and calculating the coordinates of the target point in the video image in the third coordinate system based on the inverse transformation matrix; converting the coordinates of the target trajectory point in the third coordinate system into the fifth coordinate system coordinates based on the inverse transformation matrix of the rotation matrix and the negative translation vector, and obtaining the fourth coordinate system coordinates of the target trajectory point based on the fifth coordinate system coordinates.
[0049] Step 400: Acquire the fourth coordinate system data of typical elements of natural resources, and monitor the state changes of the typical elements of natural resources in combination with the event judgment model; Specifically, it includes: obtaining the geometric information and attribute information of typical natural resource elements in the area where the scene camera's field of view is located based on field measurements, and outputting it in GeoJson format; the geometric information includes the fourth coordinate system coordinates of the elements, and the attribute information includes the element type and category; combining the identification ID of the engineering vehicle target and the fourth coordinate system coordinate set of the trajectory points, obtaining the trajectory lines of engineering vehicles with different identification IDs, calculating the intersection of the trajectory line and the element polygon through the line-polygon intersection algorithm, and judging whether the engineering vehicle has entered the typical element area. If the intersection is not empty, it means that the engineering vehicle has entered the element at least at some point in time.
[0050] Specifically, the intersection algorithm of a line and a polygon is as follows: a trajectory line is composed of multiple trajectory points , a trajectory segment is composed of adjacent trajectory points , the parametric equation of the trajectory segment can be expressed as , a polygon is composed of a series of vertices , each edge consists of two adjacent vertices , then the intersection of the trajectory line and the boundary polygon can be converted into an intersection operation between line segments.
[0051] Optionally, in some embodiments, the event judgment model introduces two threshold parameters, time and number of times. The setting of the threshold is based on the requirements and experience of actual protection of typical elements of natural resources. It stipulates that if the construction vehicle stays continuously in the element area for more than a preset time threshold, or repeatedly enters the cultivated land more than a preset number threshold within a certain period of time, it can be determined that the construction vehicle is at risk of damaging the element, thereby realizing intelligent monitoring of changes in typical elements of natural resources.
[0052] Specifically, define the characteristic function , to determine whether the vehicle is in the element area, the time when the vehicle first enters the element area is , leaving the area at , the continuous residence time ,like , determine that the engineering vehicle has a risk of destruction; within a certain time window, detect the moment when the engineering vehicle enters the element , count the number of entries ,like , and also determine that the engineering vehicles have the risk of destruction. Figure 2 In the example shown, the engineering vehicle in the image can be accurately identified and marked with a box.
[0053] The natural resource element monitoring system based on the scene event camera provided by the present invention is described below. The natural resource element monitoring system based on the scene event camera described below and the natural resource element monitoring method based on the scene event camera described above can be referenced to each other.
[0054] Figure 3 FIG is a structural diagram of a natural resource element monitoring system based on a scene event camera provided by an embodiment of the present invention. Figure 3 As shown, it includes: an acquisition module 31, a construction module 32, a conversion module 33 and a monitoring module 34, wherein: The acquisition module 31 is used to capture the target video image through the scene event camera and obtain the camera parameters of the scene event camera; the construction module 32 is used to obtain the video spatialization model construction parameters based on the preset observer viewpoint model and the geographic registration model; the conversion module 33 is used to obtain the first coordinate system coordinates of the engineering vehicle in the target video image based on the trained engineering vehicle target recognition model, and use the video spatialization model construction parameters to convert the first coordinate system coordinates of the engineering vehicle into the fourth coordinate system coordinates; the monitoring module 34 is used to obtain the fourth coordinate system data of typical elements of natural resources, and monitor the status changes of typical elements of natural resources in combination with the event judgment model.
[0055] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other via the communications bus 440. The processor 410 may call logic instructions in the memory 430 to execute a natural resource element monitoring method based on a scene event camera, the method comprising: capturing a target video image through a scene event camera and obtaining camera parameters of the scene event camera; obtaining video spatialization model construction parameters based on a preset observer viewpoint model and a geo-registration model; obtaining first coordinate system coordinates of an engineering vehicle in a target video image based on a trained engineering vehicle target recognition model, and converting the first coordinate system coordinates of the engineering vehicle into fourth coordinate system coordinates using the video spatialization model construction parameters; obtaining fourth coordinate system data of typical natural resource elements, and monitoring state changes of the typical natural resource elements in combination with an event determination model.
[0056] Furthermore, the logic instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0057] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the natural resource element monitoring method based on the scene event camera provided by the above-mentioned methods, the method comprising: collecting target video images through a scene event camera, and obtaining the camera parameters of the scene event camera; obtaining video spatialization model construction parameters based on a preset observer viewpoint model and a geographic registration model; obtaining the first coordinate system coordinates of the engineering vehicle in the target video image based on a trained engineering vehicle target recognition model, and converting the first coordinate system coordinates of the engineering vehicle into fourth coordinate system coordinates using the video spatialization model construction parameters; obtaining fourth coordinate system data of typical natural resource elements, and monitoring the state changes of typical natural resource elements in combination with the event judgment model.
[0058] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0059] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A natural resource element monitoring method based on scene event camera, characterized in that: include: Capturing a target video image through a scene event camera and obtaining camera parameters of the scene event camera; Based on the preset observer viewpoint model and geo-reference model, obtain the video spatialization model construction parameters; Based on the trained engineering vehicle target recognition model, obtaining the first coordinate system coordinates of the engineering vehicle in the target video image, and converting the first coordinate system coordinates of the engineering vehicle into the fourth coordinate system coordinates using the parameters constructed by the video spatialization model; Obtain the fourth coordinate system data of typical elements of natural resources, combine it with the event judgment model, and monitor the status changes of typical elements of natural resources.
2. The natural resource element monitoring method based on scene event camera according to claim 1 is characterized in that: The camera parameters include external parameters and internal parameters, the external parameters include azimuth and pitch angles, and the internal parameters include focal length, field of view angle, and image distortion; Correspondingly, obtain the external parameters of the scene event camera, including: Adjusting the scene event camera to a preset horizontal state; Determine the position coordinates and geographic orientation of the control point relative to the scene event camera, and calculate the spherical distance between the control point and the scene event camera; Obtaining the angle between the line of sight of the scene event camera and the geographical north direction when the horizontal azimuth angle is 0°, and calculating the azimuth angle based on the geometric angle output by the scene event camera; Calculating the straight-line distance and the actual pitch angle between the scene event camera and the control point according to the position and elevation of the scene event camera and the position and elevation of the control point, and obtaining a correction value of the pitch angle according to the pitch angle returned by the scene event camera and the correction value of the actual pitch angle; Correspondingly, obtain the internal parameters of the scene event camera, including: Obtaining the intrinsic parameter matrix and distortion parameters of the scene event camera; Based on the optimization criteria, relevant parameters are optimized and adjusted under different conditions, and the optimal parameter configuration is obtained through an iterative method.
3. The natural resource element monitoring method based on scene event camera according to claim 1 is characterized in that: Based on the preset observer viewpoint model and georeferencing model, obtain the video spatialization model construction parameters, including: Calculating a rotation matrix between the fifth coordinate system and the third coordinate system according to the external parameters and the internal parameters of the scene event camera; Establishing a projection transformation matrix based on the calibrated vertical field of view angle parameters and the calibrated horizontal field of view angle parameters, and obtaining a transformation matrix according to the rotation matrix and the projection transformation matrix; Based on the transformation matrix, an inverse transformation matrix of the transformation matrix is obtained, and the coordinates of the target point in the target video image in a third coordinate system are calculated according to the inverse transformation matrix; Based on the inverse transformation matrix of the rotation matrix and the negative translation vector, the coordinates of the target point in the third coordinate system are converted into coordinates in the fifth coordinate system, and the initial fourth coordinate system coordinates of the target point are obtained according to the coordinates in the fifth coordinate system; Based on a preset line-of-sight analysis method, the initial fourth coordinate system coordinates of the target point are positioned and corrected in line of sight in combination with the external parameters, internal parameters and digital elevation model (DEM) data of the scene event camera to obtain the true fourth coordinate system coordinates of the target point.
4. The natural resource element monitoring method based on scene event camera according to claim 3 is characterized in that: Based on a preset line-of-sight analysis method, the initial fourth coordinate system coordinates of the target point are positioned and corrected in line-of-sight in combination with the external parameters, internal parameters, and DEM data of the scene event camera to obtain the true fourth coordinate system coordinates of the target point, including: Determine the sight line direction of the scene event camera according to the external parameters and internal parameters of the scene event camera and the angle between the viewpoint of the scene event camera and the center of the scene event camera, and determine the first intersection point of the sight line direction of the scene event camera and the DEM surface; Calculating the height of the scene event camera on the ground based on the fourth coordinate system coordinates of the scene event camera and the DEM data, calculating the fourth coordinate system coordinates of the scene event camera based on the height of the scene event camera, and converting the fourth coordinate system coordinates of the target point and the scene event camera into sixth coordinate system coordinates; The real fourth coordinate system coordinates of the target point are obtained according to the first intersection point of the sight line direction of the scene event camera and the DEM surface, the first coordinate system coordinates of the target point and the sixth coordinate system coordinates.
5. The natural resource element monitoring method based on scene event camera according to claim 1 is characterized in that: Based on the trained engineering vehicle target recognition model, obtaining the first coordinate system coordinates of the engineering vehicle in the target video image, and using the video spatialization model to construct parameters to convert the first coordinate system coordinates of the engineering vehicle into the fourth coordinate system coordinates, including: Manually annotate the target video image to construct an engineering vehicle recognition dataset, and expand the engineering vehicle recognition dataset by combining it with a public vehicle detection dataset; Based on a single-stage object detection deep neural network, we set preset hyperparameters and trained it for a preset number of rounds on an expanded engineering vehicle recognition dataset to obtain an engineering vehicle recognition model for detection in natural resource scenarios. Using the engineering vehicle recognition model, a preset confidence threshold is determined, and a first coordinate system coordinate set of a vehicle center point of the engineering vehicle target in the target video image is obtained through network reasoning; Based on the coordinate set of the first coordinate system of the vehicle center point, the DeepSort algorithm is combined to perform multi-target tracking, and output the identification ID of the engineering vehicle target and the coordinate set of the first coordinate system of the trajectory point; Based on the video spatialization model construction parameters, the first coordinate system coordinate set of the vehicle center point and the first coordinate system coordinate set of the trajectory point are combined to convert into corresponding fourth coordinate system coordinates.
6. The natural resource element monitoring method based on scene event camera according to claim 5 is characterized in that: Based on the coordinate set of the first coordinate system of the vehicle center point, the DeepSort algorithm is combined to perform multi-target tracking, and the identification ID of the engineering vehicle target and the coordinate set of the first coordinate system of the trajectory point are output, including: A pre-trained convolutional neural network is used to extract features from the image within each detection frame to obtain a deep feature vector and the appearance information of the engineering vehicle target; The state of the engineering vehicle target is predicted based on the Kalman filter to obtain the estimated position of the vehicle in the current frame. A composite distance metric of motion similarity based on Mahalanobis distance and appearance similarity based on cosine distance is used to perform data association between the current detection result and the predicted trajectory according to the Hungarian algorithm to ensure that each target is correctly matched; For successfully matched targets, the corresponding Kalman filter state and appearance feature records are updated. For unsuccessfully matched targets, new trajectories are initialized, and unmatched trajectories are marked and deleted according to the preset loss threshold.
7. The natural resource element monitoring method based on scene event camera according to claim 1 is characterized in that: Obtain the fourth coordinate system data of typical natural resource elements and combine it with the event judgment model to monitor the status changes of typical natural resource elements, including: Obtain the geometric information and attribute information of typical natural resource elements in the area within the camera's field of view based on field measurements, and output them in GeoJson format. The geometric information includes the fourth coordinate system coordinates of the typical natural resource elements, and the attribute information includes the type and category of the typical natural resource elements. Combined with the identification ID of the engineering vehicle target and the coordinate set of the fourth coordinate system of the trajectory point, the trajectory lines of engineering vehicles with different identification IDs are obtained. The intersection of the trajectory line and the element polygon is calculated through the line-polygon intersection algorithm to determine whether the engineering vehicle target has entered the cultivated land area. If the intersection is not empty, it is determined that the engineering vehicle target has entered the interior of the feature at least at some point in time.
8. The natural resource element monitoring method based on scene event camera according to claim 7 is characterized in that: The event determination model includes: Determine the preset time threshold and the preset number threshold based on the requirements and experience of actual protection of typical natural resource elements; If the engineering vehicle target stays in the element area continuously for more than the preset time threshold, or repeatedly enters the cultivated land more than the preset number threshold within a certain period of time, it is determined that the engineering vehicle target has the risk of damaging the element.
9. A natural resource element monitoring system based on scene event camera, characterized in that: include: An acquisition module, configured to capture a target video image through a scene event camera and obtain camera parameters of the scene event camera; A construction module, used for obtaining video spatialization model construction parameters based on a preset observer viewpoint model and a geo-referenced model; a conversion module, configured to obtain the first coordinate system coordinates of the engineering vehicle in the target video image based on the trained engineering vehicle target recognition model, and convert the first coordinate system coordinates of the engineering vehicle into the fourth coordinate system coordinates using the parameters constructed by the video spatialization model; The monitoring module is used to obtain the fourth coordinate system data of typical elements of natural resources, and monitor the status changes of typical elements of natural resources in combination with the event judgment model.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the natural resource element monitoring method based on the scene event camera as described in any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Man-machine cooperation-oriented aerial work behavior visual identification method and system
CN121999419A
A visual recognition method and system for high-altitude operation behavior in human-machine collaboration
CN121999419B