In-field space calibration method and system based on depth field of view, and storage medium
Through the in-field space calibration method based on deep field of view, monocular vision and improved YOLOv8 neural network are used to solve the problem of efficient calibration of specific categories of objects in the field of view in natural resource monitoring, and high-precision spatial positioning and data acquisition are achieved, providing real-time decision-making support for natural resource emergency management.
Patent Information
- Application Number
- CN202510389816.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-22
AI Technical Summary
It is difficult for the prior art to quickly and inexpensively perform batch calibration and position trajectory calculations of specific categories of objects within the field of view during natural resource monitoring, especially to accurately mark the distribution coordinates and quantities of rare species in complex environments.
The in-field space calibration method based on the depth field of view is adopted, and high-altitude monitoring points, camera focal length and inclination angle data are used, combined with monocular vision principles, and target detection is performed through the improved YOLOv8 neural network, the absolute spatial coordinates of the target object are calculated, and the quadratic surface equations of the large ground in the field of view are fitted to achieve high-precision spatial positioning.
It realizes rapid response to emergencies in natural resource monitoring, provides high-precision data on coordinates, quantity and activity trajectory of target objects, provides real-time decision-making basis for natural resource emergency management, is suitable for mobile devices and high-altitude monitoring, and supports spatial data acquisition and fitting within a large range.
Smart Images

Figure CN120355779A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural resource investigation and monitoring, and particularly relates to a method, system and storage medium for calibrating the in-field space based on the depth of field of view. Background Art
[0002] In the application of natural resource animal and plant monitoring, in order to accurately obtain the growth and activity rules of rare animals and plants, a large amount of manual work is required to complete the interpretation, calculation and statistics. Quickly obtaining the spatial coordinates of preset category objects within the camera's field of view, including information such as geographical location, height and quantity, can provide a basis for the subsequent growth and activity rules of animals and plants, and can also serve relevant protection and management work. With the breakthroughs in multi-spectral imaging, AI behavior recognition and edge computing technologies, video surveillance systems have become the core tools for monitoring the growth and activities of animals and plants in the field of natural resources. The mainstream solution in 2025 (such as Huawei's "Natural Eye" system) collects multi-modal data such as species activity trajectories, breeding behaviors and habitat changes in real time by deploying a high-precision intelligent camera network, and constructs a spatio-temporal dynamic model of cross-regional species distribution based on the federated learning framework. Manual interpretation is crucial in the data closed-loop: professional personnel combine the AI pre-screening results to accurately label and correct the distribution coordinates, population quantity and activity rhythms of rare species, effectively solving the problem of false detection in complex environments. Such data provides scientific support for protection decisions - by analyzing the heat map to demarcate high-incidence poaching areas and optimize the patrol routes; formulating artificial intervention strategies based on the breeding cycle data, enabling the wild population of endangered species such as crested ibis to increase by more than 15% annually. Significance elevation: The integration of coordinate and quantity data not only helps the dynamic adjustment of the ecological red line, but also promotes the co-construction and sharing of the global biodiversity database (such as GBIF2025), providing a quantitative basis for the implementation of the goal.
[0003] To improve the recognition efficiency and accuracy of predefined targets, technologies such as super-resolution imaging, dynamic tracking, and multi-modal data fusion have been introduced into video surveillance systems. For example, the DJI "Eco Eye 4.0" drone is equipped with an 8K infrared-visible light dual-mode camera, with a night monitoring accuracy of 0.01 lux, supporting 120x digital zoom, and capable of identifying the fur texture of snow leopards 5 kilometers away. The YOLOv9-Wild model is used to fuse time series analysis to automatically correlate individual activity trajectories (e.g., the reconstruction error of the migration path of a golden snub-nosed monkey group is less than 10 meters). Based on the interpretation and protection decision-making mechanism of this type of data, extended applications can include population density heat maps, individual behavior tags, and gene sampling coordinates, etc. It can mark the overlapping areas between core habitats and human activities, mark the abnormal behaviors of rare wild animals (such as injuries, struggles in poaching traps), and can also guide non-invasive DNA collection (hair / feces). These technical methods have already had relevant successful application cases, such as the planning of the giant panda corridor in the Qinling Mountains, with a 40% increase in the road network avoidance rate, the rescue response time of Hainan gibbons shortened to 30 minutes, and the genetic diversity database of South China tigers expanded to 500 samples.
[0004] Application No. 2021107703652 discloses a method and device for geolocation of target objects based on video surveillance. Under the monitoring of tower cameras and video poles, according to the geodetic plane coordinates, elevation, vertical azimuth angle, horizontal azimuth angle, vertical field of view angle, and horizontal field of view angle of the camera, the geolocation of the target object is calculated, thus solving the problems of target positioning and automatic early warning in natural resource monitoring. This method for calculating the geolocation of the object cannot achieve batch calibration of the positions of specific types of objects in a professional scenario, such as large-scale, rapid, and inexpensive rapid classification and calibration of the position trajectories of wild animals and plants, nor can it calibrate the distance of any point within the field of view. Summary of the Invention
[0005] Aiming at the above existing technical problems, the purpose of the present invention is to provide a method, system, and storage medium for in-field space calibration based on depth of field. According to the high-altitude monitoring point positions, camera focal lengths, and declination angle data, using the basic principle of monocular vision, the distribution coordinate information of target objects and events is dynamically calculated, and the earth's surface within the high-precision field of view is fitted, providing an accurate, real-time, and rapid decision-making basis for subsequent natural resource emergency management operations.
[0006] The technical solution of the present invention is as follows:
[0007] A method for in-field space calibration based on depth of field, comprising the following steps:
[0008] S01: Obtain image data and detect the target object;
[0009] S02: Calculate the corresponding horizontal and vertical ratio values of the image according to the pixel area value of the target object in the image;
[0010] S03: Calculate the relative distance and azimuth angle between the target object in the image and the camera based on the principle of monocular vision;
[0011] S04: Obtain the absolute spatial coordinates of the target object according to the image pose information and the camera position coordinates. Fit the surface equation of the large ground within the camera's field of view through a quadratic surface equation, and obtain the absolute spatial coordinates of any point according to the surface equation of the large ground within the camera's field of view.
[0012] In the preferred technical solution, step S01 includes obtaining the camera's spatial pose and coordinate data at the moment of shooting when collecting the image, calculating the horizontal azimuth angle of the camera at the moment of shooting, the clockwise included angle between the horizontal projection line of the camera and the due north direction; the included angle between the vertical projection of the camera and the vertically upward direction, and the horizontal roll angle of the camera's y-axis pose.
[0013] In the preferred technical solution, the detection of the target object in step S01 includes:
[0014] Collect materials for the target category to be detected, label it, form a JSON-format annotation file, convert it into a YOLO-format TXT-encoded annotation file, submit it to the YOLO training environment for training, and obtain a weight file with the extension PT; if the training is successful, deploy this weight file to the periodic detection module, call the PT weight file, and detect the predetermined category in the image.
[0015] In the preferred technical solution, the periodic detection module in step S01 includes an improved YOLOv8 neural network. The improved YOLOv8 neural network includes a backbone network, a neck network, and a head network; the backbone network extracts multi-scale features through an improved CSPDarknet structure to generate a high-level feature map with high semantic information; the neck network uses a PAN-FPN structure to fuse feature maps of different levels to enhance the multi-scale detection ability; the head network includes decoupled detection heads to complete target classification, bounding box regression, and confidence prediction respectively;
[0016] The backbone network includes a C2f module and a spatial pyramid pooling fast module. The C2f module optimizes the feature extraction ability through cross convolution and dual filter design. The C2f module contains multiple Bottleneck blocks, and each block consists of a 1×1 convolution for dimensionality reduction, a 3×3 depthwise separable convolution, and a residual connection; the spatial pyramid pooling fast module realizes multi-scale feature fusion by cascading multiple max pooling layers. It retains spatial information of different granularities through parallel pooling and splicing operations, and uses depthwise separable convolution in some convolutional layers to separate the spatial filtering and channel combination processes;
[0017] The neck network enhances the multi-scale object detection ability of the model through multi-path feature fusion. Through the top-down path of the bidirectional feature pyramid, the high-level features are upsampled and added element-wise to the middle and low-level features to enhance the semantic information of the shallow features. The bottom-up path downsamples the low-level features and fuses them with the middle and high-level features. During the attention mechanism fusion process, channel attention or spatial attention is introduced to dynamically adjust the weights of different channels or spatial positions, highlighting the key feature regions, and outputting fused feature maps of three scales.
[0018] The decoupled head structure of the head network predicts the class probability distribution of each anchor point through a fully connected layer or 1×1 convolution, uses Focal Loss to alleviate the class imbalance problem, predicts the offset of the bounding box through the regression branch, uses CIoU Loss to optimize the position and size accuracy of the box, and outputs the confidence score of the existence of the target through the confidence branch, and trains in combination with the binary cross-entropy loss.
[0019] Enhanced processing is performed on the key categories. Based on the known value of the standard size of the object of this class, the distance and azimuth where the object is located are obtained, and the detection boxes are screened.
[0020] In the preferred technical solution, the method for calculating the relative distance between the target object and the camera in step S03 includes:
[0021] Screen the detection box with the largest pixel ratio;
[0022] Calculate the pixel height difference for the selected detection box: Δypx = y2 - y1;
[0023] Calculate the sensor scale factor:
[0024]
[0025] The relative distance D between the target object and the camera = k * f;
[0026] Where y2 is the y pixel of the upper left corner, y1 is the y pixel value of the lower right corner, CMOS_Length is the camera height, and f is the camera focal length;
[0027] Taking different objects as target objects respectively, multiple distance values are calculated, and then the weighted average is calculated.
[0028] In the preferred technical solution, the method for obtaining the absolute spatial coordinates of the target object in step S04 includes:
[0029] Using the plane approximation method, obtain the longitude coordinate offset between the target object and the camera position:
[0030]
[0031] Obtain the latitude coordinate offset between the target object and the camera position:
[0032]
[0033] Among them, D is the relative distance between the target object and the camera, θ is the azimuth angle, and latorigin is the original latitude;
[0034] After superimposing the obtained longitude and latitude coordinate offsets of the target object and the camera position on the longitude and latitude coordinates of the camera, the absolute spatial coordinates of the target object are obtained.
[0035] In the preferred technical solution, the method for fitting the curved surface equation of the large ground within the camera's field of view in step S04 includes:
[0036] S41: Convert the longitude and latitude of the target object into a plane rectangular coordinate system, and map the earth's curved surface to a plane;
[0037] S42: Separate the geodetic height H and the normal height Hn, and calculate the height anomaly value:
[0038] ζ = H - Hn
[0039] S43: Construct a quadratic surface equation:
[0040] ζ = a0 + a1x + a2y + a3x2 + a4y2 + a5xy
[0041] Where: x, y are plane coordinates, and a0, a1,..., a5 are coefficients to be determined;
[0042] S44: Substitute each known point i into the quadratic surface equation to construct an error equation:
[0043] V i = a0 + a1x i + a2y i + a3x i2 + a4y i2 + a5x i y i - ζ i
[0044] S45: The design matrix A is a 6×6 matrix, and each row corresponds to the polynomial expansion of a point [1, x i , y i , x i2 , y i2 , x i y i ;
[0045] The parameter vector X = [a0, a1, a2, a3, a4, a5] T ;
[0046] Observation vector L = [ζ1, ζ2,..., ζ6] T ;
[0047] S46: Solve for the parameters through least - squares solution:
[0048] X = (A T A) -1 A T L
[0049] If the matrix (A T A) is invertible, calculate directly; if it is a singular matrix, use numerical optimization methods.
[0050] In the preferred technical solution, the method for obtaining the absolute spatial coordinates of any point according to the surface equation of the large ground within the camera's field of view includes:
[0051] Obtain the distance D and azimuth angle α between the target object and the camera within the field of view;
[0052] Calculate the distance component D_north in the north - south direction = D * cos(α);
[0053] Calculate the distance component D_east in the east - west direction = D * sin(α);
[0054] Calculate the latitude change
[0055] Calculate the longitude change
[0056] Target latitude
[0057] Target longitude λ_base = λ_c + Δλ;
[0058] Wherein, the latitude of the camera is φ_c, the longitude is λ_c, and R is the radius of the earth;
[0059] When the longitude and latitude values of any point are obtained, substitute them into the quadratic surface equation to obtain the elevation value of the object.
[0060] The present invention also discloses an in - field space calibration system based on depth of field, including:
[0061] A target detection module, which acquires image data and detects the target object;
[0062] A pixel calculation module, which calculates the corresponding horizontal and vertical aspect ratio values of the image according to the pixel area value of the target object in the image;
[0063] A distance calculation module, which calculates the relative distance and azimuth angle between the target object in the image and the camera based on the principle of monocular vision;
[0064] The fitting calculation module obtains the absolute spatial coordinates of the target object based on the image pose information and the camera position coordinates, fits the surface equation of the large ground within the camera's field of view through a quadratic surface equation, and obtains the absolute spatial coordinates of any point based on the surface equation of the large ground within the camera's field of view.
[0065] The present invention also discloses a computer storage medium, on which a computer program is stored, and when the computer program is executed, the above-mentioned in-field space calibration method based on depth of field is implemented.
[0066] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0067] First, it can utilize existing high-altitude monitoring camera resources: on existing conventional video monitoring and video recording resources, periodic key target detection can be carried out on video data, real-time video streams, and photographed pictures, and the coordinates, quantity, and movement trajectories of target objects can be detected in a timely manner. It is particularly suitable for monitoring sudden type events, such as natural resource field emergencies that require rapid response, such as forest fires, debris flows, and floods.
[0068] Second, it can be deployed on mobile devices: daily inspection personnel in the natural resources department can use this method to call the mobile camera to complete a more accurate on-site natural resource census and spot check, be able to synchronously upload inspection ledger pictures, and can use the algorithm to count important on-site natural resource data.
[0069] Third, it can fit the surface of the large ground within the field of view: the target point data obtained by high-altitude monitoring or mobile cameras can be used to fit the surface of the large ground within the point coverage range, and then the absolute coordinates, quantity, trajectory, distance, and other spatial data of any point within the coverage range can be obtained, providing a reliable decision-making basis for subsequent emergency command and dispatch. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The present invention will be further described below in conjunction with the drawings and embodiments:
[0071] Figure 1 is a flowchart of the in-field space calibration method based on depth of field of the present invention;
[0072] Figure 2 is a flow node diagram of the in-field space calibration method based on depth of field of the present invention;
[0073] Figure 3 is a schematic diagram of the target detection frame of the present invention;
[0074] Figure 4 is a flowchart of a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0075] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0076] Principle description: In order to effectively utilize the high-altitude monitoring video resources that have been widely used in the field of natural resource supervision, this solution proposes a method for detecting target objects of a predetermined hot-spot category using the YOLO V8 framework, and can calculate the spatial distance of the target objects based on the absolute coordinates of the camera, thereby calculating the absolute coordinates of each predetermined category target within the field of view. At the same time, this method also supports mobile cameras and can use the mobile devices assigned to natural resource inspection personnel to achieve high-precision resource census within the supervision area. At the same time, based on this, the curved surface equation of the large ground within the camera's field of view can be derived, so that the absolute coordinates of any point can be realized. Accordingly, a complete set of calculation methods and software processes are also proposed, which are characterized by being inexpensive, efficient, and easy to deploy. When supervision events such as forest fires, debris flows, or wildlife activities occur, the coordinates, quantity, and movement trajectories of the target objects can be quickly detected, providing a basis for high-real-time data decision-making for subsequent law enforcement intervention.
[0077] As Figure 1 shown, a method for calibrating the in-field space based on the depth of field includes the following steps:
[0078] S01: Obtain image data and detect the target object;
[0079] S02: Calculate the corresponding horizontal and vertical aspect ratio values of the image according to the pixel area value of the target object in the image;
[0080] S03: Calculate the relative distance and azimuth angle between the target object in the image and the camera based on the principle of monocular vision;
[0081] S04: Obtain the absolute spatial coordinates of the target object according to the image attitude information and the camera position coordinates, fit the curved surface equation of the large ground within the camera's field of view through the quadratic surface equation, and obtain the absolute spatial coordinates of any point according to the curved surface equation of the large ground within the camera's field of view.
[0082] In this solution, it is first necessary to effectively detect specific target objects, such as being able to detect fires, vehicles, and people. Secondly, based on the pixel area value of the target object in the picture, calculate the corresponding horizontal and vertical aspect ratio values of the picture. Thirdly, according to the focal length value when the picture is taken, and based on the basic principle of monocular vision, calculate the relative position values (distance, coordinates, and azimuth angle) of the target object from the camera position. Finally, convert to the geodetic coordinate values (such as GPS coordinates) of the target object according to the camera position coordinates. After implementing the above calculation process for multiple objects in the distance within the field of view, the coordinate values of each object on the ground within the field of view can be deduced. If it is assumed that the objects are located on the ground, an approximate value of the ground space surface within the field of view can be deduced. In this way, it is possible to provide dynamic and implemented target distribution and activity trajectory values for subsequent natural resource emergency response and dispatching command. In the above process, this method will propose a supporting data organization method for the database, store the spatial attitude and geodetic coordinate values of the camera at the moment the photo is taken, and store the surrounding environment parameters at the moment of photo data acquisition, such as illuminance, temperature, geomagnetic angle, nine-axis acceleration, and other data. At the same time, a method for collecting and recording spatial attitude data parameters that can be used on mobile devices is also established, which can be transmitted to the remote server side and cooperate with the high-altitude monitoring points of natural resources to calculate the spatial position, height, and quantity data of natural resource objects collected by the mobile camera.
[0083] In this solution, the improved YOLO V8 target detection framework is used to complete the detection of targets. The specific process is as Figure 2 shown. First, define the target categories for pre-detection (Class List, Table 2). In this solution, people, vehicles, lamp posts, cameras, traffic signs, wild animals (wild boars, red-bellied squirrels, water buffalo), and natural resource events (fires, smoke, deforestation, slope collapses) are defined. The sizes of these objects are relatively fixed. Especially for road public transportation facilities, there are mandatory standards such as national standards that can be referred to for their sizes, which can be used as effective calibration objects. When collecting photos, the spatial attitude and coordinate data of the camera at the moment of shooting will be stored in the data table yolo_app_sensor, as shown in Table 1 below.
[0084] Table 1
[0085]
[0086]
[0087] Next, collect various types of image data in the Class List, perform category calibration and create training materials, and train under the YOLO V8 framework to finally obtain a weight file. Based on this weight file, establish an object detection module that periodically polls images in a fixed directory, perform object detection on each frame of the photo for the preset objects in the Class List, and store the detection results in the MYSQL data table. At the same time, start the subsequent spatial coordinate calculation module for the target object and store the calculation results in the MYSQL data table.
[0088] The improved YOLO V8 object detection algorithm directly predicts the bounding boxes and class probabilities of objects through end-to-end forward propagation. Its algorithm process can be divided into the following core stages: Input preprocessing, the input image undergoes operations such as size normalization (Letterbox processing), color space conversion (BGR→RGB), and normalization (mapping pixel values to the range of 0-1) to form a standardized tensor of a fixed size (such as 640×640).
[0089] Backbone feature extraction: YOLO V8 extracts multi-scale features through an improved CSPDarknet structure to generate high-level feature maps with high semantic information. Neck feature fusion: YOLO V8 uses the PAN-FPN structure to fuse feature maps of different levels to enhance multi-scale detection capabilities. Head prediction output: The decoupled detection head separately completes object classification, bounding box regression, and confidence prediction, and combines a dynamic label assignment strategy to optimize the training process. Post-processing (NMS): Redundant detection boxes are filtered through non-maximum suppression (NMS) to output the final detection results.
[0090] Core logic of the Backbone
[0091] The Backbone of YOLOv8 is based on CSPDarknet, and optimizes gradient flow and computational efficiency through modular design.
[0092] 1. Core components
[0093] The C2f module optimizes the feature extraction ability through cross-convolution and dual-filter design. The C2f module contains multiple Bottleneck blocks, each block consists of a 1×1 convolution for dimensionality reduction, a 3×3 depthwise separable convolution, and a residual connection to achieve lightweight and efficient feature reuse.
[0094] The structure is shown as: Input→Split→Bottleneck×N→Concat→Output
[0095] The SPPF module is the Spatial Pyramid Pooling Fast (SPPF) module that achieves multi-scale feature fusion by cascading multiple max pooling layers, significantly enhancing the receptive field. It retains spatial information at different granularities through parallel pooling and concatenation operations. Depthwise Separable Convolution is adopted in some convolutional layers to separate the spatial filtering and channel combination processes, reducing the number of parameters and improving the inference speed.
[0096] 2. Feature Hierarchy Design
[0097] The Backbone outputs feature maps at three levels: Low-level features (shallow): High resolution (e.g., 80×80), capturing detailed information (e.g., edges, textures). Middle-level features: Balancing semantic and spatial information (e.g., 40×40). High-level features (deep): Low resolution (e.g., 20×20), containing strong semantic information (e.g., object categories).
[0098] Feature Fusion Strategy of the Neck Network The Neck network enhances the model's detection ability for multi-scale targets through multi-path feature fusion, and the core structure is PAN-FPN (Path Aggregation Network + Feature Pyramid Network).
[0099] 1. Bidirectional Feature Pyramid (BiFPN) The top-down path upsamples the high-level features and adds them element-wise to the middle and low-level features to enhance the semantic information of the shallow features. The bottom-up path downsamples the low-level features and fuses them with the middle and high-level features to improve the spatial resolution of the deep features.
[0100] 2. Attention Mechanism Fusion Channel attention (e.g., SE Block) or spatial attention (e.g., CBAM) is introduced during the feature fusion process to dynamically adjust the weights of different channels or spatial positions, highlighting key feature regions.
[0101] 3. Multi-scale Output The Neck outputs fused feature maps at three scales (e.g., 80×80, 40×40, 20×20), corresponding to the detection tasks of small, medium, and large targets respectively.
[0102] Prediction Mechanism of the Head Network YOLOv8 adopts a decoupled head, separating the classification and regression tasks to improve the prediction accuracy.
[0103] 1. Decoupled Head Structure The classification branch predicts the class probability distribution of each anchor point through a fully connected layer or 1×1 convolution, and uses Focal Loss to alleviate the class imbalance problem.
[0104] Regression branch: Predicts the offsets of the bounding box (Δx, Δy, Δw, Δh), and uses CIoU Loss to optimize the position and size accuracy of the box.
[0105] Confidence branch: Outputs the confidence score of the existence of the target, and is trained in combination with Binary Cross Entropy Loss (BCE Loss).
[0106] 2. Dynamic label assignment: TaskAlignedAssigner dynamically assigns positive and negative samples according to the class scores and IoU values of the predicted boxes and the ground truth boxes, avoiding the mis-matching problems caused by traditional static threshold assignment.
[0107] The calculation formula is: Alignment Metric = Class Score × IoU
[0108] Select the sample with the highest alignment metric as the positive sample to improve the training efficiency.
[0109] 3. Anchor-free mechanism
[0110] YOLOv8 adopts an anchor-free design, directly predicting the offsets of the object center point and width / height, simplifying the model structure and reducing the dependence on hyperparameters.
[0111] Training and inference optimization techniques
[0112] The adaptive training strategy dynamically adjusts the weights of the loss function according to the target scale, enhancing the detection ability of small targets. Mixed precision training uses FP16 / FP32 mixed precision to accelerate training, balancing the calculation speed and numerical stability. Model distillation compresses the model size through the teacher-student network framework, improving the detection accuracy of the lightweight version.
[0113] There are three key fields in Data Table 1. azimuth records the horizontal azimuth angle of the camera at the moment of taking a picture, and its value is the numerical value of the clockwise included angle between the horizontal projection line of the camera and the due north direction; pitch is the included angle value between the vertical projection of the camera and the vertically upward direction; the roll field records the horizontal roll angle of the camera's y-axis attitude. The values of these three fields will be used to calculate the relative coordinate values of the objects in the picture with respect to the camera position point under the monocular vision mechanism.
[0114] According to the data requirements of the yolo_app_sensor table structure in Table 1, the image data is transmitted from the mobile device or the high-position monitoring camera to the YOLO V8 server in the Windows environment through the TCP / IP protocol. This server uses the Anaconda platform, runs in the Python environment, and installs the CUDA video inference platform. In the example environment of this solution, the NVIDIA RTX 3070 platform is used, with pytorch-cuda = 12.4 and the Python version is 3.9. Through a target detection module with a fixed time period, the image data from the mobile device or the high-position monitoring camera is subjected to the target detection task of the predefined categories in Table 2, and the image detection results are stored in Table 3.
[0115] Table 2
[0116]
[0117] Table 3
[0118]
[0119] In the above table, class is the category to which the object belongs after target detection, and its value is the target category (ClassList) in Table 2. For example, in the image to be detected captured by the daily monitoring camera, if the algorithm detects a fire event, the value of this field is 36 (the corresponding serial number 36 for fire in Table 2). When there are objects of the same type in this image, such as multiple fire points, different serial numbers need to be assigned to each object, and this serial number is stored in the label_id field. The confidence field records the confidence of the object corresponding to label_id, that is, the accuracy considered by the algorithm, and its value is a floating-point value between 0 and 1. 1 is considered 100% accurate, and 0 means it is not considered accurate. The meaning of the topleft_box_x field is the x pixel value of the upper left corner of the target detection box, the meaning of the topleft_box_y field is the y pixel value of the upper left corner of the target detection box, the meaning of the bottomright_box_x field is the x pixel value of the lower right corner of the target detection box, the meaning of the bottomright_box_y field is the y pixel value of the lower right corner of the target detection box, the meaning of the bbox_width field is the target width pixel value of the target detection box, and the meaning of the bbox_height field is the target height pixel value of the target detection box. Specifically, as follows Figure 3 As shown, topleft_box_x is the x pixel value of the upper left corner of the person box, and topleft_box_y is the y pixel value of the upper left corner of the person box.
[0120] In Table 4 Lamp_pole_run, the values of the top-left x pixel, top-left y pixel, bottom-right x pixel, and bottom-right y pixel will be read according to the fields topleft_box_x, topleft_box_y, bottomright_box_x, and bottomright_box_y in Table 3 detection_results and stored in the fields Pole_Dev_ClassID_X1, Pole_Dev_ClassID_Y1, Pole_Dev_ClassID_X2, and Pole_Dev_ClassID_Y2, and the pixel values of the length and width of the detection box will be calculated. Since the sizes of known objects such as people and traffic signs are relatively fixed, the relative position between the target object and the camera can be converted according to the camera focal length value at this time. The calculation method is as follows.
[0121] D = f * Length / (Pole_Dev_ClassID_Y1 - Pole_Dev_ClassID_X1).
[0122] D is the straight-line distance between the camera and the object, f is the camera focal length, Length is the actual length of the object, and (Pole_Dev_ClassID_Y1 - Pole_Dev_ClassID_X1) is the imaging height.
[0123] Table 4
[0124]
[0125]
[0126] Above Figure 4 In this method, a spatial positioning method for target detection of natural resource supervision objects that can be compatible with both high-altitude monitoring cameras and mobile phone cameras is proposed: First, collect materials for the target categories to be detected and label them using tools such as labelme. After the labeling is completed, a JSON-format labeling file will be formed. At this time, it needs to be converted into a YOLO-format TXT-encoded labeling file and submitted to the YOLO training environment for training to obtain a weight file with the extension PT. At this time, the training accuracy of the PT file needs to be checked. If the loss function value is not accurately fitted, it should be considered that the accuracy is abnormal. At this time, it needs to be returned to the category labeling link to re-check the labeling.
[0127] If the training is successful, deploy the weight file to the periodic detection module. Meanwhile, the image preprocessing module records the spatial attitude data at the moment of photo shooting into the data table 1yolo_app_sensor, and stores the image file in a predetermined directory. On this basis, the fixed-period target detection module calls the PT weight file to detect the items of the predetermined categories in the image. These items belong to the pre-judgment and classification components in natural resource management, so their appearance dimensions have been obtained. After completing the target detection, the horizontal and vertical pixel coordinates of the target in the image can be obtained. At this time, the length of the target object in the photo can be calculated. Based on the principle of monocular vision, the relative distance and azimuth angle between the target object in the image and the camera can be calculated. At this time, by superimposing the momentary attitude of the camera when taking the photo and its absolute coordinates, the absolute spatial coordinates (longitude, latitude, and elevation) of the object can be calculated and the calculation results are stored in the data table 4Lamp_pole_run.
[0128] In the above flowchart, the running logic is divided into several main parts: pre-training, model initialization, database connection, image processing, target detection, distance calculation, coordinate transformation, database storage, etc. Each part needs to be described in detail, and the mathematical methods or algorithms used, such as the Haversine formula, planar approximation calculation, the target detection principle of YOLO, etc., should be pointed out.
[0129] This system aims to achieve the automatic detection, distance estimation, and spatial positioning of road facilities and the attached facilities of natural resource components through computer vision technology and geographic information processing. The system integrates the following key technology modules:
[0130] Target detection: Based on the YOLOv8 model, identify the main pole and maintenance hatch in the image;
[0131] Distance calculation: Calculate the distance between the target object and the shooting point using the camera imaging principle;
[0132] Coordinate transformation: Convert the pixel space measurement values into the geographic coordinate system;
[0133] Data persistence: Store the detection results and spatial information through a relational database.
[0134] Phase 1: System initialization
[0135] Before completing the derivation of this solution, it is necessary to consider that different models of high-altitude cameras and mobile imaging sensors have different sizes, resulting in huge errors in the conversion of target distance and coordinates. Therefore, it is necessary to convert the horizontal and vertical dimensions of the sensor into the standard 35mm camera sensor size, which is specifically as follows:
[0136] Sensor parameter setting
[0137] Preset parameters according to the physical characteristics of the camera:
[0138] CMOS sensor size: 24mm (width) × 36mm (length)
[0139] Image resolution: 3000 × 4000 pixels
[0140] Initial value of equivalent focal length: 27mm (subsequently obtained dynamically from EXIF)
[0141] Loading of the target detection model
[0142] Adopt the YOLOv8 architecture and load the pre-trained weight file (best.pt). The model has been customized and trained to recognize 39 types of road facilities and natural resource event part information (including main pole class = 0, slope collapse class = 38).
[0143] Database connection and table structure initialization
[0144] Use the MySQL relational database to create a storage table containing multi-dimensional spatial information:
[0145] Geographical coordinate fields: latitude and longitude of the original shooting point, deduced coordinates, coordinate deviation values
[0146] Detection result fields: target category, bounding box coordinates, pixel height difference
[0147] Metadata fields: timestamp, image file name, device serial number
[0148] Stage 2: Image data acquisition and preprocessing
[0149] Obtain the images to be processed and their metadata. The specific steps are as follows:
[0150] File traversal and screening. Traverse the specified directory (such as D: / LTXL2) and screen standard image format files such as JPG / PNG. Use the suffix name matching algorithm to achieve fast filtering.
[0151] EXIF metadata parsing. Extract key shooting parameters from the image file:
[0152] Equivalent focal length: Obtain it by reading the EXIF data from the FocalLengthIn35mmFilm tag value;
[0153] Exception handling: Use the preset value when the tag is missing to avoid calculation interruption;
[0154] Geographical information association query. Select the image files already taken by the camera in the predetermined folder. For example, execute an SQL fuzzy query in the database for the file name (such as IMG_20230501_001.jpg):
[0155] SELECT azimuth,latitude,longitude FROM yolo_app_sensor WHERE filenameLIKE'%IMG_20230501_001%'
[0156] After the query statement is executed, the three elements at the moment when the picture was taken will be obtained:
[0157] Azimuth: The angle between the shooting direction and the due north (0° - 360°)
[0158] Latitude: In decimal format (e.g., 31.3416985970769)
[0159] Longitude: In decimal format (e.g., 119.799581757554)
[0160] The calculation method of the three elements is as follows: The azimuth uses the clockwise angle system with the geographic north as the reference, and the longitude and latitude accuracy is reserved to 10 decimal places (about 1.11 mm accuracy).
[0161] Phase 3: Object Detection and Result Analysis
[0162] Objective: Identify the target objects in the image and extract geometric features
[0163] Processing flow:
[0164] The YOLOv8 inference execution is to input the image into the neural network and perform the following operations:
[0165] Feature extraction: Extract multi-scale features through the Darknet-53 backbone network;
[0166] Object localization: Generate prediction boxes in three detection heads (80×80, 40×40, 20×20);
[0167] Non-Maximum Suppression (NMS): IoU threshold 0.45, remove redundant detection boxes.
[0168] The detection result structuring is to use the Supervision library to parse the output:
[0169] Class ID array: Identify the detected object classes (e.g., [0,31]);
[0170] Bounding box coordinates: Normalized format (x1,y1,x2,y2), unit pixel;
[0171] Confidence score: The reliability evaluation value predicted by the model (0 - 1).
[0172] The target screening strategy is to perform enhanced processing on key categories. Based on the known value of the standard size of objects in this category, the distance and orientation where the object is located can be deduced inversely. To avoid multiple target objects of the same type being detected in the same picture, thus interfering with the calculation accuracy, in this example, the main pole and the road nameplate closest to the camera can be selected as the calculation examples:
[0173] Main pole (Class 0): Select the detection box with the largest vertical span;
[0174] Road nameplate (Class 23): Select the detection box with the largest horizontal span;
[0175] Calculation method:
[0176] Calculation of the height difference of the bounding box: Δy = y2 - y1
[0177] Target screening formula: selected_index = argmax(Δy)
[0178] Stage 4: Geometric distance calculation
[0179] Objective: Convert pixel measurement values into physical space distances
[0180] The basic principle is to establish a two-dimensional projection relationship based on the pinhole camera model and the principle of similar triangles:
[0181]
[0182] Calculation steps:
[0183] 1. Pixel - physical conversion
[0184] Perform calculations on the selected detection box:
[0185] Pixel height difference: Δypx = y2 - y1
[0186] Sensor scale factor:
[0187]
[0188] CMOS_Length is the height of the CMOS sensor, which is the height of the camera;
[0189] Actual distance estimation:
[0190] Example parameters when Δy = 300px, CMOS = 36mm, focal length = 27mm, road nameplate height = 0.8m:
[0191]
[0192] Multi-target collaborative verification
[0193] Taking the lamp post and the maintenance hatch as the calculation objects respectively, after obtaining two different distance values Dpole and Ddoor, weighted averaging is then performed to improve the accuracy of the distance value. When the main pole and the hatch are detected simultaneously, the weighted average method is used to improve the accuracy:
[0194] Dfinal = 0.7 × Dpole + 0.3 × Ddoor
[0195] The weight coefficient is set based on the stability of the target size.
[0196] Phase 5: Geographic coordinate mapping
[0197] The goal of this phase is to convert the distance and azimuth angle into geographic coordinates. Since the field of view captured by the camera is very small relative to the Earth, the Earth's curved surface within the field of view can be regarded as a plane here. At this time, the plane approximation method can be used. For low-precision scenarios in a small range (<1 km), a simplified formula is used:
[0198]
[0199] Parameter description:
[0200] 111111: Arc length per degree of the equator (meters), θ: Azimuth angle (in radians), latorigin: Original latitude (in radians).
[0201] After the above process is completed, the longitude and latitude coordinate offsets of the target object and the camera position are obtained. At this time, after adding the longitude and latitude coordinates of the camera, the coordinate position of the target object can be obtained, and the calculation result of the longitude and latitude coordinates of the target can be written into the data table.
[0202] Phase 6: Data storage and visualization
[0203] Since the maximum imaging mechanism is set, it is possible to avoid detecting multiple targets of the same type in the same picture. At this time, a duplicate writing prevention mechanism can be designed based on the uniqueness of the picture name. When specifically executing, a priori query needs to be performed, and the sql statement is as follows: SELECT COUNT(*) FROM Lamp_Pole_Run WHERE Pole_PicName = 'IMG_001.jpg' AND Pole_PicName_latitude = 31.3417.
[0204] Decide the INSERT or UPDATE operation according to the return result (as shown in the flowchart in Figure 4 ).
[0205] Based on the above results, the encoding operation of the spatial data is further completed, and the quality of the calculation results is further optimized and quantified:
[0206] Coordinate difference: Elat = (latAI - latbase) × 10 6 (ppm)
[0207] Distance deviation: ΔD = ∣Dcalc - DGPS∣
[0208] Bounding box information: Store the normalized coordinates (x1, y1, x2, y2)
[0209] latAI is the longitude value obtained by the AI method, latbase is the measured longitude value, Dcalc is the distance value obtained by the AI method, and DGPS is the measured distance value.
[0210] In the last step, according to the absolute spatial coordinate points (longitude, latitude, and elevation) of the objects within the camera's field of view obtained from the data table 4Lamp_pole_run, the terrain surface within the camera's field of view can be drawn based on the basic principle of the quadratic surface. In this method, at least 6 absolute spatial coordinate points (longitude, latitude, and elevation) of the objects are required to fit the geoid surface within their coverage range through the quadratic surface equation. The specific steps are as follows:
[0211] 1. Coordinate system conversion and data preprocessing
[0212] Conversion of longitude and latitude to plane coordinates: Convert the longitude and latitude (B, L) of the six points to plane rectangular coordinates (such as UTM projection coordinates x, y). The UTM projection maps the Earth's surface to a plane, facilitating subsequent plane model calculations.
[0213] Elevation processing: If it is necessary to separate the geodetic height and the normal height, calculate the elevation anomaly value:
[0214] ζ = H - Hn
[0215] The geodetic height H is the elevation based on the reference ellipsoid surface, that is, the vertical distance from a point to the Earth's ellipsoid surface. The geodetic height is directly measured by GNSS (such as GPS) and only depends on the geometric ellipsoid model and has nothing to do with the Earth's gravity field.
[0216] The normal height Hn is the elevation based on the quasi-geoid surface, that is, the vertical distance from a point to the quasi-geoid surface. The normal height is calculated through leveling measurement combined with gravity data, reflecting the influence of the Earth's gravity field, and is a commonly used elevation reference in engineering.
[0217] 2. Quadratic surface mathematical model Quadratic surface equation:
[0218] ζ = a0 + a1x + a2y + a3x2 + a4y2 + a5xy
[0219] Where: x, y are plane coordinates, and a0, a1,..., a5 are coefficients to be determined;
[0220] 3. Coefficient Solving and Algorithm Implementation
[0221] (1) Constructing the Error Equation
[0222] For each known point i (a total of 6 points), substitute it into the quadratic surface equation to construct the error equation:
[0223] V i = a0 + a1x i + a2y i + a3x i2 + a4y i2 + a5x i y i - ζ i
[0224] Where:
[0225] Design matrix A: a 6×6 matrix, and each row corresponds to the polynomial expansion of a point [1, x i , y i , x i2 , y i2 , x i y i ;
[0226] Parameter vector X = [a0, a1, a2, a3, a4, a5] T ;
[0227] Observation vector L = [ζ1, ζ2,..., ζ6] T ;
[0228] (2) Least Squares Solution Solve the parameters through the normal equation:
[0229] X = (A T A) -1 A T L
[0230] If the matrix (A T A) is invertible, calculate directly; if the matrix is a singular matrix, use numerical optimization methods.
[0231] The above quadratic surface method is applicable to areas with gentle terrain undulation, and the fitting range is recommended to be less than 100 km 2 , and the higher the accuracy of the surface fitting as the number of measurement points increases.
[0232] Example Data
[0233] Assume that the plane coordinates (x, y) and height anomaly values (ζ) of six known points are given, and the geoid needs to be fitted through the quadratic surface equation. The data is as follows:
[0234] Dot x (m) y (m) ζ (m) 1 100 200 5.2 2 200 300 6.8 3 300 400 8.1 4 400 500 9.5 5 500 600 10.9 6 600 700 12.4
[0235] Constructing the design matrix and observation vector
[0236] According to the quadratic surface equation:
[0237] ζ=a0+a1x+a2y+a3x2+a4y2+a5xy
[0238] Construct the design matrix A and observation vector L for each point:
[0239] Design matrix A (6×6):
[0240]
[0241] Observation vector L(6×1):
[0242] 2. Least squares solution parameters
[0243] Solve the coefficient vector X through the normal equations:
[0244] X=(A T A) -1 A T L
[0245] (1) Calculate A T A
[0246]
[0247] (2) Calculate A T L
[0248]
[0249] (3) Solving the inverse matrix and coefficients
[0250] Through matrix operations (the actual calculation requires the help of numerical tools):
[0251]
[0252] 3. Substituting the coefficients into the quadratic surface equation, we get:
[0253] ζ=4.000+0.010x+0.005y+0.00002x2+0.00001y2+0.000005xy
[0254] The above formula is that the elevation outliers of the six points are perfectly fitted into a quadratic surface (the residual is close to 0) through the least squares method. This model is suitable for areas with linear and gentle changes in elevation terrain within the camera's field of view.
[0255] According to the distance and azimuth angle between the ground object and the camera within the field of view, substituting into the quadratic surface equation, the corresponding elevation value can be obtained.
[0256] Since the distance D within the camera's field of view is much smaller than the Earth's radius R (R is the Earth's radius, approximately 6,378,137 meters), a plane approximation can be used, considering the Earth's surface as a plane. This makes it simpler to calculate the latitude variable Δφ and the longitude variable Δλ. The azimuth angle α usually refers to the angle measured clockwise from the due north direction. Therefore, when calculating the changes in longitude and latitude, D needs to be decomposed into two components: north-south and east-west. The north-south component corresponds to the change in latitude, while the east-west component corresponds to the change in longitude.
[0257] The change in latitude Δφ is obtained by D * cos(α) / R because the arc length corresponding to one degree of latitude is R * π / 180 meters. Thus, the north-south distance D_north = D * cos(α), and the corresponding change in latitude is D_north / (R) radians, which is then converted to degrees. Similarly, the east-west distance D_east = D * sin(α), and the corresponding change in longitude is D_east / (R * cosφ_c) radians, which is also converted to degrees. The arc length corresponding to each degree of longitude at latitude φ is R * cosφ * π / 180 meters.
[0258] Here, the conversion between radians and meters is used. Since the units of Δφ and Δλ are radians, they need to be converted to degrees by multiplying by 180 / π.
[0259] The following example: If the latitude of the camera is φ_c, the longitude is λ_c, the azimuth angle α = 45 degrees, and the distance D = 20 meters, then Δφ = (20 * cos45°) / 6378137. Here, the obtained value is in radians and needs to be converted to degrees. Similarly, Δλ = (20 * sin45°) / (6378137 * cosφ_c), which is also converted to degrees and then added to the camera's longitude and latitude to obtain the longitude and latitude of the base.
[0260] The above process can be explained step by step, including:
[0261] 1. Convert the azimuth angle α to radians (if necessary).
[0262] 2. Calculate the north-south distance component D_north = D * cos(α)
[0263] 3. Calculate the east-west distance component D_east = D * sin(α)
[0264] 4. Calculate the change in latitude (in radians), and then convert it to degrees
[0265] 5. Calculate the change in longitude (in radians), and then convert it to degrees
[0266] 6. Target latitude
[0267] 7. Target longitude λ_base = λ_c + Δλ
[0268] After obtaining the latitude and longitude values of any point, the elevation value of the object can be obtained by substituting them into the quadratic surface equation. Thus, the coordinates, quantity, and trajectory of specific category objects within the field of view can be calibrated.
[0269] In another embodiment, a system for calibrating the in-field space based on depth of field includes:
[0270] A target detection module that acquires image data and detects target objects;
[0271] A pixel calculation module that calculates the corresponding horizontal and vertical aspect ratio values of the image based on the pixel area value of the target object in the image;
[0272] A distance calculation module that calculates the relative distance and azimuth angle between the target object in the image and the camera based on the principle of monocular vision;
[0273] A fitting calculation module that obtains the absolute space coordinates of the target object according to the image pose information and the camera position coordinates, fits the surface equation of the large ground within the camera's field of view through the quadratic surface equation, and obtains the absolute space coordinates of any point according to the surface equation of the large ground within the camera's field of view.
[0274] The specific implementation of each module is the same as the above method for calibrating the in-field space based on depth of field, and will not be elaborated here.
[0275] In yet another embodiment, a computer storage medium stores a computer program, and when the computer program is executed, it implements the above method for calibrating the in-field space based on depth of field.
[0276] The above method for calibrating the in-field space based on depth of field is adopted, and will not be elaborated here.
[0277] It should be understood that the above specific embodiments of the present invention are only used for exemplary illustration or explanation of the principle of the present invention, and do not constitute a limitation to the present invention. Therefore, any modifications, equivalent substitutions, improvements, etc. made without departing from the spirit and scope of the present invention shall be included within the protection scope of the present invention. In addition, the appended claims of the present invention are intended to cover all changes and modification examples falling within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.
Claims
1. A method for calibrating the in-field space based on the depth of field, characterized in that, It includes the following steps: S01: Obtain image data and detect the target object; S02: Calculate the corresponding horizontal and vertical ratio values of the image according to the pixel area value of the target object in the image; S03: Calculate the relative distance and azimuth angle between the target object in the image and the camera based on the monocular vision principle; S04: Obtain the absolute spatial coordinates of the target object according to the image pose information and the camera position coordinates. Fit the surface equation of the large ground in the camera's field of view through a quadratic surface equation, and obtain the absolute spatial coordinates of any point according to the surface equation of the large ground in the camera's field of view.
2. The in-field space calibration method based on depth of field according to claim 1, wherein Step S01 includes obtaining the camera's spatial pose and coordinate data at the moment of shooting when collecting the image, calculating the horizontal azimuth angle of the camera at the moment of shooting, the clockwise included angle between the horizontal projection line of the camera and the due north direction; the included angle between the vertical projection of the camera and the vertically upward direction, and the horizontal roll angle of the camera's y-axis pose.
3. The in-field space calibration method based on depth of field according to claim 1, wherein Step S01 for detecting the target object includes: Collect materials for the target category to be detected, label it, form a JSON-format annotation file, convert it into a YOLO-format TXT-encoded annotation file, submit it to the YOLO training environment for training, and obtain a weight file with the extension PT; if the training is successful, deploy this weight file to the periodic detection module, call the PT weight file, and detect the predetermined category in the image.
4. The method for calibrating the in-field space based on the depth of field according to claim 3, wherein The periodic detection module in step S01 includes an improved YOLOv8 neural network. The improved YOLOv8 neural network includes a backbone network, a neck network, and a head network; the backbone network extracts multi-scale features through an improved CSPDarknet structure to generate a high-level feature map with high semantic information; the neck network uses a PAN-FPN structure to fuse feature maps of different levels to enhance the multi-scale detection ability; the head network includes decoupled detection heads to complete target classification, bounding box regression, and confidence prediction respectively; The backbone network includes a C2f module and a spatial pyramid pooling fast module. The C2f module optimizes the feature extraction ability through cross convolution and dual filter design. The C2f module contains multiple Bottleneck blocks, and each block consists of a 1×1 convolution for dimensionality reduction, a 3×3 depthwise separable convolution, and a residual connection; the spatial pyramid pooling fast module realizes multi-scale feature fusion by cascading multiple max pooling layers. It retains spatial information of different granularities through parallel pooling and splicing operations, and uses depthwise separable convolution in some convolutional layers to separate the spatial filtering and channel combination processes; The neck network enhances the model's detection ability for multi-scale targets through multi-path feature fusion. It upsamples the high-level features through the top-down path of the bidirectional feature pyramid and adds them element by element to the middle and low-level features to enhance the semantic information of the shallow features; the bottom-up path downsamples the low-level features and fuses them with the middle and high-level features; In the process of attention mechanism fusion, channel attention or spatial attention is introduced to dynamically adjust the weights of different channels or spatial positions, highlight the key feature areas, and output fusion feature maps of three scales; The head network decoupling head structure predicts the class probability distribution of each anchor point through a fully connected layer or a 1×1 convolution, uses Focal Loss to alleviate the class imbalance problem, predicts the offset of the bounding box through the regression branch, uses CIoU Loss to optimize the position and size accuracy of the box, outputs the confidence score of the existence of the target through the confidence branch, and trains in combination with the binary cross-entropy loss; Enhanced processing is performed on key classes. Based on the known value of the standard size of the object of this class, the distance and azimuth where the object is located are obtained to filter the detection boxes.
5. The in-field space calibration method based on depth of field according to claim 3, wherein The method for calculating the relative distance between the target object and the camera in step S03 includes: Filter the detection box with the largest pixel ratio; Calculate the pixel height difference for the selected detection box: Δypx = y2 - y1; Calculate the sensor scale factor: The relative distance D between the target object and the camera = k * f; Where, y2 is the y pixel of the upper left corner, y1 is the y pixel value of the lower right corner, CMOS_Length is the height of the camera, and f is the focal length of the camera; Taking different objects as target objects respectively, multiple distance values are calculated, and then the weighted average is calculated.
6. The method for calibrating the in-field space based on the depth of field according to claim 1, characterized in that, The method for obtaining the absolute spatial coordinates of the target object in step S04 includes: Using the plane approximation method to obtain the longitude coordinate offset between the target object and the camera position: Obtain the latitude coordinate offset between the target object and the camera position: Where, D is the relative distance between the target object and the camera, θ is the azimuth angle, and latorigin is the original latitude; After adding the longitude and latitude coordinate offsets between the target object and the camera position obtained to the longitude and latitude coordinates of the camera, the absolute spatial coordinates of the target object are obtained.
7. The in-field space calibration method based on depth of field according to claim 1, characterized in that The method for fitting the curved surface equation of the large ground within the camera's field of view in step S04 through the quadratic surface equation includes: S41: Convert the longitude and latitude of the target object into a plane rectangular coordinate system, and map the earth's curved surface to a plane; S42: Separate the geodetic height H and the normal height Hn, and calculate the height anomaly value: ζ = H - Hn S43: Construct a quadratic surface equation: ζ = a0 + a1x + a2y + a3x2 + a4y2 + a5xy Where: x, y are plane coordinates, and a0, a1,..., a5 are coefficients to be determined; S44: Substitute each known point i into the quadratic surface equation to construct an error equation: V i = a0 + a1x i + a2y i + a3x i2 + a4y i2 + a5x i y i - ζ i S45: The design matrix A is a 6×6 matrix, and each row corresponds to the polynomial expansion of a point [1, x i , y i , x i2 , y i2 , x i y i ; Parameter vector X = [a0, a1, a2, a3, a4, a5] T ; Observation vector L = [ζ1, ζ2,..., ζ6] T ; S46: Solve the parameters through least squares solution: X = (A T A) -1 A T L If matrix (A T A) is invertible, calculate directly; if it is a singular matrix, use numerical optimization methods.
8. The method for calibrating the in-field space based on the depth of field according to claim 7, wherein The method for obtaining the absolute spatial coordinates of any point according to the curved surface equation of the large ground within the camera's field of view in step S04 includes: Obtain the distance D and azimuth angle α between the target object and the camera within the field of view; Calculate the distance component D_north in the north-south direction = D * cos(α); Calculate the distance component D_east in the east-west direction = D * sin(α); Calculate the latitude change Calculate the longitude change Δλ = D_east / (R * cosφ_c); Target latitude The target longitude λ_base = λ_c + Δλ; Where, the latitude of the camera is φ_c, the longitude is λ_c, and R is the radius of the earth; When the longitude and latitude values of any point are obtained, substitute them into the quadratic surface equation to obtain the elevation value of the object.
9. An in-field space calibration system based on depth of field, characterized in that, Including: A target detection module that acquires image data and detects target objects; A pixel calculation module that calculates the corresponding horizontal and vertical ratio values of an image according to the pixel area value of a target object in the image; A distance calculation module that calculates the relative distance and azimuth angle between a target object in the image and a camera based on the principle of monocular vision; A fitting calculation module that obtains the absolute spatial coordinates of a target object according to the image pose information and the camera position coordinates, fits the surface equation of the large ground within the camera's field of view through a quadratic surface equation, and obtains the absolute spatial coordinates of any point according to the surface equation of the large ground within the camera's field of view.
10. A computer storage medium, on which a computer program is stored, characterized in that, When the computer program is executed, it implements the in-field space calibration method based on the depth of field according to any one of claims 1-8.