A Rapid Method for Creating High-Standard Farmland Dots Based on Tower Top Single-Camera Photography

By using a single-camera photography technique at the top of the tower to directly project and calculate the DOM, the problem of high cost and low efficiency in generating high-resolution DOMs in traditional methods is solved. This achieves low-cost and high-efficiency DOM generation, meeting the needs of the agricultural field for detailed observation of local features.

CN120472035BActive Publication Date: 2025-11-14INST OF AGRI RESOURCES & REGIONAL PLANNING CHINESE ACADEMY OF AGRI SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510970409.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-11-14
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

Traditional DOM generation methods are difficult to use in agriculture to quickly and efficiently generate high-resolution digital orthophotos of local areas, and are also costly. In particular, due to the low resolution of satellite and aerial imagery and the requirements of field control surveys, they cannot meet the needs for detailed observation of local target features.

Method used

By employing tower-top single-camera photography technology, the camera photography model of the target camera is obtained. Combined with the ground feature elevation and the camera photography model, the DOM is directly projected and calculated, avoiding complex multi-image motion recovery structure technology and sparse/dense point cloud generation, thus simplifying the data post-processing process.

Benefits of technology

It enables the low-cost and efficient generation of high-standard farmland DOMs, reduces operating costs and streamlines the process, meets the detailed observation needs of agricultural applications for localized target features, and ensures high image resolution and a certain degree of geometric accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472035B_ABST
    Figure CN120472035B_ABST
Patent Text Reader

Abstract

This application discloses a rapid method for creating high-standard farmland DOMs based on single-camera photography from a tower top, belonging to the field of image processing technology. The method includes: obtaining a camera photography model corresponding to the target camera based on the pose angle of the target image captured by the target camera and the geographical location of the target camera; the target camera is a camera installed on the top of a tower; obtaining the ground area and grid size of the DOM based on the elevation of features within the target area and the camera photography model; and projecting each grid of the DOM onto the target image based on the distorted camera photography model corresponding to the target camera to obtain the DOM corresponding to the target image. The rapid method for creating high-standard farmland DOMs based on single-camera photography from a tower top provided by this application can significantly reduce operating costs and streamline the workflow while ensuring high image resolution and a certain degree of geometric accuracy, enabling low-cost and efficient DOM generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing technology, and in particular relates to a method for rapid production of high-standard farmland DOM based on tower top single-camera photography. Background Technology

[0002] Digital orthophoto maps (DOMs) are a true reflection of the Earth's surface information and possess both map geometry and image texture characteristics. As a result, DOMs are widely used in agriculture, forestry, land resources, and environmental monitoring.

[0003] Traditional DOM generation methods are mainly based on satellite imagery or aerial survey imagery. The specific steps mainly include field control point measurement, aerial triangulation of satellite imagery or aerial survey imagery, digital elevation model (DEM) generation and editing, and digital differential correction to generate DOM.

[0004] In agriculture, there is a need for detailed observation of target features within a localized area, as well as a need for timely observation of certain crops. These needs necessitate the rapid and efficient generation of small-scale digital orthophoto maps (DOMs) for analyzing farmland and other features within that area. However, traditional DOM generation methods utilize satellite and aerial imagery with low resolution and require field control surveys, making them unsuitable for these requirements. Therefore, how to generate farmland digital orthophotos cost-effectively and efficiently has become a pressing technical challenge in this field. Summary of the Invention

[0005] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a rapid method for producing high-standard farmland digital orthophotos (DOMs) based on tower-top single-camera photography, which can generate farmland digital orthophotos at low cost and high efficiency.

[0006] In a first aspect, this application provides a method for rapid production of high-standard farmland DOM based on tower-top single-camera photography, the method comprising:

[0007] Based on the pose angle of the target image captured by the target camera and the geographical location of the target camera, a camera photography model corresponding to the target camera is obtained; the target camera is a camera installed on the top of a tower; the camera photography model is used to indicate the correspondence between the homogeneous coordinates of pixels in the target image and the coordinates of the corresponding ground features in the three-dimensional world coordinate system; the target image is obtained by the target camera capturing high-standard farmland within the target area;

[0008] Based on the elevation of the features within the target area and the camera photography model, the ground area corresponding to the DOM and the grid size of the DOM are obtained; the elevation of the features within the target area is determined based on the height of the target camera above the ground, the maximum height of the features within the target area, and the height of the target camera in the three-dimensional world coordinate system;

[0009] Based on the distorted camera photography model corresponding to the target camera, each grid of the DOM is projected onto the target image to obtain the DOM corresponding to the target image; the grid is obtained by dividing the ground area based on the grid size of the DOM.

[0010] According to the method for rapid production of high-standard farmland DOM based on single-camera photography from a tower top, this application obtains the camera photography model corresponding to the target camera by using the pose angle of the target image captured by the target camera installed on the tower top and the geographical location of the target camera. Based on the elevation of the ground features within the target area and the camera photography model, the ground range and grid size of the DOM are obtained. Based on the distorted camera photography model corresponding to the target camera, each grid of the DOM is projected onto the target image to obtain the DOM corresponding to the target image. The DOM is calculated directly by projecting the camera photography model of the target camera. This method does not require the use of complex and time-consuming multi-image motion reconstruction structure techniques for camera position and pose estimation and iterative adjustment, and also avoids the influence of feature-related factors. This method addresses situations where camera position and pose calculations fail due to insufficient points or mismatches. It directly estimates the ground extent and grid size of the DOM using the target camera's photographic model and projects the DOM grid using 3D world coordinates constructed from elevation data. This eliminates the need for complex and time-consuming methods of constructing sparse or dense point clouds to generate digital elevation models for projection. This simplifies and improves post-processing, enabling the creation of high-resolution DOMs for localized agricultural features at a lower cost, more effectively, and faster. While maintaining high image resolution and a certain level of geometric accuracy, it significantly reduces operational costs and streamlines the workflow. This method can meet the needs of agricultural applications for detailed observation of localized target features by generating DOMs in a low-cost and efficient manner.

[0011] According to one embodiment of this application, obtaining the camera photography model corresponding to the target camera based on the pose angle of the target image captured by the target camera and the geographical location of the target camera includes:

[0012] Based on the stated attitude angle, obtain the rotation matrix;

[0013] Based on the geographical location of the target camera, obtain the coordinates of the target camera in the three-dimensional world coordinate system;

[0014] The camera photography model is obtained based on the memory matrix of the target camera, the rotation matrix, and the coordinates of the target camera in the three-dimensional world coordinate system.

[0015] According to one embodiment of this application, obtaining the ground extent and grid size of the DOM based on the elevation of the features within the target area and the camera photography model includes:

[0016] Based on the camera photography model, the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of the ground features within the target area are calculated to obtain the ground range;

[0017] The grid size of the DOM is obtained based on the camera photography model, the average elevation of the ground features within the target area, and the ground extent.

[0018] According to one embodiment of this application, the step of calculating the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of ground features within the target area based on the camera photography model to obtain the ground range includes:

[0019] Based on the camera photography model, the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of the ground features within the target area are calculated to obtain the coordinates of the corresponding eight object points in the three-dimensional world coordinate system.

[0020] Obtain the first target coordinates corresponding to the minimum value of the first coordinate axis coordinates and the minimum value of the second coordinate axis coordinates of the eight object points in the three-dimensional world coordinate system, and the second target coordinates corresponding to the maximum value of the first coordinate axis coordinates and the maximum value of the second coordinate axis coordinates;

[0021] The area on the ground corresponding to the rectangle whose diagonal lines are the points corresponding to the first target coordinates and the second target coordinates is determined as the ground range.

[0022] According to one embodiment of this application, obtaining the grid size of the DOM based on the camera photography model, the average elevation of features within the target area, and the ground extent includes:

[0023] A target object point is determined based on the first target coordinates and the average elevation, and based on the second target coordinates and the average elevation, respectively.

[0024] Based on the camera photography model and the coordinates of each target object point in the three-dimensional world coordinate system, obtain the homogeneous coordinates corresponding to each target object point;

[0025] The grid size of the DOM is obtained based on the first target coordinates, the second target coordinates, and the homogeneous coordinates corresponding to the two target object points.

[0026] According to one embodiment of this application, obtaining the grid size of the DOM based on the first target coordinates, the second target coordinates, and the homogeneous coordinates corresponding to the two target object points includes:

[0027] Based on the minimum value of the first coordinate axis coordinate, the maximum value of the first coordinate axis coordinate, and the coordinates corresponding to the first coordinate axis in the homogeneous coordinates of the two target object points, the size of the DOM grid in the direction of the first coordinate axis is obtained. Based on the minimum value of the second coordinate axis coordinate, the maximum value of the second coordinate axis coordinate, and the coordinates corresponding to the second coordinate axis in the homogeneous coordinates of the two target object points, the size of the DOM grid in the direction of the second coordinate axis is obtained.

[0028] According to one embodiment of this application, the step of projecting each grid of the DOM onto the target image based on the distorted camera photography model corresponding to the target camera, and obtaining the DOM corresponding to the target image, includes:

[0029] The ground area is divided based on the grid size of the DOM, and each grid of the DOM is obtained;

[0030] Based on the distorted camera photography model, each grid is projected onto the target image to obtain the floating-point pixel coordinates corresponding to the grid;

[0031] Interpolate the floating-point pixel coordinates to obtain the color value of each grid in the DOM, thereby obtaining the DOM corresponding to the target image.

[0032] Secondly, this application provides a device for rapid production of high-standard farmland DOM based on single-camera photography from a tower top, the device comprising:

[0033] A construction module is used to obtain a camera photography model corresponding to the target camera based on the pose angle of the target image captured by the target camera and the geographical location of the target camera; the target camera is a camera installed on the top of a tower; the camera photography model is used to indicate the correspondence between the homogeneous coordinates of pixels in the target image and the coordinates of the corresponding ground features in the three-dimensional world coordinate system; the target image is obtained by the target camera capturing high-standard farmland within the target area;

[0034] The estimation module is used to obtain the ground extent and grid size of the DOM based on the elevation of the ground features within the target area and the camera photography model; the elevation of the ground features within the target area is determined based on the height of the target camera above the ground, the maximum height of the ground features within the target area, and the height of the target camera in the three-dimensional world coordinate system;

[0035] The acquisition module is used to project each grid of the DOM onto the target image based on the distorted camera photography model corresponding to the target camera, and acquire the DOM corresponding to the target image; the grid is obtained by dividing the ground area based on the grid size of the DOM.

[0036] The high-standard farmland DOM (Domain Image) rapid production device based on tower-top single-camera photography of this application obtains the camera photography model corresponding to the target camera by using the pose angle of the target image captured by the target camera installed on the tower and the geographical location of the target camera. Based on the elevation of the ground features within the target area and the camera photography model, the ground range and grid size of the DOM are obtained. Based on the distorted camera photography model corresponding to the target camera, each grid of the DOM is projected onto the target image to obtain the DOM corresponding to the target image. The DOM is calculated by directly projecting the camera photography model of the target camera. This eliminates the need for complex and time-consuming multi-image motion reconstruction structure technology for camera position and pose estimation and iterative adjustment, and also avoids the influence of feature-related factors. This method addresses situations where camera position and pose calculations fail due to insufficient points or mismatches. It directly estimates the ground extent and grid size of the DOM using the target camera's photographic model and projects the DOM grid using 3D world coordinates constructed from elevation data. This eliminates the need for complex and time-consuming methods of constructing sparse or dense point clouds to generate digital elevation models for projection. This simplifies and improves post-processing, enabling the creation of high-resolution DOMs for localized agricultural features at a lower cost, more effectively, and faster. While maintaining high image resolution and a certain level of geometric accuracy, it significantly reduces operational costs and streamlines the workflow. This method can meet the needs of agricultural applications for detailed observation of localized target features by generating DOMs in a low-cost and efficient manner.

[0037] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for rapid production of high-standard farmland DOM based on tower-top single-camera photography as described in the first aspect above.

[0038] Fourthly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for rapid production of high-standard farmland DOM based on tower-top single-camera photography as described in the first aspect above.

[0039] Fifthly, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method for rapid production of high-standard farmland DOM based on tower-top single-camera photography as described in the first aspect.

[0040] Sixthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method for rapid production of high-standard farmland DOM based on tower-top single-camera photography as described in the first aspect above.

[0041] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0042] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0043] Figure 1 This is a flowchart illustrating the method for rapid production of high-standard farmland DOM based on single-camera photography at the top of a tower, as provided in an embodiment of this application.

[0044] Figure 2 This is a schematic diagram of the structure of the high-standard farmland DOM rapid production device based on tower top single-camera photography provided in the embodiments of this application;

[0045] Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0046] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0047] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0048] In related technologies, with the development of Global Navigation Satellite System (GNSS) and the emergence of inertial navigation technology based on Inertial Measurement Unit (IMU), the DOM creation method assisted by airborne Position & Orientation System (POS) based on UAVs (such as low-altitude UAVs) has been widely adopted.

[0049] The aforementioned UAV-based airborne POS-assisted DOM creation method may include: employing at least one of the following algorithms, including but not limited to Scale-invariant feature transform (SIFT), Speeded UpRobust Features (SURF), and Harris corner extraction, to extract feature points from each image captured by the UAV's airborne camera and perform feature point matching between images; employing Structure from Motion (SfM) technology and utilizing the corresponding POS information to estimate and optimize the image's extrinsic parameters (including position and attitude) and construct a sparse 3D point cloud; generating a dense point cloud using a Multi-View Stereo (MVS) method; generating a Digital Elevation Model (DEM) based on the dense point cloud; and projecting the image using the DEM and the image's extrinsic and extrinsic parameters to generate the DOM. This method requires only a small amount of, or even no, field control surveying, and is therefore widely used in DOM generation or creation.

[0050] It should be noted that the intrinsic and extrinsic parameters of an image refer to the intrinsic and extrinsic parameters of the camera that captured the image, respectively. The camera's intrinsic and extrinsic parameters can usually be represented by intrinsic parameter matrices and extrinsic parameter matrices, respectively.

[0051] For agricultural applications such as the creation of high-standard farmland DOMs, which require the rapid and efficient generation of small-scale DOMs, traditional DOM generation methods using satellite and aerial imagery have low resolution and require field control surveys, making them unsuitable. While the aforementioned UAV-based airborne POS-assisted DOM generation method produces higher-resolution images and requires little or no field control surveys, it requires a large number of images, typically hundreds or thousands. The resulting image data volume is substantial, and post-processing steps such as feature point extraction and matching, image extrinsic parameter estimation and optimization, and the construction of sparse and dense point clouds to generate digital elevation models are complex and time-consuming. Furthermore, the limited crop features in a given area can lead to difficulties in feature point extraction and matching, resulting in DOM generation failures. Additionally, UAV operations are significantly affected by weather conditions (such as wind speed and direction), and controlling UAV flight still incurs costs.

[0052] In summary, the DOM generation methods in related technologies are insufficient to meet the need for quickly and efficiently generating a small DOM.

[0053] The following description, in conjunction with the accompanying drawings, details the method, apparatus, electronic device, and readable storage medium for rapid production of high-standard farmland DOM based on tower-top single-camera photography provided in this application, through specific embodiments and application scenarios.

[0054] Among them, the method for rapid production of high-standard farmland DOM based on single-camera photography at the top of the tower can be applied to the terminal, specifically by the hardware or software in the terminal.

[0055] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablets with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that, in some embodiments, the terminal may not be a portable communication device, but rather a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads).

[0056] The following embodiments describe a terminal including a display and a touch-sensitive surface. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, mouse, and joystick.

[0057] The high-standard farmland DOM rapid production method based on tower top single-camera photography provided in this application embodiment can be executed by an electronic device or a functional module or entity in an electronic device that can implement the high-standard farmland DOM rapid production method based on tower top single-camera photography. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras and wearable devices. The following uses an electronic device as the execution subject to illustrate the high-standard farmland DOM rapid production method based on tower top single-camera photography provided in this application embodiment.

[0058] like Figure 1 As shown, the method for rapid production of high-standard farmland DOM based on single-camera photography from the top of the tower includes steps 110, 120 and 130.

[0059] Step 110: Based on the pose angle of the target image captured by the target camera and the geographical location of the target camera, obtain the camera photography model corresponding to the target camera; the target camera is a camera installed on the top of the tower; the camera photography model is used to indicate the correspondence between the homogeneous coordinates of pixels in the target image and the coordinates of the corresponding ground features in the three-dimensional world coordinate system; the target image is obtained by the target camera capturing high-standard farmland within the target area.

[0060] In practical implementation, this application embodiment utilizes a camera installed atop a tower near high-standard farmland to create the DOM (Document Object Model). It should be noted that, typically, one DOM can be created based on a single image captured by this camera.

[0061] This application implements a method for rapid DOM (Domain Image) production based on single-camera photography from the top of a tower. In the field of agricultural applications, it can greatly reduce operating costs and streamline the operation process. While ensuring high image resolution and a certain geometric accuracy, it can simply and efficiently produce high-standard farmland DOMs for agricultural analysis and application.

[0062] High-standard farmland (well-facilitated farmland) refers to arable land that is flat, concentrated, contiguous, well-equipped, has complete farmland facilities, fertile soil, good ecology, strong disaster resistance, and is adapted to modern agricultural production and management methods. It is guaranteed to yield high and stable harvests regardless of drought or flood and is designated as permanent basic farmland.

[0063] In some embodiments, the camera may be part of a data acquisition device. The data acquisition device may be mounted on the top of the aforementioned tower. In addition to the camera, the data acquisition device may also include a pan-tilt unit, etc.

[0064] In some embodiments, the data acquisition device may further include a positioning and measurement device. This positioning and measurement device may obtain its own geographic location based on GNSS, which is then used as the geographic location of the camera. In some embodiments, the data acquisition device may include at least one camera.

[0065] It should be noted that the aforementioned tower may not be limited to tower-type buildings or structures, but may also be other types of buildings or structures with a height greater than the target value.

[0066] It is understood that the target camera can be a camera mounted on the top of the aforementioned tower. In some embodiments, the target camera can be a high-definition camera.

[0067] In some embodiments, prior to step 110, the target camera takes a picture of the high-standard farmland within the target area, acquiring at least one image. Any one of these at least one images can be used as the target image. Furthermore, for each image obtained by the target camera from the high-standard farmland within the target area, the pose angle of the target camera when taking that image can also be obtained.

[0068] In some embodiments, the attitude angle of the target camera when it takes the image can be obtained based on the angle of the gimbal when the target camera takes the image. In some embodiments, the attitude angle may include yaw, pitch, and roll.

[0069] In some embodiments, the geographical location of the target camera can be obtained before step 110. If the geographical location of the top of the tower is known in advance, it can be used as the geographical location of the target camera. Alternatively, if the geographical location of the top of the tower is not known in advance, the geographical location of the aforementioned positioning and measurement device can be obtained using GNSS-based methods and used as the geographical location of the target camera.

[0070] In some embodiments, the aforementioned geographical location may include longitude L, latitude B, and altitude Z. c .

[0071] It should be noted that the target camera is located at the top of the tower, and this embodiment uses only one camera to capture images. The aforementioned technique for acquiring target images can be called tower-based single-camera photography technology. Tower-based single-camera photography technology can provide real-time and sustainable monitoring image information for high-standard farmland. For monitoring high-standard farmland, how to acquire digital orthophotos from a single-camera system is a key step in effectively carrying out subsequent farmland monitoring.

[0072] After obtaining the pose angle of the target image captured by the target camera and the geographical location of the target camera, a camera photography model corresponding to the target camera can be constructed based on the pose angle of the target image captured by the target camera and the geographical location of the target camera.

[0073] It should be noted that the correspondence between the homogeneous coordinates of pixels in the target image and the coordinates of the corresponding ground features in the 3D world coordinate system can be represented by the camera photography model corresponding to the target camera. Based on the pose angle of the target image captured by the target camera and the geographical location of the target camera, the camera photography model corresponding to the target camera can be constructed using any existing method. This application does not limit the specific method used to construct the camera photography model corresponding to the target camera. Similarly, this application does not limit the specific 3D world coordinate system used. Coordinates in the 3D world coordinate system can be referred to as 3D world coordinates.

[0074] It should be noted that the camera photography model corresponding to the target camera in step 110 is a camera photography model without distortion.

[0075] Step 120: Based on the elevation of the ground features within the target area and the camera photography model, obtain the ground area corresponding to the DOM and the grid size of the DOM; the elevation of the ground features within the target area is determined based on the height of the target camera from the ground, the maximum height of the ground features within the target area, and the height of the target camera in the three-dimensional world coordinate system.

[0076] In actual execution, before step 120, the height H of the target camera above the ground, the maximum height d of the ground objects within the target range, and the height of the target camera in the three-dimensional world coordinate system can be obtained.

[0077] Understandably, a reference coordinate system needs to be chosen in the environment to describe the position of the camera and objects; this coordinate system is called the world coordinate system. The world coordinate system described above is a three-dimensional orthogonal coordinate system. The two orthogonal coordinate axes located at the sea level are the X-axis and Y-axis, and the coordinate axis perpendicular to the sea level is the Z-axis, with the Z-axis coordinate representing altitude. Therefore, the coordinates of ground features in the three-dimensional world coordinate system can be represented as vectors. .

[0078] Understandably, the height of the target camera in the three-dimensional world coordinate system is generally expressed as the target camera's altitude Z. c Therefore, the height of the target camera in the three-dimensional world coordinate system can also be denoted as Z. c .

[0079] In some embodiments, the height H of the target camera above the ground, the maximum height d of ground features within the target area, and the height Z of the target camera in the three-dimensional world coordinate system are obtained. c Then, based on the target camera's height H above the ground, the maximum height d of ground features within the target area, and the target camera's height Z in the three-dimensional world coordinate system, c To obtain the elevation of features within the target area.

[0080] In some embodiments, the elevation of features within the target area may include at least one of minimum elevation, maximum elevation, and average elevation.

[0081] It should be noted that this is based on the target camera's height H above the ground, the maximum height d of ground features within the target area, and the target camera's height Z in the three-dimensional world coordinate system. c The elevation of features within the target area can be obtained using any existing method. This application does not limit the specific method used to obtain the elevation of features within the target area.

[0082] In some embodiments, the minimum elevation of ground features within the target area Maximum elevation and average elevation The results can be obtained using the following formulas: , , .

[0083] In some embodiments, after obtaining the elevation of features within the target area, the ground extent and grid size of the DOM corresponding to the target image can be estimated based on the elevation of the features within the target area and the camera photography model corresponding to the target camera, thereby obtaining the ground extent and grid size of the DOM corresponding to the target image. It is understood that to create the DOM corresponding to the target image, the ground extent of the DOM corresponding to the target image needs to be divided into multiple grids of the same size. The grid size of the DOM refers to the size of each of the aforementioned grids.

[0084] In some embodiments, the homogeneous coordinates of ground features within the target area and their coordinates in the three-dimensional world coordinate system can be transformed based on the camera photography model corresponding to the target camera. This transformation is achieved from the target area captured by the target camera under central projection to the ground area corresponding to the orthophoto projection (i.e., the ground area corresponding to the DOM of the target image). This allows the ground area corresponding to the DOM of the target image to be estimated.

[0085] In some embodiments, based on the transformation between the homogeneous coordinates of the ground features within the target range and their coordinates in the three-dimensional world coordinate system, it is also possible to determine the size of the area within the ground range corresponding to the DOM of the target image for each pixel, and the size of this area is the grid size of the DOM.

[0086] In some embodiments, the ground extent and grid size of the DOM corresponding to the target image can be stored in a tfw format file.

[0087] Step 130: Based on the distorted camera photography model corresponding to the target camera, project each grid of the DOM onto the target image to obtain the DOM corresponding to the target image; the grid is obtained by dividing the ground area based on the grid size of the DOM.

[0088] In actual execution, based on the distorted camera photography model corresponding to the target camera, inverse calculations can be performed using methods such as inverse solving. Each grid of the DOM is projected onto the target image to calculate the DOM corresponding to the target image.

[0089] It should be noted that the inverse calculation can be performed using any commonly used existing method. The specific steps for inverse calculation using inverse methods are not limited in the embodiments of this application.

[0090] In some embodiments, the DOM corresponding to the target image can be in a format such as tif. After obtaining the DOM corresponding to the target image, the DOM can be output in the aforementioned tif format.

[0091] The method for rapid production of high-standard farmland DOM based on single-camera photography from a tower, as provided in this application, obtains a camera photography model corresponding to the target camera based on the pose angle of the target image captured by the target camera installed on the tower and the geographical location of the target camera. Based on the elevation of the ground features within the target area and the camera photography model, the ground area and grid size of the DOM are obtained. Based on the distorted camera photography model corresponding to the target camera, each grid of the DOM is projected onto the target image to obtain the DOM corresponding to the target image. The DOM is calculated by directly projecting the camera photography model of the target camera. This method avoids the need for complex and time-consuming multi-image motion reconstruction structure technology for camera position and pose estimation and iterative adjustment, and also avoids... This method addresses situations where camera position and pose calculations fail due to insufficient feature points or mismatches. It directly estimates the ground extent and grid size of the DOM using the target camera's photographic model and projects the DOM grid using 3D world coordinates constructed from elevation data. This eliminates the need for complex and time-consuming methods of constructing sparse or dense point clouds to generate digital elevation models for projection. This approach simplifies and improves post-processing, enabling the creation of high-resolution DOMs for localized agricultural features at a lower cost, more effectively, and faster. While maintaining high image resolution and a certain level of geometric accuracy, it significantly reduces operational costs and streamlines the workflow. This method can meet the needs of agricultural applications for detailed observation of localized target features by generating DOMs.

[0092] In some embodiments of this application, a camera photography model corresponding to the target camera is obtained based on the pose angle of the target image captured by the target camera and the geographical location of the target camera, including: obtaining a rotation matrix based on the pose angle.

[0093] In actual execution, the rotation matrix can be obtained by calculating the pose angle of the target image captured by the target camera.

[0094] In some embodiments, the target image can be captured from three attitude angles (including yaw angle) of the target camera. Pitch angle and roll angle ), calculate the rotation matrix R.

[0095] In some embodiments, according to the yaw angle Pitch angle and roll angle The rotation order (i.e., according to) - - (Given the rotation order), the rotation matrix R can be calculated using the following formula:

[0096] .

[0097] Based on the geographical location of the target camera, obtain the coordinates of the target camera in the three-dimensional world coordinate system.

[0098] In actual execution, the geographical location of the target camera (generally represented by geographical coordinates using three coordinates: longitude, latitude, and altitude) can be transformed into a three-dimensional world coordinate system to obtain the coordinates of the target camera in the three-dimensional world coordinate system.

[0099] In some embodiments, the three-dimensional world coordinate system can be a Universal Transverse Mercator (UTM) projected coordinate system. This is used to determine the geographical location of the target camera. Transform to UTM projected coordinates .in, It is the coordinate of the target camera in the three-dimensional world coordinate system, which can be denoted as: .

[0100] The camera photography model is obtained based on the target camera's memory matrix, rotation matrix, and the target camera's coordinates in the three-dimensional world coordinate system.

[0101] In practice, the intrinsic parameter matrix of the target camera can be obtained through pre-calibration of the camera or provided by the target camera manufacturer. The intrinsic parameters of the target camera may include focal length. and Like the coordinates of the principal point and and distortion parameters , , , and .

[0102] Since the camera photography model corresponding to the target camera in step 110 is undistorted, the intrinsic parameter matrix of the target camera can be expressed as follows: .

[0103] Based on the target camera's memory matrix K, rotation matrix R, and the target camera's coordinates in the three-dimensional world coordinate system. The camera photography model corresponding to the target camera can be constructed as shown below: .

[0104] Where 'a' represents the homogeneous coordinates of a pixel in the target image. A represents the coordinates of the feature corresponding to that pixel in the three-dimensional world coordinate system. .

[0105] The method for rapid creation of high-standard farmland DOM based on single-camera photography from a tower top, as provided in this application, obtains a rotation matrix based on the pose angle of the target image captured by the target camera, obtains the coordinates of the target camera in the three-dimensional world coordinate system based on the geographical location of the target camera, and obtains the camera photography model corresponding to the target camera based on the memory matrix, rotation matrix, and coordinates of the target camera in the three-dimensional world coordinate system. This allows for direct projection calculation of the DOM using the camera photography model of the target camera, eliminating the need for complex and time-consuming multi-image motion recovery structure techniques for camera position and pose estimation and iterative adjustment. It also avoids camera issues caused by insufficient feature points or mismatches. In cases where position and attitude calculations fail, the system directly estimates the ground extent and grid size of the DOM using the target camera's photographic model. It then uses elevation data to construct the DOM grid's 3D world coordinates for projection. This eliminates the need for complex and time-consuming methods of constructing sparse or dense point clouds to generate digital elevation models for projection. This approach simplifies and improves post-processing, enabling the creation of high-resolution DOMs for localized agricultural features at a lower cost, more effectively, and faster. While maintaining high image resolution and a certain level of geometric accuracy, it significantly reduces operational costs and streamlines the workflow. This approach effectively and cost-efficiently meets the needs of agricultural applications for generating DOMs for detailed observation of localized target features.

[0106] In some embodiments of this application, the ground extent and grid size of the DOM are obtained based on the elevation of features within the target area and the camera photography model, including: calculating the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of features within the target area based on the camera photography model to obtain the ground extent.

[0107] In actual execution, the homogeneous coordinates of the four corner points of the target image and the maximum elevation of the ground features within the target area can be calculated using the camera photography model corresponding to the target camera. This yields four coordinates in the three-dimensional world coordinate system. Similarly, the homogeneous coordinates of the four corner points of the target image and the minimum elevation of the ground features within the target area can be calculated to obtain four coordinates in the three-dimensional world coordinate system. Since the four corner points of the target image are its four vertices, the eight (4+4) coordinates in the three-dimensional world coordinate system can determine a range of the ground, which is the ground range corresponding to the DOM of the target image.

[0108] The grid size of the DOM is obtained based on the camera photography model, the average elevation of ground features within the target area, and the ground extent.

[0109] In actual execution, the ground area corresponding to the DOM of the target image and the average elevation of the ground features within the target area can be converted into homogeneous coordinates of pixels in the target image using the camera photography model corresponding to the target camera. Based on the converted homogeneous coordinates and the ground area corresponding to the DOM of the target image, the grid size of the DOM used to divide the ground area corresponding to the DOM of the target image into a grid can be determined.

[0110] The method for rapid production of high-standard farmland DOMs based on single-camera photography at the top of a tower, provided in this application, calculates the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of features within the target area using the camera photography model corresponding to the target camera. This yields the ground area corresponding to the DOM of the target image. Furthermore, based on the camera photography model corresponding to the target camera, the average elevation of features within the target area, and the ground area corresponding to the DOM of the target image, the grid size of the DOM is obtained. This method enables projection of the DOM grid's three-dimensional world coordinates using elevation as a basis, based on the ground area corresponding to the DOM of the target image and the DOM grid size. It eliminates the need for complex and time-consuming methods of constructing sparse or dense point clouds to generate digital elevation models for projection. This simplifies and improves post-processing, enabling the production of high-resolution DOMs of localized agricultural features at a lower cost, more effectively, and faster. While maintaining high image resolution and a certain level of geometric accuracy, it significantly reduces operational costs and streamlines the workflow, meeting the need for detailed observation of localized target features in agricultural applications at a low cost and with high efficiency.

[0111] In some embodiments of this application, based on a camera photography model, the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of ground features within the target range are calculated to obtain the ground range, including: based on a camera photography model, calculating the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of ground features within the target range to obtain the coordinates of the corresponding eight object points in the three-dimensional world coordinate system.

[0112] In actual execution, the homogeneous coordinates of the four corner points of the target image are as follows: , , , .in, , These represent the width and height of the target image, respectively.

[0113] By using the camera imaging model corresponding to the target camera, the homogeneous coordinates of the four corner points of the target image and the maximum elevation of ground features within the target area can be calculated to obtain four coordinates in the three-dimensional world coordinate system. Similarly, the homogeneous coordinates of the four corner points of the target image and the minimum elevation of ground features within the target area can be calculated to obtain four coordinates in the three-dimensional world coordinate system. All eight (4+4) coordinates in the above three-dimensional world coordinate system correspond to object-space points, and all eight (4+4) coordinates in the above three-dimensional world coordinate system are object-space coordinates.

[0114] Obtain the first target coordinates corresponding to the minimum value of the first coordinate axis and the minimum value of the second coordinate axis in the coordinates of the eight object points in the three-dimensional world coordinate system, and the second target coordinates corresponding to the maximum value of the first coordinate axis and the maximum value of the second coordinate axis.

[0115] In actual execution, the minimum and maximum values ​​of the X-axis and Y-axis coordinates of the eight object coordinates can be obtained respectively. The X-axis and Y-axis are the first and second coordinate axes in the three-dimensional world coordinate system, respectively; the minimum value of the X-axis coordinate is denoted as... The minimum value of the Y-axis coordinate is denoted as The maximum value of the X-axis coordinate is denoted as The maximum value of the Y-axis coordinate is denoted as .

[0116] Get the minimum value of the X-axis coordinate. Minimum value of Y-axis coordinate The maximum value of the X-axis coordinate and the maximum value of the Y-axis coordinate After that, it can be determined that ( )and( , The coordinates of the first target and the second target are respectively determined.

[0117] The area on the ground corresponding to the rectangle whose diagonal lines are the points corresponding to the coordinates of the first target and the second target is defined as the ground range.

[0118] In actual execution, the ground coordinates can be based on the first target coordinates ( The corresponding point and the coordinates of the second target () , The corresponding point is the rectangular area along the diagonal, which is determined as the ground area corresponding to the DOM of the target image.

[0119] According to the high-standard farmland DOM rapid production method based on tower top single-camera photography provided in this application embodiment, the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of the ground features within the target range are calculated based on the camera photography model corresponding to the target camera. This obtains the coordinates of the corresponding eight object points in the three-dimensional world coordinate system. The method also obtains the first target coordinates corresponding to the minimum value of the first coordinate axis and the minimum value of the second coordinate axis among the eight object points in the three-dimensional world coordinate system, and the second target coordinates corresponding to the maximum value of the first coordinate axis and the maximum value of the second coordinate axis. The area on the ground corresponding to the rectangle with the points corresponding to the first and second target coordinates as diagonals is then determined as the target area. The system identifies the ground region corresponding to the DOM (Domain of Depth) in an image. It can project the DOM grid into 3D world coordinates using elevation data, based on the ground region and grid size of the DOM. This eliminates the need for complex and time-consuming methods of constructing sparse or dense point clouds to generate digital elevation models (DEMs). This simplifies and improves post-processing, enabling the creation of high-resolution DOMs for localized agricultural features at a lower cost, more effectively, and faster. While maintaining high image resolution and a certain level of geometric accuracy, it significantly reduces operational costs and streamlines the workflow. This system can efficiently and cost-effectively meet the needs of agricultural applications for detailed observation of localized target features using DOMs.

[0120] In some embodiments of this application, the grid size of the DOM is obtained based on the camera photography model, the average elevation of ground features within the target range, and the ground extent, including: determining a target object point based on the first target coordinates and average elevation, and based on the second target coordinates and average elevation, respectively.

[0121] In actual execution, it can be based on the first target coordinates ( and average elevation Determine the coordinates in the three-dimensional world coordinate system as A target object point, and based on the coordinates of the second target ( , and average elevation Determine the coordinates in the three-dimensional world coordinate system as Another target point.

[0122] Based on the camera photography model and the coordinates of each target point in the three-dimensional world coordinate system, obtain the homogeneous coordinates corresponding to each target point.

[0123] In actual execution, the coordinates of the two target object points mentioned above can be used. , Substituting these values ​​into the camera imaging model corresponding to the target camera, the homogeneous coordinates of the pixels corresponding to the two target object points are calculated. , .

[0124] Based on the coordinates of the first target, the coordinates of the second target, and the homogeneous coordinates of the two target object points, obtain the grid size of the DOM.

[0125] In actual implementation, it can be based on the two homogeneous coordinates mentioned above. , and the coordinates of the two target object points , Estimate the grid size of the DOM corresponding to the target image to obtain the grid size of the DOM corresponding to the target image.

[0126] The method for rapid production of high-standard farmland DOMs based on single-camera photography from a tower, as provided in this application, determines a target object point based on the first target coordinates and average elevation, and another target object point based on the second target coordinates and average elevation. Based on the first target coordinates, the second target coordinates, and the homogeneous coordinates corresponding to the two target object points, the grid size of the DOM corresponding to the target image is obtained. This method enables projection of the DOM grid's three-dimensional world coordinates using elevation, based on the ground area corresponding to the DOM and the DOM's grid size. It eliminates the need for complex and time-consuming methods of constructing sparse or dense point clouds to generate digital elevation models for projection. This simplifies and improves post-processing, enabling the production of high-resolution DOMs of localized agricultural features at a lower cost, more effectively, and faster. While maintaining high image resolution and a certain level of geometric accuracy, it significantly reduces operational costs and streamlines the workflow, meeting the need for detailed observation of localized target features in agricultural applications at a low cost and with high efficiency.

[0127] In some embodiments of this application, the DOM grid size is obtained based on the first target coordinates, the second target coordinates, and the homogeneous coordinates corresponding to the two target object points. This includes: obtaining the size of the DOM grid in the first coordinate axis direction based on the minimum value of the first coordinate axis coordinates, the maximum value of the first coordinate axis coordinates, and the coordinates corresponding to the first coordinate axis in the homogeneous coordinates corresponding to the two target object points; and obtaining the size of the DOM grid in the second coordinate axis direction based on the minimum value of the second coordinate axis coordinates, the maximum value of the second coordinate axis coordinates, and the coordinates corresponding to the second coordinate axis in the homogeneous coordinates corresponding to the two target object points.

[0128] In actual execution, the size of the DOM mesh in the first coordinate axis direction can be obtained based on the minimum value of the first coordinate axis coordinate, the maximum value of the first coordinate axis coordinate, and the coordinates corresponding to the first coordinate axis in the homogeneous coordinates of the two target object points.

[0129] In some embodiments, the maximum value of the first coordinate axis can be obtained. Minimum value of coordinates relative to the first coordinate axis The difference is the coordinate corresponding to the first coordinate axis in the homogeneous coordinate system of the two target object points (including the aforementioned). and The difference between the two values ​​is used as the size of the DOM grid corresponding to the target image along the first coordinate axis. Using the X-axis in the 3D world coordinate system as the first coordinate axis, this can be expressed as the formula described above. .in, This represents the size of the DOM grid corresponding to the target image along the first coordinate axis.

[0130] In actual execution, the size of the DOM mesh in the second coordinate axis direction can be obtained based on the minimum value of the second coordinate axis coordinate, the maximum value of the second coordinate axis coordinate, and the coordinates corresponding to the second coordinate axis in the homogeneous coordinates of the two target object points.

[0131] In some embodiments, the maximum value of the second coordinate axis can be obtained. Minimum value of coordinates relative to the first coordinate axis The difference is the coordinate corresponding to the second coordinate axis in the homogeneous coordinate system of the two target object points (including the aforementioned). and The difference between the two values ​​is used as the size of the DOM grid corresponding to the target image along the second coordinate axis. Using the Y-axis in the 3D world coordinate system as the second coordinate axis, this can be expressed as the formula described above. .in, This represents the size of the DOM grid corresponding to the target image along the second coordinate axis.

[0132] The method for rapid production of high-standard farmland DOMs based on single-camera photography from a tower top, as provided in this application, obtains the size of the DOM grid in the first coordinate axis direction based on the minimum and maximum values ​​of the first coordinate axis coordinates and the coordinates corresponding to the first coordinate axis in the homogeneous coordinates of the two target object points. Similarly, it obtains the size of the DOM grid in the second coordinate axis direction based on the minimum and maximum values ​​of the second coordinate axis coordinates and the coordinates corresponding to the second coordinate axis in the homogeneous coordinates of the two target object points. This method enables projection of the DOM grid's three-dimensional world coordinates using elevation data, based on the ground area corresponding to the DOM in the target image and the DOM grid size. It eliminates the need for complex and time-consuming methods of constructing sparse or dense point clouds to generate digital elevation models for projection. This simplifies and improves post-processing, enabling the production of high-resolution DOMs of localized agricultural features at a lower cost, more effectively, and faster. While maintaining high image resolution and a certain level of geometric accuracy, it significantly reduces operational costs and streamlines the workflow, meeting the need for detailed observation of localized target features in agricultural applications at a low cost and with high efficiency.

[0133] In some embodiments of this application, based on the distorted camera photography model corresponding to the target camera, each grid of the DOM is projected onto the target image to obtain the DOM corresponding to the target image, including: dividing the ground area based on the grid size of the DOM and obtaining each grid of the DOM.

[0134] In practice, the ground area corresponding to the DOM of the target image can be divided into multiple regular grids (usually rectangular grids) based on the grid size of the DOM. Each grid in the DOM corresponds to one pixel of the DOM. The number of grids, rows, and columns can be denoted as m and n, respectively.

[0135] Based on a camera photography model with distortion, each grid is projected onto the target image to obtain the floating-point pixel coordinates corresponding to the grid.

[0136] In actual execution, each of the above grids can be projected onto the target image using the distorted camera photography model corresponding to the target camera, and the floating-point pixel coordinates corresponding to that grid can be obtained.

[0137] In some embodiments, the object coordinates (i.e., three-dimensional world coordinates) of the mesh. After translation and rotation, that is, after Calculations can yield the coordinates. .make The distortion-laden camera photography model corresponding to the target camera is specifically represented by the following equation:

[0138] ;

[0139] .

[0140] in, , , , and The aforementioned distortion parameters; u and v represent the coordinates of the first and second coordinate axes in homogeneous coordinates, respectively; and These represent the intermediate results of distortion correction for u and v, respectively.

[0141] Interpolate the floating-point pixel coordinates to obtain the color value of each grid in the DOM, thereby obtaining the DOM corresponding to the target image.

[0142] In actual execution, the floating-point pixel coordinates of each grid in the DOM are interpolated to obtain the interpolated color value, which is the color value of the pixel corresponding to the grid in the DOM of the target image. At this point, the DOM corresponding to the target image is completed.

[0143] In some embodiments, the floating-point pixel coordinates of each grid cell in the DOM are interpolated using any suitable interpolation method, such as bilinear interpolation. This application does not limit the specific interpolation method used.

[0144] The method for rapid production of high-standard farmland DOM based on single-camera photography from a tower top, provided in this application, divides the ground area based on the DOM's grid size, obtains each grid of the DOM, projects each grid onto the target image based on the distorted camera photography model corresponding to the target camera, obtains the floating-point pixel coordinates corresponding to the grid, interpolates the floating-point pixel coordinates, and obtains the color value of each grid of the DOM, thereby obtaining the DOM corresponding to the target image. This method enables direct projection calculation of the DOM using the target camera's camera photography model, eliminating the need for complex and time-consuming multi-image motion recovery structure techniques for camera position and pose estimation and iterative adjustment. It also avoids issues caused by insufficient feature points or mismatches. This eliminates issues such as camera position and attitude calculation failures caused by misalignment, and enables direct estimation of the ground extent and grid size of the DOM using the target camera's photographic model. It also allows for the projection of the DOM grid's 3D world coordinates using elevation data, eliminating the need for complex and time-consuming methods of constructing sparse or dense point clouds to generate digital elevation models. This simplifies and improves post-processing, enabling the creation of high-resolution DOMs for localized agricultural features at a lower cost, more effectively, and faster. While maintaining high image resolution and a certain level of geometric accuracy, it significantly reduces operational costs and streamlines the workflow, meeting the need for detailed observation of localized target features in agricultural applications in a low-cost and efficient manner.

[0145] To facilitate understanding of the above embodiments of this application, an implementation process of a method for rapidly creating a high-standard farmland DOM based on tower-top single-camera photography is described below. In some embodiments, a method for rapidly creating a high-standard farmland DOM based on tower-top single-camera photography may include the following steps.

[0146] The first step is to use a target camera installed on the top of the tower to take pictures of the target and obtain the target camera's attitude angle, geographical location, height of the target camera above the ground, maximum height of ground objects within the target range, and internal parameters of the target camera.

[0147] The second step is to obtain the rotation matrix based on the pose angle of the target image captured by the target camera.

[0148] The third step is to obtain the coordinates of the target camera in the three-dimensional world coordinate system based on the target camera's geographical location.

[0149] Step 4: Obtain the intrinsic parameter matrix of the target camera based on its intrinsic parameters.

[0150] Step 5: Based on the intrinsic parameter matrix, rotation matrix, and coordinates of the target camera in the three-dimensional world coordinate system, obtain the camera photography model corresponding to the target camera.

[0151] Step 6: Obtain the elevation of the target camera's geographical location, the height of the target camera above the ground, and the maximum height of ground features within the target area, and obtain the elevation of ground features within the target area (including maximum elevation, minimum elevation, and average elevation).

[0152] Step 7: Based on the elevation of ground features within the target area and the camera photography model corresponding to the target camera, obtain the ground area and grid size of the DOM corresponding to the target image.

[0153] Step 8: Based on the distorted camera photography model corresponding to the target camera, the digital orthophoto (DOM) corresponding to the target image is calculated using an inverse algorithm.

[0154] The method for rapid production of high-standard farmland DOM based on tower-top single-camera photography provided in this application can be executed by a device for rapid production of high-standard farmland DOM based on tower-top single-camera photography. This application uses the example of a device for rapid production of high-standard farmland DOM based on tower-top single-camera photography executing the method for rapid production of high-standard farmland DOM based on tower-top single-camera photography to illustrate the device provided in this application.

[0155] This application also provides a device for rapid production of high-standard farmland DOM based on single-camera photography from a tower top. For example... Figure 2 As shown, the high-standard farmland DOM rapid production device based on tower top single camera photography includes: a construction module 210, an estimation module 220, and an acquisition module 230.

[0156] Module 210 is used to obtain the camera photography model corresponding to the target camera based on the pose angle of the target image captured by the target camera and the geographical location of the target camera; the target camera is a camera installed on the top of the tower; the camera photography model is used to indicate the correspondence between the homogeneous coordinates of pixels in the target image and the coordinates of the corresponding ground features in the three-dimensional world coordinate system; the target image is obtained by the target camera capturing high-standard farmland within the target area;

[0157] The estimation module 220 is used to obtain the ground extent and grid size of the DOM based on the elevation of the ground features within the target area and the camera photography model; the elevation of the ground features within the target area is determined based on the height of the target camera above the ground, the maximum height of the ground features within the target area, and the height of the target camera in the three-dimensional world coordinate system;

[0158] The acquisition module 230 is used to project each grid of the DOM onto the target image based on the distorted camera photography model corresponding to the target camera, and acquire the DOM corresponding to the target image; the grid is obtained by dividing the ground area based on the grid size of the DOM.

[0159] The high-standard farmland DOM (Domain Image) rapid production device based on single-camera photography from a tower top, as provided in this application embodiment, obtains the camera photography model corresponding to the target camera based on the pose angle of the target image captured by the target camera installed on the tower top and the geographical location of the target camera. Based on the elevation of the ground features within the target area and the camera photography model, it obtains the ground range and grid size of the DOM. Based on the distorted camera photography model corresponding to the target camera, it projects each grid of the DOM onto the target image to obtain the DOM corresponding to the target image. The DOM is calculated directly by projecting the camera photography model of the target camera. This eliminates the need for complex and time-consuming multi-image motion recovery structure technology for camera position and pose estimation and iterative adjustment, and also avoids... This method addresses situations where camera position and pose calculations fail due to insufficient feature points or mismatches. It directly estimates the ground extent and grid size of the DOM using the target camera's photographic model and projects the DOM grid using 3D world coordinates constructed from elevation data. This eliminates the need for complex and time-consuming methods of constructing sparse or dense point clouds to generate digital elevation models for projection. This approach simplifies and improves post-processing, enabling the creation of high-resolution DOMs for localized agricultural features at a lower cost, more effectively, and faster. While maintaining high image resolution and a certain level of geometric accuracy, it significantly reduces operational costs and streamlines the workflow. This method can meet the needs of agricultural applications for detailed observation of localized target features by generating DOMs.

[0160] In some embodiments, the construction module 210 may include:

[0161] The first acquisition unit is used to acquire the rotation matrix based on the attitude angle;

[0162] The second acquisition unit is used to acquire the coordinates of the target camera in the three-dimensional world coordinate system based on the geographical location of the target camera;

[0163] The third acquisition unit is used to acquire the camera photography model based on the target camera's memory matrix, rotation matrix, and the target camera's coordinates in the three-dimensional world coordinate system.

[0164] In some embodiments, the estimation module 220 may include:

[0165] The first estimation unit is used to calculate the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of ground features within the target area based on the camera photography model, so as to obtain the ground range.

[0166] The second estimation unit is used to obtain the grid size of the DOM based on the camera photography model, the average elevation of ground features within the target area, and the ground extent.

[0167] In some embodiments, the first estimation unit may be specifically used for:

[0168] Based on the camera photography model, the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of the ground features within the target range are calculated to obtain the coordinates of the corresponding eight object points in the three-dimensional world coordinate system.

[0169] Obtain the first target coordinates corresponding to the minimum value of the first coordinate axis and the minimum value of the second coordinate axis in the coordinates of the eight object points in the three-dimensional world coordinate system, and the second target coordinates corresponding to the maximum value of the first coordinate axis and the maximum value of the second coordinate axis.

[0170] The area on the ground corresponding to the rectangle whose diagonal lines are the points corresponding to the coordinates of the first target and the second target is defined as the ground range.

[0171] In some embodiments, the second estimation unit may include:

[0172] Determine sub-units to determine a target object point based on the first target coordinates and average elevation, and based on the second target coordinates and average elevation, respectively;

[0173] The first acquisition subunit is used to acquire the homogeneous coordinates of each target object point based on the camera photography model and the coordinates of each target object point in the three-dimensional world coordinate system.

[0174] The second acquisition subunit is used to obtain the grid size of the DOM based on the coordinates of the first target, the coordinates of the second target, and the homogeneous coordinates corresponding to the two target object points.

[0175] In some embodiments, the second acquisition subunit may be specifically used to acquire the size of the DOM grid in the first coordinate axis direction based on the minimum value of the first coordinate axis coordinate, the maximum value of the first coordinate axis coordinate, and the coordinate corresponding to the first coordinate axis in the homogeneous coordinates corresponding to the two target object points, and to acquire the size of the DOM grid in the second coordinate axis direction based on the minimum value of the second coordinate axis coordinate, the maximum value of the second coordinate axis coordinate, and the coordinate corresponding to the second coordinate axis in the homogeneous coordinates corresponding to the two target object points.

[0176] In some embodiments, the acquisition module 230 may be specifically used for:

[0177] The ground area is divided based on the grid size of the DOM, and each grid cell of the DOM is obtained;

[0178] Based on a camera photography model with distortion, each grid is projected onto the target image to obtain the floating-point pixel coordinates corresponding to the grid.

[0179] Interpolate the floating-point pixel coordinates to obtain the color value of each grid in the DOM, thereby obtaining the DOM corresponding to the target image.

[0180] The high-standard farmland DOM rapid production device based on tower-top single-camera photography in this application embodiment can be an electronic device or a component of an electronic device, such as an integrated circuit or chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0181] The high-standard farmland DOM rapid production device based on tower-top single-camera photography in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0182] The high-standard farmland DOM rapid production device based on tower top single-camera photography provided in this application embodiment can achieve… Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0183] In some embodiments, such as Figure 3 As shown, this application embodiment also provides an electronic device 300, including a processor 310, a memory 320, and a computer program stored in the memory 320 and executable on the processor 310. When the computer program is executed by the processor 310, it implements the various processes of the above-described embodiment of the method for rapid production of high-standard farmland DOM based on tower top single-camera photography and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0184] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0185] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described method embodiment for rapid production of high-standard farmland DOM based on tower-top single-camera photography and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0186] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0187] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method for rapid production of high-standard farmland DOM based on tower-top single-camera photography.

[0188] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0189] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described embodiment of the method for rapid production of high-standard farmland DOM based on tower-top single-camera photography, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0190] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0191] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0192] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0193] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0194] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0195] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for rapid production of high-standard farmland DOM based on tower-top single-camera photography, characterized in that, include: Based on the pose angle of the target image captured by the target camera and the geographical location of the target camera, a camera photography model corresponding to the target camera is obtained; the target camera is a camera installed on the top of a tower; the camera photography model is used to indicate the correspondence between the homogeneous coordinates of pixels in the target image and the coordinates of the corresponding ground features in the three-dimensional world coordinate system; The target image is obtained by the target camera from high-standard farmland within the target area; Based on the elevation of the features within the target area and the camera photography model, the ground area corresponding to the DOM and the grid size of the DOM are obtained; the elevation of the features within the target area is determined based on the height of the target camera above the ground, the maximum height of the features within the target area, and the height of the target camera in the three-dimensional world coordinate system; Based on the distorted camera photography model corresponding to the target camera, each grid of the DOM is projected onto the target image to obtain the DOM corresponding to the target image; The grid is obtained by dividing the ground area based on the grid size of the DOM; The process of obtaining the ground extent and grid size of the DOM based on the elevation of features within the target area and the camera photography model includes: Based on the camera photography model, the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of the ground features within the target area are calculated to obtain the ground range; Based on the camera photography model, the average elevation of ground features within the target area, and the ground extent, the grid size of the DOM is obtained; Based on the camera photography model, the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of ground features within the target area are calculated to obtain the ground range, including: Based on the camera photography model, the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of the ground features within the target area are calculated to obtain the coordinates of the corresponding eight object points in the three-dimensional world coordinate system. Obtain the first target coordinates corresponding to the minimum value of the first coordinate axis coordinates and the minimum value of the second coordinate axis coordinates of the eight object points in the three-dimensional world coordinate system, and the second target coordinates corresponding to the maximum value of the first coordinate axis coordinates and the maximum value of the second coordinate axis coordinates; The area on the ground corresponding to the rectangle whose diagonal lines are the points corresponding to the first target coordinates and the second target coordinates is defined as the ground range.

2. The method for rapid production of high-standard farmland DOM based on tower-top single-camera photography according to claim 1, characterized in that, The step of obtaining the camera photography model corresponding to the target camera based on the pose angle of the target image captured by the target camera and the geographical location of the target camera includes: Based on the stated attitude angle, obtain the rotation matrix; Based on the geographical location of the target camera, obtain the coordinates of the target camera in the three-dimensional world coordinate system; The camera photography model is obtained based on the memory matrix of the target camera, the rotation matrix, and the coordinates of the target camera in the three-dimensional world coordinate system.

3. The method for rapid production of high-standard farmland DOM based on tower-top single-camera photography according to claim 1, characterized in that, The step of obtaining the grid size of the DOM based on the camera photography model, the average elevation of ground features within the target area, and the ground extent includes: A target object point is determined based on the first target coordinates and the average elevation, and based on the second target coordinates and the average elevation, respectively. Based on the camera photography model and the coordinates of each target object point in the three-dimensional world coordinate system, obtain the homogeneous coordinates corresponding to each target object point; The grid size of the DOM is obtained based on the first target coordinates, the second target coordinates, and the homogeneous coordinates corresponding to the two target object points.

4. The method for rapid production of high-standard farmland DOM based on tower-top single-camera photography according to claim 3, characterized in that, The step of obtaining the grid size of the DOM based on the first target coordinates, the second target coordinates, and the homogeneous coordinates corresponding to the two target object points includes: Based on the minimum value of the first coordinate axis coordinate, the maximum value of the first coordinate axis coordinate, and the coordinates corresponding to the first coordinate axis in the homogeneous coordinates of the two target object points, the size of the DOM grid in the direction of the first coordinate axis is obtained. Based on the minimum value of the second coordinate axis coordinate, the maximum value of the second coordinate axis coordinate, and the coordinates corresponding to the second coordinate axis in the homogeneous coordinates of the two target object points, the size of the DOM grid in the direction of the second coordinate axis is obtained.

5. The method for rapid production of high-standard farmland DOM based on single-camera photography from a tower top, as described in any one of claims 1 to 4, is characterized in that... The step of projecting each grid of the DOM onto the target image based on the distorted camera photography model corresponding to the target camera, and obtaining the DOM corresponding to the target image, includes: The ground area is divided based on the grid size of the DOM, and each grid of the DOM is obtained; Based on the distorted camera photography model, each grid is projected onto the target image to obtain the floating-point pixel coordinates corresponding to the grid; Interpolate the floating-point pixel coordinates to obtain the color value of each grid in the DOM, thereby obtaining the DOM corresponding to the target image.

6. A device for rapid production of high-standard farmland DOM based on single-camera photography from a tower top, characterized in that, include: A construction module is used to obtain a camera photography model corresponding to the target camera based on the pose angle of the target image captured by the target camera and the geographical location of the target camera; the target camera is a camera installed on the top of a tower; the camera photography model is used to indicate the correspondence between the homogeneous coordinates of pixels in the target image and the coordinates of the corresponding ground features in the three-dimensional world coordinate system; The target image is obtained by the target camera from high-standard farmland within the target area; The estimation module is used to obtain the ground extent and grid size of the DOM based on the elevation of the ground features within the target area and the camera photography model; the elevation of the ground features within the target area is determined based on the height of the target camera above the ground, the maximum height of the ground features within the target area, and the height of the target camera in the three-dimensional world coordinate system; The acquisition module is used to project each grid of the DOM onto the target image based on the distorted camera photography model corresponding to the target camera, and acquire the DOM corresponding to the target image; The grid is obtained by dividing the ground area based on the grid size of the DOM; The estimation module includes: The first estimation unit is used to calculate the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of the ground features within the target area based on the camera photography model, so as to obtain the ground range. The second estimation unit is used to obtain the grid size of the DOM based on the camera photography model, the average elevation of the ground features within the target range, and the ground range; The first estimation unit is specifically used for: Based on the camera photography model, the homogeneous coordinates of the four corner points of the target image and the maximum and minimum elevations of the ground features within the target area are calculated to obtain the coordinates of the corresponding eight object points in the three-dimensional world coordinate system. Obtain the first target coordinates corresponding to the minimum value of the first coordinate axis coordinates and the minimum value of the second coordinate axis coordinates of the eight object points in the three-dimensional world coordinate system, and the second target coordinates corresponding to the maximum value of the first coordinate axis coordinates and the maximum value of the second coordinate axis coordinates; The area on the ground corresponding to the rectangle whose diagonal lines are the points corresponding to the first target coordinates and the second target coordinates is defined as the ground range.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for rapid production of high-standard farmland DOM based on tower-top single-camera photography as described in any one of claims 1-5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for rapid production of high-standard farmland DOM based on single-camera photography from a tower as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Ortho-image real-time generation method and system based on SLAM technology

    CN110675450A

  • Construction land surveying and mapping method and system

    CN117994463A