A method, system, and equipment for reviewing evidentiary photographs in land surveys.
By calculating the coverage of evidence photos using a 3D reconstruction model, the problems of sensor data errors and manual interpretation were solved, enabling automated, accurate, and objective review of evidence photos in land surveys and ensuring the reliability of survey results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 广东省土地调查规划院
- Filing Date
- 2026-05-12
- Publication Date
- 2026-07-10
AI Technical Summary
Existing technologies are insufficient in terms of the accuracy and consistency of evidence photo review in land surveys. This is mainly due to sensor data errors and subjective interpretation defects caused by reliance on human experience, which affect the objectivity and credibility of the survey results.
The virtual camera position, intrinsic parameter matrix, and extrinsic parameter rotation matrix of the evidence photos are obtained by reconstructing the 3D model. The model is then converted to the projection coordinate system to calculate the ground projection polygon. Based on this, the local and global coverage rates are calculated to achieve automated and quantitative coverage verification.
It improved the accuracy and consistency of the review of evidence photos, eliminated human error, and realized the full-process automation and standardization of data acquisition and result determination, thereby enhancing the data reliability of land surveys.
Smart Images

Figure CN122368638A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of three-dimensional reconstruction and image processing, and in particular to a method, system and device for reviewing evidence photos in land surveys. Background Technology
[0002] In statutory surveys such as land use change surveys, the "changed land parcels" reported by local authorities and their corresponding "field evidence photographs" are core evidence proving the legality and authenticity of changes in land use status. These photographs must clearly and completely reflect the actual condition of the features within the land parcels to ensure the objectivity and accuracy of the survey results. Therefore, reviewing the evidence photographs is essentially verifying the validity of the evidence, a crucial technical step in preventing false or misreporting, ensuring the authenticity and reliability of land survey data, and maintaining land management order.
[0003] Current technologies primarily rely on metadata inherent in the evidence photographs and vector data of changed land parcels, performing manual overlay comparisons and visual interpretations within a software platform to assess the photographs' coverage of the land parcels. This method suffers from significant accuracy flaws: First, the core sensor data it relies on (such as GPS and IMU) is susceptible to interference from factors like geomagnetic storms, signal blockage, and equipment delays, leading to errors in the photographic location and angle information itself, inevitably resulting in inaccurate coverage assessments. Second, the entire interpretation process heavily depends on the experience and subjective judgment of the reviewers, lacking a unified and quantifiable standard for coverage calculation. This not only results in low efficiency but also makes it difficult to guarantee the consistency, objectivity, and traceability of the assessment results, ultimately affecting the overall quality and credibility of land survey results. Summary of the Invention
[0004] This invention provides a method, system, and equipment for reviewing evidence photos in land surveys, which can improve the accuracy of reviewing evidence photos in land surveys.
[0005] This invention provides a method for reviewing evidentiary photographs in land surveys, including: Obtain several evidentiary photographs corresponding to the changed feature and the vector polygon data of the changed feature; Each of the aforementioned evidence photos is input into a pre-trained 3D reconstruction model to reconstruct the virtual camera position, intrinsic parameter matrix, and extrinsic parameter rotation matrix of each of the aforementioned evidence photos in the virtual world coordinate system constructed by the 3D reconstruction model. The positions of each virtual camera are transformed from the virtual world coordinate system to the projection coordinate system where the vector polygon data is located, and the target camera positions in the projection coordinate system are obtained. Based on each intrinsic parameter matrix, each extrinsic parameter rotation matrix and the target camera positions, the ground projection polygons of each evidence photo in the projection coordinate system are calculated. Based on each ground projection polygon and the vector polygon data, the local coverage rate and the global coverage rate are calculated. The evidence photos are reviewed based on the local coverage rate and the global coverage rate.
[0006] This invention provides a complete and standardized input data foundation for subsequent automated processing by acquiring several evidence photos corresponding to the changed map features and the vector polygon data of the changed map features, thus ensuring the consistency between the audit object and the data source. Based on a 3D reconstruction model, the virtual camera position, intrinsic parameter matrix, and extrinsic parameter rotation matrix are directly reconstructed from the evidence photographs. This overcomes the reliance of existing technologies on potentially erroneous sensor metadata contained in the photographs themselves, ensuring that the core parameters used for calculation are more accurate and reliable from the source. By transforming the virtual camera position to the actual projection coordinate system and calculating the ground projection polygon, a precise geometric mapping relationship is established between the virtual reconstruction scene and the actual geographic space, laying an accurate spatial foundation for the quantitative calculation of coverage. Based on the projection polygon and vector patch data, local and global coverage rates are calculated, providing objective and unified quantitative coverage indicators. This replaces subjective qualitative assessments that rely on human experience and visual interpretation, significantly improving the objectivity and consistency of coverage judgment. Review and judgment are conducted based on the aforementioned quantitative coverage rates, achieving fully automated and standardized processing from data acquisition, spatial mapping, indicator calculation to result judgment. This systematically eliminates the introduction of human error, thereby effectively improving the overall accuracy of evidence photograph review in land surveys.
[0007] Further, the step of calculating the ground projection polygon of each of the evidence photographs in the projection coordinate system based on each of the intrinsic parameter matrices, each of the extrinsic parameter rotation matrices, and the target camera position includes: Based on the aforementioned intrinsic parameter matrices, determine the direction vectors of the four image plane corner points of the evidence photograph in the camera coordinate system. Using the extrinsic rotation matrix, each of the direction vectors is transformed to the projection coordinate system to obtain the ray direction vector in the projection coordinate system; Based on the target camera position and the ray direction vector, an intersection equation with a preset ground plane is established, and the intersection equation is solved to obtain the projected coordinates of each corner point on the ground plane. Based on the aforementioned projection coordinates, the ground projection polygon is constructed.
[0008] This approach replaces the error-prone raw sensor data with precise intrinsic and extrinsic parameters output from a 3D reconstruction model. By utilizing geometric projection and angle calculation algorithms, the visual coverage area of a photograph is transformed into a quantifiable and calculable ground polygon. This achieves a shift from subjective, experience-based visual judgment to precise area calculation based on objective mathematical models, fundamentally eliminating misjudgments caused by human error and inaccurate data.
[0009] Further, determining the direction vectors of the four image plane corner points of the evidence photograph in the camera coordinate system based on each of the intrinsic parameter matrices includes: Obtain the pixel coordinates of the four image plane corner points of the evidence photo, and convert the pixel coordinates into homogeneous coordinates; The homogeneous coordinates are transformed using the inverse of the intrinsic parameter matrix to obtain the initial vector in the camera coordinate system. The initial vector is normalized to obtain the normalized direction vector.
[0010] Further, the calculation of local coverage and global coverage based on the ground projection polygons and the vector polygon data includes: Calculate the area of the first intersection between each of the ground projection polygons and the vector polygon data; The ratio of each first intersection area to the total area of the vector polygon data is calculated to obtain the local coverage rate corresponding to each of the evidence photos. Calculate the area of the second intersection of all the ground projection polygons and the vector polygon data, and calculate the area of the union of the second intersection areas; The global coverage rate is obtained by calculating the ratio of the union area to the total area of the vector polygon data.
[0011] By employing rigorous geometric calculations, subjective visual assessments of coverage are transformed into objective area ratio data, quantifying the actual coverage of map patches by individual and overall photographs. This establishes a unified and precise judgment standard, fundamentally eliminating the randomness of human experience-based judgments and the direct reliance on metadata errors, thereby significantly improving the accuracy and objectivity of the review results.
[0012] Further, the step of transforming the positions of each virtual camera from the virtual world coordinate system to the projection coordinate system where the vector polygon data is located, to obtain the target camera position in the projection coordinate system, includes: Obtain the shooting geographical location of each of the aforementioned evidence photos, and convert the shooting geographical location to a projection coordinate system to obtain the photo projection coordinates corresponding to each of the aforementioned evidence photos; Rigid body registration is performed based on the positions of each virtual camera and the projection coordinates of the photos to calculate rigid body transformation parameters that characterize the transformation relationship between the virtual world coordinate system and the projection coordinate system. Based on the rigid body transformation parameters, the positions of each virtual camera are transformed to the projection coordinate system to obtain the position of the target camera.
[0013] By using rigid body registration, the precise virtual camera position obtained from 3D reconstruction is registered and corrected with the geographical location of the original photo, which may contain errors. This accurately establishes the transformation relationship between the virtual world coordinate system and the real projected coordinate system, providing a unified and accurate coordinate benchmark for subsequent precise calculation of the actual visible range of each photo on the ground. This is a key step in fundamentally improving the accuracy of coverage calculation.
[0014] Further, the step of transforming the positions of each virtual camera to the projection coordinate system based on the rigid body transformation parameters to obtain the target camera position includes: Calculate the first intermediate result of the position of each virtual camera and the corresponding translation vector, wherein the rigid body transformation parameters include at least the rotation matrix, the translation vector and the scaling scale; Multiply the first intermediate result by the inverse of the rotation matrix to obtain the second intermediate result; The target camera position in the projection coordinate system is obtained by multiplying the second intermediate result by the reciprocal of the scaling scale.
[0015] By using rigid body transformation, the virtual coordinates of the 3D reconstruction are accurately mapped to the real geographic coordinate system, effectively correcting the errors in the original photo metadata (such as GPS). This ensures that subsequent view frustum projection and coverage calculation are based on an accurate and unified spatial reference, fundamentally improving the geometric accuracy of coverage determination.
[0016] Furthermore, the training process of the 3D reconstruction model includes: Collect initial evidence photo samples within the target area, wherein the initial evidence photo samples include a first evidence photo taken by a drone and a second evidence photo taken by a mobile phone; Using 3D reconstruction tools, target evidence photo samples that meet the requirements of multi-image shared view and complete point cloud coverage of the core area of the image patch are selected from the initial evidence photo samples; Based on the target evidence photo samples, generate the labeled data and spatial correlation parameters required for training the 3D reconstruction model; The three-dimensional reconstruction model is trained using the first and second evidence photos, wherein the three-dimensional reconstruction model includes a first reconstruction model for drone evidence photos and a second reconstruction model for mobile phone evidence photos.
[0017] By selecting high-quality samples (multiple images with complete coverage) and distinguishing between drone and mobile phone shooting characteristics, a targeted 3D reconstruction model is constructed. This ensures that the model can accurately reproduce camera parameters and scene structure from real and complex shooting data, thereby providing an accurate basis for subsequent coverage calculations and ultimately improving the objectivity and accuracy of the review results.
[0018] Furthermore, the review of each of the evidentiary photos based on the local coverage rate and the global coverage rate includes: The local coverage rate is compared with a first preset threshold to obtain a first comparison result; The global coverage rate is compared with a second preset threshold to obtain a second comparison result; Based on the first comparison result and the second comparison result, a corresponding audit judgment result is generated.
[0019] By conducting the review and judgment based on the aforementioned quantitative coverage rate, the entire process from data acquisition, spatial mapping, indicator calculation to result judgment is automated and standardized. This systematically eliminates the introduction of human error, thereby effectively improving the overall accuracy of the review of evidence photos in land surveys.
[0020] Another embodiment of the present invention provides a system for reviewing evidence photos in land surveys, comprising: The first module is used to acquire several evidence photos corresponding to the changed patch and the vector polygon data of the changed patch; The second module is used to input each of the evidence photos into a pre-trained 3D reconstruction model to reconstruct the virtual camera position, intrinsic parameter matrix, and extrinsic parameter rotation matrix of each of the evidence photos in the virtual world coordinate system constructed by the 3D reconstruction model. The third module is used to transform the positions of each virtual camera from the virtual world coordinate system to the projection coordinate system where the vector polygon data is located, to obtain the target camera position in the projection coordinate system, and to calculate the ground projection polygon of each evidence photo in the projection coordinate system based on each intrinsic parameter matrix, each extrinsic parameter rotation matrix and the target camera position, and to calculate the local coverage rate and global coverage rate based on each ground projection polygon and the vector polygon data. The fourth module is used to review each of the evidence photos based on the local coverage rate and the global coverage rate.
[0021] Another embodiment of the present invention provides a terminal device, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the steps of the land survey evidence photo review method as described in the present invention. Attached Figure Description
[0022] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating one embodiment of the land survey evidence photo review method provided in this application; Figure 2 This is a flowchart illustrating one embodiment of steps S201 to S203 provided in this application; Figure 3 This is a flowchart illustrating one embodiment of steps S301 to S304 provided in this application; Figure 4 This is a schematic diagram of one embodiment of the land survey evidence photo review system provided in this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0026] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0027] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0028] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0029] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).
[0030] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.
[0031] See Figure 1 To improve the accuracy of evidence photo review in land surveys, an embodiment of the present invention provides a method for reviewing evidence photos in land surveys, including steps S101 to S104: Step S101: Obtain several evidence photos corresponding to the changed patch and the vector polygon data of the changed patch; In some embodiments, during land use change surveys, local departments are required to report changed land parcels and supporting documents such as field photographs to demonstrate the rationality of the changes. These field photographs are taken by field personnel of the actual scenes corresponding to the changed land parcels, covering various land use type changes such as conversion of cultivated land to construction land and increase or decrease of forest land. During data reporting, the system receives multiple supporting photographs corresponding to the changed land parcels, along with the vector polygon data of the changed land parcels. This vector data, like the GPS coordinates recorded when the photographs were taken, is typically located in a geographic coordinate system, providing basic spatial information for subsequent processing.
[0032] It should be noted that the shooting methods include drone shooting and mobile phone shooting: drones are suitable for large-scale, panoramic coverage of the map area; mobile phones are used to supplement the shooting of key features of the map area at close range, such as the boundaries of the plots and the identification of land features. These photos usually record information such as geographical location and attitude angle when they are taken, and are reported along with the changed map information.
[0033] It should be noted that the vector polygon data of the changed patch is a vector area polygon existing in the form of a geographic coordinate system, representing the actual range and boundary of the patch.
[0034] Step S102: Input each of the evidence photos into a pre-trained 3D reconstruction model to reconstruct the virtual camera position, intrinsic parameter matrix, and extrinsic parameter rotation matrix of each of the evidence photos in the virtual world coordinate system constructed by the 3D reconstruction model. In some embodiments, the training process of the 3D reconstruction model includes: collecting initial evidence photo samples within the target area, wherein the initial evidence photo samples include a first evidence photo taken by a drone and a second evidence photo taken by a mobile phone; using a 3D reconstruction tool to select target evidence photo samples from the initial evidence photo samples that meet the requirements of multi-image shared view and complete point cloud coverage of the core area of the map patch; generating labeled data and spatial correlation parameters required for training the 3D reconstruction model based on the target evidence photo samples; and training the 3D reconstruction model using the first evidence photo and the second evidence photo, wherein the 3D reconstruction model includes a first reconstruction model for drone evidence photos and a second reconstruction model for mobile phone evidence photos. Specifically, firstly, initial evidence photo samples within the target area need to be collected for the land change survey business scenario, wherein these samples come from field evidence photos reported by local authorities, including a first evidence photo taken by a drone suitable for large-scale, panoramic coverage, and a second evidence photo taken by a mobile phone for close-up supplementary capture of key features of the map patch. Subsequently, due to the significant differences in perspective between drone and mobile phone photography, and the strong differences in the characteristics of changed features in different regions, 3D reconstruction tools such as Colmap are needed to automatically screen the initial samples to ensure the accuracy of 3D reconstruction. From these, target evidence photos that simultaneously meet the conditions of "multi-image co-view" (ensuring spatial consistency between photos) and "complete point cloud coverage of the core area of the map patch" (ensuring the restoration of the complete 3D structure of the map patch features) can be selected. Then, based on these selected target samples, the labeled data and core information such as spatial correlation parameters required for training the 3D reconstruction model can be further generated. Finally, using the first and second evidence photos, and considering the differences in their shooting perspectives, a first reconstruction model for drone evidence photos and a second reconstruction model for mobile phone evidence photos are trained respectively, together forming a 3D reconstruction model adapted to the business scenario, such as the Dust3r model.
[0035] In some embodiments, after the first reconstruction model for drone photos and the second reconstruction model for mobile phone photos are trained, each of the evidence photos is input into the corresponding trained 3D reconstruction model (drone photos are input into the first reconstruction model, and mobile phone photos are input into the second reconstruction model). Then, the corresponding model directly performs 3D reconstruction of the evidence photos in an end-to-end manner to obtain the virtual camera position, intrinsic parameter matrix, and extrinsic parameter rotation matrix of each photo in its own constructed virtual world coordinate system.
[0036] It should be noted that the virtual camera position corresponds to the virtual shooting position of the camera set when reconstructing the point cloud. The intrinsic parameter matrix contains internal optical parameters such as focal length, image width and height, while the extrinsic parameter rotation matrix is used to describe the spatial orientation of the camera in the virtual world coordinate system.
[0037] It should be noted that the point cloud predicted by the Dust3r model is based on the virtual world coordinate system. The virtual world camera coordinates are the core reference benchmark for reconstructing the point cloud. The two are related data under the same coordinate system. The reconstructed point cloud is the result of the model's restoration of the 3D structure of the scene, while the virtual camera position is the virtual shooting position of the camera set during model reconstruction, which can be directly extracted from the generation parameters of the reconstructed point cloud.
[0038] By selecting high-quality samples (multiple images with complete coverage) and distinguishing between drone and mobile phone shooting characteristics, a targeted 3D reconstruction model is constructed. This ensures that the model can accurately reproduce camera parameters and scene structure from real and complex shooting data, thereby providing an accurate basis for subsequent coverage calculations and ultimately improving the objectivity and accuracy of the review results.
[0039] Step S103: Transform the positions of each virtual camera from the virtual world coordinate system to the projection coordinate system where the vector polygon data is located to obtain the target camera position in the projection coordinate system. Based on each intrinsic parameter matrix, each extrinsic parameter rotation matrix and the target camera position, calculate the ground projection polygon of each evidence photo in the projection coordinate system. Based on each ground projection polygon and the vector polygon data, calculate the local coverage rate and the global coverage rate. Please refer to Figure 2 In some embodiments, the step of transforming the positions of each virtual camera from the virtual world coordinate system to the projection coordinate system where the vector polygon data is located, to obtain the target camera position in the projection coordinate system, includes steps S201 to S203: Step S201: Obtain the shooting geographical location of each of the evidence photos, and convert the shooting geographical location to the projection coordinate system to obtain the photo projection coordinates corresponding to each of the evidence photos; In some embodiments, based on the GPS positioning coordinates attached to each evidence photo record, i.e. the shooting location, the precise location of each photo in the projected coordinate system is calculated using a standard conversion algorithm from geographic coordinates to projected coordinates (such as UTM, Gauss-Kruger projection, etc.), i.e., the photo projection coordinates. In this way, the shooting location and vector polygon data can be converted to the same projected coordinate system to establish a unified planar coordinate reference.
[0040] Step S202: Perform rigid body registration based on the positions of each virtual camera and the photo projection coordinates to calculate rigid body transformation parameters that characterize the transformation relationship between the virtual world coordinate system and the projection coordinate system. In some embodiments, to align the virtual reconstructed space with the actual surveyed space, the photo projection coordinates need to be matched with the virtual camera position output by the 3D reconstruction model in the virtual world coordinate system. A set of spatial transformation parameters, i.e., rigid body transformation parameters, is calculated using a rigid body registration algorithm (such as using the least squares method to solve for the optimal spatial transformation). These parameters optimally align the two sets of points (i.e., the photo projection coordinates and the virtual camera position).
[0041] It should be noted that rigid body transformation parameters typically include a rotation matrix R, a translation vector T, and a scaling factor S, which together define the mapping relationship from the real projected coordinate system to the virtual world coordinate system. Its mathematical expression is the rigid body transformation matrix (R, T, S).
[0042] It should be noted that the core of coordinate system transformation is to realize the conversion and mapping between image coordinates, vector coordinates, and the virtual world coordinate system of the reconstructed point cloud. Step S203: Based on the rigid body transformation parameters, transform the positions of each virtual camera to the projection coordinate system to obtain the position of the target camera.
[0043] In some embodiments, step S203 includes: calculating a first intermediate result of each virtual camera position and its corresponding translation vector, wherein the rigid body transformation parameters include at least a rotation matrix, the translation vector, and a scaling factor; multiplying the first intermediate result by the inverse of the rotation matrix to obtain a second intermediate result; and calculating the product of the second intermediate result and the reciprocal of the scaling factor to obtain the target camera position in the projection coordinate system. Specifically, firstly, the position of a virtual camera in the virtual world coordinate system is calculated. The first intermediate result of the translation vector, that is ( ); then, the first intermediate result ( Multiply by the inverse of the rotation matrix Thus, the second intermediate result can be obtained. Finally, calculate the second intermediate result. The reciprocal of the scaling scale The product of the two is calculated using the following formula: Therefore, the position of a virtual camera in the virtual world coordinate system and the position of the target camera in the projected coordinate system can be calculated. .
[0044] By using rigid body transformation, the virtual coordinates of the 3D reconstruction are accurately mapped to the real geographic coordinate system, effectively correcting the errors in the original photo metadata (such as GPS). This ensures that subsequent view frustum projection and coverage calculation are based on an accurate and unified spatial reference, fundamentally improving the geometric accuracy of coverage determination.
[0045] By using rigid body registration, the precise virtual camera position obtained from 3D reconstruction is registered and corrected with the geographical location of the original photo, which may contain errors. This accurately establishes the transformation relationship between the virtual world coordinate system and the real projected coordinate system, providing a unified and accurate coordinate benchmark for subsequent precise calculation of the actual visible range of each photo on the ground. This is a key step in fundamentally improving the accuracy of coverage calculation.
[0046] Please refer to Figure 3 In some embodiments, the step of calculating the ground projection polygon of each of the evidence photographs in the projection coordinate system based on each of the intrinsic parameter matrices, each of the extrinsic parameter rotation matrices, and the target camera position includes steps S301 to S304: Step S301: Determine the direction vectors of the four image plane corner points of the evidence photo in the camera coordinate system according to the intrinsic parameter matrices. In some embodiments, step S301 includes: obtaining the pixel coordinates of the four image plane corner points of the evidence photo, and converting the pixel coordinates into homogeneous coordinates; transforming the homogeneous coordinates using the inverse of the intrinsic parameter matrix to obtain an initial vector in the camera coordinate system; and normalizing the initial vector to obtain the normalized direction vector. Specifically, firstly, the pixel coordinates of the four corner points of the evidence photo are obtained. ,in, These are the coordinates of the four corner points, where 1 represents the homogeneous coordinates; next, using the intrinsic parameter matrix... inverse matrix Transform the four corner points to the initial ray directions in the camera coordinate system to obtain the initial vector. The calculation formula is: Among them, the intrinsic parameter matrix In the formula, and These are the focal lengths output by the model in the x and y directions, respectively. and Let the width and height of the image be given. Finally, for the initial ray vector... Normalization is performed to obtain the unit direction vector. This vector is the direction vector corresponding to the lower corner point of the camera coordinate system.
[0047] Step S302: Using the extrinsic rotation matrix, transform each direction vector to the projection coordinate system to obtain the ray direction vector in the projection coordinate system; In some embodiments, the extrinsic rotation matrix obtained from the 3D reconstruction model is used. (Describe the camera's orientation in the virtual world coordinate system), using the unit direction vector. Transform from the camera coordinate system to the projected coordinate system (i.e., the unified world coordinate system after rigid body registration) to obtain the ray direction vector in the projected coordinate system. The calculation formula is: .
[0048] Through this rotational transformation, the ray direction is unified to the same projection coordinate system as the target camera position, laying the directional reference for subsequent intersection of the ray and the ground plane.
[0049] It should be noted that the camera coordinate system mentioned in this invention is a three-dimensional spatial coordinate system with the imaging center (i.e., the lens optical center) of a single evidence photograph as its origin. Each photograph has its own independent camera coordinate system, which is a local coordinate system determined by the camera's intrinsic parameters (focal length, principal point) and describes the imaging geometry. The virtual world coordinate system is a unified and continuous three-dimensional spatial reference system defined by the 3D reconstruction model (such as Dust3r) during the reconstruction process. It is an "internal" coordinate system created by the model to represent the 3D point cloud (scene structure) jointly reconstructed from all input photographs and the estimated virtual positions of each camera. The projected coordinate system is a standard planar coordinate system used for map drawing in the real world (such as UTM, Gauss-Kruger projection, etc.). It transforms geographical coordinates (latitude and longitude) on the Earth's ellipsoid onto a two-dimensional plane through map projection methods, and has a definite unit (such as meters) and geographical reference.
[0050] Step S303: Based on the target camera position and the ray direction vector, establish the intersection equation with the preset ground plane, and solve the intersection equation to obtain the projected coordinates of each corner point on the ground plane; In some embodiments, based on the coordinate transformation step... Represents the world coordinates of the camera center (i.e.) The specific coordinate values in the projected coordinate system (derived from the coordinate system transformation result) and the ray direction vector Establish the ray and the preset ground plane Establish a ray and a preset ground plane (usually set as the ground plane). Commonly used The equations for the intersection points of the rays are as follows. The parametric equations for the rays are: Where t is the scale factor. Then, its z-component is set to be equal to the ground plane height: The scale parameter can then be obtained. Finally, Substituting back into the ray equation, we can obtain the projected coordinates of the corner point on the ground plane: Continue performing the above calculations on the four corner points to obtain a set of ground projection point coordinates. .
[0051] Step S304: Construct the ground projection polygon based on each of the projection coordinates.
[0052] In some embodiments, the calculated projected coordinates of the four corner points on the ground plane Connect them in sequence to construct the polygon of the visible ground area corresponding to the evidence photo. ,Right now .
[0053] This approach replaces the error-prone raw sensor data with precise intrinsic and extrinsic parameters output from a 3D reconstruction model. By utilizing geometric projection and angle calculation algorithms, the visual coverage area of a photograph is transformed into a quantifiable and calculable ground polygon. This achieves a shift from subjective, experience-based visual judgment to precise area calculation based on objective mathematical models, fundamentally eliminating misjudgments caused by human error and inaccurate data.
[0054] In some embodiments, calculating the local coverage rate and global coverage rate based on each of the ground projection polygons and the vector polygon data includes: calculating the first intersection area of each of the ground projection polygons and the vector polygon data; calculating the ratio of each of the first intersection areas to the total area of the vector polygon data to obtain the local coverage rate corresponding to each of the evidence photos; calculating the second intersection area of all the ground projection polygons and the vector polygon data, and calculating the union area of the second intersection areas; and calculating the ratio of the union area to the total area of the vector polygon data to obtain the global coverage rate. Specifically, each ground projection polygon is calculated... With vector polygons The area of the intersection of the two points is denoted as the first intersection area. Then, the area of each first intersection... Total area of vector polygons Perform a ratio calculation to obtain the local coverage rate corresponding to each evidence photo. Next, calculate the union of the intersection areas of all ground-projected polygons and vector polygons, which is the union area of the second intersection area. Finally, the area of the union is... Total area of vector polygons The ratio was calculated to obtain the global coverage of the patch by all the evidence photos. .
[0055] By employing rigorous geometric calculations, subjective visual assessments of coverage are transformed into objective area ratio data, quantifying the actual coverage of map patches by individual and overall photographs. This establishes a unified and precise judgment standard, fundamentally eliminating the randomness of human experience-based judgments and the direct reliance on metadata errors, thereby significantly improving the accuracy and objectivity of the review results.
[0056] Step S104: Review each of the evidence photos based on the local coverage rate and the global coverage rate.
[0057] In some embodiments, step S104 includes: comparing the local coverage rate with a first preset threshold to obtain a first comparison result; comparing the global coverage rate with a second preset threshold to obtain a second comparison result; and generating a corresponding review judgment result based on the first comparison result and the second comparison result.
[0058] In some embodiments, if the local coverage rate is greater than or equal to a preset first threshold and the global coverage rate is greater than or equal to a preset second threshold, then each of the evidence photos is deemed to have passed the review; if at least one of the local coverage rates is less than the first threshold, then each of the evidence photos is deemed to be substandard, and shooting suggestion information is generated for each of the evidence photos; if the global coverage rate is less than the second threshold, then the plurality of evidence photos are deemed to be substandard as a whole, and a reshoot task list containing supplementary shooting areas and priorities is generated.
[0059] In some embodiments, the calculated local coverage of each photo is compared with a first preset threshold (e.g., 30%) to obtain a first comparison result; at the same time, the global coverage of the entire photo is compared with a second preset threshold (e.g., 70%) to obtain a second comparison result. Based on the above comparison results, the system automatically generates review judgment results and performs corresponding processing: If the local coverage rate of a single photo is <30%, the photo is automatically marked as "coverage substandard," and a feedback mechanism is triggered to inform the reviewers of the specific value and reason for the non-compliance. Simultaneously, based on the 3D structure of the map patch generated by 3D reconstruction, suggestions for areas requiring additional shooting and shooting parameter references are provided. If the global coverage rate of all photos is <70%, the evidence photo for that map patch is judged as "overall unqualified." By overlaying the surface projection polygons of each region to analyze the uncovered blank areas, a "supplementary shooting task list" containing the new shooting locations, quantities, and priorities is generated. Only when the local coverage rate of all single photos is ≥30% and the global coverage rate is ≥70%, the system automatically judges the evidence photo as "qualified," and archives the relevant coverage data, 3D reconstruction report, and changed map patch information, pushing them to subsequent review processes, achieving closed-loop review management without human intervention throughout the entire process. Furthermore, based on the coverage statistics, a field shooting quality assessment model can be constructed. By analyzing the compliance rate of different areas and shooting methods, shooting specifications and operational strategies can be dynamically optimized.
[0060] By conducting the review and judgment based on the aforementioned quantitative coverage rate, the entire process from data acquisition, spatial mapping, indicator calculation to result judgment is automated and standardized. This systematically eliminates the introduction of human error, thereby effectively improving the overall accuracy of the review of evidence photos in land surveys.
[0061] This invention provides a complete and standardized input data foundation for subsequent automated processing by acquiring several evidence photos corresponding to the changed map features and the vector polygon data of the changed map features, thus ensuring the consistency between the audit object and the data source. Based on a 3D reconstruction model, the virtual camera position, intrinsic parameter matrix, and extrinsic parameter rotation matrix are directly reconstructed from the evidence photographs. This overcomes the reliance of existing technologies on potentially erroneous sensor metadata contained in the photographs themselves, ensuring that the core parameters used for calculation are more accurate and reliable from the source. By transforming the virtual camera position to the actual projection coordinate system and calculating the ground projection polygon, a precise geometric mapping relationship is established between the virtual reconstruction scene and the actual geographic space, laying an accurate spatial foundation for the quantitative calculation of coverage. Based on the projection polygon and vector patch data, local and global coverage rates are calculated, providing objective and unified quantitative coverage indicators. This replaces subjective qualitative assessments that rely on human experience and visual interpretation, significantly improving the objectivity and consistency of coverage judgment. Review and judgment are conducted based on the aforementioned quantitative coverage rates, achieving fully automated and standardized processing from data acquisition, spatial mapping, indicator calculation to result judgment. This systematically eliminates the introduction of human error, thereby effectively improving the overall accuracy of evidence photograph review in land surveys.
[0062] like Figure 4 As shown, based on the above method embodiments, corresponding apparatus embodiments are provided; One embodiment of the present invention provides a photo verification system for land survey evidence, comprising: The first module 100 is used to acquire several evidence photos corresponding to the changed patch and the vector polygon data of the changed patch; The second module 200 is used to input each of the evidence photos into a pre-trained three-dimensional reconstruction model to reconstruct the virtual camera position, intrinsic parameter matrix and extrinsic parameter rotation matrix of each of the evidence photos in the virtual world coordinate system constructed by the three-dimensional reconstruction model. The third module 300 is used to transform the positions of each virtual camera from the virtual world coordinate system to the projection coordinate system where the vector polygon data is located, to obtain the target camera position in the projection coordinate system, and to calculate the ground projection polygon of each evidence photo in the projection coordinate system based on each intrinsic parameter matrix, each extrinsic parameter rotation matrix and the target camera position, and to calculate the local coverage rate and the global coverage rate based on each ground projection polygon and the vector polygon data. The fourth module 400 is used to review each of the evidence photos based on the local coverage rate and the global coverage rate.
[0063] It is understood that the above-described device embodiments correspond to the method embodiments of the present invention, and can implement the land survey evidence photo review method provided by any of the above-described method embodiments of the present invention.
[0064] It should be noted that the device embodiments described above are merely illustrative, and some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can specifically be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0065] Based on the above-described embodiments of the land survey evidence photo review method, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the land survey evidence photo review method of any embodiment of the present invention.
[0066] For example, in this embodiment, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the terminal device.
[0067] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0068] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.
[0069] Based on the above-described method embodiments, another embodiment of the present invention provides a computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the land survey evidence photo review method described in any of the above-described method embodiments of the present invention.
[0070] The modules / units integrated in the device / terminal equipment, if implemented as software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0071] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for reviewing evidentiary photographs in land surveys, characterized in that, include: Obtain several evidentiary photographs corresponding to the changed feature and the vector polygon data of the changed feature; Each of the aforementioned evidence photos is input into a pre-trained 3D reconstruction model to reconstruct the virtual camera position, intrinsic parameter matrix, and extrinsic parameter rotation matrix of each of the aforementioned evidence photos in the virtual world coordinate system constructed by the 3D reconstruction model. The positions of each virtual camera are transformed from the virtual world coordinate system to the projection coordinate system where the vector polygon data is located, and the target camera positions in the projection coordinate system are obtained. Based on each intrinsic parameter matrix, each extrinsic parameter rotation matrix and the target camera positions, the ground projection polygons of each evidence photo in the projection coordinate system are calculated. Based on each ground projection polygon and the vector polygon data, the local coverage rate and the global coverage rate are calculated. The evidence photos are reviewed based on the local coverage rate and the global coverage rate.
2. The method for verifying evidentiary photographs in land surveys according to claim 1, characterized in that, The calculation of the ground projection polygon of each of the evidence photographs in the projection coordinate system based on each of the intrinsic parameter matrices, each of the extrinsic parameter rotation matrices, and the target camera position includes: Based on the aforementioned intrinsic parameter matrices, determine the direction vectors of the four image plane corner points of the evidence photograph in the camera coordinate system. Using the extrinsic rotation matrix, each of the direction vectors is transformed to the projection coordinate system to obtain the ray direction vector in the projection coordinate system; Based on the target camera position and the ray direction vector, an intersection equation with a preset ground plane is established, and the intersection equation is solved to obtain the projected coordinates of each corner point on the ground plane. Based on the aforementioned projection coordinates, the ground projection polygon is constructed.
3. The method for verifying evidentiary photographs in land surveys according to claim 2, characterized in that, The step of determining the direction vectors of the four image plane corner points of the evidence photograph in the camera coordinate system based on each of the intrinsic parameter matrices includes: Obtain the pixel coordinates of the four image plane corner points of the evidence photo, and convert the pixel coordinates into homogeneous coordinates; The homogeneous coordinates are transformed using the inverse of the intrinsic parameter matrix to obtain the initial vector in the camera coordinate system. The initial vector is normalized to obtain the normalized direction vector.
4. The method for verifying evidentiary photographs in land surveys according to claim 1, characterized in that, The calculation of local coverage and global coverage based on the ground projection polygons and vector polygon data includes: Calculate the area of the first intersection between each of the ground projection polygons and the vector polygon data; The ratio of each first intersection area to the total area of the vector polygon data is calculated to obtain the local coverage rate corresponding to each of the evidence photos. Calculate the area of the second intersection of all the ground projection polygons and the vector polygon data, and calculate the area of the union of the second intersection areas; The global coverage rate is obtained by calculating the ratio of the union area to the total area of the vector polygon data.
5. The method for verifying evidentiary photographs in land surveys according to claim 1, characterized in that, The step of transforming the positions of each virtual camera from the virtual world coordinate system to the projection coordinate system where the vector polygon data is located, to obtain the target camera position in the projection coordinate system, includes: Obtain the shooting geographical location of each of the aforementioned evidence photos, and convert the shooting geographical location to a projection coordinate system to obtain the photo projection coordinates corresponding to each of the aforementioned evidence photos; Rigid body registration is performed based on the positions of each virtual camera and the projection coordinates of the photos to calculate rigid body transformation parameters that characterize the transformation relationship between the virtual world coordinate system and the projection coordinate system. Based on the rigid body transformation parameters, the positions of each virtual camera are transformed to the projection coordinate system to obtain the position of the target camera.
6. The method for verifying evidentiary photographs in land surveys according to claim 5, characterized in that, The step of transforming the positions of each virtual camera to the projection coordinate system based on the rigid body transformation parameters to obtain the target camera position includes: Calculate the first intermediate result of the position of each virtual camera and the corresponding translation vector, wherein the rigid body transformation parameters include at least the rotation matrix, the translation vector and the scaling scale; Multiply the first intermediate result by the inverse of the rotation matrix to obtain the second intermediate result; The target camera position in the projection coordinate system is obtained by multiplying the second intermediate result by the reciprocal of the scaling scale.
7. The method for verifying evidentiary photographs in land surveys according to claim 1, characterized in that, The training process of the 3D reconstruction model includes: Collect initial evidence photo samples within the target area, wherein the initial evidence photo samples include a first evidence photo taken by a drone and a second evidence photo taken by a mobile phone; Using 3D reconstruction tools, target evidence photo samples that meet the requirements of multi-image shared view and complete point cloud coverage of the core area of the image patch are selected from the initial evidence photo samples; Based on the target evidence photo samples, generate the labeled data and spatial correlation parameters required for training the 3D reconstruction model; The three-dimensional reconstruction model is trained using the first and second evidence photos, wherein the three-dimensional reconstruction model includes a first reconstruction model for drone evidence photos and a second reconstruction model for mobile phone evidence photos.
8. The method for verifying evidentiary photographs in land surveys according to any one of claims 1-7, characterized in that, The review of each of the evidence photos based on the local coverage rate and the global coverage rate includes: The local coverage rate is compared with a first preset threshold to obtain a first comparison result; The global coverage rate is compared with a second preset threshold to obtain a second comparison result; Based on the first comparison result and the second comparison result, a corresponding audit judgment result is generated.
9. A system for verifying evidence photos in land surveys, characterized in that, include: The first module is used to acquire several evidence photos corresponding to the changed patch and the vector polygon data of the changed patch; The second module is used to input each of the evidence photos into a pre-trained 3D reconstruction model to reconstruct the virtual camera position, intrinsic parameter matrix, and extrinsic parameter rotation matrix of each of the evidence photos in the virtual world coordinate system constructed by the 3D reconstruction model. The third module is used to transform the positions of each virtual camera from the virtual world coordinate system to the projection coordinate system where the vector polygon data is located, to obtain the target camera position in the projection coordinate system, and to calculate the ground projection polygon of each evidence photo in the projection coordinate system based on each intrinsic parameter matrix, each extrinsic parameter rotation matrix and the target camera position, and to calculate the local coverage rate and global coverage rate based on each ground projection polygon and the vector polygon data. The fourth module is used to review each of the evidence photos based on the local coverage rate and the global coverage rate.
10. A terminal device, characterized in that, include: One or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the land survey evidentiary photograph review method as described in any one of claims 1-8.