Three-dimensional reconstruction pricking method, system, medium and product for earthwork site

CN122597751APending Publication Date: 2026-08-18FUJIAN SHUITOU DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610572495.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-28
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0004]然而,随着工地作业范围的不断扩大,单次航拍任务产生的影像数量可达数百甚至上千张,且影像分辨率不断提高导致单张影像数据量巨大

Benefits of technology

通过采用上述技术方案,首先通过对航拍影像进行下采样生成缩略影像并记录缩放比例因子,在低分辨率层面快速完成颜色特征提取和候选掩膜生成,随后通过连通域分析在缩略影像层面筛选出目标候选区域,避免了对整张高分辨率航拍影像进行全局扫描,减少了后续需要精细处理的数据量。在确定目标候选区域后,利用缩放比例因子将其映射回航拍影像并截取目标分辨率切片,通过对目标分辨率切片进行纯净度分析将主体掩膜和编号掩膜分离,结合基于主体掩膜计算的旋转矩阵对切片、主体掩膜和编号掩膜统一进行仿射变换,消除了控制点标靶拍摄角度差异对特征提取的干扰,为后续编号识别和中心定位提供了规范化的输入。在旋正处理后,从旋正编号掩膜提取编号子图像并进行字符识别获得目标编号和置信度,同时基于旋正主体掩膜进行边缘检测与拟合计算旋正中心坐标,再利用旋转矩阵逆变换将旋正中心坐标准确映射回航拍影像坐标系得到绝对像素坐标,确保了控制点像素坐标与编号的精确对应关系。最终将绝对像素坐标、目标编号和置信度组合生成控制点观测记录,为三维重建软件提供了完整的空间约束数据。这种由粗到精的分层处理流程在保证控制点提取精度的前提下,有效应对了单次航拍任务数百至上千张高分辨率影像的处理挑战,提高了土方工地三维重建过程中控制点观测记录的生成效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122597751A_ABST
    Figure CN122597751A_ABST
Patent Text Reader

Abstract

A method, system, medium, and product for 3D reconstruction of control point observation records for earthwork construction sites are disclosed, relating to the field of data processing technology. The method involves acquiring aerial images of the site, downsampling to generate thumbnail images, and extracting color features to obtain candidate masks. Target regions are screened through connected component analysis, and target slices are extracted from the original image using a scaling factor. Purity analysis is performed on the slices to separate the subject and the numbered mask. A rotation matrix is ​​calculated based on the subject mask, and an affine transformation is performed to obtain a rotated slice and its corresponding mask. Subsequently, the rotated numbered mask is used to extract sub-images for character recognition, obtaining the target number and confidence level. Finally, the rotated subject mask is used for edge fitting to calculate the center coordinates, and an inverse transformation is performed to obtain the absolute pixel coordinates, which are then combined with the target number and confidence level to generate control point observation records. Implementing the technical solution provided in this application can improve the efficiency of generating control point observation records during 3D reconstruction of earthwork construction sites.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, specifically to a three-dimensional reconstruction method, system, medium, and product for earthwork construction sites. Background Technology

[0002] With the continuous expansion of construction projects and earthwork operations, 3D topographic surveying and earthwork volume calculation have become crucial aspects of construction management. Unmanned aerial vehicle (UAV) photogrammetry technology, with its advantages of efficiency and flexibility, is widely used in earthwork site surveying. By acquiring image data through UAV aerial photography and combining it with 3D reconstruction technology, a digital elevation model of the construction site can be quickly generated, enabling accurate calculation of earthwork volume and dynamic monitoring of construction progress.

[0003] In the 3D reconstruction process of UAV photogrammetry, several control points (also known as sniffer points or image control points) need to be established within the survey area. These control points typically consist of targets with obvious visual features and numbered markers, and their ground coordinates are accurately determined using surveying equipment such as RTK or total stations. The 3D reconstruction software needs to extract the pixel coordinates of the control points from the aerial imagery and associate them with their corresponding numbers and ground coordinates as spatial constraints to improve the absolute accuracy of the reconstructed model. This control point-constrained reconstruction method can effectively correct systematic errors in the imagery and has been widely used in large-scale topographic surveying.

[0004] However, as the scope of construction site operations continues to expand, the number of images generated in a single aerial photography mission can reach hundreds or even thousands, and the increasing image resolution leads to a massive amount of data per image. In practical applications, traditional control point extraction methods have a high computational load when processing massive amounts of high-resolution image data, making it difficult to meet the requirements of engineering projects for rapid delivery of measurement results, thus reducing the efficiency of control point observation record generation during the 3D reconstruction process. Summary of the Invention

[0005] This application provides a method, system, medium, and product for three-dimensional reconstruction of control points for earthwork construction sites, which can improve the efficiency of control point observation record generation during the three-dimensional reconstruction process for earthwork construction sites.

[0006] The first aspect of this application provides a three-dimensional reconstruction method for puncture points in earthwork construction sites, comprising: Acquire aerial images of the site to be tested; The aerial image is downsampled to generate a thumbnail image, the scaling factor is recorded, and color features are extracted from the thumbnail image to obtain a candidate mask. The candidate mask is subjected to connected component analysis to filter out the target candidate region, and the target candidate region is mapped to the aerial image according to the scaling factor to extract the target resolution slice. Purity analysis is performed on the target resolution slice to separate the main mask and the number mask. A rotation matrix is ​​calculated based on the main mask. The rotation matrix is ​​then used to perform affine transformations on the target resolution slice, the main mask, and the number mask to obtain a rotated slice, a rotated main mask, and a rotated number mask, respectively. Based on the sort number mask, numbered sub-images are extracted from the sort slice, and character recognition is performed on the numbered sub-images to obtain the target number and confidence level; Edge detection and fitting are performed on the rotated slice based on the rotated main mask, the coordinates of the rotated center are calculated, and the coordinates of the rotated center are inversely transformed to the coordinate system of the aerial image using the rotation matrix to obtain the absolute pixel coordinates. The absolute pixel coordinates, the target number and the confidence level are combined to generate control point observation records.

[0007] By adopting the above technical solution, firstly, a thumbnail image is generated by downsampling the aerial image and the scaling factor is recorded. Color feature extraction and candidate mask generation are quickly completed at the low-resolution level. Then, connected component analysis is used to filter target candidate regions at the thumbnail image level, avoiding a global scan of the entire high-resolution aerial image and reducing the amount of data requiring subsequent fine processing. After determining the target candidate region, the scaling factor is used to map it back to the aerial image and extract target resolution slices. Purity analysis of the target resolution slices separates the subject mask and the number mask. A affine transformation is then performed on the slices, subject mask, and number mask using a rotation matrix calculated based on the subject mask. This eliminates the interference of control point target shooting angle differences on feature extraction, providing standardized input for subsequent number recognition and center localization. After the rotation process, the numbered sub-images are extracted from the rotation number mask, and character recognition is performed to obtain the target number and confidence score. Simultaneously, edge detection and fitting are performed based on the rotation main mask to calculate the rotation center coordinates. Then, the rotation center coordinates are accurately mapped back to the aerial image coordinate system using the inverse rotation matrix transformation to obtain the absolute pixel coordinates, ensuring a precise correspondence between the control point pixel coordinates and their numbers. Finally, the absolute pixel coordinates, target numbers, and confidence scores are combined to generate control point observation records, providing complete spatial constraint data for the 3D reconstruction software. This coarse-to-fine layered processing flow effectively addresses the challenge of processing hundreds to thousands of high-resolution images in a single aerial photography mission while ensuring the accuracy of control point extraction, improving the efficiency of control point observation record generation during the 3D reconstruction of earthwork construction sites.

[0008] Optionally, multiple connected regions in the candidate mask are identified, and geometric feature parameters of each connected region are calculated. The geometric feature parameters include pixel area, aspect ratio of the bounding rectangle, and circularity. Based on the geometric feature parameters, a weighted calculation is performed to determine the target score of each connected region. Connected regions with target scores greater than a preset score threshold are selected as target candidate regions, and the boundary coordinates of the target candidate regions in the thumbnail image are obtained. The boundary coordinates are compensated using the scaling factor to obtain mapped boundary coordinates. The cropping range is determined in the aerial image based on the mapped boundary coordinates, and a target resolution slice is extracted based on the cropping range.

[0009] Optionally, calculate the color purity index of each pixel in the target resolution slice; divide pixels with a color purity index greater than or equal to a purity threshold into main regions and generate a main mask; divide pixels with a color purity index less than the purity threshold and a brightness greater than the brightness threshold into numbered regions and generate numbered masks; extract the outer contour of the main mask, calculate the minimum bounding rectangle of the outer contour, and obtain the center point coordinates and tilt angle of the minimum bounding rectangle; construct a rotation matrix with the center point coordinates as the rotation center and the negative value of the tilt angle as the rotation amount; use the rotation matrix to resample pixels in the target resolution slice, the main mask, and the numbered mask respectively to generate the rotated slice, the rotated main mask, and the rotated numbered mask.

[0010] Optionally, the bounding box of the foreground pixels in the swivel numbering mask is determined; a numbered sub-image is determined from the swivel slice based on the bounding box; the numbered sub-image is converted to grayscale and binarized to obtain a binarized image; the binarized image is segmented into characters to obtain a single-character image sequence; the feature vector of each character image in the single-character image sequence is extracted, and the feature vector is matched with a pre-stored standard character template library to determine the matching character and matching similarity corresponding to each character image; the matching characters corresponding to each character image are concatenated in order to generate the target number, and the minimum value is selected from the matching similarities corresponding to each character image as the confidence score.

[0011] Optionally, in the rotation slice, a set of pixels within the area covered by the rotation main mask is extracted; edge pixels of the pixel set are extracted using an edge detection algorithm, and the positions of the edge pixels are interpolated using an interpolation algorithm to obtain a sub-pixel edge point set; geometric contour fitting is performed on the sub-pixel edge point set, and the geometric center point of the fitted contour is extracted as the preliminary center coordinates; a local neighborhood is extracted in the rotation slice with the preliminary center coordinates as the center, and the grayscale weight of each pixel in the local neighborhood is calculated; the pixel coordinates in the local neighborhood are weighted and summed based on the grayscale weights, and the resulting weighted centroid coordinates are used as the rotation center coordinates.

[0012] Optionally, the rotation matrix is ​​inverted to obtain an inverse transformation matrix; the rotation center coordinates are multiplied by the inverse transformation matrix to obtain the local coordinates of the target resolution slice in the coordinate system; the local coordinates are superimposed with the cropping offset of the target resolution slice in the aerial image to obtain the absolute pixel coordinates; the image identifier of the aerial image is extracted; the image identifier, the absolute pixel coordinates, the target number, and the confidence score are encapsulated according to a preset data structure to generate the control point observation record.

[0013] Optionally, the absolute pixel coordinates in the control point observation records are matched with the pre-acquired measured 3D coordinates to generate a control point constraint file; corresponding feature points are extracted from the aerial image, and the corresponding feature points are combined with the control point constraint file to perform error minimization adjustment calculation to restore the camera pose parameters and generate a sparse point cloud; dense matching is performed based on the camera pose parameters and the sparse point cloud to generate a 3D elevation grid; the 3D elevation grid is compared with the pre-acquired benchmark elevation grid to extract the elevation difference at the corresponding coordinate positions, and the earthwork volume of the site to be measured is calculated based on the elevation difference.

[0014] Secondly, embodiments of this application provide a three-dimensional reconstruction puncture system for earthwork construction sites. The three-dimensional reconstruction puncture system for earthwork construction sites includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the three-dimensional reconstruction puncture system for earthwork construction sites to perform the method described in the first aspect and any possible implementation thereof.

[0015] Thirdly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a three-dimensional reconstruction puncture system for an earthwork site, cause the three-dimensional reconstruction puncture system for an earthwork site to perform the method described in the first aspect and any possible implementation thereof.

[0016] Fourthly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on a three-dimensional reconstruction puncture system for earthwork construction sites, causes the three-dimensional reconstruction puncture system for earthwork construction sites to perform the method described in the first aspect and any possible implementation thereof.

[0017] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages: By adopting the above technical solution, firstly, a thumbnail image is generated by downsampling the aerial image and the scaling factor is recorded. Color feature extraction and candidate mask generation are quickly completed at the low-resolution level. Then, connected component analysis is used to filter target candidate regions at the thumbnail image level, avoiding a global scan of the entire high-resolution aerial image and reducing the amount of data requiring subsequent fine processing. After determining the target candidate region, the scaling factor is used to map it back to the aerial image and extract target resolution slices. Purity analysis of the target resolution slices separates the subject mask and the number mask. A affine transformation is then performed on the slices, subject mask, and number mask using a rotation matrix calculated based on the subject mask. This eliminates the interference of control point target shooting angle differences on feature extraction, providing standardized input for subsequent number recognition and center localization. After the rotation process, the numbered sub-images are extracted from the rotation number mask, and character recognition is performed to obtain the target number and confidence score. Simultaneously, edge detection and fitting are performed based on the rotation main mask to calculate the rotation center coordinates. Then, the rotation center coordinates are accurately mapped back to the aerial image coordinate system using the inverse rotation matrix transformation to obtain the absolute pixel coordinates, ensuring a precise correspondence between the control point pixel coordinates and their numbers. Finally, the absolute pixel coordinates, target numbers, and confidence scores are combined to generate control point observation records, providing complete spatial constraint data for the 3D reconstruction software. This coarse-to-fine layered processing flow effectively addresses the challenge of processing hundreds to thousands of high-resolution images in a single aerial photography mission while ensuring the accuracy of control point extraction, improving the efficiency of control point observation record generation during the 3D reconstruction of earthwork construction sites. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a three-dimensional reconstruction puncture point method for earthwork construction sites disclosed in an embodiment of this application. Figure 2 This is another schematic flowchart of a three-dimensional reconstruction puncture point method for earthwork construction sites disclosed in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a system provided in an embodiment of this application.

[0019] Explanation of reference numerals in the attached drawings: 301, Central Processing Unit; 302, Read-Only Memory; 303, Random Access Memory; 304, Bus; 305, Input / Output Interface; 306, Input Section; 307, Output Section; 308, Storage Section; 309, Communication Section; 310, Driver; 311, Removable Media. Detailed Implementation

[0020] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0021] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0022] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0023] This application provides a method for three-dimensional reconstruction of puncture points for earthwork construction sites. In one embodiment, please refer to... Figure 1 , Figure 1 This is a flowchart illustrating a 3D reconstruction puncture method for earthwork construction sites provided in this application embodiment. This method can be implemented using a computer program, which can be integrated into an application or run as a standalone utility application. The method can also be implemented using a microcontroller or run on a 3D reconstruction puncture system for earthwork construction sites based on the von Neumann architecture. Specifically, the method may include the following steps: Step 101: Acquire aerial images of the site to be tested.

[0024] The site to be measured refers to the construction area where earthwork volume measurement or 3D reconstruction is required. This typically includes construction sites, mines, road construction sections, and other measurement objects with clearly defined geographical boundaries. Aerial imagery refers to two-dimensional raster image data obtained by vertically or obliquely capturing ground scenes using a digital camera mounted on a drone platform. Each image contains RGB three-channel color information, geographic reference information (such as GPS coordinates and shooting altitude), and camera intrinsic data (focal length, distortion coefficient). For example, for a 100m × 80m foundation pit, using a DJI Phantom 4 RTK drone at an altitude of 80 meters with 80% forward overlap and 70% lateral overlap to perform raster aerial photography, a single mission generates approximately 150 JPEG images with a resolution of 5472×3648 pixels. Each image covers a ground area of ​​approximately 60m × 40m, with a ground sampling interval of approximately 2cm / pixel.

[0025] Specifically, the drone's autopilot system executes a pre-set flight path to collect a sequence of images covering the entire survey area above the test site according to specific flight parameters. The implementation involves: first, importing the boundary coordinates of the test site (usually stored in KML or SHP format) into the ground control software. Based on the boundary range, the set aerial altitude (typically 50-150 meters), camera sensor size, and focal length, the software automatically plans a parallel flight path, calculating the 3D coordinates and camera attitude angle of each shooting point to ensure sufficient overlap between adjacent images to meet the stereoscopic vision requirements of subsequent 3D reconstruction. After takeoff, the drone sequentially reaches each shooting point along the planned flight path, triggering the camera shutter while recording precise GPS position and IMU attitude data, embedding this metadata into the EXIF ​​information of the images. After shooting, the image data is imported into a computer via an onboard memory card or wireless transmission, forming a complete aerial image set. Each image filename typically includes a sequence number, timestamp, and other identifying information for subsequent batch processing and management.

[0026] Step 102: Downsample the aerial image to generate a thumbnail image, record the scaling factor, and extract color features from the thumbnail image to obtain a candidate mask.

[0027] Downsampling refers to pixel-level dimensionality reduction of the original high-resolution image, reducing the amount of data by decreasing the spatial resolution of the image. Common algorithms include nearest neighbor interpolation, bilinear interpolation, or region averaging. A thumbnail image is a low-resolution image generated after downsampling, maintaining the aspect ratio of the original image but significantly reducing the number of pixels. The scaling factor is the ratio of the original image size to the thumbnail image size. For example, if the original image is 5400×3600 pixels and the thumbnail image is 540×360 pixels, then the scaling factor is 10. Color feature extraction refers to identifying regions with specific color characteristics by calculating the color space attributes of pixels (such as hue, saturation, and brightness in the HSV color space). Dedicated detection algorithms are designed for the high-saturation color features of spikes. A candidate mask is a binary matrix with the same size as the thumbnail image, where pixels with a value of 1 represent candidate regions whose color features match the spike features, and pixels with a value of 0 represent background regions.

[0028] Specifically, each aerial image is first downsampled using a bicubic interpolation algorithm to reduce the original image by a fixed ratio (usually 8-12 times). For example, a 5472×3648 pixel image is reduced to 684×456 pixels, and the scaling factor s=8 is recorded for subsequent coordinate mapping. The downsampled thumbnail image data size is reduced to 1 / 64 of the original, and the processing speed is increased by about 60-80 times, making rapid scanning feasible. Next, the thumbnail image is converted from the RGB color space to the HSV color space. In the HSV space, the H channel represents hue (0-360 degrees), the S channel represents saturation (0-1), and the V channel represents brightness (0-1). This representation is more in line with the human eye's color perception characteristics and has better robustness to changes in lighting. Considering that spurs typically use high-purity colors such as red (H≈0 or 360 degrees), yellow (H≈60 degrees), and blue (H≈240 degrees), a color purity index is designed: P = S×V×sin²(π×min(|H-Htarget|,360-|H-Htarget|) / 180), where Htarget is the target hue value. This index comprehensively considers saturation, brightness, and the closeness of the hue to the target color. Each pixel in the thumbnail image is traversed, and its purity value under the three target colors (red, yellow, and blue) is calculated. The maximum value is taken as the color feature response intensity of that pixel. A response threshold Pthresh = 0.6 is set. Pixels with a response intensity greater than the threshold are marked as 1, and the remaining pixels are marked as 0, generating a candidate mask matrix M. This matrix reflects the location distribution of all possible colored spurs in the thumbnail image, providing input for subsequent region selection.

[0029] Step 103: Perform connected component analysis on the candidate mask to filter out the target candidate region, and map the target candidate region to the aerial image according to the scaling factor to extract the target resolution slice.

[0030] A target resolution slice refers to a rectangular sub-image containing a single spur point extracted from the original aerial image. The slice size is usually set to 1.5-2 times the expected size of the spur point to ensure complete coverage and allowance for the edges.

[0031] Specifically, the connected component labeling algorithm is first applied to the candidate mask M using a two-pass scanning method: the first pass scans from the top left to the bottom right, assigning a temporary label to each foreground pixel and recording equivalence relationships; the second pass merges the labels based on the equivalence relationships, ultimately assigning a unique integer ID to each connected component. All labeled connected components are traversed, and the geometric feature parameters of each component are calculated, including the number of pixels (Area), the width W and height H of the minimum bounding rectangle, the aspect ratio Ratio = max(W, H) / min(W, H), and the compactness = 4π × Area / (Perimeter²). The selection criteria are set as follows: area range Amin ≤ Area ≤ Amax (e.g., 100 ≤ Area ≤ 2000 pixels, corresponding to the size of punctures in the thumbnail image), aspect ratio range 0.7 ≤ Ratio ≤ 1.3 (close to a square), and compactness > 0.6 (excluding elongated or irregular shapes). Connected components that meet all these conditions are retained as target candidate regions, while other connected components are discarded. For each retained target candidate region, obtain its bounding box coordinates (xthumb, ythumb, wthumb, hthumb) in the thumbnail image. Calculate its corresponding position in the original image (xorig, yorig) = (xthumb × s, ythumb × s) based on the scaling factor s. To ensure the slice contains complete spurs and surrounding context information, extend the margin by 20 pixels around the mapped coordinates. Calculate the final bounding box coordinates of the slice: xstart = max(0, xorig - margin), ystart = max(0, yorig - margin), xend = min(ImageWidth, xorig + wthumb × s + margin), yend = min(ImageHeight, yorig + hthumb × s + margin). Extract the pixel data of this rectangular region from the original aerial image to generate the target resolution slice image (Islice). This mapping and slicing extraction process allows subsequent fine processing to target only key areas in the original high-resolution image, avoiding full-pixel scanning of the entire image and reducing computational load while maintaining accuracy.

[0032] In one possible implementation, connected component analysis is performed on the candidate mask to filter out the target candidate region, and the target candidate region is mapped onto the aerial image according to the scaling factor to extract a target resolution slice. Specifically, steps 1031-1033 are included, as follows: Step 1031: Identify multiple connected regions in the candidate mask and calculate the geometric feature parameters of each connected region. The geometric feature parameters include pixel area, aspect ratio of the bounding rectangle, and circularity.

[0033] Specifically, a connected component labeling algorithm is performed on the candidate mask, employing a two-pass scanning method for fast labeling. The first scan starts from the top left corner and processes row by row. When a foreground pixel is encountered, the labels of its left and top neighboring pixels are checked. If the neighboring pixel has no label, a new label is assigned; if it has a label, that label is inherited. If the left and top labels are different, the equivalence relation is recorded in the disjoint-set data structure. The second scan merges the labels according to the equivalence relation of the disjoint-set data structure, assigning a unique integer ID to each connected region. After labeling, dozens to hundreds of connected regions are typically obtained. All labeled connected regions are traversed, and a mapping table is established from region IDs to a list of pixel coordinates. The pixel area of ​​each connected region is calculated, and the Area value is obtained by accumulating the number of pixels contained in the region. Calculate the circumscribed rectangle by traversing all pixel coordinates within the region, determining the minimum xmin, maximum xmax, minimum ymin, and maximum ymax of the x-coordinate. The width of the circumscribed rectangle is width = xmax - xmin + 1, the height is height = ymax - ymin + 1, and the aspect ratio is Ratio = max(width, height) / min(width, height). To calculate circularity, first extract the region outline. Use a boundary tracing algorithm to start from any starting point on the region edge, sequentially visiting boundary pixels in a clockwise or counterclockwise direction until returning to the starting point. Count the number of outline pixels as the perimeter (Perimeter), and substitute them into the formula Circularity = 4π × Area / Perimeter² to calculate circularity. Store the calculated pixel area (Area), aspect ratio (Ratio), and circularity as a geometric feature parameter triple (Area, Ratio, Circularity) for the connected region in the feature parameter table. For example, for a connected region with ID 15, the calculated feature parameters (284, 1.15, 0.81) indicate that the region contains 284 pixels, has an aspect ratio of 1.15 which is close to a square, and a circularity of 0.81 which is close to a circle or a square. These parameters will be used for subsequent target score calculations.

[0034] Step 1032: Perform weighted calculation based on geometric feature parameters to determine the target score of each connected region. Connected regions with target scores greater than a preset score threshold are selected as target candidate regions, and the boundary coordinates of the target candidate regions in the thumbnail image are obtained.

[0035] Specifically, first, the pixel area Area is scored. The ideal area range [Amin, Amax] is set to [150, 1800] pixels (corresponding to a square area with side length 12 - 42 pixels in the thumbnail image). The area scoring function is defined as follows: when Area < Amin, ScoreArea = Area / Amin; when Amin ≤ Area ≤ Amax, ScoreArea = 1; when Area > Amax, ScoreArea = exp(-(Area - Amax) / Amax). This function has a score of 1 within the ideal range and the score decreases when it is too small or too large. The aspect ratio Ratio is scored. The fiducial points are usually square or circular, and the ideal aspect ratio is close to 1. The aspect ratio scoring function ScoreRatio = exp(-α × (Ratio - 1)²) is set, where α = 2 is the sensitivity coefficient. When Ratio = 1, ScoreRatio = 1, and the scoring index decreases when the aspect ratio deviates from 1. The circularity Circularity is scored. The ideal circularity range is [0.7, 1.0]. The circularity scoring function is defined as follows: when Circularity < 0.7, ScoreCirc = Circularity / 0.7; when Circularity ≥ 0.7, ScoreCirc = 1. The target score is calculated by integrating the three scoring components. The weighted geometric mean formula Score = (ScoreArea^w1 × ScoreRatio^w2× ScoreCirc^w3)^(1 / (w1+w2+w3)) is used, and the weight coefficients are set to w1 = 0.4, w2 = 0.3, and w3 = 0.3. The geometric mean is more sensitive to low-scoring items than the arithmetic mean, and a significant deviation in any one feature will cause a significant decrease in the total score. For example, the feature parameters of a certain connected region are (284, 1.15, 0.81). Calculate the scores of each component: ScoreArea = 1 (284 is within the ideal range), ScoreRatio = exp(-2×(1.15 - 1)²) ≈ 0.956, ScoreCirc = 1 (0.81 is within the ideal range). The target score Score = (1^0.4 × 0.956^0.3 × 1^0.3)^(1 / 1) ≈ 0.986. The preset score threshold is set to 0.70. Traverse the target scores of all connected regions, filter out the connected regions with Score > 0.70, and add their IDs to the target candidate region list.For each retained target candidate region, query its pixel coordinate set and calculate the extreme values ​​of the horizontal and vertical coordinates to determine the boundary coordinates: xmin is the minimum horizontal coordinate of all pixels in the region, xmax is the maximum, ymin is the minimum vertical coordinate, and ymax is the maximum. The boundary width w = xmax - xmin + 1, and the boundary height h = ymax - ymin + 1. The boundary coordinates are represented as tuples (xmin, ymin, w, h) and stored. For example, the boundary coordinates of a target candidate region are (45, 78, 18, 16), indicating that the region lies within a rectangle with horizontal coordinates of 45-62 and vertical coordinates of 78-93 in the thumbnail image.

[0036] Step 1033: Compensate the boundary coordinates using the scaling factor to obtain the mapped boundary coordinates; determine the cropping range in the aerial image based on the mapped boundary coordinates, and extract a target resolution slice based on the cropping range.

[0037] Specifically, for the boundary coordinates (xmin, ymin, w, h) of each target candidate region, coordinate compensation calculation is performed: xmin_map = xmin × s, ymin_map = ymin × s, w_map = w × s, h_map = h × s, to obtain the mapped boundary coordinates. For example, if the boundary coordinates in the thumbnail image are (45, 78, 18, 16) and the scaling factor s = 8, then the mapped boundary coordinates are (360, 624, 144, 128), indicating that the puncture point is located in a rectangular area with horizontal coordinates of 360-503 and vertical coordinates of 624-751 in the original image. To ensure the slice includes the complete boundary and surrounding environment information of the puncture point, an extended margin is set based on the mapped boundary coordinates. The margin size is set to 20% of the mapped boundary size, i.e., margin_x = 0.2 × w_map, margin_y = 0.2 × h_map. The starting coordinates of the cropping range are calculated as: xstart = max(0, xmin_map - margin_x), ystart = max(0, ymin_map - margin_y), and the ending coordinates are: xend = min(ImageWidth, xmin_map + w_map + margin_x), yend = min(ImageHeight, ymin_map + h_map + margin_y), where ImageWidth and ImageHeight are the width and height of the original aerial image. The max and min functions ensure that the cropping range does not exceed the image boundary, avoiding out-of-bounds access. For example, if the mapped boundary coordinates are (360, 624, 144, 128), the extended margin is (29, 26), and the original image size is 5400×3600 pixels, then the cropping range is xstart = max(0, 360-29) = 331, ystart = max(0, 624-26) = 598, xend = min(5400, 360+144+29) = 533, and yend = min(3600, 624+128+26) = 778. Based on the cropping range (331, 598, 533, 778), pixel data is extracted from the original aerial image to create a target resolution slice image.In practice, the original aerial image file is read into memory. The pixel matrix is ​​accessed according to the row and column indices of the cropping range, and the slice is extracted as `Islice = Ioriginal[ystart: yend, xstart: xend, :]` (assuming the image is RGB three-channel). This generates a slice array of size (yend-ystart) × (xend-xstart) × 3 pixels; in this example, the slice size is 180×202×3 pixels. The slice data is saved as a separate image file or kept in memory for subsequent processing. Simultaneously, the starting coordinates (xstart, ystart) of the slice in the original image are recorded. These coordinates are used to convert the local coordinates within the slice back to global coordinates. Through this mapping and slice extraction process, coarse localization from the thumbnail image to precise cropping of the original image is achieved. The obtained target resolution slice contains complete high-resolution details of the spikes, providing high-quality input data for subsequent purity analysis, rotation correction, and character recognition.

[0038] Step 104: Perform purity analysis on the target resolution slice, separate the main mask and the numbered mask, calculate the rotation matrix based on the main mask, and use the rotation matrix to perform affine transformation on the target resolution slice, the main mask and the numbered mask to obtain the rotated slice, the rotated main mask and the rotated numbered mask respectively.

[0039] The main body mask refers to the binary image of the colored main body of the puncture point, where pixels with a value of 1 correspond to a highly saturated solid color region. This region is used to determine the overall orientation and position of the puncture point. The number mask refers to the binary image of the region marking the number of the puncture point, where pixels with a value of 1 correspond to the numbered label portion with black text on a white background. This region is used for subsequent character recognition. The rotation matrix refers to a 2×2 orthogonal matrix R = [[cosθ, -sinθ], [sinθ, cosθ]], used to rotate the points in the image counterclockwise by an angle θ from the origin, forming a complete affine transformation in conjunction with the translation vector.

[0040] Specifically, the target resolution slice is first converted from RGB space to HSV space, and the color purity index Ppurity = S × V × (1 - |H - Hmain| / 180) is calculated pixel by pixel, where Hmain is the dominant hue of the main color (determined by statistically analyzing the hue distribution of high-saturation pixels in the slice and taking the mode). A high purity threshold Phigh = 0.7 and a low purity threshold Plow = 0.3 are set to generate the main body mask Mmain = {Ppurity ≥ Phigh} and the numbered mask Mnum = {Ppurity ≥ Phigh}.<Plow AND V> 0.8}, the latter condition ensures the high brightness feature of the numbered region. The minimum bounding rectangle algorithm is applied to the main body mask Mmain. The rotating caliper method is used to traverse the convex hull vertices of the main body contour, calculating the minimum area rectangle enclosing all foreground pixels. The coordinates of the four vertices of this rectangle are {P1, P2, P3, P4}, and the direction vector of the long side of the rectangle is v = P2 - P1. The angle θ = arctan2(vy, vx) between this vector and the horizontal axis is calculated; this angle is the tilt angle of the puncture point relative to the horizontal direction. A rotation matrix R(-θ) = [[cos(-θ), -sin(-θ)], [sin(-θ), cos(-θ)]] is constructed, where the negative sign indicates clockwise rotation to correct the tilted puncture point to the horizontal direction. The rotation center coordinates center = (xslice / 2, yslice / 2) are calculated as the geometric center of the sliced ​​image. A complete affine transformation matrix T = [[cos(-θ), -sin(-θ), tx], [sin(-θ), cos(-θ), ty], [0, 0, 1]] is constructed, where the translation parameters tx = center.x - (center.x × cos(-θ) - center.y × sin(-θ)) and ty = center.y - (center.x × sin(-θ) + center.y × cos(-θ)) ensure that the rotation is performed around the slice center. Affine transformations are applied to the target resolution slice Islice, the main body mask Mmain, and the numbered mask Mnum, respectively. The rotated pixel values ​​are resampled using bilinear interpolation to generate the rotated slice Irotated, the rotated main body mask Mrotated_main, and the rotated numbered mask Mrotated_num. The rotated spikes are arranged in a regular horizontal or vertical pattern, and the numbered text is also adjusted to a standard upright orientation, eliminating the interference of angular deviations on subsequent character recognition and center localization. Simultaneously, the rotation matrix R(-θ) and the rotation center coordinates center are stored for the inverse transformation of the final coordinates.

[0041] In one possible implementation, a purity analysis is performed on the target resolution slice to separate the main mask and the numbered mask. A rotation matrix is ​​calculated based on the main mask, and an affine transformation is performed on the target resolution slice, the main mask, and the numbered mask using the rotation matrix to obtain a rotated slice, a rotated main mask, and a rotated numbered mask, respectively. Specifically, steps 1041-1044 are included, as follows: Step 1041: Calculate the color purity index of each pixel in the target resolution slice; divide the pixels with a color purity index greater than or equal to the purity threshold into the main region and generate the main mask; divide the pixels with a color purity index less than the purity threshold and a brightness greater than the brightness threshold into numbered regions and generate numbered masks.

[0042] Specifically, color analysis is performed on each pixel of the target resolution slice. First, the slice is converted from the RGB color space to the HSV space, and the RGB values ​​are normalized to the range of 0 to 1. The maximum, minimum, and their differences are calculated to obtain the hue, saturation, and lightness values ​​of each pixel. The hue distribution of high-saturation pixels in the slice is statistically analyzed, and the hue with the highest frequency is selected as the dominant hue of the spike. A hue matching function is defined; when the pixel hue is close to the dominant hue, the function value approaches 1, and it decays exponentially when it deviates. The saturation, lightness, and hue matching function are multiplied to obtain the color purity index. For example, a pixel with an HSV value of hue 25, saturation 0.82, and lightness 0.91, and a dominant hue of 30, has a hue matching function of approximately 0.975 and a purity index of 0.728. A purity threshold of 0.60 is set. All pixels are iterated through; when the purity index reaches or exceeds this threshold, the pixel is marked as a main region, and a value of 1 is assigned to the corresponding position in the main mask. For the identification of numbered regions, a brightness threshold of 0.75 is set. When a pixel's purity index is lower than the purity threshold but its brightness exceeds the brightness threshold, the pixel is determined to belong to the white background of the numbered region, and a value of 1 is assigned to the corresponding position in the numbered mask. After traversal, the numbered mask undergoes morphological processing. Opening operations are applied to remove isolated noise points, and closing operations are applied to fill small holes inside the region, resulting in a continuous and complete numbered region mask. This step achieves precise separation of the puncture subject and the numbered region through pixel-level color purity analysis, providing a clear region division for subsequent rotation correction and character recognition.

[0043] Step 1042: Extract the outer contour of the main body mask, calculate the minimum bounding rectangle of the outer contour, and obtain the coordinates of the center point and the tilt angle of the minimum bounding rectangle.

[0044] Specifically, a contour extraction algorithm is applied to the main body mask, and a boundary tracking method is used to identify the foreground region boundary from the mask. The mask is scanned to find the first foreground pixel as the starting point. From the starting point, eight neighboring pixels are checked clockwise, the next boundary pixel is selected and its coordinates are recorded, and this process is repeated until the starting point is reached, resulting in a complete contour point sequence. If the mask contains multiple connected regions, the region with the most pixels is selected as the main puncture point. A curve simplification algorithm is applied to the extracted outer contour, with a simplification parameter set to 2 pixels. Key inflection points are retained, reducing computation while preserving the contour shape characteristics. The convex hull of the contour is calculated to obtain the minimum convex polygon vertex sequence enclosing the contour. A rotational caliper algorithm is applied to the convex hull to calculate the minimum bounding rectangle. A pair of parallel lines are initialized to coincide with one edge of the convex hull, and they are gradually rotated to align with each edge of the convex hull. For each rotation angle, the area of ​​the rectangle enclosing the convex hull is calculated, and the parameters of the rectangle with the smallest area are recorded. For each rotation angle, all vertices of the convex hull are projected onto the rectangle's edge direction and normal direction, and the projection range determines the rectangle's length and width. After traversing all edge directions, the rectangle with the smallest area is selected as the minimum bounding rectangle. The coordinates of the four vertices are reconstructed using the center projection coordinates of the rectangle and half its length and half its width. The center point coordinates are calculated as the arithmetic mean of the four vertices. When calculating the tilt angle, the angle between the long side direction vector of the rectangle and the horizontal axis is taken, with the angle range adjusted to between -90 degrees and +90 degrees. For example, if the calculated long side direction vectors are 0.866 and 0.5, the tilt angle is approximately 30 degrees, indicating that the puncture point is tilted 30 degrees clockwise relative to the horizontal direction. The coordinates of the center point of the smallest bounding rectangle and the tilt angle are output. These two parameters fully describe the position and orientation of the puncture point, providing crucial input for constructing the rotation matrix.

[0045] Step 1043: Construct a rotation matrix with the center point coordinates as the rotation center and the negative values ​​of the tilt angle as the rotation amount.

[0046] Specifically, a rotation transformation matrix is ​​constructed around an arbitrary center point. The standard rotation matrix only applies to rotations around the origin, requiring coordinate translation to move the rotation center to the origin, perform the rotation, and then translate back to the original position. The complete transformation process includes three steps: first, translating the origin of the coordinate system to the rotation center; second, performing a rotation around the origin with a negative tilt angle; and third, translating the origin back to the original position. Homogeneous coordinate representation is used to unify the three-step transformation, extending the two-dimensional coordinates to three-dimensional homogeneous coordinates, and constructing an affine transformation matrix. The first-step translation matrix subtracts the rotation center coordinates from all point coordinates; the second-step rotation matrix is ​​constructed based on the cosine and sine values ​​of the rotation angle; and the third-step translation matrix adds the rotation center coordinates to all point coordinates. The combined transformation matrix is ​​calculated using matrix multiplication: first, the product of the rotation matrix and the first-step translation matrix is ​​calculated, then multiplied by the third-step translation matrix on the left. The simplified transformation matrix includes both rotation and translation components; the translation component is calculated by combining the rotation center coordinates and the cosine and sine values ​​of the rotation angle. For example, with center point coordinates of 137.5 and 102.5, a tilt angle of 30 degrees, a rotation amount of -30 degrees, a cosine value of 0.866, and a sine value of -0.5, a rotation matrix is ​​constructed after calculating the translation components. This matrix maps any point in the slice to positive rotation coordinates, achieving precise rotation correction around the center of the puncture point.

[0047] Step 1044: Use the rotation matrix to resample the target resolution slice, the main body mask, and the numbered mask to generate the rotated slice, the rotated main body mask, and the rotated numbered mask.

[0048] Specifically, a rotation matrix is ​​applied to the target resolution slice, the subject mask, and the numbered mask for geometric transformation. First, the size of the rectified image is determined, and the coordinates of the four corner points of the original slice after the rotation matrix transformation are calculated. The extreme values ​​of the x and y coordinates are used to determine the bounding box of the rectified image. The rectified slice, the rectified subject mask, and the rectified numbered mask are initialized as blank matrices with the same rectified size. A reverse mapping algorithm is used for pixel resampling, traversing the coordinates of each pixel in the rectified image and calculating its corresponding position in the original slice. The inverse of the rotation matrix is ​​calculated; for orthogonal rotation matrices, the inverse is equal to its transpose. The inverse transformation is applied to calculate the original coordinates, obtaining the original slice coordinates corresponding to each pixel in the rectified image. Since these coordinates are usually non-integer, an interpolation algorithm is needed to obtain the pixel value. For RGB slices, bilinear interpolation is used. The pixel values ​​of the four neighboring pixels around the coordinates in the original slice are read, and the interpolation result is calculated by weighted averaging based on the decimal part of the coordinates. The interpolation result is then assigned to the corresponding position in the rectified slice. For binary masks, nearest neighbor interpolation is used. The mask value at the rounded coordinate position is directly read and assigned to the corresponding position in the rotation mask. If the coordinates calculated by the inverse transformation exceed the original slice boundary, the pixel value at that position in the rotation image is set as the background value. After traversing all pixels, a complete rotation slice, a rotation subject mask, and a rotation number mask are generated. In the rotation slice, the puncture points present a standard horizontal or vertical orientation. The outer rectangle of the subject mask is parallel to the image coordinate axes, and the numbered text is also adjusted to be upright, providing standardized input data for subsequent character recognition and center localization. At the same time, the inverse matrix of the rotation matrix and the rotation image size parameters are saved, which are used to finally inversely transform the center point coordinates in the rotation coordinate system back to the original slice coordinate system.

[0049] Step 105: Extract the numbered sub-image from the rotated slice based on the rotated number mask, and perform character recognition on the numbered sub-image to obtain the target number and confidence level.

[0050] The numbered sub-image refers to a rectangular image block containing only the numbered text, which is cropped from the rotated slice according to the foreground region boundary of the rotated numbered mask. This image is compact in size, covers only the effective area of ​​the numbered label, and minimizes background noise.

[0051] Specifically, firstly, connected component analysis is performed on the rotated number mask `Mrotated_num` to identify all independent foreground regions, and the minimum bounding box (x, y, w, h) of each connected component is calculated. Since the puncture point numbers are usually located at specific positions of the puncture point body (as below or to the right), the relative positional relationship between the connected components and the centroid of the rotated body mask is analyzed to filter out candidate numbering regions with reasonable positions. For the retained numbering regions, the margin is extended outward by 5-10 pixels to include the complete stroke of the character. The corresponding rectangular region is extracted from the rotated slice `Irotated` to generate the numbering sub-image `Inumber`, which is typically 80-150 pixels wide and 30-50 pixels high. The numbering sub-image is preprocessed to improve the accuracy of character recognition: firstly, adaptive histogram equalization is applied to enhance contrast, and then Otsu thresholding is used to generate a binarized image with black characters in the foreground and white characters in the background. Morphological operations (opening operations) are applied to remove small noise points, and then closing operations are applied to connect broken strokes. The preprocessed binarized number image is input into the OCR engine, using either the TesseractOCR library or a custom-trained digit recognition model. The OCR algorithm first performs character segmentation, dividing the number image into individual character image blocks through projection analysis or connected component analysis. Features (such as pixel density distribution, edge orientation histogram, and HOG features) are extracted for each character image block and matched against a pre-trained character template library to calculate the similarity score for each candidate character. Dynamic programming or the Viterbi algorithm combined with a language model (constraints such as consecutive digits and letters preceding numbers) are used to determine the optimal character sequence. The OCR algorithm outputs the target number string num_str, for example, "023", along with the recognition confidence score conf. The confidence score is calculated as the arithmetic mean of the similarity scores for each recognized character: conf = (1 / N) × Σ score_i, where N is the total number of characters and score_i is the similarity score of the i-th character. A confidence score above 0.9 generally indicates highly reliable identification results, 0.7-0.9 indicates moderate reliability requiring attention, and below 0.7 indicates questionable identification results requiring manual verification. The final output is structured data containing the target number num_str and the confidence score conf, providing crucial information for subsequent control point record generation.

[0052] In one possible implementation, numbered sub-images are extracted from the rotated slice based on the rotated number mask, and character recognition is performed on the numbered sub-images to obtain the target number and confidence level. Specifically, this includes steps 1051-1054, as follows: Step 1051: Determine the bounding box of the foreground pixels in the rotated number mask; determine the numbered sub-image from the rotated slice based on the bounding box.

[0053] Specifically, the numbering region is located and the numbering sub-image is extracted from the rotated numbering mask. First, all pixels of the rotated numbering mask are traversed, recording the coordinates of foreground pixels with a value of 1. During the traversal, the minimum and maximum values ​​of the foreground pixel horizontal coordinates are counted to obtain the left boundary x_min and right boundary x_max, and the minimum and maximum values ​​of the vertical coordinates are counted to obtain the upper boundary y_min and lower boundary y_max. If the numbering mask contains multiple disconnected foreground regions, the connected region with the most pixels is selected as the main numbering region, and its bounding box is calculated separately. To avoid the bounding box being too close to the numbering edge, causing character truncation, the bounding box is appropriately expanded. The left and right boundaries are expanded outwards by 5 to 10 pixels, and the upper and lower boundaries are expanded outwards by 3 to 5 pixels, ensuring that the numbering sub-image contains the complete character outline and appropriate margins. The coordinates of the expanded bounding box need to be limited to the effective range of the rotated slice: the left boundary is not less than 0, the right boundary does not exceed the slice width minus 1, the upper boundary is not less than 0, and the lower boundary does not exceed the slice height minus 1. Based on the finally determined bounding box coordinates, the numbering sub-image is extracted from the rotated slice by cropping. The cropping operation is achieved by reading the pixel data from rows y_min to y_max and columns x_min to x_max of the rotated slice, generating a numbered sub-image. For example, if the bounding box is initially at x-coordinates 150 to 380 and y-coordinates 220 to 280, and after expanding outwards by 8 pixels, it becomes x-coordinates 142 to 388 and y-coordinates 215 to 285. Cropping the pixel data within this range from the rotated slice yields a numbered sub-image of size 247 by 71 pixels. The numbered sub-image contains black or dark-colored numbered text on a white or light-colored background, providing clear input data for subsequent grayscale conversion, binarization, and character segmentation processing.

[0054] Step 1052: Perform grayscale conversion and binarization on the numbered sub-images to obtain a binarized image; perform character segmentation on the binarized image to obtain a single-character image sequence.

[0055] Specifically, preprocessing and character segmentation are performed on the numbered sub-images. First, grayscale conversion is performed, using a weighted average method to calculate the grayscale value of each pixel. The formula is: grayscale value equals 0.299 times the red channel value plus 0.587 times the green channel value plus 0.114 times the blue channel value. This weighting distribution conforms to the human eye's perception of brightness for different colors. After obtaining the grayscale image, Gaussian blur filtering is applied using a 5x5 Gaussian kernel with a standard deviation of 1.0 to remove image noise and smooth edges. Then, binarization is performed using an adaptive thresholding method. The image is divided into multiple local regions, and a local threshold is calculated for each region based on the average grayscale value of its neighboring pixels. The local threshold equals the average grayscale value of the neighborhood pixels minus a constant offset, set to 10 to 15. The adaptive thresholding method adapts to changes in illumination in different areas of the image, ensuring clear separation between characters and the background. Pixels with grayscale values ​​below the local threshold are set to 0 to represent black characters, and pixels with grayscale values ​​above the threshold are set to 255 to represent white backgrounds, generating a binarized image. Morphological processing is performed on the binarized image. First, erosion is applied to remove small noise and burrs, then dilation is applied to restore the main structure of the character. A 3x3 rectangular kernel is used as the structuring element. Next, character segmentation is performed. The vertical projection method is used to count the number of black pixels in each column of the binarized image, and a vertical projection histogram is drawn. In the projection histogram, character regions correspond to peak values, and character gaps correspond to valley values. The segmentation threshold is set to 15% of the maximum value of the projection histogram. The projection histogram is scanned, and when the projection value of consecutive columns is lower than the segmentation threshold, it is marked as a character gap; when the projection value is higher than the segmentation threshold, it is marked as a character region. The left and right boundaries of each character are determined based on the position of the character gaps, and the complete bounding box of the character is determined by combining the upper and lower boundaries of the black pixels in each character region. The bounding box regions of each character are cropped from the binarized image from left to right to obtain a sequence of single-character images. Each single-character image contains a complete numeric or alphanumeric character, and the size depends on the character size, typically 20 to 50 pixels wide and 40 to 60 pixels high. The size of single-character images is normalized and uniformly scaled to a standard size of 32 by 32 pixels. Bicubic interpolation is used to maintain the clarity of the character shape, providing standardized input data for subsequent feature extraction and template matching.

[0056] Step 1053: Extract the feature vector of each character image in the single character image sequence, perform similarity matching between the feature vector and the pre-stored standard character template library, and determine the matching character and matching similarity corresponding to each character image.

[0057] Specifically, feature extraction and template matching are performed on each character image in the single-character image sequence. First, the feature vector of the character image is extracted using histogram of oriented gradients (HOR). The 32x32 pixel character image is divided into a 4x4 grid, with each grid cell being 8x8 pixels. For each pixel within a cell, the gradient magnitude and gradient direction are calculated. The horizontal component of the gradient equals the gray value of the right pixel minus the gray value of the left pixel, and the vertical component equals the gray value of the lower pixel minus the gray value of the upper pixel. The gradient magnitude is the square root of the sum of the squares of the horizontal and vertical components, and the gradient direction is the arctangent of the vertical component divided by the horizontal component. The gradient direction range from 0 to 180 degrees is divided into 9 directional intervals, each 20 degrees. The gradient magnitudes of each directional interval within each grid cell are summed to obtain a 9-dimensional histogram of orientation. The histograms of the 16 cells in the 4x4 grid are concatenated to form a 144-dimensional feature vector. The feature vectors are normalized by calculating their L2 norm. Each dimension value is then divided by the L2 norm to obtain a normalized feature vector of unit length. The extracted feature vectors are then matched for similarity with template feature vectors in a standard character template library containing 36 standard character feature vectors: digits 0-9 and uppercase letters A-Z. Cosine similarity is used as the similarity metric. The dot product between the feature vector to be identified and each template feature vector is calculated; the dot product is equal to the sum of the product of the corresponding elements of the two vectors. Since the feature vectors are normalized to unit vectors, the cosine similarity is directly equal to the dot product, ranging from -1 to +1, with values ​​closer to +1 indicating higher similarity. All templates in the template library are traversed, and the cosine similarity between the character to be identified and each template is calculated. The character label corresponding to the template with the highest similarity is recorded as the matching character, and this maximum similarity value is recorded as the matching similarity. For example, if the cosine similarity between the feature vector of a character image and the feature vector of the digit 7 in the template library is 0.92, and the similarity with other characters is all below 0.85, then the matching character is determined to be 7, with a matching similarity of 0.92. The matching similarity range from -1 to +1 is mapped to the range of 0 to 1 using the formula: the mapped similarity equals the original similarity plus 1, divided by 2, resulting in a standardized matching similarity. This feature extraction and matching process is repeated for each character in the single-character image sequence to obtain the corresponding matching character and matching similarity, providing basic data for final number recognition and confidence calculation.

[0058] Step 1054: Concatenate the matching characters corresponding to each character image in order to generate a target number, and select the minimum value among the matching similarities corresponding to each character image as the confidence score.

[0059] Specifically, the process generates a complete target number and calculates the recognition confidence based on the matching results of each character. First, following the original order of the single-character image sequence (left-to-right spatial order), the matching characters corresponding to each character image are read sequentially. The matching characters are then concatenated into a string without separators or spaces to form the target number. For example, if a single-character image sequence contains 5 characters and the matching characters are A, 3, 7, B, and 2, the resulting target number is A372B. The target number format is checked to ensure it conforms to the dot-matrix numbering standard, which typically consists of numbers and letters of a specific length, following fixed format rules. If the target number length does not match the expected length or contains illegal characters, it is marked as a suspected misidentification. Next, the recognition confidence is calculated by reading the matching similarity for each character in the single-character image sequence and storing these similarity values ​​in an array. The similarity array is traversed, and the minimum matching similarity value is found and used as the overall recognition confidence. The minimum value is used because the number recognition uses a serial concatenation method; any misidentification of any character results in an incorrect number. For example, if the similarity scores of the five characters are 0.95, 0.88, 0.92, 0.78, and 0.91, the minimum value of 0.78 is taken as the confidence score of the target number. A confidence threshold of 0.70 is set. When the confidence score is higher than or equal to this threshold, the number recognition result is considered reliable. When the confidence score is lower than the threshold, the number is marked as a low-confidence result, requiring manual review or the use of alternative recognition methods. The final target number string and its corresponding confidence score value are output. The recognition result is then associated with and stored with other attribute information of the puncture point, providing complete number recognition data for subsequent puncture point management and quality control.

[0060] Step 106: Based on the rotation mask, perform edge detection and fitting on the rotation slice, calculate the rotation center coordinates, and use the rotation matrix to inversely transform the rotation center coordinates to the coordinate system of the aerial image to obtain the absolute pixel coordinates. Combine the absolute pixel coordinates, target number and confidence level to generate control point observation records.

[0061] Absolute pixel coordinates refer to the two-dimensional coordinates (xabs, yabs) of the puncture point center in the original aerial image, with the origin (0, 0) at the top left corner of the image, the x-axis pointing to the right, and the y-axis pointing downwards, in pixels. Control point observation records refer to structured data containing the location, identification number, and confidence level assessment of the puncture points in the image. They are usually stored in JSON, XML, or database record format for use by 3D reconstruction software.

[0062] Specifically, the Irotated slice is first converted to a grayscale image. The Canny edge detection algorithm is applied to extract edge features, with a high threshold of 100 and a low threshold of 50. A 3×3 Sobel operator is used to calculate the gradient magnitude and direction. Non-maximum suppression is used to refine the edges. An edge image, Eedge, is generated through double-threshold filtering and concatenation. A logical AND operation is performed between the Irotated main mask Mrotated_main and the edge image Eedge to extract the set of edge pixel coordinates {(xi, yi)} belonging only to the main puncture point, eliminating edge interference from the background and numbered regions. Considering the geometric characteristics of puncture points, which are typically square or circular, a least-squares elliptic fitting algorithm is used to process the edge point set, constructing an overdetermined system of equations AX = 0, where A is an N×6 matrix (N is the number of edge points), with each row being [xi², xiyi, yi², xi, yi, 1], and X is the elliptic parameter vector to be determined [a, b, c, d, e, f]. The eigenvector corresponding to the minimum eigenvalue is solved using singular value decomposition. The formulas for the center coordinates are derived from the ellipse parameters: xc = (b×e - 2×c×d) / (4×a×c - b²), yc = (b×d - 2×a×e) / (4×a×c - b²), yielding the center coordinates (xrot, yrot) in the normalized coordinate system. To further improve the positioning accuracy to the sub-pixel level, a gray-level weighted centroid is calculated within a 5×5 pixel window around the center coordinates: xsub = Σ(I(xi, yi) × xi) / ΣI(xi, yi), ysub = Σ(I(xi, yi) × yi) / ΣI(xi, yi), where I(xi, yi) is the pixel gray value. This weighted centroid considers the influence of gray-level distribution, achieving a positioning accuracy of 0.1-0.2 pixels. After obtaining the precise normalized center coordinates (xrot_sub, yrot_sub), an inverse transformation is applied to map them back to the original slice coordinate system. Construct the inverse rotation matrix R^T = [[cosθ, sinθ], [-sinθ, cosθ]], calculate the offset vector v = (xrot_sub - center.x, yrot_sub - center.y) relative to the rotation center, apply the rotation transformation v' = R^T × v to obtain the offset vector in the original slice coordinate system, add the rotation center coordinates to obtain the center coordinates (xslice, yslice) = center + v' in the slice coordinate system. Further convert the slice coordinates to absolute coordinates of the aerial image, use the recorded slice start position (xstart, ystart) to calculate the absolute pixel coordinates (xabs, yabs) = (xstart + xslice, ystart + yslice), which is the precise position of the puncture point center in the original aerial image.The calculated absolute pixel coordinates, the identified target number num_str, the confidence level conf, and the aerial image file name image_name are combined to form a control point observation record, represented in JSON format as: {"image": "IMG_0123.jpg", "control_point_id": "023", "pixel_x": 3456.78, "pixel_y": 2134.56, "confidence": 0.94}. This record is stored in a database or exported as a text file for import into 3D reconstruction software. The reconstruction software matches these pixel coordinates with the 3D coordinates measured on the ground as constraints for bundle adjustment, thus completing a high-precision 3D reconstruction task.

[0063] In one possible implementation, edge detection and fitting are performed on the normalized slice based on the normalized main mask, and the coordinates of the normalized center are calculated. Specifically, steps 1061-1063 are included, as follows: Step 1061: In the swivel slice, extract the pixel set within the area covered by the swivel main mask; use an edge detection algorithm to extract the edge pixels of the pixel set, and use an interpolation algorithm to interpolate the position of the edge pixels to obtain the sub-pixel edge point set.

[0064] Specifically, firstly, a pixel set is extracted based on the rotated main mask. All pixel positions of the mask are traversed, and when the mask value is 1, the rotated slice pixel value and coordinates at that position are recorded, forming a pixel set. The rotated slice is then converted to grayscale using a weighted average method to convert the RGB three channels into a single-channel grayscale image. The grayscale value is equal to 0.299 times the R channel, 0.587 times the G channel, and 0.114 times the B channel. A Gaussian filter is applied to the grayscale image using a 5x5 Gaussian kernel with a standard deviation of 1.5 to smooth image noise while preserving major edge features. The Canny edge detection algorithm is used to extract edge pixels, which consists of four steps. The first step calculates the image gradient using the Sobel operator to calculate the gradient in the horizontal and vertical directions. The horizontal gradient convolution kernel has coefficients of 0 in the middle column, -1, -2, -1 in the left column, and +1, +2, +1 in the right column. The vertical gradient convolution kernel has coefficients of 0 in the middle row, -1, -2, -1 in the upper row, and +1, +2, +1 in the lower row. The gradient magnitude is equal to the square root of the sum of the squares of the horizontal and vertical gradients, and the gradient direction is equal to the arctangent of the vertical gradient divided by the horizontal gradient. The second step performs non-maximum suppression, comparing the gradient magnitude of the current pixel with its neighboring pixels along the gradient direction, retaining only local maxima as candidate edge points, and suppressing non-edge regions. The third step applies dual-threshold detection, setting a high threshold at the 70th percentile of the gradient magnitude distribution and a low threshold at 40% of the high threshold. Pixels with gradient magnitudes higher than the high threshold are identified as strong edge points, while pixels between the low and high thresholds are marked as weak edge points. The fourth step uses edge connectivity to retain weak edge points connected to strong edge points and discard isolated weak edge points, resulting in a continuous and complete set of edge pixels. Only edge pixels within the area covered by the rotated main mask are retained, filtering out background edges. Sub-pixel precision interpolation is performed on the extracted edge pixels using parabolic interpolation. For each edge pixel, the gradient magnitudes of its two neighboring pixels along the gradient direction are read, forming a three-point gradient sequence. The gradient magnitudes at three points are fitted to a parabola, with the equation y = ax² + bx + c. The coefficients a, b, and c are solved using the coordinates of the three points. The x-coordinate of the parabola's vertex is calculated using the formula: x = -b / 2a. This x-coordinate represents the offset of the true edge position relative to an integer pixel position. The original edge pixel coordinates are added to the offset along the gradient direction to obtain sub-pixel precision edge point coordinates. For example, if the edge pixel coordinates are 126 and 99, the gradient direction is 45 degrees, and the gradient magnitudes at three points along this direction are 38, 52, and 41, after fitting the parabola, the vertex offset is calculated to be 0.27 pixels, resulting in sub-pixel edge point coordinates of 126.19 and 99.19. Interpolation is performed on all edge pixels to generate a sub-pixel edge point set, providing high-precision edge data for subsequent geometric contour fitting and ensuring center positioning accuracy reaches the 0.1 pixel level.

[0065] Step 1062: Perform geometric contour fitting on the sub-pixel edge point set, and extract the geometric center point of the fitted contour as the preliminary center coordinates.

[0066] Specifically, the shape characteristics of the sub-pixel edge point set are first analyzed, and the second-order moment matrix of the edge point coordinates is calculated. This matrix describes the spatial distribution of the point set. The second-order moment matrix is ​​a 2x2 matrix. The elements in the first row and first column are equal to the sum of the squares of all edge point x-coordinates minus the mean x-coordinate. The elements in the first row and second column and the second row and first column are equal to the sum of the products of the x-coordinate deviation and the y-coordinate deviation. The elements in the second row and second column are equal to the sum of the squares of the y-coordinate deviation. The eigenvalues ​​of the second-order moment matrix are calculated. The ratio of the two eigenvalues ​​reflects the aspect ratio of the point set. When the ratio is close to 1, the point set is circularly distributed; when the ratio is greater than 1.5, the point set is elliptically distributed. The fitting model is selected based on the aspect ratio. Circular fitting is used when the ratio is less than 1.3, and elliptical fitting is used when the ratio is greater than or equal to 1.3. For circular fitting, the least squares circular fitting algorithm is used. Let the center coordinates be xc and yc, and the radius be r. For each edge point coordinate xi and yi, the distance di from the point to the center is calculated. The distance is equal to the square root of the sum of the squares of the x-coordinate difference and the squares of the y-coordinate difference. The objective function is constructed as the sum of squares of the differences between the distance *di* to all edge points and the radius *r*. This objective function is minimized using gradient descent or the Levenberg-Marquardt algorithm, iteratively solving for the optimal center coordinates *xc* and *yc*, and the radius *r*. The initial center coordinates are set as the arithmetic mean of the x and y coordinates of the edge points, and the initial radius is set as the average distance from the edge points to the initial center. During iteration, the partial derivatives of the objective function with respect to the center coordinates and radius are calculated, and the parameter values ​​are updated based on these partial derivatives. This iteration is repeated until the parameter change is less than 0.01 pixels or the number of iterations reaches 100. For ellipse fitting, a direct least squares ellipse fitting algorithm is used. The general equation of an ellipse is Ax² + Bxy + Cy² + Dx + Ey + F = 0. Substituting the coordinates of each edge point into the equation constructs a system of linear equations, with the constraint that B² - 4AC < 0 ensures the equation represents an ellipse rather than a hyperbola. The optimal values ​​of parameters A to F are obtained by solving the generalized eigenvalue problem, and further converted into the ellipse center coordinates, major axis, minor axis, and rotation angle. The x-coordinate of the ellipse center is equal to 2CD - BE divided by B squared - 4AC, and the y-coordinate is equal to 2AE - BD divided by B squared - 4AC. For example, the fitted ellipse center coordinates are 127.35 and 100.18, which are the preliminary center coordinates. For irregularly shaped spikes, the convex hull method is used to calculate the geometric center. First, the convex hull of the sub-pixel edge point set is calculated to obtain the minimum convex polygon vertex sequence surrounding the edge points. Then, the arithmetic mean of the convex hull vertex coordinates is calculated as the geometric center point. The preliminary center coordinates reflect the geometric symmetry center of the spike body, but do not consider the non-uniformity of color distribution inside the body, providing an initial position for subsequent gray-scale weighted center optimization.

[0067] Step 1063: Using the initial center coordinates as the center, extract a local neighborhood in the rotation slice and calculate the grayscale weight of each pixel in the local neighborhood; perform a weighted summation of the pixel coordinates in the local neighborhood based on the grayscale weights, and use the resulting weighted centroid coordinates as the rotation center coordinates.

[0068] Specifically, firstly, a local neighborhood is extracted from the rotated slice based on the initial center coordinates. The size of the local neighborhood is determined by the minimum bounding rectangle of the main body of the puncture point. The lengths of the long and short sides of the minimum bounding rectangle are calculated, and the average of these two lengths is taken as the feature size. The side length of the local neighborhood is set to 1.8 times the feature size. For example, if the long side of the minimum bounding rectangle is 80 pixels and the short side is 60 pixels, the feature size is 70 pixels, and the side length of the local neighborhood is 126 pixels. Using the initial center coordinates of 127.35 and 100.18 as the center, extend 63 pixels to the left and right, and 63 pixels up and down, to extract a local neighborhood region with horizontal coordinates from 64 to 190 and vertical coordinates from 37 to 163. The local neighborhood is converted from the RGB color space to the HSV color space, and the saturation (S) and brightness (V) values ​​of each pixel are calculated. A grayscale weight is calculated for each pixel, and the weight formula integrates saturation and brightness information. Let the saturation weight ws equal the square of the saturation S, and the brightness weight wv equal the brightness V multiplied by 1 minus the brightness V. This formula maximizes the weight of pixels with medium brightness and reduces the weight of pixels that are too bright or too dark. The total grayscale weight w equals the saturation weight ws multiplied by the brightness weight wv and then multiplied by the normalization coefficient. For example, a pixel with a saturation of 0.85 and a brightness of 0.72 has a saturation weight of 0.7225, a brightness weight of 0.2016, and a total weight of 0.1457. Iterate through all pixels in the local neighborhood, calculate the grayscale weight of each pixel, and store the weights in a weight matrix of the same size as the local neighborhood. Calculate the total weights, which equals the sum of all elements in the weight matrix. Perform a weighted sum of the pixel coordinates in the local neighborhood. The weighted sum of the horizontal coordinates equals the sum of all pixel horizontal coordinates multiplied by their corresponding weights, and the weighted sum of the vertical coordinates equals the sum of all pixel vertical coordinates multiplied by their corresponding weights. The weighted centroid x-coordinate is equal to the weighted sum of the x-coordinates divided by the total weights, and the weighted centroid y-coordinate is equal to the weighted sum of the y-coordinates divided by the total weights. Using the weighted centroid coordinates as the rotation center coordinates takes into account the non-uniformity of color distribution within the puncture site; high-saturation color areas contribute more to center localization, while the influence of background and transition areas is reduced. The rotation center coordinates achieve a localization accuracy of 0.05 pixels, superior to the preliminary center coordinates based on geometric fitting. This provides high-precision input data for subsequent inverse transformation of the center coordinates, ensuring that the final puncture center localization in the original slice meets the accuracy requirements for clinical applications.

[0069] In the above embodiments, automatic extraction and confidence assessment of puncture point numbers were achieved through rotation correction and character recognition. To further ensure the positioning accuracy of the recognition results in the global coordinate system of aerial imagery and support subsequent 3D reconstruction and earthwork measurement tasks based on control points, this application also provides a 3D reconstruction puncture point method for earthwork construction sites. This method restores the coordinate transformation relationship by performing inverse operations on the rotation matrix, and achieves accurate mapping from local coordinates to global coordinates by combining slice clipping offsets. The identified number information, pixel coordinates, and confidence scores are encapsulated into control point observation records according to a standardized data structure, enabling the system to efficiently and accurately convert image recognition results into control point data that can be used for photogrammetric calculations. The following section combines... Figure 2 Another method for three-dimensional reconstruction of puncture points for earthwork construction sites is described in the embodiments of this application: Please see Figure 2 This is a flowchart illustrating another three-dimensional reconstruction puncture point method for earthwork construction sites in this application embodiment.

[0070] Step 201: Invert the rotation matrix to obtain the inverse transformation matrix; multiply the rotation center coordinates with the inverse transformation matrix to obtain the local coordinates of the target resolution slice in the coordinate system.

[0071] Specifically, the inverse operation is first performed on the constructed rotation matrix. The rotation matrix is ​​a 3x3 affine transformation matrix, containing a 2x2 rotation part and a 2x1 translation part, with the third row in standard homogeneous coordinates. Since the rotation part is an orthogonal matrix, its inverse is equal to its transpose. Therefore, the inverse of the rotation part is achieved by swapping the positions of specific matrix elements while keeping the diagonal elements unchanged. Let the rotation part of the original rotation matrix consist of cosine and sine values; the inverse rotation part is obtained by changing the sign of the sine term. The inverse of the translation part needs to be recalculated; the inverse translation vector is equal to the negative inverse rotation matrix multiplied by the original translation vector. For rotation transformations around any center point in the original rotation matrix, the coordinate system is first translated to the rotation center, rotated, and then translated back to the original position. The combined matrix of these three transformations needs to be expanded into the corresponding inverse transformation sequence through matrix inversion. The calculation of the inverse translation component involves matrix multiplication of the elements of the inverse rotation matrix with the original translation component, followed by negation of the result. Assemble the complete inverse transformation matrix. The first two rows and first two columns represent the inverse rotation component, the first two rows and third column represent the inverse translation component, and the third row maintains the standard homogeneous coordinate form. Convert the rotation center coordinates to homogeneous coordinates, and add a third component to the original two-dimensional coordinates to form a three-dimensional vector. Perform matrix-vector multiplication, multiplying the inverse transformation matrix by the homogeneous coordinate vector. The first result component equals the product of the first row and first column element of the inverse rotation matrix multiplied by the rotation center x-coordinate, plus the product of the first row and second column element of the inverse rotation matrix multiplied by the rotation center y-coordinate, plus the first inverse translation component. The second result component equals the product of the second row and first column element of the inverse rotation matrix multiplied by the rotation center x-coordinate, plus the product of the second row and second column element of the inverse rotation matrix multiplied by the rotation center y-coordinate, plus the second inverse translation component. The third result component is always a unit value. Take the first two components as local coordinates, which represent the precise position of the puncture point center in the original coordinate system of the target resolution slice, eliminating the influence of rotation correction and restoring the true position of the puncture point in the cut slice. The accuracy of local coordinates inherits the sub-pixel accuracy of the rotation center coordinates, and the positioning error is controlled at the sub-pixel level, providing high-precision relative position information for subsequent overlay and cropping offsets.

[0072] Step 202: Superimpose the local coordinates with the cropping offset of the target resolution slice in the aerial image to obtain the absolute pixel coordinates.

[0073] Specifically, the cropping offset of the target resolution slice is first read, which was recorded during the slice cropping in step 103. The cropping offset includes horizontal and vertical offsets, representing the column and row indices of the top-left corner of the slice in the aerial image, respectively. A coordinate overlay operation is performed: the x-coordinate of the absolute pixel coordinates equals the x-coordinate of the local coordinates plus the horizontal component of the cropping offset, and the y-coordinate equals the y-coordinate of the local coordinates plus the vertical component of the cropping offset. This coordinate operation is a simple vector addition, mapping the relative coordinates within the slice to the global coordinate system of the aerial image. The absolute pixel coordinates uniquely identify the spatial position of the spike in the aerial image, independent of the slice cropping method and rotation correction process. The validity of the absolute pixel coordinates is verified by checking whether the x-coordinate is within the width range of the aerial image and whether the y-coordinate is within the height range of the aerial image. If the coordinates exceed the aerial image boundary, the control point is marked as a boundary anomaly, a warning message is recorded, but the coordinate data is retained. For high-precision positioning requirements, the absolute pixel coordinates retain an appropriate number of decimal places, achieving sub-pixel accuracy; the corresponding actual ground distance depends on the ground resolution of the aerial image. Absolute pixel coordinates provide the foundation for subsequent geographic coordinate transformation and control point management. By using the georeferenced information of aerial imagery, pixel coordinates are converted into geographic coordinates or actual site coordinates, achieving a global representation of puncture point locations. The output absolute pixel coordinates are stored in association with attribute information such as puncture point number and confidence level, forming complete control point observation data, supporting position comparison and accuracy evaluation of the same puncture point in multiple aerial images.

[0074] Step 203: Extract the image identifier from the aerial image; encapsulate the image identifier, absolute pixel coordinates, target number, and confidence level according to the preset data structure to generate control point observation records.

[0075] Specifically, the first step is to extract the image identifier from the aerial imagery, which is read from the imagery's metadata. Aerial images are typically stored using specific naming rules, and the filename or EXIF ​​information contains metadata such as shooting time, equipment information, and geographical location. The image filename is parsed to extract information such as the timestamp, camera serial number, and task number, which are then concatenated to form a unique image identifier string. If the image filename does not contain complete information, the image's EXIF ​​metadata is read. The shooting time is obtained from the date and time field, camera information from the equipment manufacturer and model field, and location and timestamp from the GPS-related fields, which are then used to generate the image identifier. The data structure for control point observation records is defined, using a structured format to facilitate cross-platform and cross-language data exchange. The data structure includes an image identifier field storing image source information, pixel coordinate horizontal and vertical fields storing the two components of the absolute pixel coordinates, an ID field storing the target ID, a confidence field storing the recognition reliability assessment value, a timestamp field storing the record generation time, and a detection method field indicating the detection algorithm version. Following the data structure definition, control point observation record objects are created, and values ​​are assigned to each data field. The image identifier field is assigned the extracted image identifier string. The pixel coordinate horizontal field is assigned the horizontal coordinate of the absolute pixel coordinates, and the pixel coordinate vertical field is assigned the vertical coordinate of the absolute pixel coordinates. The number field is assigned the target number, and the confidence field is assigned the recognition confidence level. The timestamp field is assigned the current system time, formatted as a standard time string. The detection method field is assigned the algorithm version identifier, used to trace the data generation method. The control point observation record objects are serialized into a structured data format, formatted into an easy-to-read format for easy manual viewing and debugging. The control point observation records are written to a database or file system. If a relational database is used, a new record row is inserted into the control point observation table, with each field corresponding to a column in the table. If a document-oriented database is used, the structured objects are directly stored as a document. If file storage is used, the records are appended to the control point observation log file, one record per line. The generated control point observation records provide complete data support for subsequent data analysis, quality assessment and control point management. They support querying all control points in a specific aerial image by image identifier, querying multiple observations of the same control point in different images by number, and filtering high-quality observation data by confidence level, thus achieving efficient management and utilization of control point data.

[0076] In one possible implementation, the image identifier, absolute pixel coordinates, target number, and confidence level are encapsulated according to a preset data structure to generate a control point observation record. Following this step, steps 2031-2033 are further included, as follows: Step 2031: Match the absolute pixel coordinates in the control point observation records with the pre-acquired measured 3D coordinates to generate a control point constraint file.

[0077] Specifically, the pre-acquired measured 3D coordinate data of the puncture points are read. Each record includes the puncture point number, eastward coordinates, northward coordinates, elevation coordinates, and measurement accuracy. The control point observation records are traversed, the target number is extracted, and records with the same number are searched in the measured data. After finding a matching record, the 3D coordinate values ​​are extracted, and the correspondence between the pixel coordinates and 3D coordinates of the puncture point is established. For multiple observations of the same puncture point in multiple images, the pixel coordinates of each observation are matched with the same set of 3D coordinates. The 3D coordinates obtained by back-projecting the pixel coordinates of the same puncture point observed in different images using the collinearity equation are checked and compared with the measured coordinates. If the deviation exceeds a threshold, it is marked as an abnormal match. Observation records without matching measured coordinates are classified as unconstrained observation points, used only for relative orientation between images. Weighting coefficients are assigned based on the spatial distribution and observation quality of the control points. Control points located in the center of the site with multiple high-confidence observations are given higher weights, while control points at the edge or with low-confidence observations are given lower weights. Generate control point constraint files according to the photogrammetry software format. The file includes header metadata such as coordinate system, units, and precision. The data section records the control point number, 3D coordinate components, image identifier, pixel coordinates, and weighting coefficients line by line. Multiple observations of the same control point occupy multiple lines. Save as text or binary format, with the file name including the task identifier and generation time.

[0078] Step 2032: Extract corresponding feature points from the aerial image, combine the corresponding feature points with the control point constraint file to perform error minimization adjustment calculation, restore the camera pose parameters and generate a sparse point cloud; perform dense matching based on the camera pose parameters and the sparse point cloud to generate a three-dimensional elevation grid.

[0079] Specifically, scale-invariant feature transformation or accelerated robust feature algorithms are used to detect keypoints in the images, searching for local extrema at different scales. For each keypoint, a feature descriptor is calculated, encoding the gradient direction and magnitude distribution of its neighborhood. For each pair of overlapping images, matching is performed using descriptor distance, and ratio testing and bidirectional consistency checks are applied to screen reliable matching pairs. A random sampling consensus algorithm is used to estimate the fundamental matrix, eliminating outliers that do not meet geometric constraints. The corresponding feature points are combined with control point constraint files to construct a bundle adjustment observation equation set. The observation equations are based on collinearity equations, describing the geometric relationship between the 3D coordinates of ground features, camera pose parameters, and pixel coordinates. Unknown parameters include the pose parameters of all cameras, intrinsic parameters, and the 3D coordinates of non-control points. The Levenberg-Marquardt algorithm is used iteratively to solve the equations, minimizing the sum of squared reprojection errors. An incremental structure restoration method is used for initialization, starting with the relative orientation of two images, gradually adding new images and optimizing global parameters. Control point constraints ensure the 3D structure is aligned with the ground coordinate system. The iteration continues until the parameter changes are less than a threshold, outputting the camera pose parameters and sparse point cloud. Based on the recovered camera pose, dense stereo matching is performed. For each pixel or sampling point in the image, a corresponding point is searched in the overlapping images. A semi-global matching algorithm is used to minimize the combined energy of matching cost and smoothing constraints, and the disparity value is calculated. 3D coordinates are calculated using triangulation formulas to generate a dense point cloud. The dense point cloud is filtered to remove noise and outliers. The point cloud is projected onto a horizontal plane, divided into a regular grid, and the elevation statistics of points within each grid cell are statistically analyzed and assigned to that cell as the ground elevation. Hollow grids are filled using bilinear interpolation or kriging interpolation. The output is a regular matrix 3D elevation grid.

[0080] Step 2033: Compare the three-dimensional elevation grid with the pre-acquired benchmark elevation grid, extract the elevation difference at the corresponding coordinate positions, and calculate the earthwork volume of the site to be measured based on the elevation difference.

[0081] Specifically, the process involves reading the baseline elevation grid data and verifying its coordinate system consistency with the 3D elevation grid. If inconsistent, a coordinate transformation is performed. The grid resolution consistency is checked; if different, the lower-resolution grid is interpolated and densified, or the higher-resolution grid is downsampled. The overlapping area of ​​the two grids is identified, and comparison is performed only within this area. Each grid cell within the overlapping area is traversed, and the elevation values ​​of the two grids at that location are extracted. The elevation difference is calculated as the 3D elevation grid elevation minus the baseline elevation grid elevation. The difference is stored in a difference matrix, with positive values ​​marked as fill and negative values ​​as cut. Abnormally large differences undergo quality checks; those exceeding reasonable ranges are marked as suspicious data and processed using neighborhood interpolation or removal. The earthwork volume is calculated based on the difference matrix. The earthwork volume of each grid cell is equal to the elevation difference multiplied by the horizontal area of ​​the grid cell, and the grid area is equal to the square of the grid spacing. The total earthwork volume change is obtained by summing the earthwork volumes of all grid cells. The fill volume is calculated as the sum of the earthwork volumes corresponding to all positive differences, and the cut volume is calculated as the sum of the earthwork volumes corresponding to the absolute values ​​of all negative differences. The net earthwork volume is calculated as the fill volume minus the cut volume. For irregular boundary sites, a boundary mask is used to mark the effective area. Considering the soil looseness coefficient and compaction coefficient, the geometric volume is adjusted to the actual project volume according to the soil type and construction technology. An earthwork calculation report is generated, including fill volume, cut volume, net earthwork volume, site area, and average fill and cut depth. An elevation difference visualization image is output, using pseudo-color mapping, with warm colors displayed for fill areas and cool colors for cut areas.

[0082] In a preferred embodiment, the specific process of identifying the puncture dot number includes: First, numbered regions are cropped using a dual-track strategy of segmentation guidance and rule-based fallback. The quality of the segmentation mask in the rotated ROI's `number_text` is used as the criterion. If the number of foreground pixels and their ratio to the main body area both meet preset thresholds, the mask quality is deemed acceptable. The foreground pixel bounding boxes are directly extracted and expanded to four sides by preset pixels as the cropping range for the numbered sub-image. Otherwise, rule-based positioning is switched to determine the coordinates of the numbered region based on the preset relative positional relationship of the minimum bounding rectangle of the rotated main body mask, and the numbered sub-image is cropped.

[0083] The numbered sub-images are preprocessed sequentially with sharpening, contrast enhancement, and binarization. Sharpening uses an unsharpened mask method, employing the Gaussian blur result as a mask to superimpose the difference signal onto the original image with preset coefficients, enhancing character edges. Contrast enhancement uses contrast-limited adaptive histogram equalization, performing histogram-limited equalization on local image blocks and eliminating block boundary artifacts through bilinear interpolation. Binarization uses an adaptive thresholding method, subtracting a preset offset from the neighborhood mean as the decision threshold for each pixel, generating a binary image with characters as the foreground and a white background, and applying morphological processing to remove minor noise.

[0084] Character segmentation and OCR recognition are then performed. A vertical projection histogram is constructed by counting the number of foreground pixels in each column of the binary image. The character gaps are defined using a preset percentage of the peak value as a threshold. The bounding boxes of each character are extracted from left to right and scaled to a standard size. For each individual character, an orienting gradient histogram feature vector is extracted, and its similarity is calculated one by one with a pre-stored standard character template. The top-k candidate characters and their confidence scores are retained in descending order of similarity, forming a candidate list for that position. From the top-k candidates at each character position, combinations are enumerated. The lowest candidate similarity for each character is used as the overall confidence score of the candidate string. Candidates below the confidence threshold are filtered out, and duplicates are removed. The resulting candidate list for the image's puncture point and its corresponding confidence score is output.

[0085] After all image processing is completed, execute multiple... Figure 1 Consistent voting. Multiple observations with adjacent ground coordinates are grouped into the same observation group. A voting pool is constructed by aggregating the candidate list of image numbers within each group. Each candidate string is weighted and voted on based on its overall confidence level. The candidate with the highest score is selected as the final target number. The final confidence level is equal to the highest vote score divided by the number of images in the group. When the final confidence level meets a preset threshold, the number and confidence level are written into the control point observation record; otherwise, the labeling requires manual verification, and a Top-k candidate list is provided for operator reference. For puncture points observed only in a single image, the Top-1 candidate number is directly used, and a reduction factor is applied to the confidence level before writing it into the record to reflect the difference in reliability between single-image identification and multi-image voting.

[0086] The following describes a 3D reconstruction puncture system for earthwork construction sites from a hardware processing perspective, according to an embodiment of this invention. Please refer to [link to relevant documentation]. Figure 3 This is a schematic diagram of the structure of a three-dimensional reconstruction puncture system for earthwork construction sites in an embodiment of this application.

[0087] It should be noted that, Figure 3 The structure of the three-dimensional reconstruction puncture system for earthwork construction sites shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0088] like Figure 3As shown, a 3D reconstruction sniping system for earthwork construction sites includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes based on a program stored in a Read-Only Memory (ROM) 302 or a program loaded from a storage section 308 into a Random Access Memory (RAM) 303, such as the method described in the above embodiment. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.

[0089] The following components are connected to I / O interface 305: input section 306 including audio input devices, push-button switches, etc.; output section 307 including a liquid crystal display (LCD) and audio output devices, indicator lights, etc.; storage section 308 including a hard disk, etc.; and communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.

[0090] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the various functions defined in the present invention.

[0091] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0092] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.

[0093] Specifically, a three-dimensional reconstruction puncture system for earthwork construction sites according to this embodiment includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, it implements a three-dimensional reconstruction puncture method for earthwork construction sites provided in the above embodiment.

[0094] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the three-dimensional reconstruction puncturing system for earthwork construction sites described in the above embodiments; or it may exist independently and not assembled into the three-dimensional reconstruction puncturing system for earthwork construction sites. The storage medium carries one or more computer programs, which, when executed by a processor of the three-dimensional reconstruction puncturing system for earthwork construction sites, cause the system to implement the three-dimensional reconstruction puncturing method for earthwork construction sites based on IoT data encryption transmission provided in the above embodiments.

Claims

1. A three-dimensional reconstruction method for puncture points in earthwork construction sites, characterized in that, The method includes: Acquire aerial images of the site to be tested; The aerial image is downsampled to generate a thumbnail image, the scaling factor is recorded, and color features are extracted from the thumbnail image to obtain a candidate mask. The candidate mask is subjected to connected component analysis to filter out the target candidate region, and the target candidate region is mapped to the aerial image according to the scaling factor to extract the target resolution slice. Purity analysis is performed on the target resolution slice to separate the main mask and the number mask. A rotation matrix is ​​calculated based on the main mask. The rotation matrix is ​​then used to perform affine transformations on the target resolution slice, the main mask, and the number mask to obtain a rotated slice, a rotated main mask, and a rotated number mask, respectively. Based on the sort number mask, numbered sub-images are extracted from the sort slice, and character recognition is performed on the numbered sub-images to obtain the target number and confidence level; Edge detection and fitting are performed on the rotated slice based on the rotated main mask, the coordinates of the rotated center are calculated, and the coordinates of the rotated center are inversely transformed to the coordinate system of the aerial image using the rotation matrix to obtain the absolute pixel coordinates. The absolute pixel coordinates, the target number and the confidence level are combined to generate control point observation records.

2. The method according to claim 1, characterized in that, The step of performing connected component analysis on the candidate mask to filter out target candidate regions, mapping the target candidate regions to the aerial image according to the scaling factor, and extracting target resolution slices includes: Identify multiple connected regions in the candidate mask, and calculate the geometric feature parameters of each connected region, including pixel area, aspect ratio of the bounding rectangle, and circularity. Based on the geometric feature parameters, a weighted calculation is performed to determine the target score of each connected region. Connected regions with target scores greater than a preset score threshold are selected as target candidate regions, and the boundary coordinates of the target candidate regions in the thumbnail image are obtained. The boundary coordinates are compensated using the scaling factor to obtain the mapped boundary coordinates; The cropping range is determined in the aerial image based on the mapped boundary coordinates, and a target resolution slice is extracted based on the cropping range.

3. The method according to claim 1, characterized in that, The process involves performing purity analysis on the target resolution slice to separate the main mask and the numbered mask, calculating a rotation matrix based on the main mask, and then performing an affine transformation on the target resolution slice, the main mask, and the numbered mask using the rotation matrix to obtain a rectified slice, a rectified main mask, and a rectified numbered mask, respectively. Calculate the color purity index of each pixel in the target resolution slice; Pixels with a color purity index greater than or equal to the purity threshold are divided into main regions and a main mask is generated. Pixels with a color purity index less than the purity threshold and a brightness greater than the brightness threshold are divided into numbered regions and a numbered mask is generated. Extract the outer contour of the main body mask, calculate the minimum bounding rectangle of the outer contour, and obtain the center point coordinates and tilt angle of the minimum bounding rectangle; A rotation matrix is ​​constructed using the coordinates of the center point as the rotation center and the negative value of the tilt angle as the rotation amount; The target resolution slice, the main body mask, and the numbered mask are resampled using the rotation matrix to generate the rotated slice, the rotated main body mask, and the rotated numbered mask.

4. The method according to claim 1, characterized in that, The step of extracting numbered sub-images from the normalized slice based on the normalized numbering mask, and performing character recognition on the numbered sub-images to obtain the target number and confidence level, includes: Determine the bounding box of the foreground pixels in the swivel number mask; Numbered sub-images are determined from the swirl slices based on the bounding box; The numbered sub-images are converted to grayscale and binarized to obtain a binarized image; The binarized image is segmented into characters to obtain a sequence of single-character images; Extract the feature vector of each character image in the single character image sequence, and perform similarity matching between the feature vector and a pre-stored standard character template library to determine the matching character and matching similarity for each character image; The target number is generated by concatenating the matching characters corresponding to each character image in sequence, and the minimum value is selected from the matching similarity values ​​corresponding to each character image as the confidence score.

5. The method according to claim 1, characterized in that, The step of performing edge detection and fitting on the normalized slice based on the normalized main mask, and calculating the normalization center coordinates, includes: In the slewd slice, extract the set of pixels within the area covered by the slewd main mask; An edge detection algorithm is used to extract the edge pixels of the pixel set, and an interpolation algorithm is used to interpolate the positions of the edge pixels to obtain a sub-pixel edge point set. Geometric contour fitting is performed on the sub-pixel edge point set, and the geometric center point of the fitted contour is extracted as the preliminary center coordinates. Using the initial center coordinates as the center, a local neighborhood is extracted from the rotated slice, and the grayscale weight of each pixel in the local neighborhood is calculated. The pixel coordinates in the local neighborhood are weighted and summed based on the grayscale weights, and the resulting weighted centroid coordinates are used as the rotation center coordinates.

6. The method according to claim 1, characterized in that, The process of inversely transforming the rotation center coordinates to the coordinate system of the aerial image using the rotation matrix to obtain absolute pixel coordinates, and combining the absolute pixel coordinates, the target number, and the confidence level to generate a control point observation record includes: Invert the rotation matrix to obtain the inverse transformation matrix; Multiplying the positive rotation center coordinates by the inverse transformation matrix yields the local coordinates of the target resolution slice in the coordinate system. The local coordinates are superimposed with the cropping offset of the target resolution slice in the aerial image to obtain the absolute pixel coordinates; Extract the image identifiers from the aerial images; The image identifier, the absolute pixel coordinates, the target number, and the confidence level are encapsulated according to a preset data structure to generate the control point observation record.

7. The method according to claim 6, characterized in that, After encapsulating the image identifier, the absolute pixel coordinates, the target number, and the confidence score according to a preset data structure to generate the control point observation record, the method further includes: The absolute pixel coordinates in the control point observation records are matched with the pre-acquired measured three-dimensional coordinates to generate a control point constraint file; Extract corresponding feature points from the aerial image, combine the corresponding feature points with the control point constraint file to perform error minimization adjustment calculation, restore the camera pose parameters and generate sparse point cloud; Dense matching is performed based on the camera pose parameters and the sparse point cloud to generate a three-dimensional elevation grid. The three-dimensional elevation grid is compared with the pre-acquired benchmark elevation grid to extract the elevation difference at the corresponding coordinate positions, and the earthwork volume of the site to be measured is calculated based on the elevation difference.

8. A three-dimensional reconstruction puncture system for earthwork construction sites, characterized in that, The three-dimensional reconstruction piercing system for earthwork sites includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the three-dimensional reconstruction piercing system for earthwork sites to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium comprising instructions, characterized in that, When the instruction is executed on a three-dimensional reconstruction puncture system facing an earthwork site, the three-dimensional reconstruction puncture system facing an earthwork site performs the method as described in any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product is run on a three-dimensional reconstruction puncture system for earthwork sites, the three-dimensional reconstruction puncture system for earthwork sites performs the method as described in any one of claims 1-7.