GIS data automatic regular and warehousing system based on multi-format file OCR recognition
Patent Information
- Application Number
- CN202611224876.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-13
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本发明的目的在于提供一种基于多格式文件OCR识别的GIS数据自动化规整与入库系统,解决以下技术问题:现有空间数据在进行规整与入库时,存在对文件扩展名或人工判读依赖性高、文件类型识别易出错、异构数据处理流程割裂且人工频繁切换、自动化程度较低,以及在批量入库时缺乏动态调度、修复拦截与空间拓扑及图文属性关联校验机制,容易导致拓扑错误或数据失准等问题,本发明提供一种能够实现多源异构多格式空间数据在物理层特征的精准识别、流程自适应动态调度修复、多格式数据智能化深度规整挂接以及严格拓扑一致性校验入库的GIS数据自动化规整与入库系统
[0036]1.本系统深入文件内部,通过提取文件头标识、空间组织信息以及重塑的灰度图块信息,并结合卷积神经网络模型进行智能识别;这种从物理层和视觉层构建多维特征的方式,有效解决了传统系统因文件扩展名缺失、被篡改或文件头局部损坏而导致读取报错、无法识别的问题,保障了数据批量入库流程的连续性;
Smart Images

Figure CN122817352A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of geographic information systems and intelligent spatial data processing technology, specifically to an automated GIS data organization and storage system based on OCR recognition of multi-format files. Background Technology
[0002] As an important infrastructure for land surveys, planning management, surveying and mapping updates and digitization of historical archives, the spatial database of geographic information systems has seen a continuous increase in data sources and increasingly complex data formats. In order to ensure the unified management and efficient use of multi-source spatial data, it is usually necessary to standardize and process heterogeneous data such as scanned map images, CAD vector files and mixed text and graphics documents, and convert them into standard spatial data that can be directly stored in the database.
[0003] However, when organizing and storing multi-format spatial data, it is necessary to accurately determine the file type and complete the corresponding processing flow. File type determination is a key link affecting the quality of subsequent denoising correction, vector analysis, image-text association, and topology storage, and it directly relates to the continuity and effectiveness of spatial database construction. Existing processing methods usually rely on file extensions, manual interpretation, or single file header information to identify data formats and process them according to fixed procedures. This method is prone to identification errors or reading failures when the extension is missing, tampered with, or the file header is partially damaged. At the same time, different types of spatial data often require different tools to complete image processing, coordinate transformation, text extraction, and attribute attachment, resulting in fragmented processing flows, multiple manual interventions, and a lack of automated processing. Especially in the scenario of batch storage of mixed data packets, if there is a lack of dynamic distribution and anomaly interception mechanisms based on identification confidence, it is easy for disguised files to be mistakenly entered into the processing chain, inaccurate image-text attribute association, and topology error data to be directly stored, thereby affecting the integrity of spatial data and the correctness of geographic logic. Summary of the Invention
[0004] The purpose of this invention is to provide an automated GIS data straightening and storage system based on OCR recognition of multi-format files, addressing the following technical problems: Existing spatial data straightening and storage suffers from high reliance on file extensions or manual interpretation, errors in file type identification, fragmented heterogeneous data processing workflows requiring frequent manual switching, low automation, and a lack of dynamic scheduling, repair interception, and spatial topology and image-text attribute correlation verification mechanisms during batch storage, easily leading to topology errors or data inaccuracies. This invention provides an automated GIS data straightening and storage system capable of accurately identifying physical layer features of multi-source heterogeneous multi-format spatial data, adaptive dynamic scheduling and repair of the workflow, intelligent deep straightening and integration of multi-format data, and strict topology consistency verification for storage. The objective of this invention can be achieved through the following technical solutions:
[0005] The GIS data automated straightening and storage system based on multi-format file OCR recognition includes a data format recognition module, a preprocessing and distribution module, a multi-format data processing module, and a data verification and storage module. The data format recognition module reads spatial data files as file data streams, extracts the file header identifier, spatial organization information, and grayscale tile information corresponding to the preset length binary stream at the beginning of the file data stream, and inputs them into a convolutional neural network recognition model to obtain the highest classification confidence and format type.
[0006] The preprocessing distribution module is configured with a high confidence threshold and a low confidence threshold; when the highest classification confidence is greater than or equal to the high confidence threshold, the file data stream is connected to the regular preprocessing process to generate regular routing instructions; when the highest classification confidence is greater than or equal to the low confidence threshold and less than the high confidence threshold, it is switched to the supplementary preprocessing process to generate supplementary routing instructions; when the highest classification confidence is less than the low confidence threshold, an interception instruction is generated.
[0007] The multi-format data processing module responds to the regular routing instruction or the supplementary routing instruction and routes according to the format type: for scanned map images, it calls the image denoising and correction unit; for CAD vector files, it calls the vector parsing and normalization unit; for PDF documents, it calls the hybrid document element extraction unit. It then normalizes the file data stream, calls the OCR recognition engine to extract text annotation information and links it with the normalized spatial geometric data attributes to output standard spatial data.
[0008] The data verification and database entry module performs a topology consistency check on the standard spatial data. If the check passes, the data is converted to a standard format and written to the spatial database. If the check fails, an exception record is output.
[0009] Optionally, when extracting the spatial organization information, the data format recognition module searches whether the file data stream contains a preset GIS identifier byte, which includes layer, projection, or reference plane identifier information;
[0010] The convolutional neural network recognition model is trained based on a pre-built sample database containing files in various known GIS formats.
[0011] Optionally, the high confidence threshold and low confidence threshold in the preprocessing distribution module are set based on the historical error rate statistics of spatial data classification;
[0012] When the file data stream is transferred to the supplementary preprocessing process, the system triggers a program to perform integrity verification and repair on the file header identifier and data blocks of the file data stream, and then hands it over to the multi-format data processing module for processing.
[0013] Optionally, when the format type is scanned map image, the multi-format data processing module calls the image denoising and correction unit to perform noise assessment, specifically:
[0014] The scanned map image is divided into multiple local blocks, the local pixel variance of each local block is calculated, and the global pixel variance of the scanned map image is calculated.
[0015] The following logical judgment is executed, and the judgment condition is as follows: calculate the ratio of the global pixel variance to the mean local pixel variance of all the local blocks to generate a noise intensity value; dynamically determine the window size of the denoising filter based on the noise intensity value, wherein the window size is equal to the floor value of the product of the scaling factor and the noise intensity value plus three, and when the calculated window size is an even number, add one to constrain it to an odd dimension; wherein the scaling factor is a preset weight constant.
[0016] Optionally, the image denoising and correction unit further performs tilt detection and synchronous correction, specifically:
[0017] The Hough transform is used to detect the inner and outer contour lines of the scanned map image, and the tilt angle of the inner and outer contour lines relative to the horizontal reference is calculated.
[0018] An affine transformation matrix containing the tilt angle is constructed. During the process of spatial filtering and denoising the pixels traversed by the scanned map image, the coordinates of the target pixel are mapped to the original image through the inverse matrix of the affine transformation matrix to find neighboring pixels.
[0019] The grayscale values of the neighboring pixels are calculated using a filter with a window size equal to the window size. Noise reduction and tilt correction are completed simultaneously through one bilinear interpolation, and a normalized orthophoto image is output. The OCR recognition engine is then called to extract the image annotation text.
[0020] Optionally, when the format type is a CAD vector file, the multi-format data processing module calls the vector parsing and normalization unit to perform layer parsing and coordinate system inference, specifically:
[0021] The internal entity object tree of the CAD vector file is parsed to extract the source layer name, and the elements are classified using a pre-configured layer name mapping table that contains the correspondence between the source layer name and the standard layer name.
[0022] Extract the spatial bounding box coordinate extreme values of the CAD vector file. When the CAD vector file lacks a coordinate reference system definition, infer the source projection zone and source coordinate system based on the range of the spatial bounding box coordinate extreme values. The specific logic of the inference is as follows: extract the total length of the integer digits of the coordinate maxima in the spatial bounding box coordinate extreme values and the characteristic digits of the preset number of digits at the beginning. Input the total length of the integer digits and the characteristic digits into a pre-configured coordinate system rule discrimination tree for traversal and comparison. If a preset national standard or industry standard coordinate system numerical range rule is matched, the corresponding source projection zone and source coordinate system are output. If the numerical range rule is not matched, a coordinate system unknown anomaly mark is output and the coordinate system inference process of the current file is terminated.
[0023] Optionally, after inferring the source projection zone and source coordinate system, the vector analysis and normalization unit establishes a coordinate transformation model:
[0024] Based on the inferred source coordinate system and the system's preset standard target coordinate system, a coordinate transformation matrix is constructed;
[0025] The vertex coordinates of all extracted spatial entities are batch-transformed based on the coordinate transformation matrix to output standard vector data under a unified geographic framework.
[0026] Optionally, when the format type is a PDF document, the multi-format data processing module calls the hybrid document feature extraction unit to perform synchronous feature extraction, specifically as follows:
[0027] The underlying page content flow of the PDF document is parsed, a document object tree is constructed, and vector graphics operation instructions and text drawing instructions are separated.
[0028] The vector graphics operation instructions are restored to GIS standard geometric objects, and the coordinates of the geometric center point of the GIS standard geometric objects in the coordinate system of the PDF document page are extracted;
[0029] The text stream content and its bounding box center coordinates are extracted synchronously using the text drawing instructions.
[0030] Optionally, after extracting the coordinates, the hybrid document element extraction unit performs spatial location association, specifically as follows:
[0031] The following logical judgment is performed, and the judgment condition is: calculate the Euclidean distance between the coordinates of the bounding box center and the coordinates of the geometric center point; when the Euclidean distance is less than or equal to the association distance threshold, determine that the corresponding text stream content is the spatial attribute annotation of the corresponding GIS standard geometric object; when the Euclidean distance is greater than the association distance threshold, determine that the corresponding text stream content is independent text.
[0032] When determined to be a spatial attribute annotation, the text stream content is written as an attribute field into the attribute table of the GIS standard geometric object; wherein, the association distance threshold is calculated by multiplying the DPI parameter parsed from the PDF document by a preset tolerance physical size in inches.
[0033] Optionally, the topology consistency check performed by the data verification and storage module includes polygon self-intersection detection and dangling node detection;
[0034] The standard format includes GeoJSON or WKT format.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1. This system delves deep into the file, extracting file header identifiers, spatial organization information, and reconstructed grayscale patch information, and combining this with a convolutional neural network model for intelligent recognition. This approach of constructing multi-dimensional features from both physical and visual layers effectively solves the problems of traditional systems encountering reading errors and being unable to recognize files due to missing file extensions, tampering, or partial damage to the file header, thus ensuring the continuity of the batch data entry process.
[0037] 2. This system sets high and low confidence thresholds based on historical error rate statistics, and dynamically distributes files based on the highest confidence level of the identified categories. Files with high confidence directly enter the regular process; files with medium confidence are automatically triggered to perform supplementary preprocessing for integrity verification and damage repair, avoiding secondary manual intervention caused by direct interception; and files with low confidence are intercepted, thus constructing a deeply coupled intelligent distribution and defense mechanism.
[0038] 3. This system can automatically recognize and accurately route various formats such as scanned map images, CAD vector files, and PDF documents. For CAD files, it can automatically classify features through a layer name mapping table, and automatically infer the source projection zone and source coordinate system based on the range characteristics of the spatial bounding box coordinate extreme values. By constructing a coordinate transformation matrix, it can batch transform vertex coordinates and output standard vector data under a unified geographic framework, thus solving the tedious process of frequently switching tools manually.
[0039] 4. For scanned images, the system can adaptively adjust the denoising filter window based on the pixel variance ratio, and simultaneously complete denoising and tilt correction during the filtering process through inverse matrix mapping, avoiding blurring caused by secondary resampling and improving the accuracy of subsequent OCR text recognition; for PDF documents, the system parses the underlying content stream and calculates the Euclidean distance between the geometric center and the center of the text bounding box, automatically realizes the structured binding of image and text attributes based on the correlation distance threshold, and finally ensures the integrity and geographic logic correctness of the spatial data in the database through strict topological consistency checks. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0041] Figure 1 This is a schematic diagram of the process of the GIS data automated organization and storage system based on multi-format file OCR recognition of the present invention. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0043] like Figure 1 As shown, the GIS data automated straightening and storage system based on multi-format file OCR recognition includes a data format recognition module, a preprocessing and distribution module, a multi-format data processing module, and a data verification and storage module. The data format recognition module reads spatial data files as file data streams, extracts the file header identifier, spatial organization information, and grayscale tile information corresponding to the preset length binary stream at the beginning of the file data stream, and inputs them into a convolutional neural network recognition model to obtain the highest classification confidence and format type.
[0044] The preprocessing and distribution module configures high and low confidence thresholds; when the highest category confidence is greater than or equal to the high confidence threshold, the file data stream is connected to the regular preprocessing process to generate regular routing instructions; when the highest category confidence is greater than or equal to the low confidence threshold and less than the high confidence threshold, it is switched to the supplementary preprocessing process to generate supplementary routing instructions; when the highest category confidence is less than the low confidence threshold, an interception instruction is generated.
[0045] The multi-format data processing module responds to regular or supplementary routing instructions and routes according to format type: scanned map images call the image denoising and correction unit, CAD vector files call the vector parsing and normalization unit, and PDF documents call the hybrid document element extraction unit; it normalizes the file data stream, calls the OCR recognition engine to extract text annotation information and links it with the normalized spatial geometric data attributes, and outputs standard spatial data;
[0046] The data verification and database entry module performs a topology consistency check on standard spatial data. If the check passes, the data is converted to a standard format and written to the spatial database. If the check fails, an exception record is output.
[0047] Physical layer features are extracted from received multi-source spatial data files, and the features are input into a convolutional neural network recognition model to generate the highest classification confidence and format type. In the construction and updating of GIS spatial databases, data is collected from a wide range of sources. Traditional systems rely on file extensions or manual interpretation to determine file formats. When the extension is missing, modified, or the file header is damaged, reading errors will occur. By extracting file header identifiers and spatial organization information, and reshaping the binary stream at the beginning of the file into grayscale tiles, specifically, the one-dimensional binary file data stream is truncated to a preset fixed width and folded line by line into a two-dimensional matrix. The hexadecimal value of each byte is directly mapped to the pixel grayscale value, thereby generating a two-dimensional grayscale tile.
[0048] The system constructs multi-dimensional features from the binary physical layer and the visual layer; based on the highest classification confidence and preset threshold logic, the system dynamically schedules the file data stream and directly couples the format judgment result with the preprocessing pipeline; when the confidence meets different conditions, the system automatically selects regular preprocessing, triggers a supplementary preprocessing repair mechanism, or intercepts manual review.
[0049] The multi-format data processing module routes the data stream to the corresponding dedicated processing unit for normalization processing based on the identified format type, and extracts text annotations through the OCR recognition engine to achieve the connection between spatial geometric data and attribute text; it performs topology consistency checks on standard spatial data and completes the data entry; when operators import mixed data packages containing surveying CAD, historical scanned paper maps and planning description PDFs in batches, the system automatically identifies the file formats and dynamically distributes them to the corresponding preprocessing pipelines to complete noise reduction, parsing, image and text extraction and topology verification. Data that passes the verification is directly written to the underlying GIS spatial database.
[0050] This process reduces reliance on explicit file extensions, effectively addresses files with abnormal extensions, solves the problems of fragmented multi-tool processing and frequent manual switching in traditional GIS data processing, and achieves automated data organization and storage.
[0051] In this embodiment, when extracting spatial organization information, the data format recognition module searches whether the file data stream contains a preset GIS identifier byte. The GIS identifier byte contains layer, projection, or reference plane identifier information.
[0052] The convolutional neural network recognition model is trained based on a pre-built sample database containing files in various known GIS formats.
[0053] The low-level semantic retrieval of file data streams is used to extract spatial organization information, and a convolutional neural network recognition model is trained based on a sample database containing various known GIS file formats. In the construction and updating of conventional GIS spatial databases, file reading tools usually rely on file extensions to determine the format. When encountering files with missing extensions or corrupted file headers, they will directly report an error.
[0054] To overcome the limitations of explicit format recognition, the system delves into the file's binary stream and performs format determination at the physical layer. Specifically, the system searches the file's binary stream for predefined GIS identifier bytes, such as whether it contains identifier byte segments representing layers, projections, or reference planes. The extracted spatial structure semantic features, binary header file feature identifier numbers, and reconstructed grayscale image visual features are then concatenated. Before concatenation, the spatial structure semantic features and binary header file feature identifier numbers are numerically encoded and normalized, and the grayscale image visual features are flattened into a one-dimensional vector. The three are then concatenated in the same dimension to construct a multi-format file feature fingerprint database as a feature vector.
[0055] The feature vector is input into a convolutional neural network model pre-trained based on real data from a historical GIS archive. The model outputs probability distribution vectors for various spatial data formats. The convolutional neural network recognition model does not replace rule matching based on file header signatures. Instead, when the file header identifier cannot be matched due to local damage, it uses spatial organization information and the byte distribution statistics of grayscale tiles to perform redundant and fault-tolerant judgment on the format type, thereby ensuring the continuity of format determination in the case of damaged file headers. When operators import historical spatial data from a batch of historical archives, some files have lost their file extensions and have damaged bytes in the header. The system does not rely on external attributes, directly reads the internal binary stream, successfully retrieves the layer and projection identifier byte segments, and combines the grayscale tile features reconstructed from the preceding bytes, the convolutional neural network determines that the file is a standard graphic file format.
[0056] This process effectively addresses files with abnormal extensions, improves format recognition accuracy and the system's anti-interference capabilities, and ensures the continuity of the batch data entry process.
[0057] The high and low confidence thresholds in the preprocessing and distribution module are set based on the historical error rate statistics of spatial data classification;
[0058] When the file data stream is transferred to the supplementary preprocessing process, the system triggers a program to verify and repair the integrity of the file header identifier and data blocks of the file data stream, and then hands it over to the multi-format data processing module for processing.
[0059] Based on historical error rate statistics for spatial data classification, high and low confidence thresholds are set, and integrity checks and repairs are performed when supplementary preprocessing is triggered. To achieve unattended dynamic scheduling throughout the entire process, the system deeply couples format recognition confidence with the preprocessing pipeline. By setting reasonable threshold ranges, the system can dynamically grade and evaluate file quality. The high and low confidence thresholds are set based on historical error rate statistics for spatial data classification in this field; for example, the high confidence threshold is set to 0.85, and the low confidence threshold is set to 0.60. When the highest classification confidence output by the convolutional neural network is greater than or equal to the low confidence threshold but less than the high confidence threshold, it indicates that the file format features are basically matched but there is local damage or feature ambiguity.
[0060] At this point, the system automatically transfers the file data stream into the supplementary preprocessing process, triggering a preset secondary automatic file integrity detection and repair mechanism to perform targeted repair and completion of damaged file header identifiers and data blocks. The specific technical means of targeted repair and completion are as follows: extract the undamaged valid feature segments in the file data stream, search for the standard template with the highest matching degree in the pre-built standard GIS format file header template library, and use the corresponding standard bytes of the standard template to overwrite and replace the damaged file header identifier. At the same time, fill the missing data blocks with preset security placeholder bytes to complete the data structure.
[0061] After the repair is completed, the data stream is then processed by the multi-format data processing module. When the system processes a cross-project aggregated DXF file, the highest classification confidence of the model output is 0.72 because its file header has been modified by non-standard software. The system determines that the value is in the range of 0.60 to 0.85 and automatically triggers a supplementary preprocessing process to repair and complete the damaged file header identifier.
[0062] The repaired data stream was transferred to the vector analysis and normalization unit, avoiding secondary manual intervention caused by direct interception;
[0063] In this embodiment, when the format type is scanned map image, the multi-format data processing module calls the image denoising and correction unit to perform noise assessment, specifically:
[0064] The scanned map image is divided into multiple local blocks, the local pixel variance of each local block is calculated, and the global pixel variance of the scanned map image is calculated.
[0065] The following logic is executed, with the following conditions: calculate the ratio of the global pixel variance to the mean of the local pixel variances of all local blocks to generate a noise intensity value; dynamically determine the window size of the denoising filter based on the noise intensity value. The window size is equal to the floor value of the product of the scaling factor and the noise intensity value plus three. When the calculated window size is even, add one to constrain it to an odd dimension. The scaling factor is a preset weight constant.
[0066] The system divides the scanned map image into blocks to calculate the local and global pixel variances, and dynamically determines the window size of the denoising filter based on the variance ratio. Scanned map images generally contain paper crease noise, and fixed-window denoising methods cannot simultaneously achieve denoising effect and map line edge clarity. The system divides the scanned map image into multiple local blocks, calculates the local pixel variance of each local block, and calculates the global pixel variance of the scanned map image. The local pixel variance reflects the smoothness or texture change of the local area, while the global pixel variance reflects the overall fluctuation of the entire image.
[0067] By calculating the global pixel variance Mean local pixel variance of all local blocks The ratio generates the noise intensity value. Its calculation logic is as follows:
[0068]
[0069] in, Represents a mathematical function for calculating the mean, and when When the value is 0, it is determined that there is no noise in the scanned map image, and the noise intensity value is directly set to 0. 0; the larger this ratio, the stronger the global structural noise of the image; combined with the scaling factor fitted based on the quality calibration library of historical map scans. The window size of the denoising filter is dynamically determined based on the noise intensity value. The calculation formula is: , where the symbol This indicates a floor function; when processing a historical scanned map containing paper creases with noise levels exceeding a preset threshold, the system calculates the mean of local pixel variances when the global pixel variance is greater than a preset proportion, generating a noise intensity value. The value is 4.5; at the preset scaling factor. When the value is 2, the system automatically calculates the filter window size to be 13, which is the floor value of the product of 2 and 4.5 plus 3. Since 12 is an even number, we add one to constrain it to an odd dimension, and then we get... This allows for the adaptive allocation of a large 13×13 filtering window for the high-noise image;
[0070] This adaptive evaluation mechanism effectively avoids the defects of a uniform small window failing to smooth crease noise or a uniform large window causing blurred lines, ensuring effective smoothing of crease noise and preserving the clarity of map line edges to the greatest extent.
[0071] The image denoising and correction unit also performs tilt detection and synchronous correction, specifically:
[0072] The Hough transform is used to detect the inner and outer map contours of the scanned map image and to calculate the tilt angle of the inner and outer map contours relative to the horizontal reference.
[0073] Construct an affine transformation matrix containing the tilt angle. During the process of spatial filtering and denoising the pixels traversed by the scanned map image, map the coordinates of the target pixel to the original image through the inverse matrix of the affine transformation matrix to find neighboring pixels.
[0074] A filter with a window size equal to the window size is used to calculate the pixel grayscale value of the neighboring pixels. Noise reduction and tilt correction are completed simultaneously through one bilinear interpolation, and a normalized orthophoto is output. The OCR recognition engine is then called to extract the image annotation text.
[0075] The system employs Hough transform to detect tilt angles and constructs an affine transformation matrix. Denoising and tilt correction are then performed simultaneously during the filtering process via inverse matrix mapping. Traditional serial denoising and tilt correction methods cause secondary resampling of the image, resulting in blurred map line edges and reducing the accuracy of subsequent OCR text recognition. The system uses Hough transform to detect the inner and outer map contours of the scanned map image and calculates the spatial tilt angles of these contours relative to the horizontal reference. ; Constructing a structure that includes tilt angles affine transformation matrix ;
[0076] During the process of performing spatial filtering and denoising on the scanned map image pixels, the system transforms the target pixel coordinates using the inverse of the affine transformation matrix. The neighboring pixels are found by directly mapping back to the original oblique image; within this neighborhood, the dynamic window size calculated in the previous steps is used. The filter calculates the pixel grayscale value for neighboring pixels, and denoising and tilt correction can be completed jointly with only one bilinear interpolation. When processing a paper map image with a 3-degree scanning tilt, the system uses Hough transform to detect the map outline and calculate the 3-degree tilt angle, and constructs an affine transformation inverse matrix containing a 3-degree rotation. During pixel-by-pixel filtering, the system directly extracts the pixel grayscale values of the corresponding neighboring pixels from the original tilted image for calculation, and outputs a regularized orthophoto image with creases removed and the map outline completely horizontal in one go.
[0077] This mechanism of combined noise reduction and tilt correction preserves the accuracy of spatial data and avoids the line blurring problem caused by traditional secondary interpolation. This reduces the character recognition error rate caused by line edge blurring when the built-in OCR recognition engine is called to extract the text annotations on the map.
[0078] In this embodiment, when the format type is a CAD vector file, the multi-format data processing module calls the vector parsing and normalization unit to perform layer parsing and coordinate system inference, specifically:
[0079] Parse the internal entity object tree of the CAD vector file, extract the source layer name, and classify the features using a pre-configured layer name mapping table that contains the correspondence between the source layer name and the standard layer name;
[0080] Extract the spatial bounding box coordinate extrema of the CAD vector file. When the CAD vector file lacks a coordinate reference system definition, infer the source projection zone and source coordinate system based on the range of the spatial bounding box coordinate extrema. The specific logic of the inference is as follows: extract the total length of the integer digits of the coordinate maxima in the spatial bounding box coordinate extrema and the characteristic digits of the preset number of digits at the beginning. Input the total length of the integer digits and the characteristic digits into a pre-configured coordinate system rule discrimination tree for traversal and comparison. If a preset national standard or industry standard coordinate system numerical range rule is matched, the corresponding source projection zone and source coordinate system are output. If no numerical range rule is matched, a coordinate system unknown anomaly mark is output and the coordinate system inference process of the current file is terminated.
[0081] This system parses the internal entity tree of CAD files to extract source layer names and uses a layer name mapping table for preliminary feature classification. It also extracts spatial bounding box coordinate extrema to infer the source projection zone and source coordinate system when a coordinate reference system definition is missing. In GIS data processing scenarios where CAD vector files are aggregated across projects, issues such as non-standard layer naming and missing coordinate systems often arise, making it difficult to directly perform spatial overlay analysis on heterogeneous data. The system parses the internal entity tree of CAD vector files, extracts source layer names, and automatically maps and classifies them using a pre-configured layer name mapping table that includes the correspondence between source layer names and standard layer names.
[0082] The system extracts the spatial bounding box coordinate extreme values of the CAD vector file. When the file itself lacks a coordinate reference system definition, the system infers its source projection zone and source coordinate system based on the order of magnitude and range characteristics of the spatial bounding box coordinate extreme values. When an operator imports a computer-aided design vector file in a standard graphic format provided by an early surveying department, the file does not define a coordinate reference system and the layer naming does not conform to the preset specifications.
[0083] After parsing the internal entity object tree, the system automatically unified the source layers named Source Layer A and Source Layer B into the standard residential map layer according to the layer name mapping table, completing the initial unification of feature classification; at the same time, the system read that the maximum X-coordinate of the bounding box of this file space is 38543210, where 38543210 is an 8-digit integer and starts with the band number 38, and the maximum Y-coordinate is 4321098, where 4321098 is a 7-digit integer;
[0084] Based on the range characteristics of this coordinate extreme value, the system accurately infers that the source projection zone of the file is the 38-degree zone and the source coordinate system is the Gauss-Kruger projection; this process effectively solves the problems of layer semantic confusion and coordinate system loss.
[0085] After inferring the source projection zone and source coordinate system, the vector analytical and normalized unit establishes a coordinate transformation model:
[0086] Construct a coordinate transformation matrix based on the inferred source coordinate system and the system's preset standard target coordinate system;
[0087] Perform batch transformations based on coordinate transformation matrices on the vertex coordinates of all extracted spatial entities, and output standard vector data under a unified geographic framework.
[0088] Based on the inferred source coordinate system and standard target coordinate system, a coordinate transformation matrix is constructed, and the transformation based on this matrix is performed on the vertex coordinates of all spatial entities in batches. In order to solve the problem of spatial coordinate accuracy loss and inability to overlay analysis caused by the inconsistency of coordinate systems of CAD vector files from different sources, the system establishes a coordinate transformation model after inferring the source projection zone and source coordinate system.
[0089] Based on the inferred source coordinate system and the system's preset standard target coordinate system, a coordinate transformation matrix containing affine and projection transformations is constructed. Using this coordinate transformation matrix, the vertex coordinates of all extracted lines, polygons, and other spatial entities are batch mapped and calculated, outputting standard vector data under a unified geographic framework. When the system infers that the source coordinate system of a certain DXF format CAD vector file is the Beijing 54 coordinate system Gaussian projection, the system automatically retrieves the preset CGCS2000 latitude and longitude coordinate system as the standard target coordinate system and constructs the corresponding projection transformation combination matrix as the coordinate transformation matrix.
[0090] The system extracts the vertex coordinates of all polygonal spatial entities from this file. Batch execution of transformations based on coordinate transformation matrices, where, and The original x and y coordinates of the vertices of spatial entities are represented respectively, and standard vector data under a unified geographic framework is output. The output standard vector data can be seamlessly overlaid and analyzed with the existing CGCS2000 standard database, which solves the problem of coordinate system confusion between different files and ensures the geographic logic correctness and spatial location accuracy of the data in the database.
[0091] In this embodiment, when the format type is PDF document, the multi-format data processing module calls the hybrid document feature extraction unit to perform synchronous feature extraction, specifically:
[0092] Parse the underlying page content flow of the PDF document, construct the document object tree, and separate vector graphics operation instructions from text drawing instructions;
[0093] The vector graphics manipulation commands are restored to GIS standard geometric objects, and the coordinates of the geometric center point of the GIS standard geometric objects in the coordinate system of the PDF document page are extracted.
[0094] Simultaneously extract the text stream content and its bounding box center coordinates using text drawing commands.
[0095] The system parses the underlying page content flow of portable documents to separate graphic and text instructions, restores graphic instructions to standard geometric objects and extracts the coordinates of their geometric center points, and simultaneously extracts the text flow content and its bounding box center coordinates. When processing portable planning documents containing mixed text and graphics, spatial vector line segments and annotation text are usually stored together. Traditional extraction methods often can only extract graphics or text separately, making it difficult to maintain the geographical location association between text and spatial entities. The system deeply parses the underlying page content flow of portable documents, constructs a document object tree, and separates vector graphics operation instructions from text drawing instructions.
[0096] The system restores the extracted vector graphics operation instructions into standard GIS geometric objects such as points, lines, and polygons, and calculates and extracts the coordinates of the geometric center points of these standard GIS geometric objects in the portable document page coordinate system. Simultaneously, the system extracts the specific content of the text stream and its bounding box center coordinates through the extracted text drawing instructions. When an operator imports a portable document containing a mix of planning drawings and explanatory text, the system reads the graphic construction operation instructions in the underlying page content stream, restores it into a polygonal geometric object representing an industrial land plot, and extracts the coordinates of the polygon's geometric center point in the page coordinate system. Simultaneously, the system identifies the text drawing operation instructions in the content stream and extracts the annotation text "industrial land" and its corresponding bounding box center coordinates.
[0097] This process enables the independent and synchronous structured extraction of vector and text elements in a hybrid document, laying the foundation for subsequent spatial association between images and text.
[0098] After extracting coordinates, the hybrid document feature extraction unit performs spatial location association, specifically as follows:
[0099] The following logical judgment is executed, with the judgment condition being: calculate the Euclidean distance between the coordinates of the bounding box center and the coordinates of the geometric center point; when the Euclidean distance is less than or equal to the association distance threshold, determine that the corresponding text stream content is the spatial attribute annotation of the corresponding GIS standard geometric object; when the Euclidean distance is greater than the association distance threshold, determine that the corresponding text stream content is independent text.
[0100] When determined to be a spatial attribute annotation, the text stream content is written as an attribute field into the attribute table of the GIS standard geometric object; the associated distance threshold is calculated by multiplying the DPI parameter parsed from the PDF document by the preset tolerance physical size in inches.
[0101] The system calculates the Euclidean distance between the bounding box center coordinates and the geometric center point coordinates, and determines the attribute annotation affixation based on the association distance threshold to achieve structured binding of text and image data. To solve the technical problem of separating spatial graphics and attribute annotations in planning portable documents, the system performs spatial location association after extracting the coordinates.
[0102] The system extracts the center coordinates of the bounding box of the text stream content. Geometric center point coordinates of GIS standard geometric objects And calculate the Euclidean distance between them. Its calculation formula is ,in, and These represent the x and y coordinates of the center point of the text stream's bounding box, respectively. and These represent the x and y coordinates of the geometric center point of a standard GIS geometric object, respectively.
[0103] The system multiplies the image resolution parameters parsed from the portable document with the physical page size to calculate the association distance threshold, which represents the effective annotation tolerance distance on the page. When the calculated Euclidean distance is less than or equal to the association distance threshold, it indicates that the text meets the spatial association conditions with the corresponding geometric entity, and the system determines that the corresponding text stream content is a spatial attribute annotation of the corresponding GIS standard geometric object. When the Euclidean distance is greater than the association distance threshold, it indicates that the text is outside the association range of the geometric entity, and the system determines that the corresponding text stream content is irrelevant independent text. When determined to be a spatial attribute annotation, the system writes the text stream content as an attribute field into the attribute table of the corresponding GIS standard geometric object, realizing the structured binding of text and image data.
[0104] When the operator was processing the aforementioned planning drawing PDF document, the system calculated that the Euclidean distance between the bounding box center coordinates of the extracted annotation text "industrial land" and the geometric center point coordinates of the polygon representing the industrial land plot was 12 pixels; the system retrieved the association distance threshold obtained by multiplying the current PDF document's DPI parameter by the page's physical size, which was 18 pixels; since 12 is less than 18, the system determined that the industrial land text was the spatial attribute annotation of the polygon plot, and automatically wrote the industrial land text as an attribute field into the attribute table of the polygon vector object;
[0105] This spatial anchoring and association mechanism successfully solved the problem of separating graphics and annotations in PDF documents, ensuring the integrity of the data and the correctness of the geographic logic.
[0106] In this embodiment, the topology consistency check performed by the data verification and database entry module includes polygon self-intersection detection and dangling node detection;
[0107] Standard formats include GeoJSON or WKT.
[0108] The received standard spatial data undergoes a topology consistency check, including polygon self-intersection detection and dangling node detection. After passing the check, the data is converted to GeoJSON or WKT format and written to the spatial database. To ensure the quality and spatial logic validity of the final data, the data verification and database entry module receives the standard spatial data output by the multi-format data processing module and performs a strict GIS topology consistency check. The topology consistency check specifically includes polygon self-intersection detection and dangling node detection, which are used to identify topology errors such as boundary intersections or unclosed line segments in the geometric structure of spatial entities.
[0109] When standard spatial data passes the topology consistency check, the system converts it into a preset standard format, which may include GeoJSON or WKT. The data is then automatically written into the underlying GIS spatial database through the spatial database middleware. When the topology consistency check fails, the system intercepts the corresponding data and outputs an error record. When operators process surveying data aggregated across projects in batches, the data verification and data entry module receives a batch of standard spatial data of residential polygons that has been processed by vector parsing and normalization units.
[0110] The system automatically performs a topology consistency check on the batch of data, identifying three polygons with boundary self-intersection errors and two independent dangling node segments. These failed data are then intercepted and output as anomaly records for later traceability. For the remaining valid polygon data that successfully pass the topology consistency check, the system converts them into geographic object symbol format text and automatically writes them in batches to the underlying spatial object relational database through the spatial database middleware interface.
[0111] This mechanism ensures the spatial logical correctness of the data entering the database and realizes the process of identifying multi-source heterogeneous spatial data from the physical layer, dynamically scheduling and organizing it, and then automating its entry into the database.
[0112] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A GIS data automated standardization and warehousing system based on multi-format file OCR recognition, comprising a data format recognition module, a preprocessing and distribution module, a multi-format data processing module, and a data verification and warehousing module, characterized in that, The data format recognition module reads the spatial data file as a file data stream, extracts the file header identifier, spatial organization information and grayscale block information corresponding to the preset length binary stream of the file data stream, and inputs them into the convolutional neural network recognition model to obtain the highest classification confidence and format type. The preprocessing distribution module is configured with a high confidence threshold and a low confidence threshold; when the highest classification confidence is greater than or equal to the high confidence threshold, the file data stream is connected to the regular preprocessing process to generate regular routing instructions; when the highest classification confidence is greater than or equal to the low confidence threshold and less than the high confidence threshold, it is switched to the supplementary preprocessing process to generate supplementary routing instructions; when the highest classification confidence is less than the low confidence threshold, an interception instruction is generated. The multi-format data processing module responds to the regular routing instruction or the supplementary routing instruction and routes according to the format type: for scanned map images, it calls the image denoising and correction unit; for CAD vector files, it calls the vector parsing and normalization unit; for PDF documents, it calls the hybrid document element extraction unit. It then normalizes the file data stream, calls the OCR recognition engine to extract text annotation information and links it with the normalized spatial geometric data attributes to output standard spatial data. The data verification and database entry module performs a topology consistency check on the standard spatial data. If the check passes, the data is converted to a standard format and written to the spatial database. If the check fails, an exception record is output.
2. The GIS data automated organization and storage system based on multi-format file OCR recognition according to claim 1, characterized in that, When extracting the spatial organization information, the data format recognition module searches whether the file data stream contains a preset GIS identifier byte, which includes layer, projection, or reference plane identifier information. The convolutional neural network recognition model is trained based on a pre-built sample database containing files in various known GIS formats.
3. The GIS data automated organization and storage system based on multi-format file OCR recognition according to claim 1, characterized in that, The high confidence threshold and low confidence threshold in the preprocessing distribution module are set based on the historical error rate statistics of spatial data classification; When the file data stream is transferred to the supplementary preprocessing process, the system triggers a program to perform integrity verification and repair on the file header identifier and data blocks of the file data stream, and then hands it over to the multi-format data processing module for processing.
4. The GIS data automated organization and storage system based on multi-format file OCR recognition according to claim 1, characterized in that, When the format type is scanned map image, the multi-format data processing module calls the image denoising and correction unit to perform noise assessment, specifically: The scanned map image is divided into multiple local blocks, the local pixel variance of each local block is calculated, and the global pixel variance of the scanned map image is calculated. The following logical judgment is executed, and the judgment condition is as follows: calculate the ratio of the global pixel variance to the mean local pixel variance of all the local blocks to generate a noise intensity value; dynamically determine the window size of the denoising filter based on the noise intensity value, wherein the window size is equal to the floor value of the product of the scaling factor and the noise intensity value plus three, and when the calculated window size is an even number, add one to constrain it to an odd dimension; wherein the scaling factor is a preset weight constant.
5. The GIS data automated organization and storage system based on multi-format file OCR recognition according to claim 4, characterized in that, The image denoising and correction unit also performs tilt detection and synchronous correction, specifically: The Hough transform is used to detect the inner and outer contour lines of the scanned map image, and the tilt angle of the inner and outer contour lines relative to the horizontal reference is calculated. An affine transformation matrix containing the tilt angle is constructed. During the process of spatial filtering and denoising the pixels traversed by the scanned map image, the coordinates of the target pixel are mapped to the original image through the inverse matrix of the affine transformation matrix to find neighboring pixels. The grayscale values of the neighboring pixels are calculated using a filter with a window size equal to the window size. Noise reduction and tilt correction are completed simultaneously through one bilinear interpolation, and a normalized orthophoto image is output. The OCR recognition engine is then called to extract the image annotation text.
6. The GIS data automated organization and storage system based on multi-format file OCR recognition according to claim 1, characterized in that, When the format type is a CAD vector file, the multi-format data processing module calls the vector parsing and normalization unit to perform layer parsing and coordinate system inference, specifically: The internal entity object tree of the CAD vector file is parsed to extract the source layer name, and the elements are classified using a pre-configured layer name mapping table that contains the correspondence between the source layer name and the standard layer name. Extract the spatial bounding box coordinate extreme values of the CAD vector file. When the CAD vector file lacks a coordinate reference system definition, infer the source projection zone and source coordinate system based on the range of the spatial bounding box coordinate extreme values. The specific logic of the inference is as follows: extract the total length of the integer digits of the coordinate maxima in the spatial bounding box coordinate extreme values and the characteristic digits of the preset number of digits at the beginning. Input the total length of the integer digits and the characteristic digits into a pre-configured coordinate system rule discrimination tree for traversal and comparison. If a preset national standard or industry standard coordinate system numerical range rule is matched, the corresponding source projection zone and source coordinate system are output. If the numerical range rule is not matched, a coordinate system unknown anomaly mark is output and the coordinate system inference process of the current file is terminated.
7. The GIS data automated straightening and storage system based on multi-format file OCR recognition according to claim 6, characterized in that, After inferring the source projection zone and source coordinate system, the vector analysis and normalization unit establishes a coordinate transformation model: Based on the inferred source coordinate system and the system's preset standard target coordinate system, a coordinate transformation matrix is constructed; The vertex coordinates of all extracted spatial entities are batch-transformed based on the coordinate transformation matrix to output standard vector data under a unified geographic framework.
8. The GIS data automated straightening and storage system based on multi-format file OCR recognition according to claim 1, characterized in that, When the format type is PDF document, the multi-format data processing module calls the hybrid document element extraction unit to perform synchronous element extraction, specifically: The underlying page content flow of the PDF document is parsed, a document object tree is constructed, and vector graphics operation instructions and text drawing instructions are separated. The vector graphics operation instructions are restored to GIS standard geometric objects, and the coordinates of the geometric center point of the GIS standard geometric objects in the coordinate system of the PDF document page are extracted; The text stream content and its bounding box center coordinates are extracted synchronously using the text drawing instructions.
9. The GIS data automated straightening and storage system based on multi-format file OCR recognition according to claim 8, characterized in that, After extracting the coordinates, the hybrid document element extraction unit performs spatial location association, specifically as follows: The following logical judgment is performed, and the judgment condition is: calculate the Euclidean distance between the coordinates of the bounding box center and the coordinates of the geometric center point; when the Euclidean distance is less than or equal to the association distance threshold, determine that the corresponding text stream content is the spatial attribute annotation of the corresponding GIS standard geometric object; when the Euclidean distance is greater than the association distance threshold, determine that the corresponding text stream content is independent text. When determined to be a spatial attribute annotation, the text stream content is written as an attribute field into the attribute table of the GIS standard geometric object; wherein, the association distance threshold is calculated by multiplying the DPI parameter parsed from the PDF document by a preset tolerance physical size in inches.
10. The GIS data automated organization and storage system based on multi-format file OCR recognition according to claim 1, characterized in that, The topology consistency check performed by the data verification and storage module includes polygon self-intersection detection and dangling node detection; The standard format includes GeoJSON or WKT format.