Multi-village government affair service optimization method and system based on Internet

By combining satellite imagery, drone mapping, and mobile data collection with OCR and AI technologies, the problem of low data collection efficiency and poor data quality in rural government services has been solved. This has enabled data standardization and real-time updates, thereby improving the efficiency and accuracy of rural government services.

CN121960841APending Publication Date: 2026-05-01HUBEI EDONG DIGITAL GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511661436.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In rural government services, data collection is inefficient, data quality is poor, there is a lack of unified standards, evidence storage is unreliable, and dynamic maintenance is insufficient, resulting in inefficient land ownership confirmation and population management, and easily leading to disputes.

Method used

By combining satellite imagery, drone mapping, and mobile offline data collection apps, and using OCR technology to batch process historical paper archives, we identify abnormal data through AI and RAG intelligent cleaning models, construct knowledge graphs to achieve data standardization and correlation, use blockchain to store and confirm ownership results, deploy IoT sensors for dynamic monitoring, and combine AI models for data quality inspection.

Benefits of technology

It has enabled the efficient aggregation and digital transformation of rural government data, improved data integrity and accuracy, ensured data security and real-time updates, reduced the workload of grassroots staff, and improved service efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121960841A_ABST
    Figure CN121960841A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of multi-village government affair service optimization methods, and particularly relates to an internet-based multi-village government affair service optimization method and system, which adopts a mode of combining satellite images, unmanned aerial vehicle surveying and mapping and mobile terminal off-line acquisition APP, grabs land boundary and terrain space data, acquires population household registration and contract business data, and provides a multi-village government affair service optimization method and system. Historical paper archives are identified through OCR batch scanning, and information is extracted and input into an information management system; the method comprises the following steps: constructing an AI and RAG intelligent cleaning model, filling data missing values by applying a generative adversarial network, recognizing population repeated registration and land area conflict abnormal data through a local abnormal factor algorithm, training an entity recognition model by relying on a rural government affair vertical RAG library, and automatically labeling data attributes and association relationships.
Need to check novelty before this filing date? Find Prior Art

Description

Internet-based methods and systems for optimizing multi-village government services Technical Field

[0001] This invention belongs to the technical field of multi-village government service optimization methods, and particularly relates to multi-village government service optimization methods and systems based on the Internet. Background Technology

[0002] Rural government services have long faced significant bottlenecks in data collection and integration. Traditional methods rely heavily on manual on-site recording and paper archives. Due to weak network coverage in remote areas, spatial data such as land boundaries and topography are often fragmented with operational data such as population registration and contracting agreements, making simultaneous aggregation difficult. Furthermore, many historical paper archives suffer from yellowing, blurred handwriting, and inconsistent formats due to varying preservation environments. Digital processing requires manual entry of each document, which is not only inefficient but also prone to human error, compromising data integrity and accuracy. This directly impacts the efficiency and quality of core government services such as land ownership confirmation and contracting transfer.

[0003] In terms of data processing, security, and dynamic maintenance, rural government data is scattered and lacks unified standards. Data cleaning relies heavily on manual screening, making it difficult to efficiently identify anomalies such as duplicate population registrations and land area conflicts. The cross-departmental data sharing mechanism is inadequate, and data from departments such as civil affairs and natural resources cannot be effectively cross-verified, easily leading to data contradictions. Furthermore, data modification records lack credible documentation, making key information such as land ownership confirmation results prone to disputes. Dynamic information such as changes in land use status and population shifts cannot be updated in real time, resulting in insufficient data timeliness. This not only hinders the standardization and intelligent development of multi-village government services but also increases the workload at the grassroots level and the difficulty of dispute resolution. Summary of the Invention

[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide an internet-based method for optimizing multi-village government services. The method includes: using a combination of satellite imagery, UAV mapping, and a mobile offline data collection app to capture land boundary and topographic spatial data; collecting population registration and contract data; using OCR to batch scan and identify historical paper archives; extracting information and entering it into an information management system; constructing an AI and RAG intelligent cleaning model; using generative adversarial networks to fill in missing data values; using a local anomaly factor algorithm to identify abnormal data such as duplicate population registrations and land area conflicts; training an entity recognition model based on a rural government vertical RAG library; and automatically labeling data attributes and relationships; and referring to the agricultural data standard system, compiling unified data characters for land classification codes and population information fields. The system clearly defines data formats, accuracy requirements, and association rules. Tools are used to batch convert raw data from different sources and formats into a standard format, constructing a knowledge graph of land, population, and assets to achieve interconnectivity and interoperability of land, people, and assets data. Cross-departmental cross-verification and blockchain evidence storage are implemented, combining data source cross-verification with time-series verification to compare and verify core data such as population identity information and land ownership boundaries. Key information such as confirmation results and data modification records are stored on the blockchain, and data verification rules are automatically executed through smart contracts. IoT sensors are deployed to monitor changes in land use status in real time, establishing a real-time data synchronization channel with the civil affairs system. AI models are used to regularly conduct full-scale data quality checks, generating yellow, orange, and red early warning reports. Rectification processes are automatically triggered for data missing or errors, continuously maintaining data quality.

[0005] In another aspect, embodiments of the present invention also provide an Internet-based multi-village government service optimization system, comprising: a processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the aforementioned Internet-based multi-village government service optimization system by executing the machine-executable instructions.

[0006] In another aspect, embodiments of the present invention also provide a computer program product, the computer program product including machine-executable instructions, the machine-executable instructions being stored in a computer-readable storage medium, the processor of a computer device reading the machine-executable instructions from the computer-readable storage medium, the processor executing the machine-executable instructions, causing the computer device to execute the aforementioned Internet-based multi-village government service optimization system.

[0007] Based on the above, and addressing the core pain points of data collection difficulties, poor quality, inconsistent standards, unreliable evidence storage, and insufficient dynamic maintenance in rural government services, comprehensive optimization is achieved through multi-dimensional technological innovation. In the data collection phase, a combination of satellite imagery, drone mapping, and a mobile offline data collection app is employed, along with OCR technology for batch processing of historical paper archives. This effectively overcomes the limitations of weak network coverage and scattered archives in remote rural areas, enabling efficient aggregation and digital transformation of land spatial data with business data such as population registration and contracting agreements, thus solving the problems of low efficiency and incomplete coverage associated with traditional collection methods.

[0008] In the data processing stage, relying on AI and RAG intelligent cleaning models, generative adversarial networks are used to accurately fill in missing data values. Local anomaly factor algorithms are combined to identify abnormal data such as duplicate population registrations and land area conflicts. Simultaneously, an entity recognition model is trained using a rural government affairs vertical RAG library to automatically label data attributes and relationships, significantly reducing manual intervention and greatly improving data integrity and accuracy. Regarding data standardization and correlation, a unified data dictionary is compiled with reference to the agricultural data standard system, and a knowledge graph of land, population, and assets is constructed. This breaks down barriers between data from different sources and formats, enabling the interconnection and interoperability of land, people, and assets data, solving the problem of integrating cross-type data. In terms of data security and credibility, cross-departmental verification ensures the accuracy of core data. A lightweight consortium blockchain is used to record key information such as rights confirmation results and modification records on the blockchain for evidence storage. Smart contracts solidify verification rules, strengthening data credibility and traceability, and avoiding the risk of data tampering. In the dynamic data maintenance stage, IoT sensors are deployed to monitor land use status in real time, establishing a real-time synchronization channel with the civil affairs system. Combined with AI models, regular full-scale inspections are conducted and tiered early warnings are generated, automatically triggering rectification processes to ensure the long-term freshness and effectiveness of the data. Overall, the solution effectively reduces the labor costs and operational difficulties of rural grassroots government services, improves service efficiency and accuracy, and fully adapts to the needs of complex rural scenarios. Attached Figure Description

[0009] Figure 1 is a flowchart illustrating the multi-village government service optimization method based on the Internet according to the present invention.

[0010] Figure 2 is a schematic block diagram of the multi-village government service optimization method based on the Internet of the present invention. Detailed Implementation

[0011] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 is a flowchart illustrating an embodiment of the Internet-based multi-village government service optimization method provided by the present invention. The Internet-based multi-village government service optimization method will be described in detail below.

[0012] Step S110: Using a combination of satellite imagery, UAV mapping, and mobile offline data collection APP, land boundary and topographic spatial data are captured, population registration and contract business data are collected, historical paper archives are scanned and identified in batches through OCR, information is extracted and entered into the information management system.

[0013] In this embodiment, for remote mountainous areas with weak network coverage, a mobile offline data collection app can be pre-installed with high-definition satellite base maps and pre-marked land parcel information. When staff conduct on-site verification with their equipment, they can use GPS to locate and match land parcels, manually correct boundary deviations, and take on-site photos. After data collection, the data is automatically synchronized in the village committee's WiFi environment. A long-endurance multi-rotor drone equipped with a LiDAR module is used to conduct low-altitude aerial surveys of complex terrains such as mountains and forests, generating a 1:500 high-precision digital elevation model. This model accurately extracts spatial data such as land parcel boundaries and slopes, compensating for blind spots in satellite imagery in obscured areas. For traditional villages with abundant historical paper archives, a dedicated archive digitization workstation is built, equipped with a high-speed scanner and enhanced OCR equipment. For archives from different eras and in different preservation conditions (yellowed paper, handwritten contracts, documents with blurred seals), image enhancement technology is first used to digitize them. The system optimizes clarity and then activates multimodal OCR recognition to automatically extract key fields such as land parcel number, area, contract period, and contractor identity information from the contract, and name, ID number, kinship, and residential address from the population registration file. It also supports batch processing and breakpoint resume to avoid repetitive work. After recognition, it automatically associates the corresponding spatial data with the personnel ID. For villages with frequent land transfers and large population mobility, the mobile APP adds self-declaration and family / friend proxy functions. Farmers can take photos of their ID card and new transfer contract, and the OCR will automatically extract the transfer information. After filling in the land change description, they can submit the application. Migrant workers can authorize family members to upload relevant materials on their behalf. The system verifies the kinship between the proxy and the applicant through facial recognition and simultaneously connects to the household registration system and the township agricultural economic department database to automatically verify the authenticity of identity information and the validity of the contract, achieving real-time updates of transfer data. For data collection scenarios involving collective assets (village collective construction land, forest land, ponds, etc.), UAV oblique photography technology is used to generate 3D reality models. Combined with satellite imagery time-series data, the asset utilization status and change trajectory are identified. OCR technology is used to extract core information from forest tenure certificates, construction land approval documents, and asset lease agreements in batches, linking them to spatial coordinates in the 3D model to form integrated map and ownership data. Simultaneously, low-power IoT positioning terminals are deployed in key asset areas to monitor asset usage status in real time, with data automatically synchronized to the resource pool for dynamic tracking. To address sudden scenarios such as post-disaster land boundary changes and population information shifts, an emergency data collection mechanism is activated: UAVs conduct emergency aerial surveys to acquire post-disaster terrain data, comparing it with pre-disaster data to identify boundary offsets and land damage. A mobile app is enabled in offline priority mode to collect information on affected land and population relocation and resettlement in environments without network access, transmitting it to the village committee terminal via Bluetooth, and automatically synchronizing after network recovery. OCR technology quickly identifies post-disaster land ownership certificates, temporary resettlement agreements, and other materials, rapidly updating land ownership and population residence information in the resource pool, providing data support for post-disaster reconstruction and policy implementation.After data collection is completed, the system automatically starts the preliminary cleaning process: duplicate collection records of the same plot are removed by comparing spatial coordinates, duplicate entries in population data are removed by verifying ID card numbers, and fields with large OCR recognition errors (such as fuzzy numbers and variant characters) are automatically marked as pending review and pushed to village collective grid members and township data administrators.

[0014] Step S111: Select satellite imagery covering the entire target area, preprocess the images using remote sensing image processing software to generate digital orthophoto maps, automatically identify land use types such as cultivated land, forest land, and construction land based on a deep learning semantic segmentation model, mark the boundary contours and topographic relief features of the plots, and form a pre-labeled plot dataset.

[0015] In this embodiment, based on the actual needs of rural collective land ownership confirmation, satellite images with a resolution of 0.5-1 meters from Gaofen-2 and Sentinel-2 are prioritized. The coverage area is delineated based on the administrative boundaries of townships to ensure that all cultivated land, forest land, homesteads, collective construction land, and corner plots such as pits and ditches owned by the village collective are included without omission. Then, using ENVI or ArcGIS remote sensing processing software, cloud shadows and atmospheric scattering noise in the images are first targeted to be removed. Then, orthorectification under the UTM projection coordinate system is completed through ground control points. During stitching, color equalization processing is performed on overlapping areas to avoid stitching traces, generating a unified digital orthorectified image map of the entire region. Based on local confirmed land ownership samples, the U-Net semantic segmentation model is finely adjusted. The pre-processed images are input to automatically identify various land use types. The continuous boundaries of plots are delineated through pixel-level classification. At the same time, the differences in image gray values ​​are used to help judge the terrain undulation characteristics (plain plots are marked as gentle, and mountainous and hilly plots are marked as undulating). Finally, a pre-labeled dataset in shp format containing plot numbers, land type codes, boundary coordinate strings, and terrain descriptions is output.

[0016] Step S112: Based on the satellite image annotation results and the village terrain, use UAV aerial survey routes; use grid-like routes for plain areas and strip-like routes for mountainous and hilly areas.

[0017] In this embodiment, considering the actual differences in rural terrain, the distribution of land parcels pre-marked in satellite imagery (with a focus on scattered corner parcels and terrain transition areas) is first assessed in conjunction with the village committee to determine the complexity of the village's terrain. When planning grid-like flight routes for villages in plains, a flight altitude of 120-150 meters, a forward overlap rate of 80%, and a lateral overlap rate of 60% are set to ensure that no land parcel boundaries are missed. For villages in mountainous and hilly areas, a strip-like flight route is adopted, with flight paths planned along contour lines, and the flight altitude is increased to 180-220 meters, with the overlap rate increased by 10%. To avoid data loss due to mountain obstruction, industrial-grade drones such as the DJI M300RTK or Huace P700 are selected, equipped with a 16-line LiDAR and a 20-megapixel full-frame camera. Flight plans are reported to the local air traffic control department in advance before the aerial survey, avoiding obstacles such as high-voltage lines and communication base stations. Low-altitude aerial surveys are carried out in clear weather with visibility of ≥5 kilometers and wind force ≤3. During the flight, the flight path deviation is corrected in real time through the drone flight control system to ensure stable flight along the planned path, and high-precision terrain point cloud data and multi-angle real-scene images are collected simultaneously.

[0018] Step S113: Using point cloud data obtained from UAV aerial surveys, digital elevation models and digital surface models are generated using 3D modeling software. Topographic parameters such as slope, aspect, and altitude of the plots are extracted. The boundaries are pre-annotated by combining real-world images and satellite images. A multi-source image registration algorithm is used to correct the plot boundary deviations, generating plot vector data and forming preliminary spatial data results.

[0019] In this embodiment, the point cloud data acquired by UAV aerial survey is first preprocessed using CloudCompare software to remove non-ground points such as vegetation and buildings, as well as noise points (retaining ground point density ≥ 30 points / square meter). Then, it is imported into ContextCapture or Agisoft Metashape 3D modeling software, and a digital elevation model and digital surface model are generated by setting a 0.5-meter grid size. Topographic parameters such as slope, aspect, and elevation of the plots are extracted from the DEM using ArcGIS spatial analysis tools. Simultaneously, special topographic features such as abrupt slope changes and low-lying water accumulation are labeled. Subsequently, the pre-labeled boundaries from satellite imagery are used as the initial reference. Using SIFT and RANSAC multi-source image registration algorithms, UAV real-world images and satellite orthophotos are aligned pixel-level. By comparing the detailed features of natural boundaries such as field ridges, roads, and rivers in the real-world images, boundary deviations caused by scale or occlusion in the satellite images are corrected (ensuring the deviation is controlled within 0.5 meters). For scattered corner plots and overlapping areas of plots, boundaries are supplemented by combining multi-view real-world images. Finally, shp format plot vector data containing plot boundary coordinates (WGS84 coordinate system), land use type, terrain parameters, and area (accurate to 0.01 mu) are automatically generated, forming preliminary spatial data results that can be directly used for subsequent field verification.

[0020] Step S114: Pre-install the above-mentioned satellite image base map, land parcel vector data generated by the drone, and pre-labeled information in the mobile APP. After the staff arrives at the site with the equipment, turn on the offline mode of the mobile APP and match the current location with the land parcel data in real time through GPS / BeiDou dual-mode positioning function. Correct the deviation area by manually dragging the boundary node or drawing a new boundary. At the same time, take photos of the current status of the land parcel as supporting materials.

[0021] In this embodiment, satellite imagery can be compressed in advance to a mobile-friendly tile format at a 1:1000 scale. The drone-generated SHP format land parcel vector data simplifies coordinate accuracy and is associated with pre-labeled information such as parcel number and land type. This data is then pre-installed in batches on an offline data collection app on rugged tablets running Android 10 or later. Simultaneously, an offline map package of the target area is downloaded in advance to ensure positioning stability in offline environments. Once staff arrive on-site with their equipment, they can activate the app's offline mode with a single click. Using the device's built-in GPS / BeiDou dual-mode positioning module (positioning accuracy ≤1 meter; in obstructed areas such as forests, base station-assisted positioning correction can be used), they can continuously compare their location with the pre-installed land parcel vector data. Data overlay and matching are performed, and the current plot's boundary outline, pre-annotation information, and satellite image base map are displayed simultaneously on the APP interface. Then, the plot boundary is checked on-site by walking along it, focusing on comparing the consistency between natural boundaries such as field ridges, stone walls, and waterways and vector data. If boundary deviations are found (such as newly built field ridges not captured by satellite imagery, or boundary changes caused by plot merging and splitting), the boundary nodes can be manually dragged and fine-tuned using the APP's built-in editing tools, or the new boundary can be drawn directly in hand-drawn mode. At the same time, more than 5 current status photos are taken according to the standard of the four corners and center of the plot (the boundary markings, crops / buildings within the plot, and surrounding reference objects must be clearly shown). The photos are automatically associated with the current plot number and location information and stored locally on the device.

[0022] Step S115: For remote areas without network coverage, the mobile APP stores the corrected land boundary data, supplementary terrain information and on-site photos to the local cache. After the staff returns to the village committee or an area with network coverage, the mobile APP automatically triggers data synchronization, uploads the offline collected data to the cloud server, merges it with the spatial data generated by UAV aerial survey, and updates the land vector database.

[0023] For remote areas with no mobile network signal, such as mountainous and hilly areas, the mobile offline data collection app uses an AES-256 encrypted local caching mechanism. It stores the corrected plot boundary coordinates, supplemented terrain information such as steep slopes / low-lying areas / woodland obstructions, and supporting photos named by plot number and shooting time (automatically compressed to 1.5MB / photo, preserving clear details while saving device storage) in a separate encrypted partition on the device's built-in SD card. A local log containing the number of data entries, checksum, and collection time is generated simultaneously to prevent data tampering or loss. When staff return to areas with WiFi or 4G / 5G networks, such as village committees or township service centers, the app automatically detects network connectivity and triggers data synchronization without manual intervention. During synchronization, text-based data (boundary, terrain information, etc.) is uploaded first. The system first collects information and then uploads image data in batches according to the shooting order. If the network is interrupted due to fluctuations, the system automatically records the upload progress and resumes the upload from the breakpoint after the network is restored, avoiding duplicate uploads and wasting bandwidth. After the data is uploaded to the cloud server, the system accurately associates the offline collected correction data with the spatial data generated by UAV aerial survey using the unique identification code of the land parcel. Following the fusion rule that the priority of on-site verification data is higher than that of aerial survey data, the system automatically replaces the boundary coordinates of the deviation, updates the terrain parameters, and simultaneously mounts the supporting photos to the corresponding land parcel data file, completing the real-time update of the land parcel vector database. At the same time, the system retains the original version of the aerial survey data and marks it with the on-site correction mark of 202X X month, which facilitates subsequent data traceability and verification. The entire process is suitable for grassroots staff without professional technical backgrounds, and the security, integrity and accuracy of data synchronization can be ensured without additional operation.

[0024] Step S116: While collecting spatial data, business data is collected synchronously through the mobile APP. The APP's built-in OCR module automatically extracts the business data. After the farmers confirm that there are no errors on site, they sign electronically. The original contract is photographed or a scanned copy is uploaded through the APP. The OCR technology automatically recognizes the information and automatically associates it with the corresponding land parcel vector data.

[0025] In this embodiment, while simultaneously collecting land spatial data, staff can use a mobile app to collect business data such as population registration and contract information. When collecting population registration information, addressing common issues in rural areas such as worn edges on ID cards, illegible handwriting on household registration books, and blurred seals, the app's OCR module uses multispectral recognition technology to enhance character edges, supports mixed recognition of simplified / traditional / variant characters, and automatically extracts fields such as name, ID number, registered address, and kinship. Content with a recognition confidence level below 90% is highlighted in orange. After staff make corrections on-site by comparing with the original documents, farmers complete electronic signatures through the app's elderly mode, supporting touchscreen handwriting, fingerprint input, or proxy signing. Signature information is timestamped in real-time and appended to the data. When collecting contract information, for example, in 1998... In cases where the contract pages were yellowed, folded, or had manually altered area figures from the second round of contracting, the APP automatically initiates image restoration after staff photographs the contract. OCR technology prioritizes identifying the red-stamped area to confirm the contract's validity, then extracts information such as the plot number, contracted area, contract period, and signatures and seals of both parties. Alterations are automatically marked as requiring verification, and staff manually confirm these by questioning the farmers on-site. By scanning the plot code QR code on the contract, the APP instantly matches it with the collected plot vector data, displaying the corresponding plot's boundary graphic for verification. This ensures consistency between farmer information, contract content, and plot spatial data. All business data is stored using edge encryption to prevent leakage in village environments without network access, meeting the practical needs of diverse rural documents and complex operational scenarios.

[0026] Step S117: For historical paper archives that are stored in a scattered manner, they are centralized in the township government service center, and electronic images are generated by batch scanning using a high-speed scanner. The images are then imported into the OCR processing system. First, the image quality is optimized using image enhancement technology, and then the batch recognition function is activated to extract the core business fields in the archives. After being organized in a unified format, the data is matched with existing spatial data and real-time collected business data to fill in the gaps in historical data.

[0027] In this embodiment, village committees and grid workers can first organize a village-by-village search for historical paper archives scattered in village office metal cabinets, farmers' homes, and old accountant's records. The focus is on collecting archives with mold, insect damage, obscured binding, or illegible handwriting. These archives are then transported to the township government service center and sorted by type and year. Metal bindings are removed, wrinkles are smoothed, and severely damaged archives are sealed with transparent protective film. A high-speed A3 scanner is used to generate 300DPI TIFF format electronic images in batches. Archives with text obscured by binding are scanned double-sided and then stitched together. After importing into the OCR processing system, image enhancement technology is used for targeted optimization: removing mold, noise, correcting skewed images, and enhancing the contrast between handwriting and paper. Super-resolution reconstruction technology is used to improve the clarity of blurred text. Then, the batch recognition function is activated. The system is compatible with common rural archive fonts such as Song and Kai fonts (printed fonts) and Xing and Li fonts (handwritten fonts), supports the recognition of variant characters and mixed simplified and traditional characters, and automatically... The system extracts core business fields such as names, ID numbers, and registered addresses from population files; plot numbers, contracted areas, and names of contracting parties from land files; and lease terms and amounts from asset files. Fields with a recognition confidence level below 85% are automatically marked and prompted with pop-up windows, allowing staff to manually correct them by comparing them with the original paper documents. After recognition, the data is organized according to the village code, file type code, and naming rules of year and serial number uniformly formulated by the township, and exported as CSV format data. The system intelligently matches key fields such as plot numbers and ID numbers with existing UAV aerial survey spatial data and business data collected in real time by mobile devices to supplement missing data items in historical files, such as plot boundary association information, population kinship, and contract expiration time. At the same time, the original electronic images are bound and stored with the recognized structured data to form a dual backup of images and data, filling the gaps in early undigitized historical data. The entire process balances the security of archives with the practicality of data and is suitable for the limited technical manpower configuration of township government service centers.

[0028] Step S118: After the data is uploaded to the cloud, the system initiates the correlation verification of spatial data and business data to form an integrated data record of spatial attributes and business attributes; for duplicate land data, the system automatically removes duplicates based on the order of collection time and accuracy priority; fields with large OCR recognition errors are marked as pending review and pushed to staff to complete the correction by comparing the original image online.

[0029] In this embodiment, after data is uploaded to the cloud, the system uses an 18-digit unique plot number as the core association key to precisely bind business data such as population registration (name, ID number, kinship) and contract information (contract term, area, signature information) collected by the mobile terminal with spatial data such as land boundary coordinates and slope aspect generated by the drone. This forms an integrated data record with one file per plot. Each record integrates spatial graphics, business fields, supporting photos, electronic signatures, and other full-dimensional information. For duplicate plots caused by multiple rounds of data collection in rural areas (collected by villages and communities themselves and by townships), and overlap between historical archives and real-time collected data, the system prioritizes data based on accuracy: on-site verification data > none. The system prioritizes aerial survey data over pre-annotated satellite imagery data, prioritizing data collected in the last three months over historical data. It automatically removes duplicates, retaining the best data entries and indicating the reasons for deduplication. For fields with significant OCR recognition errors (such as handwritten contracted area "5.3 mu" being recognized as "53 mu" in historical archives, or the last digit of an ID number being recognized as "0"), the system marks them as pending review based on a confidence level below 85%. This information is then pushed to township data administrators via the cloud management platform. Administrators can retrieve the corresponding electronic images of paper archives online, directly modifying the erroneous fields by comparing them to the originals. After correction, the system automatically generates a log containing the modifier, modification time, and original data, which is then synchronously updated to the integrated data record.

[0030] Step S119: Finally, the verified and corrected spatial data such as land boundaries and topography, along with business data such as population registration and contracting agreements, are uniformly integrated into the basic data resource pool (using OCR technology to batch scan and recognize historical paper archives, extract key information and automatically enter it into the system to achieve comprehensive aggregation of land and population data), and spatial indexes and business field indexes are established to support fast queries by multiple dimensions such as plot number, ID card number, and contract number, so as to achieve comprehensive aggregation and interoperability of the two types of data.

[0031] Spatial data such as land boundary coordinates, slope and aspect, and altitude, after verification and correction, along with ancillary data such as population registration, contracting agreements and supporting photos, electronic signatures, and electronic images of archives, are uniformly integrated into a rural basic data resource pool using a distributed architecture. The resource pool establishes R-tree spatial indexes and B+ tree business field indexes, covering not only plot numbers, ID card numbers, and contract numbers, but also additionally adding frequently used rural query fields such as village / group names, contract start year, and land use type, supporting multi-condition combined queries. Grassroots staff can retrieve corresponding data within one second by entering any index field through the township government platform or mobile app. Entering an ID card number simultaneously displays spatial graphics of all contracted plots, the original contract text, and family member information for each farmer. Entering a plot number allows viewing the household registration details and historical transfer records of the contracting household, achieving seamless interoperability between spatial and business data.

[0032] Step S1191: Collect historical paper archives scattered in village offices, farmers' homes, and old accountants' records in each village. Organize them by land type / population type / asset type + year. Remove metal bindings. Seal moldy, insect-damaged, or torn archives with transparent protective film. Smooth out creases to avoid obstructing the scan. Use an A3 high-speed scanner, set to 300 DPI resolution, and TIFF format for batch scanning. Scan archives with text obscured by binding lines on both sides and then stitch them together to restore the original image. Scan large-format archives in sections and then synthesize complete electronic images. Import the scanned electronic images into the OCR processing system. First, optimize the quality using image enhancement technology: remove mold spots, reduce noise, correct skewed images, enhance the contrast between handwritten text and paper pages, use super-resolution reconstruction technology to improve the clarity of blurred text, and perform image restoration on archives with severe creases.

[0033] In this embodiment, scanned TIFF format electronic images can be batch imported into an OCR processing system adapted for rural historical archives. The system first automatically detects quality issues in each image. For mold spots, a common problem in rural archives, a deep learning-based semantic segmentation model is used to accurately locate the mold area. Pixel-level replacement technology is then used to remove the mold spots without damaging the surrounding text. For noise generated during the scanning process due to paper aging and scanning equipment interference, an adaptive median filtering algorithm is used to filter high-frequency noise while preserving text edge details. For image tilt, Hough transform is used to automatically detect the tilt direction and angle, accurately... The text is leveled to ensure alignment. For handwritten text, local histogram equalization is used to enhance the grayscale difference between the handwritten text and the paper, making the strokes clearer while avoiding over-enhancement that could cause background noise to bounce back. For blurry text, ESRGAN-based super-resolution reconstruction technology is used to double the image resolution, and texture restoration algorithms are used to restore the details of blurry strokes. For files with severe creases, edge detection algorithms are first used to identify the direction and width of the creases, and then image interpolation filling technology is used to repair the text in the crease-covered area while preserving the original texture of the paper to avoid text distortion after repair.

[0034] Step S1192: Load the OCR model adapted to rural archives, start the batch recognition function, and extract fields such as plot number, contracted area, contract period, and name of the contracting party for land-related archives, and extract fields such as name, ID number, registered address, and kinship for population-related archives.

[0035] Load the OCR model pre-trained based on rural historical archive samples. This model optimizes the character feature extraction algorithm for the characteristics of scribbled handwriting and inconsistent formats in rural archives, supports accurate recognition of blurred handwriting, faint handwritten characters, and text covered by red seals. After activating the batch recognition function, the system automatically classifies and processes according to the "land type / population type" archive labels; for land type archives, it first locates the contract title and red seal area to confirm the document attributes, and then accurately extracts core fields such as plot number, contracted area, contracted term, and name of the contracting party; for population type archives, it focuses on extracting fields such as name (compatible with variant characters and rare characters, such as "弢" and "赟"), ID number (with error tolerance for the last digit X and automatic verification of the logical correctness of 18-digit numbers), household registration address (accurately split to the "township, village, group" level, identify handwritten abbreviations of village and group names and complete them), and kinship (adapt to common abbreviations in rural areas, such as "daughter-in-law", "grandson", "nephew", distinguish the association relationship between the household head and family members), etc. During the recognition process, fields with a confidence level lower than 85% (such as scribbled handwritten village and group names, worn ID number digits) are automatically marked, and at the same time, the coordinate positions of the fields in the image are retained.

[0036] Step S1193: The system automatically screens out fields with a recognition confidence level lower than 85%, marks them in orange highlight and generates a review list, and the staff guides and corrects them one by one in accordance with the paper original to ensure that there are no omissions or errors in the key information.

[0037] Conduct a confidence screening on the structured data after OCR recognition. Fields with a confidence level lower than 85% (such as scribbled handwritten village and group names in rural archives, worn and blurred ID number digits, variant characters and rare characters, the name of the contracting party covered by red seals, etc.) are marked in eye-catching orange highlight. At the same time, a review list is generated, which includes the unique archive number, field name, original recognition result, confidence value, corresponding electronic image thumbnail, and the coordinate position of the field in the image. The list is sorted by land type / population type and preferentially displays key fields such as plot number and ID number; at the review workstation in the township government service center, the staff operates in split screen - the electronic review list and the electronically imaged and jump-located image are displayed on the left, and the corresponding paper original is placed on the right. The staff guides and corrects them one by one in accordance with the original. When correcting, the APP supports handwritten input, voice input, and联想 of common rural fields (such as village and group names, kinship abbreviations). It automatically verifies the 18-digit logic correctness of the ID number and automatically matches the unit format of the contracted area. After the correction is completed, the confirmation correction button needs to be clicked. The system automatically records the content before and after the correction, the name of the corrector, and the correction time, forming a traceable operation log; for batch-occurring similar errors (such as "Group 3" in a certain village being uniformly recognized as "3 groups"), batch replacement correction is supported, which greatly improves efficiency. At the same time, a secondary sampling inspection mechanism is set up, and the team leader randomly selects 10% of the corrected data to verify against the original to ensure that there are no omissions or errors in key information such as plot number, ID number, and contracted term.

[0038] Step S1194: Process the identified structured data according to a unified standard: convert the area unit to mu (a Chinese unit of area), make the address accurate to the township, village, and group, and make the date uniformly in YYYY-MM-DD format, and assign a unique identification code to each data entry.

[0039] Regarding area units, it is compatible with common rural archives using mu (a unit of land area), fen (a unit of area), square meters, and handwritten expressions. It automatically converts to 1 mu = 666.67 square meters and 1 fen = 0.1 mu, ultimately retaining two decimal places in mu. For address information, it automatically completes the township, village, and group levels by referencing the township administrative code database, standardizing abbreviations to standard names, and supplementing missing village / group information by linking collected household registration or land parcel information. The date format covers various expressions found in rural archives, such as 1998.5.10, May 10, 1998, October 5, 2000, and 2000-10-5. After dynamic identification, the data is uniformly converted to the YYYY-MM-DD standard format. For data that only indicates the year but not the month and day, it is temporarily stored as YYYY-01-01 and marked as to be supplemented later. At the same time, a unique identification code is assigned to each data entry. The coding rule is 6-digit township administrative code + 3-digit village code + 1-digit archive type code (land = 1, population = 2, asset = 3) + 4-digit year of formation + 5-digit serial number. This ensures that each data entry is uniquely identified and can be traced back to the archive source. The entire formatting process is completed automatically by the system. Only information that cannot be automatically completed is marked with a prompt. Staff members supplement the information with paper originals to complete the final formatting, ensuring that the data format is uniform and standardized.

[0040] Step S1195: Using key fields such as land parcel number and ID number, the formatted historical data is intelligently matched with the spatial data generated by UAV aerial survey and the business data collected in real time by mobile terminal to complete the missing data items such as land parcel boundary association information and population kinship in the historical archives.

[0041] Using land parcel numbers (compatible with mapping and conversion between historical handwritten codes and current standard codes, e.g., mapping No. 12, Group 3, XX Village to a 6-digit township code + 3-digit village code + 003 + 0012) and ID card numbers (historical population data without ID card numbers, using the combination of household registration number + name + registered address as the association key) as the core, a multi-dimensional intelligent matching engine is built. First, the formatted historical data is accurately correlated with spatial data generated by UAV aerial surveying. By matching the corresponding boundary coordinates, slope, altitude, and other spatial information through the land parcel number, missing land parcel vector data and terrain attributes in historical archives are supplemented. Then, it connects with real-time business data collected from mobile terminals, linking the latest household registration information, electronic signatures, and on-site supporting photos through the ID card number, simultaneously verifying the deviation between historical contracted area and aerial survey measured area; targeting historical population... For missing kinship relationships in the archives, the household head's identity information is extracted through ID card numbers. Combined with family member lists and household registration information collected by mobile devices, the system intelligently completes the kinship relationships of the household head, spouse, children, parents, etc. Multiple historical records under the same household address are automatically merged into a complete household file. For special data such as fuzzy plot numbers and missing ID card numbers, a fuzzy matching method is used with the combination of "name + household address + contract years". Combined with historical ledgers provided by the village committee, the system assists in confirming the relationship and completes data items such as unmarked plot boundaries, incomplete kinship relationships, and unlinked spatial locations in the contract in the historical archives. At the same time, the system automatically marks the matching confidence level (≥90% directly entered into the database, 60%-90% manually confirmed, <60% returned for supplementary verification) to ensure the logical consistency between historical data and existing spatial and business data.

[0042] Step S1196: Bind the structured data to the corresponding electronic image originals one by one, encrypt them with AES-256 and import them in batches into the rural basic data resource pool, simultaneously establish data logs, generate offline backup packages and store them on the local server to ensure that the data is traceable and not lost.

[0043] By using a unique identifier for each data entry, a strong association is established between the formatted structured data and the corresponding TIFF format electronic image original. Clicking on any data entry allows for real-time retrieval of the complete electronic image, facilitating subsequent verification and tracing. The associated data packets are then fully encrypted using the AES-256 symmetric encryption algorithm (the encryption key is managed and kept by two designated personnel at the township government service center, and changed quarterly). After encryption, the data is uploaded in batches via dedicated line or encrypted Wi-Fi to the rural basic data resource pool. The resource pool automatically adapts to importing data in multiple formats such as CSV and shp, simultaneously verifying data integrity and relevance. Data packets lacking relevant information will trigger a pop-up prompt for staff to complete them. Meanwhile... The system automatically establishes an immutable data log, recording detailed information such as data import time, operator's name and employee number, data source (specific village / group, year of record), encryption status, and import result (reason for success / failure), facilitating subsequent auditing and verification. In response to unstable networks and sudden power outages in rural areas, it synchronously generates encrypted offline backup packages (compressed files with the .enc extension, named according to the backup date and batch number), storing them on a local dual-hard drive RAID5 array server in the township government center. These backup packages are automatically synchronized to village-level backup storage devices weekly, and the integrity of the backup packages is manually verified and recorded monthly by designated personnel, ensuring the security and controllability of sensitive data such as land ownership confirmation and population registration throughout the entire transmission, storage, and backup process.

[0044] Step S1197: Perform final verification on the incoming data. After confirming that there are no errors, the data is aggregated to form a complete archive system of electronic images, structured data, and associated spatial data through field logic verification and deduplication.

[0045] A final verification mechanism is initiated for the incoming data. First, outliers are screened through field logic checks: Contract periods are automatically compared to rural land contracting policies (e.g., second-round contracts do not exceed 30 years), and those exceeding this range are marked as anomalies; plot area values ​​are verified for rationality (excluding negative numbers and extreme values ​​exceeding the total area of ​​the region); ID numbers are automatically verified according to the 18-digit coding rules (first 6 digits are the administrative region code, the middle 8 digits are the birth date, and the last digit is the check digit); logical contradictions in kinship relationships are investigated (e.g., conflicting identities between a son and father within the same household); for duplicate data, duplicate entries are identified by matching key field combinations (plot number + ID number / contract number), and then prioritized by accuracy: field-collected data > UAV aerial survey data > historical files. Case data; time priority: data from the last 3 years > earlier data. The best record is automatically retained. After deduplication, historical versions are archived synchronously and the reason for deduplication is marked. After the data that has passed the verification is confirmed by the township data administrator by randomly checking 10% of the samples, the final aggregation is completed to form a complete archive system with a unique identification code as the link. Each file is simultaneously linked to TIFF format electronic images (which can view the details of the original paper file), standardized structured data (including population, contracts, and land parcel attributes), and spatial vector data generated by UAV aerial survey (which can view boundary graphics and terrain parameters). The three are linked and queried in real time, which not only meets the business needs of rural land rights confirmation and population management, but also provides full-chain support for subsequent data updates and traceability.

[0046] Step S120: Construct an AI and RAG intelligent cleaning model, use generative adversarial networks to fill in missing data values, identify abnormal data such as duplicate population registration and land area conflict through the local anomaly factor algorithm, train an entity recognition model based on the rural government affairs vertical RAG library, and automatically label data attributes and relationships.

[0047] In this embodiment, an AI and RAG intelligent cleaning model tailored to the characteristics of rural government data is constructed. Supported by a rural government vertical RAG library covering land contracting policies, household registration management regulations, village and group administrative codes, and historical archive format standards, the model first uses a generative adversarial network to selectively fill in missing data values. For missing fields such as contract years and kinship in historical archives, it references the distribution characteristics and policy constraints of similar archives from the same village and group in the same period (e.g., the second round of contracting is filled with a default of 30 years) to generate reasonable values ​​that conform to rural realities, avoiding results that are out of touch with business scenarios. Then, the Local Anomaly Factor (LOF) algorithm is used, combined with the distribution patterns of rural data to set an adaptation threshold, automatically identifying abnormal data such as duplicate registration of the same ID number across villages / groups, area conflicts caused by multiple confirmations of the same plot, contracting periods exceeding policy limits (e.g., exceeding 30 years), and mismatches between household registration addresses and the village / group to which the plot belongs. Simultaneously, the model labels the anomaly type and confidence level. Finally, an entity recognition model trained on the RAG library... It can accurately identify abbreviations, variant characters, and handwritten variations unique to rural areas, automatically label core data attributes such as plot numbers, ID numbers, and contracted areas. Based on policies, rules, and historical cases, it can automatically establish the ownership relationship between farmers and contracted plots, the kinship relationship of family members, and even the corresponding rural government policy clauses. The entire process transforms the cleaning steps that originally required more than 80% manual intervention, such as missing value filling, anomaly screening, attribute labeling, and relationship association, into automated algorithm processing. Only less than 5% of complex issues are retained, such as scattered historical files without any key fields, conflicts of interest among multiple households fighting for the same plot without policy basis, and cases with serious data logic contradictions that cannot be judged by the algorithm. These issues are manually reviewed, and the algorithm automatically classifies and labels complex issues, synchronously linking reference policies, similar processing cases, and original electronic images in the RAG library. This reduces the difficulty of manual review, adapts to the actual scenario of insufficient technical manpower at the grassroots level in rural areas, and significantly improves the efficiency and accuracy of data cleaning.

[0048] Step S121: Construct a rural government affairs vertical RAG database, collect rural land contracting case data sources, remove outdated and duplicate content, and classify and organize them into policy, coding, case, and format categories. Perform text preprocessing using jieba word segmentation, and then use the Sentence-BERT model to convert the text into vector embeddings and store them in the Milvus vector database. At the same time, build a keyword search engine to support fast association queries by policy clauses, village / group names, and data types.

[0049] First, collect data sources such as rural land contract policy documents (including policies related to the second round of contracting, confirmation of rights and issuance of certificates, and local implementation rules), household registration management norms (household registration, standards for determining kinship), the latest administrative coding table of village and group levels, historical file format standards, and case files on land right confirmation disputes over the years from units such as township government service centers, county agriculture and rural affairs bureaus, and natural resources bureaus. At the same time, supplement and collect the对照表 of common abbreviations, variant Chinese characters, rare Chinese characters, and local dialect vocabulary libraries in rural areas. Subsequently, manually screen and remove obsolete policies that have been repealed (such as early regulations on land contract years), duplicate coding tables and cases, and conduct refined classification and sorting according to policy categories (including four levels of central, provincial, municipal, and county levels), coding categories (current coding + historical coding mapping), case categories (subdivided by dispute types), and format categories (classified by file years and types). Mark the effective date and applicable regions for policy documents, and supplement key information such as village and group affiliations and handling basis for cases. Then, use the jieba word segmentation tool adapted to rural government affairs scenarios to segment the text, remove meaningless stop words such as "的" and "了", and extract core semantic units. Next, use the pre-trained Sentence-BERT model to convert the processed text into 768-dimensional vector embeddings and store them in batches in the lightweight Milvus vector database according to classification labels. Finally, build a keyword retrieval engine that supports multi-dimensional combined queries. In addition to basic retrievals of policy clauses, village and group names, and data types, add fuzzy query functions and regional screening functions, and set sorting rules for retrieval results to ensure that grass-roots staff can quickly associate and query the required information without professional retrieval skills.

[0050] Step S122: Preprocess and standardize the original data, unify the field names, align the data formats, remove redundant characters, extract key information segments through regular expressions, form a standardized data set recognizable by the model, and at the same time divide it into training set, validation set, and test set.

[0051] First, we integrated raw data from multiple sources, including structured historical archives, UAV aerial survey data, and business data collected from mobile devices. Addressing the dispersed nature and disorganized formats of rural data, we standardized core field names, renaming land parcel codes, parcel serial numbers, and land parcel numbers as "land parcel number," and unifying contractor and household head names as "contractor household head." Contract period and contract term were merged into "contract term," with field types clearly defined simultaneously. When aligning data formats, date fields (e.g., "1998.5.10," "May 1998," "October 2000") were uniformly converted to the YYYY-MM-DD standard format. Missing month and day fields were temporarily stored and marked as YYYY-01-01. Area fields were uniformly converted to values ​​in "mu" (acres). Spatial data coordinates were standardized to the WGS84 coordinate system, and redundant decimal places were removed. Batch processing tools were used to remove erasures, irrelevant notes, leading and trailing spaces, and special redundant characters such as "***" and "—" from handwritten archives. For unclear characters, mark them as unrecognizable; use regular expressions to accurately extract key information fragments—use "^\d{18}|\d{17}(\d|X|x)$" to extract ID card numbers, use "[village group]\d{1,3}[plot number]?" to extract plot numbers, and use "\d+(.\d+)?(mu|fen|mianfang)" to extract area information, ensuring that no key information is omitted; integrate the processed data into spatial data (shp format) and business data categories, unify the field order, fill in missing markers for required fields, and form a standardized dataset that the model can directly recognize; finally, divide the training set, validation set, and test set into a 7:2:1 ratio using random sampling by village group to ensure that each dataset covers samples from different eras, different village groups, and different data types, avoiding model training bias due to uneven data distribution, while keeping a dataset partitioning log to record the sampling ratio, the scope of village groups involved, and the amount of data.

[0052] Step S123: Train a generative adversarial network adapted to rural data. Input samples with missing values ​​and construct a loss function by combining statistical characteristics and policy constraints of archives from the same village group and the same era in the RAG library. The discriminator is responsible for distinguishing the rationality of generated data from real data. The model parameters are optimized through iterative training. After training, missing fields are filled in.

[0053] When training a generative adversarial network adapted to rural data, the generator adopts a lightweight U-Net architecture. Input samples containing missing values ​​such as contract period, kinship, and land parcel attachment information are used. By accessing the rural government affairs vertical RAG database, statistical characteristics of archives from the same village group and era are extracted (e.g., the second round of contract periods are concentrated at 30 years, land parcels are larger by 0.5-10 mu, and kinship is mainly "head of household-spouse-children"). A composite loss function is constructed based on land contracting policy constraints to ensure that the generated values ​​closely match the data distribution and comply with business rules. The discriminator uses real and complete rural archive data as training samples, and... The model learns the rationality characteristics of rural data and distinguishes the difference between generated data and real data through probability output. During training, the model parameters are iteratively optimized in batches. Every 10 rounds, the accuracy of filling in missing data is tested with a validation set. If the accuracy is lower than 90%, the model is backtracked to the RAG library to supplement samples and retrained until the model is stable. After training, missing fields are automatically filled in. For example, missing contract period is filled in according to the mainstream years of the same village group, and missing kinship is deduced by referring to the information of other members of the same household. The filling results need to be verified by the RAG library policy to ensure that the generated data does not deviate from the actual business scenario in rural areas and avoid logical contradictions or policy conflicts.

[0054] Step S124: Adapt the local anomaly factor algorithm for anomaly identification. First, calculate the local density threshold based on the normal data in the training set. Then, input the preprocessed dataset into the algorithm to automatically identify abnormal data and label the risk level according to the severity of the anomaly.

[0055] When using a local anomaly factor algorithm adapted to the distribution characteristics of rural data for anomaly identification, it first uses normal data without logical contradictions in the training set as a foundation. It then calculates a local density threshold based on actual rural business scenarios, referencing the distribution patterns of core data such as plot area, contract period, and household registration affiliation to eliminate interference from extreme reasonable values, ensuring the threshold aligns with the characteristics of rural data. The neighborhood size is set to 50 to suit the scale of village-level data. After inputting the standardized dataset into the algorithm, it automatically screens for multiple types of anomalies: duplicate population registration, land area conflicts, abnormal contract periods, and village / group affiliation conflicts. Simultaneously, it labels the risk level according to the severity of the anomaly: high risk (area deviation exceeding 20%, duplicate registration involving multiple plot disputes, contract period exceeding 5 years), medium risk (deviation 5%-20%, village / group coding errors, contract period exceeding 1-5 years), and low risk (inconsistent abbreviations, minor format errors that can be automatically corrected). This facilitates priority processing by staff, reducing false positives and accurately identifying anomalies requiring key verification, thus meeting the needs of rural grassroots data verification.

[0056] Step S125: Train the entity recognition model based on the RAG library, add labeled samples to the training set, fine-tune the model parameters through transfer learning, input standardized data after training, automatically label the core entity attributes, the labeling accuracy needs to reach a set ratio, and feed back samples that do not meet the standard to the RAG library to supplement labeled samples, and iteratively optimize the entity recognition model.

[0057] When training the entity recognition model based on the RAG library for rural government affairs, a lightweight BERT-base model is selected as the basic architecture. The training set is supplemented with labeled samples of variant characters, rare characters, village group abbreviations ("three groups", "three teams", "village committee" corresponds to "village committee"), handwritten variants (cursive script, simplified characters), and commonly used local expressions, to build a dedicated labeled dataset that fits the rural archive scenario. Relying on the policy terms, encoding rules and other knowledge in the RAG library, the model parameters are fine-tuned through transfer learning, focusing on optimizing the feature extraction capabilities of handwritten characters and variant characters, and solving the core problems of non-standard handwriting and inconsistent expressions in rural archives. After training, standardized data is input, and the model can automatically and accurately label core entity attributes such as plot number, ID number, contracted area, and kinship. It can also initially identify implicit relationships between fields (such as the matching between the household registration area implied by the ID number and the village group of the plot). The labeling accuracy needs to be consistently above 95%. Samples that do not meet the standard (such as extremely illegible handwriting or rare variant characters) will be automatically fed back to the RAG library to supplement the labeled samples and update the model training set in sync. Through multiple rounds of iteration, the model's recognition accuracy in special rural data scenarios will be continuously improved to meet the actual operational needs of grassroots units without professional technical personnel.

[0058] Step S126: Construct a graph of relationships between people, land, and contracts using a graph neural network, linking farmers' ID numbers with contracted land plot numbers, land plot numbers with spatial boundary data, and family members' ID numbers with kinship tags.

[0059] When establishing an automatic matching mechanism for related relationships, the mechanism is based on entity recognition results and deeply integrates rural land contracting policies and rules and historical cases of land rights confirmation in the RAG database. A three-dimensional relationship map of people, land, and contracts is constructed through a lightweight graph neural network. Farmers, contracted land plots, and contracts are used as core nodes, and contracting, ownership, and kinship are used as related edges to automatically achieve multi-dimensional and accurate association: the farmer's ID number is associated with the contracted land plot number according to the coding mapping rules in the RAG database; the land plot number is bound to the spatial boundary data of the WGS84 coordinate system of UAV aerial survey according to the coordinate matching logic; and the ID numbers of family members are associated with the corresponding kinship labels (head of household, spouse, children) according to the kinship identification standards (such as the same household registration address + age gradient). Meanwhile, the system automatically verifies the rationality of the association logic: it investigates issues such as minors being contractors, mismatches between ID numbers and registered areas and village groups, and conflicts in kinship (such as the same person being marked as both "father" and "son"). Samples that fail to be associated (such as those without matching key fields or irreconcilable logical conflicts) are directly marked as pending review and are simultaneously associated with similar cases in the RAG library for subsequent manual reference, adapting to the actual needs of rural archive data that is not standardized and has complex association scenarios.

[0060] Step S127: Integrate AI models and RAG libraries to build an automated cleaning system. Set up an automated execution chain according to the process of data preprocessing, GAN missing value imputation, LOF anomaly identification, entity annotation, and relationship establishment. The system automatically reads input data, calls the corresponding model, outputs cleaning results, and generates a cleaning report simultaneously.

[0061] A lightweight, automated data cleaning system is built by integrating AI models with a rural government affairs vertical RAG library. The system is adapted to the computing power of township-level servers and supports batch import of multi-source format files, including historical archive structured data, UAV aerial survey spatial data, and mobile terminal business data. It automatically executes data in a fixed chain: data preprocessing, GAN missing value imputation, LOF anomaly identification, entity labeling, and relationship establishment. First, it automatically completes preprocessing such as field unification, format alignment, and redundant character removal. Then, it calls the trained GAN model to impute missing values ​​such as contract period and kinship. Next, it uses the LOF algorithm to screen for abnormal data such as duplicate registration and area conflicts and labels the risk level. Then, it starts the entity recognition model to label the core field attributes. Finally, it relies on graph neural networks to establish relationships between people, land, and contracts and verify the logic. The entire process requires no manual intervention. The system automatically reads data, calls the corresponding model interface, integrates the cleaning results, and generates a visual cleaning report simultaneously. It clearly records key information such as the amount of data processed at each stage, the number and accuracy of missing value filling, the types and proportion of abnormal data, entity annotation accuracy, and the success rate of association. The report supports exporting to PDF format, which makes it easy for grassroots staff to quickly grasp the cleaning results. It is suitable for the actual needs of rural grassroots areas with limited technical manpower and the need to efficiently process massive amounts of archival data.

[0062] Step S128: Set up a manual review mechanism. The system automatically filters out complex problem samples, sorts them by risk level, and pushes them to staff. At the same time, the review interface displays reference policy clauses, similar handling cases, and original electronic images from the RAG library. The data corrected by the staff is automatically fed back to the model training set.

[0063] The system automatically filters out complex problem samples in rural data scenarios—including scattered historical files without plot numbers / ID numbers, conflicts of interest involving multiple households vying for the same plot without clear policy basis, land area deviations exceeding 20%, or serious contradictions in kinship logic that the algorithm cannot determine. These are prioritized as high-risk (disputes, significant data deviations), medium-risk (fuzzy coding mapping, questionable policy compatibility), and low-risk (multiple fields missing requiring manual supplementation) before being pushed to township data review staff. The review interface uses a split-screen design; the left side clearly displays the anomaly type and current data status of the problem sample, while the middle... The system retrieves relevant rural land contracting policy clauses and similar cases from the RAG database in real time, and loads original electronic images on the right to reduce the difficulty of judgment for grassroots staff. After the staff completes the correction, the system automatically feeds back the corrected data and problem labels to the model training set, supplements exclusive labeled samples, and continuously iterates and optimizes the adaptability of models such as GAN missing value imputation and LOF anomaly recognition to rural scenarios. At the same time, it regularly counts the proportion of data with manual intervention, and gradually reduces the scope of intervention by improving model accuracy, ultimately keeping the proportion of manual intervention stable below 5%, balancing data cleaning efficiency and the feasibility of actual operation at the grassroots level.

[0064] Step S1261: Define the map entities and relationship types, set people, land, and contracts as core entities, clarify the four types of relationships: contractor (people, land), bearer (land, contract), signer (people, contract), and kinship (people, people), and incorporate the unique historical coding and current coding mapping rules of rural areas to address the problem of inconsistent coding in the archives.

[0065] When defining the core entities and relationship types of the map, the "Person" entity explicitly includes attributes such as ID number, name, registered address, date of birth, and contract qualification status, adapting to the actual situation of some elderly people without ID cards or whose registered addresses have changed; the "Land" entity covers attributes such as plot number (recording both historical handwritten codes, such as No. 12, Group 3, XX Village, and current standard codes), WGS84 coordinate range, measured area, plot use, and village group, compatible with differences in plot descriptions from different eras; the "Contract" entity includes attributes such as contract number, contract period, contracting party (full name of the village collective), signing date, and contract status (valid / expired), covering the information recording characteristics of printed and handwritten contracts. Four types of relationships are precisely adapted to rural business scenarios: the contracting relationship clarifies the ownership of the contracting rights between people and land, based on the rights and obligations stipulated in the contract; the bearing relationship defines the correspondence between land and contract, i.e., a certain plot is contracted in detail by a certain contract; the signing relationship reflects the signing relationship between people and contracts, marking the head of the contracting household as the contracting party; and the kinship relationship is further subdivided into specific types such as head of household, spouse, children, parents, and siblings, adapting to the characteristics of rural family structures. At the same time, a coding mapping rule library unique to rural areas was established, which includes the correspondence between historical codes and current codes, embeds graph association logic, and automatically matches entities with different coding formats, effectively solving the core problems of inconsistent coding and the inability to directly associate historical codes with current codes in rural archives.

[0066] Step S1262: Extract entity features and association clues. Extract key attributes of people, places and contracts from structured data, extract implicit association clues through regular matching and entity recognition models, and transform the features into vector forms that can be recognized by graph neural networks.

[0067] This research precisely extracts key attributes of people, land, and contracts from multi-source structured rural data. For people, core information includes ID number (or household registration number + name combination if no ID card is available), name (compatible with variant characters and rare characters), registered address (refined to village / group level), and date of birth. For land, attributes include historical handwritten codes, current standard plot numbers, WGS84 coordinate range, measured area (in mu), and village / group affiliation. For contracts, key fields include contract number, contractor's name / ID number, associated plot number, contract period, and contracting party's name. Furthermore, addressing the implicit associations in rural archives, regular expression matching is used to uncover textual clues, such as matching contract remarks to phrases like "contracted plot No. 5, Group 3, Village XX," "head of household XXX," and "children XXX," thus identifying relationships between people, land, and other people. An entity recognition model adapted to rural scenarios is used to identify implicit associations arising from handwritten variant characters and village / group abbreviations, such as extracting the correspondence between contractors and plots from fuzzy contract text. Finally, a Sentence-BERT model fine-tuned with rural government affairs corpus was adopted to transform all entity attributes and association clues into 768-dimensional vectors, ensuring that the feature vectors can accurately reflect the semantic associations and business logic of rural data and adapt to the input requirements of graph neural networks.

[0068] Step S1263: Construct an initial graph structure, with entities as nodes and association clues as edges, and automatically establish initial connections: associate the person with the contract through the contractor's ID number in the contract, associate the contract with the land through the land parcel number in the contract, and associate the kinship between people through the registered address and age difference, forming graph data containing node attributes and initial edges.

[0069] When constructing the initial graph structure, people, land, and contracts are first transformed into nodes with rural-specific attributes: People nodes are supplemented with household registration numbers and name combinations for farmers without ID cards; land nodes simultaneously carry historical handwritten codes and current standard codes; and contract nodes indicate the handwritten / printed source of the contractor's information. Next, connections are automatically established according to the characteristics of rural data: when associating people and contracts, if there are handwritten errors in the contractor's ID number in the contract, the association is confirmed by matching the first 17 digits and using the name as an auxiliary method; when associating contracts and land, the abbreviation of the land parcel in the contract is automatically mapped to the current code before establishing the connection, resolving the incompatibility issue between historical and current codes; when associating people and their kinship, people nodes with the same village / group household registration address are first filtered, and then determined according to common age differences in rural families (age difference between head of household and children 20-40 years, age difference between head of household and spouse ≤10 years), initially establishing head of household-children and head of household-spouse relationships, with abnormal age differences temporarily marked pending confirmation. The resulting graph data includes nodes with complete attributes (such as a person's household registration and the coordinates of the land), and edges labeled with association types (signature, contract, etc.) and initial confidence levels (0.9 for exact matching and 0.6 for fuzzy matching), adapting to the real-world scenario where rural data is not standardized.

[0070] Step S1264: Train a graph neural network to optimize associations. Use GCN or GAT architecture, input initial graph data, train the model with correctly labeled association samples, and strengthen the semantic associations between entities by updating node embedding vectors.

[0071] When training the graph neural network to optimize associations, considering the computing power limitations of rural grassroots servers, the lightweight GAT architecture is prioritized. Its attention mechanism can accurately focus on the core association attributes of people, land, and contracts, avoiding redundant calculations. After inputting the initial graph data, the model is trained using rural-specific labeled samples: the labeled samples cover correct association cases such as relatives in the same village group jointly contracting land, matching historical handwritten codes with current codes, and minors having no contracting rights, as well as abnormal cases such as erroneous associations caused by ambiguous ID numbers and conflicts in the village group ownership of land plots, ensuring coverage of common rural data scenarios. During training, the node embedding vector is updated to incorporate rural-specific attributes such as people's household registration village group, historical codes of land, and the contracting party into the vector representation, strengthening the semantic association between entities; at the same time, the model autonomously learns rural business patterns, such as relatives in the same village group often jointly sharing contracting rights, the first 9 digits of the land plot code corresponding to the township, village, and group levels, and the contract signing dates mostly concentrated in the policy implementation year, improving the ability to judge fuzzy associations. After each round of training, the model is validated using real archive data from a certain township. If the historical code matching accuracy is lower than 90%, the code mapping samples of that township are added and the model is retrained.

[0072] Step S1265: Use the trained model to predict new associations and verify the logic. Mark abnormal associations for manual review and correction. At the same time, support the addition of new data to the graph through incremental learning.

[0073] During the graph verification phase, after the trained model predicts new associations, it automatically verifies them using rural administrative rules: it checks for irregular associations such as minors being contractors, discrepancies between the confirmed land area and the aerial survey area exceeding 15%, and mismatches between the registered village group and the village group to which the land belongs. It also identifies kinship conflicts where the same person is both the head of household and a child, marking these as high-risk (involving land disputes), medium-risk (coding mapping deviation), and low-risk (format errors). During manual review, the system pushes corresponding policy clauses from the RAG library (such as the contract qualification provisions of the Rural Land Contract Law), similar association cases from the same village group, and original electronic images (supporting zooming to view handwritten contract details) to help grassroots staff quickly determine and correct errors. In the dynamic update phase, newly added data (such as changes in kinship after household division, newly signed contracting agreements, and land transfer records) does not require retraining the model; it is directly integrated into the graph through incremental learning, automatically updating the association edges of people, land, and contracts. For supplementary historical archive data, the system matches existing node attributes to complete the associations; for example, it maps the old well land in XX village in the old archives to the current land parcel number through coding mapping.

[0074] Step S1271: Build a parsing tool adapted to the rural data format, read the input data of Excel account books and OCR results of paper scans, split the fields by type, and uniformly convert them into UTF-8 encoded JSON format; for handwritten variant characters and date abbreviations, call the pre-trained text normalization model to batch correct and output standardized intermediate data.

[0075] The built parsing tool is deeply adapted to the chaotic characteristics of rural data: for the Excel account books manually entered by village accountants, it supports automatic recognition of multiple version formats (such as 2003 version.xls, 2016 version.xlsx), and can handle issues such as merged cells (e.g., the merged item of household head + family members is automatically split into independent fields) and disordered headers (e.g., the plot acreage and field acreage are unified into the contract area field). For the OCR results of paper file scans, a noise reduction module is built-in to correct the recognition errors caused by paper wrinkles and scribbled handwriting (e.g., correcting Li Renyao to Li Yao, and automatically calibrating the misrecognition of DK502 as DK5O2). When splitting fields, it accurately matches the rural scenario, extracts core items such as ID card numbers, names, and household registration village groups for population data, and splits exclusive fields such as plot numbers, four boundaries, and contract years for land data. The text normalization model is fine-tuned with local corpus (importing 20,000 villagers' names and 500 village group names in the township), and batch corrects issues such as Zhang Erzhuan to Zhang Xiao, 23.8.15 to 2023-08-15, and 05 years being completed as 2005. Finally, UTF-8 encoded JSON is output, and "to be supplemented" is marked when fields are missing, adapting to the requirements of the subsequent cleaning module.

[0076] Step S1272: Train a rural scenario GAN model based on historical complete data. For the missing kinship in population data and the partially missing four boundaries in land data, input the associated features of the missing fields, and the model generates logical filling values, which are verified by rules after filling.

[0077] Using the complete confirmation files and village collective account books of the township in the past 5 years as training corpus, optimize the GAN model to adapt to the rural data logic. When kinship is missing, input associated features such as the household head's ID card number, household registration village group, and age difference between family members. For example, if the household head is 45 years old, the missing member is 20 years old and has the same household registration, the model refers to the common relationships of the same age group in the same village and generates reasonable results such as "father / son" or "father / daughter". When the four boundaries are partially missing, associate the coordinates of adjacent plots and the boundary range of the village group. For example, if the east boundary of plot DK301 is missing, combine the boundary trends of DK302 on the west side and DK303 on the north side to generate coordinate values that fit the terrain. After filling, trigger rule verification: the kinship needs to meet the age difference threshold (grandparent-grandchild ≥ 25 years), the four boundaries must not exceed the administrative scope of the village, and the coordinate deviation ≤ 0.5 meters to ensure that the filling results conform to the actual rural confirmation scenario.

[0078] Step S1273: Set the rural data anomaly threshold, call the LOF algorithm to calculate the outlier score of the data, mark the scores that exceed the threshold as anomalies, and synchronously compare with historical data to label the anomaly type.

[0079] Specific anomaly thresholds are set based on the characteristics of rural data: land area is based on the average of 3-10 mu per plot in the village; exceeding 5 times or falling below 0.3 mu is set as a threshold; the first 6 digits of the ID number must match the administrative code of the registered residence; mismatch triggers an alert; contract periods exceeding 30 years (the legal upper limit for rural contracted land) and age differences between relatives less than 15 years (excluding unreasonable parent-child relationships) are also included in the threshold range. The LOF algorithm is used to calculate outlier scores; scores ≥0.8 are marked as anomalies, and the data is simultaneously compared with historical land ownership data from the past 3 years: for example, if plot DK207 is currently entered as 80 mu, but historical data shows 6 mu, the value is marked as entered incorrectly; if a villager's ID number has the first 6 digits from another city, the logic is marked as conflicting, providing a clear direction for subsequent manual verification.

[0080] Step S1274: Deploy the entity recognition model to automatically label core entities such as owners, land parcels, and contract numbers; establish association relationships based on the rule engine; for conflicts in multi-source data, call historical name comparison records in the RAG library to assist in matching and generate association confidence.

[0081] A BERT entity recognition model, finely tuned from a township corpus (20,000 villager names, 800 land parcel codes, and 5,000 contract texts), is deployed to accurately label core entities: property owners are distinguished as head of household / co-owners; land parcels are associated with a number + village / group location (e.g., DK012-Li Village Group 3); and contract numbers are extracted using a year + village / group + sequence number structure (e.g., 2024-LC3-008). The rule engine pre-sets association logic (e.g., binding contract numbers to corresponding property owners and land parcels). In case of multi-source conflicts (e.g., Zhang San in the system versus Zhang San and Zhang Lao San in the land ownership records), it calls historical comparison records from the RAG database (including records of former names and aliases signed by villagers during previous land ownership confirmations) to assist in matching and generate association confidence scores: those above 95% are automatically merged; those below the threshold are supplemented with scanned copies of historical signatures, village committee certificates, and other evidence for manual verification.

[0082] Step S1275: Connect the modules through the workflow engine, set the trigger conditions, and automatically call the model in the order of preprocessing, filling, identification, annotation, and association; after cleaning is completed, generate a report, which includes the total amount of cleaning, missing filling rate, number of exceptions handled, association matching results, and typical cases.

[0083] The lightweight chemical workflow engine connects various modules, with the trigger condition set to automatically read newly confirmed land rights data at 3:00 AM daily and trigger immediately when village administrators manually upload ledgers. This adapts to the scenario of scattered rural data reporting. The model is automatically called in the order of preprocessing, filling, identification, labeling, and association, without manual intervention. After cleaning, a visual report is generated, which clearly shows the total amount of data cleaned (e.g., 820 population data entries and 510 land data entries), the missing data filling rate (e.g., kinship filling rate of 85% and boundary filling rate of 78%), the number of exceptions handled (e.g., 23 data entry errors and 8 logical conflicts), and the association matching results. Typical cases are attached (e.g., "The area of ​​plot DK091 was mistakenly recorded as 50 mu, which was corrected to 5 mu after comparison with historical data; Zhang San and Zhang San were merged with 96% confidence"). The report can be exported to Excel for easy archiving and verification by the township land rights office.

[0084] Step S130: Referring to the agricultural data standard system, compile a unified data dictionary for land classification coding and population information fields, clarify the data format, accuracy requirements and association rules, and use tools to batch convert raw data from different sources and formats into standard formats, construct a knowledge graph of land, population and assets, and realize the interconnection and interoperability of land, people and things data.

[0085] In this embodiment, in the rural land ownership confirmation scenario, after processing according to a unified data standard, township staff no longer need to manually check data in different formats. The original handwritten land ledgers (Excel spreadsheets, paper scans), UAV aerial survey data in shp format, and provided household registration Excel data are all converted into a standard format in batches using an ETL tool: the land classification code is coded according to the dictionary to correspond to agricultural cultivated land-sloping cultivated land and cultivated land-terraced fields, the population information fields are unified with the head of household name and family member relationship format, the ID number is forced to have 18 digits for verification, and the land area is retained to 2 decimal places. When villagers apply for homestead approval, the land-population-asset knowledge graph automatically links to their family members, the location and area of ​​their existing contracted land plots by entering the applicant's ID number, verifying compliance with the "one household, one homestead" policy and avoiding duplicate applications. In village collective asset inventory scenarios, the knowledge graph can link collective forest land codes, information on leasing farmers, and rent payment records. During the inventory, the system automatically compares the forest land area with the lease agreement and verifies whether rent has been paid on time, eliminating the need for staff to manually verify multiple sets of land, financial, and population data across departments. Furthermore, for farmers' land transfer needs, the standardized data-linked graph allows for quick queries of land ownership, transfer history, and transferee credit records, shortening the processing time for transfer procedures.

[0086] Step S131: Compare with higher-level standards, collect documents issued by the agricultural sector, extract land classification codes, core fields of population information, and data format benchmarks, establish a standard benchmarking list, and clarify the mandatory clauses to be followed and the flexible space for adjustment.

[0087] Download the latest versions of the "Database Specifications for Confirmation and Registration of Rural Land Contractual Management Rights" and the "Data Collection Specifications for Rural Collective Economic Organizations" from the official agricultural website, and simultaneously collect the supporting implementation rules issued by the Provincial Department of Agriculture and Rural Affairs. Focus on extracting land classification codes (e.g., primary code 200 for cultivated land, secondary code 210 corresponding to paddy fields), core fields of population information (12 mandatory fields including ID number, name, and registered address), and data format standards (date in YYYY-MM-DD format, area in mu with two decimal places). Establish a benchmark list based on field names, superior standards, and local adaptation instructions, clarifying mandatory clauses (e.g., 18-digit ID number verification, 19-digit land parcel code rules), and defining flexible space (e.g., former names are optional fields for superiors, but are made mandatory when there are many local archives; for reclaimed land without superior codes, a temporary code of 599 is used and marked as pending superior updates). Simultaneously verify with township archivists to ensure the list is suitable for grassroots operations.

[0088] Step S132: Conduct a survey on the characteristics of rural data, organize discussions with township agricultural technicians, archivists, and village cadres, identify special data scenarios in the local area, and compile a list of special rural data scenarios.

[0089] Three thematic seminars were organized for township agricultural technicians (focusing on practical land classification), archivists (focusing on historical data preservation), and village secretaries and accountants (focusing on population and ownership records). These seminars were combined with on-site reviews of land ledgers from 2000 onwards and visits to five different types of villages (plains, mountainous areas, and suburbs). Three core special scenarios were identified: Land-related issues included the lack of corresponding codes for local colloquial terms like "private plot" and "reclaimed land," and inconsistent use of acreage in old ledgers (e.g., 3.2 mu not being uniformly converted); Population-related issues included elderly people without ID cards using household registration numbers + names instead, and abbreviations like "daughter-in-law" and "son-in-law" for kinship; and Archives-related issues included handwritten plot codes (e.g., three groups of old well land) without standard mappings, and missing fields in contracts from different eras (e.g., early contracts lacking signatures of the contracting party). A "List of Special Rural Data Scenarios" was compiled, clearly defining scenario descriptions, problem impacts, and adaptation requirements, serving as the core basis for localizing standards.

[0090] Step S133: Compile a unified data dictionary. Based on the superior standard framework and local characteristics, refine three core dictionaries: the land classification coding dictionary clarifies local colloquial names, standard names, and coding mappings; the population information dictionary standardizes field types, required fields, and optional fields; and the association rule dictionary defines the business logic of people, land, land, and village groups.

[0091] When compiling the unified data dictionary, three core contents were refined in close accordance with the superior standards and local scenarios: In the land classification coding dictionary, private plots are mapped to other agricultural land-510, reclaimed land is marked with temporary code 599 and noted that it will be updated by the superior standards, and sloping land is corresponding to cultivated land-211; The population information dictionary clearly states that the ID number is a mandatory 18-digit string, and for elderly people without ID cards, a field combining household registration number and name is added as an identifier; for kinship, daughter-in-law and son-in-law are standardized as daughter-in-law and son-in-law and set as mandatory fields; The association rule dictionary defines that people-land follow the principle of one household, one residence, and minors do not contract separately, and land-village groups are matched with the township-village administrative code by matching the first 6 digits of the land plot code, ensuring that the logic fits the actual rural land rights confirmation business.

[0092] Step S134: Develop detailed rules for data format and precision, and clarify standards for commonly used data types in rural areas: area data should be uniformly expressed in mu (a Chinese unit of area); date fields should distinguish between the contract signing date and the start date of the contract period; spatial coordinates should uniformly adopt the 2000 National Geodetic Coordinate System; and redundant suffixes for village and group should be removed from text fields to ensure format consistency.

[0093] Step S135: Build a standard verification tool, configure automated verification scripts based on data dictionary rules, support batch checks of field integrity, format compliance, and logical rationality, mark error types for abnormal data and provide correction suggestions.

[0094] When formulating detailed rules for data format and accuracy, specific requirements were made to align with the actual scenarios of rural data: area data was uniformly converted to mu (a Chinese unit of area), and 3 mu 2 fen (a Chinese unit of area) and 500 square meters were converted to decimal forms (3.20 mu and 0.75 mu), retaining two decimal places to adapt to the accuracy of land rights measurement; date fields were standardized, with the contract signing date accurate to the day (e.g., May 10, 1998 converted to 1998-05-10), and the contract term start date accurate to the year (marked as 2000-01-01 with a note on the start year); spatial coordinates were uniformly adopted using the 2000 National Geodetic Coordinate System, and old data in the Beijing 54 coordinate system needed to be converted in batches, with coordinate values ​​retaining six decimal places; text fields were standardized, removing redundant suffixes such as "village committee" and "village group", and "XX village group 3" and "XX village team 3" were uniformly standardized to "XX village group 3", and handwritten abbreviations were standardized according to this rule to ensure format consistency.

[0095] Step S140: Construct cross-departmental cross-verification and blockchain evidence storage. Use a combination of data source cross-verification and time series verification to compare and verify core data such as population identity information and land ownership boundaries. Store the confirmation results and key information of data modification records on the blockchain for evidence storage, and automatically execute data verification rules through smart contracts.

[0096] In this embodiment, in the scenario of rural land ownership confirmation and registration, this solution can efficiently resolve data verification and ownership disputes. When staff initiate the ownership confirmation process, the system automatically connects to multiple departmental interfaces for cross-verification: comparing farmers' ID numbers and registered addresses through the household registration interface to eliminate duplicate ownership confirmations caused by one person having multiple households; retrieving drone aerial survey data from the natural resources department to verify whether the boundaries of the land parcel declared by the farmer are consistent with the measured coordinates, triggering an alert if the area deviation exceeds 15%; linking with civil affairs marriage registration data to verify whether the kinship label matches the marital status, avoiding false spouse associations. At the same time, combining time series verification, comparing the 2018 ownership confirmation file of the land parcel with the current data, checking whether there is any unregistered transfer of ownership. After the ownership confirmation result is generated, the system puts the land parcel number, owner information, boundary coordinates, and review records on the blockchain for evidence storage. Data modifications require authorization from the village, township, and county-level review nodes, and modification records are uploaded to the blockchain in real time. When farmers have disputes with their neighbors over land boundaries, staff can quickly access aerial survey comparison records stored on the blockchain and verification logs from multiple departments to trace the source of the data throughout the process. During mediation, the immutable blockchain evidence can be directly used as proof of ownership, improving the efficiency of handling disputes related to land rights at the grassroots level by 60%, and gaining recognition from both farmers and government departments for the credibility of the data.

[0097] Step S141: When connecting with the household registration, natural resources, and civil affairs departments, a standardized API gateway is used to adapt to the interface protocols of each department, and read-only permissions are applied for to avoid the risk of data tampering; a data caching module is configured for rural grassroots networks, and a departmental key management system is established, with each department assigned an independent access key and operation logs recorded in real time.

[0098] When connecting with the natural resources and civil affairs departments, a lightweight, standardized API gateway is deployed to adapt to the protocols of each department: the household registration interface adapts to the GB / T35678 government data protocol, the natural resources land boundary interface adapts to the WFS spatial data protocol, and the civil affairs marriage and kinship interface adapts to the REST protocol, ensuring cross-departmental data interaction compatibility. Only read-only permissions are requested from each department for fields necessary for land rights confirmation, such as only obtaining ID card numbers and household registration addresses, avoiding access to sensitive information and mitigating the risk of tampering. To address network fluctuations in rural areas, a data caching module is configured on township servers to cache frequently queried data from the past three years (such as resident household registration and core land boundaries), supporting data retrieval for up to three days offline. A department-specific key management system is established, with independent access keys assigned to natural resources and civil affairs respectively. Keys are forcibly changed every 90 days, and operation logs record the access subject, time, retrieved data items, and purpose in real time. Abnormal access immediately triggers SMS alerts, adapting to the needs of grassroots data security management.

[0099] Step S142: After extracting the original data of each department, the format is unified through an ETL tool. For household registration data, the household register number and name combination identifier are associated with the ID card number. For natural resources data, the land boundaries in the Beijing 54 coordinate system are converted to the 2000 National Geodetic Coordinate System. For civil affairs marriage data, the date format is unified as YYYY-MM-DD; outliers are synchronously cleaned.

[0100] A lightweight ETL tool adapted to the computing power of grass-roots servers is used to process data, taking into account both efficiency and practicality. When processing household registration data, in the case of some elderly people without ID cards, the "household register number + name" combination identifier is associated and bound with the ID card number of relatives, and variant Chinese characters in names are synchronously corrected (such as "弢" unified as "涛"); in the conversion of natural resources data, the land boundaries in the Beijing 54 coordinate system are batch-converted to the 2000 National Geodetic Coordinate System through the CoordinateSharp tool, and combined with on-site calibration of village group landmarks to ensure that the boundary deviation ≤ 0.5 meters; for civil affairs marriage data, irregular formats such as handwritten "08.3.15", March 2008, etc. are unified as 2008-03-15, and the missing years in early archives are supplemented. In the outlier cleaning link, problems such as 15 / 17-digit ID card numbers, land boundaries exceeding the administrative scope of the village, and marriage registration dates later than the birth dates of children are automatically screened, types such as format errors and logical conflicts are marked, and correction suggestions are pushed, adapting to the actual scenario of irregular rural historical data.

[0101] Step S143: When cross-verifying data sources, for population information, the consistency of names and kinship between household registration and civil affairs marriage registration is compared. For land ownership boundaries, the coordinate deviation between natural resources aerial survey data and the coordinates in the village collective ledger is compared.

[0102] When cross-verifying data sources, for population information, the names (including those after the correction of variant Chinese characters) and kinship (such as whether the spouse Zhang San registered in civil affairs is consistent with the household registration) between household registration and civil affairs marriage registration are mainly compared; for land ownership boundaries, the precise coordinates of natural resources UAV aerial surveys are compared with the rough coordinates in the village collective handwritten ledger. For time series verification, the confirmation files and land transfer agreements in the past 10 years are retrieved. When there is no transfer record, whether the owners are consistent is verified. If the land area or boundary deviation exceeds 10% or the owner changes abnormally, manual review is immediately triggered.

[0103] Step S144: A lightweight consortium blockchain is adopted to upload the confirmation results, modification records, and verification logs to the chain, and a unique hash value is generated for each piece of data; the verification rules are solidified in a smart contract.

[0104] Employing the FISCOBCOS lightweight consortium blockchain, adapted to the computing power of rural grassroots servers, the on-chain information focuses on the core of land rights confirmation. Confirmation results include the owner's ID number, land parcel number, and boundary coordinates. Modification records indicate the operator (village / township reviewer), time, and modification items (such as area correction). Verification logs include cross-verification results from civil affairs and natural resources departments. Each data entry generates a unique hash value; farmers can retrieve the entire process information by entering the hash when checking land rights confirmation records. Smart contracts solidify verification rules, such as automatically blocking confirmations for minors and rejecting on-chain data with boundary deviations exceeding 15%. This supports rapid data tracing from collection and verification to modification via hash, adapting to the needs of convenient grassroots operations.

[0105] Step S145: Build a real-time monitoring module to monitor interface connectivity and data verification pass rate, and automatically switch to backup interface when connection is lost; conduct data synchronization calibration with various departments every month, and for historical data, scan paper files to generate electronic copies, manually enter the data, and have multiple departments jointly sign and confirm before uploading it to the blockchain.

[0106] The established real-time monitoring module is adapted for grassroots operations. The interface displays the connectivity status and data verification pass rate of the natural resources and civil affairs interfaces. A red alert is triggered when the connection is interrupted or the pass rate is below 90%, and the system automatically switches to the department's backup interface (such as a backup VPN interface). Data synchronization and calibration are conducted with the three departments on the 10th of each month to verify newly added household registration, land transfer, and other data. For historical paper archives without electronic copies, high-speed scanners are used to generate electronic copies, which are then entered by both the village accountant and the township archivist. After verification and signature by the relevant departments, the data is uploaded to the blockchain to ensure the compliance and traceability of legacy data.

[0107] Step S1411: First, investigate the high-frequency data needs at the rural grassroots level, determining the cache scope to include data such as the household registration of permanent residents in each village within the township for the past 3 years, the boundaries of core contracted land plots, and active contracted contracts. A lightweight caching tool is selected and deployed on the township-level government server, with cache partitions divided by village group and data type. Incremental synchronization is triggered upon network recovery. During the investigation, visits to the township land rights confirmation office and village committees revealed that high-frequency data at the grassroots level is concentrated in land rights registration and dispute mediation scenarios. The final cache scope is determined to include the household registration of permanent residents in each village within the township for the past 3 years (including combinations of household registration numbers for elderly people without ID cards), the boundaries of core contracted land plots (marked with historical and current codes), and active contracted contracts (signed within the past 5 years). A RedisCompact version adapted to low-configuration township servers is selected and deployed in the township government data center, with partitions set by village group and data type (e.g., Li Village - population, Wang Village - land plots). Upon network recovery, only newly added / modified data is synchronized (e.g., newly added household registration, land transfer records), avoiding full synchronization that consumes limited bandwidth. Offline querying of cached data is also supported to adapt to the actual situation of large fluctuations in rural networks.

[0108] Step S1412: Create an offline adaptation function. When the grassroots land rights confirmation software detects a network interruption, it automatically switches to cached data mode, supports data query and temporary entry, and automatically re-uploads temporary data after the network is restored.

[0109] The grassroots land rights confirmation software has a built-in real-time network status detection module. When rural areas experience network outages due to thunderstorms or other weather conditions, it automatically prompts a pop-up window within one second to switch to cache mode, requiring no manual operation. In this mode, cached household registration data (including elderly combination identifiers) and land boundary data can be queried normally. It also supports the temporary entry of land rights confirmation information declared by farmers (such as land area and kinship). The data is encrypted and stored locally and marked as pending synchronization. After the network is restored, the software automatically compares the temporary data with the cloud data. If there are no conflicts, it silently re-uploads the data. If conflicts occur, such as the modification of information for the same land parcel, a pop-up window prompts the administrator to verify, adapting to the actual scenario of unstable grassroots networks and frequent unattended operation.

[0110] Step S1413: Using the AES-256 symmetric encryption algorithm, independent access keys are generated for Natural Resources and Civil Affairs respectively. The keys are delivered offline by the county-level competent department to the township data administrator via an encrypted USB flash drive, and the "Key Receipt Registration Form" is filled out for record-keeping at the same time; key rotation rules are set; operation log recording fields are designed, including access subject, access time, retrieved data items, operation type, and operation result. The logs are stored locally on the township server, and log monitoring rules are configured. When data retrieval exceeds the scope or key verification fails three times consecutively, the system immediately sends an SMS warning to the township administrator and locks the access permissions of the corresponding department, which requires manual unlocking.

[0111] Step S150: Deploy IoT sensors to monitor changes in land use status in real time, establish a real-time data synchronization channel with the civil affairs system, regularly conduct full-scale inspections of data quality through AI models, generate yellow, orange, and red early warning reports, automatically trigger rectification processes for data missing or erroneous issues, and continuously maintain data quality.

[0112] In this embodiment, in a land ownership confirmation scenario in a certain township, this mechanism effectively solves the problems of data lag and errors. Low-power IoT sensors deployed in the fields (with a one-year battery life, suitable for rural environments without mains power) monitor the land use status in real time. When a plot of land changes from farmland to greenhouse, the sensor triggers a land use change signal, and the system automatically marks it as needing to update the transfer information. At the same time, data is synchronized with the system in real time: after a newborn is registered in a villager's household, the family population base is updated within one hour; when an elderly person dies and cancels their household registration, the ownership of their contracted land is automatically linked to the verification process. Synchronization with civil affairs marriage data ensures that after divorce and separation of households, the co-ownership relationship of the original family's contracted land is promptly separated. During the monthly full-scale inspection by the AI ​​model, issues such as "the contracting party is not filled in (yellow warning), the land area is marked in both mu and hectares (orange warning), and the same plot is confirmed by two companies (red warning)" are screened out, generating a report with rectification steps. After receiving a red alert, the system automatically pushes historical land ownership records and sensor monitoring records to assist in rapid verification. An orange alert that is not corrected within the specified time will trigger township-level supervision to ensure that the data always reflects the actual situation of rural population changes and land transfer, providing reliable support for land ownership confirmation and dispute mediation.

[0113] The specific steps for deploying low-power IoT sensors and monitoring land conditions include: Step S151: Selecting sensors that are suitable for the rural field environment, prioritizing multi-parameter sensors that integrate soil moisture, vegetation coverage, and surface temperature; deploying them at a density of one sensor per 50 mu of contracted land at the boundaries of the plots, areas with frequent transfers, and disputed plots, avoiding areas with shade, steep slopes, etc., and fixing them with cement bases to prevent movement.

[0114] When selecting sensors, priority is given to multi-parameter sensors with an 18-month battery life, IP68 waterproof and corrosion resistant (suitable for fields with rainwater and pesticide residues), and integrated functions for collecting soil moisture, vegetation coverage, and surface temperature, while balancing low power consumption and monitoring accuracy. Deployment is based on one sensor per 50 mu (approximately 3.3 hectares), with increased density to one sensor per 30 mu (approximately 2 hectares) in areas with concentrated vegetable greenhouses, plots that have been transferred more than twice in the past three years, and historically disputed boundary plots. Avoid areas with trees (to prevent misjudgment of vegetation coverage) and steep slopes (to prevent base tipping). Use precast cement bases, buried 30 cm deep, for fixation. Label the bases with the sensor number and corresponding plot code for easy future maintenance and verification.

[0115] Step S152: LoRaWAN low-power wide area network communication is adopted. Data is aggregated through the township-level edge gateway, which is deployed near the village committee signal tower and supports local caching of 3 days of data.

[0116] Employing LoRaWAN low-power wide-area network communication, its strong penetration capability makes it suitable for rural forests and residential buildings with obstructed views, ensuring stable transmission of sensor data to the township-level edge gateway. The gateway is a lightweight, wall-mounted model, fixed to the village committee's signal tower bracket, leveraging the tower to extend coverage (covering the entire village). It is powered by mains electricity and equipped with a 12-hour backup battery (for temporary power outages). The gateway has 16GB of built-in storage, automatically caching data in case of network disconnection due to thunderstorms, retaining it for up to 3 days. After network recovery, only newly added data is synchronized, avoiding full data transmission that consumes bandwidth.

[0117] Step S153: Embed a lightweight AI model at the edge to analyze sensor data in real time: when vegetation coverage drops sharply and soil moisture remains below the threshold, it is determined that the farmland has been converted to non-farmland; when the daily fluctuation of surface temperature exceeds 15°C and regular mechanical vibration occurs, it is marked as possible construction occupation, and the abnormal status is pushed to the township land rights confirmation system in real time.

[0118] An optimized, lightweight random forest AI model is embedded at the edge (adapted to the limited computing power of township gateways). When analyzing sensor data, judgment thresholds are set to fit the realities of rural areas: if soil moisture remains ≤25% for 72 consecutive hours (the lower limit of suitable moisture for rural arable land), and vegetation coverage decreases by more than 30% within 24 hours, it is judged as arable land being converted to non-arable land (such as abandonment or conversion to non-agricultural crops); if the daily fluctuation of surface temperature exceeds 15℃, and regular mechanical vibration of 5-15Hz is detected (corresponding to bulldozer or excavator operation), it is marked as potentially occupied by construction. Abnormal statuses are simultaneously pushed to the township land ownership confirmation system, notifying the administrator via SMS and pop-up window, along with the sensor number and corresponding plot coordinates, facilitating rapid on-site verification.

[0119] The steps for establishing a real-time synchronization channel with the civil affairs system include: Step S161: For population changes, the civil affairs side records changes in marital status and kinship.

[0120] Population changes include newborn registration, cancellation of household registration for deceased elderly, relocation of villagers (such as marriage into the village), and correction of ID numbers; marital status on the civil affairs side includes marriage registration and divorce registration, and changes in kinship include household splitting and adoption of new relatives, which are closely related to the actual scenarios of rural population changes.

[0121] Step S162: The system binds to the township land rights confirmation platform through a WebSocket long connection and pushes the data proactively at regular intervals after population data changes; the civil affairs system adapts to its RESTful interface and configures a dual mechanism of timed polling and change-triggered push.

[0122] After establishing a long WebSocket connection with the township land ownership confirmation platform, the system is configured with an automatic reconnection function (retrying within 3 seconds after disconnection to avoid network fluctuations in rural areas causing disconnection). Population data changes (such as newborn registration, elderly household registration cancellation, and villagers relocating) trigger push notifications. Within 10 minutes, a data packet containing the ID number, change type, and village / group code is encrypted and transmitted to ensure the land ownership confirmation platform updates the family population base in a timely manner. The civil affairs system adapts to its RESTful interface, employing a 5-minute timed polling + key change-triggered push notification: daily polling synchronizes ordinary marital status; key changes such as divorce, household separation, and adoption registration trigger immediate push notifications. The data is in JSON format, including the party's ID number, marital / kinship change time, and village / group affiliation. If the push fails (e.g., due to a grassroots network interruption), the system retains the data and prioritizes resending it after network recovery, ensuring accurate synchronization of information such as land ownership relationships and owner associations during land ownership confirmation.

[0123] Step S163: The transport layer uses TLS1.3 encryption, and the data format is uniformly JSON, including the change ID and original data, new data and operator signature; deploy data verification middleware, which automatically compares the field integrity after receiving the data. If the verification is successful, it triggers an incremental update of the core database. If it fails, it is cached and the administrator is notified via SMS. If the problem is not resolved within 2 hours, it will automatically switch to the backup channel.

[0124] The transport layer uses TLS 1.3 encryption to adapt to low-configuration township servers, avoiding excessive computing power consumption and potential lag. Data is encapsulated in JSON format: change IDs are labeled by department (e.g., "G20240508001"), and the original / new data is exemplified as "Original household registration: Li Village Group 1 → New household registration: Wang Village Group 2". The operator's signature is the village administrator's electronic signature. The data verification middleware automatically checks key fields (population data must include ID card number and village / group code, marriage data must include registration date). If verification passes, it triggers an incremental update to the core database; if it fails, it is cached to the township server, and the township land rights confirmation office administrator is notified via SMS. If not processed within 2 hours, it automatically switches to the FTP backup channel, adapting to rural network fluctuation scenarios.

[0125] Based on the same inventive concept, please refer to Figure 2, which shows a schematic block diagram of an Internet-based multi-village government service optimization system 100 provided in this application embodiment for executing the above-described Internet-based multi-village government service optimization method. The Internet-based multi-village government service optimization system 100 may include a communication unit 110, a machine-readable storage medium 120, and a processor 130. Alternatively, the machine-readable storage medium 120 may also be integrated into the processor 130 and can communicate and interact with external systems through the communication unit 110. The machine-readable storage medium 120 is used to store machine-executable instructions for executing the scheme of this application, and the processor 130 is used to execute the machine-executable instructions stored in the machine-readable storage medium 120 to implement the Internet-based multi-village government service optimization method provided in the aforementioned method embodiment.

[0126] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.

Claims

1. A method for optimizing multi-village government services based on the Internet, characterized by: The process includes the following steps: First, using a combination of satellite imagery, UAV mapping, and a mobile offline data collection app to capture land boundary and topographical spatial data; collecting population registration and contract data; batch scanning and recognizing historical paper archives using OCR; extracting information and entering it into the information management system; second, constructing AI and RAG intelligent cleaning models; using generative adversarial networks to fill in missing data values; identifying abnormal data such as duplicate population registrations and land area conflicts using local anomaly factor algorithms; training entity recognition models based on the rural government affairs vertical RAG library; and automatically labeling data attributes and relationships; third, referencing the agricultural data standard system to compile a unified data dictionary for land classification coding and population information fields, clarifying data formats, accuracy requirements, and association rules; using tools to batch convert raw data from different sources and formats into standard formats; constructing a knowledge graph of land, population, and assets to achieve interconnection and interoperability of land, people, and assets data; and fourth, constructing cross-departmental cross-verification and blockchain evidence storage, using a combination of data source cross-verification and time series verification to compare and verify core data such as population identity information and land ownership boundaries; storing key information such as the confirmation results and data modification records on the blockchain; and automatically executing data verification rules through smart contracts. Deploy IoT sensors to monitor changes in land use status in real time, establish a real-time data synchronization channel with the civil affairs system, regularly conduct full-scale inspections of data quality using AI models, generate yellow, orange, and red early warning reports, and automatically trigger rectification processes for data missing or errors to continuously maintain data quality.

2. The method for optimizing multi-village government services based on the Internet according to claim 1, characterized in that: The method employs a combination of satellite imagery, UAV mapping, and mobile offline data collection apps to capture land boundary and topographic spatial data, collect population registration and contract data, and batch scan and identify historical paper archives using OCR to extract information and input it into the information management system. This includes: selecting satellite imagery covering the entire target area; preprocessing the images using remote sensing image processing software to generate digital orthophoto maps; automatically identifying land use types such as cultivated land, forest land, and construction land based on a deep learning semantic segmentation model; annotating plot boundary contours and topographic relief features to form a pre-annotated plot dataset; and then, based on the satellite image annotation results... The drone aerial survey route is designed to integrate village terrain; a grid-like route is used for plains areas, and a strip-like route is used for mountainous and hilly areas; point cloud data acquired by drone aerial survey is used to generate digital elevation models and digital surface models using 3D modeling software, extracting terrain parameters such as slope, aspect, and elevation of plots, combining real-world imagery with pre-annotated boundaries from satellite imagery, and using a multi-source image registration algorithm to correct plot boundary deviations, generating plot vector data and forming preliminary spatial data results; the aforementioned satellite imagery base map, drone-generated plot vector data, and pre-annotated information are pre-installed in a mobile app, allowing staff to access the app offline after arriving at the site with their equipment. The system uses GPS / BeiDou dual-mode positioning to match the current location with the land parcel data in real time. Deviations are corrected by manually dragging boundary nodes or drawing new boundaries, while simultaneously taking photos of the land parcel's current condition as supporting evidence. For remote areas without network coverage, the mobile app stores the corrected land parcel boundary data, supplementary terrain information, and on-site photos in a local cache. When staff return to the village committee or an area with network access, the mobile app automatically triggers data synchronization, uploading the offline-collected data to the cloud server. This data is then merged with spatial data generated by UAV aerial surveying to update the land parcel vector database. Simultaneously with spatial data collection, the system also utilizes mobile... The mobile app synchronously collects business data, and the app's built-in OCR module automatically extracts the business data. After farmers confirm the data on-site, they sign electronically. They can also take a picture of the original contract or upload a scanned copy through the app. OCR technology automatically recognizes the information and links it to the corresponding land parcel vector data. After the data is uploaded to the cloud, the system initiates a correlation verification between spatial data and business data to form an integrated data record of spatial and business attributes. For duplicate land parcel data, duplicates are automatically removed based on the collection time and accuracy priority. Fields with large OCR recognition errors are marked as pending review and pushed to staff for online comparison with the original image to complete the correction.

3. The method for optimizing multi-village government services based on the Internet according to claim 1, characterized in that: The construction of the AI ​​and RAG intelligent cleaning model utilizes generative adversarial networks to fill in missing data values, identifies abnormal data such as duplicate population registrations and land area conflicts through the local anomaly factor algorithm, and trains an entity recognition model based on the rural government affairs vertical RAG library to automatically label data attributes and relationships. This includes: constructing a rural government affairs vertical RAG library, collecting rural land contracting case data sources, removing outdated and duplicate content, classifying and organizing it into policy / coding / case / format categories, performing text preprocessing using jieba word segmentation, and then using the Sentence-BERT model to convert the text into vector embeddings, storing them in the Milvus vector database. A keyword search engine was built to support rapid, related queries based on policy clauses, village / group names, and data types. Raw data was preprocessed and standardized, unifying field names, aligning data formats, removing redundant characters, and extracting key information fragments using regular expressions to form a standardized dataset recognizable by the model. Training, validation, and test sets were also created. A generative adversarial network adapted to rural data was trained, using samples with missing values ​​as input. A loss function was constructed by combining statistical characteristics of archives from the same village / group and era in the RAG library with policy constraints. A discriminator was responsible for distinguishing the reasonableness of generated data from real data. Model parameters were optimized through iterative training, and missing fields were filled after training. The algorithm adapts to the local anomaly factor for anomaly identification. First, it calculates a local density threshold based on normal data in the training set. The preprocessed dataset is then input into the algorithm to automatically identify anomalous data and label risk levels according to the severity of the anomalies. An entity recognition model is trained based on the RAG library. Labeled samples are added to the training set, and model parameters are fine-tuned through transfer learning. After training, standardized data is input, and core entity attributes are automatically labeled. The labeling accuracy must reach a set percentage; samples that do not meet the standard are fed back to the RAG library to supplement labeled samples, iteratively optimizing the entity recognition model. A graph neural network is used to construct a graph linking people, land, and contracts, linking farmer ID numbers with contracted land plot numbers and plot numbers. The system integrates spatial boundary data, family member ID numbers, and kinship tags; it also integrates AI models and the RAG library to build an automated cleaning system. The system is set up with an automated execution chain for data preprocessing, GAN missing value imputation, LOF anomaly identification, entity annotation, and relationship establishment. The system automatically reads input data, calls the corresponding model, outputs cleaning results, and simultaneously generates a cleaning report. A manual review mechanism is also included. The system automatically filters out complex problem samples, sorts them by risk level, and pushes them to staff. Simultaneously, the review interface displays reference policy clauses, similar processing cases, and original electronic images from the RAG library. The corrected data is automatically fed back to the model training set.

4. The method for optimizing multi-village government services based on the Internet according to claim 3, characterized in that: The method constructs a graph of relationships between people, land, and contracts using a graph neural network, linking farmers' ID numbers with contracted land plot numbers, land plot numbers with spatial boundary data, and family members' ID numbers with kinship tags. This includes: defining graph entities and relationship types, setting people, land, and contracts as core entities, clarifying four types of relationships: contracting, undertaking, signing, and kinship, while incorporating historical and current coding mapping rules for rural areas; extracting entity features and association clues, extracting key attributes of people, land, and contracts from structured data, extracting implicit association clues through regular expression matching and entity recognition models, and converting features into vector forms recognizable by graph neural networks; constructing an initial graph structure, with entities as nodes and association clues as edges, automatically establishing initial connections, specifically associating people with contracts through the contractor's ID number in the contract, associating contracts with land through the land parcel number in the contract, and associating kinship relationships between people through registered address and age difference, forming graph data containing node attributes and initial edges; training graph neural networks to optimize associations, using GCN or GAT architecture, inputting the initial graph data, and training the graph neural network model with correctly labeled association samples, strengthening semantic associations between entities by updating node embedding vectors; using the trained graph neural network model to predict new associations and verify logic, marking abnormal associations for manual review and correction, and supporting the addition of new data to the graph through incremental learning.

5. The method for optimizing multi-village government services based on the Internet according to claim 3, characterized in that: The integrated AI model and RAG library construct an automated data cleaning system. The system is set up with an automated execution chain for data preprocessing, GAN missing value imputation, LOF anomaly detection, entity annotation, and relationship establishment. The system automatically reads input data, calls the corresponding model, outputs cleaning results, and synchronously generates a cleaning report. This includes: building a parsing tool adapted to rural data formats; reading input data from Excel ledgers and OCR results of scanned paper documents; splitting fields by type; and uniformly converting them to UTF-8 encoded JSON format; batch correcting handwritten variant characters and date abbreviations using a pre-trained text normalization model and outputting standardized intermediate data; and training a rural scene GAN model based on complete historical data to address issues such as missing kinship in population data and missing boundaries in land data. For missing data, the model generates logically valid imputation values ​​based on the input missing field's associated features. After imputation, the values ​​are validated according to rules. Anomaly thresholds are set for rural data, and the LOF algorithm is used to calculate outlier scores. Data exceeding the threshold is marked as anomalies, and historical data is compared and anomaly types are labeled. An entity recognition model is deployed to automatically label core entities such as property owners, land parcels, and contract numbers. Relationships are established based on a rule engine. For conflicts in multi-source data, historical name comparison records from the RAG library are used for auxiliary matching to generate association confidence. A workflow engine connects the modules, setting trigger conditions to automatically call the model in the order of preprocessing, imputation, recognition, labeling, and association. After cleaning, a report is generated, including the total cleaning volume, missing data imputation rate, number of anomaly cases, association matching results, and typical case studies.

6. The method for optimizing multi-village government services based on the Internet according to claim 1, characterized in that: Referring to the aforementioned agricultural data standard system, a unified data dictionary for land classification codes and population information fields was compiled, clarifying data formats, accuracy requirements, and association rules. Tools were used to batch convert raw data from different sources and formats into the standard format, constructing a knowledge graph of land, population, and assets to achieve interconnectivity and interoperability of land, people, and assets. This included: reviewing and aligning with higher-level standards, collecting agricultural documents, extracting core fields for land classification codes and population information, and establishing a standard alignment list, clarifying mandatory clauses and adjustable flexibility; conducting surveys on rural data characteristics, and organizing township agricultural technicians, archivists, and village cadres. The meeting addressed specific issues related to local data scenarios, resulting in a list of such scenarios for rural data. A unified data dictionary was developed, refining three core categories based on higher-level standards and local characteristics: a land classification coding dictionary specifying local colloquialisms, standard names, and coding mappings; a population information dictionary standardizing field types, required fields, and optional fields; and an association rule dictionary defining the business logic for people, land, and village groups. Detailed rules for data format and precision were established, clarifying standards for rural data types. Automated verification scripts were configured based on the data dictionary rules to support batch checks of field integrity, format compliance, and logical rationality, marking error types for abnormal data and providing correction suggestions.

7. The method for optimizing multi-village government services based on the Internet according to claim 1, characterized in that: The construction of cross-departmental cross-verification and blockchain evidence storage adopts a combination of data source cross-verification and time series verification to compare and verify core data such as population identity information and land ownership boundaries. The results of rights confirmation and key information such as data modification records are stored on the blockchain. Data verification rules are automatically executed through smart contracts, including: when connecting with household registration, natural resources, and civil affairs departments, a standardized API gateway is used to adapt to the interface protocols of each department, and read-only permissions are requested to avoid the risk of data tampering; a data caching module is configured for rural grassroots networks, and a departmental key management system is established, with each department assigned an independent access key and operation logs recorded in real time; after extracting the original data from each department, an ETL tool is used to unify the format. For household registration data, the household registration number and name combination identifier are associated with the ID card number; for natural resources data, the format is... The land boundaries of the Beijing 54 coordinate system were converted to the 2000 National Geodetic Coordinate System; the date format for civil affairs marriage data was unified as YYYY-MM-DD; outliers were cleaned synchronously; during cross-validation of data sources, population information was compared with the names and kinship records of household registration and civil affairs marriage registration; and the coordinate deviation of land ownership boundaries was compared with the coordinates of natural resource aerial survey data and village collective ledgers; a lightweight consortium blockchain was adopted to put the confirmation results, modification records, and verification logs on the chain, and a unique hash value was generated for each piece of data; smart contracts solidified the verification rules; a real-time monitoring module was built to monitor interface connectivity and data verification pass rate, and automatically switch to a backup interface when the connection is lost; monthly data synchronization calibration was carried out with various departments; for historical data, after scanning paper archives to generate electronic copies, manual entry was performed and confirmed by multiple departments before being put on the chain.

8. The method for optimizing multi-village government services based on the Internet according to claim 7, characterized in that: When connecting with the household registration, natural resources and civil affairs departments, a standardized API gateway is used to adapt to the interface protocols of each department, and read-only permissions are applied for to avoid the risk of data tampering. A data caching module was configured for rural grassroots networks, and a departmental key management system was established, with each department assigned an independent access key and operation logs recorded in real time. This included: surveying the high-frequency data needs of rural grassroots areas and determining the caching scope to include the household registration of permanent residents in each village within the township for the past three years, the boundaries of core contracted land plots, and active contract data; selecting a lightweight caching tool and deploying it on the township-level government server, dividing the cache into partitions according to village groups and data types, and triggering incremental synchronization when the network recovers; and creating an offline adaptation function so that when the grassroots land rights confirmation software detects a network interruption, it automatically switches to cached data mode, supports data querying and temporary data entry, and automatically re-uploads temporary data after the network recovers. The AES-256 symmetric encryption algorithm is used to generate independent access keys for natural resources and civil affairs respectively. The keys are delivered offline by the county-level competent department to the township data administrator via an encrypted USB flash drive, and the key requisition registration form is filled out and filed at the same time. Key rotation rules are set up; and operation log recording fields are designed.

9. The method for optimizing multi-village government services based on the Internet according to claim 1, characterized in that: The specific steps for establishing a real-time synchronization channel with the civil affairs system include: for population changes, the civil affairs side handles changes in marital status and kinship; the system binds to the township land rights confirmation platform via a WebSocket long connection, and actively pushes population data changes on a regular basis; the civil affairs system adapts to its RESTful interface, configuring a dual mechanism of timed polling and change-triggered push; the transport layer uses TLS 1.3 encryption, and the data format is uniformly JSON, including the change ID and original data, new data, and operator's signature; a data verification middleware is deployed, which automatically compares the field integrity upon receipt, triggers incremental updates to the core database upon successful verification, and caches and notifies the administrator via SMS if it fails, automatically switching to the backup channel if the issue is not resolved within 2 hours.

10. A multi-village government service optimization system based on the Internet, characterized in that: include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the Internet-based multi-village government service optimization method according to any one of claims 1 to 9 by executing the machine-executable instructions.