A method for detecting consistency of provincial, municipal and county real estate registration data
Through the layered fingerprint comparison method and the information fingerprint generation method combined with SHA-256 and the autoencoder, the problems of low data consistency detection efficiency and poor accuracy in the prior art are solved, and efficient and accurate data consistency detection and data protection are achieved.
Patent Information
- Application Number
- CN202410705735.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-03
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-06-03
AI Technical Summary
When the prior art detects the consistency of real estate registration data in provinces, cities and counties, the calculation amount is large and data comparison inconsistent problems are prone to problems, especially when the data amount is large, efficiency and accuracy are difficult to guarantee.
The layered fingerprint comparison method is used to compare the information fingerprints of the three-level real estate registration data at the province, city and county level layer by layer, quickly locate and repair data differences to ensure the consistency of the data. The information fingerprint generation method combined with SHA-256 algorithm and autoencoder is used to improve the representativeness and robustness of fingerprints.
It greatly improves the efficiency and accuracy of data consistency comparison, is suitable for large and complex data sets, reduces unnecessary calculations, enhances the protection layer of data, and improves security.
Smart Images

Figure CN118643045B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of consistency of real estate registration data, and specifically to a method for detecting consistency of provincial, municipal and county real estate registration data. Background Art
[0002] Ensuring the synchronization and consistency of real estate registration data in provinces, cities and counties is the basic support for realizing the unified collection and real-time online access of real estate registration data across the province, and is of great significance to further improving the ability to provide convenient services. The provincial real estate registration database is formed by the collection and access of cities and counties. The stock data of real estate registration in cities and counties are collected and submitted to the provincial level at a unified time point, and the daily business processing data of cities and counties are reported to the provincial level through real-time access. Due to reasons such as data standardization, reporting channels, and supervision mechanisms, the real estate registration data of provinces, cities, and counties have not been completely consistent.
[0003] The document with the prior art publication number CN116894193A provides a method for consistency detection of provincial, municipal and county real estate registration data based on information fingerprints. The business number is used as a unique identifier to extract the business data corresponding to each business from the provincial and municipal and county real estate registration databases. According to the format required by the message, it is organized as a string and stored in the MDB file. The MDB file also stores the business number corresponding to each business data string. The string format is unified, including deleting nodes and fields that do not participate in the consistency comparison, sorting nodes and fields, and deleting spaces and line breaks. The information fingerprint of each business data is extracted according to the SHA-256 algorithm. The information fingerprint extracted from the same business identifier of the provincial and municipal and county real estate registration data is compared. If it is consistent, it means that the business data is consistent. If it is inconsistent, it means that the business data is inconsistent, which solves the problems of inadequate integration of existing data and local non-standardized business.
[0004] However, the method in the prior art still has some problems in practical application. This method needs to compare the hash codes generated by all data rows when comparing data consistency, which will generate unnecessary calculations and is prone to problems when dealing with large amounts of data. The simple use of the SHA-256 algorithm will generate different hash codes due to its avalanche effect due to the format, encoding, and similar semantics of the data, resulting in inconsistent data comparisons for data that was originally consistent.
[0005] In view of the above problems, the present invention proposes a method for detecting consistency of provincial, municipal and county real estate registration data.
[0006] Technical Solution
[0007] The present invention provides a method for consistency detection of real estate registration data at the provincial, municipal and county levels, which aims to perform consistency detection on real estate registration data at the provincial, municipal and county levels through a hierarchical fingerprint comparison method to ensure the accuracy and consistency of data at all levels. By comparing data fingerprints layer by layer, potential data differences can be quickly located and repaired, data quality can be improved, and reliable data support can be provided for real estate registration work. The technical solution includes the following aspects.
[0008] According to one aspect of an embodiment of the present application, a method for detecting consistency of provincial, municipal and county real estate registration data is provided, the method comprising the following steps:
[0009] S1. Extract the fields and tables that need to be compared for consistency and stored in provinces, cities and counties to create a new database;
[0010] S2. Extract information fingerprints from each row of data in the table for the provincial and municipal level summary, municipal and county level storage data;
[0011] S3. Compare the information fingerprints summarized at the provincial level and the city level to see if they are consistent. If they are consistent, it means that all data are consistent;
[0012] S4. If there is inconsistency, the information fingerprint of each city in the provincial data storage is compared with the corresponding city-level storage data, and the city with inconsistent information fingerprint is found for further comparison;
[0013] S5. Compare the information fingerprints of counties in the city where the provincial data is inconsistent with the information fingerprints of counties in the city data stored at the municipal level to find the counties where the information fingerprints are inconsistent;
[0014] S6. Compare the data information fingerprints of each row of the county with inconsistent data stored at the provincial level with the data information fingerprints of each row of the county with inconsistent data stored at the municipal level, and find the rows with inconsistent information fingerprints, that is, find the rows with inconsistent data;
[0015] S7. Repair or adjust the difference data according to the comparison results to ensure data consistency and record all operations on the data, including fingerprint generation, comparison and repair process, for tracking and auditing.
[0016] S8. Perform consistency comparison on the data stored at the city level and the data stored at the county level using the same approach.
[0017] In this way, the scope can be narrowed down layer by layer, the differences in the data can be found quickly and accurately, and detailed comparison and analysis can be performed. The layered fingerprint comparison method can greatly improve the efficiency and accuracy of data consistency comparison and is suitable for large and complex data sets. At the same time, it has good scalability and can be customized and optimized according to specific needs and data characteristics.
[0018] Further, as an optional solution of the present application, the information fingerprint generation method (taking the consistency comparison of provincial and municipal data as an example) can be:
[0019] S201, compile the provincial-level stored data that needs to be compared for consistency into a master table, city-level data tables, and county-level data tables;
[0020] S202, summarizing the data stored at the municipal level that need to be compared for consistency into a master table, each municipal-level data table and the corresponding county-level data table;
[0021] S203. Use the SHA-256 algorithm to extract information fingerprints from each table, and obtain the total table stored at the provincial and municipal levels, the municipal data table, the county data table, and the hash value of each row data table as the information fingerprint.
[0022] Through the above scheme, when the consistency comparison result of the upper layer is inconsistent, the information fingerprint can be extracted from the lower layer, avoiding unnecessary calculations and reducing time complexity;
[0023] Further, as an optional solution of the present application, the information fingerprint generation method (taking the consistency comparison of provincial and municipal data as an example) can also be:
[0024] S201', use the SHA-256 algorithm to extract information fingerprints from each row of data that needs to be compared for consistency at the provincial and municipal levels, and store the obtained hash value as the information fingerprint in the county table according to the county corresponding to the data;
[0025] S202', use the SHA-256 algorithm to extract the information fingerprint from the county table, and store the obtained hash value as the information fingerprint into the city table according to the city corresponding to the county;
[0026] S203', use the SHA-256 algorithm to extract the information fingerprint from the city-level table, and store the obtained hash value as the information fingerprint in the provincial-level table;
[0027] S204', using the SHA-256 algorithm to extract the information fingerprint from the provincial table to obtain a hash value as the information fingerprint;
[0028] Through the above scheme, when performing consistency comparison, the provincial information fingerprints are compared first, and when there is inconsistency, the next layer of information fingerprints are compared, which can reduce the spatial complexity to a certain extent.
[0029] Further, as an optional solution of the present application, the algorithm used for information fingerprint generation can also be a fingerprint generation algorithm based on deep learning. In one embodiment, an autoencoder is used to extract features from a table or a data row in a table to obtain a feature vector, and then the SHA-256 algorithm is used to encode the feature vector to form an information fingerprint (taking data table information fingerprint generation as an example). The specific steps are:
[0030] The data table is preprocessed and converted into a format that the neural network can process, such as a numerical matrix or tensor.
[0031] Construct an autoencoder network structure, including an encoder and a decoder. The encoder compresses the input data into a low-dimensional feature representation, and the decoder attempts to reconstruct the original data from the feature representation.
[0032] Use tabular data for training and optimize the parameters of the autoencoder by minimizing the reconstruction error.
[0033] After the autoencoder training is completed, the decoder is removed and only the encoder is retained for feature extraction. Each row of data in the table is input into the encoder to obtain the feature vector representation of each row of data.
[0034] For each feature vector, the SHA-256 hash algorithm is applied to generate a fixed-length hash value. This hash value will serve as the digital fingerprint of the row of data.
[0035] Through the above scheme, a fingerprint generation algorithm combined with an autoencoder can be used to replace the encoding method of simply using the SHA-256 hash algorithm. By using this algorithm, the generated fingerprint will be more compact and meaningful, which improves the representativeness of the fingerprint. It can solve the problem of different encodings generated by the SHA-256 hash algorithm caused by the fact that the content of the data has not changed, but the format or encoding has changed, or the content of the data has changed slightly, but the semantics has not changed. It improves the robustness of the algorithm, can provide an additional layer of protection for the data, and increases the difficulty of tampering or forging the data.
[0036] In one embodiment of the present application, access rights to fingerprint data and raw data are strictly controlled to prevent data leakage, and all operations on the data, including fingerprint generation, comparison and repair processes, are recorded for tracking and auditing. Specifically, staff with operating authority obtain access to fingerprint data, raw data, fingerprint generation, comparison and repair rights through facial recognition. All access and operations to the data and the image path captured by facial recognition are stored in the log to facilitate auditing and tracing.
[0037] The above technical solution improves data security, and facilitates tracing of data access operation records through recorded logs.
[0038] One or more technical solutions provided in the technical solution of this application have at least the following technical effects or advantages:
[0039] This application uses a hierarchical comparison method to reduce the amount of data comparison, narrow the scope layer by layer, and quickly and accurately find the differences in the data, which can greatly improve the efficiency and accuracy of data consistency comparison. It combines the autoencoder and the SHA-256 hash algorithm to generate information fingerprints, improves the representativeness of the fingerprint, and can solve the problem of different encodings generated by the SHA-256 hash algorithm when the content of the data has not changed but the format or encoding has changed, or when the content of the data has changed slightly but the semantics has not changed. It improves the robustness of the algorithm, provides an additional layer of protection for the data, increases the difficulty of data tampering or forgery, obtains data access operation permissions through face recognition, and records them in the log, thereby improving data security. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 A flow chart of a method for detecting consistency of provincial, municipal and county real estate registration data provided for one embodiment of the present application;
[0041] Figure 2 A flow chart of an information fingerprint generation method in a method for detecting consistency of provincial, municipal and county real estate registration data provided in an embodiment of the present application;
[0042] Figure 3 A flow chart of an information fingerprint generation method in another method for detecting consistency of provincial, municipal and county real estate registration data provided in one embodiment of the present application;
[0043] Figure 4 This is a flowchart of the algorithm required for the information fingerprint generation method provided in one embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be described in further detail below in conjunction with the accompanying drawings.
[0045] Please refer to Figure 1 , which shows a flow chart of a method for consistency detection of provincial, municipal and county real estate registration data provided by an embodiment of the present application;
[0046] S1. Extract the fields and tables that need to be compared for consistency and stored in provinces, cities and counties to create a new database;
[0047] S2. Extract information fingerprints from each row of data in the table for the provincial and municipal level summary, municipal and county level storage data;
[0048] S3. Compare the information fingerprints summarized at the provincial level and the city level to see if they are consistent. If they are consistent, it means that all data are consistent;
[0049] S4. If there is inconsistency, the information fingerprint of each city in the provincial data storage is compared with the corresponding city-level storage data, and the city with inconsistent information fingerprint is found for further comparison;
[0050] S5. Compare the information fingerprints of counties in the city where the provincial data is inconsistent with the information fingerprints of counties in the city data stored, and find the counties where the information fingerprints are inconsistent;
[0051] S6. Compare the data information fingerprints of each row of the county with inconsistent data stored at the provincial level with the data information fingerprints of each row of the county with inconsistent data stored at the municipal level, and find the rows with inconsistent information fingerprints, that is, find the rows with inconsistent data;
[0052] S7. Repair or adjust the difference data according to the comparison results to ensure data consistency and record all operations on the data, including fingerprint generation, comparison and repair process, for tracking and auditing.
[0053] S8. Perform consistency comparison on the data stored at the city level and the data stored at the county level using the same approach.
[0054] In this way, the scope can be narrowed down layer by layer, the differences in the data can be found quickly and accurately, and detailed comparison and analysis can be performed. The layered fingerprint comparison method can greatly improve the efficiency and accuracy of data consistency comparison and is suitable for large and complex data sets. At the same time, it has good scalability and can be customized and optimized according to specific needs and data characteristics.
[0055] Please refer to Figure 2 , which shows a flow chart of an information fingerprint generation method in a method for detecting consistency of provincial, municipal and county real estate registration data provided by an embodiment of the present application;
[0056] S201, compile the provincial-level stored data that needs to be compared for consistency into a master table, city-level data tables, and county-level data tables;
[0057] S202, summarizing the data stored at the municipal level that need to be compared for consistency into a master table, each municipal-level data table and the corresponding county-level data table;
[0058] S203. Use the SHA-256 algorithm to extract information fingerprints from each table, and obtain the total table stored at the provincial and municipal levels, the municipal data table, the county data table, and the hash value of each row data table as the information fingerprint.
[0059] Through the above scheme, when the consistency comparison result of the upper layer is inconsistent, the information fingerprint can be extracted from the lower layer, avoiding unnecessary calculations and reducing time complexity;
[0060] Please refer to Figure 3 , which shows a flow chart of an information fingerprint generation method in another method for detecting consistency of provincial, municipal and county real estate registration data provided by an embodiment of the present application;
[0061] S201', use the SHA-256 algorithm to extract information fingerprints from each row of data that needs to be compared for consistency at the provincial and municipal levels, and store the obtained hash value as the information fingerprint in the county table according to the county corresponding to the data;
[0062] S202', use the SHA-256 algorithm to extract the information fingerprint from the county table, and store the obtained hash value as the information fingerprint into the city table according to the city corresponding to the county;
[0063] S203', use the SHA-256 algorithm to extract the information fingerprint from the city-level table, and store the obtained hash value as the information fingerprint in the provincial-level table;
[0064] S204', using the SHA-256 algorithm to extract the information fingerprint from the provincial table to obtain a hash value as the information fingerprint;
[0065] Through the above scheme, when performing consistency comparison, the provincial information fingerprints are compared first, and when there is inconsistency, the next layer of information fingerprints are compared, which can reduce the spatial complexity to a certain extent.
[0066] Please refer to Figure 3 , which shows the algorithm flow chart required for the information fingerprint generation method provided by one embodiment of the present application;
[0067] The data table is preprocessed and converted into a format that the neural network can process, such as a numerical matrix or tensor.
[0068] Construct an autoencoder network structure, including an encoder and a decoder. The encoder compresses the input data into a low-dimensional feature representation, and the decoder attempts to reconstruct the original data from the feature representation.
[0069] Use tabular data for training and optimize the parameters of the autoencoder by minimizing the reconstruction error.
[0070] After the autoencoder training is completed, the decoder is removed and only the encoder is retained for feature extraction. Each row of data in the table is input into the encoder to obtain the feature vector representation of each row of data.
[0071] For each feature vector, the SHA-256 hash algorithm is applied to generate a fixed-length hash value. This hash value will serve as the digital fingerprint of the row of data.
[0072] Among them, the network structure of the autoencoder consists of two parts: the encoder and the decoder. The encoder is responsible for converting the input data into a low-dimensional code, which consists of multiple fully connected layers. These layers will gradually reduce the dimension of the data and extract the most important features. The output of the encoder is a low-dimensional feature vector. The decoder is responsible for restoring the code to an output similar to the input data. The structure is opposite to that of the encoder. It consists of multiple fully connected layers, gradually increasing the dimension of the data until it matches the dimension of the original input data. The output of the decoder is the reconstructed data. The training goal of the autoencoder is to minimize the reconstruction error, that is, to make the difference between the input data and the data reconstructed by the encoder and decoder as small as possible. The reconstruction error metric is the mean square error (MSE), and its formula is: L(x,x^)=1 / n∑ n i=1 (x i -x^ i ) 2 ,x is the original input data, x^ is the data reconstructed by the autoencoder, n is the dimension of the data (for a row of data in a table, n can be the number of features in the row), during the training process, the weights of the autoencoder are updated through the backpropagation algorithm and optimizer (such as gradient descent) to minimize the reconstruction error. In this way, the autoencoder can learn the compressed representation of the input data and try to retain important information so that the original data can be accurately reconstructed during decoding. In order to prevent the autoencoder from learning the identity mapping, sparsity constraints and regularization terms are added between the encoder and decoder to encourage the autoencoder to learn more meaningful feature representations. The autoencoder uses the PyTorch deep learning framework for training and feature extraction.
[0073] Through the above scheme, a fingerprint generation algorithm combined with an autoencoder can be used to replace the encoding method of simply using the SHA-256 hash algorithm. By using this algorithm, the generated fingerprint will be more compact and meaningful, which improves the representativeness of the fingerprint. It can solve the problem of different encodings generated by the SHA-256 hash algorithm caused by the fact that the content of the data has not changed, but the format or encoding has changed, or the content of the data has changed slightly, but the semantics has not changed. It improves the robustness of the algorithm, can provide an additional layer of protection for the data, and increases the difficulty of tampering or forging the data.
[0074] In one embodiment of the present application, access rights to fingerprint data and raw data are strictly controlled to prevent data leakage, and all operations on the data, including fingerprint generation, comparison and repair processes, are recorded for tracking and auditing. Specifically, staff with operating authority obtain access to fingerprint data, raw data, fingerprint generation, comparison and repair rights through facial recognition. All access and operations to the data and the image path captured by facial recognition are stored in the log to facilitate auditing and tracing.
[0075] The above technical solution improves data security, and facilitates tracing of data access operation records through recorded logs.
Claims
1. A method for detecting consistency of provincial, municipal and county real estate registration data, characterized in that: The following steps are involved: S1. Extract the fields and tables that need to be compared for consistency and stored in provinces, cities and counties to create a new database; S2. Extract information fingerprints from each row of data in the table for the provincial and municipal level summary, municipal and county level storage data; The information fingerprint generation method in S2 is: S201, compile the provincial-level stored data that needs to be compared for consistency into a master table, city-level data tables, and county-level data tables; S202, summarizing the data stored at the municipal level that need to be compared for consistency into a master table, each municipal-level data table and the corresponding county-level data table; S203, using the SHA-256 algorithm to extract information fingerprints from each table, and obtaining the total table stored at the provincial and municipal levels, the municipal data table, the county data table, and the hash values of each row of the data table as information fingerprints; The algorithm used to generate information fingerprints is a fingerprint generation algorithm based on deep learning: Use the autoencoder to extract features from the table or the data rows in the table to obtain feature vectors, and then use the SHA-256 algorithm to encode the feature vectors to form information fingerprints. The specific steps are as follows: Preprocess the data table and convert it into a format that the neural network can process, a numerical matrix or a tensor; Construct an autoencoder network structure, including an encoder and a decoder; the encoder compresses the input data into a low-dimensional feature representation, and the decoder attempts to reconstruct the original data from the feature representation; Use tabular data for training and optimize the parameters of the autoencoder by minimizing the reconstruction error; After the autoencoder training is completed, the decoder is removed and only the encoder is retained for feature extraction; each row of data in the table is input into the encoder to obtain the feature vector representation of each row of data; For each feature vector, apply the SHA-256 hash algorithm to generate a fixed-length hash value; This hash value will serve as the digital fingerprint of the row of data; S3. Compare the information fingerprints summarized at the provincial level and the city level to see if they are consistent. If they are consistent, it means that all data are consistent; S4. If there is inconsistency, the information fingerprint of each city in the provincial data storage is compared with the corresponding city-level storage data, and the city with inconsistent information fingerprint is found for further comparison; S5. Compare the county-level information fingerprints of the city where the provincial-level stored data information fingerprints are inconsistent with the county-level information fingerprints of the data stored at the municipal level to find the counties where the information fingerprints are inconsistent; S6. Compare the data information fingerprints of each row of the county with inconsistent data stored at the provincial level with the data information fingerprints of each row of the county with inconsistent data stored at the municipal level, and find the rows with inconsistent information fingerprints, that is, find the rows with inconsistent data; S7. Repair or adjust the difference data according to the comparison results to ensure data consistency and record all operations on the data, including fingerprint generation, comparison and repair process, for tracking and auditing.
2. The detection method according to claim 1, characterized in that: Also includes: S8. Perform consistency comparison on city-level storage data and county-level storage data using the same approach.
3. The detection method according to claim 1, characterized in that: The information fingerprint generation method in S2 is: S201', use the SHA-256 algorithm to extract information fingerprints from each row of data that needs to be compared for consistency at the provincial and municipal levels, and store the obtained hash value as the information fingerprint in the county table according to the county corresponding to the data; S202', use the SHA-256 algorithm to extract the information fingerprint from the county table, and store the obtained hash value as the information fingerprint into the city table according to the city corresponding to the county; S203', use the SHA-256 algorithm to extract the information fingerprint from the city-level table, and store the obtained hash value as the information fingerprint in the provincial-level table; S204', use the SHA-256 algorithm to extract the information fingerprint from the provincial table and obtain a hash value as the information fingerprint.
4. The detection method according to claim 2, characterized in that: Strictly control access rights to fingerprint data and original data to prevent data leakage, and record all operations on data, including fingerprint generation, comparison and repair processes, for tracking and auditing.
5. The detection method according to claim 4, characterized in that: Access rights to fingerprint data and raw data are strictly controlled. Personnel with operational authority can obtain access to fingerprint data, raw data, fingerprint generation, comparison and repair through face recognition.
6. The detection method according to claim 5, characterized in that: All access and operations to data and the image paths captured by facial recognition are stored in logs to facilitate auditing and tracing.
7. The detection method according to claim 6, characterized in that: All access and operations to data and the image paths captured by facial recognition are stored in logs to facilitate auditing and tracing.
Citation Information
Patent Citations
Water conservancy general survey industry capacity data merging method based on k-nearest neighborhood
CN104657441A
Province-city-county real estate registration data consistency detection method based on information fingerprints
CN116894193A