A method and system for cleaning massive heterogeneous surveying and mapping big data
By generating and comparing feature codes, the problems of data redundancy and chaotic management in the storage of surveying and mapping big data are solved, and efficient cleaning of surveying and mapping big data and optimized management of storage space are achieved.
Patent Information
- Application Number
- CN202510255352.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-03-05
AI Technical Summary
In the surveying and mapping field and big data storage systems, there are problems of data redundancy, chaotic management, and inability to manage in a unified manner. In particular, the diversity of storage brands and methods for heterogeneous data leads to inefficient data integration and utilization.
By generating a feature code and comparing it, we determine whether the new feature code is the same as the old feature code. If they are the same, the database content is deleted. If they are not the same, the database content is not deleted. A deep decision learning algorithm model is used to generate the feature code, and the abnormal fields are trimmed through the feature information adjustment method.
It achieves efficient cleaning of surveying and mapping big data, optimizes storage space management, reduces data redundancy, and improves data utilization efficiency.
Smart Images

Figure CN120234304B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data and heterogeneous storage management, and in particular relates to a method and a data processing system for cleaning massive heterogeneous surveying and mapping big data. Background Art
[0002] In the field of surveying and mapping and in big data storage systems, there are many problems with data management.
[0003] First, in surveying and mapping, production and output data include large amounts of heterogeneous data, such as images and vectors. In daily work, in addition to raw and output data, process data is retained for varying project requirements. This data is highly continuous and has an irregular lifespan. As time passes and users or projects change, subsequent users struggle to determine whether the data still needs to be retained and hesitate to delete it. This leads to a continuous increase in stored data and an increasing shortage of storage space. Furthermore, process data is named by users without a unified standard, making it difficult to manage the long-term accumulation.
[0004] Secondly, for systems that store big data, storage brands and methods are diverse. For example, currently used storage includes five domestic and international brands: NetApp, EMC, VNX, Tongyou, and Sangfor. Four storage methods are used: distributed storage, object storage, block storage, and file storage. Unified planning and management cannot be achieved by relying on a single vendor's storage software, posing a significant challenge to the effective integration and efficient use of data. Summary of the Invention
[0005] The present invention provides a method and a data processing system for cleaning up massive heterogeneous surveying and mapping big data, which are used to solve the technical problems of data redundancy, chaotic data management and inability to uniformly manage data in surveying and mapping and big data storage management. By comparing new feature codes with old feature codes one by one, it is determined whether the new feature codes are the same as the old feature codes. If the new feature codes are the same as the old feature codes, the content in the database is deleted. If the new feature codes are different from the old feature codes, the content in the database is not deleted.
[0006] In order to achieve the above object, the present invention is implemented by the following technical solutions:
[0007] A method for cleaning up massive heterogeneous surveying and mapping big data processing includes the following steps:
[0008] Step 1): Scan a database within a short period of time and generate a plurality of feature codes corresponding to the database; wherein the plurality of feature codes corresponding to the database are recorded and the plurality of feature codes corresponding to the database are aggregated;
[0009] Step 2): Scan a database within a second time and generate multiple feature codes corresponding to the database; wherein the multiple feature codes corresponding to the database are recorded and the multiple feature codes corresponding to the database are aggregated;
[0010] Step 3): Compare the new signature code with the old signature code one by one to determine whether the new signature code is the same as the old signature code;
[0011] Step 4): If the new signature code is the same as the old signature code, the content in the database will be deleted. If the new signature code is different from the old signature code, the content in the database will not be deleted.
[0012] Optionally, in step 1) or in step 2), the feature code is generated based on a combination of three elements: the file name, the file size, and the access time;
[0013] After a database is scanned, multiple feature codes corresponding to the database are different.
[0014] Furthermore, the feature code is generated using the deep decision learning algorithm model of formula (1):
[0015] (1);
[0016] in, 、 and They are the file name of the variable of the signature, the file size of the variable, and the access time of the variable. To record multiple signatures, For a database, a database It is a database that is randomly scanned. For a database Scan and generate signatures. For a database After scanning, multiple signature codes are generated;
[0017] For data connection operation, multiple feature codes will be generated Transport to middle, To record multiple signatures, 、 and They are the determined file name, the determined file size, and the determined access time.
[0018] Optionally, in step 1), multiple feature codes are detected, and the specific detection steps are as follows:
[0019] Step a): Arranging the plurality of feature codes row by row in a row-by-row order; wherein the plurality of feature codes have serial numbers arranged row by row;
[0020] Step b): Scan the target object of each line of feature code to determine whether the target object of each line of feature code has an abnormal field;
[0021] Step c): when an abnormal field appears in the target object of the feature code of a row, the target object of the feature code of the row is trimmed;
[0022] In step c), a feature information adjustment method is used to adjust or modify the target object of the feature code. The feature information adjustment method is the following formula (2):
[0023] (2);
[0024] In formula (2), The location coordinates of the determined file name, To determine the file size positioning coordinates, To determine the location coordinates of the access time, To check the file name after it is determined, To determine the size of the file after the detection operation, To determine the access time after the detection operation, 、 and They are the determined file name, the determined file size, and the determined access time. To separate the determined file name from the positioning coordinates and detection operations, To determine the file size and position coordinates, separate the detection operation, To separate the access time after determination from the positioning coordinates and detection operations;
[0025] To detect the file name, To detect the file size after determination, The time of visit after the test is confirmed;
[0026] or This is an operation to trim the target object of a row's feature code.
[0027] Optionally, in step 2), multiple feature codes are detected, and the specific detection steps are as follows:
[0028] Step a'): multiple feature codes are arranged row by row in a row-by-row order; wherein the multiple feature codes have serial numbers arranged row by row;
[0029] Step b'): based on each serial number, detecting whether each row of feature codes is arranged in a row-by-row manner, and determining whether the target object of each row of feature codes has an abnormal field;
[0030] Step c'): when an abnormal field appears in the target object of the feature code of a row, the target object of the feature code of the row is adjusted.
[0031] Optionally, in step 3), the new signature code is compared with the old signature code by the following steps:
[0032] Step I): Compare the new feature code of each serial number with the old feature code of each serial number one by one to determine whether the new feature code of each serial number is the same as the old feature code of each serial number;
[0033] Step II): Compare the new feature code of each serial number with the old feature code of each serial number to see if they are the same. If they are the same, it is determined that the data is in use; if they are not the same, it is determined that the data is not in use.
[0034] Furthermore, in step II), the numerical model of the following formula (3) is used for comparison:
[0035] (3);
[0036] in, is the stored serial number, For the determination of the serial number, is the result of determining the serial number. is the new feature code, is the old feature code. is the distinguishing operation of the feature code, To distinguish the new feature code , To distinguish old signature codes , Represents a new feature code With the old signature same, Represents a new feature code With the old signature Not the same, is the feature code determination operation, To determine the new feature code With the old signature Are they the same?
[0037] or Represents the new feature code The serial number and the old signature code After determining the serial number, determine the new feature code With the old signature Are they the same?
[0038] Optionally, in step 4), if each new feature code is identical to each corresponding old feature code, the content in the database is completely deleted; if each new feature code is different from each old feature code, the content in the database is not deleted at all.
[0039] Optionally, in step 4), if part of the new feature code is identical to part of the corresponding old feature code or if part of the new feature code is different from part of the corresponding old feature code, the content in the database is not completely deleted.
[0040] A system for processing massive heterogeneous surveying and mapping big data, comprising:
[0041] A scanning module is used to scan a database in different time periods and generate multiple signature codes corresponding to the database in different time periods;
[0042] A feature code storage module is used to store multiple feature codes corresponding to different time periods, that is, to store multiple feature codes at the first rapid scan and multiple feature codes at the second rapid scan;
[0043] A signature analysis module is used to analyze whether multiple signatures corresponding to different time periods are the same;
[0044] An execution unit, configured to determine whether to delete data based on whether multiple feature codes corresponding to different time periods are the same;
[0045] The scanning module, the signature code storage module, the signature code analysis module and the execution unit are connected in sequence.
[0046] Beneficial effects of the present invention:
[0047] The present invention compares the new feature code with the old feature code one by one to determine whether the new feature code is the same as the old feature code. If the new feature code is the same as the old feature code, the content in the database will be deleted. If the new feature code is different from the old feature code, the content in the database will not be deleted. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] Figure 1 A schematic diagram of the storage principle of the multi-source heterogeneous data structure of the present invention;
[0050] Figure 2 A schematic diagram of the principle of comparing multiple feature codes corresponding to different time periods of the present invention;
[0051] Figure 3 Schematic diagram of the system structure of the present invention;
[0052] Figure 4 It is the workflow diagram of the present invention. DETAILED DESCRIPTION
[0053] The embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0054] Example 1;
[0055] like Figure 1 As shown in the figure, the storage status of multi-source heterogeneous data structure is that there are multiple databases for data storage (i.e. Figure 1 Multiple data storage), each data storage database is in a different server, FC storage mainly uses FC switches to establish a direct connection between FC storage and servers, achieving high-speed data transmission and reliable storage access, NAS storage uses NAS switches to establish a direct connection between NAS storage and servers, and then completely separates NAS storage from servers, and NAS storage centrally manages data.
[0056] The specific storage method of multi-source heterogeneous data is as follows Figure 1 As shown, data from two data storage databases are randomly stored in one FC storage, data from four data storage databases are randomly stored in another FC storage, and data from three data storage databases are randomly stored in a NAS storage.
[0057] Regardless of the storage method, multi-source data is randomly stored in a single FC storage or NAS storage, forming a multi-source heterogeneous data structure. Multi-source refers to the data in each data storage database, and heterogeneous refers to the databases in which the data is stored.
[0058] Example 2;
[0059] Based on Example 1, Figure 2 As shown, a database (i.e. Figure 1 There are multiple storage files in a certain data storage), and a certain storage file is scanned in the multiple storage files, and the certain storage file contains different contents.
[0060] Specifically, a storage file in a certain database is scanned. Since a certain storage file contains different contents, after scanning a certain storage file, multiple feature codes are generated at the same time according to the different contents contained in the certain storage file. Each feature code corresponds to the different contents contained in a certain storage file. The feature code is composed of the file name, file size (that is, the sizes of different contents in a certain storage file) and access time. Each feature code includes its own file name (that is, File01 to Filexx), and each feature code contains its own file size (and each file size is determined according to the sizes of different contents in a certain storage file, and the sizes of different contents are generally different). In addition, each feature code has an access time. The access time is the time each time a storage file in a certain database is scanned, and the time of each access or scan is the same in each feature code.
[0061] To further illustrate, if Figure 2 As shown in , a storage file in a database is scanned at a previous access time to obtain a scan file or multiple feature codes at the first quick scan (the scan file is obtained by HASH calculation). Each feature code at the first quick scan contains the respective file name (i.e., File01 to Filexx), each feature code contains the size of the respective different content, and each feature code contains a previous access time.
[0062] At a subsequent access time, a stored file in a database is scanned to obtain a scan file or multiple feature codes at a second quick scan (the scan file is obtained through HASH calculation). Each feature code at the second quick scan contains the respective file name (i.e., File01 to Filexx), each feature code contains the size of its respective different content, and each feature code contains a subsequent access time.
[0063] like Figure 2As shown, by comparing a feature code of a previous access time (which may be the feature code of the file name File01) with a feature code of a subsequent access time (i.e., the feature code of the file name File01) to see if they are the same, it is possible to determine whether a certain content corresponding to the file name File01 in a certain stored file has been modified. If a feature code of a previous access time (the feature code of the file name File01) is the same as a feature code of a subsequent access time (the feature code of the file name File01), then the certain content in the certain stored file has not been modified. If a feature code of a previous access time is different from a feature code of a subsequent access time, then the certain content in the certain stored file has been modified. In this way, by comparing a feature code of a previous access time with a feature code of a subsequent access time line by line in the order of the file names to see if they are the same, it is possible to determine whether the various contents in a certain stored file have been modified.
[0064] Example 3;
[0065] Based on Example 1-Example 2, as Figure 3 As shown, this embodiment provides a system for processing massive heterogeneous surveying and mapping big data, including:
[0066] A scanning module is used to scan a database in different time periods and generate multiple signature codes corresponding to the database in different time periods;
[0067] A feature code storage module is used to store multiple feature codes corresponding to different time periods, that is, to store multiple feature codes at the first rapid scan and multiple feature codes at the second rapid scan;
[0068] A signature analysis module is used to analyze whether multiple signatures corresponding to different time periods are the same;
[0069] An execution unit, configured to determine whether to delete data based on whether multiple feature codes corresponding to different time periods are the same;
[0070] The scanning module, the signature code storage module, the signature code analysis module and the execution unit are connected in sequence.
[0071] Example 4;
[0072] Based on Example 1-Example 3, as Figure 4 As shown, this embodiment provides a method for cleaning up massive heterogeneous surveying and mapping big data, including the following steps:
[0073] Step 1): Scan a database within a first time (a previous access time) and generate multiple feature codes corresponding to the database; wherein the multiple feature codes corresponding to the database are recorded and the multiple feature codes corresponding to the database are aggregated;
[0074] Step 2): Scan a database within a second time (a subsequent access time) and generate multiple feature codes corresponding to the database; wherein the multiple feature codes corresponding to the database are recorded and the multiple feature codes corresponding to the database are aggregated;
[0075] Step 3): The new feature code (i.e. Figure 2 The second time the file is scanned, the file name is File02) and the old signature code (i.e. Figure 2 The new signature code is compared with the old signature code one by one to determine whether the new signature code is the same as the old signature code;
[0076] In step 3), the signature code of the file named File01 after the first scan is compared with the signature code of the file named File01 after the second scan, the signature code of the file named File02 after the first scan is compared with the signature code of the file named File02 after the second scan, and so on.
[0077] Step 4): If the new signature code is the same as the old signature code, the content in the database will be deleted. If the new signature code is different from the old signature code, the content in the database will not be deleted.
[0078] Example 5;
[0079] Based on Example 4, Figure 2 As shown, in step 1) or in step 2), the feature code is generated based on a combination of three elements: the file name, the file size (i.e., the size of each different content in a certain stored file), and the access time.
[0080] After scanning a database, the corresponding multiple feature codes in the database are different (because the file names and file sizes of the feature codes are different, that is, the sizes of different contents in a storage file are different).
[0081] Furthermore, the signature code is generated by using the deep decision learning algorithm model of formula (1). Formula (1) is the operation model for generating a signature code each time a database is scanned:
[0082] (1);
[0083] in, 、 and They are the file name of the variable of the signature (because the file name of each signature is different), the file size of the variable (the size of each content in a storage file is not necessarily the same) and the access time of the variable (because the access time is different each time). To record multiple signatures (i.e., multiple signatures for different file names), For a database, a database It is a database that is randomly scanned. For a database Scan and generate signatures. For a database After scanning, multiple signature codes are generated;
[0084] For data connection operation, multiple feature codes will be generated Transport to middle, To record multiple signatures, 、 and They are the determined file name (i.e. the file name of each signature code is arranged in sequence after each scan), the determined file size (i.e. the different sizes of each content in the stored file after each scan correspond to their respective file sizes) and the determined access time (i.e. the determined scanning time for each time).
[0085] Specifically, in step 1), multiple feature codes are detected, and the specific detection steps are as follows:
[0086] Step a): Arranging the plurality of signature codes line by line (i.e., in the order of the file name); wherein the plurality of signature codes have serial numbers arranged line by line (i.e., the serial numbers of the file names, such as File01, File02...Filexx);
[0087] Step b): Scan the target object of each line of feature code (i.e., the file name, file size, and access time of each feature code) to determine whether the target object of each line of feature code (i.e., each feature code) has an abnormal field;
[0088] Step c): when an abnormal field appears in the target object of the feature code of a certain row, the target object of the feature code of the certain row (i.e., a certain feature code) is trimmed;
[0089] In step c), a feature information adjustment method is used to adjust or trim the target object of the feature code. The feature information adjustment method is the following formula (2). Formula (2) is an adjustment or trimming performed on a feature code that has an abnormality:
[0090] (2);
[0091] In formula (2), The location coordinates of the determined file name (i.e. the file name where an abnormal field of a certain feature code appears). The location coordinates of the determined file size (i.e. the file size where an abnormal field of a certain feature code appears) are located. The location coordinates of the access time after determination (i.e., the access time of the field where an abnormality occurs in a certain feature code). The detection operation for the determined file name (that is, the file name of each signature code is arranged in order after each scan) To determine the file size (i.e. the different sizes of the contents in the stored file correspond to the corresponding file sizes after each scan), To determine the subsequent access time detection operation (i.e., to determine the scanning time for each time), 、 and They are the determined file name, the determined file size, and the determined access time. To separate the determined file name from the positioning coordinates and detection operations, To separate the file size from the positioning coordinates and detection operations, To separate the access time after determination from the positioning coordinates and detection operations;
[0092] To detect the file name, To detect the file size after determination, The time of visit after the test is confirmed;
[0093] After each scan, the file names of each signature are arranged in order. Locate the file name of a feature code in the abnormal field (because the file name of the abnormal field can be modified), After each scan, the different sizes of each content in the file are stored and mapped to their respective file sizes. Locate the file size of a feature code in the abnormal field (because the abnormal field of the file size contains text content, the abnormal field of the file size is modified, the abnormal field of the file size changes, and the file size also changes). To pass Locate the access time of a feature code where an abnormal field appears (because an abnormal field appears at the access time of a feature code).
[0094] or To modify the target object of a row's feature code, The file name of a feature code located in an abnormal field is trimmed, or the file size of a feature code located in an abnormal field is trimmed, or the access time of a feature code located in an abnormal field is trimmed.
[0095] In step 2), multiple feature codes are detected, and the specific detection steps are as follows:
[0096] Step a'): multiple signatures are arranged line by line (i.e., in the order of the file name); wherein the multiple signatures have serial numbers arranged line by line (i.e., the serial numbers of the file names, such as File01, File02...Filexx);
[0097] Step b'): according to each serial number, detect whether each row of feature codes is arranged in a row-by-row manner, and determine whether the target object of each row of feature codes has an abnormal field, that is, this step b') is the same as step b);
[0098] Step c'): when an abnormal field appears in the target object of the feature code of a row, the target object of the feature code of the row is adjusted, that is, this step c') is the same as step c).
[0099] In step 3), the new signature code is compared with the old signature code by the following steps:
[0100] Step I): Compare the new feature code of each serial number with the old feature code of each serial number one by one to determine whether the new feature code of each serial number is the same as the old feature code of each serial number;
[0101] Step II): Compare the new feature code of each serial number with the old feature code of each serial number to see if they are the same. If they are the same, it is determined that the data is in use; if they are not the same, it is determined that the data is not in use.
[0102] Furthermore, in step II), the numerical model of the following formula (3) is used for comparison:
[0103] (3);
[0104] in, The stored serial number is (the serial number of File01, File02...Filexx of the new signature code and the old signature code). For the determination of the serial number (i.e. the determination of the new feature code File01 and the old feature code File01), is the result of the serial number determination (i.e. the new feature code File01 corresponds to the old feature code File01). is the new feature code, is the old feature code. It is the distinguishing operation of the signature code (distinguished by scanning time), To distinguish the new feature code , To distinguish old signature codes , Represents a new feature code With the old signature same, Represents a new feature code With the old signature Not the same, is the feature code determination operation, To determine the new feature code With the old signature Are they the same?
[0105] or Represents the new feature code The serial number and the old signature code After determining the serial number, determine the new feature code With the old signature Are they the same?
[0106] Formula (3) still uses the feature code of the file named File01 after the first scan to compare with the feature code of the file named File01 after the second scan, the feature code of the file named File02 after the first scan to compare with the feature code of the file named File02 after the second scan, and so on.
[0107] Example 6;
[0108] In step 4), if each new feature code is identical to each corresponding old feature code, the content in the database is completely deleted; if each new feature code is different from each old feature code, the content in the database is not deleted at all.
[0109] In step 4), if part of the new feature code is identical to part of the corresponding old feature code or if part of the new feature code is different from part of the corresponding old feature code, the content in the database is not completely deleted.
[0110] In both steps 4) of this embodiment, the following operations are performed: if the signature code of the file named File01 scanned at the first time is compared with the signature code of the file named File01 scanned at the second time, and they are the same, then the content corresponding to the signature code of File01 in the certain stored file is deleted; if the signature code of the file named File01 scanned at the first time is compared with the signature code of the file named File01 scanned at the second time, and they are different, then the content corresponding to the signature code of File01 in the certain stored file is not deleted. The operation is repeated in this manner.
[0111] The present invention accurately compares the characteristic codes obtained by scanning at different times, and determines whether changes such as addition, deletion or modification have occurred to files in a certain database through the comparison results, and further infers whether the files in a certain database are still in use. Specifically, a comprehensive traversal scan and comparison is performed on all files in a certain database or a certain volume to obtain the data changes of the volume, so as to determine whether the volume is still in use. If the characteristic code of a certain volume has not changed after multiple scans, the volume is determined to be a "dead volume", that is, it is no longer in use. At this time, through human-computer collaboration, the server where the volume is mounted and the corresponding usage location are targeted to be found, and then the volume data is deleted after confirmation, thereby freeing up storage space and achieving optimized management of data storage.
[0112] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope of the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A method for cleaning up massive heterogeneous surveying and mapping big data, characterized in that: The steps include: Step 1): Scan a database within a short period of time and generate a plurality of feature codes corresponding to the database; wherein the plurality of feature codes corresponding to the database are recorded and the plurality of feature codes corresponding to the database are aggregated; Step 2): Scan a database within a second time and generate multiple feature codes corresponding to the database; wherein the multiple feature codes corresponding to the database are recorded and the multiple feature codes corresponding to the database are aggregated; In step 1) or step 2), the feature code is generated by using a deep decision learning algorithm model of formula (1): (1); in, 、 and are the file name of the variable of the feature code, the file size of the variable and the access time of the variable, To record multiple signatures, For a database, a database It is a database that is randomly scanned. For a database Scan and generate signatures. For a database After scanning, multiple signature codes were generated; For data connection operation, multiple feature codes will be generated Transport to middle, To record multiple signatures, 、 and They are respectively the determined file name, the determined file size and the determined access time; Step 3): Compare the new signature code with the old signature code one by one to determine whether the new signature code is the same as the old signature code; In step 3), the comparison between the new feature code and the old feature code is performed by the following steps: Step I): Compare the new feature code of each serial number with the old feature code of each serial number one by one to determine whether the new feature code of each serial number is the same as the old feature code of each serial number; Step II): Compare the new signature code of each serial number with the old signature code of each serial number to see if they are the same. If they are the same, it is determined that the data is in use; if they are not the same, it is determined that the data is not in use; In step II), the numerical model of the following formula (3) is used for comparison: (3); in, is the stored serial number, For the determination of the serial number, is the result of determining the serial number. is the new feature code, is the old feature code, is the distinguishing operation of the feature code, To distinguish the new feature code , To distinguish old signature codes , Represents a new feature code With the old signature same, Represents a new feature code With the old signature Different, is the feature code determination operation, To determine the new feature code With the old signature Are they the same? or Represents the new feature code The serial number and the old signature code After determining the serial number, determine the new feature code With the old signature Are they the same? Step 4): If the new signature code is the same as the old signature code, the content in the database will be deleted. If the new signature code is different from the old signature code, the content in the database will not be deleted.
2. A method for cleaning up massive heterogeneous surveying and mapping big data according to claim 1, characterized in that: In step 1) or step 2), the feature code is generated based on a combination of three elements: file name, file size, and access time; After scanning a certain database, the corresponding multiple feature codes in the certain database are different.
3. The method for cleaning up massive heterogeneous surveying and mapping big data according to claim 1, characterized in that: In step 1), the multiple feature codes are detected, and the specific detection steps are as follows: Step a): Arranging the plurality of feature codes row by row in a row-by-row order; wherein the plurality of feature codes have serial numbers arranged row by row; Step b): Scan the target object of each line of feature code to determine whether the target object of each line of feature code has an abnormal field; Step c): when an abnormal field appears in the target object of the feature code of a row, the target object of the feature code of the row is trimmed; In step c), a feature information adjustment method is used to adjust or modify the target object of the feature code. The feature information adjustment method is the following formula (2): (2); In formula (2), The location coordinates of the determined file name, To determine the file size positioning coordinates, To determine the location coordinates of the access time, To check the file name after it is determined, To determine the size of the file after the detection operation, To determine the access time after the detection operation, 、 and They are the determined file name, the determined file size, and the determined access time. To separate the determined file name from the positioning coordinates and detection operations, To determine the file size and position coordinates, separate the detection operation, To separate the access time after determination from the positioning coordinates and detection operations; To detect the file name, To detect the file size after determination, The time of visit after the test is confirmed; or This is an operation to trim the target object of a row's feature code.
4. A method for cleaning up massive heterogeneous surveying and mapping big data according to claim 1, characterized in that: In step 2), the multiple feature codes are detected, and the specific detection steps are as follows: Step a'): multiple feature codes are arranged row by row in a row-by-row order; wherein the multiple feature codes have serial numbers arranged row by row; Step b'): based on each serial number, detecting whether each row of feature codes is arranged in a row-by-row manner, and determining whether the target object of each row of feature codes has an abnormal field; Step c'): when an abnormal field appears in the target object of the feature code of a row, the target object of the feature code of the row is adjusted.
5. The method for cleaning up massive heterogeneous surveying and mapping big data according to claim 1, characterized in that: In step 4), if each new feature code is identical to each corresponding old feature code, the content in the database is completely deleted; if each new feature code is different from each old feature code, the content in the database is not deleted at all.
6. The method for cleaning up massive heterogeneous surveying and mapping big data according to claim 1, characterized in that: In step 4), if part of the new feature codes is identical to part of the corresponding old feature codes or if part of the new feature codes is different from part of the corresponding old feature codes, the content in the database is not completely deleted.
7. A system for processing massive heterogeneous surveying and mapping big data, configured to execute a method for processing massive heterogeneous surveying and mapping big data according to any one of claims 1 to 6, characterized in that: include: A scanning module is used to scan a database in different time periods and generate multiple signature codes corresponding to the database in different time periods; A feature code storage module is used to store multiple feature codes corresponding to different time periods, that is, to store multiple feature codes at the first rapid scan and multiple feature codes at the second rapid scan; A signature analysis module is used to analyze whether multiple signatures corresponding to different time periods are the same; An execution unit, configured to determine whether to delete data based on whether multiple feature codes corresponding to different time periods are the same; The scanning module is connected sequentially through the feature code storage module, the feature code analysis module and the execution unit in a circuit manner.
Citation Information
Patent Citations
Full distributed duplicate copy positioning method for data grid
CN101251843A
Method for deleting duplicated data in file system in real time
CN101908073A