A method, device, computer equipment and medium for deduplicating watershed business data

By splicing and fusion feature selection of the header attributes of multi-table databases, combining hash value calculation and normal distribution processing, efficient deduplication of massive multi-source data is achieved, solving the problem of low deduplication efficiency, saving storage overhead and improving data quality.

CN119441209BActive Publication Date: 2025-05-16THREE GORGES GROUP IND DEVELOPMENT (BEIJING) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510039746.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-16
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

When processing massive multi-source data, the efficiency of the deduplication process decreases significantly, resulting in an increase in storage overhead.

Method used

By obtaining the header attributes of the multi-table database for splicing, the attribute document is obtained; the attribute document and the data table are selected in a fusion feature to obtain the target attribute combination; the target attribute combination is hashed, and the hash distribution is approximately processed to a normal distribution; the hash judgment is used for hashing to achieve data deduplication.

Benefits of technology

It greatly improves the deduplication efficiency of multi-source data, eliminates redundant data, saves system storage overhead, and reduces the overall data volume, making it easier to retrieve and analyze subsequent data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119441209B_ABST
    Figure CN119441209B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data deduplication, and discloses a method, device, computer equipment and medium for deduplication of watershed business data, the method comprising: obtaining a multi-table database corresponding to watershed business data, splicing the header attributes of the multi-table database to obtain an attribute document; performing fusion feature selection on the attribute document and the data table in the multi-table database to obtain a target attribute combination corresponding to each data table; performing hash value calculation on the target attribute combination corresponding to each data table to obtain a hash value distribution corresponding to each data table, and approximating the hash value distribution to a normal distribution; obtaining watershed business data to be processed, performing hash determination on the watershed business data to be processed using a normal distribution, performing data deduplication based on the hash determination result, and obtaining a deduplication result of the watershed business data. The present invention greatly improves the deduplication efficiency of watershed business data, eliminates redundant data, and saves system storage overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data deduplication, and in particular to a method, device, computer equipment and medium for deduplication of watershed business data. Background Art

[0002] In the process of building a watershed business management system, the sources of watershed business data are wide-ranging and the amount of data is huge, forming a multi-table database. The redundancy of massive multi-source data causes a sharp increase in data storage overhead. Therefore, it is urgent to deduplicate data to achieve storage efficiency and improve data quality of the watershed business management system.

[0003] When processing massive amounts of multi-source data, the deduplication process requires a lot of time and computing resources, resulting in a significant decrease in deduplication efficiency. Summary of the invention

[0004] In view of this, the present invention provides a method, apparatus, computer equipment and medium for deduplication of watershed business data to solve the problem of significantly reduced deduplication efficiency in processing massive multi-source data.

[0005] In a first aspect, the present invention provides a method for deduplicating watershed service data, the method comprising:

[0006] Obtain the multi-table database corresponding to the watershed business data, and concatenate the header attributes of the multi-table database to obtain the attribute document;

[0007] Perform fusion feature selection on attribute documents and data tables in multi-table databases to obtain the target attribute combination corresponding to each data table;

[0008] Calculate the hash value of the target attribute combination corresponding to each data table to obtain the hash value distribution corresponding to each data table, and approximate the hash value distribution to a normal distribution;

[0009] Obtain the watershed business data to be processed, perform hash judgment on the watershed business data to be processed using normal distribution, deduplicate the data based on the hash judgment result, and obtain the deduplication result of the watershed business data.

[0010] The present embodiment provides a method for deduplicating watershed business data, which splices header attributes of a multi-table database to obtain an attribute document; performs fusion feature selection on the attribute document and the data tables in the multi-table database to obtain a target attribute combination corresponding to each data table; performs hash value calculation on the target attribute combination corresponding to each data table to obtain a hash value distribution corresponding to each data table, and approximates the hash value distribution to a normal distribution; performs hash judgment on the watershed business data to be processed using the normal distribution, and performs data deduplication based on the hash judgment result to obtain a watershed business data deduplication result; by splicing header attributes into an attribute document, and by fusing algorithms and characteristics such as feature selection, distribution conversion, and normal distribution, the deduplication efficiency of multi-source data is greatly improved, redundant data is eliminated, system storage overhead is saved, and the overall data volume is reduced to facilitate subsequent data retrieval and data analysis.

[0011] In an optional implementation, header attributes of a multi-table database are concatenated to obtain an attribute document, including:

[0012] The field names of the same header attributes in multi-table databases are processed uniformly to obtain the attribute names corresponding to the data tables;

[0013] Concatenate the attribute names corresponding to the data table one by one to obtain the attribute document.

[0014] The present embodiment provides a method for deduplicating watershed business data. The method unifies the field names of the same table header attributes so that the attribute names of all database tables remain consistent, thereby facilitating item-by-item splicing of the attribute names corresponding to the data tables, and converting the key attribute selection into a document feature selection problem, thereby simplifying the data deduplication process and improving the deduplication efficiency of watershed business data.

[0015] In an optional implementation, fusion feature selection is performed on the attribute document and the data table in the multi-table database to obtain a target attribute combination corresponding to each data table, including:

[0016] Use keyword extraction algorithm to select key attributes of attribute documents and obtain key attribute name combinations;

[0017] For the data table that does not contain the key attribute name, randomly select an attribute name and add it to the key attribute name combination to obtain a combination containing the attribute names of all data tables;

[0018] The attribute name combination of each data table in the multi-table database is obtained, and the target attribute combination corresponding to each data table is determined based on the attribute name combination of each data table and the combination of attribute names of all data tables.

[0019] The present embodiment provides a method for deduplicating watershed business data. By performing key attribute selection on attribute documents, key features of documents are extracted. For data tables that do not contain key attribute names, an attribute name is randomly selected and added to the key attribute name combination to obtain a combination containing attribute names of all data tables, thereby obtaining attribute name combinations of all data tables. Finally, based on the attribute name combination of each data table and the combination containing attribute names of all data tables, a target attribute combination corresponding to each data table is determined. Overlapping data between the attribute name combination of each data table and the combination of attribute names of all data tables is extracted, thereby calculating the target attribute combination corresponding to each data table. This further reduces the amount of data in each data table, extracts key attributes in each data table, lays a foundation for subsequent hash value calculation, and greatly improves the deduplication efficiency of multi-source data.

[0020] In an optional implementation, a hash value is calculated for a target attribute combination corresponding to each data table to obtain a hash value distribution corresponding to each data table, and the hash value distribution is approximated to a normal distribution, including:

[0021] Set hash functions for target attribute combinations corresponding to each data table respectively;

[0022] Using a hash function to calculate hash values ​​of the data in the target attribute combination corresponding to each data table, respectively, to obtain hash values ​​of multiple data entries corresponding to each data table;

[0023] Use a single linked list to link data entries with the same hash value to obtain a hash value linked list, and store the head node of the single linked list in the hash table;

[0024] The hash value distribution corresponding to each data table is determined based on the hash values ​​of the plurality of data entries, and the hash value distribution is approximated as a normal distribution.

[0025] The present embodiment provides a method for deduplication of watershed business data. A hash function is set for the target attribute combination corresponding to each data table, and then the hash function is used to calculate the hash value of the data in the target attribute combination corresponding to each data table, to obtain the hash values ​​of multiple data entries corresponding to each data table. The hash values ​​of the multiple data entries are used as discrete random variables, the deduplication problem is converted into a probability distribution problem, and the hash value distribution is approximated as a normal distribution, which lays a foundation for subsequent hash judgment of the watershed business data to be processed using the normal distribution, and greatly improves the deduplication efficiency of multi-source data.

[0026] In an optional implementation, hash determination is performed on the watershed service data to be processed using normal distribution, and data deduplication is performed based on the hash determination result to obtain a watershed service data deduplication result, including:

[0027] Obtain the hash value of the attribute column of the watershed business data to be processed, and determine whether the hash value of the attribute column of the watershed business data to be processed is in the horizontal axis interval corresponding to the normal distribution;

[0028] If the hash value of the attribute column of the watershed business data to be processed is within the horizontal axis interval, the Bloom filter is used to query each data in the hash value linked list in turn;

[0029] If there are identical data entries in the hash value linked list, the basin business data to be processed is deleted to obtain the deduplication result of the basin business data.

[0030] The present embodiment provides a method for deduplicating watershed business data. By judging whether the hash value of the attribute column of the watershed business data to be processed is in the horizontal axis interval corresponding to the normal distribution, it is determined whether there is data identical to the watershed business data to be processed in the multi-table database, and then the watershed business data to be processed is deleted according to the judgment result. The "3σ" principle of the normal distribution is utilized. Since the data falling in the horizontal axis interval corresponding to the normal distribution is judged to be a low-probability event, the hash judgment is performed using the horizontal axis interval corresponding to the normal distribution to achieve the purpose of efficient deduplication and save system storage overhead.

[0031] In an optional implementation, using normal distribution to perform hash determination on the watershed service data to be processed, performing data deduplication based on the hash determination result, and obtaining a watershed service data deduplication result, further comprising:

[0032] If the hash value of the attribute column of the watershed business data to be processed is not in the horizontal axis interval, or if the same data entry does not exist in the hash value linked list, the watershed business data to be processed is added to the multi-table database, hash table and normal distribution data set.

[0033] In a second aspect, the present invention provides a device for deduplicating watershed service data, the device comprising:

[0034] The splicing module is used to obtain the multi-table database corresponding to the watershed business data, splice the header attributes of the multi-table database, and obtain the attribute document;

[0035] A fusion feature selection module is used to perform fusion feature selection on attribute documents and data tables in a multi-table database to obtain a target attribute combination corresponding to each data table;

[0036] A hash value calculation module is used to calculate the hash value of the target attribute combination corresponding to each data table, obtain the hash value distribution corresponding to each data table, and approximate the hash value distribution to a normal distribution;

[0037] The hash determination module is used to obtain the watershed business data to be processed, perform hash determination on the watershed business data to be processed using normal distribution, deduplicate the data based on the hash determination result, and obtain the deduplication result of the watershed business data.

[0038] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, computer instructions being stored in the memory, and the processor executing the computer instructions to execute a method for deduplicating watershed business data according to the first aspect or any corresponding embodiment thereof.

[0039] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute a method for deduplicating watershed business data according to the first aspect or any corresponding embodiment thereof.

[0040] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, wherein the computer instructions are used to enable a computer to execute a method for deduplicating watershed business data according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0042] Figure 1 It is a flow chart of a method for deduplication of watershed service data according to an embodiment of the present invention;

[0043] Figure 2 is a flow chart of another method for deduplication of watershed service data according to an embodiment of the present invention;

[0044] Figure 3 is a flow chart of another method for deduplication of watershed service data according to an embodiment of the present invention;

[0045] Figure 4 is a structural block diagram of a device for deduplicating watershed service data according to an embodiment of the present invention;

[0046] Figure 5 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0048] An embodiment of the present invention provides a method for deduplication of watershed business data. By concatenating the header attribute names into documents, key attribute selection is converted into a document feature selection problem, and the hash value of the extracted feature is calculated as a discrete random variable. The deduplication problem is converted into a probability distribution problem, and according to the "3σ" principle of normal distribution, the deduplication efficiency of watershed business data based on a multi-table database is improved, and the storage overhead problem caused by multi-source data redundancy is solved.

[0049] The embodiment of the present invention provides a method for deduplicating watershed business data. It should be noted that the method for deduplicating watershed business data provided by the embodiment of the present invention may be executed by a device for deduplicating watershed business data. The device for deduplicating watershed business data may be implemented as part or all of an electronic device through software, hardware, or a combination of software and hardware. The electronic device may be a server or a terminal. The server in the embodiment of the present application may be a single server or a server cluster composed of multiple servers. The terminal in the embodiment of the present application may be a system server. The system server is integrated in other intelligent hardware devices such as smart phones, personal computers, tablet computers, wearable devices, and intelligent robots. In the following method embodiments, the execution subject is an electronic device as an example for explanation.

[0050] According to an embodiment of the present invention, an embodiment of a method for deduplication of watershed business data is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0051] In this embodiment, a method for deduplicating flow domain service data is provided, which can be used in the above-mentioned electronic device. Figure 1 is a flow chart of a method for deduplicating watershed service data according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0052] Step S101, obtaining a multi-table database corresponding to the watershed business data, concatenating the header attributes of the multi-table database to obtain an attribute document.

[0053] Specifically, the header attributes of the multi-table database include: river code, name, administrative division, section code, station code, station name, station type, water level, flow, rainfall, warning water level, warning flow, guaranteed water level, guaranteed flow, etc.

[0054] Step S102: performing fusion feature selection on the attribute document and the data table in the multi-table database to obtain a target attribute combination corresponding to each data table.

[0055] Step S103 , performing hash value calculation on the target attribute combination corresponding to each data table to obtain the hash value distribution corresponding to each data table, and approximating the hash value distribution to a normal distribution.

[0056] Specifically, the Central Limit Theorem (CLT) is applied to convert all hash values ​​in the same data table into ( ) The corresponding hash value distribution is approximately treated as a normal distribution ,in, is the expected value of the normal distribution, is the standard deviation of the normal distribution.

[0057] Step S104, obtaining the watershed business data to be processed, performing hash determination on the watershed business data to be processed using normal distribution, performing data deduplication based on the hash determination result, and obtaining the watershed business data deduplication result.

[0058] Specifically, according to the "3σ" principle of normal distribution, it is determined whether the watershed business data to be processed falls within the horizontal axis interval corresponding to the normal distribution, and then it is determined whether to delete the watershed business data to be processed.

[0059] The present embodiment provides a method for deduplicating watershed business data, which splices header attributes of a multi-table database to obtain an attribute document; performs fusion feature selection on the attribute document and the data tables in the multi-table database to obtain a target attribute combination corresponding to each data table; performs hash value calculation on the target attribute combination corresponding to each data table to obtain a hash value distribution corresponding to each data table, and approximates the hash value distribution to a normal distribution; performs hash judgment on the watershed business data to be processed using the normal distribution, and performs data deduplication based on the hash judgment result to obtain a watershed business data deduplication result; by splicing header attributes into an attribute document, and by fusing algorithms and characteristics such as feature selection, distribution conversion, and normal distribution, the deduplication efficiency of multi-source data is greatly improved, redundant data is eliminated, system storage overhead is saved, and the overall data volume is reduced to facilitate subsequent data retrieval and data analysis.

[0060] In this embodiment, a method for deduplicating flow domain service data is provided, which can be used in the above-mentioned electronic device. Figure 2 is a flow chart of a method for deduplicating watershed service data according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:

[0061] Step S201, obtaining a multi-table database corresponding to the watershed business data, and concatenating the header attributes of the multi-table database to obtain an attribute document.

[0062] Specifically, the above step S201 includes:

[0063] Step S2011, uniformly process the field names of the same table header attribute in the multi-table database to obtain the attribute name corresponding to the data table.

[0064] Specifically, the header attributes of a multi-table database are processed through manual screening and modification, and the field names of the same header attributes in different data tables are unified.

[0065] Step S2012, concatenate the attribute names corresponding to the data table item by item to obtain an attribute document.

[0066] Specifically, the attribute names of each data item in the data table are concatenated one by one, and spaces are used as separators to integrate them into a document, namely, an attribute document.

[0067] Step S202: performing fusion feature selection on the attribute document and the data table in the multi-table database to obtain a target attribute combination corresponding to each data table.

[0068] Specifically, the above step S202 includes:

[0069] Step S2021, using a keyword extraction algorithm to select key attributes of the attribute document to obtain a key attribute name combination.

[0070] Specifically, the keyword extraction algorithm uses TF-IDF (Term Frequency-Inverse Document Frequency, a commonly used weighting technology for information retrieval and text mining) to select the key attributes of attribute documents through TF-IDF, and obtain the key attribute name combination, which is recorded as .

[0071] Step S2022: for a data table that does not contain a key attribute name, randomly select an attribute name and add it to the key attribute name combination to obtain a combination containing attribute names of all data tables.

[0072] Specifically, for a data table that does not contain a key attribute name, a random attribute name is selected and added to the key attribute name combination C to obtain a combination containing all data table attribute names. The combination containing all data table attribute names is recorded as .

[0073] Step S2023, obtaining the attribute name combination of each data table in the multi-table database, and determining the target attribute combination corresponding to each data table based on the attribute name combination of each data table and the combination of attribute names of all data tables.

[0074] Specifically, assuming that the data table in the multi-table database The attribute name combination is ,in, , L is the number of data tables in the multi-table database, k is the number of data tables The number of attribute columns.

[0075] Furthermore, the data table Corresponding target attribute combination It can be expressed as:

[0076] (1)

[0077] Step S203, perform hash value calculation on the target attribute combination corresponding to each data table, obtain the hash value distribution corresponding to each data table, and approximate the hash value distribution to a normal distribution. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.

[0078] Step S204, obtain the watershed service data to be processed, perform hash determination on the watershed service data to be processed using normal distribution, perform data deduplication based on the hash determination result, and obtain the deduplication result of the watershed service data. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.

[0079] The present embodiment provides a method for deduplication of watershed business data. The method uniformly processes the field names of the same header attributes so that the attribute names of all database tables are consistent, which facilitates the item-by-item splicing of the attribute names corresponding to the data tables, converts the key attribute selection into a document feature selection problem, simplifies the data deduplication process, and improves the deduplication efficiency of watershed business data. Secondly, by performing key attribute selection on the attribute document, the key features of the document are extracted, and for the data table that does not contain the key attribute name, a random attribute name is selected and added to the key attribute name combination to obtain a combination containing the attribute names of all data tables, thereby obtaining the attribute name combinations of all data tables. Finally, based on the attribute name combination of each data table and the combination containing the attribute names of all data tables, the target attribute combination corresponding to each data table is determined, the overlapping data of the attribute name combination of each data table and the combination of the attribute names of all data tables is extracted, and the calculation of the target attribute combination corresponding to each data table is realized, thereby further reducing the data volume of each data table, realizing the extraction of key attributes in each data table, laying a foundation for subsequent hash value calculation, and greatly improving the deduplication efficiency of multi-source data.

[0080] In this embodiment, a method for deduplicating flow domain service data is provided, which can be used in the above-mentioned electronic device. Figure 3 is a flow chart of a method for deduplicating watershed service data according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:

[0081] Step S301, obtain the multi-table database corresponding to the watershed business data, and concatenate the header attributes of the multi-table database to obtain an attribute document. Figure 2 Step S201 of the illustrated embodiment will not be described in detail here.

[0082] Step S302: perform fusion feature selection on the attribute document and the data table in the multi-table database to obtain the target attribute combination corresponding to each data table. Figure 2 Step S202 of the illustrated embodiment will not be described in detail here.

[0083] Step S303 , performing hash value calculation on the target attribute combination corresponding to each data table, obtaining the hash value distribution corresponding to each data table, and approximating the hash value distribution to a normal distribution.

[0084] Specifically, the above step S303 includes:

[0085] Step S3031, setting a hash function for each target attribute combination corresponding to each data table.

[0086] Specifically, a separate hash function is set for each target attribute combination. , that is, set a separate hash function for each data table .

[0087] Step S3032: Use a hash function to calculate hash values ​​of the data in the target attribute combination corresponding to each data table, and obtain hash values ​​of multiple data items corresponding to each data table.

[0088] Specifically, for the data table , combining the target attributes Each row of data ( j The data row number) is taken as a whole parameter into the hash function In the above example, the hash value of each piece of data is calculated item by item. ( ).

[0089] For example, suppose =[River code, river name, administrative division, section code, station code, station name, station type, water level, flow, rainfall, warning water level, warning flow, guaranteed water level, guaranteed flow, sediment carrying capacity], =[section code, sediment carrying capacity, date], then =[section code, sediment carrying capacity]; Set the hash function to: =5i+3, data table The data in is: [(101,12,20240101),(102,11,20240101),(101,11,20240102),(102,13,20240102)], then The hash value is: [(101,12),(102,11),(101,11),(102,13)] The calculation formula is: h([101,12])=5 10112+3=50563, and the number of other data entries is similar.

[0090] Step S3033: Use a single linked list to link data entries with the same hash value to obtain a hash value linked list, and store the head node of the single linked list in the hash table.

[0091] Specifically, the same hash value ( ) are linked together through a single linked list, and the head node of each linked list is stored in the hash table. The hash value linked list at this time is also called a hash bucket.

[0092] Step S3034, determining the hash value distribution corresponding to each data table based on the hash values ​​of the plurality of data entries, and approximating the hash value distribution to a normal distribution.

[0093] Step S304, obtaining the watershed business data to be processed, performing hash determination on the watershed business data to be processed using normal distribution, performing data deduplication based on the hash determination result, and obtaining the watershed business data deduplication result.

[0094] Specifically, the above step S304 includes:

[0095] Step S3041, obtaining the hash value of the attribute column of the watershed business data to be processed, and determining whether the hash value of the attribute column of the watershed business data to be processed is in the horizontal axis interval corresponding to the normal distribution.

[0096] Specifically, the horizontal axis interval corresponding to the normal distribution is (μ-3σ,μ+3σ).

[0097] Step S3042: If the hash value of the attribute column of the watershed service data to be processed is within the horizontal axis interval, the Bloom filter is used to query each piece of data in the hash value linked list in sequence.

[0098] Specifically, if the hash value of the attribute column of the watershed business data to be processed is in the horizontal axis interval, it is considered that the watershed business data to be processed already exists, and then the watershed business data to be processed is input into the Bloom filter, and each data in the corresponding hash value linked list (hash bucket) is queried in turn.

[0099] Step S3043: If the same data entry exists in the hash value linked list, the watershed business data to be processed is deleted to obtain the deduplication result of the watershed business data.

[0100] Specifically, if the hash value of the attribute column of the watershed business data to be processed is not in the horizontal axis interval, or if the same data entry does not exist in the hash value linked list, the watershed business data to be processed is added to the multi-table database, hash table and normal distribution data set.

[0101] Furthermore, if the hash value of the attribute column corresponding to the watershed business data to be processed does not fall within the horizontal axis interval (μ-3σ,μ+3σ), it is considered that the watershed business data to be processed does not exist, and the complete record of the watershed business data to be processed is added to the multi-table database, hash table and normal distribution data set.

[0102] Furthermore, if the same data entry does not exist in the hash value linked list, a complete record of the watershed service data to be processed is added to the multi-table database, the hash table and the normally distributed data set.

[0103] The present embodiment provides a method for deduplicating watershed business data, which sets a hash function for a target attribute combination corresponding to each data table, and then uses the hash function to calculate the hash value of the data in the target attribute combination corresponding to each data table, to obtain the hash values ​​of multiple data entries corresponding to each data table, and uses the hash values ​​of the multiple data entries as discrete random variables, converts the deduplication problem into a probability distribution problem, and approximates the hash value distribution to a normal distribution, which lays a foundation for subsequent hash judgment of the watershed business data to be processed using the normal distribution, and greatly improves the deduplication efficiency of multi-source data; secondly, by judging whether the hash value of the attribute column of the watershed business data to be processed is in the horizontal axis interval corresponding to the normal distribution, it is determined whether there is data identical to the watershed business data to be processed in the multi-table database, and then the watershed business data to be processed is deleted according to the judgment result, and the "3σ" principle of the normal distribution is utilized. Since the data falling in the horizontal axis interval corresponding to the normal distribution is judged to be a low-probability event, the hash judgment is performed using the horizontal axis interval corresponding to the normal distribution to achieve the purpose of efficient deduplication and save system storage overhead.

[0104] In this embodiment, a device for deduplicating data of a watershed service is also provided, and the device is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0105] This embodiment provides a device for deduplicating data of a watershed service. Figure 4 As shown, including:

[0106] The splicing module 401 is used to obtain a multi-table database corresponding to the watershed business data, and splice the header attributes of the multi-table database to obtain an attribute document;

[0107] A fusion feature selection module 402 is used to perform fusion feature selection on the attribute document and the data table in the multi-table database to obtain a target attribute combination corresponding to each data table;

[0108] The hash value calculation module 403 is used to calculate the hash value of the target attribute combination corresponding to each data table, obtain the hash value distribution corresponding to each data table, and approximate the hash value distribution to a normal distribution;

[0109] The hash determination module 404 is used to obtain the watershed business data to be processed, perform hash determination on the watershed business data to be processed using normal distribution, perform data deduplication based on the hash determination result, and obtain the watershed business data deduplication result.

[0110] In some optional implementations, the splicing module 401 includes:

[0111] A unified processing unit is used to uniformly process the field names of the same header attribute in a multi-table database to obtain the attribute name corresponding to the data table;

[0112] The splicing unit is used to splice the attribute names corresponding to the data table item by item to obtain an attribute document.

[0113] In some optional implementations, the fusion feature selection module 402 includes:

[0114] A first selection unit is used to select key attributes of the attribute document using a keyword extraction algorithm to obtain a key attribute name combination;

[0115] The second selection unit is used to randomly select an attribute name for a data table that does not contain a key attribute name and add it to the key attribute name combination to obtain a combination containing attribute names of all data tables;

[0116] The determination unit is used to obtain the attribute name combination of each data table in the multi-table database, and determine the target attribute combination corresponding to each data table based on the attribute name combination of each data table and the combination of attribute names of all data tables.

[0117] In some optional implementations, the hash value calculation module 403 includes:

[0118] A setting unit, used to set a hash function for each target attribute combination corresponding to each data table;

[0119] A calculation unit, used to use a hash function to calculate hash values ​​of data in a target attribute combination corresponding to each data table, to obtain hash values ​​of multiple data entries corresponding to each data table;

[0120] A linking unit, used to link data entries with the same hash value using a single linked list to obtain a hash value linked list, and store a head node of the single linked list into the hash table;

[0121] The approximate processing unit is used to determine the hash value distribution corresponding to each data table based on the hash values ​​of multiple data entries, and approximate the hash value distribution into a normal distribution.

[0122] In some optional implementations, the hash determination module 404 includes:

[0123] A judgment unit, used to obtain a hash value of an attribute column of the watershed business data to be processed, and judge whether the hash value of the attribute column of the watershed business data to be processed is in a horizontal axis interval corresponding to a normal distribution;

[0124] A query unit, used for querying each piece of data in the hash value linked list in sequence by using a Bloom filter if the hash value of the attribute column of the watershed business data to be processed is within the horizontal axis interval;

[0125] The deletion unit is used to delete the to-be-processed basin business data if the same data entries exist in the hash value linked list, so as to obtain the deduplication result of the basin business data.

[0126] In some optional implementations, the hash determination module 404 further includes:

[0127] An adding unit is used to add the watershed business data to be processed to the multi-table database, hash table and normal distribution data set if the hash value of the attribute column of the watershed business data to be processed is not in the horizontal axis interval, or if the same data entry does not exist in the hash value linked list.

[0128] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0129] A watershed business data deduplication device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0130] The embodiment of the present invention also provides a computer device having the above Figure 4 A device for deduplicating watershed business data is shown.

[0131] See also Figure 5 , Figure 5 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 5 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 5 A processor 10 is taken as an example.

[0132] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0133] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0134] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0135] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0136] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 5 The example of connecting through bus is taken in the following.

[0137] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device may be a touch screen.

[0138] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0139] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0140] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for deduplicating watershed service data, characterized in that: The method comprises: Acquire a multi-table database corresponding to the watershed business data, and concatenate the header attributes of the multi-table database to obtain an attribute document; Performing fusion feature selection on the attribute document and the data table in the multi-table database to obtain a target attribute combination corresponding to each data table; Performing hash value calculation on the target attribute combination corresponding to each data table to obtain the hash value distribution corresponding to each data table, and approximating the hash value distribution to a normal distribution; Obtain the watershed business data to be processed, perform hash judgment on the watershed business data to be processed using the normal distribution, deduplicate the data based on the hash judgment result, and obtain the deduplication result of the watershed business data; wherein, the hash judgment is to determine whether the hash value of the attribute column of the watershed business data to be processed is in the horizontal axis interval corresponding to the normal distribution.

2. The method according to claim 1, characterized in that The header attributes of the multi-table database are spliced ​​to obtain an attribute document, including: The field names of the same table header attributes in the multi-table database are uniformly processed to obtain the attribute names corresponding to the data table; The attribute names corresponding to the data table are concatenated item by item to obtain the attribute document.

3. The method according to claim 1, characterized in that The step of performing fusion feature selection on the attribute document and the data table in the multi-table database to obtain a target attribute combination corresponding to each data table includes: Using a keyword extraction algorithm to select key attributes of the attribute document to obtain a key attribute name combination; For a data table that does not contain a key attribute name, randomly select an attribute name and add it to the key attribute name combination to obtain a combination containing attribute names of all data tables; The attribute name combination of each data table in the multi-table database is obtained, and the target attribute combination corresponding to each data table is determined based on the attribute name combination of each data table and the combination of attribute names of all data tables.

4. The method according to claim 1, characterized in that: The step of performing hash value calculation on the target attribute combination corresponding to each data table to obtain a hash value distribution corresponding to each data table, and approximating the hash value distribution to a normal distribution includes: Setting a hash function for each target attribute combination corresponding to each data table; Using the hash function, respectively calculate hash values ​​of the data in the target attribute combination corresponding to each data table to obtain hash values ​​of multiple data entries corresponding to each data table; Use a single linked list to link data entries with the same hash value to obtain a hash value linked list, and store the head node of the single linked list in the hash table; A hash value distribution corresponding to each data table is determined based on the hash values ​​of the plurality of data entries, and the hash value distribution is approximated to a normal distribution.

5. The method according to claim 4, characterized in that The method of performing hash determination on the to-be-processed basin service data by using the normal distribution, and performing data deduplication based on the hash determination result to obtain the basin service data deduplication result includes: Obtaining a hash value of an attribute column of the watershed service data to be processed, and determining whether the hash value of the attribute column of the watershed service data to be processed is in a horizontal axis interval corresponding to the normal distribution; If the hash value of the attribute column of the to-be-processed basin service data is within the horizontal axis interval, then use the Bloom filter to query each piece of data in the hash value linked list in turn; If the same data entries exist in the hash value linked list, the to-be-processed basin business data is deleted to obtain the deduplication result of the basin business data.

6. The method according to claim 5, characterized in that The method of performing hash determination on the to-be-processed basin service data by using the normal distribution, performing data deduplication based on the hash determination result, and obtaining the basin service data deduplication result further includes: If the hash value of the attribute column of the watershed business data to be processed is not in the horizontal axis interval, or if the same data entry does not exist in the hash value linked list, the watershed business data to be processed is added to the multi-table database, the hash table and the normally distributed data set.

7. A watershed service data deduplication device, characterized in that: The device comprises: A splicing module is used to obtain a multi-table database corresponding to the watershed business data, and splice the header attributes of the multi-table database to obtain an attribute document; A fusion feature selection module is used to perform fusion feature selection on the attribute document and the data table in the multi-table database to obtain a target attribute combination corresponding to each data table; A hash value calculation module, used to perform hash value calculation on the target attribute combination corresponding to each data table, obtain the hash value distribution corresponding to each data table, and approximate the hash value distribution to a normal distribution; A hash judgment module is used to obtain the watershed business data to be processed, perform hash judgment on the watershed business data to be processed using the normal distribution, deduplicate the data based on the hash judgment result, and obtain the deduplication result of the watershed business data; wherein, the hash judgment is to determine whether the hash value of the attribute column of the watershed business data to be processed is in the horizontal axis interval corresponding to the normal distribution.

8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for deduplication of watershed business data according to any one of claims 1 to 6 by executing the computer instructions.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method for deduplicating watershed business data according to any one of claims 1 to 6.

10. A computer program product, characterized in that It includes computer instructions, and the computer instructions are used to enable a computer to execute the method for deduplication of watershed business data according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Database uniqueness constraint processing method and device, equipment and medium

    CN115617809A

  • One-person multi-number identification method based on online locality sensitive hashing algorithm

    CN118965081A