Data cleaning rule optimization method and device, equipment and storage medium
By automatically identifying and repairing deep logical contradictions in data cleaning rules and optimizing initial cleaning rules, the problem that static rules cannot cope with dynamic changes is solved, the efficiency and accuracy of data cleaning are improved, and the losses of enterprises are avoided.
Patent Information
- Application Number
- CN202510357167.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, the cleaning rules in the cleaning rule library are static rules, which are difficult to deal with dynamic changes in business logic, resulting in the inability to identify deep logic contradictions and the manual processing efficiency is inefficient, so deep logic contradictions cannot be effectively handled before applying cleaning data.
By obtaining the data to be cleaned, performing preliminary cleaning based on the initial cleaning rules, calculating conflict values, matching the target data repair strategy and repairing target conflict data, optimizing the initial cleaning rules, and automatically realizing dynamic optimization of the cleaning rules.
It realizes automatic identification and repair of deep logical contradictions before cleaning data applications, avoids corporate losses, improves processing efficiency, and avoids inefficiency and inconsistency in manual processing.
Smart Images

Figure CN120371812A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data cleaning, and in particular to a method, device, equipment and storage medium for optimizing data cleaning rules. Background Art
[0002] In the field of truck express delivery, logistics companies rely on multi-source data such as order management systems, on-board sensors, and road condition monitoring platforms to achieve real-time scheduling and optimization of transportation resources. These data cover order information (such as cargo type, delivery time), vehicle status (such as location, load), road condition dynamics (such as congestion index, weather), etc., and have the characteristics of diverse sources, heterogeneous formats, and frequent updates. However, due to the complexity of the data collection process (such as equipment signal interference, manual input errors) and the dynamic nature of business scenarios (such as temporary traffic control, sudden weather events), there are generally problematic data in the original data. For example, the vehicle assigned to an order has insufficient load capacity, or the positioning of a truck shows that it cannot arrive at the destination within the scheduled time. If such data quality problems are not effectively cleaned, it will directly lead to the failure of scheduling instructions, such as empty vehicle driving and redundant path planning, thereby increasing operating costs and reducing service reliability.
[0003] In order to process the problematic data in the original data, a cleaning rule library is set in the related technology. The cleaning rule library is provided with preset cleaning rules and a rule cleaning engine, wherein the rule cleaning engine can use the preset cleaning rules to clean the problematic data in the original data.
[0004] However, the cleaning rules in the cleaning rule library are static rules. Specifically, predefined cleaning rules (such as field format verification and numerical threshold filtering) can only handle surface errors in fixed patterns (such as abnormal date formats), but it is difficult to cope with dynamic changes in business logic. For example, after the new city restriction policy was added, the system could not automatically identify compliance conflicts in route planning. Moreover, since static rules need to be set manually, due to the limitations of manpower itself, it is impossible to set rules that cover all data logics. Static rules can only cover basic rules, which leads to the inability to identify deep logical contradictions between data (such as vehicles being repeatedly assigned to multiple orders). As a result, the data finally cleaned still has problems, and more serious problems arise when using these data.
[0005] In order to solve the above technical problems, the industry mainly adopts the following technical solutions.
[0006] Manually add cleaning rules in the cleaning rule library to adapt to the dynamic changes of business logic. And after problems occur due to deep logical contradictions when the cleaning rules in the cleaning rule library are applied, manually analyze the deep logical contradictions that exist, and then manually modify, add, or subtract the cleaning rules in the cleaning rule library according to the analysis results.
[0007] However, when manually processing the cleaning rules in the cleaning rule library, the processing efficiency is too low. And when manually analyzing the deep logical contradictions existing in the cleaning rules, on the one hand, due to different technical experiences mastered by different technicians, the deep logical contradictions analyzed in the cleaning rules are different, resulting in different modifications, additions, or subtractions to the cleaning rules for the same deep logical contradiction in the end; on the other hand, since the above solution modifies, adds, or subtracts the rules with deep logical contradictions only when problems occur after the data with deep logical contradictions is applied, this will cause the enterprise to suffer losses due to problems. Summary of the Invention
[0008] The present invention provides an optimization method, device, equipment, and storage medium for data cleaning rules to solve the problems of low processing efficiency and inability to effectively handle the deep logical contradictions in the cleaning rules before applying the cleaning data in the related art.
[0009] To solve the above technical problems, in a first aspect, the present invention provides an optimization method for data cleaning rules, and the method includes:
[0010] Obtain the data to be cleaned;
[0011] Based on a preset initial cleaning rule, clean the data to be cleaned to obtain preliminary cleaned data;
[0012] Calculate the conflict value of the preliminary cleaned data to obtain target conflict data with the conflict value greater than a first preset value;
[0013] Match a target data repair strategy corresponding to the target conflict data, and repair the target conflict data based on the target data repair strategy;
[0014] Optimize the initial cleaning rule based on the repair result of the target conflict data.
[0015] Optionally, the optimizing the initial cleaning rule based on the repair result of the target conflict data includes:
[0016] In the initial cleaning rule, match at least one target cleaning rule corresponding to the target conflict data;
[0017] Among at least one of the target cleaning rules, reward the first target cleaning rule that the repair result of the target conflict data conforms to;
[0018] Among at least one of the target cleaning rules, punish the second target cleaning rule that the repair result of the target conflict data does not conform to.
[0019] Optionally, calculating the conflict value of the preliminary cleaning data includes:
[0020] Analyze the preliminary cleaning data to obtain the target conflict category corresponding to the target conflict data set in the preliminary cleaning data and the target time cluster corresponding to the target conflict data set;
[0021] Based on the target conflict category and the target time cluster, calculate the conflict value of the conflict data in the target conflict data set.
[0022] Optionally, based on the target conflict category and the target time cluster, calculating the conflict value of the conflict data in the target conflict data set includes:
[0023] Based on the preset weight corresponding to the target conflict category, determine the first influence coefficient of the conflict value, where the value of the first influence coefficient is proportional to the value of the preset weight;
[0024] Based on the first influence coefficient and the target time cluster, calculate the conflict value of the conflict data in the target conflict data set.
[0025] Optionally, based on the first influence coefficient and the target time cluster, calculating the conflict value of the conflict data in the target conflict data set includes:
[0026] Based on the target time cluster, analyze the density of the target conflict data set within the target time cluster;
[0027] Based on the density of the target conflict data set within the target time cluster, determine the second influence coefficient of the conflict value, where the value of the second influence coefficient is proportional to the density;
[0028] Based on the first influence coefficient and the second influence coefficient, calculate the conflict value of the conflict data in the target conflict data set, where both the first influence coefficient and the second influence coefficient are proportional to the conflict value of the conflict data in the target conflict data set.
[0029] Optionally, analyzing the preliminary cleaned data to obtain a target conflict category corresponding to a target conflict data set in the preliminary cleaned data and a target time cluster corresponding to the target conflict data set includes:
[0030] Inputting the preliminary cleaned data into a data conflict model to obtain the target conflict data set in the preliminary cleaned data, a target conflict category corresponding to the target conflict data set, and a target time cluster corresponding to the target conflict data set;
[0031] Calculating the data similarity of target conflict data groups in the target conflict data set and calculating the data conflict value of the target conflict data groups;
[0032] Matching the data similarity and the data conflict value of the target conflict data groups. When the data similarity and the data conflict value match, rewarding the parameters of the data conflict model, and when the data similarity and the data conflict value do not match, punishing the parameters of the data conflict model.
[0033] Optionally, the matching value of the data similarity and the data conflict value is obtained through the following calculation method:
[0034] Δ = ∣data conflict value - (1 - data conflict value)∣, where Δ is the matching value.
[0035] In a second aspect, the present invention provides an optimization device for data cleaning rules, including:
[0036] An acquisition module for acquiring data to be cleaned;
[0037] A cleaning module for cleaning the data to be cleaned based on preset initial cleaning rules to obtain preliminary cleaned data;
[0038] A calculation module for calculating the conflict value of the preliminary cleaned data to obtain target conflict data with a conflict value greater than a first preset value;
[0039] A matching module for matching a target data repair strategy corresponding to the target conflict data and repairing the target conflict data based on the target data repair strategy;
[0040] An optimization module for optimizing the initial cleaning rules based on the repair result of the target conflict data.
[0041] In a third aspect, the present invention provides an optimization device for data cleaning rules, including a memory and a processor, where:
[0042] The memory is used to store a computer program;
[0043] The processor is used to read the program in the memory and execute the steps of the optimization method of the data cleaning rule provided in the first aspect above.
[0044] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a readable computer program is stored. When the program is executed by a processor, the steps of the optimization method of the data cleaning rule provided in the first aspect above are implemented.
[0045] Compared with the prior art, an optimization method, device, equipment and storage medium of a data cleaning rule provided by the present invention have the following beneficial effects:
[0046] After the target conflict data in the initially cleaned data is screened out in the embodiment of the present invention, the target conflict data is repaired, and the initial cleaning rule is optimized according to the repair result of the target conflict data, so that the dynamic optimization of the initial cleaning rule can be automatically realized without manual dynamic optimization of the initial cleaning rule. Moreover, before the cleaning data is applied in the embodiment of the present invention, the conflicting data is obtained first, and it is corrected based on the target data repair strategy corresponding to the target conflict data, and the cleaning rule is adjusted and optimized based on the correction result (that is, the cleaning rule with deep logical contradictions is repaired), thus avoiding repairing the cleaning rule with deep logical contradictions according to the errors generated by the cleaning data after the cleaning data is applied, thereby avoiding the losses of the enterprise. In addition, according to the above statement, the embodiment of the present invention avoids analyzing the deep logical contradictions existing in the cleaning rule by manual means, thereby avoiding the problems of modification or addition and subtraction of the cleaning rule for the same deep logical contradiction by manual means. Description of the Drawings
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only a part of the embodiments of the present invention, rather than all the embodiments. For those of ordinary skill in the art, without creative efforts, other drawings obtained based on these drawings all belong to the scope protected by this application.
[0048] Figure 1 It is a flowchart of an optimization method of a data cleaning rule provided by an embodiment of the present invention.
[0049] Figure 2 It is a flowchart of a method for calculating the conflict value of initially cleaned data provided by an embodiment of the present invention.
[0050] Figure 3 It is a flowchart of a parameter optimization method of a data conflict model provided by an embodiment of the present invention.
[0051] Figure 4 It is a flowchart of a method for calculating the conflict values of conflict data in a conflict data set provided by an embodiment of the present invention.
[0052] Figure 5 It is an optimization device for data cleaning rules provided by an embodiment of the present invention.
[0053] Figure 6 It is a schematic structural diagram of an optimization device for data cleaning rules provided by an embodiment of the present invention.
[0054] Figure 7 It is a schematic structural diagram of a computer-readable storage medium provided by an embodiment of the present invention. Detailed implementation manners
[0055] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0056] In order to make the description of the present disclosure more detailed and complete, the following provides an illustrative description of the embodiments and specific examples of the present invention; however, this is not the only form for implementing or applying the specific examples of the present invention. The embodiments cover the features of multiple specific examples and the method steps and their sequences for constructing and operating these specific examples. However, other specific examples can also be used to achieve the same or equivalent functions and step sequences. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of this application.
[0057] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here.
[0058] In the description of the embodiments of the present invention, unless otherwise specified, " / " means "or". For example, A / B can mean A or B; "and / or" in the text is only a relationship description of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. The optional embodiments described here are only used to illustrate and explain the present invention and are not used to limit the present invention. And without conflict, the embodiments of this application and the features in the embodiments can be combined with each other.
[0059] Example 1
[0060] As Figure 1 described, it is a flowchart of an optimization method for data cleaning rules provided by an embodiment of the present invention, including the following steps.
[0061] Step S101, obtain the data to be cleaned.
[0062] In step S101, the data to be cleaned can be any data that needs to be cleaned. For example, in the field of truck express delivery, the data to be cleaned can be order information data, vehicle operation data, road conditions and environment data, scheduling and operation data, etc.
[0063] Step S102, clean the data to be cleaned based on a preset initial cleaning rule to obtain preliminarily cleaned data.
[0064] Specifically, the preset initial cleaning rule can be a cleaning rule for the data to be cleaned set according to business requirements. For example, in the field of truck express delivery, when transporting goods, it is necessary to transport according to the customer's address. To ensure the accuracy of the transport address data, an initial cleaning rule needs to be set for the transport address data to clean the transport address data. For this purpose, the administrative division standard format of the address data (usually province / autonomous region / municipality directly under the Central Government - city - county / district - town - township, etc.) can be used as the initial cleaning rule, and the data in the transport address data that does not conform to the administrative division standard format can be cleaned to obtain preliminarily cleaned data.
[0065] Step S103, calculate the conflict value of the preliminarily cleaned data to obtain target conflict data whose conflict value is greater than a first preset value.
[0066] It should be noted that the conflict value of the preliminarily cleaned data is a value representing the conflicts existing in the preliminarily cleaned data itself and the conflicts existing between the preliminarily cleaned data and other preliminarily cleaned data. For example, if there is a first preliminarily cleaned data "The customer address is Dingbian County, Yan'an City, Shaanxi Province", but Dingbian County is actually under the jurisdiction of Yulin City, Shaanxi Province, and Yan'an City does not govern Dingbian County, so it can be understood that there are conflicts in the first preliminarily cleaned data itself. In addition, there is a second preliminarily cleaned data "The transport time from Xi'an City, Shaanxi Province to Yan'an City, Shaanxi Province is one hour", but the shortest time from Xi'an City, Shaanxi Province to the nearest county in Yan'an City is three hours, so it can be understood that there are conflicts between the first preliminarily cleaned data and the second preliminarily cleaned data. In this way, by calculating the conflict values of each preliminarily cleaned data itself and the conflict values between each preliminarily cleaned data and other preliminarily cleaned data, the conflict value of the preliminarily cleaned data can be obtained.
[0067] When setting the first preset value, it can be set according to specific requirements during application. For example, the first preset value can be a conflict value that affects the normal use of the preliminary cleaning data. That is, when the conflict value is greater than the first preset value, the preliminary cleaning data cannot be used normally. When the conflict value is not greater than the first preset value, the preliminary cleaning data only has a small error compared to the standard data and can still be put into use normally.
[0068] In an alternative implementation, as Figure 2 described, it is a flowchart of a method for calculating the conflict value of preliminary cleaning data provided by an embodiment of the present invention, including the following steps.
[0069] Step S1031, analyze the preliminary cleaning data to obtain the target conflict category corresponding to the target conflict data set in the preliminary cleaning data and the target time cluster corresponding to the target conflict data set.
[0070] In an alternative implementation, as Figure 3 described, it is a flowchart of a method for optimizing the parameters of a data conflict model provided by an embodiment of the present invention, including the following steps.
[0071] Step S10311, input the preliminary cleaning data into the data conflict model to obtain the target conflict data set in the preliminary cleaning data, the target conflict category corresponding to the target conflict data set, and the target time cluster corresponding to the target conflict data set.
[0072] Among them, the target conflict data set is a set of conflict data with the same or similar conflict attributes and aggregated in time. The target conflict category corresponding to the target conflict data set is the conflict category of the data in the data set, and the target time cluster corresponding to the target conflict data set is the time span corresponding to all conflict data in the target conflict data set.
[0073] Specifically, after inputting the preliminary cleaning data into the data conflict model, the data conflict model can extract conflict-related fields from the preliminary cleaning data to construct a feature vector; further, the data conflict model can cluster the feature vectors, merge similar clusters and assign semantic labels as conflict categories; further, the data conflict model can extract the timestamps of the conflict data, group them by conflict category; further, the data conflict model can perform density clustering (such as the OPTICS algorithm) on each group of timestamp sequences, identify dense intervals, and then can output the time cluster range (such as "2024-05-10 14:00 to 16:00").
[0074] Step S10312, calculate the data similarity of the target conflict data group in the target conflict data set, and calculate the data conflict value of the target conflict data group.
[0075] Specifically, when calculating the data similarity between target conflict data groups in the target conflict data set, the key fields of each conflict data in the target conflict data group can be extracted (such as time deviation, distance deviation, vehicle type, cargo priority, etc.), and the numerical key fields can be standardized, and the categorical key fields can be encoded. Further, the weighted similarity of the key fields of each conflict data can be calculated based on a preset similarity algorithm, and then the weighted similarity can be used as the final data similarity. For example, the preset similarity algorithm can be a hybrid similarity algorithm, and its calculation formula can be: data similarity = a * cosine similarity (numerical key fields) + b * Jaccard similarity (categorical key fields) + c * time window overlap similarity = a * cosine similarity (numerical key fields) + b * Jaccard similarity (categorical key fields) + c * time window overlap similarity.
[0076] Specifically, when calculating the data conflict value, a conflict severity formula can be defined based on a preset conflict quantification rule, and then the data conflict value can be calculated based on the conflict severity formula.
[0077] Step S10313: Match the data similarity of the target conflict data group and the data conflict value. When the data similarity and the data conflict value match, reward the parameters of the data conflict model. When the data similarity and the data conflict value do not match, punish the parameters of the data conflict model.
[0078] It can be understood that based on the above statements, for the conflict data in the same conflict data group, their conflict attributes are relatively similar, that is, their similarity is relatively high, and at the same time, the conflict value between them is relatively low. Therefore, by matching whether the data similarity and the data conflict value match (that is, whether the data similarity and the data conflict value are inversely proportional), it can be judged whether the clustering result is accurate when the data conflict model obtains the target conflict data group for clustering. When the data similarity and the data conflict value match, it proves that the clustering result of the data conflict model for the target data group is good. Therefore, the parameters can be rewarded. Otherwise, the parameters of the data conflict model are punished, so that the prediction result of the data conflict model is more accurate.
[0079] In an alternative implementation, the matching value of the data similarity and the data conflict value is obtained through the following calculation method:
[0080] Δ = ∣data conflict value - (1 - data conflict value)∣, where Δ is the matching value.
[0081] Specifically, when calculating the data conflict value in this implementation, the initial data conflict value needs to be divided by its corresponding maximum value so that the finally obtained data conflict value is less than 1.
[0082] It can be understood that in the formula given in this implementation, the matching value is equal to the data conflict value minus the complementary value of the data conflict value. In this way, the larger the matching value is, the larger the gap between the data conflict value and the complementary value of the data conflict value is. The complementary value of the data conflict value can represent the similarity between data. Therefore, through the above formula, it can be obtained that the larger the gap between the data conflict value and the data similarity value is, the higher the matching degree between the data is. Therefore, the above formula can be used to intuitively judge whether the data similarity and the data conflict value match.
[0083] Step S1032: Calculate the conflict value of the conflict data in the target conflict data set based on the target conflict category and the target time cluster.
[0084] In an alternative implementation, as Figure 4 shown, it is a flowchart of a method for calculating the conflict value of conflict data in a conflict data set provided by an embodiment of the present invention, including the following steps.
[0085] Step S10321: Determine the first influence coefficient of the conflict value based on the preset weight corresponding to the target conflict category.
[0086] Wherein, the value of the first influence coefficient is directly proportional to the value of the preset weight.
[0087] Specifically, when setting the preset weight corresponding to the target conflict category, it can be determined according to the degree of influence of the conflict category on the service. The deeper the degree of influence of the conflict category on the service is, the larger the weight value corresponding to the conflict category can be set. For example, if the conflict category is path planning conflict, its influence on the service is cost increase and delay in timeliness, which has a greater impact on truck express transportation, and its corresponding weight can be set larger.
[0088] Step S10322: Calculate the conflict value of the conflict data in the target conflict data set based on the first influence coefficient and the target time cluster.
[0089] In an alternative implementation, step S10322 includes:
[0090] Analyze the density of the target conflict data set within the target time cluster based on the target time cluster.
[0091] It can be understood that the shorter the time span of the time cluster is and the more data there is in the target conflict data set, the greater the density of the target conflict data set within the target time cluster is. Based on the above explanation, when determining the density of the target conflict data set within the target time cluster, it can be obtained by dividing the number of conflict data in the target conflict data set by the time span value of the target time cluster.
[0092] Determine a second influence coefficient of the conflict value based on the density of the target conflict data set within the target time cluster.
[0093] Wherein, the value of the second influence coefficient is directly proportional to the density.
[0094] Calculate the conflict value of the conflict data in the target conflict data set based on the first influence coefficient and the second influence coefficient.
[0095] Wherein, both the first influence coefficient and the second influence coefficient are directly proportional to the conflict value of the conflict data in the target conflict data set.
[0096] It can be understood that the larger the preset weight of the target conflict category, the greater the impact caused by the conflict of the conflict data in the target conflict data set belonging to the target conflict category. Therefore, the first influence coefficient of the target conflict category on the conflict value of the conflict data is also larger. The denser the conflict data in the target conflict data set, the greater the impact caused by the target conflict data set. Therefore, the second influence coefficient on the conflict value of the conflict data is also larger. Therefore, according to the first influence coefficient and the second influence coefficient, it can be known the influence degree of the target conflict category and the target time cluster on the conflict value of the conflict data in the target conflict data set, so that the conflict value of the conflict data can be calculated.
[0097] Step S104, match a target data repair strategy corresponding to the target conflict data, and repair the target conflict data based on the target data repair strategy.
[0098] Specifically, when matching a target data repair strategy corresponding to the target conflict data, semantic analysis can be performed on the target conflict data to obtain a semantic analysis result corresponding to the target conflict data. Then, the semantic analysis result of the target conflict data is input into a pre-established database to obtain the target data repair strategy corresponding to the semantic analysis result of the target conflict data in the database. This database can be set with various conflict data semantics and data repair strategies corresponding to various conflict data semantics.
[0099] Step S105, optimize the initial cleaning rule based on the repair result of the target conflict data.
[0100] Specifically, the target conflict data can be repaired into correct data without conflict by using the corresponding target data repair strategy. In this way, by comparing the target conflict data and its corresponding correct data, the reason why the target conflict data is not cleaned by the initial cleaning rule can be analyzed. Then, the initial cleaning rule can be optimized according to this reason, so that the optimized initial cleaning rule can cover the cleaning of the target conflict data.
[0101] In an alternative implementation, step S105 includes:
[0102] In the initial cleaning rule, at least one target cleaning rule corresponding to the target conflict data is matched.
[0103] Specifically, the semantics of the target conflict data can be analyzed to generate semantic fields corresponding to the target conflict data, and the semantic fields of the target conflict data are matched with the fields of each initial cleaning rule. Then, when the number of semantic fields of the initial cleaning rule that match the target conflict data is greater than a first value, and when the matching value of the semantic fields of the initial cleaning rule that match the target conflict data is greater than a second value, the initial cleaning rule is determined as the target cleaning rule corresponding to the target conflict data.
[0104] Among at least one of the target cleaning rules, the first target cleaning rule that the repair result of the target conflict data conforms to is rewarded.
[0105] Among at least one of the target cleaning rules, the second target cleaning rule that the repair result of the target conflict data does not conform to is punished.
[0106] Specifically, for the repair result of the target conflict data, the target conflict data is evaluated using the target cleaning rule. Since the repair result of the target conflict data is correct data, the rule it satisfies is naturally the correct rule, and the correct rule (i.e., the first target cleaning rule) can be rewarded. The rule it does not satisfy is naturally the incorrect rule, and the incorrect rule (i.e., the second target cleaning rule) can be punished.
[0107] More specifically, when rewarding the first target cleaning rule, the first target cleaning rule can be multiplied by a positive value greater than 1. When punishing the second target cleaning rule, the second target cleaning rule can be multiplied by a negative value greater than 1. In this way, the larger the coefficient value multiplied by the cleaning rule, the more correct it is, and the smaller the coefficient value multiplied by the cleaning rule, the more incorrect it is. Finally, the cleaning rules with a coefficient less than the preset value can be eliminated to achieve the optimization of the cleaning rules.
[0108] It can be understood that after screening out the target conflict data in the preliminary cleaning data in the embodiments of the present invention, the target conflict data is repaired, and based on the repair result of the target conflict data, the initial cleaning rule is optimized, so that the dynamic optimization of the initial cleaning rule can be automatically realized without manually performing dynamic optimization on the initial cleaning rule. Moreover, before the cleaning data is applied in the embodiments of the present invention, data with conflicts is first obtained, and it is corrected based on the target data repair strategy corresponding to the target conflict data, and the cleaning rule is adjusted and optimized based on the correction result (that is, the cleaning rule with deep logical contradictions is repaired), thus avoiding repairing the cleaning rule with deep logical contradictions according to the errors generated by the cleaning data after the cleaning data is applied, thereby avoiding losses to the enterprise. In addition, according to the above statement, the embodiments of the present invention avoid analyzing the deep logical contradictions existing in the cleaning rule through manual means, thus avoiding the problems of manual modification or addition and subtraction of the same deep logical contradiction to the cleaning rule.
[0109] Embodiment 2
[0110] Based on the above method for optimizing data cleaning rules, an embodiment of the present invention provides an apparatus for optimizing data cleaning rules, as Figure 5 shown, the apparatus 50 for optimizing data cleaning rules includes the following modules.
[0111] An obtaining module 51, configured to obtain data to be cleaned;
[0112] A cleaning module 52, configured to clean the data to be cleaned based on a preset initial cleaning rule to obtain preliminary cleaning data;
[0113] A calculating module 53, configured to calculate a conflict value of the preliminary cleaning data to obtain target conflict data whose conflict value is greater than a first preset value;
[0114] A matching module 54, configured to match a target data repair strategy corresponding to the target conflict data and repair the target conflict data based on the target data repair strategy;
[0115] An optimizing module 55, configured to optimize the initial cleaning rule based on the repair result of the target conflict data.
[0116] For other details of the implementation of the above technical solutions by each module in the above apparatus for optimizing data cleaning rules, reference can be made to the description in the method for optimizing data cleaning rules provided in the above embodiments of the present invention, which will not be elaborated here.
[0117] Based on the above method for optimizing data cleaning rules, as Figure 6As shown in the figure, an embodiment of the present invention further provides a structural schematic diagram of an optimization device for data cleaning rules. The optimization device includes a processor 61 and a memory 62 coupled to the processor 61. The memory 62 stores a computer program. When the computer program is executed by the processor 61, the processor 61 is caused to execute the steps of the optimization method for data cleaning rules in the above embodiment.
[0118] For other details of the implementation of the above technical solution by the processor 61 in the optimization device for data cleaning rules, reference may be made to the description in the optimization method for data cleaning rules provided in the above embodiment of the invention, which will not be elaborated here.
[0119] Among them, the processor 61 can also be referred to as a CPU (Central Processing Unit, central processing unit). The processor 61 may be an integrated circuit chip with signal processing capabilities; the processor 61 may also be a general-purpose processor, DSP (Digital Signal Process, digital signal processor), ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), FPGA (Field Programmable Gata Array, field programmable gate array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Among them, the general-purpose processor may be a microprocessor or the processor 61 may also be any conventional processor, etc.
[0120] As Figure 7 shown in the figure, an embodiment of the present invention further provides a structural schematic diagram of a computer-readable storage medium. A readable computer program 71 is stored on the storage medium; among them, the computer program 71 can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which may be a personal computer, server, or network device, etc.) or a processor (processor) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, magnetic disks or optical discs, ROM (Read-Only Memory, read-only memory), RAM (Random Access Memory, random access memory), etc., or terminal devices such as computers, servers, mobile phones, and tablets.
[0121] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces, devices, or modules, and can be in electrical, mechanical, or other forms.
[0122] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0123] In addition, in each embodiment of the present application, the various functional modules can be integrated in one processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0124] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0125] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, they implement all or part of the processes or functions described in the embodiments of the present application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0126] The technical solutions provided in the present application have been described in detail above. Specific examples are used in the present application to illustrate the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
[0127] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0128] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing the processes Figure 1one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks
[0129] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions in the process Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks
[0130] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the process Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks
[0131] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. An optimization method for data cleaning rules, characterized in that Including: Obtain the data to be cleaned; Clean the data to be cleaned based on a preset initial cleaning rule to obtain preliminarily cleaned data; Calculate the conflict value of the preliminarily cleaned data to obtain target conflict data with the conflict value greater than a first preset value; Match to obtain a target data repair strategy corresponding to the target conflict data, and repair the target conflict data based on the target data repair strategy; Optimize the initial cleaning rule based on the repair result of the target conflict data.
2. The optimization method of the data cleaning rule according to claim 1, characterized in that The optimizing the initial cleaning rule based on the repair result of the target conflict data includes: In the initial cleaning rule, match to obtain at least one target cleaning rule corresponding to the target conflict data; Among at least one of the target cleaning rules, reward the first target cleaning rule that the repair result of the target conflict data conforms to; Among at least one of the target cleaning rules, punish the second target cleaning rule that the repair result of the target conflict data does not conform to.
3. The optimization method of the data cleaning rule according to claim 1, characterized in that The calculating the conflict value of the preliminarily cleaned data includes: Analyze the preliminarily cleaned data to obtain a target conflict category corresponding to a target conflict data set in the preliminarily cleaned data and a target time cluster corresponding to the target conflict data set; Based on the target conflict category and the target time cluster, calculate the conflict value of the conflict data in the target conflict data set.
4. The optimization method of the data cleaning rule according to claim 3, wherein The calculating the conflict value of the conflict data in the target conflict data set based on the target conflict category and the target time cluster includes: Based on a preset weight corresponding to the target conflict category, determine a first influence coefficient of the conflict value, where the value of the first influence coefficient is proportional to the value of the preset weight; Based on the first influence coefficient and the target time cluster, calculate the conflict value of the conflict data in the target conflict data set.
5. The optimization method of the data cleaning rule according to claim 4, characterized in that, The calculating the conflict value of the conflict data in the target conflict data set based on the first influence coefficient and the target time cluster includes: Based on the target time cluster, analyze the density of the target conflict data set within the target time cluster; Based on the density of the target conflict data set within the target time cluster, determine a second influence coefficient of the conflict value, where the value of the second influence coefficient is proportional to the density; Based on the first influence coefficient and the second influence coefficient, calculate the conflict value of the conflict data in the target conflict data set, where both the first influence coefficient and the second influence coefficient are proportional to the conflict value of the conflict data in the target conflict data set.
6. The optimization method of the data cleaning rule according to claim 3, characterized in that The analyzing the preliminarily cleaned data to obtain a target conflict category corresponding to a target conflict data set in the preliminarily cleaned data and a target time cluster corresponding to the target conflict data set includes: Input the preliminarily cleaned data into a data conflict model to obtain the target conflict data set in the preliminarily cleaned data, the target conflict category corresponding to the target conflict data set, and the target time cluster corresponding to the target conflict data set; Calculate the data similarity of the target conflict data groups in the target conflict data set, and calculate the data conflict value of the target conflict data groups; Match the data similarity and the data conflict value of the target conflict data groups. When the data similarity and the data conflict value match, reward the parameters of the data conflict model. When the data similarity and the data conflict value do not match, punish the parameters of the data conflict model.
7. The optimization method of the data cleaning rule according to claim 6, characterized in that, Obtain the matching value of the data similarity and the data conflict value through the following calculation method: Δ = |data conflict value - (1 - data conflict value)|, where Δ is the matching value.
8. An optimization device for data cleaning rules, characterized in that, Includes: An acquisition module for acquiring data to be cleaned; A cleaning module for cleaning the data to be cleaned based on a preset initial cleaning rule to obtain preliminarily cleaned data; A calculation module for calculating the conflict value of the preliminarily cleaned data to obtain target conflict data whose conflict value is greater than a first preset value; A matching module for matching and obtaining a target data repair strategy corresponding to the target conflict data, and repairing the target conflict data based on the target data repair strategy; An optimization module for optimizing the initial cleaning rule based on the repair result of the target conflict data.
9. An optimization device for data cleaning rules, characterized in that, Includes a memory and a processor, where: The memory is used to store computer programs; The processor is used to read the computer program in the memory and execute the steps of any data cleaning rule optimization method as described in claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Stored thereon is a readable computer program, and when the program is executed by the processor, it implements the steps of any data cleaning rule optimization method as described in claims 1 to 7.