Data processing optimization method and device, electronic equipment and storage medium

Through clustering and multi-objective optimization models, combined with non-dominant sorting genetic algorithms, data processing time is optimized, and the problem of optimizing data processing time in a limited information environment is solved, and data processing time optimization with the minimum resource consumption is achieved.

CN120067865AActive Publication Date: 2025-05-30CIVIL AVIATION UNIV OF CHINA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510519643.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-30
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

In a limited information environment where data processing information characteristics are unclear, how to optimize data processing time to achieve minimum resource consumption.

Method used

By clustering the original data processing information into K clusters, and adding the data processing method identification value of each original data processing information to it, a multi-objective optimization model is constructed, and the number of file shards, total threads and data processing method identification value is optimized to optimize the data processing time using the non-dominant sorting genetic algorithm.

Benefits of technology

In an environment where the data processing method is unclear, the optimal values ​​of parameters affecting the data processing time are quickly and accurately obtained, effectively optimizing the data processing time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067865A_ABST
    Figure CN120067865A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a data processing optimization method and device, electronic equipment and a storage medium. Obtaining an original value of a data processing method identification value used by each piece of original data processing information, and adding the original value of the data processing method identification value used by each piece of original data processing information into the original data processing information, taking the original data processing information of the original value added with the data processing method identification value as to-be-processed information; constructing a multi-objective optimization model based on all the to-be-processed information; and based on a non-dominated sorting genetic algorithm, optimizing the multi-objective optimization model to obtain the optimal values of the file fragment number, the total thread number and the data processing method identification value corresponding to each piece of to-be-processed information. According to the method, the data processing time can be optimized to the maximum extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data desensitization, and in particular to a data processing optimization method, apparatus, electronic device, and storage medium. Background Art

[0002] In some application scenarios, it is necessary to optimize the data processing time of data processing information, such as optimizing the desensitization time, to ensure that data processing can be achieved with the least resources as much as possible. In some actual application scenarios, it is necessary to optimize the data processing time of data processing information without informing the data processing method used for each piece of data processing information. Therefore, how to optimize the data processing time in an environment of limited information with unclear data processing information characteristics is a topic worthy of discussion. Summary of the Invention

[0003] For the above technical problems, the technical solution adopted by the present invention is as follows: According to a first aspect of the present invention, there is provided a data processing optimization method, the method comprising the following steps: S100, based on the set of original data processing information to be processed and K data processing methods used by the set of original data processing information, clustering all the original data processing information into K clusters, where the original value of the data processing method identification value used by each original data processing information belonging to the r-th cluster is r; where each original data processing information includes the file ID corresponding to the information and the original value of the data processing parameter, and the data processing parameter includes file size, total number of file lines, number of file shards, number of lines per file shard, total number of threads, and data processing time; r takes values from 1 to K.

[0004] S200, adding the original value of the data processing method identification value used by each original data processing information to the original data processing information, and using the original data processing information with the original value of the data processing method identification value added as the information to be processed.

[0005] S300, based on all the information to be processed, constructing a multi-objective optimization model; where the multi-objective optimization model includes constraint conditions and objective functions, and the objective functions include a first objective function corresponding to minimizing the data processing time, a second objective function corresponding to minimizing the data preprocessing time, and a third objective function corresponding to minimizing the cluster center distance, where the independent variables of the first objective function and the third objective function are the number of file shards, the total number of threads, and the data processing method identification value, and the independent variables of the second objective function are the number of file shards and the total number of threads.

[0006] The S400 optimizes the multi-objective optimization model based on the non-dominated sorting genetic algorithm to obtain the optimal values of the number of file shards, the total number of threads, and the data processing method identification value corresponding to each piece of information to be processed.

[0007] According to a second aspect of the present invention, there is provided a data desensitization optimization device, and the device includes the following steps: A data processing module, configured to cluster all the original data processing information into K clusters based on the original data processing information set to be processed and the K data processing methods used by the original data processing information set. Among them, the original value of the data processing method identification value used by each original data processing information belonging to the r-th cluster is r; among them, each original data processing information includes the file ID corresponding to the information and the original value of the data processing parameter, and the data processing parameter includes the file size, the total number of file rows, the number of file shards, the number of file shard rows, the total number of threads, and the data processing time; the value of r ranges from 1 to K.

[0008] A multi-objective optimization model construction module, configured to construct a multi-objective optimization model based on all the information to be processed; among them, the multi-objective optimization model includes constraint conditions and objective functions, and the objective functions include a first objective function corresponding to minimizing the data processing time, a second objective function corresponding to minimizing the data preprocessing time, and a third objective function corresponding to minimizing the cluster center distance. Among them, the independent variables of the first objective function and the third objective function are the number of file shards, the total number of threads, and the data processing method identification value, and the independent variables of the second objective function are the number of file shards and the total number of threads.

[0009] An optimization module, configured to optimize the multi-objective optimization model based on the non-dominated sorting genetic algorithm to obtain the optimal values of the number of file shards, the total number of threads, and the data processing method identification value corresponding to each piece of information to be processed.

[0010] According to a third aspect of the present invention, there is provided an electronic device, including a processor and a memory; the processor is configured to execute the steps of the method according to the first aspect of the present invention by calling a program or instruction stored in the memory.

[0011] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium, and the computer-readable storage medium stores a program or instruction, and the program or instruction causes a computer to execute the steps of the method according to the first aspect of the present invention.

[0012] The present invention has at least the following beneficial effects: The data processing optimization method provided by the embodiments of the present invention can quickly and accurately obtain the optimal values of the parameters that affect the data processing time in a limited information environment where the data processing method is not clear, and thus can effectively optimize the data processing time.

[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0015] Figure 1 It is a flowchart of the data processing optimization method provided for the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of the present invention.

[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art belonging to the technical field of the present invention. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0018] It should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the steps as sequential processes, many of the steps can be performed in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0019] The embodiments of the present invention provide a data processing optimization method, as Figure 1 shown, the method may include the following steps: S100, clustering all the original data processing information into K clusters based on the original data processing information set and the K data processing methods used by the original data processing information set.

[0020] Among them, the original value of the data processing method identification value used for each piece of original data processing information belonging to the r-th cluster is r; among them, each piece of original data processing information includes the file ID corresponding to this information and the original value of the data processing parameters, and the data processing parameters include file size, total number of lines in the file, number of file shards, number of lines in each file shard, total number of threads, and data processing time; the data processing method identification value is used to represent which of the K data processing methods it belongs to; the value of r ranges from 1 to K.

[0021] In an embodiment of the present invention, the data processing method is a method for processing the file corresponding to the data processing information. The data processing method identification value is used to represent which of the K data processing methods it belongs to. For example, when the data processing method identification value is 1, it means that the first data processing method is used.

[0022] In an embodiment of the present invention, both the original data processing information set and the K data processing methods used by the original data processing information set can be information provided by the user. Among them, the data processing method used for each piece of original data processing information is not specified, that is, the data processing method used for each piece of original data processing information is unknown. It is only known that when performing data processing on the file corresponding to the original data processing information set, K data processing methods are used, but the specific data processing method used for each piece of original data processing information is not known.

[0023] In a practical application scenario, the original data processing information set can be an information set obtained by performing desensitization processing on civil aviation passenger sensitive information. In this scenario, the data processing method can be a desensitization method.

[0024] In an embodiment of the present invention, the file ID refers to the file name. The file ID, file size, and total number of lines in the file are the inherent attributes of each file and will not change during subsequent processing. The total number of threads refers to the number of threads used to perform data processing on the file. In the original data processing information set, the total number of threads for all pieces of original data processing information is the same. The number of file shards refers to the number of sub-files obtained by splitting the file, and the number of lines in each file shard refers to the number of lines in each file shard. Theoretically, the number of file shards is equal to the total number of lines in the file divided by the number of lines in each file shard.

[0025] In an embodiment of the present invention, since it is impossible to know the data processing method used for each piece of original data processing information, the data processing information can only be clustered by an unsupervised clustering method. Further, S100 may specifically include: S110, randomly select K pieces of original data processing information from the original data processing information set as K initial cluster centers.

[0026] S120. For any original data processing information i, obtain the similarity S between the original data processing information i and the j-th current cluster center in the current cluster center set, ij , and obtain the similarity set S corresponding to the original data processing information i. i , and add the original data processing information i to the data set corresponding to the current cluster center corresponding to the minimum similarity in S; i The value range of i is from 1 to n, where n is the number of original data processing information, and the value range of j is from 1 to K; the initial value of the current cluster center set is K initial cluster centers, and the initial value of the data set corresponding to the current cluster center is the original data processing information corresponding to the current cluster center.

[0027] In the embodiments of the present invention, the similarity between two original data processing information may be the Euclidean norm between the two original data processing information.

[0028] S130. Obtain the within-cluster dispersion corresponding to the data set corresponding to each current cluster center. If the within-cluster dispersion corresponding to any current cluster center is less than the set dispersion, use the data sets corresponding to the K current cluster centers as the K clusters and execute S140; otherwise, update the current cluster center based on the data set corresponding to each current cluster center and execute S120.

[0029] Among them, the within-cluster dispersion D corresponding to the j-th current cluster center j satisfies the following conditions: D j = (1 / 4) ∑ 4 v=1 ((∑ z(j) u=1 (N ju v - N0 j v )) / z(j)), where N ju v is the original value of the v-th associated parameter in the u-th data in the data set corresponding to the j-th current cluster center, and N0 j v is the original value of the v-th associated parameter corresponding to the j-th current cluster center. The value range of u is from 1 to z(j), where z(j) is the number of data in the data set corresponding to the j-th current cluster center, and the value range of v is from 1 to 4. The associated parameters include file size, total number of lines of the file, number of file shards, and total number of threads.

[0030] In the embodiments of the present invention, the parameter value of the v-th associated parameter of the updated current cluster center corresponding to each current cluster center is the mean value of the original values of the v-th associated parameter of all the to-be-processed information in the data set corresponding to the current cluster center.

[0031] S140. Set the data processing method identification value of the original data processing information belonging to the r-th cluster to r, where r ranges from 1 to K.

[0032] S200. Add the original value of the data processing method identification value used in each original data processing information to the original data processing information, and use the original data processing information with the original value of the data processing method identification value added as the information to be processed.

[0033] S300. Based on all the information to be processed, construct a multi-objective optimization model. The multi-objective optimization model includes constraint conditions and objective functions. The objective functions include a first objective function corresponding to minimizing the data processing time, a second objective function corresponding to minimizing the data preprocessing time, and a third objective function corresponding to minimizing the cluster center distance. Among them, the independent variables of the first objective function and the third objective function are the number of file shards, the total number of threads, and the data processing method identification value, and the independent variables of the second objective function are the number of file shards and the total number of threads.

[0034] In the embodiments of the present invention, by analyzing the relationships between the file size and data processing time, the total number of file lines and data processing time, and the number of file shards and data processing time in the original data processing information set, the following conclusions are obtained: The relationship between the file size and the data processing time includes the following situations: (1) The file size is large and the content is sensitive, resulting in a long data processing time. (2) The file size is small and the content is sensitive, resulting in a long data processing time. (3) The file size is large and the content is not sensitive, resulting in a short data processing time. (4) The file size is small and the content is not sensitive, resulting in a short data processing time.

[0035] In the embodiments of the present invention, content sensitivity means that the data processing time of the file is greatly affected by the file size, the total number of file lines, the number of file shards, and the total number of threads, and internal insensitivity means that the data processing time of the file is little affected by the file size, the total number of file lines, the number of file shards, and the total number of threads.

[0036] The relationship between the number of file lines and the data processing time includes the following situations: (1) The number of file lines is large and the content is sensitive, resulting in a long data processing time. (2) The number of file lines is small and the content is sensitive, resulting in a long data processing time. (3) The number of file lines is large and the content is not sensitive, resulting in a short data processing time. (4) The number of file lines is small and the content is not sensitive, resulting in a short data processing time.

[0037] The relationship between the number of file shards and the data processing time includes: there are cases where the number of shards is large and the data processing time is still high. The possible reasons may be that the file size is large, the content is sensitive, the data processing method is inappropriate, the number of thread allocations is small, etc.; there are also cases where the number of shards is small and the data processing time is low, and the reasons are also related to factors such as file size and content.

[0038] Through the above data analysis, it can be seen that there is no obvious linear relationship in this problem, and the various parameters affect each other. All kinds of special situations need to be considered comprehensively. In the present invention, the data processing time is optimized by adjusting three decision variables: the number of file shards, the number of threads, and the data processing method, and the sharding and thread allocation need to be within a reasonable and acceptable range as much as possible, thus constituting the above multi-objective optimization model. Among them, the first objective function aims to find three decision variables that can minimize the data processing time under the condition of minimizing the consumption of resources, the second objective function aims to find the most reasonable total number of threads, and the third objective function aims to find the most suitable data processing method.

[0039] In the embodiment of the present invention, the data preprocessing time refers to the time consumed before data processing of the file, including the time consumed for file sharding and thread allocation, etc.

[0040] Further, the first objective function Y1 corresponding to the gth piece of information to be processed g satisfies the following conditions: Y1 g =min T, where T is the data processing time calculation function, T = [(1 + min(β×NF g +γ×NL g ))×α g ×(Size g / (NF g +NL g )) + min(β×NF g +γ×NL g )].

[0041] Among them, NF g is the number of file shards of the gth piece of information to be processed during the optimization process, NL g is the total number of threads of the gth piece of information to be processed during the optimization process, α g is the time-consuming influence coefficient of the gth piece of information to be processed, Size g is the original value of the file size of the gth piece of information to be processed; the value of g ranges from 1 to n, and n is the number of pieces of information to be processed.

[0042] In the embodiment of the present invention, the time-consuming influence coefficient is used to characterize the influence degree of the data processing time by the file size, the total number of file lines, the number of file shards, and the total number of threads. The larger the time-consuming influence coefficient is, the greater the influence degree of the data processing time by the file size, the total number of file lines, the number of file shards, and the total number of threads is, and vice versa. That is, for files 1 and 2 with the same file size, total number of file lines, number of file shards, and total number of threads, if the data processing time of file 1 is greater than that of file 2, it means that the data processing time of file 1 is more easily affected by the file size, the total number of file lines, the number of file shards, and the total number of threads. For example, a file with a large size, many lines, many shards, a large total number of threads, and a long processing time indicates that its content sensitivity will naturally be high, and the time-consuming influence coefficient will be large; in contrast, a file with a large size, many lines, many shards, a large total number of threads, but a short processing time indicates that its content sensitivity is poor, and the time-consuming influence coefficient is low.

[0043] In the embodiment of the present invention, the time-consuming influence coefficient of the information to be processed can be obtained based on the original values of the file size, the total number of file lines, the number of file shards, and the total number of threads corresponding to the information. Specifically, α g Satisfies the following conditions: α g =T g ×(Size g -1 +N g -1 )×NL g org ×NF g org -λ.

[0044] Among them, T g Is the original value of the data processing time corresponding to the gth information to be processed, N g Is the total number of file lines corresponding to the gth information to be processed, NL g org Is the original value of the total number of threads corresponding to the gth information to be processed, NF g org Is the original value of the number of file shards corresponding to the gth information to be processed, and λ is an adjustment factor.

[0045] The second objective function Y2 corresponding to the gth information to be processed g Satisfies the following conditions: Y2 g =min (β×NF g +γ×NL g ); β represents the file shard weight, and γ represents the thread weight.

[0046] The third objective function Y3 corresponding to the gth information to be processed gMeet the following conditions: Y3 g = min(∑ 4 v=1 (N gv - N0 gv ) 2 / (N0 gv ) 2 ), where N gv is the parameter value of the v-th associated parameter of the g-th information to be processed during the optimization process. The value range of v is from 1 to 4, and N0 gv is the parameter value of the v-th data processing parameter of the cluster center corresponding to the data processing method identification value of the g-th information to be processed during the optimization process.

[0047] The constraint conditions corresponding to the g-th information to be processed meet the following conditions: (1) 1 ≤ C g ≤ K; C g is the data processing method identification value of the g-th information to be processed.

[0048] (2) 1 < NL g < NL0; NL0 is the total number of threads threshold, which can be an empirical value. For example, 448; (3) n1 ≤ NR g / NF g ≤ n2, where NR g is the total number of lines of the file of the g-th information to be processed, n1 is the minimum number of lines for file sharding threshold, n2 is the maximum number of lines for file sharding threshold, and both n1 and n2 can be empirical values.

[0049] Furthermore, in the embodiments of the present invention, β, γ, and λ are parameters to be optimized so that the constructed data processing time calculation function can be more accurate. In a schematic embodiment, β, γ, and λ can be obtained through the following steps: S10, substitute the original values of the file size, total number of lines of the file, number of file shards, total number of threads, and data processing method identification value of each information to be processed into the current data processing time calculation function to obtain the current predicted data processing time corresponding to each information to be processed, and obtain n current predicted data processing times.

[0050] In the embodiments of the present invention, the initial value of the current data processing time calculation function is a function for initializing β, γ, and λ.

[0051] The data processing time calculation function satisfies the following conditions: T = [(1 + (β × NF + γ × NL)) × α × (Size / (NF + NL)) + (β × NF + γ × NL)], where T is the data processing time calculation function, α is the time-consuming influence coefficient, Size is the file size, NF is the number of file shards, and NL is the total number of threads.

[0052] S20. Obtain the deviation between the n currently predicted data processing times and the corresponding n actual data processing times. If the deviation is less than or equal to the set deviation, use the current β, γ, and λ as the final β, γ, and λ. Otherwise, adjust β, γ, and λ in the current data processing time calculation function based on the deviation, and execute S10.

[0053] The actual data processing time of the information to be processed is the original value of the data processing time of the information to be processed. The deviation between the n currently predicted data processing times and the corresponding n actual data processing times can be calculated through the t-statistic, p-value, R-squared value, etc. The t-statistic is used to measure the degree of variation of the difference between the means of two samples relative to the sample data. The larger the absolute value, the more significant the difference between the means of the two samples. The p-value represents the probability of observing the statistic or a more extreme one under the null hypothesis being true. The smaller the value, the stronger the evidence to reject the null hypothesis. The R-squared value is an index to measure the goodness of fit of the model to the data, and its value ranges from 0 to 1. The closer the value is to 1, the stronger the explanatory power of the model to the data.

[0054] In the embodiments of the present invention, the deviation results of the finally obtained data processing time calculation function are shown in Table 1 below: Table 1

[0055] As can be seen from Table 1 above, the t-statistic is close to 0, indicating that the difference between the means of the two samples is very small. Here, the p-value is much greater than 0.05, indicating that there is not enough evidence to reject the null hypothesis, that is, the difference between the means of the two samples is not significant. Here, the R-squared value is very close to 1, representing that the model well explains the variance of the data. The test results show that the data processing time calculation function established in the embodiments of the present invention can better predict the data processing time of files under different strategy combinations.

[0056] S400. Based on the non-dominated sorting genetic algorithm, optimize the multi-objective optimization model to obtain the optimal values of the number of file shards, total number of threads, and data processing method identification value corresponding to each information to be processed.

[0057] The present invention uses the non-dominated sorting genetic algorithm to optimize the multi-objective optimization model to solve the problem of limited information with unclear decision variables in the real business scenario.

[0058] Further, S300 may specifically include: S301, randomly select P1 pieces of information to be processed from n pieces of information to be processed as the initial population.

[0059] P1 can be set based on actual needs. In an exemplary embodiment, P1 = 100.

[0060] S302, perform non-dominated sorting operation on the current population to be processed, obtain the non-dominated solutions in the current population. If non-dominated solutions are obtained, use the obtained non-dominated solutions as the current parents and execute S303; if no non-dominated solutions are obtained, execute S301; the initial value of the current population to be processed is the initial population.

[0061] In the embodiments of the present invention, a non-dominated solution refers to at least one piece of information whose objective function value is not the smallest in the current population to be processed. Those skilled in the art know that the method for obtaining non-dominated solutions can be the prior art.

[0062] S303, perform crossover operation on the current parents, randomly exchange the number of file shards, the total number of threads, and the data processing method identification value between any two individuals in the current parents to obtain the corresponding crossover operation result as the current crossover operation result; where an individual represents a piece of information to be processed.

[0063] In the embodiments of the present invention, the parents can be subjected to the crossover operation with a probability of 70%.

[0064] S304, perform mutation operation on the current crossover operation result, randomly mutate the number of file shards, the total number of threads, and the data processing method identification value of the individuals in the current crossover operation result to obtain the corresponding mutation result as the current offspring.

[0065] In the embodiments of the present invention, the individuals can be randomly mutated with a probability of 20%.

[0066] S305, merge the current parents and the current offspring to obtain the current merged population, and obtain the data processing time and the data processing time optimization value of each individual in the current merged population as the current data processing time and the current data processing time optimization value of the individual respectively; the current data processing time optimization value of each individual is equal to the difference obtained by subtracting the previous data processing time from the current data processing time of the individual.

[0067] In the embodiments of the present invention, the data processing time of each individual can be calculated by substituting the parameter value of the current data processing time influence parameter of the individual into the data processing time calculation function.

[0068] S306. Sort all the current data processing time optimization values corresponding to the current merged population in descending order, and use the individuals corresponding to the top P1 optimization values in the sorted data processing time optimization values as the current population to be processed. In addition, add the current data processing time of each individual in the current merged population to the corresponding current data processing time record set. Set the iteration count counter C = C + 1, where the initial value of C is 0, and the initial value of the current data processing time record set is empty.

[0069] S307. If C ≤ C0, execute S302. If C > C0, use the file shard number, total number of threads, and data processing method identification value corresponding to the maximum data processing time in the current data processing time record set of each individual as the optimal values of the file shard number, total number of threads, and data processing method identification value corresponding to the information to be processed. C0 is a preset iteration count threshold, which can be an empirical value. For example, C0 = 100.

[0070] In summary, the data desensitization optimization algorithm provided by the embodiments of the present invention aims at the limited information problem environment with a single thread number feature and an unclear data processing method feature in the original data set. It establishes a mathematical model of file data processing time based on personalized time-consuming influence coefficients and clustering, and verifies the effectiveness of the mathematical model through statistical tests. According to the constructed mathematical model, the research goal is established as a multi-objective optimization problem, an evaluation function is designed according to the optimization goal, and the non-dominated sorting genetic algorithm is used to solve the objective function, so as to effectively optimize the data processing time of the information to be processed with the least resource consumption.

[0071] Another embodiment of the present invention provides a data desensitization optimization device, and the device includes the following steps: A data processing module, configured to cluster all the original data processing information into K clusters based on the original data processing information set to be processed and the K data processing methods used by the original data processing information set. Among them, the original value of the data processing method identification value used by each original data processing information belonging to the r-th cluster is r. Each original data processing information includes the file ID corresponding to the information and the original values of the data processing parameters, and the data processing parameters include file size, total number of file rows, file shard number, number of rows in each file shard, total number of threads, and data processing time. The value range of r is from 1 to K.

[0072] A multi-objective optimization model construction module, configured to construct a multi-objective optimization model based on all the information to be processed; wherein, the multi-objective optimization model includes constraint conditions and objective functions, and the objective functions include a first objective function corresponding to minimizing the data processing time, a second objective function corresponding to minimizing the data preprocessing time, and a third objective function corresponding to minimizing the cluster center distance. Among them, the independent variables of the first objective function and the third objective function are the number of file shards, the total number of threads, and the data processing method identification value, and the independent variables of the second objective function are the number of file shards and the total number of threads.

[0073] An optimization module, configured to optimize the multi-objective optimization model based on the non-dominated sorting genetic algorithm to obtain the optimal values of the number of file shards, the total number of threads, and the data processing method identification value corresponding to each piece of information to be processed.

[0074] This device can be used to execute Figure 1 the method shown in the embodiments shown, therefore, for the functions that can be realized by each functional module of this device, reference can be made to Figure 1 the description of the embodiments shown, and details will not be repeated here.

[0075] An embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are set to execute the method of the embodiment of the present invention.

[0076] An embodiment of the present invention further provides a computer-readable storage medium, storing computer-executable instructions, and the computer instructions are used to execute the method of the embodiment of the present invention.

[0077] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in the present invention can be achieved, and no limitations are imposed herein.

[0078] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data processing optimization method, characterized in that: The method comprises the following steps: S100, based on the original data processing information set to be processed and the K data processing methods used by the original data processing information set, cluster all the original data processing information into K clusters, wherein the original value of the data processing method identification value used by each original data processing information belonging to the rth cluster is r; wherein each original data processing information includes the file ID corresponding to the information and the original value of the data processing parameter, wherein the data processing parameter includes the file size, the total number of file lines, the number of file fragments, the number of file fragment lines, the total number of threads and the data processing time; the value of r ranges from 1 to K; S200, adding the original value of the data processing method identification value used by each original data processing information to the original data processing information, and using the original data processing information to which the original value of the data processing method identification value is added as information to be processed; S300, constructing a multi-objective optimization model based on all the information to be processed; wherein the multi-objective optimization model includes constraints and objective functions, wherein the objective functions include a first objective function corresponding to minimizing data processing time, a second objective function corresponding to minimizing data preprocessing time, and a third objective function minimizing cluster center distance, wherein the independent variables of the first objective function and the third objective function are the number of file fragments, the total number of threads, and the data processing method identification value, and the independent variables of the second objective function are the number of file fragments and the total number of threads; S400, optimizing the multi-objective optimization model based on a non-dominated sorting genetic algorithm to obtain optimal values ​​of the number of file fragments, the total number of threads and the data processing method identification value corresponding to each information to be processed.

2. The method according to claim 1, characterized in that: S100 specifically includes: S110, randomly selecting K pieces of original data processing information from the original data processing information set as K initial cluster centers; S120, for any original data processing information i, obtain the similarity S between the original data processing information i and the jth current cluster center in the current cluster center set. ij , get the similarity set S corresponding to the original data processing information i i , and add the original data processing information i to S i The data set corresponding to the current cluster center corresponding to the minimum similarity in ; the value of i ranges from 1 to n, n is the number of original data processing information, and the value of j ranges from 1 to K; the initial value of the current cluster center set is K initial cluster centers, and the initial value of the data set corresponding to the current cluster center is the original data processing information corresponding to the current cluster center; S130, obtaining the intra-cluster discreteness corresponding to the data set corresponding to each current cluster center. If the intra-cluster discreteness corresponding to any current cluster center is less than the set discreteness, the data sets corresponding to the K current cluster centers are combined as the K clusters, and S140 is executed; otherwise, the current cluster center is updated based on the data set corresponding to each current cluster center, and S120 is executed; S140, setting the data processing method identification value of the original data processing information belonging to the rth cluster to r.

3. The method according to claim 2, characterized in that The intra-cluster dispersion D corresponding to the jth current cluster center j The following conditions must be met: D j =(1 / 4)∑ 4 v=1 ∑ z(j) u=1 (N ju v -N0 j v )) / z(j)),N ju v is the original value of the vth associated parameter in the uth data in the data set corresponding to the jth current cluster center, N0 j v is the original value of the vth associated parameter corresponding to the jth current cluster center, u ranges from 1 to z(j), z(j) is the number of data in the data set corresponding to the jth current cluster center, v ranges from 1 to 4, and the associated parameters include file size, total number of file lines, number of file shards, and total number of threads.

4. The method according to claim 3, characterized in that The first objective function Y1 corresponding to the g-th information to be processed g The following conditions must be met: Y1 g =min [(1+ min (β×NF g +γ×NL g ))×α g ×(Size g / (NF g +NL g ))+min(β×NF g +γ×NL g )]; NF g is the number of file fragments of the g-th information to be processed during the optimization process, NL g is the total number of threads for the gth information to be processed during the optimization process, α g is the time-consuming impact coefficient of the g-th information to be processed, Size g is the original value of the file size of the gth information to be processed; the value of g ranges from 1 to n, where n is the number of information to be processed; β represents the file shard weight, and γ represents the thread weight; The second objective function Y2 corresponding to the g-th information to be processed g The following conditions must be met: Y2 g =min (β×NF g +γ×NL g ); β represents the file shard weight, γ represents the thread weight; The third objective function Y3 corresponding to the g-th information to be processed g The following conditions must be met: Y3 g =min (∑ 4 v=1 (N gv -N0 gv ) 2 / (N0 gv ) 2 ), N gv is the parameter value of the vth associated parameter of the gth information to be processed in the optimization process, N0 gv is the parameter value of the vth data processing parameter of the cluster center to which the data processing method identification value corresponding to the gth information to be processed belongs in the optimization process; The constraints corresponding to the g-th information to be processed satisfy the following conditions: (1) 1 ≤ C g ≤K; C g is the data processing method identification value of the g-th information to be processed; (2) 1<NL g <NL0; NL0 is the total thread number threshold; (3) n1≤NR g / NF g ≤n2,NR g is the total number of lines in the file of the gth information to be processed, n1 is the minimum file segment line number threshold, and n2 is the maximum file segment line number threshold.

5. The method according to claim 4, characterized in that α g The following conditions must be met: α g =T g ×(Size g -1 +N g -1 )×NL g org ×NF g org -λ; Among them, T g is the original value of the data processing time corresponding to the g-th information to be processed, N g is the total number of lines in the file corresponding to the g-th information to be processed, NL g org is the original value of the total number of threads corresponding to the g-th information to be processed, NF g org is the original value of the number of file fragments corresponding to the g-th information to be processed, and λ is the adjustment factor.

6. The method according to claim 4, characterized in that β, γ and λ are obtained by the following steps: S10, substitute the original values ​​of the file size, the total number of file lines, the number of file fragments, the total number of threads and the data processing method identification value of each information to be processed into the current data processing time calculation function to obtain the current predicted data processing time corresponding to each information to be processed, and obtain n currently predicted data processing times; the initial value of the current data processing time calculation function is a function for initializing β, γ and λ; the data processing time calculation function satisfies the following conditions: T=[(1+(β×NF+γ×NL))×α×(Size / (NF+NL))+(β×NF+γ×NL)], T is the data processing time calculation function, α is the time consumption influence coefficient, Size is the file size, NF is the number of file fragments, and NL is the total number of threads; S20, obtain the deviations between the n currently predicted data processing times and the corresponding n actual data processing times. If the deviations are less than or equal to the set deviations, use the current β, γ and λ as the final β, γ and λ. Otherwise, adjust the β, γ and λ in the current data processing time calculation function based on the deviations and execute S10.

7. The method according to claim 1, characterized in that S300 specifically includes: S301, randomly selecting P1 pieces of information to be processed from n pieces of information to be processed as an initial population; S302, perform a non-dominated sorting operation on the current population to be processed, obtain a non-dominated solution in the current population, if a non-dominated solution is obtained, use the obtained non-dominated solution as the current parent generation, and execute S303; if no non-dominated solution is obtained, execute S301; the initial value of the current population to be processed is the initial population; S303, performing a crossover operation on the current parent generation, so that any two individuals in the current parent generation randomly exchange the number of file shards, the total number of threads and the data processing method identification value, and obtaining a corresponding crossover operation result as the current crossover operation result; wherein one individual represents one piece of information to be processed; S304, performing a mutation operation on the current crossover operation result, so that the number of file fragments, the total number of threads, and the data processing method identification value of the individual in the current crossover operation result are randomly mutated, and a corresponding mutation result is obtained as the current offspring; S305, merging the current parent generation and the current child generation to obtain the current merged population, and obtaining the data processing time and the optimized data processing time value of each individual in the current merged population, which are used as the current data processing time and the optimized data processing time value of the individual respectively; the optimized current data processing time value of each individual is equal to the difference between the current data processing time of the individual and the previous data processing time; S306, sort all current data processing time optimization values ​​corresponding to the current merged population in descending order, and use the individuals corresponding to the first P1 optimization values ​​in the sorted data processing time optimization values ​​as the current population to be processed, and add the current data processing time of each individual in the current merged population to the corresponding current data processing time record set; set the iteration counter C=C+1, and the initial value of C is 0; the initial value of the current data processing time record set is empty; S307, if C≤C0, execute S302; if C>C0, use the number of file fragments, total number of threads and data processing method identification value corresponding to the maximum data processing time in the current data processing time record set corresponding to each individual as the optimal value of the number of file fragments, total number of threads and data processing method identification value corresponding to the information to be processed; C0 is the preset iteration number threshold.

8. A data desensitization optimization device, characterized in that: The device comprises the following steps: A data processing module, for clustering all raw data processing information into K clusters based on the raw data processing information set to be processed and the K data processing methods used by the raw data processing information set, wherein the original value of the data processing method identification value used by each raw data processing information belonging to the rth cluster is r; wherein each raw data processing information includes a file ID corresponding to the information and an original value of a data processing parameter, wherein the data processing parameter includes a file size, a total number of file lines, a number of file fragments, a number of file fragment lines, a total number of threads, and a data processing time; and the value of r ranges from 1 to K; A multi-objective optimization model construction module is used to construct a multi-objective optimization model based on all the information to be processed; wherein the multi-objective optimization model includes constraints and objective functions, and the objective functions include a first objective function corresponding to minimizing data processing time, a second objective function corresponding to minimizing data preprocessing time, and a third objective function minimizing cluster center distance, wherein the independent variables of the first objective function and the third objective function are the number of file fragments, the total number of threads, and the data processing method identification value, and the independent variables of the second objective function are the number of file fragments and the total number of threads; The optimization module is used to optimize the multi-objective optimization model based on a non-dominated sorting genetic algorithm to obtain the optimal values ​​of the number of file fragments, the total number of threads and the data processing method identification value corresponding to each information to be processed.

9. An electronic device, characterized in that: including a processor and a memory; The processor is used to execute the steps of the method according to any one of claims 1 to 7 by calling the program or instruction stored in the memory.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a program or an instruction, wherein the program or the instruction enables a computer to execute the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data processing time consumption optimization method and system

    CN114518914A

  • A Multi-Objective Optimization Method for Rotor Systems Based on Cluster Analysis

    CN116805095A

  • High-throughput genome sequence data compression parallel optimization method

    CN117059181A

  • Information processing method and device and electronic equipment

    CN118195242A

  • Parameter set determination for clustering of datasets

    US20170255688A1