Data processing method and system for synchronous auditing
By constructing audit sharding based mapping equations and optimizing sharding operations and thread count, the problems of high server pressure and low efficiency during the audit process are solved, and efficient audit processing is achieved.
Patent Information
- Application Number
- CN202510647122.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-15
AI Technical Summary
There is a risk of missing key differences through sampling auditing during the existing audit process, and large data audits have caused great pressure on server operation, affecting efficiency.
By setting the audit data sharding based on dimensions and sharding based selection parameters, constructing the audit sharding based mapping equation, and using the Firefly optimization algorithm to adjust the constant coefficient, optimize the sharding operation and thread count to reduce server pressure and improve audit efficiency.
It reduces the operating pressure of the server, improves the audit efficiency, and meets the mapping accuracy and time-consuming requirements of the audit process.
Smart Images

Figure CN120492443A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of audit data processing, and in particular, relates to a data processing method and system for synchronous auditing. Background Art
[0002] A Chinese patent with announcement number CN117992441B discloses a data processing method and system for synchronization auditing. During data synchronization from a data source address to a data target address, part of the sent data sent by the data source address and part of the received data received by the data target address are obtained according to a preset ratio; according to a data volume matching rule, the part of the sent data and the part of the received data are matched to obtain a matching result; according to a data sampling audit algorithm, a sampling audit is performed on the part of the sent data and the part of the received data to obtain an audit result; based on the matching result and the audit result, the synchronization success result corresponding to the part of the sent data and the part of the received data is determined.
[0003] The existing audit process uses sampling audits to audit large amounts of data. This audit method is risky, as sampling comparisons may miss key differences and make it impossible to audit all data. However, large-scale data comparisons require loading data into memory or intermediate storage, which puts a lot of pressure on the server and affects audit efficiency. Summary of the Invention
[0004] In response to the problems in the related art, the present invention proposes a data processing method and system for synchronous auditing to overcome the above technical problems existing in the existing related art.
[0005] To solve the above technical problems, the present invention is achieved through the following technical solutions: The present invention is a data processing method for synchronous auditing, comprising the following steps: S1. Set several types of audit data sharding basis dimensions and corresponding several types of sharding basis selection and determination parameter types to obtain an audit data sharding basis dimension type set and an audit data sharding basis selection parameter type set; collect several sets of data to be audited and corresponding audit data sharding basis selection parameters in conjunction with the audit data sharding basis dimension type set and the audit data sharding basis selection parameter type set to obtain a current data set to be audited and a current data sharding basis selection parameter matrix; S2. Construct a final audit sharding basis mapping equation based on several sets of historical audit data, an audit sharding basis dimension type set, and an audit sharding basis selection parameter type set; S3. Shard the current dataset to be audited according to the mapping equation and the parameter matrix selected based on the final audit sharding, and obtain the audit data of the current dataset to be audited after sharding; record the total time consumed for each audit operation on the current dataset to be audited, and obtain the current total audit time dataset; S4. Adjust the number of threads for sharding the current audit data corresponding to the data in the current audit total time consumption data set that does not meet the requirements; This solution optimizes the audit efficiency by sharding the current data to be audited and adjusting the number of threads for the sharding operation. The sharding operation helps reduce the operating pressure of the server; adjusting the number of threads for the sharding operation reduces the time consumed by the sharding operation, thereby improving the audit efficiency while reducing the pressure on the server.
[0006] Preferably, the S1 comprises the following steps: S11. Set several groups of data to be audited to obtain a current set of data to be audited; then set several types of audit data sharding basis dimensions to obtain an audit sharding basis dimension type set; and, in conjunction with the audit sharding basis dimension type set, set several sharding basis selection determination parameter types corresponding to each audit sharding basis dimension type to obtain an audit sharding basis selection parameter type set; S12. Acquire the audit sharding selection parameters corresponding to each current data set to be audited based on the current data set to be audited and the audit sharding selection parameter type set, and obtain the current audit sharding selection parameter matrix; By setting the audit sharding basis dimension type set and the audit sharding basis parameter type set, it provides a collection basis for the subsequent collection of corresponding historical data, and determines the parameter type and mapping data type for the subsequent construction of the mapping equation. When the amount of audit data is large, the appropriate data sharding rules can be selected according to the adaptability of different types of audit data, thereby sharding the audit data. In this way, data from different shards can be distributedly processed, which improves processing efficiency and reduces server pressure.
[0007] Preferably, said S2 comprises the following steps: S21. Collect several sets of historical audit data to obtain a historical audit data set; S22. Construct a final audit sharding basis mapping equation based on the historical audit data set, the audit sharding basis dimension type set, and the audit sharding basis selection parameter type set; By collecting historical audit data sets, data support is provided for constructing the final audit sharding mapping equation; by constructing the final audit sharding mapping equation, corresponding sharding rules are set for the audit data according to the different feature types of the audit data, thereby improving data processing efficiency.
[0008] Preferably, the S22 includes the following steps: S221. Collect several sets of historical audit data to obtain historical audit data sets. , a i Indicates the collected i The data of group history audit, Indicates the total number of groups of historical audited data collected; then, based on the audit sharding based dimension type set and the audit sharding based selection parameter type set, collect the sharding based dimension data and sharding based selection parameters corresponding to each historical audit data in the historical audit data set to obtain the historical audit sharding based dimension data set. And the historical audit sharding is based on the selection parameter matrix ;as follows, ; in, Indicates the collected i The data corresponding to the audited group history j Types of audit sharding are based on selected parameters. Indicates the total number of audit shards based on the selected parameter types; S222: Construct an initial audit sharding mapping equation based on the historical audit sharding dimension data set; as follows: ; in, b 1 is the dependent variable of the mapping equation based on the initial audit sharding, indicating that the audit sharding is based on dimensional data; b 2 represents the mapping relationship of the initial audit shards according to the mapping equation; b 3i The initial audit segment is based on the mapping equation i The independent variable represents the i Types of sharding are based on selection parameters; S223. Set a numerical encoding rule for selecting parameters based on the sharding basis; numerically encode each historical audit sharding basis selection parameter in the historical audit sharding basis selection parameter matrix according to the numerical encoding rule for selecting parameters based on the sharding basis selection, and obtain a historically encoded audit sharding basis selection parameter matrix; Substitute each row of data in the selected parameter matrix of the historically coded audit shard into the initial audit shard mapping equation for mapping, and obtain the historical audit shard initial mapping dataset based on the dimension. , Indicates that the audit fragment after historical coding is selected according to the parameter matrix i Substitute the row data into the initial audit shard and map it according to the mapping equation to obtain the data; S224, set the slice selection parameter mapping error threshold; calculate the error data between the historical audit slice dimension initial mapping data set and the historical audit slice dimension data set, and obtain the slice selection parameter mapping error data. The calculation formula is as follows: ; When the slice basis selection parameter mapping error data is greater than or equal to the slice basis selection parameter mapping error threshold, the initial audit slice basis mapping equation is adjusted until the slice basis selection parameter mapping error data is less than the slice basis selection parameter mapping error threshold; otherwise, no adjustment is required; By setting the number of independent variables of the initial audit sharding basis mapping equation to the total number of the set audit sharding basis selection parameter types, it is convenient to directly map multiple audit sharding basis selection parameters to corresponding audit sharding basis dimensional data in the future, and then use the audit sharding basis dimensional data obtained by the mapping to perform sharding processing on the audit data; wherein, since non-numerical data may exist in the audit sharding basis selection parameters, the audit sharding basis selection parameters are numerically encoded by setting the sharding basis selection parameter numerical encoding rules, so as to facilitate subsequent input into the initial audit sharding basis mapping equation for mapping; by inputting the historically encoded audit sharding basis selection parameter matrix into the initial audit sharding basis mapping equation for pre-mapping, it is possible to detect whether the mapping accuracy of the initial audit sharding basis mapping equation meets the requirements, and then determine whether the initial audit sharding basis mapping equation needs to be adjusted; by setting the sharding basis selection parameter mapping error threshold, a quantitative judgment basis is provided for determining whether the mapping accuracy of the initial audit sharding basis mapping equation meets the requirements.
[0009] Adjusting the initial audit shards according to the mapping equation in S224 includes the following steps: S2241, set the initial audit segmentation based on the value intervals of several constant coefficients in the mapping equation, and obtain the segmentation based on the mapping equation constant coefficient value interval set ;as follows, ; in, 、 Respectively represent the initial audit sharding according to the mapping equation i The lower and upper limits of the constant coefficients, Represents the total number of constant coefficients of the mapping equation for the initial audit shard; Construct audit shards to adjust the firefly population according to the mapping system; set the maximum number of iterations of the audit shards to adjust the firefly population according to the mapping system c 1 and the current number of iterations is c 2, respectively recorded as the maximum number of iterations for shard adjustment and the current number of iterations for shard adjustment; the audit shard adjusts the search space dimension of the firefly population according to the mapping system. same; S2242, according to the slices, set the audit slices according to the mapping equation constant coefficient value interval set, adjust the initial position of each firefly in the firefly population according to the mapping constant, and obtain the first initial position matrix ;as follows, ; in, Indicates that the audit sharding adjusts the number of fireflies in the firefly population according to the mapping system. j The initial position of the firefly is in the initial audit segment according to the mapping equation i The position component in the constant coefficient dimension, Indicates that the audit shard adjusts the size of the firefly population according to the mapping system; The generation formula is as follows: ; Where, rand 1ji Indicates that Generate a random number between 0 and 1; S2243, construct the fitness function of the audit segment to adjust the firefly population according to the mapping system ;as follows, ; in, d Indicates that a set of constant coefficients obtained in each round of iteration is substituted into the initial audit sharding basis mapping equation, and then the parameter matrix of the audit sharding basis selected after historical coding is input into the initial audit sharding basis mapping equation to map the obtained data set to the error data between the historical audit sharding basis dimension data set; S2244, start iteration, before iteration, set the current iteration number of the shard adjustment to 1; in the first round of iteration, use the audit shard to adjust the fitness function of the firefly population according to the mapping constant Calculating the fitness value of the initial position of each firefly in the first initial position matrix to obtain a first fitness value set; using the maximum fitness value in the first fitness value set and the initial position of the corresponding firefly as the first global optimal fitness and the first global optimal position, respectively; updating the initial position of each firefly in the first initial position matrix according to the first global optimal fitness and the first global optimal position; after the update is completed, adding 1 to the current iteration number of the slice adjustment and entering the next iteration; In each other round of iteration, the fitness function of the firefly population is adjusted according to the mapping system by using the audit sharding Calculate the fitness value of each firefly position in the firefly population adjusted by the mapping constant for the audit shard obtained during the previous iteration to obtain a second fitness value set; use the maximum fitness value and the corresponding firefly position in the second fitness value set as the second global optimal fitness and the second global optimal position, respectively; update the audit shard obtained during the previous iteration by adjusting the position of each firefly in the firefly population according to the mapping constant based on the second global optimal fitness and the second global optimal position; after the update is completed, increase the current iteration number of the shard adjustment by 1 and enter the next iteration; S2245, when When , stop the iteration and get the first final global best position and the first final global best fitness; otherwise, continue to iterate until ; using the first final global optimal fitness as the optimized slice basis selection parameter mapping error data; when the optimized slice basis selection parameter mapping error data is less than the slice basis selection parameter mapping error threshold, substituting each position component of the first final global optimal position into the initial audit slice basis mapping equation to obtain the final audit slice basis mapping equation; otherwise, returning to S2244 to continue iterating until the optimized slice basis selection parameter mapping error data is less than the slice basis selection parameter mapping error threshold; The firefly optimization algorithm can quickly move towards the optimal solution in the search space by simulating the flashing behavior and mutual attraction mechanism of fireflies, so that a better solution can be found within a relatively small number of iterations, thereby improving the efficiency of the algorithm; through information exchange and mutual attraction between fireflies, the algorithm can quickly converge to the global optimal solution or a region close to the global optimal solution; it has strong robustness and does not require excessive adjustment and optimization of the initial values; it has a certain adaptability to parameter changes, and when the parameter values are adjusted within a certain range, the algorithm can still maintain a good search effect; based on the above advantages, this scheme adopts the firefly optimization algorithm to simultaneously perform multiple iterative adjustments on the initial audit partition based on several constant coefficients in the mapping equation, and uses the mapping accuracy of the initial audit partition based on the mapping equation as the fitness function; as the iteration proceeds, the mapping accuracy of the initial audit partition based on the mapping equation becomes higher and higher, and finally meets the mapping requirements.
[0010] Preferably, the step S3 includes the following steps: S31. Set a data sharding data volume threshold. When the volume of the data currently to be audited in the current data set to be audited is greater than or equal to the data sharding data volume threshold, the current data to be audited is used as the current data to be audited and the process proceeds to S32. Otherwise, there is no need to perform sharding on the data in the current data set to be audited. S32. Numerically encode the parameter set for selecting the sharding basis to be audited corresponding to the audit data for the current sharding based on the selected parameter matrix and the numerical encoding rule for selecting the sharding basis to be audited, to obtain the currently encoded parameter set for selecting the sharding basis to be audited. Input the selected parameter set of the current encoded sharding basis into the final audit sharding basis mapping equation for mapping to obtain the current audit sharding basis dimension data; S33: Slice the current audit data to be sliced according to the dimension data of the current audit slice; after the processing is completed, obtain the current audit data to be sliced after slicing; and perform an audit operation on the current audit data to be sliced after slicing; S34. Repeat S11, S12, S31, S32, and S33 to record the total time consumed by each audit operation on the current dataset to be audited, and obtain the current total audit time dataset. By setting a data sharding data volume threshold, when the data volume of the data to be audited exceeds the data sharding data volume threshold, if sharding is not performed, the data needs to be read into the server memory all at once, which puts greater operating pressure on the server; through data sharding, different shards can be read into the memory of different servers through distributed technology for distributed processing, reducing the operating pressure of each server; due to the addition of sharding operations, by recording the total time data of each audit, the impact of sharding operations on the efficiency of the entire audit process can be determined so that corresponding measures can be taken.
[0011] Preferably, the S4 comprises the following steps: S41, setting a current audit total time threshold according to the total amount of data in the current data set to be audited; S42. When the current audit total time consumption data set contains data of the current audit total time consumption that is greater than or equal to the current audit total time consumption threshold, the number of threads corresponding to the sharding operation on the current audit data is adjusted until the current audit total time consumption data set contains data of the current audit total time consumption that is greater than or equal to the current audit total time consumption threshold; otherwise, the process proceeds to S43. S43. Set several future time points to obtain a future time point set; use the future time point set and the current audit total time data set to predict the audit total time data of the future time points to obtain the future audit total time data set; when there is future audit total time data in the future audit total time data set that is greater than or equal to the current audit total time threshold, adjust the number of threads corresponding to the sharding operation on the current audit data until there is no future audit total time data in the future audit total time data set that is greater than or equal to the current audit total time threshold; otherwise, there is no need to adjust the number of threads corresponding to the sharding operation on the current audit data; By setting the current total audit time threshold, a quantitative basis is provided for determining whether the total audit time data caused by the sharding operation on the current audit data meets the requirements; since the number of threads for data sharding directly affects the time consumed for sharding the data; and the number of threads for data sharding is not the larger the better, since the total number of threads of the server is fixed, it is necessary to adjust the number of threads for the sharding operation on the current audit data to an appropriate number so that the time data of the entire audit operation meets the requirements; since the time data may change as the audit time progresses, through prediction, the situation where the audit time data exceeds the threshold in the future can be prevented in advance, thereby ensuring the efficiency of the audit.
[0012] Preferably, in S43, the total audit time data at a future time point is predicted using a BP neural network model; During training, the BP neural network can automatically extract "reasonable rules" between input and output data through learning, and adaptively memorize the learning content in the network's weights. It can be adjusted and optimized according to different problems and data sets, and has strong adaptability. Through training, the BP neural network can learn some common features, so that it has good generalization performance when performing classification or prediction, can apply learning results to new knowledge, and has a certain ability to correctly classify unseen patterns or patterns with noise pollution.
[0013] Preferably, the step of adjusting the number of threads corresponding to the sharding operation on the current audit data in S42 and S43 includes the following steps: S421. Set the value range of the number of threads corresponding to the sharding operation on the current audit data to obtain the value range of the number of threads for the current audit sharding operation. , 、 Respectively represent the lower limit and upper limit of the number of threads corresponding to the sharding operation on the current audit data; Construct the current audit shard thread number to adjust the firefly population; set the maximum number of iterations of the current audit shard thread number to adjust the firefly population c 3 and the current number of iterations is c 4, respectively recorded as the maximum number of iterations adjusted by the thread and the current number of iterations adjusted by the thread; the search space dimension of the firefly population is adjusted to one dimension by the number of threads in the current audit shard; S422, according to the current audit shard thread number value range, set the current audit shard thread number to adjust the initial position of each firefly in the firefly population, and obtain the second initial position set , Indicates that the current audit shard thread number is adjusted in the firefly population. i The initial position of the fireflies, Indicates that the current number of audit shard threads adjusts the size of the firefly population; The generation formula is as follows: ; Where, rand 2i Indicates that Generate a random number between 0 and 1; S423: Construct the fitness function of the firefly population by adjusting the number of threads in the current audit shard. ;as follows, ; Where, Indicates the total time consumed by the current audit process after the number of threads obtained in each iteration is set to the operation of sharding the current audit data; S424, start iteration, before iteration, set the thread to adjust the current iteration number to 1; in the first round of iteration, use the current audit shard thread number to adjust the firefly population fitness function Calculating the fitness value of the initial position of each firefly in the second initial position set to obtain a third fitness value set; using the maximum fitness value in the third fitness value set and the initial position of the corresponding firefly as the third global optimal fitness value and the third global optimal position, respectively; updating the initial position of each firefly in the second initial position set according to the third global optimal fitness value and the third global optimal position; after the update is completed, adjusting the current iteration number of the thread by 1 and entering the next iteration; In each other round of iteration, the fitness function of the firefly population is adjusted by using the current number of audit shard threads. Calculate the fitness value of each firefly position in the firefly population adjusted by the current audit shard thread number obtained during the previous iteration to obtain a fourth fitness value set; use the maximum fitness value and the corresponding firefly position in the fourth fitness value set as the fourth global optimal fitness and the fourth global optimal position, respectively; update the position of each firefly in the firefly population adjusted by the current audit shard thread number obtained during the previous iteration according to the fourth global optimal fitness and the fourth global optimal position; after the update is completed, increase the current iteration number of the thread adjustment by 1 and enter the next iteration; S425, when When , stop the iteration and get the second final global best fitness and the second final global best position; otherwise, continue to iterate until ; using the second final global optimal fitness as the optimized current audit total time data; when the optimized current audit total time data is less than the current audit total time threshold, setting the second final global optimal position as the optimized number of threads to the current audit data for sharding operation; otherwise, returning to S424 to continue iterating until the optimized current audit total time data is less than the current audit total time threshold; The number of threads for the current audit data sharding is iteratively adjusted multiple times by using the Firefly optimization algorithm, and the total time consumed by the current audit process is used as the fitness function; therefore, as the iteration proceeds, the total time consumed by the current audit process becomes shorter and shorter, ultimately meeting the efficiency requirements of the audit.
[0014] A data processing system for synchronous auditing includes an audit data sharding basis setting module, a current audit data collection module, an audit sharding basis mapping equation construction module, a current audit data sharding module, a current audit total time consumption data recording module, and a current audit sharding thread number adjustment module.
[0015] The present invention has the following beneficial effects: 1. The present invention optimizes the audit efficiency by sharding the current data to be audited and adjusting the number of threads for the sharding operation. The sharding operation is beneficial to reducing the operating pressure of the server; adjusting the number of threads for the sharding operation reduces the time consumption caused by the sharding operation, thereby reducing the pressure on the server while improving the audit efficiency.
[0016] 2. In the present invention, by constructing the final audit sharding based on the mapping equation, corresponding sharding rules are set for the audit data according to different feature types of the audit data, thereby improving data processing efficiency.
[0017] 3. In the present invention, the firefly optimization algorithm is used to perform multiple iterative adjustments on the initial audit segment based on several constant coefficients in the mapping equation, and the mapping accuracy of the initial audit segment based on the mapping equation is used as the fitness function; as the iteration proceeds, the mapping accuracy of the initial audit segment based on the mapping equation becomes higher and higher, and finally meets the mapping requirements.
[0018] 4. In the present invention, the number of threads for the current audit data sharding is iteratively adjusted multiple times by using the firefly optimization algorithm, and the total time consumption data of the current audit process is used as the fitness function; therefore, as the iteration proceeds, the total time consumption of the current audit process becomes shorter and shorter, ultimately meeting the efficiency requirements of the audit.
[0019] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 The figure is a flow chart of a data processing method for synchronous auditing according to the present invention. DETAILED DESCRIPTION
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0023] Example 1
[0024] See also Figure 1 This embodiment is a data processing method for synchronous auditing, comprising the following steps: S1. Set several types of audit data sharding basis dimensions and corresponding several types of sharding basis selection and determination parameter types to obtain an audit data sharding basis dimension type set and an audit data sharding basis selection parameter type set; collect several sets of data to be audited and corresponding audit data sharding basis selection parameters in conjunction with the audit data sharding basis dimension type set and the audit data sharding basis selection parameter type set to obtain a current data set to be audited and a current data sharding basis selection parameter matrix; Said S1 comprises the following steps: S11. Set several groups of data to be audited to obtain the current data set to be audited; then set several types of audit data sharding dimensions to obtain an audit sharding dimension type set; the audit sharding dimension type set includes time range, primary key or business ID, file hash value, and data access frequency, etc. Each audit sharding dimension type is expressed using a natural number, such as 1 for time range, etc.; in conjunction with the audit sharding dimension type set, set several sharding basis selection and determination parameter types corresponding to each audit sharding dimension type to obtain an audit sharding basis selection parameter type set; the audit sharding basis selection parameter type set includes audit data field type, the most queried field in the most recent time period, and the query efficiency of each audit data field type, etc. These data can be obtained based on the data query log; S12. Acquire the audit sharding selection parameters corresponding to each current data set to be audited based on the current data set to be audited and the audit sharding selection parameter type set, and obtain the current audit sharding selection parameter matrix; S2. Construct a final audit sharding basis mapping equation based on several sets of historical audit data, an audit sharding basis dimension type set, and an audit sharding basis selection parameter type set; The S2 comprises the following steps: S21. Collect several sets of historical audit data to obtain a historical audit data set; S22. Construct a final audit sharding basis mapping equation based on the historical audit data set, the audit sharding basis dimension type set, and the audit sharding basis selection parameter type set; The S22 includes the following steps: S221. Collect several sets of historical audit data to obtain historical audit data sets. , a i Indicates the collected i The data of group history audit, Indicates the total number of groups of historical audited data collected; then, based on the audit sharding based dimension type set and the audit sharding based selection parameter type set, collect the sharding based dimension data and sharding based selection parameters corresponding to each historical audit data in the historical audit data set to obtain the historical audit sharding based dimension data set. And the historical audit sharding is based on the selection parameter matrix ;as follows, ; in, Indicates the collected i The data corresponding to the audited group history j Types of audit sharding are based on selected parameters. Indicates the total number of audit shards based on the selected parameter types; S222: Construct an initial audit sharding mapping equation based on the historical audit sharding dimension data set; as follows: ; in, b 1 is the dependent variable of the mapping equation based on the initial audit sharding, indicating that the audit sharding is based on dimensional data; b 2 represents the mapping relationship of the initial audit shards according to the mapping equation; such as the proportional relationship or inverse proportional relationship of different shards according to the selected parameters, which is used to combine various types of shards according to the selected parameters by mathematical equations; b 3i The initial audit segment is based on the mapping equation i The independent variable represents the i Types of sharding are based on selection parameters; S223. Set a numerical encoding rule for selecting parameters based on the sharding basis; numerically encode each historical audit sharding basis selection parameter in the historical audit sharding basis selection parameter matrix according to the numerical encoding rule for selecting parameters based on the sharding basis selection, and obtain a historically encoded audit sharding basis selection parameter matrix; Substitute each row of data in the selected parameter matrix of the historically coded audit shard into the initial audit shard mapping equation for mapping, and obtain the historical audit shard initial mapping dataset based on the dimension. , Indicates that the audit fragment after historical coding is selected according to the parameter matrix i Substitute the row data into the initial audit shard and map it according to the mapping equation to obtain the data; S224, set the slice selection parameter mapping error threshold; calculate the error data between the historical audit slice dimension initial mapping data set and the historical audit slice dimension data set, and obtain the slice selection parameter mapping error data. The calculation formula is as follows: ; When the slice basis selection parameter mapping error data is greater than or equal to the slice basis selection parameter mapping error threshold, the initial audit slice basis mapping equation is adjusted until the slice basis selection parameter mapping error data is less than the slice basis selection parameter mapping error threshold; otherwise, no adjustment is required; Adjusting the initial audit shards according to the mapping equation in S224 includes the following steps: S2241, set the initial audit segmentation based on the value intervals of several constant coefficients in the mapping equation, and obtain the segmentation based on the mapping equation constant coefficient value interval set ;as follows, ; in, 、 Respectively represent the initial audit sharding according to the mapping equation i The lower and upper limits of the constant coefficients, Represents the total number of constant coefficients of the mapping equation for the initial audit shard; Construct audit shards to adjust the firefly population according to the mapping system; set the maximum number of iterations of the audit shards to adjust the firefly population according to the mapping system c 1 and the current number of iterations is c 2, respectively recorded as the maximum number of iterations for shard adjustment and the current number of iterations for shard adjustment; the audit shard adjusts the search space dimension of the firefly population according to the mapping system. same; S2242, according to the slices, set the audit slices according to the mapping equation constant coefficient value interval set, adjust the initial position of each firefly in the firefly population according to the mapping constant, and obtain the first initial position matrix ;as follows, ; in, Indicates that the audit sharding adjusts the number of fireflies in the firefly population according to the mapping system. j The initial position of the firefly is in the initial audit segment according to the mapping equation i The position component in the constant coefficient dimension, Indicates that the audit shard adjusts the size of the firefly population according to the mapping system; The generation formula is as follows: ; Where, rand 1ji Indicates that Generate a random number between 0 and 1; S2243, construct the fitness function of the audit segment to adjust the firefly population according to the mapping system ;as follows, ; in, d Indicates that a set of constant coefficients obtained in each round of iteration is substituted into the initial audit sharding basis mapping equation, and then the parameter matrix of the audit sharding basis selected after historical coding is input into the initial audit sharding basis mapping equation to map the obtained data set to the error data between the historical audit sharding basis dimension data set; S2244, start iteration, before iteration, set the current iteration number of the shard adjustment to 1; in the first round of iteration, use the audit shard to adjust the fitness function of the firefly population according to the mapping constant Calculating the fitness value of the initial position of each firefly in the first initial position matrix to obtain a first fitness value set; using the maximum fitness value in the first fitness value set and the initial position of the corresponding firefly as the first global optimal fitness and the first global optimal position, respectively; updating the initial position of each firefly in the first initial position matrix according to the first global optimal fitness and the first global optimal position; after the update is completed, adding 1 to the current iteration number of the slice adjustment and entering the next iteration; In each other round of iteration, the fitness function of the firefly population is adjusted according to the mapping system by using the audit sharding Calculate the fitness value of each firefly position in the firefly population adjusted by the mapping constant for the audit shard obtained during the previous iteration to obtain a second fitness value set; use the maximum fitness value and the corresponding firefly position in the second fitness value set as the second global optimal fitness and the second global optimal position, respectively; update the audit shard obtained during the previous iteration by adjusting the position of each firefly in the firefly population according to the mapping constant based on the second global optimal fitness and the second global optimal position; after the update is completed, increase the current iteration number of the shard adjustment by 1 and enter the next iteration; S2245, when When , stop the iteration and get the first final global best position and the first final global best fitness; otherwise, continue to iterate until ; using the first final global optimal fitness as the optimized slice basis selection parameter mapping error data; when the optimized slice basis selection parameter mapping error data is less than the slice basis selection parameter mapping error threshold, substituting each position component of the first final global optimal position into the initial audit slice basis mapping equation to obtain the final audit slice basis mapping equation; otherwise, returning to S2244 to continue iterating until the optimized slice basis selection parameter mapping error data is less than the slice basis selection parameter mapping error threshold; S3. Shard the current dataset to be audited according to the mapping equation and the parameter matrix selected based on the final audit sharding, and obtain the audit data of the current dataset to be audited after sharding; record the total time consumed for each audit operation on the current dataset to be audited, and obtain the current total audit time dataset; The S3 includes the following steps: S31. Set a data sharding data volume threshold. When the volume of the data currently to be audited in the current data set to be audited is greater than or equal to the data sharding data volume threshold, the current data to be audited is used as the current data to be audited and the process proceeds to S32. Otherwise, there is no need to perform sharding on the data in the current data set to be audited. S32. Numerically encode the parameter set for selecting the sharding basis to be audited corresponding to the audit data for the current sharding based on the selected parameter matrix and the numerical encoding rule for selecting the sharding basis to be audited, to obtain the currently encoded parameter set for selecting the sharding basis to be audited. Input the selected parameter set of the current encoded sharding basis into the final audit sharding basis mapping equation for mapping to obtain the current audit sharding basis dimension data; S33: Slice the current audit data to be sliced according to the dimension data of the current audit slice; after the processing is completed, obtain the current audit data to be sliced after slicing; and perform an audit operation on the current audit data to be sliced after slicing; S34. Repeat S11, S12, S31, S32, and S33 to record the total time consumed by each audit operation on the current dataset to be audited, and obtain the current total audit time dataset. S4. Adjust the number of threads for sharding the current audit data corresponding to the data in the current audit total time consumption data set that does not meet the requirements; The S4 comprises the following steps: S41, setting a current audit total time threshold according to the total amount of data in the current data set to be audited; S42. When the current audit total time consumption data set contains data of the current audit total time consumption that is greater than or equal to the current audit total time consumption threshold, the number of threads corresponding to the sharding operation on the current audit data is adjusted until the current audit total time consumption data set contains data of the current audit total time consumption that is greater than or equal to the current audit total time consumption threshold; otherwise, the process proceeds to S43. S43. Set several future time points to obtain a future time point set; use the future time point set and the current audit total time data set to predict the audit total time data of the future time points to obtain the future audit total time data set; when there is future audit total time data in the future audit total time data set that is greater than or equal to the current audit total time threshold, adjust the number of threads corresponding to the sharding operation on the current audit data until there is no future audit total time data in the future audit total time data set that is greater than or equal to the current audit total time threshold; otherwise, there is no need to adjust the number of threads corresponding to the sharding operation on the current audit data; In S43, the BP neural network model is used to predict the total audit time data at future time points; Adjusting the number of threads corresponding to the sharding operation on the current audit data in S42 and S43 includes the following steps: S421. Set the value range of the number of threads corresponding to the sharding operation on the current audit data to obtain the value range of the number of threads for the current audit sharding operation. , 、 Respectively represent the lower limit and upper limit of the number of threads corresponding to the sharding operation on the current audit data; Construct the current audit shard thread number to adjust the firefly population; set the maximum number of iterations of the current audit shard thread number to adjust the firefly population c 3 and the current number of iterations is c4, respectively recorded as the maximum number of iterations adjusted by the thread and the current number of iterations adjusted by the thread; the search space dimension of the firefly population is adjusted to one dimension by the number of threads in the current audit shard; S422, according to the current audit shard thread number value range, set the current audit shard thread number to adjust the initial position of each firefly in the firefly population, and obtain the second initial position set , Indicates that the current audit shard thread number is adjusted in the firefly population. i The initial position of the fireflies, Indicates that the current number of audit shard threads adjusts the size of the firefly population; The generation formula is as follows: ; Where, rand 2i Indicates that Generate a random number between 0 and 1; S423: Construct the fitness function of the firefly population by adjusting the number of threads in the current audit shard. ;as follows, ; Where, Indicates the total time consumed by the current audit process after the number of threads obtained in each iteration is set to the operation of sharding the current audit data; S424, start iteration, before iteration, set the thread to adjust the current iteration number to 1; in the first round of iteration, use the current audit shard thread number to adjust the firefly population fitness function Calculating the fitness value of the initial position of each firefly in the second initial position set to obtain a third fitness value set; using the maximum fitness value in the third fitness value set and the initial position of the corresponding firefly as the third global optimal fitness value and the third global optimal position, respectively; updating the initial position of each firefly in the second initial position set according to the third global optimal fitness value and the third global optimal position; after the update is completed, adjusting the current iteration number of the thread by 1 and entering the next iteration; In each other round of iteration, the fitness function of the firefly population is adjusted by using the current number of audit shard threads. Calculate the fitness value of each firefly position in the firefly population adjusted by the current audit shard thread number obtained during the previous iteration to obtain a fourth fitness value set; use the maximum fitness value and the corresponding firefly position in the fourth fitness value set as the fourth global optimal fitness and the fourth global optimal position, respectively; update the position of each firefly in the firefly population adjusted by the current audit shard thread number obtained during the previous iteration according to the fourth global optimal fitness and the fourth global optimal position; after the update is completed, increase the current iteration number of the thread adjustment by 1 and enter the next iteration; S425, when When , stop the iteration and get the second final global best fitness and the second final global best position; otherwise, continue to iterate until until the second final global optimal fitness is used as the optimized current audit total time data; when the optimized current audit total time data is less than the current audit total time threshold, the second final global optimal position is used as the optimized number of threads to be set to the current audit data during the sharding operation; otherwise, return to S424 and continue to iterate until the optimized current audit total time data is less than the current audit total time threshold.
[0025] Example 2
[0026] This embodiment discloses a data processing system for synchronous auditing, which can implement the method of the above embodiment and includes an audit data sharding basis setting module, a current audit data collection module, an audit sharding basis mapping equation construction module, a current audit data sharding module, a current audit total time consumption data recording module, and a current audit shard thread number adjustment module; The audit data sharding basis setting module sets several types of audit data sharding basis dimensions and corresponding several types of sharding basis selection and determination parameter types, and obtains an audit sharding basis dimension type set and an audit sharding basis selection parameter type set; The current audit data collection module cooperates with the audit sharding based on the dimension type set and the audit sharding based on the selection parameter type set to collect several groups of data to be audited and the corresponding audit sharding based on the selection parameters to obtain the current data set to be audited and the current sharding based on the selection parameter matrix; The audit sharding is based on a mapping equation construction module, which cooperates with several groups of historical audited data, an audit sharding dimension type set, and an audit sharding based on a selected parameter type set to construct a final audit sharding based mapping equation; The current audit data sharding module cooperates with the final audit sharding to perform a sharding operation on the current data set to be audited based on the mapping equation and the parameter matrix selected based on the current sharding to be audited, and obtains the current sharded audit data after sharding; The current audit total time consumption data recording module records the total time consumption data of each audit operation on the current data set to be audited, and obtains the current audit total time consumption data set; The current audit sharding thread number adjustment module adjusts the number of threads for sharding the current audit data corresponding to the data that does not meet the requirements in the current audit total time consumption data set.
[0027] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0028] The preferred embodiments of the invention disclosed above are intended only to help illustrate the invention. These preferred embodiments do not exhaust all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.
Claims
1. A data processing method for synchronous auditing, characterized in that: The following steps are involved: S1. Set several types of audit data sharding basis dimensions and corresponding several types of sharding basis selection and determination parameter types to obtain an audit sharding basis dimension type set and an audit sharding basis selection parameter type set; Cooperate with the audit sharding based dimension type set and the audit sharding based selection parameter type set to collect several groups of data to be audited and the corresponding audit sharding based selection parameters, and obtain the current data set to be audited and the current sharding based selection parameter matrix; S2. Construct a final audit sharding basis mapping equation based on several sets of historical audit data, an audit sharding basis dimension type set, and an audit sharding basis selection parameter type set; S3. Perform a sharding operation on the current data set to be audited based on the mapping equation of the final audit sharding and the parameter matrix selected based on the current sharding to be audited, and obtain the audit data to be sharded after sharding; Record the total time consumed by each audit operation on the current audit data set to obtain the current audit total time consumed data set; S4. Adjust the number of threads for sharding the current audit data corresponding to the data that does not meet the requirements in the current audit total time consumption data set.
2. A data processing method for synchronous auditing according to claim 1, characterized in that: Said S1 comprises the following steps: S11. Set several groups of data to be audited to obtain a current set of data to be audited; then set several types of audit data sharding basis dimensions to obtain an audit sharding basis dimension type set; and, in conjunction with the audit sharding basis dimension type set, set several sharding basis selection determination parameter types corresponding to each audit sharding basis dimension type to obtain an audit sharding basis selection parameter type set; S12. Acquire the audit sharding basis selection parameters corresponding to each current data set to be audited in conjunction with the current data set to be audited and the audit sharding basis selection parameter type set, and obtain the current audit sharding basis selection parameter matrix.
3. A data processing method for synchronous auditing according to claim 2, characterized in that: The S2 comprises the following steps: S21. Collect several sets of historical audit data to obtain a historical audit data set; S22. Construct a final audit sharding basis mapping equation in conjunction with the historical audit data set, the audit sharding basis dimension type set, and the audit sharding basis selection parameter type set.
4. A data processing method for synchronous auditing according to claim 3, characterized in that: The S22 includes the following steps: S221. Collect several sets of historical audit data to obtain a historical audit data set; then, based on the audit sharding basis dimension type set and the audit sharding basis selection parameter type set, collect the sharding basis dimension data and sharding basis selection parameters corresponding to each historical audit data in the historical audit data set to obtain a historical audit sharding basis dimension data set and a historical audit sharding basis selection parameter matrix; S222: Constructing an initial audit sharding basis mapping equation based on the historical audit sharding basis dimensional data set; S223. Set a numerical encoding rule for selecting parameters based on the sharding basis; numerically encode each historical audit sharding basis selection parameter in the historical audit sharding basis selection parameter matrix according to the numerical encoding rule for selecting parameters based on the sharding basis selection, and obtain a historically encoded audit sharding basis selection parameter matrix; Substituting each row of data in the selected parameter matrix of the historically coded audit shard into the mapping equation of the initial audit shard for mapping, thereby obtaining an initial mapping data set of the historical audit shard based on the dimension; S224. Set a sharding basis selection parameter mapping error threshold; calculate the error data between the historical audit sharding basis dimension initial mapping dataset and the historical audit sharding basis dimension dataset to obtain sharding basis selection parameter mapping error data; When the error data of the slice based on the selected parameters mapping is greater than or equal to the error threshold of the slice based on the selected parameters mapping, the initial audit slice based on the mapping equation is adjusted until the error data of the slice based on the selected parameters mapping is less than the error threshold of the slice based on the selected parameters mapping; otherwise, no adjustment is required.
5. A data processing method for synchronous auditing according to claim 4, characterized in that: In S224, the initial audit fragment is adjusted according to the mapping equation using the firefly optimization algorithm.
6. A data processing method for synchronous auditing according to claim 5, characterized in that: The S3 includes the following steps: S31. Set a data sharding data volume threshold. When the volume of the data currently to be audited in the current data set to be audited is greater than or equal to the data sharding data volume threshold, the current data to be audited is used as the current data to be audited and the process proceeds to S32. Otherwise, there is no need to perform sharding on the data in the current data set to be audited. S32. Numerically encode the parameter set for selecting the sharding basis to be audited corresponding to the audit data for the current sharding based on the selected parameter matrix and the numerical encoding rule for selecting the sharding basis to be audited, to obtain the currently encoded parameter set for selecting the sharding basis to be audited. Input the selected parameter set of the current encoded sharding basis into the final audit sharding basis mapping equation for mapping to obtain the current audit sharding basis dimension data; S33: Slice the current audit data to be sliced according to the dimension data of the current audit slice; after the processing is completed, obtain the current audit data to be sliced after slicing; and perform an audit operation on the current audit data to be sliced after slicing; S34. Repeat S11, S12, S31, S32 and S33 to record the total time consumed by each audit operation on the current data set to be audited, and obtain the current audit total time consumed data set.
7. A data processing method for synchronous auditing according to claim 6, characterized in that: The S4 comprises the following steps: S41, setting a current audit total time threshold according to the total amount of data in the current data set to be audited; S42. When the current audit total time consumption data set contains data of the current audit total time consumption that is greater than or equal to the current audit total time consumption threshold, the number of threads corresponding to the sharding operation on the current audit data is adjusted until the current audit total time consumption data set contains data of the current audit total time consumption that is greater than or equal to the current audit total time consumption threshold; otherwise, the process proceeds to S43. S43. Set several future time points to obtain a set of future time points; use the set of future time points and the current total audit time data set to predict the total audit time data of future time points to obtain a future total audit time data set; when there is future total audit time data in the future total audit time data set that is greater than or equal to the current total audit time threshold, adjust the number of threads corresponding to the sharding operation on the current audit data until there is no future total audit time data in the future total audit time data set that is greater than or equal to the current total audit time threshold; otherwise, there is no need to adjust the number of threads corresponding to the sharding operation on the current audit data.
8. A data processing method for synchronous auditing according to claim 7, characterized in that: In S43, the BP neural network model is used to predict the total audit time data at future time points.
9. A data processing method for synchronous auditing according to claim 8, characterized in that: Adjusting the number of threads corresponding to the sharding operation on the current audit data in S42 and S43 includes the following steps: S421, set the value range of the number of threads corresponding to the sharding operation on the current audit data to obtain the value range of the current audit sharding thread number; construct the current audit sharding thread number adjustment firefly population; set the maximum number of iterations of the current audit sharding thread number adjustment firefly population to c 3 and the current number of iterations is c 4, respectively recorded as the maximum number of iterations adjusted by the thread and the current number of iterations adjusted by the thread; the search space dimension of the firefly population is adjusted to one dimension by the number of threads in the current audit shard; S422, setting the current audit shard thread number according to the current audit shard thread number value range to adjust the initial position of each firefly in the firefly population to obtain a second initial position set; S423: Constructing a fitness function for adjusting the number of firefly populations based on the current number of audit shard threads; S424, start iteration, before iteration, set the thread adjustment current iteration number to 1; in each iteration, use the current audit shard thread number to adjust the firefly population fitness function to calculate the fitness value of each firefly position in the current audit shard thread number adjustment firefly population updated in the previous iteration, and update the current audit shard thread number adjustment firefly population position updated in the previous iteration; after the update is completed, add 1 to the thread adjustment current iteration number and enter the next iteration; S425, when When , stop the iteration and get the second final global best fitness and the second final global best position; otherwise, continue to iterate until until the second final global optimal fitness is used as the optimized current audit total time data; when the optimized current audit total time data is less than the current audit total time threshold, the second final global optimal position is used as the optimized number of threads to be set to the current audit data during the sharding operation; otherwise, return to S424 and continue to iterate until the optimized current audit total time data is less than the current audit total time threshold.
10. A system for implementing the data processing method for synchronous auditing according to any one of claims 1 to 9.
Citation Information
Patent Citations
Data processing method and system for synchronous auditing
CN117992441B
Data fragment numerical value determination method and device, equipment and storage medium
CN116010526A
Machine learning-based data security protection system and education data security method
CN119961989A