Intelligent data optimization method, device and computer-readable storage medium

Through intelligent data optimization methods and devices, using exception removal, gray prediction and data range modification, the problem of insufficient automated data optimization in existing technologies is solved, and efficient and low-cost data optimization is achieved.

CN111259318BActive Publication Date: 2025-09-16PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010068234.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-01-19
Publication Date
2025-09-16
Estimated Expiration
2040-01-19

AI Technical Summary

Technical Problem

Existing technologies lack automated data optimization mechanisms, making it difficult to scientifically optimize data based on manual experience. In addition, big data optimization models have high hardware requirements and low cost-effectiveness.

Method used

An intelligent data optimization method is adopted, including outlier processing, grey prediction, cost value calculation and data range modification, to achieve automated data optimization through intelligent data optimization devices and computer-readable storage media.

Benefits of technology

It reduces manual intervention, improves data optimization efficiency, reduces hardware requirements, and achieves efficient data optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111259318B_ABST
    Figure CN111259318B_ABST
Patent Text Reader

Abstract

The present invention relates to artificial intelligence technology and discloses an intelligent data optimization method, including: receiving a data optimization instruction input by a user, extracting an original data set from a big data storage platform, performing anomaly removal processing on the original data set to obtain a standard data set, performing gray prediction on the standard data set to obtain a statistical information set, calculating the cost value of the statistical information set to obtain a cost data set, eliminating data in the cost data set that is greater than or equal to a preset cost threshold to obtain an optimized cost data set, performing a data range modification operation on the optimized cost data set to obtain an optimal data set, storing the optimal data set in the big data storage platform, and completing the data optimization operation. The present invention also proposes an intelligent data optimization device and a computer-readable storage medium. The present invention can realize efficient and intelligent data optimization functions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent data optimization method, device, and computer-readable storage medium. Background Art

[0002] Currently, data optimization relies heavily on manual experience and big data optimization models such as Hadoop. However, manual experience makes it difficult to scientifically optimize models based on a company's data. In other words, there is a lack of an automatic optimization mechanism and optimization model to help developers complete data optimization more quickly. Big data optimization models require powerful hardware capabilities for data expansion and support for unstructured data, and therefore have high hardware requirements. Therefore, a cost-effective data optimization method is urgently needed. Summary of the Invention

[0003] The present invention provides an intelligent data optimization method, device and computer-readable storage medium, the main purpose of which is to perform intelligent data optimization according to user optimization requirements.

[0004] To achieve the above objectives, the present invention provides an intelligent data optimization method, comprising:

[0005] Receive data optimization instructions input by the user, extract the original data set from the big data storage platform, and perform anomaly removal processing on the original data set to obtain a standard data set;

[0006] Performing grey prediction on the standard data set to obtain a statistical information set;

[0007] Calculating the cost value of the statistical information set to obtain a cost data set;

[0008] Eliminating data in the cost data set that is greater than or equal to a preset cost threshold to obtain an optimized cost data set;

[0009] The optimized cost data set is subjected to a data range modification operation to obtain an optimal data set, and the optimal data set is stored in the big data storage platform to complete the data optimization operation.

[0010] Optionally, the exception removal process includes a bilateral test rejection process and a unilateral test rejection process, and the unilateral test rejection process includes a minimum value test rejection process and a maximum value test rejection process;

[0011] The calculation method of the two-sided test elimination process is:

[0012]

[0013] Where i is a positive integer, represents the mean value of the original data set, S represents the standard deviation of the original data set, and Y i represents the data in the original data set, and G1 is the value of the two-sided test elimination process.

[0014] The calculation method of the minimum value test elimination process is:

[0015] G2 is the value after the minimum value test is eliminated;

[0016] The calculation method of the maximum value test elimination process is:

[0017]

[0018] G3 is the value after the maximum value test elimination process.

[0019] Optionally, performing grey prediction on the standard data set to obtain a statistical information set includes:

[0020] Counting historical data of the standard data set according to a sampling statistical method to obtain a historical data set;

[0021] Adding the historical data set and the standard data set to obtain a total data set;

[0022] A differential equation is established according to the total data set, and the differential equation is solved to obtain a statistical information set.

[0023] Optionally, the differential equation is:

[0024]

[0025] Among them, X (2) represents the total data set, s is the data number of the total data set, a is the constraint factor of the differential equation, and u is the target value of the differential equation.

[0026] Optionally, calculating the cost value of the statistical information set to obtain a cost data set includes:

[0027] Performing full permutation on the statistical information set to obtain a plurality of full permutation values;

[0028] Calculating cost values ​​of the plurality of full permutation values ​​according to a pre-constructed cost function;

[0029] The permutation data set corresponding to the full permutation value with the smallest cost value is selected to obtain the cost data set.

[0030] In addition, to achieve the above-mentioned object, the present invention further provides an intelligent data optimization device, which includes a memory and a processor, wherein the memory stores an intelligent data optimization program that can be run on the processor, and when the intelligent data optimization program is executed by the processor, the following steps are implemented:

[0031] Receive data optimization instructions input by the user, extract the original data set from the big data storage platform, and perform anomaly removal processing on the original data set to obtain a standard data set;

[0032] Performing grey prediction on the standard data set to obtain a statistical information set;

[0033] Calculating the cost value of the statistical information set to obtain a cost data set;

[0034] Eliminating data in the cost data set that is greater than or equal to a preset cost threshold to obtain an optimized cost data set;

[0035] The optimized cost data set is subjected to a data range modification operation to obtain an optimal data set, and the optimal data set is stored in the big data storage platform to complete the data optimization operation.

[0036] Optionally, the exception removal process includes a bilateral test rejection process and a unilateral test rejection process, and the unilateral test rejection process includes a minimum value test rejection process and a maximum value test rejection process;

[0037] The calculation method of the two-sided test elimination process is:

[0038]

[0039] Where i is a positive integer, represents the mean value of the original data set, S represents the standard deviation of the original data set, and Y i represents the data in the original data set, and G1 is the value of the two-sided test elimination process.

[0040] The calculation method of the minimum value test elimination process is:

[0041] G2 is the value after the minimum value test is eliminated;

[0042] The calculation method of the maximum value test elimination process is:

[0043]

[0044] G3 is the value after the maximum value test elimination process.

[0045] Optionally, performing grey prediction on the standard data set to obtain a statistical information set includes:

[0046] Counting historical data of the standard data set according to a sampling statistical method to obtain a historical data set;

[0047] Adding the historical data set and the standard data set to obtain a total data set;

[0048] A differential equation is established according to the total data set, and the differential equation is solved to obtain a statistical information set.

[0049] Optionally, calculating the cost value of the statistical information set to obtain a cost data set includes:

[0050] Performing full permutation on the statistical information set to obtain a plurality of full permutation values;

[0051] Calculating cost values ​​of the plurality of full permutation values ​​according to a pre-constructed cost function;

[0052] The permutation data set corresponding to the full permutation value with the smallest cost value is selected to obtain the cost data set.

[0053] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which an intelligent data optimization program is stored. The intelligent data optimization program can be executed by one or more processors to implement the steps of the intelligent data optimization method as described above.

[0054] The present invention uses gray prediction to obtain a statistical information set, calculates the cost value of the data set to obtain a cost data set, and then uses data range modification operations to obtain the optimal data set. Because it uses automatic optimization mechanisms such as gray prediction and data range modification operations, it reduces the need for manual intervention and helps developers complete data optimization more quickly. Furthermore, the calculation method of each optimizer is relatively simple, requiring neither very powerful hardware capabilities nor support for unstructured data. Therefore, the intelligent data optimization method, device, and computer-readable storage medium proposed by the present invention can achieve efficient data optimization functions. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 A flow chart of an intelligent data optimization method provided by one embodiment of the present invention;

[0056] Figure 2 A schematic diagram of the internal structure of an intelligent data optimization device provided by one embodiment of the present invention;

[0057] Figure 3 A schematic diagram of a module of an intelligent data optimization program in an intelligent data optimization device provided in one embodiment of the present invention.

[0058] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0059] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0060] The present invention provides an intelligent data optimization method. Figure 1 FIG2 is a flow chart of an intelligent data optimization method according to an embodiment of the present invention. The method can be executed by a device, which can be implemented by software and / or hardware.

[0061] In this embodiment, the intelligent data optimization method includes:

[0062] S1. Receive a data optimization instruction input by a user, extract an original data set from a big data storage platform, and perform anomaly removal processing on the original data set to obtain a standard data set.

[0063] The big data storage platform is a framework or platform that stores and processes a large amount of data, such as MapReduce, Hive, Spark, etc.

[0064] The original data set refers to the data set that needs to be optimized by the present invention, such as the life insurance data input by the user. Since the specifications, data volume, and final use of life insurance data are different, the storage method, data calculation method, etc. are different, so data optimization is required.

[0065] The anomaly removal process is an operation to remove abnormal data such as missing and duplicate data from the original data set to obtain standard data. The anomaly removal process includes bilateral test elimination and unilateral test elimination. The unilateral test elimination includes minimum test elimination and maximum test elimination. Furthermore, the bilateral test elimination data formula is as follows:

[0066]

[0067] Where i is a positive integer, represents the mean value of the original data set, S represents the standard deviation of the original data set, and Y i Represents the data in the original data set.

[0068] The formula for minimum value test elimination is as follows:

[0069]

[0070] The formula for maximum value test elimination is as follows:

[0071]

[0072] S2. Perform grey prediction on the standard data set to obtain a statistical information set.

[0073] Preferably, the purpose of the grey prediction is to evaluate the concurrency and optimal resource allocation of processing the standard data set based on the running status and resource usage of the currently input standard data set and historical data tasks, such as the utilization and running time of CPU, memory, disk and network IO, so as to obtain a statistical data set.

[0074] Furthermore, S2 includes: collecting historical data of the standard data set according to a sampling statistical method to obtain a historical data set, adding the historical data set and the standard data set to obtain a total data set, establishing a differential equation based on the total data set, and solving the differential equation to obtain a statistical information set.

[0075] In detail, the process of establishing the differential equation is as follows:

[0076] X (0) ={X (0) (i), i=1,2,3,…,n}

[0077] Among them, X (0) represents the standard dataset, is the data volume of the standard dataset, and the historical dataset is:

[0078] X (1) ={X (1) (k), k=1,2,3,…,t}

[0079] The total data set is X (2) (k)

[0080]

[0081] For the total data set X (2) (k) Establish the differential equation:

[0082]

[0083] Where s is the data number of the total data set, a is the constraint factor of the differential equation, and u is the target value of the differential equation. The solution to the above differential equation is:

[0084]

[0085] or

[0086] Wherein, k represents the data number of the standard data set.

[0087] S3. Calculate the cost value of the statistical information set to obtain a cost data set.

[0088] The S3 mainly calculates the price (ie, cost) of each execution mode according to the statistical information set, and then selects the optimal execution mode, such as storage mode, data calculation mode, etc.

[0089] Furthermore, the S3 includes: receiving the statistical information set, fully arranging the statistical information set to obtain multiple full arrangement values, calculating the cost values ​​of the multiple full arrangement values ​​according to a pre-constructed cost function, and selecting the arrangement data set corresponding to the full arrangement value with the smallest cost value to obtain the cost data set.

[0090] In detail, the full permutation value y is:

[0091]

[0092] Where n! represents the permutation and combination of the statistical information set, r k ! means traversing the data of the statistical information set and sorting them.

[0093] Preferably, the cost function is:

[0094]

[0095] Wherein, N represents the specific number of the multiple full permutation values, y goal Indicates the target value of the preset full arrangement value, y i represents the multiple full permutation values, L represents the objective function, preferably a gradient descent algorithm can be used, J(y i ) represents the penalty function, and ρ represents the adjustment factor.

[0096] S4. Eliminate data in the cost data set that is greater than or equal to a preset cost threshold to obtain an optimized cost data set.

[0097] Preferably, if the preset cost threshold is 0.8, if the data in the cost data set is greater than or equal to the cost threshold 0.8, the data is discarded; if the data in the cost data set is less than the cost threshold 0.8, the data is retained.

[0098] S5. Perform a data range modification operation on the optimized cost data set to obtain an optimal data set, and store the optimal data set in the big data storage platform to complete the data optimization operation.

[0099] Preferably, the data range modification operation includes methods such as partition pruning, distribution pull-up, distribution push-down, and distribution alignment.

[0100] Furthermore, if the user feels that the data distribution of the optimized cost data set is relatively complex, this solution can perform the partition pruning on the optimized cost data set according to the CART algorithm or other pruning algorithms to make the data distribution simpler; if the data distribution of the optimized cost data set is relatively scattered and the user needs to concentrate the data, the distribution pull-up operation can be performed to map the optimized cost data set within a data interval; if the data distribution of the optimized cost data set is relatively large, the distribution push-down operation can be performed to map the optimized cost data set to a reasonable data interval; if the data of the optimized cost data set is incomplete in structure in terms of data arrangement, the distribution alignment can be performed to make the structure of the data distribution more complete.

[0101] The invention also provides an intelligent data optimization device. Figure 2 FIG. 1 is a schematic diagram of the internal structure of an intelligent data optimization device provided by an embodiment of the present invention.

[0102] In this embodiment, the intelligent data optimization device 1 can be a PC (Personal Computer), or a terminal device such as a smart phone, a tablet computer, a portable computer, or a server. The intelligent data optimization device 1 includes at least a memory 11, a processor 12, a communication bus 13, and a network interface 14.

[0103] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of the intelligent data optimization device 1, such as the hard disk of the intelligent data optimization device 1. In other embodiments, the memory 11 can also be an external storage device of the intelligent data optimization device 1, such as a plug-in hard disk equipped on the intelligent data optimization device 1, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Further, the memory 11 can also include both the internal storage unit of the intelligent data optimization device 1 and an external storage device. The memory 11 can not only be used to store application software and various types of data installed in the intelligent data optimization device 1, such as the code of the intelligent data optimization program 01, etc., but can also be used to temporarily store data that has been output or is to be output.

[0104] In some embodiments, the processor 12 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip, configured to execute program codes stored in the memory 11 or process data, such as executing an intelligent data optimization program 01.

[0105] The communication bus 13 is used to realize the connection and communication between these components.

[0106] The network interface 14 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface), and is generally used to establish a communication connection between the device 1 and other electronic devices.

[0107] Optionally, the device 1 may further include a user interface, which may include a display and an input unit such as a keyboard. The optional user interface may also include a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touch device. The display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the intelligent data optimization device 1 and to display a visual user interface.

[0108] Figure 2 Only the intelligent data optimization device 1 having components 11-14 and the intelligent data optimization program 01 is shown. It can be understood by those skilled in the art that Figure 1 The structure shown does not constitute a limitation on the intelligent data optimization device 1, and may include fewer or more components than shown in the figure, or a combination of certain components, or a different arrangement of components.

[0109] exist Figure 2 In the embodiment of the device 1 shown, the memory 11 stores an intelligent data optimization program 01; when the processor 12 executes the intelligent data optimization program 01 stored in the memory 11, the following steps are implemented:

[0110] Step 1: Receive a data optimization instruction input by a user, extract the original data set from the big data storage platform, and perform anomaly removal processing on the original data set to obtain a standard data set.

[0111] The big data storage platform is a framework or platform that stores and processes a large amount of data, such as MapReduce, Hive, Spark, etc.

[0112] The original data set refers to the data set that needs to be optimized by the present invention, such as the life insurance data input by the user. Since the specifications, data volume, and final use of life insurance data are different, the storage method, data calculation method, etc. are different, so data optimization is required.

[0113] The anomaly removal process is an operation to remove abnormal data such as missing and duplicate data from the original data set to obtain standard data. The anomaly removal process includes bilateral test elimination and unilateral test elimination. The unilateral test elimination includes minimum test elimination and maximum test elimination. Furthermore, the bilateral test elimination data formula is as follows:

[0114]

[0115] Where i is a positive integer, represents the mean value of the original data set, S represents the standard deviation of the original data set, and Y i Represents the data in the original data set.

[0116] The formula for minimum value test elimination is as follows:

[0117]

[0118] The formula for maximum value test elimination is as follows:

[0119]

[0120] Step 2: Perform grey prediction on the standard data set to obtain a statistical information set.

[0121] Preferably, the purpose of the grey prediction is to evaluate the concurrency and optimal resource allocation of processing the standard data set based on the running status and resource usage of the currently input standard data set and historical data tasks, such as the utilization and running time of CPU, memory, disk and network IO, so as to obtain a statistical data set.

[0122] Furthermore, the step 2 includes: performing statistical analysis of historical data on the standard data set according to a sampling statistical method to obtain a historical data set, adding the historical data set and the standard data set to obtain a total data set, establishing a differential equation based on the total data set, and solving the differential equation to obtain a statistical information set.

[0123] In detail, the process of establishing the differential equation is as follows:

[0124] X (0) ={X (0) (i), i=1,2,3,…,n}

[0125] Among them, X (0) represents the standard dataset, is the data volume of the standard dataset, and the historical dataset is:

[0126] X (1) ={X (1) (k), k=1,2,3,…,t}

[0127] The total data set is X (2) (k)

[0128]

[0129] For the total data set X (2) (k) Establish the differential equation:

[0130]

[0131] Where s is the data number of the total data set, a is the constraint factor of the differential equation, and u is the target value of the differential equation. The solution to the above differential equation is:

[0132]

[0133] or

[0134] Wherein, k represents the data number of the standard data set.

[0135] Step 3: Calculate the cost value of the statistical information set to obtain a cost data set.

[0136] The step three is mainly to calculate the price (ie, cost) of each execution mode according to the statistical information set, and then select the optimal execution mode, such as storage mode, data calculation mode, etc.

[0137] Furthermore, step three includes: receiving the statistical information set, fully arranging the statistical information set to obtain multiple full arrangement values, calculating the cost values ​​of the multiple full arrangement values ​​according to a pre-constructed cost function, and selecting the arrangement data set corresponding to the full arrangement value with the smallest cost value to obtain the cost data set.

[0138] In detail, the full permutation value y is:

[0139]

[0140] Where n! represents the permutation and combination of the statistical information set, r k ! means traversing the data of the statistical information set and sorting them.

[0141] Preferably, the cost function is:

[0142]

[0143] Wherein, N represents the specific number of the multiple full permutation values, y goal Indicates the target value of the preset full arrangement value, y i represents the multiple full permutation values, L represents the objective function, preferably a gradient descent algorithm can be used, J(y i ) represents the penalty function, and ρ represents the adjustment factor.

[0144] Step 4: Eliminate data in the cost data set that is greater than or equal to a preset cost threshold to obtain an optimized cost data set.

[0145] Preferably, if the preset cost threshold is 0.8, if the data in the cost data set is greater than or equal to the cost threshold 0.8, the data is discarded; if the data in the cost data set is less than the cost threshold 0.8, the data is retained.

[0146] Step 5: Perform a data range modification operation on the optimized cost data set to obtain an optimal data set, and store the optimal data set in the big data storage platform to complete the data optimization operation.

[0147] Preferably, the data range modification operation includes methods such as partition pruning, distribution pull-up, distribution push-down, and distribution alignment.

[0148] Furthermore, if the user feels that the data distribution of the optimized cost data set is relatively complex, this solution can perform the partition pruning on the optimized cost data set according to the CART algorithm or other pruning algorithms to make the data distribution simpler; if the data distribution of the optimized cost data set is relatively scattered and the user needs to concentrate the data, the distribution pull-up operation can be performed to map the optimized cost data set within a data interval; if the data distribution of the optimized cost data set is relatively large, the distribution push-down operation can be performed to map the optimized cost data set to a reasonable data interval; if the data of the optimized cost data set is incomplete in structure in terms of data arrangement, the distribution alignment can be performed to make the structure of the data distribution more complete.

[0149] Optionally, in other embodiments, the intelligent data optimization program can also be divided into one or more modules, one or more modules are stored in the memory 11, and executed by one or more processors (processor 12 in this embodiment) to complete the present invention. The module referred to in the present invention refers to a series of computer program instruction segments that can perform specific functions, which are used to describe the execution process of the intelligent data optimization program in the intelligent data optimization device.

[0150] For example, refer to Figure 3FIG. 1 is a schematic diagram of program modules of an intelligent data optimization program in an embodiment of an intelligent data optimization device according to the present invention. In this embodiment, the intelligent data optimization program can be divided into a data receiving and processing module 10, a gray prediction module 20, a cost optimization module 30, and a data optimization module 40. For example:

[0151] The data receiving and processing module 10 is used to receive a data optimization instruction input by a user, extract an original data set from a big data storage platform, and perform anomaly removal processing on the original data set to obtain a standard data set.

[0152] The grey prediction module 20 is used to perform grey prediction on the standard data set to obtain a statistical information set.

[0153] The cost optimization 30 is used to calculate the cost value of the statistical information set to obtain a cost data set, and eliminate data in the cost data set that is greater than or equal to a preset cost threshold to obtain an optimized cost data set.

[0154] The data optimization module 40 is used to perform a data range modification operation on the optimized cost data set to obtain an optimal data set, store the optimal data set in the big data storage platform, and complete the data optimization operation.

[0155] The functions or operation steps implemented when the above-mentioned program modules such as the data receiving and processing module 10, the gray prediction module 20, the cost optimization module 30, and the data optimization module 40 are executed are substantially the same as those in the above-mentioned embodiment and will not be repeated here.

[0156] In addition, an embodiment of the present invention further provides a computer-readable storage medium having an intelligent data optimization program stored thereon. The intelligent data optimization program can be executed by one or more processors to implement the following operations:

[0157] The data optimization instruction input by the user is received, the original data set is extracted from the big data storage platform, and the original data set is processed to remove anomalies to obtain a standard data set.

[0158] The standard data set is subjected to grey prediction to obtain a statistical information set.

[0159] The cost value of the statistical information set is calculated to obtain a cost data set, and data greater than or equal to a preset cost threshold in the cost data set is eliminated to obtain an optimized cost data set.

[0160] The optimized cost data set is subjected to a data range modification operation to obtain an optimal data set, and the optimal data set is stored in the big data storage platform to complete the data optimization operation.

[0161] It should be noted that the serial numbers of the above-mentioned embodiments of the present invention are for descriptive purposes only and do not represent the advantages or disadvantages of the embodiments. In addition, the terms "including", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "including a ..." does not exclude the presence of other identical elements in the process, device, article or method comprising the element.

[0162] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0163] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. An intelligent data optimization method, characterized in that: The method comprises: Receive data optimization instructions input by the user, extract the original insurance data set from the big data storage platform, remove missing information and duplicate information in the original insurance data set, and obtain a standard insurance data set; Performing grey prediction based on the operation status and resource usage of the standard insurance data set and the historical data set to obtain a statistical information set of the standard insurance data set, wherein the operation status includes the operation time and the resource usage includes the resource usage rate; Calculating the cost value of the statistical information set to obtain a cost data set, including: performing a full permutation on the statistical information set to obtain a plurality of full permutation values, calculating the cost values ​​of the plurality of full permutation values ​​according to a pre-constructed cost function, and selecting the permutation data set corresponding to the full permutation value having the smallest cost value to obtain the cost data set; Eliminating data in the cost data set that is greater than or equal to a preset cost threshold to obtain an optimized cost data set; The optimized cost data set is subjected to a data range modification operation to obtain an optimal data set, and the optimal data set is stored in the big data storage platform to complete the data optimization operation.

2. The intelligent data optimization method according to claim 1, characterized in that: The rejection process includes a bilateral test rejection process and a unilateral test rejection process, and the unilateral test rejection process includes a minimum value test rejection process and a maximum value test rejection process; The calculation method of the two-sided test elimination process is: in, i is a positive integer, represents the average value of the original insurance data set, represents the standard deviation of the original insurance data set, represents the data in the original insurance dataset, a value for culling processing for the two-sided test; The calculation method of the minimum value test elimination process is: in, The value after the minimum value test is eliminated; The calculation method of the maximum value test elimination process is: in, The value after culling is tested for the maximum value.

3. The intelligent data optimization method according to claim 1, characterized in that: Gray prediction is performed based on the operation status and resource usage of the standard insurance dataset and the historical dataset to obtain a statistical information set of the standard insurance dataset, including: Counting historical data of the standard insurance data set according to a sampling statistical method to obtain a historical data set; Adding the historical data set and the standard insurance data set to obtain a total data set; A differential equation is established according to the total data set, and the statistical information set is obtained by solving the differential equation.

4. The intelligent data optimization method according to claim 3, characterized in that: The differential equation is: in, represents the total data set, is the data number of the total data set, is the constraint factor of the differential equation, is the target value of the differential equation.

5. An intelligent data optimization device, characterized in that: The device includes a memory and a processor. The memory stores an intelligent data optimization program that can be run on the processor. When the intelligent data optimization program is executed by the processor, the following steps are implemented: Receive data optimization instructions input by the user, extract the original insurance data set from the big data storage platform, remove missing information and duplicate information in the original insurance data set, and obtain a standard insurance data set; Performing grey prediction based on the operation status and resource usage of the standard insurance data set and the historical data set to obtain a statistical information set of the standard insurance data set, wherein the operation status includes the operation time and the resource usage includes the resource usage rate; Calculating the cost value of the statistical information set to obtain a cost data set, including: performing a full permutation on the statistical information set to obtain a plurality of full permutation values, calculating the cost values ​​of the plurality of full permutation values ​​according to a pre-constructed cost function, and selecting the permutation data set corresponding to the full permutation value having the smallest cost value to obtain the cost data set; Eliminating data in the cost data set that is greater than or equal to a preset cost threshold to obtain an optimized cost data set; The optimized cost data set is subjected to a data range modification operation to obtain an optimal data set, and the optimal data set is stored in the big data storage platform to complete the data optimization operation.

6. The intelligent data optimization device according to claim 5, characterized in that: The rejection process includes a bilateral test rejection process and a unilateral test rejection process, and the unilateral test rejection process includes a minimum value test rejection process and a maximum value test rejection process; The calculation method of the two-sided test elimination process is: in, i is a positive integer, represents the average value of the original insurance data set, represents the standard deviation of the original insurance data set, represents the data in the original insurance dataset, a value for culling processing for the two-sided test; The calculation method of the minimum value test elimination process is: in, The value after the minimum value test is eliminated; The calculation method of the maximum value test elimination process is: in, The value after culling is tested for the maximum value.

7. The intelligent data optimization device according to claim 5, characterized in that: The gray prediction is performed based on the operation status and resource usage of the standard insurance data set and the historical data set to obtain the statistical information set of the standard insurance data set, including: Counting historical data of the standard insurance data set according to a sampling statistical method to obtain a historical data set; Adding the historical data set and the standard insurance data set to obtain a total data set; A differential equation is established according to the total data set, and the differential equation is solved to obtain a statistical information set.

8. A computer-readable storage medium, characterized in that An intelligent data optimization program is stored on the computer-readable storage medium, and the intelligent data optimization program can be executed by one or more processors to implement the steps of the intelligent data optimization method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Electrode arrangement optimization method and device for ultra-high density resistivity method

    CN106570227A

  • Virtual machine on-line migration optimizing method capable of sensing compound application characteristics and network bandwidth

    CN106775949A