Big data cleaning method, system and device based on AI training and storage medium
By using AI to analyze data sets in big data cleaning, filtering out abnormal duplicate data and redundant data, the problem that AI cannot accurately judge duplicate items during cleaning is solved, and data quality improvement and data loss are achieved.
Patent Information
- Application Number
- CN202411960204.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When using AI to clean big data, AI cannot accurately determine whether duplicates are removed, resulting in the backup data being deleted, resulting in the problem of missing data in the cleaned database.
By performing data analysis on the data set to be cleaned based on AI, abnormal duplicate data and abnormal regular data are obtained, and abnormal duplicate data is filtered using redundant quality screening method, redundant data is eliminated, and only abnormal duplicate data is trained to ensure that AI can accurately determine whether to remove duplicate items.
It effectively avoids the problem of data missing, ensures the data quality after data cleaning, and trains AI to accurately judge duplicates, avoiding the error deletion of backup data.
Smart Images

Figure CN119918694A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data cleaning technology, and specifically to a big data cleaning method, system, device and storage medium based on AI training. Background Art
[0002] Big data cleaning refers to the process of screening, converting and correcting data in big data analysis. Since big data usually comes from multiple different sources, there may be errors, duplications, missing and other problems, so data cleaning is needed to fix these problems to ensure the accuracy and reliability of the analysis results.
[0003] In the existing data cleaning methods, AI analysis is usually used to automate big data cleaning, and the data quality is identified and sorted, so as to assist users in evaluating the quality of the cleaned data. In the process of using AI for data cleaning, duplicate items in the data are usually removed to achieve redundancy. Although this processing method can optimize storage space, AI cannot accurately determine whether to remove duplicate items during training, which may result in the deletion of backup data, resulting in data missing in the cleaned database. For example, in the Chinese patent with publication number CN118210791A, a big data cleaning method and a big data acquisition system based on AI training are disclosed. This solution is to automatically Perform data quality analysis, identify and report data quality issues, and help users better understand and evaluate data quality through the application of AI models. Although the above method can automate big data cleaning based on AI training and assist users in quality assessment, the data for quality assessment are all cleaned and high-quality data. For data in the backup folder during data cleaning, AI cannot accurately determine whether to remove duplicate items during the training process, resulting in duplicate data in the backup folder being identified by AI as low-quality redundant data, which is removed and cannot be submitted to the user, resulting in data missing but cannot be discovered in time. In view of this, it is necessary to improve the existing big data cleaning based on AI training. Summary of the invention
[0004] The present invention aims to solve one of the technical problems in the prior art to at least a certain extent. By proposing a big data cleaning method, system, device and storage medium based on AI training, it is used to solve the problem in the prior art that in the process of using AI for data cleaning, duplicate items in the data are usually removed to achieve redundancy removal. Although this processing method can optimize the storage space, AI cannot accurately determine whether to remove duplicate items during the training process, which may result in the deletion of backup data, resulting in the problem of missing data in the cleaned database.
[0005] To achieve the above objectives, in a first aspect, the present application provides a big data cleaning method based on AI training, comprising the following steps:
[0006] Perform data analysis on the data set to be cleaned based on AI, and obtain abnormal duplicate data and abnormal regular data in the data set based on data analysis;
[0007] Carry out routine cleaning of abnormal routine data;
[0008] The abnormal duplicate data is screened using the redundant quality screening method, and the redundant data in the abnormal duplicate data is eliminated based on the screening results. The AI is trained using the data that has been eliminated and the data that has not been eliminated from the abnormal duplicate data.
[0009] Furthermore, data analysis is performed on the data set to be cleaned based on AI, and abnormal duplicate data and abnormal regular data in the data set are obtained based on the data analysis, including:
[0010] Obtain the enterprise type of the data set to be analyzed, which is recorded as the enterprise to be analyzed; use AI to obtain all data types related to the enterprise to be analyzed based on big data;
[0011] Use data anomaly analysis methods to analyze all data types, and obtain standard data graphs corresponding to each data type based on the analysis results;
[0012] Based on the standard data graph corresponding to each data type, abnormal duplicate data and abnormal regular data in the data set to be cleaned are obtained.
[0013] Furthermore, the data anomaly analysis method includes:
[0014] For any data type α, all data in the data type α are recorded as data to be extracted DS1 to data to be extracted DS m ; For any data DS to be extracted m1 , the data to be extracted DS m1 The amount of data manually filled in is recorded as the filled amount; the data to be extracted DS m1 The data manually filled in are recorded as filled data TR1 to filled data TR n , will fill in data TR1 to fill in data TR n The corresponding character length is recorded as character length ZC1 to character length ZC n , the interval consisting of the minimum and maximum values of all character lengths ZC is recorded as the filled character interval, where m1 is a positive integer less than or equal to m and greater than or equal to 1;
[0015] Obtain the filling quantity and filling character interval corresponding to all the data to be extracted DS; establish a plane rectangular coordinate system, recorded as the standard sub-one coordinate system, wherein the units of the X-axis and the Y-axis of the standard sub-one coordinate system are both quantity; perform punctuation in the standard sub-one coordinate system based on the number of times each filling quantity appears in all the data to be extracted DS and the numerical value corresponding to the filling quantity, wherein, for any marked point X0, the horizontal coordinate of X0 is a filling quantity corresponding to the data to be extracted DS, and the vertical coordinate of X0 is the number of times the filling quantity appears in the filling quantities of all the data to be extracted DS; fit all the punctuation points in the standard sub-one coordinate system into a curve, recorded as the standard sub-one curve, and record the curve after the standard sub-one curve is mirrored based on the X-axis as the mirror sub-one curve;
[0016] A plane rectangular coordinate system is established, recorded as the standard sub-two coordinate system, wherein the units of the X-axis and the Y-axis of the standard sub-two coordinate system are both numbers; for any positive integer C in the filled-in character interval corresponding to all data DS to be extracted, the number of filled-in character intervals containing the positive integer C is recorded as the existence value of the positive integer C, and a bar graph is established at X = positive integer C in the standard sub-two coordinate system, the height of the bar graph is the existence value of the positive integer C, and the width of the bar graph is half of the distance between adjacent coordinate points in the X axis; the existence values of all numbers in the filled-in character intervals corresponding to all DS to be extracted are obtained, and the corresponding bar graphs are drawn in the standard sub-two coordinate system; a mirror sub-one curve is drawn in the fourth quadrant of the standard sub-two coordinate system, and the standard sub-two coordinate system at this time is recorded as the standard data graph of data type α.
[0017] Further, obtaining abnormal duplicate data and abnormal regular data in the data set to be cleaned based on the standard data graph corresponding to each data type includes:
[0018] Use AI to obtain all data types in the data set to be cleaned based on the acquisition method when obtaining the data type of the enterprise to be analyzed; for any data type β in the data set to be cleaned, obtain the standard sub-one coordinate system and the standard sub-two coordinate system of the data type β based on the data anomaly method, and record the standard sub-two coordinate system after drawing the mirror sub-one curve in the fourth quadrant of the standard sub-two coordinate system as the data graph to be cleaned of the data type β; overlap the data graph to be cleaned of the data type β with the standard data graph in the same standard sub-two coordinate system, record the area where the two curves do not overlap in the first quadrant of the overlapped standard sub-two coordinate system as the filled-in differential area, record the point in the filled-in differential area whose horizontal coordinate is equal to the filled-in amount as the differential point, and mark the horizontal coordinates of all the differential points as abnormal filled-in amounts; record the data with abnormal filled-in amounts in the data type β as quantitative abnormal data;
[0019] The bar graphs with the same bottom sides and not overlapping in the fourth quadrant of the standard sub-two coordinate system after the overlap are recorded as differential bar graphs, the horizontal coordinates corresponding to the differential bar graphs are marked as differential filling values, and the heights of the non-overlapping areas in the differential bar graphs are recorded as the differential quantity; all data in the data type β are analyzed using the differential comparison method, and the differential comparison method includes: when the positive integers in the filling character interval of any data in the data type β are all differential filling values and the differential data quantities corresponding to all differential filling values are not 0, the data are recorded as abnormal filling data, and the differential filling value of the positive integer in the filling character interval corresponding to the data is reduced by 1; when one abnormal filling data is obtained, the differential comparison method is used again to analyze all data in the data type β until the abnormal filling data cannot be extracted;
[0020] When any data in the data type β is recorded as both quantity abnormal data and filling abnormal data, the data is recorded as abnormal duplicate data; the data in the data type β that is only recorded as quantity abnormal data or filling abnormal data is recorded as abnormal regular data.
[0021] Furthermore, the routine cleaning process for abnormal routine data includes:
[0022] Analyze abnormal and regular data based on AI. When there are gaps in the abnormal and regular data that have not been filled manually, use AI to fill in the gaps based on interpolation.
[0023] AI is used to analyze the values of manually filled data in data that does not contain gaps. When the value of the data is not within the standard range of the data retrieved by AI, the value of the data is recorded as an abnormal value, and the data with the abnormal value is deleted based on the discard method.
[0024] Furthermore, the abnormal duplicate data is screened using a redundant quality screening method, and redundant data in the abnormal duplicate data is removed based on the screening result. The AI is trained using the data removed from the abnormal duplicate data and the data not removed, including:
[0025] The abnormal duplicate data are screened using the redundant quality screening method, and the redundant data in the abnormal duplicate data are removed based on the screening results;
[0026] The AI is trained using the data that has not been eliminated from the abnormal duplicate data. The training method is: use AI to record the format corresponding to the data that has not been eliminated, and record it as a repeatable format. When AI performs big data cleaning on the data of the enterprise to be analyzed and the format corresponding to any abnormal duplicate data in the cleaning process is equal to the repeatable format, the abnormal duplicate data will be eliminated from all abnormal duplicate data.
[0027] Furthermore, the redundant quality screening method includes:
[0028] Set a data duplication standard, which is: when the data not filled in manually in two data are exactly the same and the values of the data in the manually filled position are the same, the two data are recorded as data duplication, where the exact same means the characters are the same and the positions of the characters in the data are the same;
[0029] The data related to the enterprise to be analyzed acquired by AI and having the same data type as the data set to be cleaned are recorded as data to be judged, and the arrangement of the data types in the data to be judged is consistent with the arrangement of the data types in the data set to be cleaned;
[0030] For any abnormal repetitive data, an abnormal repetitive algorithm is used to obtain the repetitive characteristic value of the abnormal repetitive data. The abnormal repetitive algorithm includes: Among them, F is the repeated feature value, p is the number of data types in the data set to be cleaned, and Q i is the number of data that meets the data duplication standard in the i-th data type of the data set to be cleaned and the abnormal duplicate data, U i Q i The ratio of the total number of data in the i-th data type of the data set to be cleaned, W i is the number of data in the i-th data type of the data to be judged that meets the data duplication standard and the abnormal duplication data, U i W i The ratio of the total number of data in the i-th data type of the data to be judged;
[0031] When the repetition characteristic value of abnormal duplicate data is greater than 1, the data in the data set to be analyzed that meets the data duplication standard with the abnormal duplicate data will be eliminated; when the repetition characteristic value of abnormal duplicate data is less than or equal to 0, the data in the data set to be analyzed that meets the data duplication standard with the abnormal duplicate data will be retained.
[0032] On the second aspect, the present application also provides a big data cleaning system based on AI training, including an abnormal screening module, a regular cleaning module, and a repeated cleaning module;
[0033] The anomaly screening module is used to perform data analysis on the data set to be cleaned based on AI, and obtain abnormal duplicate data and abnormal regular data in the data set based on data analysis;
[0034] The routine cleaning module is used to perform routine cleaning on abnormal routine data;
[0035] The duplicate cleaning module is used to screen abnormal duplicate data using the redundant quality screening method, and to eliminate redundant data in the abnormal duplicate data based on the screening results, and to train AI using the data eliminated and the data not eliminated from the abnormal duplicate data.
[0036] In a third aspect, the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the steps in the above method are performed.
[0037] In a fourth aspect, the present application provides a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps in the above method are performed.
[0038] Beneficial effects of the present invention: The present invention performs data analysis on the data set to be cleaned based on AI, obtains abnormal duplicate data and abnormal regular data in the data set based on data analysis, and performs conventional cleaning processing on the abnormal regular data. The advantage of this is that by dividing the data in the data set into abnormal duplicate data and abnormal regular data based on data analysis, the data that appears abnormal during the data cleaning process can be divided based on whether it is duplicate data, so that only the abnormal duplicate data is trained during AI training, ensuring that after training, the AI can accurately determine whether to remove duplicate items; through conventional cleaning processing, the data that is not duplicate data in the abnormal data can be effectively cleaned, so as to ensure that while the abnormal duplicate data is screened and judged, the remaining abnormal data is filled with missing values, processed with abnormal values, and format conversion and other conventional data cleaning operations are performed, so as to ensure that the data set after data cleaning is all high-quality data;
[0039] The present invention also uses a redundant quality screening method to screen the abnormal duplicate data, and based on the screening results, eliminates redundant data in the abnormal duplicate data, and uses the data eliminated and the data not eliminated in the abnormal duplicate data to train AI. The advantage of this is that by using the redundant quality screening method to screen the abnormal duplicate data, the duplicate data in the abnormal duplicate data, such as those in the backup file, can be retained, and the meaningless duplicate data can be deleted, thereby ensuring that the backup data will not be deleted during actual data cleaning, resulting in data missing problems in the cleaned database; at the same time, based on the data eliminated and the data not eliminated in the abnormal duplicate data, the AI is trained, so that the AI can accurately determine whether the abnormal duplicate data should be deleted after training. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a functional block diagram of the system of the present invention;
[0041] Figure 2 It is a schematic diagram of obtaining the standard sub-curve and the mirror sub-curve of the present invention;
[0042] Figure 3is a schematic diagram of the standard sub-two coordinate system of the present invention;
[0043] Figure 4 is a schematic diagram of a standard data graph of the present invention;
[0044] Figure 5 is a flow chart of the steps of the method of the present invention;
[0045] Figure 6 It is a structural block diagram of the electronic device of the present invention. DETAILED DESCRIPTION
[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0047] Example 1, please refer to Figure 1 As shown, the present application provides a big data cleaning system based on AI training, including an abnormal screening module, a regular cleaning module and a repeated cleaning module;
[0048] The abnormal screening module is used to perform data analysis on the data set to be cleaned based on AI, and obtain abnormal duplicate data and abnormal regular data in the data set based on data analysis. The abnormal screening module includes an abnormal data screening unit, and the abnormal data screening unit is configured with an abnormal data screening strategy. The abnormal data screening strategy includes:
[0049] Obtain the enterprise type of the data set to be analyzed, which is recorded as the enterprise to be analyzed; use AI to obtain all data types related to the enterprise to be analyzed based on big data;
[0050] In the specific implementation process, for example, in a data processing process, the data in the data set to be analyzed is logistics management data, then the enterprise type of the data set to be analyzed, that is, the enterprise to be analyzed is logistics management; and the data types corresponding to logistics management include order data, inventory data, transportation data, procurement logistics data, production logistics data, sales logistics data and customer data;
[0051] The data anomaly analysis method is used to analyze all data types, and based on the analysis results, a standard data graph corresponding to each data type is obtained. The data anomaly analysis method includes: for any data type α, all data in the data type α are recorded as data to be extracted DS1 to data to be extracted DS m ; For any data DS to be extracted m1 , the data to be extracted DS m1The amount of data manually filled in is recorded as the filled amount; the data to be extracted DS m1 The data manually filled in are recorded as filled data TR1 to filled data TR n , will fill in data TR1 to fill in data TR n The corresponding character length is recorded as character length ZC1 to character length ZC n , the interval consisting of the minimum and maximum values of all character lengths ZC is recorded as the filled character interval, where m1 is a positive integer less than or equal to m and greater than or equal to 1;
[0052] In the specific implementation process, the data type α is the customer data in logistics management. All the data in the customer data are basic information, contact information, purchase information, transaction records and customer feedback. The value of m is 5, and the data to be extracted DS1 to the data to be extracted DS m They are basic information, contact information, purchase information, transaction records and customer feedback respectively; for any data DS2 to be extracted, i.e. contact information, the data manually filled in the contact information is telephone number, email address and address, then the value of the filled amount, i.e. n, is 3, and the filled data TR1 to filled data TR3 are telephone number, email address and address respectively; the character lengths corresponding to telephone number, email address and address are 11, 10 and 20 respectively, then the filled character range is [10,20]; by obtaining the filled character range, it is helpful to provide a screening basis for the subsequent screening of abnormal data in the data;
[0053] Obtain the filling quantity and filling character interval corresponding to all the data DS to be extracted; establish a plane rectangular coordinate system, recorded as the standard sub-one coordinate system, wherein the units of the X-axis and the Y-axis of the standard sub-one coordinate system are both quantity; perform punctuation in the standard sub-one coordinate system based on the number of occurrences of each filling quantity in all the data DS to be extracted and the numerical value corresponding to the filling quantity, wherein, for any marked point X0, the horizontal coordinate of X0 is a filling quantity corresponding to the data DS to be extracted, and the vertical coordinate of X0 is the number of occurrences of the filling quantity in the filling quantities of all the data DS to be extracted; fit all the punctuation points in the standard sub-one coordinate system into a curve, recorded as the standard sub-one curve, and record the curve after the standard sub-one curve is mirrored based on the X-axis as the mirror sub-one curve; in a data processing process, all the data DS to be extracted obtained are basic information, contact information, purchase information, transaction records and customer feedback, and the filling quantity corresponding to each data DS to be extracted is 10, 3, 3, 2 and 3, then the standard sub-one coordinate system established by analysis can be referred to Figure 2As shown, TT1 to TT6 are scales 2, 4, 6, 8, 10 and 12 respectively, SS1 to SS4 are scales 1, 2, 3 and 4 respectively, points XX1 to XX3 are points marked based on the filling amount of all data to be extracted DS, curve QQ1 is a standard sub-curve, and QQ2 is a mirror sub-curve;
[0054] A plane rectangular coordinate system is established, recorded as the standard sub-two coordinate system, wherein the units of the X-axis and the Y-axis of the standard sub-two coordinate system are both numbers; for any positive integer C in the filled-in character interval corresponding to all data DS to be extracted, the number of filled-in character intervals containing the positive integer C is recorded as the existence value of the positive integer C, and a bar graph is established at X = positive integer C in the standard sub-two coordinate system, the height of the bar graph is the existence value of the positive integer C, and the width of the bar graph is half of the distance between adjacent coordinate points in the X axis; the existence values of all numbers in the filled-in character intervals corresponding to all DS to be extracted are obtained, and the corresponding bar graphs are drawn in the standard sub-two coordinate system; a mirror sub-one curve is drawn in the fourth quadrant of the standard sub-two coordinate system, and the standard sub-two coordinate system at this time is recorded as the standard data graph of data type α.
[0055] In a data processing process, all the data to be extracted DS are basic information, contact information, purchase information, transaction records and customer feedback, and the character intervals corresponding to each data to be extracted DS are [2,5], [4,10], [1,3], [5,10] and [10,20]. The standard sub-two coordinate system established by analysis can be found in Figure 3 As shown, where C1 to C10 are positive integers 2, 4, 6, 8, 10, 12, 14, 16, 18 and 20 respectively, then the value corresponding to half of the distance between adjacent coordinate points is 1, then all the bar graphs obtained can be referred to Figure 3 As shown, Figure 2 The standard data diagram obtained by plotting the mirror image sub-1 curve in the fourth quadrant of the standard sub-2 coordinate system can be found in Figure 4 As shown; by obtaining the standard data graph, the standard features of each data type that can be used as a reference when performing abnormal analysis can be obtained, so that the abnormal data screened out during analysis is more accurate;
[0056] Based on the standard data graph corresponding to each data type, obtain the abnormal duplicate data and abnormal regular data in the data set to be cleaned. Specifically: use AI to obtain all data types in the data set to be cleaned based on the acquisition method when obtaining the data type of the enterprise to be analyzed; for any data type β in the data set to be cleaned, obtain the standard sub-one coordinate system and standard sub-two coordinate system of the data type β based on the data anomaly method, and record the standard sub-two coordinate system after drawing the mirror sub-one curve in the fourth quadrant of the standard sub-two coordinate system as the data graph to be cleaned of the data type β; overlap the data graph to be cleaned of the data type β with the standard data graph in the same standard sub-two coordinate system, record the area where the two curves do not overlap in the first quadrant of the overlapped standard sub-two coordinate system as the filled-in differential area, record the point in the filled-in differential area whose horizontal coordinate is equal to the filled-in amount as the differential point, and mark the horizontal coordinates of all differential points as abnormal filled-in amounts; record the data with abnormal filled-in amounts in the data type β as quantitative abnormal data;
[0057] In the specific implementation process, when the two curves in the first quadrant of the standard sub-two coordinate system do not overlap, it means that there is data with inaccurate character length in the data type β. When there is no point equal to the filled-in amount in the filled-in differential area, it means that the error may be caused by errors in other filled-in amounts. Therefore, only the point whose abscissa is equal to the filled-in amount in the filled-in differential area is recorded as a differential point.
[0058] The bar graphs with the same bottom sides and no overlap in the fourth quadrant of the standard sub-two coordinate system after overlap are recorded as differential bar graphs, the horizontal coordinates corresponding to the differential bar graphs are marked as differential fill-in values, and the heights of the non-overlapping areas in the differential bar graphs are recorded as differential quantities; all data in the data type β are analyzed using the differential comparison method, and the differential comparison method includes: when the positive integers in the fill-in character interval of any data in the data type β are all differential fill-in values and the differential data quantities corresponding to all differential fill-in values are not 0, the data are recorded as abnormal fill-in data, and the differential fill-in value of the positive integer in the fill-in character interval corresponding to the data is reduced by 1; when an abnormal fill-in data is obtained, the differential comparison method is used again to analyze all data in the data type β until the fill-in data cannot be extracted. Abnormal data; for example, the differential filling values in the data type β are 2, 3, 4, 5, 9 and 10, and the corresponding differential quantities are 2, 2, 3, 4 and 4, respectively. The filling character intervals in the data type β are [2, 5], [2, 6], [10, 15] and [9, 10]. Then, through analysis, the abnormal filling data are the data corresponding to [2, 5] and [9, 10]; in this embodiment, only when the positive integers in the filling character interval are all differential filling values, it indicates that the data corresponding to the filling character interval is abnormal. When there is a positive integer in the filling character interval that is not a differential filling value, it indicates that the differential histogram may be caused by abnormal data in other filling character intervals and has nothing to do with the filling character interval;
[0059] When any data in the data type β is recorded as both quantity abnormal data and filling abnormal data, the data is recorded as abnormal duplicate data; the data in the data type β that is only recorded as quantity abnormal data or filling abnormal data is recorded as abnormal regular data.
[0060] The conventional cleaning module is used to perform conventional cleaning processing on abnormal conventional data; the conventional cleaning module includes a conventional cleaning unit, and the conventional cleaning unit includes: analyzing abnormal conventional data based on AI, and when there are gaps in the abnormal conventional data that are not filled in manually, using AI to fill the gaps based on interpolation;
[0061] AI is used to analyze the values of manually filled data in data that does not contain gaps. When the value of the data is not within the standard range of the data retrieved by AI, the value of the data is recorded as an abnormal value, and the data with the abnormal value is deleted based on the discard method.
[0062] The conventional cleaning unit in this embodiment processes abnormal data that is not duplicate data by using existing data cleaning methods, wherein the interpolation method is to reasonably estimate the missing values by using the information of known data points, thereby achieving data integrity and continuity. For example, if the known data has a cargo retention quantity between 10 and 20, then for the data whose missing value is the cargo retention quantity, data between 10 and 20 can be filled in to achieve data integrity and continuity; the discard method is to delete data with invalid values or missing values after checking data consistency. For example, if the known data has a cargo retention quantity between 10 and 20, then for the data with a cargo retention quantity of 54, 54 can be recorded as an abnormal value, and the data corresponding to the cargo retention quantity where 54 is located is deleted based on the discard method to ensure the accuracy and validity of the data;
[0063] The duplicate cleaning module is used to screen abnormal duplicate data using the redundant quality screening method, and to eliminate redundant data in the abnormal duplicate data based on the screening results, and to train AI using the data eliminated and the data not eliminated from the abnormal duplicate data.
[0064] The duplicate cleaning module includes a duplicate data screening unit, which is configured with a duplicate data screening strategy. The duplicate data screening strategy includes: screening abnormal duplicate data using a redundant quality screening method, and removing redundant data in the abnormal duplicate data based on the screening result;
[0065] The redundant quality screening method includes: setting a data duplication standard, the data duplication standard is: when the data not filled in manually in two data are exactly the same and the values of the data in the manually filled position are the same, the two data are recorded as data duplication, wherein the exact same means the characters are the same and the positions of the characters in the data are the same;
[0066] The data related to the enterprise to be analyzed acquired by AI and having the same data type as the data set to be cleaned are recorded as data to be judged, and the arrangement of the data types in the data to be judged is consistent with the arrangement of the data types in the data set to be cleaned;
[0067] For any abnormal repetitive data, an abnormal repetitive algorithm is used to obtain the repetitive characteristic value of the abnormal repetitive data. The abnormal repetitive algorithm includes: Among them, F is the repeated feature value, p is the number of data types in the data set to be cleaned, and Q i is the number of data that meets the data duplication standard in the i-th data type of the data set to be cleaned and the abnormal duplicate data, U i Q i The ratio of the total number of data in the i-th data type of the data set to be cleaned, W iis the number of data in the i-th data type of the data to be judged that meets the data duplication standard and the abnormal duplication data, U i W i The ratio of the total number of data in the i-th data type of the data to be judged;
[0068] In the specific implementation process, the abnormal duplicate data obtained is purchase record data, then the number of data that meets the data duplication standard with the purchase record data in all data types in the data set is 5, 10, 2, 4, 0, 0, 0 and 3, respectively, and the ratios of 5, 10, 2, 4, 0, 0, 0 and 3 to the total number of data in all data types in the data set are 0.5, 0.1, 0.2, 0.4, 0, 0, 0 and 0.3 respectively; at the same time, the number of data that meets the data duplication standard with the purchase record data in all data types in the data to be judged is 50, 100, 20, 40, 0, 0, 0 and 30, and the ratios of 50, 100, 20, 40, 0, 0, 0 and 30 to the total number of data types in the data to be judged are 0.8, 0.2, 0.5, 0.8, 0, 0, 0 and 0.6 respectively; the repeated characteristic value obtained by calculation is about 0.068; the data in the data set to be analyzed that meets the data duplication standard with the purchase record data is retained;
[0069] When the repetition characteristic value of abnormal duplicate data is greater than 1, the data in the data set to be analyzed that meets the data duplication standard with the abnormal duplicate data are eliminated; when the repetition characteristic value of abnormal duplicate data is less than or equal to 0, the data in the data set to be analyzed that meets the data duplication standard with the abnormal duplicate data are retained; in the specific implementation process, when the repetition characteristic value of abnormal duplicate data is greater than 1, it means that the proportion of abnormal duplicate data in the data set to be analyzed is greater than the proportion in the conventional enterprise to be analyzed, so the abnormal duplicate data is invalid data and should be eliminated; and when the repetition characteristic value of abnormal duplicate data is less than or equal to 1, it means that the proportion of abnormal duplicate data in the data set to be analyzed is less than or equal to the proportion in the conventional enterprise to be analyzed, so the abnormal duplicate data should be valid data and should not be eliminated;
[0070] The AI is trained using the data that has not been eliminated from the abnormal duplicate data. The training method is: use AI to record the format corresponding to the data that has not been eliminated, and record it as a repeatable format. When AI performs big data cleaning on the data of the enterprise to be analyzed and the format corresponding to any abnormal duplicate data in the cleaning process is equal to the repeatable format, the abnormal duplicate data will be eliminated from all abnormal duplicate data.
[0071] In the specific implementation process, the repeatable format is the format corresponding to the data that is allowed to be repeated in the data set. Therefore, when AI performs big data cleaning and the format corresponding to the abnormal repeated data is equal to the repeatable format, the abnormal repeated data can be eliminated from all abnormal repeated data. This can avoid the abnormal repeated data being identified as redundant items and deleted during the de-redundancy process, causing data missing problems.
[0072] Example 2, please refer to Figure 5 As shown, the present application also provides a big data cleaning method based on AI training, comprising the following steps: Step S1, performing data analysis on the data set to be cleaned based on AI, and obtaining abnormal duplicate data and abnormal regular data in the data set based on the data analysis.
[0073] Step S1 includes the following sub-steps: Step S101, obtaining the type of enterprise in which the data set to be analyzed is located, recorded as the enterprise to be analyzed; using AI based on big data to obtain all data types related to the enterprise to be analyzed;
[0074] Step S102: Analyze all data types using a data anomaly analysis method, and obtain a standard data graph corresponding to each data type based on the analysis result. The data anomaly analysis method includes:
[0075] Step S1021: for any data type α, all data in the data type α are recorded as to-be-extracted data DS1 to to-be-extracted data DS2. m ; For any data DS to be extracted m1 , the data to be extracted DS m1 The amount of data manually filled in is recorded as the filled amount; the data to be extracted DS m1 The data manually filled in are recorded as filled data TR1 to filled data TR n , will fill in data TR1 to fill in data TR n The corresponding character length is recorded as character length ZC1 to character length ZC n , the interval consisting of the minimum and maximum values of all character lengths ZC is recorded as the filled character interval, where m1 is a positive integer less than or equal to m and greater than or equal to 1;
[0076] Step S1022, obtaining the filling quantity and filling character interval corresponding to all the data to be extracted DS; establishing a plane rectangular coordinate system, recorded as a standard sub-one coordinate system, wherein the units of the X-axis and the Y-axis of the standard sub-one coordinate system are both quantity; performing punctuation in the standard sub-one coordinate system based on the number of times each filling quantity appears in all the data to be extracted DS and the numerical value corresponding to the filling quantity, wherein, for any marked point X0, the horizontal coordinate of X0 is a filling quantity corresponding to the data to be extracted DS, and the vertical coordinate of X0 is the number of times the filling quantity appears in the filling quantities of all the data to be extracted DS; fitting all the punctuation points in the standard sub-one coordinate system into a curve, recorded as a standard sub-one curve, and recording the curve after the standard sub-one curve is mirrored based on the X-axis as a mirror sub-one curve;
[0077] Step S1023, establish a plane rectangular coordinate system, recorded as the standard sub-two coordinate system, wherein the units of the X-axis and the Y-axis of the standard sub-two coordinate system are both numbers; for any positive integer C in the filled-in character interval corresponding to all the data DS to be extracted, the number of filled-in character intervals containing the positive integer C is recorded as the existence value of the positive integer C, and in the standard sub-two coordinate system, a bar graph is established at X = positive integer C, the height of the bar graph is the existence value of the positive integer C, and the width of the bar graph is half of the distance between adjacent coordinate points in the X axis; obtain the existence values of all numbers in the filled-in character intervals corresponding to all the DS to be extracted, and draw the corresponding bar graph in the standard sub-two coordinate system; draw a mirror sub-one curve in the fourth quadrant of the standard sub-two coordinate system, and record the standard sub-two coordinate system at this time as the standard data graph of data type α.
[0078] Step S103, obtaining abnormal duplicate data and abnormal regular data in the data set to be cleaned based on the standard data graph corresponding to each data type;
[0079] Step S103 includes: step S1031, using AI to obtain all data types in the data set to be cleaned based on the acquisition method when obtaining the data type of the enterprise to be analyzed; for any data type β in the data set to be cleaned, obtaining the standard sub-one coordinate system and the standard sub-two coordinate system of the data type β based on the data anomaly method, and recording the standard sub-two coordinate system after drawing the mirror sub-one curve in the fourth quadrant of the standard sub-two coordinate system as the data graph to be cleaned of the data type β; overlapping the data graph to be cleaned of the data type β with the standard data graph in the same standard sub-two coordinate system, recording the area where the two curves do not overlap in the first quadrant of the overlapped standard sub-two coordinate system as the filled-in differential area, recording the point in the filled-in differential area whose abscissa is equal to the filled-in amount as the differential point, and marking the abscissa of all the differential points as abnormal filled-in amounts; recording the data in the data type β whose filled-in amount is abnormal filled-in amount as quantity abnormal data;
[0080] Step S1032, record the bar graphs with the same bottom sides and non-overlapping in the fourth quadrant of the overlapped standard sub-two coordinate system as differential bar graphs, mark the horizontal coordinates corresponding to the differential bar graphs as differential fill-in values, and record the heights of the non-overlapping areas in the differential bar graphs as differential quantities; use the differential comparison method to analyze all data in the data type β, the differential comparison method includes: when the positive integers in the fill-in character interval of any data in the data type β are all differential fill-in values and the differential data quantities corresponding to all differential fill-in values are not 0, record the data as abnormal fill-in data, and subtract 1 from the differential fill-in value of the positive integer in the fill-in character interval corresponding to the data; when one abnormal fill-in data is obtained, use the differential comparison method again to analyze all data in the data type β until the abnormal fill-in data cannot be extracted;
[0081] Step S1033, when any data in the data type β is recorded as both quantity abnormal data and filling abnormal data, the data is recorded as abnormal duplicate data; the data in the data type β that is only recorded as quantity abnormal data or filling abnormal data is recorded as abnormal regular data.
[0082] Step S2, performing routine cleaning processing on the abnormal routine data; Step S2 includes: Step S201, analyzing the abnormal routine data based on AI, when there are vacancies in the abnormal routine data that are not filled in manually, using AI to fill the vacant data based on interpolation method;
[0083] Step S202, using AI to analyze the values of the manually filled data in the data that does not contain gaps. When the value of the data is not within the standard range of the data retrieved by AI, the value of the data is recorded as an abnormal value, and the data containing the abnormal value is deleted based on the discard method.
[0084] Step S3, screening the abnormal duplicate data using a redundant quality screening method, and removing redundant data in the abnormal duplicate data based on the screening result, and using the data removed from the abnormal duplicate data and the data not removed to train AI; Step S3 includes: Step S301, screening the abnormal duplicate data using a redundant quality screening method, and removing redundant data in the abnormal duplicate data based on the screening result;
[0085] The redundant quality screening method includes: step S3011, setting a data duplication standard, the data duplication standard is: when the data not manually filled in two data are exactly the same and the values of the data in the manually filled position are the same, the two data are recorded as data duplication, wherein the completely identical means that the characters are the same and the positions of the characters in the data are the same;
[0086] The data related to the enterprise to be analyzed acquired by AI and having the same data type as the data set to be cleaned are recorded as data to be judged, and the arrangement of the data types in the data to be judged is consistent with the arrangement of the data types in the data set to be cleaned;
[0087] For any abnormal repetitive data, an abnormal repetitive algorithm is used to obtain the repetitive characteristic value of the abnormal repetitive data. The abnormal repetitive algorithm includes: Among them, F is the repeated feature value, p is the number of data types in the data set to be cleaned, and Q i is the number of data that meets the data duplication standard in the i-th data type of the data set to be cleaned and the abnormal duplicate data, U i Q i The ratio of the total number of data in the i-th data type of the data set to be cleaned, W i is the number of data in the i-th data type of the data to be judged that meets the data duplication standard and the abnormal duplication data, U i W i The ratio of the total number of data in the i-th data type of the data to be judged;
[0088] Step S3012, when the repetition characteristic value of the abnormal repetitive data is greater than 1, the data in the data set to be analyzed that meets the data repetition standard with the abnormal repetitive data is removed; when the repetition characteristic value of the abnormal repetitive data is less than or equal to 0, the data in the data set to be analyzed that meets the data repetition standard with the abnormal repetitive data is retained.
[0089] Step S302, using the data that has not been eliminated from the abnormal duplicate data to train AI, the training method is: use AI to record the format corresponding to the data that has not been eliminated, record it as a repeatable format, when AI performs big data cleaning on the data of the enterprise to be analyzed and the format corresponding to any abnormal duplicate data in the cleaning process is equal to the repeatable format, the abnormal duplicate data will be eliminated from all abnormal duplicate data.
[0090] Example 3, please refer to Figure 6 As shown, Figure 6The structural schematic diagram of an electronic device is illustrated, and the electronic device may include: a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus. The memory stores computer-readable instructions, and the processor can call the instructions in the memory. When the computer-readable instructions are executed by the processor, the steps in the big data cleaning method based on AI training are executed to achieve the following functions: firstly, data analysis is performed on the data set to be cleaned based on AI, and abnormal duplicate data and abnormal regular data in the data set are obtained based on the data analysis, and then the abnormal regular data is cleaned conventionally, and finally the abnormal duplicate data is screened using a redundant quality screening method, and redundant data in the abnormal duplicate data is eliminated based on the screening results, and the AI is trained using the data eliminated and the data not eliminated in the abnormal duplicate data.
[0091] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.
[0092] Embodiment 4, the present application also provides a computer-readable storage medium, the present application provides a storage medium, on which a computer program is stored. When the computer program is executed by the processor, the steps in the big data cleaning method based on AI training are executed to achieve the following functions: first, data analysis is performed on the data set to be cleaned based on AI, and abnormal duplicate data and abnormal regular data in the data set are obtained based on the data analysis, and then the abnormal regular data is cleaned routinely, and finally, the abnormal duplicate data is screened using a redundant quality screening method, and redundant data in the abnormal duplicate data is eliminated based on the screening results, and the AI is trained using the data eliminated and the data not eliminated in the abnormal duplicate data.
[0093] Through the description of the above implementation methods, the embodiments of the present invention can be provided as methods, systems or computer program products. Based on such an understanding, the above technical solutions can be essentially or partly contributed to the prior art in the form of software products, which can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and include several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0094] In the embodiments provided in the present application, it should be understood that the disclosed system or method can be implemented in other ways. The embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of systems, modules and units can be electrical, mechanical or other forms.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A big data cleaning method based on AI training, characterized in that: The steps include: Perform data analysis on the data set to be cleaned based on AI, and obtain abnormal duplicate data and abnormal regular data in the data set based on data analysis; Carry out routine cleaning of abnormal routine data; The abnormal duplicate data is screened using the redundant quality screening method, and the redundant data in the abnormal duplicate data is eliminated based on the screening results. The AI is trained using the data that has been eliminated and the data that has not been eliminated from the abnormal duplicate data.
2. The big data cleaning method based on AI training according to claim 1 is characterized in that: Perform data analysis on the data set to be cleaned based on AI, and obtain abnormal duplicate data and abnormal regular data in the data set based on data analysis, including: Obtain the enterprise type of the data set to be analyzed, which is recorded as the enterprise to be analyzed; use AI to obtain all data types related to the enterprise to be analyzed based on big data; Use data anomaly analysis methods to analyze all data types, and obtain standard data graphs corresponding to each data type based on the analysis results; Based on the standard data graph corresponding to each data type, abnormal duplicate data and abnormal regular data in the data set to be cleaned are obtained.
3. The big data cleaning method based on AI training according to claim 2 is characterized in that: Data anomaly analysis methods include: For any data type α, all data in the data type α are recorded as data to be extracted DS1 to data to be extracted DS m ; For any data DS to be extracted m1 , the data to be extracted DS m1 The amount of data manually filled in is recorded as the filled amount; the data to be extracted DS m1 The data manually filled in are recorded as filled data TR1 to filled data TR n , will fill in data TR1 to fill in data TR n The corresponding character length is recorded as character length ZC1 to character length ZC n , the interval consisting of the minimum and maximum values of all character lengths ZC is recorded as the filled character interval, where m1 is a positive integer less than or equal to m and greater than or equal to 1; Obtain the filling quantity and filling character interval corresponding to all the data to be extracted DS; establish a plane rectangular coordinate system, recorded as the standard sub-one coordinate system, wherein the units of the X-axis and the Y-axis of the standard sub-one coordinate system are both quantity; perform punctuation in the standard sub-one coordinate system based on the number of times each filling quantity appears in all the data to be extracted DS and the numerical value corresponding to the filling quantity, wherein, for any marked point X0, the horizontal coordinate of X0 is a filling quantity corresponding to the data to be extracted DS, and the vertical coordinate of X0 is the number of times the filling quantity appears in the filling quantities of all the data to be extracted DS; fit all the punctuation points in the standard sub-one coordinate system into a curve, recorded as the standard sub-one curve, and record the curve after the standard sub-one curve is mirrored based on the X-axis as the mirror sub-one curve; A plane rectangular coordinate system is established, recorded as the standard sub-two coordinate system, wherein the units of the X-axis and the Y-axis of the standard sub-two coordinate system are both numbers; for any positive integer C in the filled-in character interval corresponding to all data DS to be extracted, the number of filled-in character intervals containing the positive integer C is recorded as the existence value of the positive integer C, and a bar graph is established at X = positive integer C in the standard sub-two coordinate system, the height of the bar graph is the existence value of the positive integer C, and the width of the bar graph is half of the distance between adjacent coordinate points in the X axis; the existence values of all numbers in the filled-in character intervals corresponding to all DS to be extracted are obtained, and the corresponding bar graphs are drawn in the standard sub-two coordinate system; a mirror sub-one curve is drawn in the fourth quadrant of the standard sub-two coordinate system, and the standard sub-two coordinate system at this time is recorded as the standard data graph of data type α.
4. The big data cleaning method based on AI training according to claim 3 is characterized in that: Based on the standard data graph corresponding to each data type, the abnormal duplicate data and abnormal regular data in the data set to be cleaned are obtained, including: Use AI to obtain all data types in the data set to be cleaned based on the acquisition method when obtaining the data type of the enterprise to be analyzed; for any data type β in the data set to be cleaned, obtain the standard sub-one coordinate system and the standard sub-two coordinate system of the data type β based on the data anomaly method, and record the standard sub-two coordinate system after drawing the mirror sub-one curve in the fourth quadrant of the standard sub-two coordinate system as the data graph to be cleaned of the data type β; overlap the data graph to be cleaned of the data type β with the standard data graph in the same standard sub-two coordinate system, record the area where the two curves do not overlap in the first quadrant of the overlapped standard sub-two coordinate system as the filled-in differential area, record the point in the filled-in differential area whose horizontal coordinate is equal to the filled-in amount as the differential point, and mark the horizontal coordinates of all the differential points as abnormal filled-in amounts; record the data with abnormal filled-in amounts in the data type β as quantitative abnormal data; The bar graphs with the same bottom sides and not overlapping in the fourth quadrant of the standard sub-two coordinate system after the overlap are recorded as differential bar graphs, the horizontal coordinates corresponding to the differential bar graphs are marked as differential filling values, and the heights of the non-overlapping areas in the differential bar graphs are recorded as the differential quantity; all data in the data type β are analyzed using the differential comparison method, and the differential comparison method includes: when the positive integers in the filling character interval of any data in the data type β are all differential filling values and the differential data quantities corresponding to all differential filling values are not 0, the data are recorded as abnormal filling data, and the differential filling value of the positive integer in the filling character interval corresponding to the data is reduced by 1; when one abnormal filling data is obtained, the differential comparison method is used again to analyze all data in the data type β until the abnormal filling data cannot be extracted; When any data in the data type β is recorded as both quantity abnormal data and filling abnormal data, the data is recorded as abnormal duplicate data; the data in the data type β that is only recorded as quantity abnormal data or filling abnormal data is recorded as abnormal regular data.
5. The big data cleaning method based on AI training according to claim 4 is characterized in that: Routine cleaning of abnormal routine data includes: Analyze abnormal and regular data based on AI. When there are gaps in the abnormal and regular data that have not been filled manually, use AI to fill in the gaps based on interpolation. AI is used to analyze the values of manually filled data in data that does not contain gaps. When the value of the data is not within the standard range of the data retrieved by AI, the value of the data is recorded as an abnormal value, and the data with the abnormal value is deleted based on the discard method.
6. The big data cleaning method based on AI training according to claim 5 is characterized in that: The abnormal duplicate data is screened using a redundant quality screening method, and the redundant data in the abnormal duplicate data is removed based on the screening results. The AI is trained using the removed data and the non-removed data in the abnormal duplicate data, including: The abnormal duplicate data are screened using the redundant quality screening method, and the redundant data in the abnormal duplicate data are removed based on the screening results; The AI is trained using the data that has not been eliminated from the abnormal duplicate data. The training method is: use AI to record the format corresponding to the data that has not been eliminated, and record it as a repeatable format. When AI performs big data cleaning on the data of the enterprise to be analyzed and the format corresponding to any abnormal duplicate data in the cleaning process is equal to the repeatable format, the abnormal duplicate data will be eliminated from all abnormal duplicate data.
7. The big data cleaning method based on AI training according to claim 6, characterized in that: Redundant quality screening methods include: Set a data duplication standard, which is: when the data not filled in manually in two data are exactly the same and the values of the data in the manually filled position are the same, the two data are recorded as data duplication, where the exact same means the characters are the same and the positions of the characters in the data are the same; The data related to the enterprise to be analyzed acquired by AI and having the same data type as the data set to be cleaned are recorded as data to be judged, and the arrangement of the data types in the data to be judged is consistent with the arrangement of the data types in the data set to be cleaned; For any abnormal repetitive data, an abnormal repetitive algorithm is used to obtain the repetitive characteristic value of the abnormal repetitive data. The abnormal repetitive algorithm includes: Among them, F is the repeated feature value, p is the number of data types in the data set to be cleaned, and Q i is the number of data that meets the data duplication standard in the i-th data type of the data set to be cleaned and the abnormal duplicate data, U i Q i The ratio of the total number of data in the i-th data type of the data set to be cleaned, W i is the number of data in the i-th data type of the data to be judged that meets the data duplication standard and the abnormal duplication data, U i W i The ratio of the total number of data in the i-th data type of the data to be judged; When the repetition characteristic value of abnormal duplicate data is greater than 1, the data in the data set to be analyzed that meets the data duplication standard with the abnormal duplicate data will be eliminated; when the repetition characteristic value of abnormal duplicate data is less than or equal to 0, the data in the data set to be analyzed that meets the data duplication standard with the abnormal duplicate data will be retained.
8. A big data cleaning system based on AI training, used to implement the big data cleaning method based on AI training according to any one of claims 1 to 7, characterized in that: Includes abnormal screening module, regular cleaning module and repeated cleaning module; The anomaly screening module is used to perform data analysis on the data set to be cleaned based on AI, and obtain abnormal duplicate data and abnormal regular data in the data set based on data analysis; The routine cleaning module is used to perform routine cleaning on abnormal routine data; The duplicate cleaning module is used to screen abnormal duplicate data using the redundant quality screening method, and to eliminate redundant data in the abnormal duplicate data based on the screening results, and to train AI using the data eliminated and the data not eliminated from the abnormal duplicate data.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps in the method according to any one of claims 1 to 7 are executed.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 7 are executed.