Data analysis management method and system based on digital intelligent engineering
By performing position and anomaly analysis of the digital engineering data sets, using European computing and K nearest neighbor algorithms to fill the data, the data incompleteness problem is solved, the data call time is shortened and the analysis accuracy is improved.
Patent Information
- Application Number
- CN202510483916.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-17
AI Technical Summary
There are missing values and outliers in the data in the digital intelligence project, which leads to incomplete data and extends the data call time.
Through the analysis and management terminal, the data set is analyzed and abnormally analyzed, the location of the data to be filled is determined, and the data is filled using European calculations and K nearest neighbor algorithms to generate a normal data set.
The filling and modification of missing data and abnormal data is realized, which shortens the data call time of digital intelligence engineering, and improves the integrity and analysis accuracy of data.
Smart Images

Figure CN120336306A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis management, and specifically relates to a data analysis management method and system based on digital intelligence engineering. Background Art
[0002] Digital intelligence engineering is a systematic project that deeply integrates digitization and intelligence, aiming to optimize and innovate the entire business process through new generation information technologies such as big data, artificial intelligence, and the Internet of Things, and promote the reconstruction of the industrial ecosystem. Its core lies in using data to drive decision-making and intelligent empowerment to improve efficiency, forming "digital intelligence" capabilities to serve social and economic development and industrial upgrading.
[0003] There are missing values and outliers in the data of digital intelligence engineering. If the missing values are not supplemented and the outliers are not modified, it will lead to incomplete data in digital intelligence engineering. When the data of digital intelligence engineering needs to be called later, it is also necessary to re-supplement and modify the missing values and outliers, which prolongs the data call time of digital intelligence engineering. Summary of the Invention
[0004] To solve the above technical problems, a data analysis management method and system based on digital intelligence engineering are provided. The technical solution of the present invention solves the problem that there are missing values and outliers in the data of digital intelligence engineering mentioned in the above background art. If the missing values are not supplemented and the outliers are not modified, it will lead to incomplete data in digital intelligence engineering. When the data of digital intelligence engineering needs to be called later, it is also necessary to re-supplement and modify the missing values and outliers, which prolongs the data call time of digital intelligence engineering.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows: A data analysis management method based on digital intelligence engineering, comprising: Obtain a dataset to be managed, and based on an analysis management terminal, perform location analysis on the dataset to be managed to determine the position of the first data to be filled; The analysis management terminal performs outlier analysis on the dataset to be managed according to the position of the first data to be filled to determine the position of the second data to be filled and the dataset to be filled; Based on the analysis management terminal, perform data filling on the dataset to be filled according to the position of the first data to be filled and the position of the second data to be filled to obtain a normal dataset.
[0006] Preferably, the step of obtaining the dataset to be managed and performing location analysis on the dataset to be managed based on an analysis management terminal to determine the position of the first data to be filled specifically includes the following steps: Based on the analysis management terminal, perform data reading processing on the database system to obtain the dataset to be managed; Based on the analysis management terminal, traverse and process the dataset to be managed to determine the data missing positions; Based on the analysis management terminal, perform information reading processing on the dataset to be managed with the data missing positions as features, and obtain the row information and column information of the data missing positions; The analysis management terminal generates missing position labels according to the row information and column information of the data missing positions; Based on the analysis management terminal, embed the missing position labels into the data missing positions, and set the data missing positions containing the missing position labels as the first data positions to be filled.
[0007] Preferably, the analysis management terminal performs anomaly analysis on the dataset to be managed according to the first data positions to be filled, and determines the second data positions to be filled and the dataset to be filled, which specifically includes the following steps: Obtain the dataset to be analyzed for anomalies; Based on the analysis management terminal, perform calculation processing on the dataset to be analyzed for anomalies to determine the anomaly data positions; The analysis management terminal deletes data from the dataset to be managed according to the anomaly data positions, and determines the second data positions to be filled and the dataset to be filled.
[0008] Preferably, the obtaining of the dataset to be analyzed for anomalies specifically includes the following steps: Based on the analysis management terminal, perform information reading processing on the missing position labels of each first data position to be filled, and obtain the row information of each missing data; Based on the analysis management terminal, perform counting processing on the row information of each missing data to obtain the missing quantity of each row of data; Based on the analysis management terminal, perform judgment processing on the missing quantity of each row of data to determine the first dataset to be deleted; Based on the analysis management terminal, delete the missing data in the first dataset to be deleted, and set the remaining data as the dataset to be analyzed for anomalies.
[0009] Preferably, the based on the analysis management terminal, performing judgment processing on the missing quantity of each row of data to determine the first dataset to be deleted specifically includes the following steps: Based on the analysis management terminal, perform judgment processing on the missing quantity of each row of data and the set missing quantity threshold; If the missing quantity of each row of data is greater than or equal to the set missing quantity threshold, the data in this row is missing too much and cannot be filled with data. Delete the data in this row, and set the remaining data as the first dataset to be deleted; If the number of missing values in each row of data is less than the set missing value threshold, the missing data in that row is less, and subsequent data filling can be performed. All rows of data with a number of missing values less than the set missing value threshold are set as the first deleted data set.
[0010] Preferably, the analysis management terminal performs calculation processing on the data set to be abnormally analyzed to determine the position of abnormal data, which specifically includes the following steps: Based on the analysis management terminal, perform multiple calculation processes on the data set to be abnormally analyzed to obtain the standard deviation and mean value of the data set to be abnormally analyzed; Based on the analysis management terminal, perform calculation processing on each data, standard deviation, and mean value in the data set to be abnormally analyzed to obtain the information to be verified for each data; Based on the analysis management terminal, perform judgment processing on the information to be verified for each data to determine the position of abnormal data.
[0011] Preferably, the analysis management terminal performs judgment processing on the information to be verified for each data to determine the position of abnormal data, which specifically includes the following steps: Based on the analysis management terminal, perform judgment processing on the information to be verified for each data and the set verification data threshold; If the absolute value of the information to be verified is greater than the set verification data threshold, the data corresponding to the information to be verified is abnormal data. Record the row information and column information of this data, and set the row information and column information of this data as the position of abnormal data; If the absolute value of the information to be verified is less than or equal to the set verification data threshold, the data corresponding to the information to be verified is normal data.
[0012] Preferably, the analysis management terminal performs data filling on the data set to be filled according to the first data position to be filled and the second data position to be filled to obtain a normal data set, which specifically includes the following steps: Based on the analysis management terminal, perform data standardization processing on the data set to be filled; Based on the first data position to be filled and the second data position to be filled, perform calculation processing on the remaining data respectively to obtain a first data similarity set and a second data similarity set; Based on the minimum value function, perform sorting processing on the data in the first data similarity set and the second data similarity set respectively; Based on the random number function, perform length selection on the first data similarity set and the second data similarity set respectively to obtain a first data length and a second data length; Based on the analysis management terminal, perform data selection on the first data similarity set and the second data similarity set respectively according to the first data length and the second data length to obtain a first data similarity subset and a second data similarity subset; Based on the analysis management terminal, calculate the mean values of the data within the first data similarity subset and the second data similarity subset respectively to obtain the first filled data and the second filled data; Based on the analysis management terminal, fill the first filled data and the second filled data into the first data position to be filled and the second data position to be filled respectively to obtain the dataset to be verified; Based on the analysis management terminal, perform verification processing on the dataset to be verified to obtain the first data adjustment length and the second data adjustment length; Based on the analysis management terminal, recalculate and fill the data of the first data similarity set and the second data similarity set according to the first data adjustment length and the second data adjustment length to obtain the normal dataset.
[0013] Preferably, the step of performing verification processing on the dataset to be verified based on the analysis management terminal to obtain the first data adjustment length and the second data adjustment length specifically includes the following steps: Based on the analysis management terminal, perform calculation processing on the dataset to be verified respectively to obtain the mean value of the row where the first filled data is located and the mean value of the row where the second filled data is located; Based on the analysis management terminal, perform judgment processing on the mean value of the row where the first filled data is located and the first filled data, and the mean value of the row where the first filled data and the second filled data are located and the second filled data respectively; If the mean value of the row where the first filled data is located is much greater than or much less than the first filled data, based on the analysis management terminal, adjust the length of the first data to obtain the first data adjustment length; If the mean value of the row where the first filled data is located is approximately equal to the first filled data, there is no need to adjust the length of the first data; If the mean value of the row where the second filled data is located is much greater than or much less than the second filled data, based on the analysis management terminal, adjust the length of the second data to obtain the second data adjustment length; If the mean value of the row where the second filled data is located is approximately equal to the second filled data, there is no need to adjust the length of the second data.
[0014] Furthermore, a data analysis management system based on digital intelligence engineering is proposed, which is used to implement a data analysis management method based on digital intelligence engineering as described above, including: An analysis management terminal, which is used to control each module to perform position analysis, anomaly analysis and data filling on the dataset to be managed to obtain a normal dataset, and the analysis management terminal is used to control information interaction and data transmission between each module; A database system, which is used to store the dataset to be managed; A data missing position determination module, which is used to traverse and read information from the dataset to be managed to determine the position of the first data to be filled. A data anomaly position determination module, which is used to delete and calculate data from the dataset to be managed to determine the position of the second data to be filled and the dataset to be filled. A filled data calculation module, which calculates the similarity, selects the length, and calculates the data mean of the dataset to be filled according to the position of the first data to be filled and the position of the second data to be filled to obtain the dataset to be verified. A filled data verification module, which is used to verify the data of the dataset to be verified and perform secondary calculation of the filled data to obtain a normal dataset.
[0015] Furthermore, a storage medium is proposed, on which a computer program is stored. When the computer program is called and run, it executes a data analysis and management method based on digital intelligence engineering as described above.
[0016] Compared with the prior art, the present invention provides a data analysis and management method and system based on digital intelligence engineering, which has the following beneficial effects: The present invention first performs position analysis on the dataset to be managed to determine the position of the first data to be filled. Then, it performs anomaly analysis on the dataset to be managed to determine the position of the abnormal data and sets the position of the abnormal data as the position of the second data to be filled. Finally, through Euclidean calculation, the dataset to be filled is calculated to obtain the first data similarity set and the second data similarity set, and the first data similarity set and the second data similarity set are selected and calculated to obtain the first filled data and the second filled data. To avoid the first filled data and the second filled data obtained by calculation being abnormal data, the first filled data and the second filled data are verified to determine a normal dataset. The above method not only realizes the filling and modification of missing data and abnormal data, but also shortens the data call time of digital intelligence engineering. Brief Description of the Drawings
[0017] Figure 1 It is a schematic flow chart of steps S100 - S300 in a data analysis and management method and system based on digital intelligence engineering proposed by the present invention; Figure 2 It is a structural block diagram of a data analysis and management method and system based on digital intelligence engineering proposed by the present invention. Detailed Embodiment
[0018] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are only examples, and those skilled in the art can think of other obvious variations.
[0019] Reference Figure 1 As shown, a data analysis and management method based on digital intelligence engineering includes: S100. Obtain the data set to be managed, and based on the analysis and management terminal, perform location analysis on the data set to be managed to determine the first data position to be filled; It can be understood that the data processed here are all numerical data, and a series of subsequent data processing methods are designed for numerical data; S200. The analysis and management terminal performs anomaly analysis on the data set to be managed according to the first data position to be filled, and determines the second data position to be filled and the data set to be filled; S300. Based on the analysis and management terminal, perform data filling on the data set to be filled according to the first data position to be filled and the second data position to be filled, and obtain a normal data set; Those skilled in the art can understand that data may be missing and abnormal during collection. If the data is stored without being processed, it will cause the missing data and abnormal data to be stored in the storage space as well. When the data needs to be called later, it is necessary to reprocess the missing data and abnormal data in the data, which prolongs the data call time. If the missing data and abnormal data are not processed during data call, it will increase the error rate of subsequent data analysis. The same is true for the data of digital intelligence engineering. Therefore, before storing the data of digital intelligence engineering, fill in the missing data and modify the abnormal data to ensure the integrity and correctness of the data of digital intelligence engineering, shorten the subsequent data call time of digital intelligence engineering, and at the same time, improve the accuracy of subsequent data analysis.
[0020] Embodiment 1 Obtaining the data set to be managed, and based on the analysis and management terminal, performing location analysis on the data set to be managed to determine the first data position to be filled specifically includes the following steps: S101. Based on the analysis and management terminal, perform data reading processing on the database system to obtain the data set to be managed; S102. Based on the analysis and management terminal, perform traversal processing on the data set to be managed to determine the data missing positions; S103. Based on the analysis and management terminal, perform information reading processing on the data set to be managed with the data missing positions as features, and obtain the row information and column information of the data missing positions; S104. The analysis and management terminal generates missing position labels according to the row information and column information of the data missing positions; S105. Based on the analysis and management terminal, embed the missing position labels into the data missing positions, and set the data missing positions containing the missing position labels as the first data positions to be filled; In this embodiment, when collecting numerical data, it is recorded according to a certain rule. Data with the same characteristics or the same information is placed in the same row. After data collection, it needs to be recorded. During the recording process, there may be cases of missing or incorrect recording. Therefore, there will be missing data and abnormal data in the data. Therefore, it is necessary to fill in the missing data. First, it is necessary to determine the position of the missing data. Therefore, by traversing the dataset to be managed, the position of the missing data is determined. Then, according to the row information and column information of the missing data position, a missing position label is generated, and then the missing position label is placed at the position of the missing data. Because when the data is missing, the staff may think that there is an extra space in the missing data position, resulting in an extra blank position. Therefore, the missing position label is used to remind the staff that this position is missing data. The above method is to prevent the staff from performing other tasks after determining the position of the missing data, and another staff member performs the subsequent tasks. However, this staff member may not know the processing method of the previous staff member and will delete the missing data position, which will lead to incomplete data in the digital intelligence project.
[0021] Embodiment 2 The analysis and management terminal performs anomaly analysis on the dataset to be managed according to the first data position to be filled, and determines the second data position to be filled and the dataset to be filled, which specifically includes the following steps: S201. Obtain the dataset to be analyzed for anomalies; S202. Based on the analysis and management terminal, perform calculation processing on the dataset to be analyzed for anomalies to determine the position of the abnormal data; S203. The analysis and management terminal deletes data from the dataset to be managed according to the position of the abnormal data, and determines the second data position to be filled and the dataset to be filled; Among them, S203. The analysis and management terminal deletes data from the dataset to be managed according to the position of the abnormal data, and determines the second data position to be filled and the dataset to be filled, which specifically includes the following steps: S2031. The analysis and management terminal deletes the abnormal data from the dataset to be managed according to the position of the abnormal data, and obtains the second trimmed dataset; S2032. Repeat steps S2011 - S2014 and steps S20131 - S20133 to obtain the third trimmed dataset; S2033. Repeat steps S103 - S105 to obtain the second data position to be filled, and set the third trimmed dataset as the dataset to be filled; In this embodiment, after the abnormal data is determined, it is necessary to replace the abnormal data. Since the abnormal data is the result of incorrect recording, the abnormal data has no relevance to other data in the row where the abnormal data is located. Therefore, the abnormal data is deleted. However, after the abnormal data is deleted, a blank data bit will be left. At this time, the blank data bit left by the abnormal data is equivalent to data loss. Therefore, the same processing method for the data loss position is used to process the blank data bit to determine the second data position to be filled and the data set to be filled.
[0022] Embodiment 3 Obtaining the data set to be analyzed for abnormalities specifically includes the following steps: S2011. Based on the analysis management terminal, perform information reading processing on the missing position tags of each first data position to be filled, and obtain the row information of each missing data; S2012. Based on the analysis management terminal, perform counting processing on the row information of each missing data to obtain the missing quantity of each row of data; It can be understood that first, each row of data is traversed. When it is determined that the data at a certain position is missing, counting is performed through a counter; S2013. Based on the analysis management terminal, perform judgment processing on the missing quantity of each row of data to determine the first data set to be deleted; S2014. Based on the analysis management terminal, delete the missing data in the first data set to be deleted, and set the remaining data as the data set to be analyzed for abnormalities; Among them, S2013. Based on the analysis management terminal, perform judgment processing on the missing quantity of each row of data to determine the first data set to be deleted specifically includes the following steps: S20131. Based on the analysis management terminal, perform judgment processing on the missing quantity of each row of data and the set missing quantity threshold; S20132. If the missing quantity of each row of data is greater than or equal to the set missing quantity threshold, the data in this row is missing too much and cannot be filled with data. Delete the data in this row, and set the remaining data as the first data set to be deleted; S20133. If the missing quantity of each row of data is less than the set missing quantity threshold, the data in this row is missing less and can be filled with subsequent data. Set all the row data with a missing quantity less than the set missing quantity threshold as the first data set to be deleted; In this embodiment, when there is too much data missing in a certain row, the missing data cannot be supplemented based on the remaining data. Because when there is too much data missing, the correlation between the data cannot be obtained, and subsequent supplementation cannot be carried out. Moreover, the data in the row with too much data missing has little reference significance in the subsequent analysis process. Therefore, in order to reduce the usage rate of the storage space, the row of data with too much data missing is deleted.
[0023] Embodiment 4 Based on the analysis management terminal, performing calculation processing on the dataset to be abnormally analyzed to determine the abnormal data location specifically includes the following steps: S2021. Based on the analysis management terminal, performing multiple calculation processes on the dataset to be abnormally analyzed to obtain the standard deviation and mean value of the dataset to be abnormally analyzed; S2022. Based on the analysis management terminal, performing calculation processing on each data, the standard deviation, and the mean value in the dataset to be abnormally analyzed to obtain the information to be verified for each data; Among them, the calculation formula for obtaining the information to be verified for each data is:
[0024] In the formula, is the information to be verified for each data in each row; is each data in each row of the dataset to be abnormally analyzed; is the mean value of each row in the dataset to be abnormally analyzed; is the standard deviation of each row in the dataset to be abnormally analyzed; It can be understood that the above calculation formula is calculated for each data in each row, and the standard deviation and mean value are both the standard deviation and mean value of that row. After the calculation of each data in that row is completed, it is necessary to calculate the standard deviation and mean value corresponding to the information to be verified for each data in the next row. The standard deviation and mean value of the data in the next row may be the same as those in the previous row or may be different; S2023. Based on the analysis management terminal, performing judgment processing on the information to be verified for each data to determine the abnormal data location; It can be understood that the Z-Score method can be used to determine the abnormal data, that is, calculating the dataset to be abnormally analyzed through the Z-Score method to obtain the information to be verified for each data, and then judging the information to be verified for each data to determine the abnormal data; Among them, S2023. Based on the analysis management terminal, performing judgment processing on the information to be verified for each data to determine the abnormal data location specifically includes the following steps: S20231. Based on the analysis management terminal, judge and process the information to be verified of each data and the set verification data threshold; S20232. If the absolute value of the information to be verified is greater than the set verification data threshold, the data corresponding to the information to be verified is abnormal data. Record the row information and column information of this data, and set the row information and column information of this data as the abnormal data position; S20233. If the absolute value of the information to be verified is less than or equal to the set verification data threshold, the data corresponding to the information to be verified is normal data; In this embodiment, the specific data value of the set verification data threshold is 3. When the absolute value of the information to be verified is greater than 3, it indicates that the data corresponding to the information to be verified is abnormal data, and the row information and column information of this data need to be recorded, and the row information and column information of this data are set as the abnormal data position. When the absolute value of the information to be verified is less than or equal to 3, the data corresponding to the information to be verified is normal data. It can be understood that in the Z-Score method, data greater than 3 or less than -3 can be regarded as abnormal data. In order to avoid setting two comparison thresholds, the absolute value of the information to be verified is calculated, making the comparison and judgment process more convenient and fast.
[0025] Embodiment 5 Based on the analysis management terminal, data filling is performed on the data set to be filled according to the first data position to be filled and the second data position to be filled, and obtaining the normal data set specifically includes the following steps: S301. Based on the analysis management terminal, perform data standardization processing on the data set to be filled; It can be understood that performing data standardization processing on the data set to be filled is to map the data to the interval. Because the subsequent similarity calculation is performed by the K-nearest neighbor algorithm, and the K-nearest neighbor algorithm is calculated based on distance metrics, therefore, it is necessary to perform data standardization processing on the data set to be filled; S302. Based on the first data position to be filled and the second data position to be filled, calculate and process the remaining data respectively to obtain the first data similarity set and the second data similarity set; It can be understood that the calculation method of this step is calculated by the Euclidean distance; Among them, the specific calculation formula for obtaining the first data similarity set is:
[0026] In the formula, is the similarity data in the first data similarity set; and are the values of the i-th row and all other rows j on the k-th feature; n is the number of features; S303. Sort the data in the first data similarity set and the second data similarity set respectively based on the minimum value function; S304. Select the lengths of the first data similarity set and the second data similarity set respectively based on the random number function to obtain the first data length and the second data length; S305. Based on the analysis management terminal, select the data in the first data similarity set and the second data similarity set respectively according to the first data length and the second data length to obtain the first data similarity subset and the second data similarity subset; It can be understood that the smaller the data in the first data similarity set and the second data similarity set, the higher the correlation between the data in the missing row and the missing data. Therefore, sort them through the minimum value function, and then select the lengths through the random number function. The first data similarity subset and the second data similarity subset are selected from small to large. The first data length and the second data length specifically refer to the number of data selections; S306. Based on the analysis management terminal, calculate the mean values of the data within the first data similarity subset and the second data similarity subset respectively to obtain the first filled data and the second filled data; S307. Based on the analysis management terminal, fill the first filled data and the second filled data into the first data position to be filled and the second data position to be filled respectively to obtain the dataset to be verified; S308. Based on the analysis management terminal, perform verification processing on the dataset to be verified to obtain the first data adjustment length and the second data adjustment length; S309. Based on the analysis management terminal, recalculate and fill the data in the first data similarity set and the second data similarity set according to the first data adjustment length and the second data adjustment length to obtain the normal dataset; In this embodiment, the specific steps for obtaining the first filled data and the second filled data by the K-nearest neighbor algorithm are as follows: The first step. Perform data standardization processing on the data to be filled so that the data to be filled can be mapped to the interval. It can be understood that the K-nearest neighbor algorithm is for numerical data, and the data within the data to be filled are all numerical data; The second step. Calculate the distances between the data in the missing row and the data in the remaining other rows through the Euclidean distance to obtain the first data similarity set and the second data similarity set. It can be understood that the similarity data specifically refers to the distance between two data; Step 3: Sort the first data similarity set and the second data similarity set through the minimum value function. Because the smaller the values in the first data similarity set and the second data similarity set, the smaller the distance between the two data, which indirectly indicates a stronger correlation between the two data; Step 4: Select an appropriate data length through the random number function. The specific data length is the number of selected data, and the data selection rule is to select in ascending order; Step 5: Calculate the mean of the selected data to obtain the first filled data and the second filled data; It can be understood that when the data length selected by the random number function is too long, some larger data will be filtered out. Because larger data represents a smaller correlation between the two data, it will cause the filled data to be too large. When subsequent data analysis is required, it will indirectly increase the error rate. Therefore, it is necessary to verify the filled data and adjust the first data length and the second data length; Among them, S308: Based on the analysis management terminal, verify the dataset to be verified and obtain the first data adjustment length and the second data adjustment length, which specifically includes the following steps: S3081: Based on the analysis management terminal, perform calculation processing on the dataset to be verified respectively to obtain the mean of the row where the first filled data is located and the mean of the row where the second filled data is located; S3082: Based on the analysis management terminal, perform judgment processing on the mean of the row where the first filled data is located and the first filled data, and the mean of the row where the first filled data and the second filled data are located and the second filled data respectively; S3083: If the mean of the row where the first filled data is located is much greater than or much less than the first filled data, based on the analysis management terminal, adjust the first data length to obtain the first data adjustment length; S3084: If the mean of the row where the first filled data is located is approximately equal to the first filled data, there is no need to adjust the first data length; S3085: If the mean of the row where the second filled data is located is much greater than or much less than the second filled data, based on the analysis management terminal, adjust the second data length to obtain the second data adjustment length; S3086: If the mean of the row where the second filled data is located is approximately equal to the second filled data, there is no need to adjust the second data length; In this embodiment, if the data length selected by the random number function is too long, some larger data will be selected, resulting in too large padding data. If the data length selected by the random number function is too small, too little data will be selected, resulting in too small padding data. Therefore, by verifying the first padding data and the second padding data, it is determined whether the first padding data and the second padding data meet the standards. If they do not meet the standards, the data length selected by the random number function is adjusted, and then the judgment is made again until appropriate data is selected. The above method reduces the error rate generated by subsequent data analysis, and at the same time ensures the integrity of the data, making the samples for subsequent data analysis sufficient.
[0027] Referring Figure 2 As shown, a data analysis management system based on digital intelligence engineering is used to implement a data analysis management method based on digital intelligence engineering as described above, including: An analysis management terminal, which is used to control each module to perform position analysis, anomaly analysis, and data filling on the dataset to be managed, obtain a normal dataset, and control information interaction and data transmission between each module; A database system, which is used to store the dataset to be managed; A data missing position determination module, which is used to traverse the data and read information of the dataset to be managed to determine the position of the first data to be filled; A data anomaly position determination module, which is used to delete data and perform data calculations on the dataset to be managed to determine the position of the second data to be filled and the dataset to be filled; A padding data calculation module, which calculates the similarity, selects the length, and calculates the data mean of the dataset to be filled according to the position of the first data to be filled and the position of the second data to be filled to obtain a dataset to be verified; A padding data verification module, which is used to verify the data of the dataset to be verified and perform secondary calculations on the padding data to obtain a normal dataset.
[0028] Furthermore, a storage medium is proposed, on which a computer program is stored. When the computer program is called and run, it executes a data analysis management method based on digital intelligence engineering as described above. Among them, the storage medium can be a magnetic medium, such as a floppy disk, a hard disk, a magnetic tape; an optical medium, such as a DVD; or a semiconductor medium, such as a solid state disk (SSD).
[0029] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, various changes and improvements will occur to the present invention, and all these changes and improvements fall within the scope of the present invention claimed. The scope of protection required by the present invention is defined by the appended claims and their equivalents.
Claims
1. A data analysis and management method based on digital intelligence engineering, characterized in that, Including: Obtain the dataset to be managed, perform location analysis on the dataset to be managed based on the analysis management terminal, and determine the first data position to be filled; The analysis management terminal performs anomaly analysis on the dataset to be managed according to the first data position to be filled, and determines the second data position to be filled and the dataset to be filled; Based on the analysis management terminal, perform data filling on the dataset to be filled according to the first data position to be filled and the second data position to be filled, and obtain a normal dataset.
2. The data analysis and management method based on digital intelligence engineering according to claim 1, characterized in that The step of obtaining the dataset to be managed, performing location analysis on the dataset to be managed based on the analysis management terminal, and determining the first data position to be filled specifically includes the following steps: Based on the analysis management terminal, perform data reading processing on the database system to obtain the dataset to be managed; Based on the analysis management terminal, perform traversal processing on the dataset to be managed to determine the data missing positions; Based on the analysis management terminal, perform information reading processing on the dataset to be managed with the data missing positions as features, and obtain the row information and column information of the data missing positions; The analysis management terminal generates a missing position label according to the row information and column information of the data missing positions; Based on the analysis management terminal, embed the missing position label into the data missing position, and set the data missing position containing the missing position label as the first data position to be filled.
3. A data analysis and management method based on digital intelligence engineering according to claim 1, characterized in that, The step that the analysis management terminal performs anomaly analysis on the dataset to be managed according to the first data position to be filled, and determines the second data position to be filled and the dataset to be filled specifically includes the following steps: Obtain the dataset to be analyzed for anomalies; Based on the analysis management terminal, perform calculation processing on the dataset to be analyzed for anomalies to determine the anomaly data positions; The analysis management terminal deletes data from the dataset to be managed according to the anomaly data positions, and determines the second data position to be filled and the dataset to be filled.
4. A data analysis and management method based on digital intelligence engineering according to claim 3, characterized in that, The step of obtaining the dataset to be analyzed for anomalies specifically includes the following steps: Based on the analysis management terminal, perform information reading processing on the missing position labels of each first data position to be filled, and obtain the row information of each missing data; Based on the analysis management terminal, perform counting processing on the row information of each missing data, and obtain the missing quantity of each row of data; Based on the analysis management terminal, perform judgment processing on the missing quantity of each row of data to determine the first dataset to be trimmed; Based on the analysis management terminal, delete the missing data in the first dataset to be trimmed, and set the remaining data as the dataset to be analyzed for anomalies.
5. A data analysis and management method based on digital intelligence engineering according to claim 4, characterized in that, The step that based on the analysis management terminal, perform judgment processing on the missing quantity of each row of data to determine the first dataset to be trimmed specifically includes the following steps: Based on the analysis management terminal, perform judgment processing on the missing quantity of each row of data and the set missing quantity threshold; If the missing quantity of each row of data is greater than or equal to the set missing quantity threshold, the data in this row is missing too much and cannot be filled, delete the data in this row, and set the remaining data as the first dataset to be trimmed; If the missing quantity of each row of data is less than the set missing quantity threshold, the data in this row is missing less and can be filled with subsequent data, and set all the row data with a missing quantity less than the set missing quantity threshold as the first dataset to be trimmed.
6. A data analysis and management method based on digital intelligence engineering according to claim 3, characterized in that, The above-mentioned analysis management terminal performs calculation processing on the data set to be abnormally analyzed to determine the abnormal data position, which specifically includes the following steps: Based on the analysis management terminal, perform multiple calculation processes on the data set to be abnormally analyzed to obtain the standard deviation and mean value of the data set to be abnormally analyzed; Based on the analysis management terminal, perform calculation processing on each data, standard deviation, and mean value in the data set to be abnormally analyzed to obtain the information to be verified for each data; Based on the analysis management terminal, perform judgment processing on the information to be verified for each data to determine the abnormal data position.
7. A data analysis and management method based on digital intelligence engineering according to claim 6, characterized in that, The above-mentioned analysis management terminal performs judgment processing on the information to be verified for each data to determine the abnormal data position, which specifically includes the following steps: Based on the analysis management terminal, perform judgment processing on the information to be verified for each data and the set verification data threshold; If the absolute value of the information to be verified is greater than the set verification data threshold, the data corresponding to the information to be verified is abnormal data, record the row information and column information of this data, and set the row information and column information of this data as the abnormal data position; If the absolute value of the information to be verified is less than or equal to the set verification data threshold, the data corresponding to the information to be verified is normal data.
8. A data analysis and management method based on digital intelligence engineering according to claim 1, characterized in that The above-mentioned analysis management terminal performs data filling on the data set to be filled according to the first data position to be filled and the second data position to be filled to obtain a normal data set, which specifically includes the following steps: Based on the analysis management terminal, perform data standardization processing on the data set to be filled; Based on the first data position to be filled and the second data position to be filled, perform calculation processing on the remaining data respectively to obtain the first data similarity set and the second data similarity set; Based on the minimum value function, perform sorting processing on the data in the first data similarity set and the second data similarity set respectively; Based on the random number function, perform length selection on the first data similarity set and the second data similarity set respectively to obtain the first data length and the second data length; Based on the analysis management terminal, perform data selection on the first data similarity set and the second data similarity set respectively according to the first data length and the second data length to obtain the first data similarity subset and the second data similarity subset; Based on the analysis management terminal, perform mean value calculation on the data inside the first data similarity subset and the second data similarity subset respectively to obtain the first filling data and the second filling data; Based on the analysis management terminal, fill the first filling data and the second filling data into the first data position to be filled and the second data position to be filled respectively to obtain the data set to be verified; Based on the analysis management terminal, perform verification processing on the data set to be verified to obtain the first data adjusted length and the second data adjusted length; Based on the analysis management terminal, re-perform data calculation and data filling processing on the first data similarity set and the second data similarity set according to the first data adjusted length and the second data adjusted length to obtain a normal data set.
9. A data analysis and management method based on digital intelligence engineering according to claim 8, characterized in that The above-mentioned analysis management terminal performs verification processing on the data set to be verified to obtain the first data adjusted length and the second data adjusted length, which specifically includes the following steps: Based on the analysis management terminal, calculate and process the to-be-verified data set respectively to obtain the mean value of the row where the first filled data is located and the mean value of the row where the second filled data is located; Based on the analysis management terminal, judge and process the mean value of the row where the first filled data is located and the first filled data, and the mean value of the row where the first filled data and the second filled data are located and the second filled data respectively; If the mean value of the row where the first filled data is located is much larger or much smaller than the first filled data, based on the analysis management terminal, adjust the length of the first data to obtain the adjusted length of the first data; If the mean value of the row where the first filled data is located is approximately equal to the first filled data, there is no need to adjust the length of the first data; If the mean value of the row where the second filled data is located is much larger or much smaller than the second filled data, based on the analysis management terminal, adjust the length of the second data to obtain the adjusted length of the second data; If the mean value of the row where the second filled data is located is approximately equal to the second filled data, there is no need to adjust the length of the second data.
10. A data analysis and management system based on digital intelligence engineering is used to implement a data analysis and management method based on digital intelligence engineering as described in any one of claims 1-9, characterized in that, It includes: An analysis management terminal, which is used to control each module to perform position analysis, anomaly analysis and data filling on the data set to be managed, obtain a normal data set, and the analysis management terminal is used to control information interaction and data transmission between each module; A database system, which is used to store the data set to be managed; A data missing position determination module, which is used to traverse the data and read information of the data set to be managed to determine the position of the first data to be filled; A data anomaly position determination module, which is used to delete data and perform data calculation on the data set to be managed to determine the position of the second data to be filled and the data set to be filled; A filled data calculation module, which calculates the similarity, selects the length, and calculates the data mean value of the data set to be filled according to the position of the first data to be filled and the position of the second data to be filled to obtain a to-be-verified data set; A filled data verification module, which is used to verify the data of the to-be-verified data set and perform secondary calculation of the filled data to obtain a normal data set.
Citation Information
Patent Citations
Data analysis method based on big data
CN117708636A
Method for intelligently realizing big data early warning based on model strategy
CN119168788A