An industrial big data data governance method
By acquiring, cleaning, transforming, and auditing industrial data, improving it according to data standards, and sharing it openly via API, the problems of low quality and difficulty in using industrial data have been solved, improving data applicability and reducing the waste of manual verification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-24
- Publication Date
- 2026-03-24
AI Technical Summary
Industrial data suffers from problems such as inconsistent standards, low quality, and difficulty in utilization in the industrial sector, which limits its application.
By acquiring multi-source heterogeneous industrial data, we clean, transform, and audit it, improve it according to published data standards, and open up data application services through API sharing.
It improves the quality and applicability of industrial data, solves the problem of data being difficult to utilize, and reduces the waste of manual data quality verification.
Smart Images

Figure CN116628059B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of processing application of industrial data, and particularly relates to an industrial big data data management method. BACKGROUND
[0002] At present, with the rapid development of the industrial internet, the rapid growth of industrial data brings great challenges to industrial enterprises.
[0003] With the wide application of big data in various fields, the use of industrial data has inevitably become one of the research directions of technical personnel in the industrial field. However, unlike other fields, in the industrial field, due to the lack of unified data standards and data management tools, the problems of non-uniform standards, low quality and difficult utilization of industrial data are highlighted, so that the industrial data does not fully play its value, and the industrial data has limitations in the application process.
[0004] Therefore, how to solve the problems of low data quality and difficult utilization of industrial data for the complex and diverse data collection and data management needs in the industrial field has become a problem to be solved. SUMMARY
[0005] In view of the above technical problems, the present application provides an industrial big data data management method, which can solve the problems of low data quality and difficult utilization of industrial data.
[0006] In order to solve the above technical problems, the present application adopts the following technical scheme:
[0007] An industrial big data data management method comprises the following steps:
[0008] S1, obtaining multi-source heterogeneous industrial data;
[0009] S2, cleaning and converting the data obtained in S1;
[0010] S3, auditing and improving the quality of the data according to the published data standards;
[0011] S4, opening the audited and improved data to the outside through the API sharing mode, and providing data application services.
[0012] Preferably, in S1, the obtained data includes device basic data, device running data and business data. The business data refers to product sales data, logistics data, production data and other business-related data.
[0013] Preferably, in S2, the data cleaning content includes repeated data processing, missing data processing and abnormal value processing.
[0014] Preferably, the data missing processing includes deletion method, replacement method and interpolation method; the deletion method includes deleting the observation with missing data when the proportion of missing observation is lower than a preset x value, or deleting the variable with missing data when the proportion of missing data of a variable is higher than a preset y value; the replacement method includes using mean or median to replace the missing value for continuous variable, and using mode to replace the missing value for discrete variable; the interpolation method includes predicting the missing value according to other non-missing variables or observations.
[0015] Preferably, the outlier processing includes deleting the outlier, not processing, mean value replacement or regarding as missing value.
[0016] Preferably, in S2, the data conversion is performed by script or ETL tool.
[0017] Preferably, S3 includes:
[0018] In the first step, data standard is prepared according to data requirement, data item is determined, and data attribute information of the data item is determined from data standard management execution group;
[0019] In the second step, the prepared data standard is reviewed to determine whether the data standard meets application and management requirement and whether the data standard meets data strategy requirement;
[0020] In the third step, after the data standard review passes, the data standard is released;
[0021] In the fourth step, after the data standard is released, data quality auditing task is constructed according to the data standard by using tool, data quality problem is comprehensively investigated, and specific data index is located;
[0022] In the fifth step, according to the data quality auditing result, the quality problem data is modified, and the data quality is improved.
[0023] Preferably, in S1, when the industrial data is acquired, the data acquisition speed of each device is also counted;
[0024] In S2, after the data is cleaned and converted, it is also determined whether the data amount of each device is lower than a preset data amount;
[0025] If the data amount of a device is lower than the preset amount, and the data acquisition speed of the device is less than a preset waiting value, then the data of the device after cleaning and conversion is proportionally expanded, so that the expanded data amount is greater than or equal to the preset data amount; and the expanded data is recorded;
[0026] After S4, it further includes:
[0027] S5, judging whether there is expansion data in the data of each device opened to the outside; if there is expansion data for a device, continuously obtaining actual data of the device, and performing cleaning and conversion to obtain precision adjustment data; and when the data amount of the precision adjustment data reaches a first percentage threshold and a second percentage threshold of the expansion data, respectively performing a precision adjustment, and after auditing and improving the quality of the data according to the published data standard, re-opening to the outside, wherein the first percentage threshold is 20-25%, and the second percentage threshold is 60-65%; and when the data amount of the precision adjustment data reaches the total amount of the expansion data, completely replacing the expansion data with the precision adjustment data, and after auditing and improving the quality of the data according to the published data standard, re-opening to the outside.
[0028] Preferably, the precision adjustment comprises: comparing the deviation degree of the expansion data and the precision adjustment data, if the deviation degree is less than or equal to a preset first deviation value, performing a preset correction; the preset correction comprises: replacing an equal amount of expansion data with precision adjustment data, and recording the remaining expansion data;
[0029] If the deviation degree is greater than the preset first deviation value and less than or equal to a preset second deviation value, deleting all the expansion data, and performing equal proportion expansion with the data obtained at the beginning and the precision adjustment data to obtain corrected expansion data, so that the data amount after expansion is greater than or equal to a preset data amount, and recording the expansion data at this time; wherein the second deviation value is greater than the first deviation value;
[0030] If the deviation degree is greater than the second deviation value, withdrawing the data opened to the outside by the device; after the total amount of the data obtained at the beginning and the precision adjustment data reaches a preset data amount, performing processing according to the steps of S3-S4, and re-opening to the outside.
[0031] Preferably, in S4, when opening the data of each device, the data of each device is also marked with a precision level; if expansion data is used and precision adjustment is not performed, the precision level is weak; after one precision adjustment, the precision level is relatively weak; after two precision adjustments, the precision level is relatively strong; and if expansion data is not used, the precision level is strong.
[0032] Compared with the prior art, the present application has the following beneficial effects:
[0033] 1. Using the method, after obtaining industrial data and performing data cleaning and conversion, the quality of the data is audited and improved according to the published data standard, so that the effectiveness, format consistency and applicability of the data can be ensured. Then, the audited and improved data is opened to the outside through API sharing to provide data application services. In this way, the audited and improved data can be fully utilized. Thus, the problems of low data quality and difficult data utilization in the industrial field are solved.
[0034] In summary, the present application can solve the problems of low data quality and difficulty in utilizing industrial data.
[0035] 2. The method can reduce the waste of manpower caused by manual data quality checking.
[0036] 3. If the data volume of a certain device is insufficient and the data acquisition speed is low, the method will expand the data to make it available for sufficient data volume immediately, and provide an accuracy level of one when opened to the outside, so as to meet the usage needs of data users who do not have high requirements for data accuracy as soon as possible. Then, in order to meet the usage needs of as many data users as possible, the present application will continuously acquire the actual data of the device, and clean and convert the data to obtain precision adjustment data, and when the data volume of the precision adjustment data reaches the first and second percentage threshold values of the expanded data, the precision adjustment data is adjusted once respectively. By comparing with the preset first and second deviation values, the accuracy of the expanded data is understood, so that the expanded data is adjusted in accuracy accordingly. When the data volume of the precision adjustment data reaches the total amount of the expanded data, the expanded data is completely replaced by the precision adjustment data, and after the quality of the data is audited and improved according to the published data standard, the data is re-opened to the outside. Thus, the accuracy of the data opened to the outside is gradually improved. Moreover, when opening the data of each device, the present application also identifies the accuracy level of the data of each device. Data users can choose whether to use the corresponding data immediately or to use it after the accuracy of the corresponding data is improved according to their actual needs. In this way, the needs of various types of data users can be met in a timely manner. BRIEF DESCRIPTION OF DRAWINGS
[0037] In order to make the purpose, technical scheme and advantages of the application more clear, the application will be further described in detail below with reference to the drawings, in which:
[0038] Figure 1 is a flowchart of the present application;
[0039] Figure 2 is a flowchart of the second embodiment of the present application. DETAILED DESCRIPTION
[0040] The application will be further described in detail below through specific embodiments:
[0041] Embodiment One
[0042] As shown in the figure, the present embodiment discloses an industrial big data data governance method, which comprises the following steps: Figure 1
[0043] S1, acquire multi-source heterogeneous industrial data; in specific implementation, the acquired data includes device basic data, device running data and business data. The business data refers to product sales data, logistics data, production data and other business-related data.
[0044] S2, clean and convert the data acquired in S1.
[0045] In specific implementation, the data cleaning content includes repeated data processing, missing data processing and abnormal value processing.
[0046] Among them, the missing data processing includes deletion method, replacement method and interpolation method; the deletion method includes deleting the observation with missing data when the missing observation proportion is lower than the preset x value (such as within 10%), or deleting the variable with missing data when the missing proportion of a variable is higher than the preset y value (such as more than 90%); the replacement method includes using the mean or median to replace the missing value for continuous variables, and using the mode to replace the missing value for discrete variables; the interpolation method includes predicting the missing value according to other non-missing variables or observations; in specific implementation, the interpolation method can adopt regression interpolation method, K-nearest neighbor interpolation method or Lagrange interpolation method.
[0047] The abnormal value processing includes deleting abnormal value, not processing, average value substitution or regarding as missing value. The abnormal value refers to the data far away from the normal value, i.e. the "out-group" data. ① Delete abnormal value - directly delete if obviously abnormal and the number is small; ② Do not process - if the algorithm is not sensitive to abnormal values, it can be processed, but if the algorithm is sensitive to abnormal values, it is best not to use this method, such as some algorithms based on distance calculation, including kmeans, knn, etc.; ③ Average value substitution - small loss of information, simple and efficient; ④ Regard as missing value - can be processed according to the method of processing missing values.
[0048] When data conversion is performed, data conversion is performed through scripts or ETL tools. Among them, scripts: use SQL or Python to perform data conversion through scripts, write code to extract and convert data. ETL tool: ETL (extraction, transformation, loading) tool can complete most of the pain of script conversion through automatic process.
[0049] S3, according to the published data standard, the quality of data is audited and improved;
[0050] In specific implementation, S3 includes:
[0051] First, according to the data requirement, compile the data standard, determine the data item, and determine the data attribute information of the data item from the data standard management execution group;
[0052] Secondly, the prepared data standard is reviewed to determine whether the data standard meets the application and management requirements and whether the data standard meets the data strategy requirements;
[0053] Thirdly, after the data standard review is passed, the data standard is released;
[0054] Fourthly, after the data standard is released, the data quality audit task is constructed according to the data standard through the tool, and the data quality problem is fully investigated and the specific data index is located;
[0055] Fifthly, according to the data quality audit result, the quality problem data is modified, and the data quality is improved.
[0056] S4, the audited and improved data is opened to the outside through API sharing, and data application service is provided.
[0057] The data after governance is analyzed and utilized, including data analysis, data cockpit construction, and predictive maintenance of equipment. After data collection and governance are completed, data application service capability is constructed. Various business statistical analysis indexes are uniformly managed, relying on rich algorithm support, visual data model design service is provided to meet the development needs of various front-end applications, and data subscription and application access service are provided to the outside through API sharing, and API is managed and supported throughout the process.
[0058] Using the method, after obtaining the industrial data and performing data cleaning and conversion, the quality of the data is audited and improved according to the published data standard, so that the effectiveness, format consistency and applicability of the data can be guaranteed. Then, the audited and improved data is opened to the outside through API sharing, and data application service is provided. In this way, the audited and improved data can be fully utilized. Thus, the problem of low data quality and difficult data utilization in the industrial field is solved. In addition, the method can reduce the waste of manpower caused by manual data quality checking.
[0059] Embodiment Two
[0060] Because the working frequency, single working time, maintenance time, and equipment quantity of different equipment are different, the data acquisition speed of each equipment is also different. When the data of a certain equipment is utilized according to the method of embodiment one, the data quantity of the equipment may be too small to meet the use requirements, and if the data acquisition speed of the equipment is slow, it will take a very long time to wait until the data quantity of the equipment meets the use requirements. On the other hand, different enterprises have different data use purposes, and their requirements for data accuracy are also different. Some only need to be roughly accurate, and some need to be very accurate.
[0061] Therefore, in order to meet the needs of various types of data users in a timely manner, unlike in Embodiment 1, in S1 of this embodiment, when acquiring industrial data, the acquisition speed of data from each device is also statistically analyzed.
[0062] In S2, after cleaning and transforming the data, it is also determined whether the data volume of each device is lower than the preset data volume.
[0063] If the data volume of a certain device is lower than the preset amount, and the data acquisition speed of the device is lower than the preset waiting value, then the data cleaned and converted by the device will be expanded proportionally so that the expanded data volume is greater than or equal to the preset data volume; and the expanded data will be recorded.
[0064] like Figure 2 As shown, after S4, it also includes:
[0065] S5. Determine whether there is supplementary data in the data of each device that is open to the public. If there is supplementary data for a certain device, continuously acquire the actual data of that device, clean and transform it to obtain precision-adjusted data. When the amount of precision-adjusted data reaches the first and second percentage thresholds of the supplementary data, perform a precision adjustment once each time. After auditing and improving the quality of the data according to the published data standards, the data is reopened to the public. The first percentage threshold is 20-25%, and the second percentage threshold is 60-65%. When the amount of precision-adjusted data reaches the total amount of supplementary data, completely replace the supplementary data with precision-adjusted data. After auditing and improving the quality of the data according to the published data standards, the data is reopened to the public.
[0066] The precision adjustment includes: comparing the deviation between the expanded data and the precision adjustment data; if the deviation is less than or equal to a preset first deviation value, then a preset correction is performed; the preset correction includes replacing an equal amount of expanded data with the precision adjustment data and recording the remaining expanded data.
[0067] If the deviation is greater than the preset first deviation value and less than or equal to the preset second deviation value, then all the expanded data is deleted, and the data acquired at the beginning and the precision adjustment data are expanded proportionally to obtain the corrected expanded data, so that the amount of expanded data is greater than or equal to the preset data amount, and the expanded data at this time is recorded; wherein, the second deviation value is greater than the first deviation value.
[0068] If the deviation exceeds the second deviation value, the data released by the device will be withdrawn. After the total amount of data acquired and the precision adjustment data reaches the preset data amount, the data will be processed according to steps S3-S4 and then released to the public again.
[0069] In S4, the data of each device is also marked with precision level when the data is opened; the precision level is weak when the extended data is used without precision adjustment; the precision level is relatively weak when the precision is adjusted once; the precision level is relatively strong when the precision is adjusted twice; and the precision level is strong when the extended data is not used.
[0070] Using the method, if the data amount of a device is insufficient and the data acquisition speed is low, the method can extend the data so that it can be used immediately with sufficient data amount, and the precision level is marked when the data is opened, so that the use demand of the data user with low data precision requirement can be met as soon as possible. Then, in order to meet the use demand of as many data users as possible, the actual data of the device is continuously acquired, cleaned and converted to obtain the precision adjustment data, and when the data amount of the precision adjustment data reaches the first percentage threshold and the second percentage threshold of the extended data, the precision is adjusted once respectively. By comparing with the preset first deviation value and the second deviation value, the precision of the extended data is known, so that the extended data is adjusted in precision accordingly. When the data amount of the precision adjustment data reaches the total amount of the extended data, the extended data is completely replaced by the precision adjustment data, and after the quality of the data is audited and improved according to the published data standard, the data is opened again. Thus, the precision of the opened data is gradually improved.
[0071] Moreover, when the data of each device is opened, the data of each device is also marked with precision level. The data user can choose whether to use the corresponding data immediately or to use the data after the precision of the corresponding data is improved according to the actual demand. In this way, the demand of each type of data user can be met in time.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application but not to limit the technical solutions. Those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the purpose and scope of the technical solutions, which should be covered in the scope of the claims of the present application.
Claims
1. A data governance method based on industrial big data, characterized in that, Includes the following steps: S1. Acquire multi-source heterogeneous industrial data; S2. Clean and transform the data obtained in S1; S3. Audit and improve the quality of the data according to the published data standards; S4. Share audited and improved data with external parties through APIs to provide data application services; In S1, when acquiring industrial data, the acquisition speed of each device is also statistically analyzed. In S2, after cleaning and transforming the data, it is also determined whether the data volume of each device is lower than the preset data volume. If the data volume of a certain device is lower than the preset amount, and the data acquisition speed of the device is lower than the preset waiting value, then the data cleaned and converted by the device will be expanded proportionally so that the expanded data volume is greater than or equal to the preset data volume; and the expanded data will be recorded. Following S4 are: S5. Determine whether there is any supplementary data in the data of each device that is publicly accessible. If supplementary data exists for a certain device, continuously acquire the actual data of that device, clean and transform it to obtain precision-adjusted data. When the amount of precision-adjusted data reaches the first and second percentage thresholds of the supplementary data, perform a precision adjustment once each time. After auditing and improving the data quality according to the published data standards, the data is then reopened to the public. The first percentage threshold is 20-25%, and the second percentage threshold is 60-65%. When the amount of precision-adjusted data reaches the total amount of supplementary data, completely replace the supplementary data with precision-adjusted data. After auditing and improving the data quality according to the published data standards, the data is then reopened to the public. The precision adjustment includes: comparing the deviation between the expanded data and the precision adjustment data; if the deviation is less than or equal to a preset first deviation value, then a preset correction is performed; the preset correction includes replacing an equal amount of expanded data with precision adjustment data and recording the remaining expanded data. If the deviation is greater than the preset first deviation value and less than or equal to the preset second deviation value, then all the expanded data is deleted, and the data acquired at the beginning and the precision adjustment data are expanded proportionally to obtain the corrected expanded data, so that the amount of expanded data is greater than or equal to the preset data amount, and the expanded data at this time is recorded; wherein, the second deviation value is greater than the first deviation value. If the deviation exceeds the second deviation value, the data released by the device will be withdrawn. After the total amount of data acquired and the precision adjustment data reaches the preset data amount, the data will be processed according to steps S3-S4 and then released to the public again.
2. The data governance method based on industrial big data as described in claim 1, characterized in that: In S1, the acquired data includes basic equipment data, equipment operation data, and business data.
3. The data governance method based on industrial big data as described in claim 1, characterized in that: In S2, data cleaning includes handling duplicate data, missing data, and outlier handling.
4. The industrial big data governance method as described in claim 3, characterized in that: Data missing value handling includes deletion, replacement, and imputation. Deletion methods include deleting missing observations when the proportion of missing observations is lower than a preset x-value, or deleting a missing variable when the proportion of missing observations is higher than a preset y-value. Replacement methods include replacing missing values with the mean or median for continuous variables, and replacing missing values with the mode for discrete variables. Imputation methods include predicting missing values based on other non-missing variables or observations.
5. The data governance method based on industrial big data as described in claim 3, characterized in that: Outlier handling includes deleting outliers, not handling them, replacing them with the average value, or treating them as missing values.
6. The data governance method based on industrial big data as described in claim 1, characterized in that: In S2, data transformation is performed using scripts or ETL tools.
7. The data governance method based on industrial big data as described in claim 1, characterized in that: S3 includes: The first step is to develop data standards based on data requirements, identify data items, and determine the data attribute information of the data items from the data standard management and execution group. The second step is to review the data standards to determine whether they meet application and management needs and whether they comply with data strategy requirements. The third step is to publish the data standards after they have been reviewed and approved. The fourth step is to use tools to build data quality audit tasks according to the data standards after the data standards are released, comprehensively investigate data quality issues, and identify specific data indicators. The fifth step is to modify the data with quality issues based on the data quality audit results to improve data quality.
8. The data governance method based on industrial big data as described in claim 1, characterized in that: In S4, when opening data from each device, the accuracy level of the data from each device is also identified; if extended data is used but no accuracy adjustment is made, the accuracy level is weak; if one accuracy adjustment is made, the accuracy level is relatively weak; if two accuracy adjustments are made, the accuracy level is relatively strong; if extended data is not used, the accuracy level is strong.
Citation Information
Patent Citations
Nuclear power industrial data warehouse system
CN114357088A