A method and device for automatic function mining based on industrial big data
By building a wide data table, performing data portraits and obtaining key data in industrial big data, and combining with symbol regression prediction models, the problem of low efficiency and accuracy of function mining in the industrial field is solved, and more efficient function mining is achieved.
Patent Information
- Application Number
- CN202111289512.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-02
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-11-02
AI Technical Summary
In the industrial field, the search space for function mining is huge, the search efficiency is extremely low, and the data interference is strong, making it difficult to guarantee efficiency and accuracy.
By obtaining industrial big data for preprocessing, building a data wide table, performing data portraits to obtain a weight-based function mapping set, and obtaining key data from the wide table based on the target parameters, inputting a symbol regression prediction model to automatically obtain function mining results.
It effectively narrows the search space of data, avoids the influence of interference factors, and improves the efficiency and accuracy of function mining.
Smart Images

Figure CN114238429B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data mining, and in particular to an automatic function mining method and device based on industrial big data. Background Art
[0002] In the industrial field, such as large-scale engineering machinery manufacturing (aircraft, ships, automobiles and large engineering machinery), energy production environment (mining, drilling), construction environment (such as the construction of large buildings), etc., highly complex machines, equipment, systems and highly complex workflows are involved. The emergence of the Internet of Things (IoT) enables people to continuously connect a wider range of devices and achieve data acquisition between these devices. However, in a highly complex industrial environment, devices generate a large amount of data, while the range of available data is often very limited, and processing data from multiple devices is very cumbersome and complex.
[0003] Function mining is a supervised learning method that uses symbolic regression algorithms to try to discover some hidden mathematical formulas, so as to use feature variables to predict target variables. Compared with traditional regression methods such as linear regression, polynomial regression, artificial neural networks, etc., function mining does not require a specific function form in advance, does not require any prior knowledge and models, and can provide an intuitive and explicit function expression model, which helps researchers understand and analyze the internal mechanism of the system to be studied. Therefore, function mining has been applied to fields such as time series prediction, data mining, pattern classification, and system design optimization in recent years, and has become a research hotspot in the field of intelligent computing. However, due to the complexity and high dimensionality of industrial big data, the search space of function mining is very large, the search efficiency is extremely low, and the data interference is strong, so the efficiency and accuracy of function mining cannot be guaranteed. Summary of the invention
[0004] The present invention provides a method and device for automatic function mining based on industrial big data, which are used to solve the defects of low efficiency and accuracy of function mining in the prior art and effectively improve the efficiency and accuracy of function mining.
[0005] The present invention provides a method for automatic function mining based on industrial big data, comprising:
[0006] Acquire industrial big data and preprocess the industrial big data;
[0007] Based on the preprocessed industrial big data, a data wide table is obtained; wherein the data wide table is used to store data to be mined for functions;
[0008] Profiling the data in the wide table to obtain data distribution of the data in the wide table, and obtaining a weight-based function mapping set according to the data distribution;
[0009] Based on a target parameter, key data is obtained from the wide table; wherein the target parameter is a dependent variable of the function to be mined;
[0010] The function mapping set and the key data are input into a symbolic regression prediction model to automatically obtain the function mining results of the industrial big data.
[0011] According to a method for automatic function mining based on industrial big data provided by the present invention, obtaining a data wide table based on the preprocessed industrial big data includes:
[0012] Based on the preprocessed industrial big data, a logical model and a physical model are constructed; wherein the logical model is used to indicate the format of the storage table of the industrial big data; and the physical model is used to indicate the data entities in the industrial big data to be stored in the storage table;
[0013] Based on the preprocessed industrial big data, a mapping relationship between the business model and the data model is established, and based on the mapping relationship, the data wide table is obtained from the storage table; wherein the business model is used to indicate the parameters corresponding to each work task in the industrial big data, and the relationship between each parameter; the data model is used to indicate the parameter values of the parameters.
[0014] According to a method for automatic function mining based on industrial big data provided by the present invention, after obtaining the industrial big data, the method further includes:
[0015] The industrial big data is stored in a data warehouse; wherein the data warehouse performs hierarchical storage of the industrial big data based on different storage granularities.
[0016] According to an automatic function mining method based on industrial big data provided by the present invention, the function mapping set based on weights is obtained according to the data distribution, including:
[0017] The data distribution is visualized, and based on a preset function set and the visualization result of the data distribution, a weight-based function mapping set is obtained from the preset function set through a deep learning model.
[0018] According to an automatic function mining method based on industrial big data provided by the present invention, the function mapping set based on weights is obtained from the preset function set by using a deep learning model, including:
[0019] The preset function set and the visualization display results of the data distribution are input into the deep learning model, and the visualization display results of the data distribution and each candidate function in the preset function set are subjected to curve or surface fitting by the deep learning model to obtain the degree of fit between the visualization display results of the data distribution and each candidate function, and the weight-based function mapping set and the weight of each function in the function mapping set are obtained according to the degree of fit.
[0020] According to an automatic function mining method based on industrial big data provided by the present invention, the step of obtaining key data from the wide table based on target parameters includes:
[0021] Based on the target parameter, correlation calculation is performed on the data in the wide table to obtain correlation values between each parameter in the wide table and the target parameter, key parameters are obtained based on the correlation values, and key data are obtained from the wide table based on the target parameter and the key parameters; wherein the key parameter is an independent variable of the function to be mined.
[0022] According to a method for automatic function mining based on industrial big data provided by the present invention, before calculating the correlation of the data in the wide table based on the target parameter, the method further includes:
[0023] The data in the wide table are sequentially subjected to standardization and encoding processing.
[0024] The present invention also provides a function automatic mining device based on industrial big data, comprising:
[0025] A data acquisition module, used to acquire industrial big data and pre-process the industrial big data;
[0026] A wide table extraction module, used to obtain a data wide table based on the preprocessed industrial big data; wherein the data wide table is used to store data to be mined for functions;
[0027] A data profiling module, used to perform data profiling on the data in the wide table, obtain data distribution of the data in the wide table, and obtain a weight-based function mapping set according to the data distribution;
[0028] A key data extraction module, used to obtain key data from the wide table based on a target parameter; wherein the target parameter is a dependent variable of the function to be mined;
[0029] The model prediction module is used to input the function mapping set and the key data into the symbolic regression prediction model to automatically obtain the function mining results of the industrial big data.
[0030] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of any of the above-mentioned automatic function mining methods based on industrial big data are implemented.
[0031] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described automatic function mining methods based on industrial big data.
[0032] The automatic function mining method and device based on industrial big data provided by the present invention obtain a wide data table based on preprocessed industrial big data, obtain a weighted function mapping set by performing data profiling on the data in the wide table, select key data from the wide table according to target parameters, and input the weighted function mapping set and key data into a symbolic regression prediction model to automatically obtain function mining results, which can effectively narrow the data search space, avoid the influence of interference factors, and ensure the efficiency and accuracy of function mining. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0034] Figure 1 It is a flow chart of the automatic function mining method based on industrial big data provided by the present invention;
[0035] Figure 2 It is a structural schematic diagram of a function automatic mining device based on industrial big data provided by the present invention;
[0036] Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0038] Function mining is a supervised learning method that uses symbolic regression algorithms to try to discover some hidden mathematical formulas, so as to use feature variables to predict target variables. Compared with traditional regression methods such as linear regression, polynomial regression, artificial neural networks, etc., function mining does not require a specific function form in advance, does not require any prior knowledge and models, and can provide an intuitive and explicit function expression model, which helps researchers understand and analyze the internal mechanism of the system to be studied. Therefore, function mining has been applied to fields such as time series prediction, data mining, pattern classification, and system design optimization in recent years, and has become a research hotspot in the field of intelligent computing. However, in the industrial field, in a highly complex industrial environment, equipment will generate a large amount of data, while the range of available data is often very limited, and processing data from multiple devices is very cumbersome and complicated, resulting in a very large search space for function mining, extremely low search efficiency, and strong data interference, which makes it impossible to guarantee the efficiency and accuracy of function mining.
[0039] To this end, the present invention provides an automatic function mining method based on industrial big data. Figure 1 It is a flow chart of the automatic function mining method based on industrial big data of the present invention. Figure 1 As shown, the method comprises the following steps:
[0040] S100: Acquire industrial big data and pre-process the industrial big data.
[0041] Specifically, industrial big data refers to a large amount of data generated at high speed by various industrial equipment in the Internet of Things, corresponding to the equipment status at different times. There are many ways to obtain industrial big data. For example, a large amount of data generated by industrial equipment and machine sensors can be collected in real time through flink, and the collected industrial big data can be merged according to the order of processing batches, and different data sources can be stored in a unified data storage platform. Among them, flink is an open source stream processing framework that executes any stream data program in a data parallel and pipeline manner. When the flink pipeline is running, it can execute batch and stream processing programs. The data storage platform can be selected according to actual needs. For example, HDFS (Hadoop Distributed File System) is used, which has high fault tolerance and can provide high-throughput data access, which is very suitable for applications on large-scale data sets. Hadoop is a distributed system basic framework. HDFS can also compress the stored data to reduce the storage space occupied.
[0042] In addition, in an industrial environment, equipment will generate a large amount of data, and in the process of collecting industrial big data, the collected data is greatly affected by the environment or equipment status, and there are many outliers and / or missing values. Therefore, before function mining, the collected industrial big data needs to be preprocessed. The specific method of data preprocessing can be set according to actual needs, for example, outlier removal and missing value filling to ensure the accuracy and effectiveness of industrial big data. At the same time, data preprocessing can also include data sampling. Through data sampling, the amount of data can be effectively reduced while ensuring the comprehensiveness and effectiveness of the data.
[0043] S200. Based on the preprocessed industrial big data, a data wide table is obtained; wherein the data wide table is used to store data to be mined for functions.
[0044] Specifically, since industrial big data includes data from a variety of different industrial equipment and machine sensors in the Internet of Things, the data is extremely cumbersome and complex, and the data volume is extremely large and has strong interference. In the function mining process, it is difficult to quickly and accurately find useful data associated with the function mining target. The embodiment of the present invention obtains a data wide table based on pre-processed industrial big data, and uses the data wide table to store data to be mined, that is, data related to work tasks, so that in the process of function mining, useful data related to the target can be quickly found, the search space of function mining can be reduced, the accuracy and effectiveness of data can be guaranteed, and a technical basis for improving the accuracy and efficiency of function mining is provided.
[0045] S300: Profiling the data in the wide table to obtain the data distribution of the data in the wide table, and obtaining a weight-based function mapping set according to the data distribution.
[0046] Specifically, a method for performing data profiling on the data in a wide table and obtaining the data distribution of the data in the wide table includes: using statistical methods to perform exploratory analysis on the data in the wide table to obtain the data distribution of the data in the wide table. Among them, exploratory analysis is used to obtain hidden relationships in the data, for example, to explore the central tendency, degree of dispersion, and extreme values of the data in the wide table to obtain the distribution results of the data, such as the mean, median, variance, standard deviation, coefficient of dispersion, maximum value, minimum value, and total sample size, thereby obtaining a mapping relationship between the data model and the mathematical function based on the data distribution, and obtaining a weighted function set based on the mapping relationship, thereby effectively reducing the space of the function set required in function mining and further improving the efficiency of function mining.
[0047] S400, based on the target parameter, obtain key data from the wide table; wherein the target parameter is the dependent variable of the function to be mined.
[0048] Specifically, the data stored in the wide table is data associated with the function mining target, and because industrial big data is extremely cumbersome and complex, these high-dimensional data are often mixed with a large amount of redundant, noisy and irrelevant features and other interfering data, which not only brings great challenges to model learning, but also increases storage and computing costs, making it difficult to form an effective "intelligent" solution for the industrial sector. In order to solve this problem, the embodiment of the present invention obtains key data from the wide table through target parameters, which can effectively eliminate interfering data, further reduce the search space of function mining, and provide a technical basis for the accuracy and efficiency of function mining.
[0049] S500, inputting the function mapping set and key data into the symbolic regression prediction model to automatically obtain the function mining results of the industrial big data.
[0050] Specifically, the symbolic regression prediction model is a model used for function prediction. The model is built based on a neural network algorithm. For example, the symbolic regression prediction model uses the existing PySR and DSO frameworks. By inputting the function mapping set and key data into the symbolic regression prediction model for iterative search, functions with a relatively high degree of fit can be automatically mined.
[0051] For function mining, the two most important factors are the selection of feature variables and target functions. The embodiment of the present invention obtains data to be mined from industrial big data by constructing a wide table, and can effectively reduce irrelevant variables by extracting key data, eliminate irrelevant and redundant features to reduce the data input dimension, and select the most important features as the input of the symbolic regression prediction model; at the same time, a function mapping set with high similarity is obtained through data profiling, and the function mapping set is input into the symbolic regression prediction model according to the weight, thereby effectively narrowing the data search space, avoiding the influence of interference factors, and ensuring the efficiency and accuracy of function mining.
[0052] It can be seen that the embodiment of the present invention obtains a wide table of data based on preprocessed industrial big data, obtains a weighted function mapping set by performing data profiling on the data in the wide table, selects key data from the wide table according to target parameters, and inputs the weighted function mapping set and key data into the symbolic regression prediction model to automatically obtain the function mining results, which can effectively narrow the data search space, avoid the influence of interference factors, and ensure the efficiency and accuracy of function mining.
[0053] Based on the above embodiment, based on the preprocessed industrial big data, a data wide table is obtained, including:
[0054] Based on the preprocessed industrial big data, a logical model and a physical model are constructed; wherein the logical model is used to indicate the format of the storage table of the industrial big data; and the physical model is used to indicate the data entities in the industrial big data to be stored in the storage table;
[0055] Based on the preprocessed industrial big data, a mapping relationship between the business model and the data model is established, and based on the mapping relationship, the data wide table is obtained from the storage table; wherein, the business model is used to indicate the parameters corresponding to each work task in the industrial big data, as well as the relationship between the parameters; the data model is used to indicate the parameter values of the parameters.
[0056] Specifically, after preprocessing the industrial big data, a logical model and a physical model are constructed, wherein the logical model is used to indicate the format of the storage table of the industrial big data, for example, which storage tables are specifically used to store the industrial big data, as well as the specific format of each storage table, the fields included in the table, etc. The physical model is used to indicate the data entities in the industrial big data to be stored in the storage table, that is, according to the format of the storage table, the corresponding data in the industrial big data is stored in the storage table.
[0057] After storing the corresponding data in the industrial big data into the storage table, a mapping relationship between the business model and the data model is established. The business model is the parameters corresponding to each work task in the industrial big data, and the relationship between each parameter, such as the power and power consumption of the crane, and the relationship between power and power consumption; the data model is the specific value corresponding to each parameter in the work task. Through the mapping relationship between the business model and the data model, the data corresponding to the business model can be obtained from the storage table of the industrial big data, and a wide table can be built to store the data.
[0058] It can be seen that the embodiment of the present invention obtains a wide table from the storage table of industrial big data through the mapping relationship between the business model and the data model, and stores business-related data and the relationship between the data through the wide table. The relevant data of different businesses in the industrial big data can be clearly and completely stored independently, and then in the function mining process, the data related to the target can be quickly and accurately obtained according to the target, which reduces the search space of function mining, ensures the accuracy and effectiveness of the data, and provides a technical basis for improving the accuracy and efficiency of function mining.
[0059] Based on any of the above embodiments, after obtaining industrial big data, the method further includes:
[0060] Industrial big data is stored through a data warehouse, wherein the data warehouse performs hierarchical storage of industrial big data based on different storage granularities.
[0061] Specifically, a data warehouse is a subject-oriented, integrated, non-volatile and time-varying data set that can store industrial big data according to actual needs. Different storage granularities refer to the level of refinement or integration of data stored in the data unit of the data warehouse. The higher the refinement, the smaller the granularity level. On the contrary, the lower the refinement, the larger the granularity level. For example, in an embodiment of the present invention, industrial big data of different granularities are stored in layers through 4 different levels. The 4 different levels are: ODS (Operational Data Store, original data layer), DWD (DataWarehouse Detail, data detail layer), DWS (Data Warehouse Service, service data layer), ADS (Application Data Store, application data layer); among them, ODS is used to store original industrial big data, that is, data that has not been processed in any way, and only synchronizes the acquired industrial big data to hive, among which hive is a data warehouse tool used for data extraction, transformation, and loading, and can map structured data files to a database table. DWD is used to store data detail tables, that is, to store the most detailed data. It is a fine-grained aggregation of data for business topics, that is, a storage table stored according to a logical model. DWS is used to store light summary tables of tables with different topics for the same business, that is, to store all data of the same business in one large table. ADS is used to store wide tables, that is, to associate and merge data in light summary tables according to the business model to form wide tables.
[0062] It can be seen that the embodiment of the present invention stores industrial big data in layers through a data warehouse. During data processing or use, the corresponding data can be quickly and accurately acquired according to actual needs without traversing all data, which greatly reduces the data search space and improves the efficiency of function mining.
[0063] Based on any of the above embodiments, a weight-based function mapping set is obtained according to data distribution, including:
[0064] The data distribution is visualized. Based on the visualization results of the preset function set and data distribution, a weight-based function mapping set is obtained from the preset function set through a deep learning model.
[0065] Specifically, the data distribution is visualized, that is, the results of the data distribution are expressed in the form of graphics, such as box plots, bar charts, scatter plots, histograms, and pie charts. The visualization results of the data distribution and the preset function set are input into the deep learning model to obtain the similarity function of each group of data.
[0066] It can be seen that the embodiment of the present invention can pre-screen functions similar to the data through the visual display of data distribution results, that is, complete the rough screening of functions, thereby greatly reducing the search space of functions in the function mining process, providing a technical basis for improving the efficiency of function mining.
[0067] Based on any of the above embodiments, obtaining a weight-based function mapping set from a preset function set through a deep learning model includes:
[0068] The preset function set and the visualization results of the data distribution are input into the deep learning model. The deep learning model is used to perform curve or surface fitting on the visualization results of the data distribution and each candidate function in the preset function set to obtain the degree of fit between the visualization results of the data distribution and each candidate function. According to the degree of fit, a weight-based function mapping set and the weight of each function in the function mapping set are obtained.
[0069] Specifically, the deep learning model is trained using the TensorFlow framework. The higher the degree of fit, the closer the candidate function is to the visualization result of the data distribution. The candidate functions are sorted from high to low according to the degree of fit, and the candidate functions with higher degree of fit are selected according to the threshold to form a function mapping set. The weight of each function in the function mapping set is obtained based on the degree of fit. The higher the degree of fit, the greater the corresponding weight of the candidate function.
[0070] It can be seen that the embodiment of the present invention fits the visualization results of candidate functions and data distribution through a deep learning model, and can quickly and accurately screen out functions that are relatively close to the data distribution. At the same time, it effectively narrows the search space of the function in the function mining process and improves the function mining efficiency.
[0071] Based on any of the above embodiments, obtaining key data from the wide table based on the target parameter includes:
[0072] Based on the target parameters, the correlation of the data in the wide table is calculated to obtain the correlation values of each parameter in the wide table and the target parameters, the key parameters are obtained based on the correlation values, and the key data are obtained from the wide table based on the target parameters and the key parameters; wherein the key parameters are the independent variables of the function to be mined.
[0073] Specifically, the target parameter is the independent variable of the function to be mined. The correlation calculation is performed on the data in the wide table to obtain the correlation value between each parameter and the target parameter. The larger the correlation value, the higher the correlation between the parameter and the target parameter, and the greater the possibility of a functional relationship. Therefore, the correlation values of each parameter and the target parameter are sorted from high to low. According to the sorting results, the key parameters that may have a functional relationship with the target parameter can be obtained. The data corresponding to the target parameter and the key parameter are extracted from the wide table to obtain the key data. Among them, the method of correlation calculation can be selected according to actual needs, for example, PCA, Pearson, and Chi-square.
[0074] It can be seen that the embodiment of the present invention mines the correlation of features, removes irrelevant features, retains important features, and reduces feature dimensions by calculating the correlation between the data corresponding to each parameter in the wide table and the data corresponding to the target parameter, thereby realizing the screening of industrial big data, further reducing the amount of data involved in function mining, and ensuring the efficiency and accuracy of function mining.
[0075] Based on any of the above embodiments, before calculating the correlation of the data in the wide table based on the target parameter, the method further includes:
[0076] The data in the wide table are standardized and encoded in turn.
[0077] Specifically, since industrial big data includes data from different industrial equipment and machine sensors, the magnitude of the data varies greatly. Therefore, before performing correlation calculations, the data in the wide table is also standardized. Through standardization, the data of different magnitudes in the wide table are converted into dimensionless pure values, which facilitates the correlation calculation of data of different magnitudes. The standardization method can be selected according to actual needs, for example, normalization.
[0078] In addition, the data in the wide table is not necessarily a specific number, but may also be text, symbols, etc. Therefore, by encoding the data in the wide table, data of different formats can be mapped to a unified space to achieve the calculation of the correlation between different data. The encoding method can be selected according to actual needs. For example, one-hot encoding can be used to expand the value of discrete features to Euclidean space. A certain value of the discrete feature corresponds to a certain point in the Euclidean space, which greatly facilitates the calculation of the correlation between data.
[0079] It can be seen that the embodiment of the present invention maps data of different dimensions and formats into a unified space by standardizing and encoding the data in the wide table, thereby facilitating the calculation of data correlation, so that key data can be more accurately obtained through the correlation calculation results.
[0080] As a preferred implementation, the automatic function mining method based on industrial big data of the present invention can be implemented through a visualization system. The visualization system adopts a user UI graphical interface system, and its overall framework adopts a browser-server architecture, including a task management module, a data asset module, a data subscription module, a display module, an analysis report module, and a message push module; wherein the functions of the task management module, the data asset module, the data subscription module, the display module, the analysis report module, and the message push module are shown in Tables 1 to 6, respectively.
[0081] Table 1
[0082]
[0083] Table 2
[0084]
[0085] Table 3
[0086]
[0087] Table 4
[0088]
[0089] Table 5
[0090]
[0091] Table 6
[0092]
[0093] Through the implementation of the visualization system, the specific process of function mining can be displayed to users or R&D personnel in a fine-grained manner, supporting the R&D team to automatically process data.
[0094] The following is a description of the automatic function mining device based on industrial big data provided by the present invention. The automatic function mining device based on industrial big data described below and the automatic function mining method based on industrial big data described above can be referred to each other. Figure 2 As shown, the automatic function mining device based on industrial big data of the present invention includes:
[0095] The data acquisition module 210 is used to acquire industrial big data and pre-process the industrial big data;
[0096] The wide table extraction module 220 is used to obtain a data wide table based on the preprocessed industrial big data; wherein the data wide table is used to store data to be mined for functions;
[0097] A data profiling module 230 is used to perform data profiling on the data in the wide table, obtain data distribution of the data in the wide table, and obtain a weight-based function mapping set according to the data distribution;
[0098] The key data extraction module 240 is used to obtain key data from the wide table based on a target parameter; wherein the target parameter is a dependent variable of the function to be mined;
[0099] The model prediction module 250 is used to input the function mapping set and key data into the symbolic regression prediction model to automatically obtain the function mining results of industrial big data.
[0100] Based on the above embodiment, the wide table extraction module 220 establishes a mapping relationship between the business model and the data model based on the preprocessed industrial big data, and obtains the data wide table from the storage table based on the mapping relationship; wherein the business model is used to indicate the parameters corresponding to each work task in the industrial big data, as well as the relationship between each parameter; the data model is used to indicate the parameter value of the parameter.
[0101] Based on any of the above embodiments, it also includes a storage module, which is used to store industrial big data through a data warehouse; wherein the data warehouse performs hierarchical storage of industrial big data based on different storage granularities.
[0102] Based on any of the above embodiments, the data profiling module 230 is used to visualize the data distribution, and based on the preset function set and the visualization result of the data distribution, a weight-based function mapping set is obtained from the preset function set through a deep learning model.
[0103] Based on any of the above embodiments, the data profiling module 230 is used to input the preset function set and the visualization display results of the data distribution into the deep learning model, and perform curve or surface fitting on the visualization display results of the data distribution and each candidate function in the preset function set through the deep learning model to obtain the degree of fit between the visualization display results of the data distribution and each candidate function, and obtain a weight-based function mapping set and the weight of each function in the function mapping set based on the degree of fit.
[0104] Based on any of the above embodiments, the key data extraction module 240 performs correlation calculation on the data in the wide table based on the target parameters, obtains the correlation values of each parameter in the wide table and the target parameters, obtains the key parameters based on the correlation values, and obtains the key data from the wide table based on the target parameters and the key parameters; wherein the key parameters are the independent variables of the function to be mined.
[0105] Based on any of the above embodiments, it also includes a data processing module, which is used to perform standardization and encoding processing on the data in the wide table in sequence.
[0106] Figure 3 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 3 As shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330 and a communication bus 340, wherein the processor 310, the communication interface 320 and the memory 330 communicate with each other through the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute the automatic function mining method based on industrial big data, the method comprising: obtaining industrial big data and preprocessing the industrial big data;
[0107] Based on the preprocessed industrial big data, a data wide table is obtained; wherein the data wide table is used to store data to be mined for functions;
[0108] Profiling the data in the wide table to obtain the data distribution of the data in the wide table, and obtaining a weight-based function mapping set based on the data distribution;
[0109] Based on the target parameter, key data is obtained from the wide table; the target parameter is the dependent variable of the function to be mined;
[0110] The function mapping set and key data are input into the symbolic regression prediction model to automatically obtain the function mining results of industrial big data.
[0111] In addition, the logic instructions in the above-mentioned memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0112] On the other hand, the present invention further provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, when the program instructions are executed by a computer, the computer can execute the automatic function mining method based on industrial big data provided by the above methods, the method comprising: obtaining industrial big data, and preprocessing the industrial big data;
[0113] Based on the preprocessed industrial big data, a data wide table is obtained; wherein the data wide table is used to store data to be mined for functions;
[0114] Profiling the data in the wide table to obtain the data distribution of the data in the wide table, and obtaining a weight-based function mapping set based on the data distribution;
[0115] Based on the target parameter, key data is obtained from the wide table; the target parameter is the dependent variable of the function to be mined;
[0116] The function mapping set and key data are input into the symbolic regression prediction model to automatically obtain the function mining results of industrial big data.
[0117] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to perform the above-mentioned automatic function mining method based on industrial big data, the method comprising: obtaining industrial big data, and preprocessing the industrial big data;
[0118] Based on the preprocessed industrial big data, a data wide table is obtained; wherein the data wide table is used to store data to be mined for functions;
[0119] Profiling the data in the wide table to obtain the data distribution of the data in the wide table, and obtaining a weight-based function mapping set based on the data distribution;
[0120] Based on the target parameter, key data is obtained from the wide table; the target parameter is the dependent variable of the function to be mined;
[0121] The function mapping set and key data are input into the symbolic regression prediction model to automatically obtain the function mining results of industrial big data.
[0122] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0123] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for automatic function mining based on industrial big data, It is characterized in that include: Acquire industrial big data and preprocess the industrial big data; Based on the preprocessed industrial big data, a data wide table is obtained; wherein the data wide table is used to store data to be mined for functions; Profiling the data in the wide table to obtain data distribution of the data in the wide table, and obtaining a weight-based function mapping set according to the data distribution; Based on a target parameter, key data is obtained from the wide table; wherein the target parameter is a dependent variable of the function to be mined; Inputting the function mapping set and the key data into a symbolic regression prediction model to automatically obtain the function mining result of the industrial big data; Wherein, obtaining key data from the wide table based on the target parameter includes: Based on the target parameter, a correlation calculation is performed on the data in the wide table to obtain a correlation value between each parameter in the wide table and the target parameter; The correlation values of each parameter with the target parameter are sorted from high to low, and based on the sorting results, key parameters that may have a functional relationship with the target parameter are obtained; the data corresponding to the target parameter and the key parameter are extracted from the wide table to obtain the key data; wherein the key parameter is the independent variable of the function to be mined.
2. According to the method for automatic function mining based on industrial big data according to claim 1, It is characterized in that The step of obtaining a data wide table based on the pre-processed industrial big data includes: Based on the preprocessed industrial big data, a logical model and a physical model are constructed; wherein the logical model is used to indicate the format of the storage table of the industrial big data; and the physical model is used to indicate the data entities in the industrial big data to be stored in the storage table; Based on the preprocessed industrial big data, a mapping relationship between the business model and the data model is established, and based on the mapping relationship, the data wide table is obtained from the storage table; wherein the business model is used to indicate the parameters corresponding to each work task in the industrial big data, and the relationship between each parameter; the data model is used to indicate the parameter values of the parameters.
3. According to the method for automatic function mining based on industrial big data in claim 1, It is characterized in that After obtaining the industrial big data, the following steps are also included: The industrial big data is stored in a data warehouse; wherein the data warehouse performs hierarchical storage of the industrial big data based on different storage granularities.
4. According to the method for automatic function mining based on industrial big data in claim 1, It is characterized in that The step of obtaining a weight-based function mapping set according to the data distribution includes: The data distribution is visualized, and based on a preset function set and the visualization result of the data distribution, a weight-based function mapping set is obtained from the preset function set through a deep learning model.
5. According to the method for automatic function mining based on industrial big data according to claim 4, It is characterized in that The step of obtaining the weight-based function mapping set from the preset function set by using a deep learning model includes: The preset function set and the visualization display results of the data distribution are input into the deep learning model, and the visualization display results of the data distribution and each candidate function in the preset function set are subjected to curve or surface fitting by the deep learning model to obtain the degree of fit between the visualization display results of the data distribution and each candidate function, and the weight-based function mapping set and the weight of each function in the function mapping set are obtained according to the degree of fit.
6. According to the method for automatic function mining based on industrial big data in claim 1, It is characterized in that Before calculating the correlation of the data in the wide table based on the target parameter, the method further includes: The data in the wide table are sequentially subjected to standardization and encoding processing.
7. An automatic function mining device based on industrial big data, It is characterized in that include: A data acquisition module, used to acquire industrial big data and pre-process the industrial big data; A wide table extraction module, used to obtain a data wide table based on the preprocessed industrial big data; wherein the data wide table is used to store data to be mined for functions; A data profiling module, used to perform data profiling on the data in the wide table, obtain data distribution of the data in the wide table, and obtain a weight-based function mapping set according to the data distribution; A key data extraction module, used to obtain key data from the wide table based on a target parameter; wherein the target parameter is a dependent variable of the function to be mined; A model prediction module, used for inputting the function mapping set and the key data into a symbolic regression prediction model to automatically obtain the function mining result of the industrial big data; The key data extraction module is used to obtain key data from the wide table based on target parameters, including: Based on the target parameter, a correlation calculation is performed on the data in the wide table to obtain a correlation value between each parameter in the wide table and the target parameter; The correlation values of each parameter with the target parameter are sorted from high to low, and based on the sorting results, key parameters that may have a functional relationship with the target parameter are obtained; the data corresponding to the target parameter and the key parameter are extracted from the wide table to obtain the key data; wherein the key parameter is the independent variable of the function to be mined.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the steps of the automatic function mining method based on industrial big data as described in any one of claims 1 to 6 are implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the automatic function mining method based on industrial big data as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Big data modeling-based BI application system
CN111126852A
Business processing method applied to big data portrait pushing and machine learning server
CN112839102A