Industrial equipment data model generation method based on automatic machine learning
By collecting and integrating various data types from industrial equipment, and using automated machine learning methods to build models, the problem of poor adaptability of traditional models is solved, enabling efficient diagnosis and optimization in complex scenarios.
Patent Information
- Application Number
- CN202510992020.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-11-04
AI Technical Summary
Traditional machine learning-generated industrial equipment models are constructed using only a single dataset, resulting in poor adaptability and difficulty in meeting diagnostic needs in complex scenarios. They are also unable to effectively predict faults, detect anomalies, and optimize processes.
We collect and integrate time-series data, image data, and text data from industrial equipment, and build models using automated machine learning methods, including data preprocessing, feature extraction, model matching, and hyperparameter tuning, to ensure the applicability and robustness of the models during training, validation, and testing.
It improves the model's adaptability to complex scenarios, enhances the accuracy of fault prediction, anomaly detection, and process optimization, and ensures the model's stability and applicability on unknown data.
Smart Images

Figure CN120892811A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically, to a method for generating industrial equipment data models based on automated machine learning. Background Technology
[0002] Industrial equipment refers to industrial production equipment and various machine tools, such as lathes, milling machines, grinding machines, and planers. Automated Machine Learning (AutoML) is a technological system that lowers the barrier to entry for machine learning applications by automating feature engineering, model selection, and hyperparameter tuning. The core technologies of AutoML encompass neural architecture search, Bayesian optimization, and evolutionary algorithms, enabling even non-expert users to quickly build high-performance models.
[0003] Currently, enterprises utilize industrial equipment to continuously monitor its status, images, and audio data during production processes. After acquiring this data, enterprises then construct data models for the industrial equipment. This model building process requires engineers to perform highly complex machine learning modeling work, including data processing, model selection, and parameter tuning. Artificial intelligence or big data technologies can be used to conduct industrial data analysis and business modeling to achieve intelligent equipment operation and maintenance. However, industrial data is characterized by multimodality, high throughput, and high noise levels. Traditional machine learning models, built using only single types of industrial data, have poor adaptability and struggle to meet the diagnostic needs of industrial equipment in complex scenarios. They are unsuitable for applications such as fault prediction, anomaly detection, and process optimization.
[0004] Therefore, it is necessary to design an automatic machine learning-based method for generating industrial equipment data models to address the problems existing in current technologies. Summary of the Invention
[0005] In view of this, the present invention proposes an automatic machine learning-based method for generating industrial equipment data models, which aims to solve the problem that traditional machine learning-generated models are constructed using only a single industrial data source, resulting in poor model adaptability.
[0006] In one aspect, the present invention proposes a method for generating industrial equipment data models based on automated machine learning, comprising:
[0007] Collect time-series data, image data, and text data of industrial equipment during operation;
[0008] The time-series data, image data, and text data are fused together to obtain fused data;
[0009] Based on the fused data, a preliminary learning model was confirmed;
[0010] The fused data is categorized to identify the training set, validation set, and test set.
[0011] The preliminary learning model is trained and validated using the training set and validation set, respectively, to obtain a qualified preliminary model of device data.
[0012] After the preliminary equipment data model is tested and approved using the test set, the industrial equipment data model is obtained.
[0013] Furthermore, the steps for collecting time-series data, image data, and text data during the operation of industrial equipment include:
[0014] Confirm that all sensors used by industrial equipment to acquire timing data, image data, and text data share the same clock reference during operation;
[0015] After all sensors share the same clock reference, preliminary time-series data, preliminary image data, and preliminary text data are collected during the operation of industrial equipment.
[0016] The preliminary time-series data, preliminary image data, and preliminary text data are timestamped to obtain the time-series data, image data, and text data of the industrial equipment during operation.
[0017] Furthermore, the step of fusing the time-series data, image data, and text data to obtain fused data includes:
[0018] After preprocessing the time-series data, image data, and text data respectively, the fusion features of the time-series data, image data, and text data are extracted.
[0019] The fusion features of the time-series data, the image data, and the text data are concatenated and fused to obtain fused data.
[0020] Further, the steps of preprocessing the time-series data, image data, and text data respectively, and then extracting the fusion features of the time-series data, image data, and text data include:
[0021] The time series data is denoised and standardized to obtain the processed time series data;
[0022] The image data is subjected to noise reduction and enhancement processing to obtain the processed image data;
[0023] The text data is cleaned and vectorized to obtain the processed text data;
[0024] The processed time-series data, processed image data, and processed text data were extracted separately to obtain the fusion features of the time-series data, the fusion features of the image data, and the fusion features of the text data.
[0025] Furthermore, based on the fused data, the steps for confirming the preliminary learning model include:
[0026] Feature extraction is performed on the fused data to obtain feature ratio parameters;
[0027] The feature ratio parameters are matched with the model in the preset feature ratio scheme to confirm the adapted model.
[0028] Based on the adapted model, the initial learning model was confirmed.
[0029] Furthermore, based on the adapted model, the steps for confirming the initial learning model include:
[0030] The adapted model is then automatically tuned to obtain the optimized model.
[0031] The optimized model was validated using industrial rule alignment, and the validation results were obtained.
[0032] After the verification results are satisfactory, the optimized model will be used as the initial learning model.
[0033] Further, the steps of training and validating the preliminary learning model using the training set and validation set respectively to obtain a valid preliminary model of device data include:
[0034] The initial learning model is trained using the training set to obtain the trained learning model;
[0035] The trained learning model is validated using a validation set to obtain validation coefficients;
[0036] The verification coefficient is compared with the preset verification threshold to obtain a preliminary model of the equipment data that has passed verification.
[0037] Further, the step of comparing the verification coefficient with a preset verification threshold to obtain a preliminary model of the verified equipment data includes:
[0038] The verification coefficient is compared with a preset verification threshold. If the verification coefficient is lower than the preset verification threshold, the ratio of the training set, verification set and test set of the fused data is reconfirmed.
[0039] The verification coefficient is compared with the preset verification threshold. If the verification coefficient is equal to or greater than the preset verification threshold, a preliminary model of the equipment data that has passed verification is obtained.
[0040] Further, the step of comparing the verification coefficient with a preset verification threshold, and if the verification coefficient is lower than the preset verification threshold, then reconfirming the ratio of the training set, verification set, and test set of the fused data includes:
[0041] The verification coefficient is compared with the preset verification threshold. If the verification coefficient is between the preset verification threshold and 0.9 times the preset verification threshold, the ratio of the training set, verification set and test set of the fused data is adjusted to 8:1:1.
[0042] The verification coefficient is compared with the preset verification threshold. If the verification coefficient is less than 0.9 times the preset verification threshold, the amount of fused data is increased by at least 20%, and the ratio of the training set, verification set and test set of the increased fused data is adjusted to 8:1:1.
[0043] Furthermore, after the preliminary equipment data model is tested and approved using the test set, the steps to obtain the industrial equipment data model include:
[0044] The preliminary model of the device data was tested using the test set, and the test results were obtained.
[0045] The test results are compared with a preset test threshold. If the test results are greater than the preset test threshold, the preliminary model of the equipment data that passes the test is used as the industrial equipment data model.
[0046] Compared with existing technologies, the beneficial effects of this invention are as follows: By collecting time-series data, image data, and text data from industrial equipment during operation, it can comprehensively reflect the state of the industrial equipment. The time-series data, image data, and text data are fused to obtain fused data. By fusing different data, the fused data provides more comprehensive information, eliminates data redundancy, and mines cross-modal correlation features. Based on the fused data, a preliminary learning model is confirmed. The fused data is classified to confirm the training set, validation set, and test set. By classifying the data, the objectivity and robustness of the model evaluation are ensured. The preliminary learning model is trained and validated separately using the training set and validation set to obtain a qualified preliminary model of equipment data. Not only is overfitting prevented by training and validating the preliminary learning model separately using the training set and validation set, but the stability of the model on unknown data is also ensured. Furthermore, after the preliminary equipment data model is tested and approved using the test set, an industrial equipment data model is obtained. This verifies the applicability of the industrial equipment data model in the actual industrial environment, thereby improving the adaptability of the industrial equipment data model. At the same time, it meets the requirements for diagnosing industrial equipment in complex scenarios, improving the accuracy of fault prediction, anomaly detection, and process optimization of industrial equipment. Attached Figure Description
[0047] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0048] Figure 1 A flowchart illustrating the method for generating industrial equipment data models based on automatic machine learning, provided in an embodiment of the present invention.
[0049] Figure 2 This is a flowchart of step S100 in the automatic machine learning-based industrial equipment data model generation method provided in an embodiment of the present invention.
[0050] Figure 3 This is a flowchart of step S200 in the method for generating industrial equipment data models based on automatic machine learning provided in an embodiment of the present invention.
[0051] Figure 4 This is a flowchart of step S300 in the automatic machine learning-based industrial equipment data model generation method provided in an embodiment of the present invention.
[0052] Figure 5 The flowchart is shown in step S500 of the method for generating industrial equipment data models based on automatic machine learning provided in the embodiments of the present invention. Detailed Implementation
[0053] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0054] See Figure 1 As shown, this application provides a method for generating industrial equipment data models based on automatic machine learning, including the following steps:
[0055] S100: Collects time-series data, image data, and text data from industrial equipment during operation. Time-series data is data recorded sequentially over time during equipment operation, such as data collected by sensors measuring temperature, pressure, and speed. Image data includes photographs or videos of the equipment in operation, such as monitoring footage or infrared thermal imaging data. Text data includes operation logs, maintenance records, and fault reports. By collecting these data, a comprehensive picture of the industrial equipment's status can be obtained, enabling the generated industrial equipment data model to meet the needs of complex scenarios.
[0056] S200: The time-series data, image data, and text data are fused to obtain fused data. In this embodiment, feature extraction is performed on the time-series data, image data, and text data respectively before fusion processing, which eliminates data redundancy and can also uncover cross-modal correlation features, such as the correlation between vibration anomalies and bearing cracks.
[0057] S300: Based on the fused data, a preliminary learning model is confirmed. By fusing the data and considering its characteristics, AutoML is used to select a model from a candidate model library, thereby confirming the preliminary learning model. This constructs an adaptive learning framework capable of handling multimodal data, providing an initial structure for subsequent optimization.
[0058] S400: Classify the fused data to identify the training, validation, and test sets. Specifically, the ratio of the training, validation, and test sets should be 7:1.5:1.5. This ensures the objectivity and robustness of the model evaluation.
[0059] S500: The preliminary learning model is trained and validated using the training set and validation set respectively to obtain a qualified preliminary model of device data. In this embodiment, the preliminary learning model is trained using the training set and validated using the validation set. AutoML is used to achieve efficient parameter tuning through Bayesian optimization, thereby preventing overfitting and ensuring the stability of the model on unknown data.
[0060] S600: After the preliminary equipment data model is tested and approved using the test set, the industrial equipment data model is obtained. Specifically, the preliminary equipment data model is evaluated using the test set, and the final industrial equipment data model is obtained upon passing the evaluation. Furthermore, the applicability of the preliminary equipment data model in a real industrial environment is tested using the test set, ensuring its suitability for diagnosing industrial equipment in complex scenarios.
[0061] Specifically, this application collects time-series data, image data, and text data from industrial equipment during operation; this comprehensively reflects the state of the industrial equipment. The time-series data, image data, and text data are fused to obtain fused data. By fusing different data, the fused data provides more comprehensive information, eliminates data redundancy, and mines cross-modal correlation features. Based on the fused data, a preliminary learning model is confirmed. The fused data is classified to confirm the training set, validation set, and test set. Classifying the data ensures the objectivity and robustness of the model evaluation. The preliminary learning model is trained and validated using the training set and validation set respectively, resulting in a validated preliminary model of the equipment data. Training and validating the preliminary learning model using the training set and validation set respectively prevents overfitting and ensures the model's stability on unknown data. After the preliminary model of the equipment data passes the test set, the industrial equipment data model is obtained. This further verifies the applicability of the industrial equipment data model in actual industrial environments, improves the adaptability of the industrial equipment data model, meets the needs of diagnosing industrial equipment in complex scenarios, and improves the accuracy of fault prediction, anomaly detection, and process optimization of industrial equipment.
[0062] Please refer to some embodiments of this application as well. Figure 2 This embodiment proposes step S100: the step of collecting time-series data, image data, and text data during the operation of industrial equipment includes:
[0063] S110: Ensure that all sensors used to acquire time-series, image, and text data during industrial equipment operation share the same clock reference. The clock reference is a unified time standard (such as GPS time or Network Time Protocol NTP), ensuring that all sensor timestamps are based on the same reference frame. Sharing the same clock reference ensures consistency of all data across the time dimension. This resolves the issue of clock asynchrony, where sensors may be supplied by different manufacturers. In cases of clock asynchrony, acquired data cannot be aligned in time, leading to inaccurate correlation of data from different modes during subsequent analysis. For example, inconsistencies between the timestamps of vibration data and image data may prevent the correct identification of the equipment's state at a given moment.
[0064] S120: After all sensors share the same clock reference, preliminary time-series data, preliminary image data, and preliminary text data are collected during the operation of industrial equipment. Specifically, data collection after clock synchronization ensures the time validity of the raw data. If data is not collected synchronously, even with subsequent alignment, data distortion may occur due to sensor response delays (such as camera exposure time and text log writing delays). Preliminary time-series data consists of raw time-series signals collected by sensors at fixed intervals. Preliminary image data consists of raw image frames captured by cameras or infrared thermal imagers. Preliminary text data consists of raw logs generated by the equipment control system, such as PLC alarm messages and operator maintenance records. By acquiring unprocessed raw data, the most complete equipment operation information is preserved.
[0065] S130: The preliminary time-series data, preliminary image data, and preliminary text data are timestamped to obtain the time-series data, image data, and text data of the industrial equipment during operation. Timestamp alignment maps the timestamps of different modal data to the same time axis, ensuring the correspondence between multi-source data in the time dimension. This eliminates subtle time deviations in sensor acquisition and transmission, achieving strict correspondence of multi-modal data in the time dimension and providing accurate input for subsequent feature fusion. This avoids inconsistencies in data arrival times caused by transmission delays.
[0066] In this embodiment, by confirming that all sensors used to acquire time-series data, image data, and text data during the operation of industrial equipment share the same clock reference, and by aligning the timestamps of preliminary time-series data, preliminary image data, and preliminary text data, the temporal consistency of multimodal data is ensured, preventing the model from learning incorrect associations and guaranteeing data quality. Simultaneously, high-quality data reduces the model's sensitivity to noise, improving the efficiency of AutoML parameter tuning and the model's generalization ability.
[0067] Please refer to some embodiments of this application as well. Figure 3 This embodiment proposes step S200: fusing the time-series data, image data, and text data to obtain fused data, including the following steps:
[0068] S210: After preprocessing the time-series data, image data, and text data respectively, extract the fusion features of the time-series data, image data, and text data. The preprocessing involves using existing algorithms to address issues such as high-frequency noise (e.g., sensor electromagnetic interference) or missing values (e.g., data packet loss due to communication interruptions) in the time-series data; blurriness (e.g., camera shake) or uneven illumination (e.g., overexposure in infrared thermal imaging) in the image data; and unstructured noise (e.g., spelling errors or irrelevant symbols in operation logs) in the text data, thereby improving data quality. The fusion features of the time-series data capture feature vectors representing potential patterns of the time-series data that are periodic (e.g., motor rotation frequency) and transient (e.g., impact vibration). The fusion features of the image data identify feature vectors representing potential patterns of the image data that are local anomalies (e.g., bearing surface cracks) and global patterns (e.g., equipment thermal distribution). The fusion features of the text data extract feature vectors representing potential patterns of the text data that are key events (e.g., "motor overload alarm") and status descriptions (e.g., "normal operating temperature").
[0069] S220: The fusion features of the time-series data, image data, and text data are concatenated and fused to obtain fused data. The concatenation and fusion process involves connecting feature vectors from different modalities along their dimensions to form a joint feature vector. For example, concatenating and fusion features of 32-dimensional time-series data, 256-dimensional image data, and 768-dimensional text data yields fused data with 1056-dimensional joint features.
[0070] In this embodiment, fusion features are extracted from time-series data, image data, and text data, and then spliced and fused to obtain fused data. This avoids the potential for missing crack information, such as in vibration data, which may not be directly observable, in single-modal data. The gaps can be filled through cross-modal correlation after splicing. Simultaneously, the synergistic effect of different modal features after splicing enhances the features in the fused data. Furthermore, feature extraction compresses the dimensionality of the original data, and splicing and fusion avoids direct processing of high-dimensional original data, thus reducing computational costs.
[0071] In some embodiments of this application, the steps of preprocessing the time-series data, image data, and text data respectively, and then extracting the fusion features of the time-series data, image data, and text data include:
[0072] The time-series data is denoised and standardized to obtain processed time-series data. Denoising can be performed using existing wavelet denoising algorithms, which suppress noise, retain useful signals, and improve the signal-to-noise ratio; for example, the characteristic frequencies of bearing faults in vibration signals become clearer. Standardization is performed using the Z-score standardization method, which maps the data to a uniform scale, avoiding model training bias caused by differences in units.
[0073] The image data is then subjected to denoising and enhancement processing to obtain processed image data. Specifically, existing Gaussian blurring and nonlocal mean denoising methods are used to denoise the image data, thereby suppressing image noise, preserving edge information, and reducing the interference of image noise on subsequent feature extraction. Furthermore, existing histogram equalization is used to expand image contrast, making details in dark areas visible and improving the identifiability of target areas (such as cracks and oil leaks), for example, making overheated components more prominent in infrared thermal imaging.
[0074] The text data is cleaned and vectorized to obtain processed text data. Regular expressions are used for cleaning to standardize the text format, remove irrelevant information, and improve text consistency; for example, unifying "motor overload" and "motor superload" as the same expression. Vectorization converts the text into numerical vectors, thus transforming unstructured text into structured vectors, facilitating fusion with numerical features (such as time series or images).
[0075] The processed time-series data, processed image data, and processed text data are extracted separately to obtain the fusion features of the time-series data, the image data, and the text data. Specifically, the fusion feature extraction uses existing algorithms to extract representative and low-dimensional features from the preprocessed data.
[0076] Preprocessing steps such as denoising, noise reduction, and cleaning significantly improve data quality and reduce model overfitting to noise. The feature extraction step compresses data dimensions and extracts semantic information, enabling AutoML to search for model structures more efficiently. The data format after preprocessing and feature extraction is compatible with AutoML frameworks (such as TPOT and H2O), supports automated model tuning, and improves adaptability.
[0077] Please refer to some embodiments of this application as well. Figure 4 This embodiment proposes step S300: Based on the fused data, the step of confirming the preliminary learning model includes:
[0078] S310: Perform feature extraction on the fused data to obtain feature ratio parameters. These feature ratio parameters quantify the distribution proportion of each modality feature in the overall feature space. For example, they can be calculated based on the dimensionality of each modality feature, such as temporal features accounting for 30%, image features for 50%, and text features for 20%. Feature selection (e.g., L1 regularization) can compress the feature dimension from 1056 dimensions to 200 dimensions, reducing model training costs.
[0079] S320: Perform model matching on the feature ratio parameters within a preset feature ratio scheme to confirm the adapted model. The preset feature ratio scheme is a predefined model selection rule base based on domain knowledge (such as experience in industrial equipment fault diagnosis). In this embodiment, if the temporal feature ratio is greater than 50%, a model capable of capturing temporal dependencies is selected; if the image feature ratio is greater than 50%, a model capable of extracting spatial features is selected; if the text feature ratio is greater than 50%, a model capable of handling semantics is selected; when the multimodal features are relatively balanced, a model supporting multiple inputs, such as a branched neural network or a cross-modal attention model, is selected. If the temporal feature ratio is greater than 60% and the target is a regression task, an LSTM model is matched; if the image feature ratio is greater than 70% and the target is a classification task, a ResNet-18 model is matched.
[0080] S330: Based on the adapted model, confirm the preliminary learning model. The adapted model is the model adjusted according to the feature ratio parameters, possessing an input structure, hyperparameters, and interaction mechanisms that match the fused data. The preliminary learning model is the adapted and adjusted model, possessing basic learning capabilities and can directly enter the training and validation phases. Quantifying data characteristics through feature ratio parameters transforms model selection from a black-box search to feature-driven rule matching. The preset feature ratio scheme injects industrial experience into the AutoML process, improving the rationality and interpretability of model matching. Simultaneously, rule matching replaces global search, significantly shortening model selection time and meeting the real-time requirements of industrial scenarios.
[0081] In some embodiments of this application, this embodiment proposes a step for confirming the preliminary learning model based on the adapted model, including:
[0082] The adapted model undergoes automated hyperparameter tuning to obtain a tuned model. Specifically, automated hyperparameter tuning is the process of automatically searching for the optimal combination of hyperparameters using existing algorithms such as Bayesian optimization. Automated tuning efficiently explores the hyperparameter space, avoids overfitting, and improves the model's generalization ability. The tuned model is a model with optimized hyperparameters, possessing better performance metrics. Through tuning, the model's performance on the validation set is significantly improved.
[0083] The optimized model was validated using industrial rule alignment to obtain validation results. This industrial rule alignment validation involves comparing the model output with known industrial rules or expert knowledge to ensure consistency. For example, a comparative analysis method can be used to directly compare the model output with actual data (such as fault types and safety thresholds).
[0084] After the verification results are satisfactory, the optimized model is used as the initial learning model. Specifically, the degree of matching between the model output and the industry rules must reach a preset threshold; for example, an error rate of less than 5% indicates a satisfactory verification result. Automated optimization improves performance, while industry rule verification ensures reliability, making the initial learning model a reliable model and providing assurance for subsequent testing and deployment.
[0085] Please refer to some embodiments of this application as well. Figure 5 This embodiment proposes step S500: training and validating the preliminary learning model using the training set and validation set respectively to obtain a qualified preliminary model of device data.
[0086] S510: The initial learning model is trained using the training set to obtain the trained learning model. In this embodiment, the training set is labeled historical data used for model parameter optimization (e.g., equipment operation data of a factory over the past 6 months). The training set provides labeled data including normal and fault labels, enabling the model to capture equipment features, learn the periodicity of time-series data (e.g., motor rotation frequency), the spatial patterns of image data (e.g., crack shape), and the semantic relationships of text data (e.g., "overload" and "current surge"). Parameters are optimized by adjusting model weights through backpropagation (e.g., convolutional kernel parameters of CNNs or gating unit parameters of LSTMs) and minimizing the loss function (e.g., cross-entropy loss or mean squared error). The trained learning model, with optimized weights from the training set, possesses the ability to predict equipment status and can output normal or fault labels.
[0087] S520: Validate the trained learning model using the validation set to obtain validation coefficients. Specifically, the validation set is data that is independent of but identically distributed to the training set, used for model selection and hyperparameter tuning. If the training set accuracy is 95% but the validation set accuracy is 80%, it indicates that the model is overfitting and needs regularization; at the same time, determine the model's optimal performance on the validation set to provide a basis for comparison in subsequent tests.
[0088] S530: The verification coefficients are compared with a preset verification threshold to obtain a preliminary model of the validated equipment data. The preset verification threshold is the minimum performance requirement of the model, determined by business needs or industry standards, such as a fault detection rate greater than or equal to 90%. A preliminary model of validated equipment data is a model whose verification coefficients exceed the threshold and is qualified for deployment in a production environment. In this embodiment, by optimizing parameters using the training set, evaluating generalization ability using the validation set, and comparing the preset threshold with the verification coefficients, the model qualification determination is automated, reducing manual intervention.
[0089] In some embodiments of this application, the following steps are proposed: comparing the verification coefficient with a preset verification threshold to obtain a preliminary model of the verified equipment data:
[0090] The validation coefficient is compared with a preset validation threshold. If the validation coefficient is lower than the preset threshold, the ratio of the training set, validation set, and test set of the fused data is reconfirmed. For example, the ratio can be changed from 7:1.5:1.5 to 8:1:1. Increasing the proportion of the training set (e.g., from 70% to 80%) reduces the size of the validation set and lowers the risk of overfitting.
[0091] The verification coefficient is compared with a preset verification threshold. If the verification coefficient is equal to or greater than the preset verification threshold, a preliminary model of the equipment data that has passed verification is obtained. Specifically, a preliminary model of the equipment data that has passed verification by the verification coefficient threshold has basic reliability and can enter the testing phase. The dataset ratio is adjusted by triggering verification failures, thus achieving a closed loop between model development and data adaptation. Simultaneously, the automatic comparison of the verification coefficient and the threshold reduces manual intervention and accelerates model iteration.
[0092] In some embodiments of this application, this embodiment proposes a step of comparing the verification coefficient with a preset verification threshold, and if the verification coefficient is lower than the preset verification threshold, then reconfirming the ratio of the training set, verification set, and test set of the fused data, including:
[0093] The validation coefficient is compared with a preset validation threshold. If the validation coefficient is between the preset validation threshold and 0.9 times the preset validation threshold, the ratio of the training set, validation set, and test set of the fused data is adjusted to 8:1:1. When the validation coefficient is between the preset threshold and 0.9 times the threshold, it indicates that the model performance is close to the expected target, but the model may be slightly underfitting or overfitting. The model's performance on the training and validation sets fluctuates slightly, possibly due to an unreasonable data partitioning ratio leading to evaluation bias. This also addresses the issue of insufficient data utilization efficiency. The original data partitioning ratio (e.g., 7:1.5:1.5) may not have fully utilized the data volume, resulting in a limited training set size and affecting the model's learning potential. This solution, without changing the total amount of data, adjusts the ratio of the training set, validation set, and test set of the fused data to 8:1:1. That is, 80% of the fused data is used as the training set, providing more training samples and significantly enhancing the model's learning ability for marginal cases (such as rare categories or anomalous features). Using 10% of the fused data as a validation set can provide a reliable basis for hyperparameter tuning even when the amount of data is sufficient. Using 10% of the fused data as a test set can ensure the objectivity of the final evaluation and avoid performance evaluation bias caused by an insufficiently small test set.
[0094] Specifically, by increasing the amount of training data, the proportion of the training set was increased from 70% to 80%, thereby increasing the number of samples that the model can learn.
[0095] The validation coefficient is compared with a preset validation threshold. If the validation coefficient is lower than 0.9 times the preset validation threshold, the amount of fused data is increased by at least 20%, and the ratio of the training set, validation set, and test set of the increased fused data is adjusted to 8:1:1. Specifically, adding 20% more fused data can significantly improve model performance. That is, by setting at least a 20% incremental data amount, model performance can be significantly improved while avoiding excessive increases in computational costs. At the same time, the added data may contain more fault types or operating states, covering more scenarios. By increasing the amount of data, the distribution of the training set and validation set is made closer to the real scenario, optimizing the data distribution. In this embodiment, the dataset ratio or amount is adjusted in stages according to the difference between the validation coefficient and the threshold to achieve precise optimization and also to balance resource efficiency.
[0096] In some embodiments of this application, the steps for obtaining an industrial equipment data model after passing the test of the preliminary equipment data model using the test set include:
[0097] The preliminary model of the device data is tested using the test set, and the test results are obtained. The test set is a dataset independent of the training and validation sets, used for the final evaluation of the model's performance. The test results are quantitative performance metrics of the model on the test set, including accuracy, fault detection rate, and false alarm rate. Evaluation using an independent test set ensures the model's reliability in real-world scenarios. If the test set performance is significantly lower than the validation set performance, it indicates that the preliminary model of the device data is overfitted and needs to be readjusted.
[0098] The test results are compared with a preset test threshold. If the test result is greater than the preset test threshold, the preliminary data model of the qualified equipment is taken as the industrial equipment data model. The preset test threshold is the minimum performance value that the industrial equipment data model must achieve, determined by expert experience or industry standards.
[0099] In this embodiment, model testing is conducted using a test set to automatically evaluate model performance. Combined with preset thresholds, automated judgment of model qualification is achieved, reducing manual intervention. Simultaneously, by comparing the test results with preset test thresholds, it is ensured that the preliminary model of qualified equipment data is introduced into the production environment, avoiding production accidents caused by insufficient model performance. This enables the industrial equipment data model to meet the needs of diagnosing industrial equipment in complex scenarios, improving the accuracy of fault prediction, anomaly detection, and process optimization for industrial equipment.
[0100] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program goods. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program goods embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0101] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program goods according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0102] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0103] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for generating industrial equipment data models based on automated machine learning, characterized in that, include: Collect time-series data, image data, and text data of industrial equipment during operation; The time-series data, image data, and text data are fused together to obtain fused data; Based on the fused data, a preliminary learning model was confirmed; The fused data is categorized to identify the training set, validation set, and test set. The preliminary learning model is trained and validated using the training set and validation set, respectively, to obtain a qualified preliminary model of device data. After the preliminary equipment data model is tested and approved using the test set, the industrial equipment data model is obtained.
2. The method for generating industrial equipment data models based on automatic machine learning according to claim 1, characterized in that, The steps for collecting time-series data, image data, and text data during the operation of industrial equipment include: Confirm that all sensors used by industrial equipment to acquire timing data, image data, and text data share the same clock reference during operation; After all sensors share the same clock reference, preliminary time-series data, preliminary image data, and preliminary text data are collected during the operation of industrial equipment. The preliminary time-series data, preliminary image data, and preliminary text data are timestamped to obtain the time-series data, image data, and text data of the industrial equipment during operation.
3. The method for generating industrial equipment data models based on automatic machine learning according to claim 1, characterized in that, The steps for fusing the time-series data, image data, and text data to obtain fused data include: After preprocessing the time-series data, image data, and text data respectively, the fusion features of the time-series data, image data, and text data are extracted. The fusion features of the time-series data, the image data, and the text data are concatenated and fused to obtain fused data.
4. The method for generating industrial equipment data models based on automatic machine learning according to claim 3, characterized in that, The steps of preprocessing the time-series data, image data, and text data respectively, and then extracting the fusion features of the time-series data, image data, and text data include: The time series data is denoised and standardized to obtain the processed time series data; The image data is subjected to noise reduction and enhancement processing to obtain the processed image data; The text data is cleaned and vectorized to obtain the processed text data; The processed time-series data, processed image data, and processed text data were extracted separately to obtain the fusion features of the time-series data, the fusion features of the image data, and the fusion features of the text data.
5. The method for generating industrial equipment data models based on automatic machine learning according to claim 1, characterized in that, Based on the fused data, the steps to confirm the preliminary learning model include: Feature extraction is performed on the fused data to obtain feature ratio parameters; The feature ratio parameters are matched with the model in the preset feature ratio scheme to confirm the adapted model. Based on the adapted model, the initial learning model was confirmed.
6. The method for generating industrial equipment data models based on automatic machine learning according to claim 5, characterized in that, Based on the adapted model, the steps to confirm the initial learning model include: The adapted model is then automatically tuned to obtain the optimized model. The optimized model was validated using industrial rule alignment, and the validation results were obtained. After the verification results are satisfactory, the optimized model will be used as the initial learning model.
7. The method for generating industrial equipment data models based on automatic machine learning according to claim 1, characterized in that, The steps for training and validating the preliminary learning model using the training set and validation set respectively to obtain a valid preliminary model of device data include: The initial learning model is trained using the training set to obtain the trained learning model; The trained learning model is validated using a validation set to obtain validation coefficients; The verification coefficient is compared with the preset verification threshold to obtain a preliminary model of the equipment data that has passed verification.
8. The method for generating industrial equipment data models based on automatic machine learning according to claim 7, characterized in that, The steps for comparing the verification coefficient with a preset verification threshold to obtain a preliminary model of the verified equipment data include: The verification coefficient is compared with a preset verification threshold. If the verification coefficient is lower than the preset verification threshold, the ratio of the training set, verification set and test set of the fused data is reconfirmed. The verification coefficient is compared with the preset verification threshold. If the verification coefficient is equal to or greater than the preset verification threshold, a preliminary model of the equipment data that has passed verification is obtained.
9. The method for generating industrial equipment data models based on automatic machine learning according to claim 8, characterized in that, The step of comparing the verification coefficient with a preset verification threshold, and if the verification coefficient is lower than the preset verification threshold, then reconfirming the ratio of the training set, validation set, and test set of the fused data includes: The verification coefficient is compared with the preset verification threshold. If the verification coefficient is between the preset verification threshold and 0.9 times the preset verification threshold, the ratio of the training set, verification set and test set of the fused data is adjusted to 8:1:
1. The verification coefficient is compared with the preset verification threshold. If the verification coefficient is less than 0.9 times the preset verification threshold, the amount of fused data is increased by at least 20%, and the ratio of the training set, verification set and test set of the increased fused data is adjusted to 8:1:
1.
10. The method for generating industrial equipment data models based on automatic machine learning according to claim 1, characterized in that, After the preliminary equipment data model is tested and approved using the test set, the steps to obtain the industrial equipment data model include: The preliminary model of the device data was tested using the test set, and the test results were obtained. The test results are compared with a preset test threshold. If the test results are greater than the preset test threshold, the preliminary model of the equipment data that passes the test is used as the industrial equipment data model.
Citation Information
Cited By
Cross-modal industrial data efficient cleaning and semantic unification method and device
CN121524492A