Data processing method and device, computer equipment and storage medium
By storing and processing data in a distributed file system, combining feature extraction and parallel training prediction models, the problems of information lag and abnormal identification in traditional quality inspection methods are solved, efficient and accurate quality prediction and dynamic early warning are achieved, and defective rate and production costs are reduced.
Patent Information
- Application Number
- CN202511061655.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-08-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional quality inspection methods have lag in information and untimely abnormal identification, resulting in low accuracy of quality inspection results, which can easily lead to batch defective products and high rework costs.
The original data is stored in a distributed file system, and after preprocessing and feature extraction, the trained prediction model is input to predict. The Spark MLlib/TensorFlow On Spark framework of the Hadoop cluster is used for parallel distributed training, and the features are filtered by sequence feature selection and principal component analysis algorithm are combined to establish a dynamic early warning mechanism.
It improves the accuracy of data processing and the accuracy of predicted results, reduces defective rates and production costs, and improves the efficiency of production processes and resource allocation.
Smart Images

Figure CN120561567A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a data processing method, apparatus, computer equipment, and storage medium. Background Art
[0002] With increasingly fierce competition in the manufacturing industry and the continuous improvement of customers' requirements for product quality, traditional quality control methods can no longer meet the needs of modern production for efficiency and precision.
[0003] The quality inspection methods currently widely used mostly rely on post-inspection and manual experience judgment. Since these quality inspection methods are prone to problems such as delayed information feedback, untimely anomaly identification, and low prediction accuracy, the final quality inspection results are relatively inaccurate, which in turn easily leads to batch defective products and high rework costs. Summary of the Invention
[0004] The embodiments of the present application provide a data processing method, apparatus, computer equipment, and storage medium, which aim to solve the problem that traditional quality inspection methods have low accuracy of quality inspection results due to problems such as information lag and untimely anomaly identification.
[0005] In a first aspect, an embodiment of the present application provides a data processing method, the data processing method comprising: Acquiring raw data to be processed, wherein the raw data to be processed includes quality inspection data, raw material batch information, production process parameters, and equipment operating status information; Storing the raw data to be processed in a preset distributed file system; Reading target data from the preset distributed file system; Preprocessing and feature extraction of the target data to obtain target feature data; The target feature data is input into a trained prediction model for prediction, and a prediction result output by the trained prediction model is obtained.
[0006] In some possible implementations, preprocessing and feature extracting the target data to obtain target feature data includes: Cleaning the target data to obtain target valid data; Feature extraction is performed on the target valid data to obtain target feature data.
[0007] In some possible implementations, extracting features from the target valid data to obtain target feature data includes: Calculating statistical characteristics of the target valid data to obtain target statistical characteristics; The target statistical features are screened using a sequence feature selection algorithm and / or a principal component analysis algorithm to obtain target feature data.
[0008] In some possible implementations, after preprocessing and extracting features from the target data to obtain target feature data, the method further includes: Performing encoding conversion processing and / or feature scaling processing on the target feature data to obtain converted target feature data; Inputting the target feature data into a trained prediction model for prediction, and obtaining a prediction result output by the trained prediction model, includes: The converted target feature data is input into a trained prediction model for prediction, and a prediction result output by the trained prediction model is obtained.
[0009] In some possible implementations, after obtaining the prediction result output by the trained prediction model, the method further includes: Determining whether the prediction result meets a preset condition; If the prediction result does not meet the preset conditions, an alarm will be automatically triggered.
[0010] In some possible implementations, reading the target data from the preset distributed file system includes: The target data is read from the preset distributed file system using the preset Hive data processing tool and the HBase data processing tool.
[0011] In some possible implementations, the prediction model is trained in parallel and distributed manner using the Spark MLlib / TensorFlow On Spark framework of a Hadoop cluster.
[0012] In a second aspect, an embodiment of the present application further provides a data processing device, which includes a unit for executing the above method.
[0013] In a third aspect, an embodiment of the present application further provides a computer device, which includes a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.
[0014] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program can implement the above method when executed by a processor.
[0015] Embodiments of the present application provide a data processing method, apparatus, computer device, and storage medium. The method includes obtaining raw data to be processed, wherein the raw data to be processed includes quality inspection data, raw material batch information, production process parameters, and equipment operating status information; storing the raw data to be processed in a preset distributed file system; reading target data from the preset distributed file system; preprocessing and feature extraction of the target data to obtain target feature data; and inputting the target feature data into a trained prediction model for prediction, thereby obtaining a prediction result output by the trained prediction model.
[0016] The embodiment of the present application first stores the raw data to be processed (such as quality inspection data, production process parameters, equipment operating status, etc.) in a preset distributed file system, and then reads the target data from the preset distributed file system. This can avoid direct real-time judgment of the sensor's instantaneous readings, which is prone to false alarms or missed alarms due to occasional fluctuations, thereby making the read data more accurate and reliable. In addition, the read target data is first processed (such as effective processing, noise processing, etc.) and feature extracted to obtain target feature data, and then the processed target feature data is input into the trained prediction model for prediction to obtain the prediction result output by the trained prediction model, which can make the accuracy of the final output prediction result relatively high. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0020] Figure 1 A flowchart of a first embodiment of a data processing method provided by this application; Figure 2 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0021] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0022] The disclosure below provides many different embodiments or examples for implementing different structures of the present application. In order to simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, these are merely examples and are not intended to limit the present application. In addition, the present application may repeat reference numbers and / or letters in different examples. Such repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or settings discussed.
[0023] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0024] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0025] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0026] As used in this specification and the appended claims, the term “if” can be interpreted as “when” or “upon” or “in response to determining” or “in response to detecting,” depending on the context. Similarly, the phrase “if it is determined” or “if [described condition or event] is detected” can be interpreted as meaning “upon determination” or “in response to determining” or “upon detection of [described condition or event]” or “in response to detecting [described condition or event],” depending on the context.
[0027] In order to solve the above problems, the present application provides a data processing method that can improve the accuracy of data processing and thereby improve the accuracy of quality prediction results.
[0028] See Figure 1 , Figure 1 This is a flow chart of a first embodiment of a data processing method provided by the present application, wherein the data processing method comprises the following steps: Step 110: Obtain the original data to be processed.
[0029] The raw data to be processed include quality inspection data, raw material batch information, production process parameters and equipment operation status information.
[0030] Step 120: Store the raw data to be processed in a preset distributed file system.
[0031] Step 130: Read target data from the preset distributed file system.
[0032] In some possible implementations, reading the target data from the preset distributed file system includes: Use the preset Hive data processing tools and HBase data processing tools to read target data from the preset distributed file system (Hadoop Distributed File System, Hadoop).
[0033] Hive is used for data warehouse operations and provides SQL query capabilities for large-scale data. HBase is used for real-time reading and writing of large-scale data and provides efficient random access capabilities.
[0034] Step 140: Preprocessing and feature extraction are performed on the target data to obtain target feature data.
[0035] Step 150: Input the target feature data into the trained prediction model for prediction, and obtain the prediction result output by the trained prediction model.
[0036] This embodiment first stores the raw data to be processed (such as quality inspection data, production process parameters, equipment operating status, etc.) in a preset distributed file system, and then reads the target data from the preset distributed file system. This can avoid direct real-time judgment of the sensor's instantaneous readings, which is prone to false alarms or missed alarms due to occasional fluctuations, thereby making the read data more accurate and reliable. In addition, the read target data is first processed (such as effective processing, noise processing, etc.) and feature extracted to obtain target feature data, and then the processed target feature data is input into the trained prediction model for prediction to obtain the prediction result output by the trained prediction model, which can make the accuracy of the final output prediction result relatively high.
[0037] In some possible implementations, preprocessing and feature extracting the target data to obtain target feature data includes: Step 141: Clean the target data to obtain target valid data.
[0038] Step 142: Extract features from the target valid data to obtain target feature data.
[0039] In some possible implementations, extracting features from the target valid data to obtain target feature data includes: Step 1421: Calculate statistical features of the target valid data to obtain target statistical features; Step 1422: Utilize a sequence feature selection algorithm and / or a principal component analysis algorithm to perform feature screening on the target statistical features to obtain target feature data.
[0040] In some possible implementations, after preprocessing and extracting features from the target data to obtain target feature data, the method further includes: Performing encoding conversion processing and / or feature scaling processing on the target feature data to obtain converted target feature data; Inputting the target feature data into a trained prediction model for prediction, and obtaining a prediction result output by the trained prediction model, includes: The converted target feature data is input into a trained prediction model for prediction, and a prediction result output by the trained prediction model is obtained.
[0041] In some possible implementations, after obtaining the prediction result output by the trained prediction model, the method further includes: Determining whether the prediction result meets a preset condition; If the prediction result does not meet the preset conditions, an alarm will be automatically triggered.
[0042] In some possible implementations, the prediction model is trained in parallel and distributed manner using the Spark MLlib / TensorFlow On Spark framework of a Hadoop cluster.
[0043] Based on the above embodiments, the data processing method provided by this application mainly includes the following processes: 1. Data Collection and Storage 1) Obtain the raw data to be processed.
[0044] Among them, the raw data to be processed include quality inspection data, raw material batch information, production process parameters and equipment operation status information.
[0045] Specifically, quality inspection data can be data such as product size, appearance, and performance collected in real time from inspection equipment deployed in various links of the production line.
[0046] Raw material batch information can record detailed information such as the storage time, supplier, batch number, quality inspection report, etc. of each batch of raw materials.
[0047] Production process parameters can be key process parameters such as temperature, pressure, humidity, production speed, etc. collected during the production process.
[0048] Equipment operating status information may include the operating status, fault codes, maintenance records, etc. of the monitored production equipment.
[0049] 2) Storing the raw data to be processed in a preset distributed file system.
[0050] When storing data, a tiered storage strategy can be adopted to store data in different layers of HDFS based on the access frequency and importance of the data.
[0051] In addition, efficient and reliable message queue technologies such as Nifi or Kafka can be used to ensure that data is transmitted to the Hadoop cluster in real time and accurately, and then the HDFS distributed file system can be used to store large amounts of raw data to be processed.
[0052] In some embodiments, during the data transmission process, a data verification mechanism may be designed to ensure the integrity and accuracy of the data during the transmission process.
[0053] (2) Data processing and feature engineering 1) Cleaning the target data to obtain target valid data.
[0054] Among them, data cleaning processing can include operations such as missing value processing, outlier detection and data standardization processing.
[0055] For example, for numerical data, you can use the mean, median, or mode to fill missing values; for time series data, you can use interpolation methods (such as linear interpolation and polynomial interpolation) to fill missing values; for categorical data, you can use the most frequent value to fill missing values.
[0056] Outlier detection can be done by using a boxplot to identify outliers, where data points outside the upper and lower bounds of the boxplot are considered outliers. Alternatively, the 3σ rule can be used, where a data point is considered an outlier if it deviates from the mean by more than three standard deviations.
[0057] Detected outliers can be eliminated or corrected according to business needs to ensure the accuracy and reliability of the data.
[0058] In some embodiments, data normalization may be mapping the data to the interval [0, 1], or converting the data into a distribution with a mean of 0 and a standard deviation of 1.
[0059] In this way, the data is normalized or standardized to eliminate dimension and magnitude differences, thereby improving the stability and accuracy of model training.
[0060] 2) Extracting features from the target valid data to obtain target feature data.
[0061] 2-1) Calculating statistical characteristics of the target valid data to obtain target statistical characteristics (such as mean, variance, range, kurtosis, skewness, etc.); Alternatively, the correlation coefficients between the features may be calculated, highly correlated features may be identified, and the highly correlated features may be used as target feature data.
[0062] 2-2) Using methods such as sequence feature selection algorithm (SBS) and principal component analysis algorithm (PCA) to perform feature screening on the target statistical features to obtain target feature data.
[0063] Among them, the target feature data can be a key feature that has a greater impact on product quality.
[0064] Specifically, using the sequential feature selection algorithm (SBS), features that contribute less to the model can be gradually eliminated, and key features can be retained.
[0065] The principal component analysis (PCA) algorithm can be used to reduce the dimensionality of high-dimensional data, retain the main components, and reduce the number of features.
[0066] In some embodiments, the experience of manufacturing industry experts can be combined to extract key features that have a significant impact on product quality, such as the range of variation of specific process parameters and certain characteristics of raw materials. Alternatively, new features can be generated based on business needs, such as sliding window statistical features of time series data.
[0067] 2-3) Performing encoding conversion processing and / or feature scaling processing on the target feature data to obtain converted target feature data.
[0068] The encoding conversion process can be one-hot encoding of categorical variables, converting each category into a binary vector, or label encoding of ordered categorical variables, converting categories into numerical labels.
[0069] Feature scaling can be done by normalizing the data, mapping the data to the [0, 1] interval, and eliminating dimensional differences.
[0070] Alternatively, the data can be standardized and converted into a distribution with a mean of 0 and a standard deviation of 1, so that the prediction model can learn and converge better, thereby improving the stability of model training.
[0071] In addition, the target feature data can be smoothed before feature scaling (for example, moving average or exponential smoothing of time series data) to eliminate noise in the target feature data and improve its stability.
[0072] 4) Inputting the converted target feature data into a trained prediction model for prediction, and obtaining a prediction result output by the trained prediction model.
[0073] 5) Determine whether the prediction result meets the preset conditions.
[0074] If the prediction result does not meet the preset conditions, an alarm will be automatically triggered.
[0075] For example, the prediction result can be compared with a preset threshold (such as a critical value of defective rate), and an alarm will be automatically triggered if the prediction result exceeds the threshold.
[0076] In addition, relevant personnel can be notified in a timely manner through various means such as email, text messages, system pop-ups, etc. to take measures to deal with the situation.
[0077] Alternatively, it can display forecast results and historical trend charts to users through a visual interface, or provide users with customized reports and data analysis functions to help corporate management better understand the quality status of the production process and make scientific decisions.
[0078] By storing batches of data within a specific time window, aggregate statistics (such as calculating mean and standard deviation) can be performed on the data, reducing the impact of random noise. Combined with time series analysis (such as sliding windows), trend anomalies can be captured. After storage, sliding averaging can filter out noise, ensuring that the data input into the trained forecasting model is more accurate and reliable.
[0079] In addition, by predicting quality problems through trained prediction models, companies can take timely measures to intervene and adjust according to the prediction results, effectively reduce the defective rate and rework rate, improve the overall quality level of products, reduce additional expenses such as scrap, rework and repairs caused by quality problems, optimize production processes and resource allocation, and reduce production costs.
[0080] In some embodiments, the training of the above prediction model may include the following process: (1) Algorithm selection For classification problems, you can choose Random Forest, Gradient Boosting Trees, and Support Vector Machine (SVM).
[0081] Random Forest is suitable for high-dimensional data and has good generalization and overfitting resistance. Gradient Boosting Trees improves model prediction accuracy by gradually optimizing the loss function. Support Vector Machines (SVM) are suitable for small sample data and have good classification performance.
[0082] For regression problems, you can choose Linear Regression, Neural Networks, Decision Tree Regression, etc.
[0083] Among them, linear regression is suitable for data with strong linear relationships, and the model is simple and easy to interpret. Neural networks are suitable for complex nonlinear relationships and have high prediction accuracy. Decision tree regression is suitable for processing nonlinear relationships, and the model is easy to interpret.
[0084] (2) Model training (2-1) Distributed training 1) In the Hadoop environment, use distributed computing frameworks such as Spark MLlib or TensorFlow on Spark to perform large-scale parallel model training.
[0085] 2) Distribute the dataset across multiple nodes, calculate gradients in parallel, and accelerate the model training process.
[0086] In this way, compared with traditional single-machine training, this application can significantly accelerate the model convergence speed by utilizing the Spark MLlib / TensorFlow On Spark framework of the Hadoop cluster to implement parallel training of large-scale data (such as TB-level data stored in HDFS).
[0087] (2-2) Hyperparameter Tuning 1) Use Grid Search or Random Search to iterate over the hyperparameter combinations and find the optimal parameters.
[0088] 2) Use Bayesian Optimization to intelligently select the next set of hyperparameters to try based on existing evaluation results.
[0089] In this way, Bayesian optimization is used instead of traditional grid search / random search, and the optimal parameter combination is intelligently screened based on historical evaluation results. The learning rate and tree depth can be dynamically adjusted in the gradient boosting tree model, avoiding the high cost of manual trial and error.
[0090] (2-3) Cross-validation: Cross-validation can be K-fold cross-validation, for example, where the dataset is divided into K subsets, with K-1 subsets used as the training set and the remaining subset used as the validation set. This is repeated K times, with a different subset used as the validation set each time. The average of the K validation results is then used to evaluate the model performance.
[0091] Exemplarily, K-fold cross validation may include the following: 1) Randomly divide the dataset into K subsets of similar size.
[0092] 2) For each subset, train the model using the other K-1 subsets and validate on that subset.
[0093] 3) Calculate the evaluation indicators (such as accuracy, recall, etc.) for each verification.
[0094] 4) Take the average of K verification results as the final evaluation result.
[0095] (3) Model evaluation When evaluating a model, you can use metrics such as accuracy, recall, and F1 score to assess the performance of a classification model. You can use metrics such as mean squared error (MSE) and mean absolute error (MAE) to assess the performance of a regression model.
[0096] For accuracy, recall, and F1 scores, please refer to the following introduction: Accuracy: The proportion of samples correctly predicted by the model to the total samples.
[0097] Recall: The ratio of positive samples correctly predicted by the model to all actual positive samples.
[0098] F1 score: The harmonic mean of precision and recall, used to balance the relationship between the two.
[0099] The specific operation may include the following steps: 1) Calculate the confusion matrix of the model, including true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN).
[0100] 2) Calculate the accuracy, recall and F1 score based on the confusion matrix: For example, precision = (TP + TN) / (TP + FP + TN + FN); recall = TP / (TP +FN); F1 score = 2 * (precision * recall) / (precision + recall).
[0101] In some embodiments, the model can be evaluated using ROC curves and AUC values.
[0102] Specifically, the ROC curve can be used to evaluate the classification performance of the model at different thresholds by plotting the relationship between the true positive rate (TPR) and the false positive rate (FPR).
[0103] The UC value can be the area under the ROC curve, which is used to measure the overall classification performance of the model. The closer the AUC value is to 1, the better the model performance.
[0104] Exemplarily, the calculation of the ROC curve and the AUC value may include the following: 1) Calculate the TPR and FPR of the model at different thresholds.
[0105] 2) Draw the ROC curve with FPR on the horizontal axis and TPR on the vertical axis.
[0106] 3) Calculate the area under the ROC curve (AUC value) as an evaluation indicator of the model classification performance.
[0107] In this way, compared with the traditional evaluation method that uses a single indicator (such as only accuracy), this application comprehensively evaluates the performance of the model through comprehensive cross-validation, ROC curve, AUC value, accuracy, recall rate and other indicators, which can ensure that the model performs as expected in prediction and classification tasks.
[0108] In other words, this application uses the distributed computing framework of the Hadoop cluster to achieve efficient storage and parallel processing of massive heterogeneous data, and combines the Bayesian optimization algorithm to intelligently optimize hyperparameters of complex models, breaking through the efficiency bottleneck of traditional single-machine training. It innovatively integrates the experience of manufacturing experts with data-driven feature engineering, extracts key process features through PCA dimensionality reduction and sequence feature selection algorithms, and significantly improves model interpretability. It uses a multi-dimensional evaluation system (cross-validation + ROC curve + F1 score) to ensure the generalization ability of the model, and establishes a dynamic early warning mechanism and a visual decision-making platform to achieve full-link intelligence from data collection to quality control. Compared with traditional quality control methods, this system will increase the prediction response speed by 80%, reduce the defective rate by 45%, and support management to obtain real-time production quality situation awareness.
[0109] Corresponding to the above data processing method, the present application also provides a data processing device. The data processing device includes a unit for executing the above data processing method, and the data processing device can be configured in a desktop computer, tablet computer, laptop computer, or other terminal.
[0110] like Figure 2 As shown, an embodiment of the present application provides a computer device, including a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114. Memory 113, for storing computer programs; In one embodiment of the present application, the processor 111 is configured to execute a program stored in the memory 113 to implement the data processing method provided by any of the aforementioned method embodiments, including: Acquiring raw data to be processed, wherein the raw data to be processed includes quality inspection data, raw material batch information, production process parameters, and equipment operating status information; Storing the raw data to be processed in a preset distributed file system; Reading target data from the preset distributed file system; Preprocessing and feature extraction of the target data to obtain target feature data; The target feature data is input into a trained prediction model for prediction, and a prediction result output by the trained prediction model is obtained.
[0111] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.
[0112] Therefore, an embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the data processing method provided in any of the aforementioned method embodiments are implemented, including: Acquiring raw data to be processed, wherein the raw data to be processed includes quality inspection data, raw material batch information, production process parameters, and equipment operating status information; Storing the raw data to be processed in a preset distributed file system; Reading target data from the preset distributed file system; Preprocessing and feature extraction of the target data to obtain target feature data; The target feature data is input into a trained prediction model for prediction, and a prediction result output by the trained prediction model is obtained.
[0113] The storage medium is a physical, non-transient storage medium, such as a USB flash drive, a removable hard drive, a read-only memory (ROM), a magnetic disk, or an optical disk, among other physical storage media capable of storing program code. The computer-readable storage medium may be either non-volatile or volatile.
[0114] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0115] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and other division methods may be used in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not implemented.
[0116] The steps in the method of the embodiment of the present application can be adjusted in order, combined, and deleted according to actual needs. The units in the device of the embodiment of the present application can be combined, divided, and deleted according to actual needs. In addition, the functional units in the various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit.
[0117] If this integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this application.
[0118] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0119] Obviously, those skilled in the art may make various modifications and variations to this application without departing from the spirit and scope of this application. Thus, as long as these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
[0120] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A data processing method, characterized in that: The data processing method includes: Acquiring raw data to be processed, wherein the raw data to be processed includes quality inspection data, raw material batch information, production process parameters, and equipment operating status information; Storing the raw data to be processed in a preset distributed file system; Reading target data from the preset distributed file system; Preprocessing and feature extraction of the target data to obtain target feature data; The target feature data is input into a trained prediction model for prediction, and a prediction result output by the trained prediction model is obtained.
2. The method according to claim 1, characterized in that The target data is preprocessed and feature extracted to obtain target feature data, including: Cleaning the target data to obtain target valid data; Feature extraction is performed on the target valid data to obtain target feature data.
3. The method according to claim 2, characterized in that The step of extracting features from the target valid data to obtain target feature data includes: Calculating statistical characteristics of the target valid data to obtain target statistical characteristics; The target statistical features are screened using a sequence feature selection algorithm and / or a principal component analysis algorithm to obtain target feature data.
4. The method according to claim 2, characterized in that After preprocessing and feature extraction of the target data to obtain target feature data, the method further includes: Performing encoding conversion processing and / or feature scaling processing on the target feature data to obtain converted target feature data; Inputting the target feature data into a trained prediction model for prediction, and obtaining a prediction result output by the trained prediction model, includes: The converted target feature data is input into a trained prediction model for prediction, and a prediction result output by the trained prediction model is obtained.
5. The method according to claim 1, wherein After obtaining the prediction result output by the trained prediction model, the method further includes: Determining whether the prediction result meets a preset condition; If the prediction result does not meet the preset conditions, an alarm will be automatically triggered.
6. The method according to claim 1, characterized in that The target data is read from the preset distributed file system, including: The target data is read from the preset distributed file system using the preset Hive data processing tool and the HBase data processing tool.
7. The method according to claim 1, characterized in that The prediction model is trained in parallel and distributed manner using the SparkMLlib / TensorFlow On Spark framework of the Hadoop cluster.
8. A data processing device, characterized in that: The method comprises a unit for executing the method according to any one of claims 1 to 7.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the computer program can implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data management platform of industrial internet
CN119166715A
Data processing method based on artificial intelligence
CN119988834A