Model training method, device and equipment for overflow concentration in ore grinding process and medium
By constructing labeled and unlabeled datasets and using pseudo-labels to expand the training set for semi-supervised learning, the time-consuming and labor-intensive problems of traditional supervised learning methods are solved, and the model training efficiency and performance are improved.
Patent Information
- Application Number
- CN202510813594.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Traditional supervised learning methods rely on large amounts of labeled data in model training, which is time-consuming, labor-intensive and costly, reducing model training efficiency.
By constructing labeled and unlabeled datasets, using labeled data to train the initial model to generate pseudo labels, expanding the unlabeled dataset, constructing an extended training set, and performing semi-supervised learning to reduce dependence on labeled data.
It improves the efficiency and performance of model training, reduces time and resource consumption, and enhances the model's generalization ability and data utilization.
Smart Images

Figure CN120670852A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a model training method, device, equipment and medium for overflow concentration in a grinding process. Background Art
[0002] In today's artificial intelligence field, model training is undoubtedly a crucial and core component. It directly impacts whether AI systems can accurately and efficiently perform tasks and provide satisfactory services to users. With the advent of the big data era, the scale and complexity of data are increasing at an unprecedented rate. This data contains a wealth of information and value, but it also presents unprecedented challenges for model training. How to efficiently utilize this massive and complex data to train high-performance, stable, and reliable models has become a key concern and pressing issue for researchers and engineers.
[0003] Traditional supervised learning methods play a crucial role in model training, but they rely heavily on large amounts of labeled data. Labeled data refers to data that has been manually annotated or categorized, providing the model with clear learning objectives and directions. However, in practical applications, obtaining labeled data is often an extremely time-consuming, labor-intensive, and costly task. This requires not only specialized personnel to perform the labeling but also rigorous review and verification of the labeled data to ensure accuracy and consistency. This consumes significant time and resources, reducing the efficiency of model training. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a model training method, device, equipment and medium for overflow concentration in the grinding process, so as to reduce the time and resources required for model training and improve the efficiency of model training.
[0005] In a first aspect, an embodiment of the present application provides a model training method for overflow concentration in a grinding process, the method comprising: Constructing a labeled data set and an unlabeled data set based on the input variables and output variables of each labeled data and the input variables of each unlabeled data; constructing a candidate list based on the labeled dataset and the unlabeled dataset; Using the labeled data set to train a first initial model to obtain a first target model, and using the candidate list to expand the unlabeled data set to obtain a target data set; Determining a first pseudo label for each target data in the target data set by using the first target model, and constructing an extended training set according to the first pseudo label for each target data; The second initial model is trained using the extended training set to obtain a second target model.
[0006] Optionally, the method further includes: The overflow concentration during the grinding process is predicted using the second target model.
[0007] Optionally, constructing the labeled data set and the unlabeled data set based on the input variables and output variables of each labeled data and the input variables of each unlabeled data includes: Determine the weight value of each input variable according to the variable characteristics of the input variables and output variables of each labeled data; Perform weighted processing on each labeled data and each unlabeled data based on the weight value of each input variable; A labeled dataset and an unlabeled dataset are constructed based on the weighted labeled data and unlabeled data respectively.
[0008] Optionally, constructing a candidate list based on the labeled dataset and the unlabeled dataset includes: Performing clustering processing on the labeled data set to obtain a number of clusters; Calculating the distance between the centroid of each cluster and the unlabeled data set in the feature space; The candidate list is constructed based on the minimum distance corresponding to the centroid of each cluster.
[0009] Optionally, constructing an extended training set according to the first pseudo label of each target data includes: Calculate the confidence of the first pseudo label of each target data; The extended training set is constructed according to the confidence of the first pseudo label of each target data.
[0010] Optionally, constructing the extended training set according to the confidence of the first pseudo label of each target data includes: Determine whether the confidence of the first pseudo label of each target data exceeds a preset threshold; If there is valid target data whose confidence level of the first pseudo label exceeds the preset threshold, constructing the extended training set based on each valid target data; If the valid target data does not exist, updating the model parameters of the first initial model to obtain a second initial model, and training the second initial model using the labeled data set to obtain a second target model; The second target model is used to determine a second pseudo label for each target data in the target data set, and the extended training set is constructed according to the second pseudo label for each target data.
[0011] Optionally, constructing the extended training set based on each valid target data includes: Construct auxiliary training sets based on each valid target data; The extended training set is constructed based on the auxiliary training set and the labeled dataset.
[0012] In a second aspect, an embodiment of the present application provides a model training device for overflow concentration in a grinding process, the device comprising: A data set construction module, configured to construct a labeled data set and an unlabeled data set based on the input variables and output variables of each labeled data and the input variables of each unlabeled data; a candidate list building module, configured to build a candidate list based on the labeled data set and the unlabeled data set; a data set expansion module, configured to train a first initial model using the labeled data set to obtain a first target model, and to expand the unlabeled data set using the candidate list to obtain a target data set; a training set construction module, configured to determine a first pseudo label for each target data in the target data set by using the first target model, and construct an extended training set according to the first pseudo label for each target data; The model training module is used to train the second initial model using the extended training set to obtain a second target model.
[0013] Optionally, the device further comprises: The overflow concentration prediction module is used to predict the overflow concentration during the grinding process using the second target model.
[0014] Optionally, constructing the labeled data set and the unlabeled data set based on the input variables and output variables of each labeled data and the input variables of each unlabeled data includes: Determine the weight value of each input variable according to the variable characteristics of the input variables and output variables of each labeled data; Perform weighted processing on each labeled data and each unlabeled data based on the weight value of each input variable; A labeled dataset and an unlabeled dataset are constructed based on the weighted labeled data and unlabeled data respectively.
[0015] Optionally, constructing a candidate list based on the labeled dataset and the unlabeled dataset includes: Performing clustering processing on the labeled data set to obtain a number of clusters; Calculating the distance between the centroid of each cluster and the unlabeled data set in the feature space; The candidate list is constructed based on the minimum distance corresponding to the centroid of each cluster.
[0016] Optionally, constructing an extended training set according to the first pseudo label of each target data includes: Calculate the confidence of the first pseudo label of each target data; The extended training set is constructed according to the confidence of the first pseudo label of each target data.
[0017] Optionally, constructing the extended training set according to the confidence of the first pseudo label of each target data includes: Determine whether the confidence of the first pseudo label of each target data exceeds a preset threshold; If there is valid target data whose confidence level of the first pseudo label exceeds the preset threshold, constructing the extended training set based on each valid target data; If the valid target data does not exist, updating the model parameters of the first initial model to obtain a second initial model, and training the second initial model using the labeled data set to obtain a second target model; The second target model is used to determine a second pseudo label for each target data in the target data set, and the extended training set is constructed according to the second pseudo label for each target data.
[0018] Optionally, constructing the extended training set based on each valid target data includes: Construct auxiliary training sets based on each valid target data; The extended training set is constructed based on the auxiliary training set and the labeled dataset.
[0019] In a third aspect, an embodiment of the present application provides a computer device comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, the steps of the model training method for overflow concentration of the grinding process described in any optional embodiment of the first aspect are performed.
[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, executes the steps of the model training method for overflow concentration of a grinding process described in any optional embodiment of the first aspect above.
[0021] The technical solutions provided by this application include but are not limited to the following beneficial effects: This application first provides a clear data basis for subsequent steps by dividing the data set into two parts: labeled and unlabeled. The labeled data set ensures the accuracy of model learning, while the unlabeled data set provides potential and mineable data resources, which helps to improve the generalization ability of the model. Then, a candidate list is constructed based on the labeled data set and the unlabeled data set, and the labeled data and the unlabeled data are screened based on their similarity or other characteristics, which provides the possibility for subsequent data expansion and pseudo-label generation, helps to efficiently utilize unlabeled data and improve the training efficiency of the model. Next, the first initial model is trained using the labeled data set to obtain the first target model. Through supervised learning, the basic model is trained using the labeled data set, which provides a reliable prediction tool for subsequent pseudo-label generation. The accuracy of the first target model directly affects the quality of subsequent pseudo-labels, and thus affects the performance of the final model. Then, the unlabeled data set is expanded using the candidate list to obtain the target data set. The unlabeled data similar to the labeled data is screened out through the candidate list and expanded, which increases the diversity of the training data, helps to improve the generalization ability of the model, and enables it to better adapt to different data distributions. The first target model is then used to determine the first pseudo-label for each target data point in the target dataset. An extended training set is then constructed based on the first pseudo-label for each target data point. This enables transfer learning from limited labeled data to a large amount of unlabeled data, reducing reliance on labeled data while improving model training efficiency and performance. Finally, the second initial model is trained using the extended training set to obtain the second target model. This helps the model learn more complex data distributions and feature representations, resulting in a high-performance model.
[0022] By adopting the above scheme, combining labeled data and unlabeled data and utilizing a semi-supervised learning strategy, data utilization is improved, thereby reducing the time and resources required for model training and improving the efficiency of model training.
[0023] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without making any creative efforts.
[0025] Figure 1 A flow chart of a model training method for overflow concentration in a grinding process provided by the first embodiment of the present invention is shown; Figure 2 A flowchart of a data set construction method provided in the first embodiment of the present invention is shown; Figure 3 A flowchart of a candidate list construction method provided by the first embodiment of the present invention is shown; Figure 4 A flowchart of a method for constructing an extended training set provided in the first embodiment of the present invention is shown; Figure 5 A flowchart of a specific method for constructing an extended training set provided by the first embodiment of the present invention is shown; Figure 6 A flowchart of a second specific method for constructing an extended training set provided in the first embodiment of the present invention is shown; Figure 7 A schematic structural diagram of a model training device for overflow concentration in a grinding process provided by a second embodiment of the present invention is shown; Figure 8 A schematic structural diagram of a computer device provided in the third embodiment of the present invention is shown. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0027] Example 1 To facilitate understanding of this application, Figure 1 The flowchart of the model training method for overflow concentration in a grinding process provided by the first embodiment of the present invention is shown to describe the content of the first embodiment of the present application in detail.
[0028] See also Figure 1 As shown, Figure 1 A flowchart of a model training method for overflow concentration in a grinding process provided by the first embodiment of the present invention is shown, wherein the method includes steps S101 to S105: S101: Constructing a labeled dataset and an unlabeled dataset based on the input variables and output variables of each labeled data and the input variables of each unlabeled data.
[0029] Specifically, datasets (also known as sample sets) are divided into labeled datasets (i.e., labeled samples) and unlabeled datasets (i.e., unlabeled samples). Labeled datasets contain a series of data points whose true labels are known. These data points consist of input variables (also known as features, which describe the various attributes of the data points) and corresponding output variables (also known as labels or target values, which are the values we hope to predict based on the input variables). Unlabeled datasets, on the other hand, contain only input variables and do not have corresponding output variables or labels.
[0030] When training a model, the execution process includes: The collected data are divided into Labeled data for input variables and one output variable and only Unlabeled data for input variables , and Represents the number of labeled data and unlabeled data respectively. By calculating the input variables and the target variable The correlation of , is a diagonal matrix, where . The calculation formula is as follows: (1) in, Indicates the The first sample Features, It is The mean of the features; Indicates the The target variable for each sample, represents the mean of the target variable, Indicates the The weight of the feature, Represents all diagonal elements in a diagonal matrix.
[0031] for and Assign the corresponding , obtain the weighted adjusted labeled dataset and unlabeled dataset; Among them, the labeled dataset is: ; The unlabeled dataset is: .
[0032] S102: Construct a candidate list based on the labeled dataset and the unlabeled dataset.
[0033] Specifically, after obtaining labeled and unlabeled datasets, a candidate list is constructed to select data points from the unlabeled dataset that may help improve model performance. This list can be constructed based on various strategies, such as the similarity between unlabeled and labeled data, the distribution characteristics of unlabeled data, or specific heuristic rules. The constructed candidate list will be used in the subsequent data augmentation step.
[0034] When training a model, the execution process also includes: Using Gaussian mixture model Clustering, first calculate a data point through Bayes theorem The posterior probability of belonging to the kth cluster , select the cluster with the maximum probability as the final cluster of the point, and finally determine the centroid of each cluster , the relevant formula is as follows:
[0035] in, is the number of clusters; Indicates the The probability density of a Gaussian distribution; Indicates the The weight of a cluster, Indicates the The weight of a cluster, .
[0036] By formula calculate Weighted distance from the centroid in feature space , and according to the formula Get the nearest centroid to the weighted point and return the corresponding distance By formula Generate candidate list based on distance for unlabeled , candidate list Contains samples x1, x2, x3, ..., x u-1 , x u , in preparation for further marking.
[0037]
[0038] in, Represents the position of the centroid in the jth dimension.
[0039] S103: Using the labeled data set to train a first initial model to obtain a first target model, and using the candidate list to expand the unlabeled data set to obtain a target data set.
[0040] Specifically, we first use the labeled dataset to train an initial model (the SCN model can be selected), referred to as the first initial model. This model predicts the unlabeled data, generating pseudo-labels. Then, using the constructed candidate list and the predictions of the first initial model, we select data points with high confidence from the unlabeled dataset and assign them pseudo-labels. These pseudo-labeled data points are then added to the labeled dataset, forming a larger target dataset containing more diverse data. When training a model, the execution process also includes: use The first initial model (SCN model) is trained to obtain the first target model. At the same time, the first batch of data in the candidate list C is moved into the unlabeled data pool. See also Figure 4 As shown, Figure 4 A schematic diagram of a specific random configuration network framework provided by the first embodiment of the present invention is shown. The random configuration network framework includes an input layer, a hidden layer, and an output layer. Its working mechanism can be described as follows:
[0041]
[0042]
[0043] in, Indicates the number of nodes in the SCN hidden layer is When the network output; Represents the output weight of the hidden layer node r; Activation function Relu; and Denote the input weight and bias of the rth node in the hidden layer respectively; is the current network residual; Represents the hidden layer nodes Output; and Represents nodes respectively Candidate parameters for ; represents a non-real number sequence, where , ,satisfy The candidate node parameter with the maximum value is taken as the first Node parameters; Represented as the output weight of the hidden layer node; =[ , ,..., ].
[0044] S104: Determine a first pseudo label for each target data in the target data set by using the first target model, and construct an extended training set according to the first pseudo label for each target data.
[0045] Specifically, the first target model (i.e., a model trained on a labeled dataset) is used to predict each data point in the target dataset. Several first pseudo-labels are sequentially placed into the labeled data pool. The first target model generates first pseudo-labels to obtain pseudo-labeled samples. The confidence levels of these pseudo-labeled samples are calculated. High-confidence samples are used to expand the labeled samples to construct an extended training set. For low-confidence samples, the model is retrained and pseudo-label predictions are repeated.
[0046] When training a model, the execution process also includes: Apply the trained model to , generate pseudo labels for each unlabeled sample, and form a pseudo label dataset based on the pseudo labels of each unlabeled sample .
[0047] By calculating the confidence of the pseudo-label dataset The quality of each pseudo-label sample in is evaluated, and the confidence calculation process is as follows: Calculate pseudo-labeled samples The weighted Euclidean distance to all labeled samples is sorted and the nearest one is selected Neighbors As a comparison. Its weighted distance is .
[0048]
[0049] 、 denote the pseudo-labeled samples and the labeled samples, respectively. Features.
[0050] In order to enhance the importance of closer samples and suppress the interference of distant samples in the pseudo-label evaluation process, a distance-based weighted absolute error is proposed. :
[0051] is the minimum value; the standard deviation of the pseudo-label sample and the nearest neighbor is normalized:
[0052] in, Represents a pseudo-labeled dataset The largest standard deviation.
[0053] Confidence The formula is:
[0054] The confidence levels are sorted from high to low, with C1 being high confidence, C2 being medium confidence, and C3 being low confidence. The samples whose confidence levels in the pseudo-label dataset exceed a preset threshold (preferably 0.9) are considered valid data samples. , and use it as an auxiliary training sample With the original labeled data Combine to form an expanded training set , and execute steps S103 to S104 in a loop. If there is no valid data sample, that is , then increase hidden layer nodes to improve the model's fitting ability, that is, the current maximum number of hidden layer nodes is , then retrain the model and use it on unlabeled data Regenerate the label. Increase to the preset upper limit , and there is still no valid data sample, delete the unlabeled sample with confidence C3 and return to step S104. When there is no unlabeled data in the candidate list C, and If there is still no valid data sample, the iterative training process is stopped.
[0055] S105: Using the extended training set to train the second initial model to obtain a second target model.
[0056] Specifically, after constructing the expanded training set, it is used to train a second initial model, which is then tested on the test set. This training and testing process yields a second target model with improved performance. This model leverages not only the original labeled dataset but also the additional data points obtained through data augmentation, making it better able to handle unseen data and produce more accurate predictions.
[0057] When training a model, the execution process also includes: Using the expanded dataset (including labeled data and auxiliary training data) to train the new model.
[0058] In summary, this series of steps constitutes a semi-supervised learning framework, which enhances the performance of the model by utilizing unlabeled data, thereby improving the accuracy and reliability of data prediction.
[0059] In an optional embodiment, the method further comprises: The overflow concentration during the grinding process is predicted using the second target model.
[0060] Specifically, overflow concentration is a critical parameter in the grinding process, directly reflecting the quality of grinding and the processing capacity of subsequent processes. Accurately predicting overflow concentration allows for timely adjustment of grinding parameters, optimization of the grinding process, and improved production efficiency. Real-time data from the grinding process (including various input variables such as ore properties, mill speed, and feed rate) is fed into the secondary objective model. Based on these input variables and previously learned knowledge, the model calculates a predicted overflow concentration for the current grinding process.
[0061] In an alternative embodiment, see Figure 2 As shown, Figure 2 A flowchart of a dataset construction method provided in the first embodiment of the present invention is shown, wherein the method of constructing a labeled dataset and an unlabeled dataset based on the input variables and output variables of each labeled data and the input variables of each unlabeled data includes steps S201 to S203: S201: Determine the weight value of each input variable according to the variable characteristics of the input variable and output variable of each labeled data.
[0062] Specifically, we analyze the characteristics of the input and output variables in the labeled data. These characteristics may include the variable's numerical range, distribution pattern, and correlation with other variables. Based on these characteristics, we can use various methods to determine the weight of each input variable. The weight reflects the variable's importance in predicting the output variable. Common methods include statistical methods (such as correlation coefficients), machine learning methods (such as feature selection algorithms), and the empirical judgment of domain experts.
[0063] S202: performing weighted processing on each labeled data and each unlabeled data based on the weight value of each input variable.
[0064] Specifically, the labeled and unlabeled data are weighted according to these weights. The goal of this weighting is to emphasize those input variables that are more important for predicting the output variable, while relatively deemphasizing those that are less important. This can be achieved by adjusting the contribution of data points during training, for example, by multiplying each input variable by a corresponding weight in the loss function.
[0065] S203: Constructing a labeled dataset and an unlabeled dataset based on the weighted labeled data and unlabeled data respectively.
[0066] Specifically, labeled and unlabeled datasets are constructed based on the weighted labeled and unlabeled data. These weighted datasets are used in subsequent steps such as model training and candidate list construction. Because the weighted processing allows the data points in the dataset to reflect the varying importance of the input variables, it is expected to improve the model's sensitivity to key features and predictive accuracy.
[0067] In an alternative embodiment, see Figure 3 As shown, Figure 3 A flowchart of a candidate list construction method provided by the first embodiment of the present invention is shown, wherein the candidate list is constructed based on the labeled data set and the unlabeled data set, including steps S301 to S303: S301: performing clustering processing on the labeled data set to obtain a plurality of clusters.
[0068] Specifically, a clustering algorithm is used to process the labeled dataset. The goal of clustering is to group the sample points in the labeled dataset into clusters based on their feature similarities. Sample points within each cluster are relatively close in feature space, while sample points between different clusters are relatively far apart. Common clustering algorithms include K-means, hierarchical clustering, and DBSCAN. The choice of clustering algorithm depends on the characteristics of the data, the shape of the clusters, and the desired clustering results.
[0069] S302: Calculate the distance between the centroid of each cluster and the unlabeled data set in the feature space.
[0070] Specifically, the distance between the centroid of each cluster (i.e., the mean vector of all points within the cluster) and each point in the unlabeled dataset is calculated in feature space. This distance can be expressed using various metrics, such as Euclidean distance, Manhattan distance, and cosine similarity. The goal of calculating distance is to identify the unlabeled data points closest to each cluster; these points are likely to have similar characteristics to the points within the cluster.
[0071] S303: Construct the candidate list based on the minimum distance corresponding to the centroid of each cluster.
[0072] Specifically, we construct a candidate list based on the minimum distance between each cluster's centroid and a sample point in the unlabeled dataset. Specifically, for each cluster, we find the unlabeled data points with the minimum distance to its centroid and add these points to the candidate list. Points in the candidate list may have similar characteristics to some sample points in the labeled dataset, making them potential unlabeled data points worthy of further attention.
[0073] In an alternative embodiment, see Figure 4 As shown, Figure 4 A flowchart of a method for constructing an extended training set provided by the first embodiment of the present invention is shown, wherein the method for constructing the extended training set according to the first pseudo label of each target data includes steps S401 to S402: S401: Calculate the confidence of the first pseudo label of each target data.
[0074] Specifically, the reliability of each first pseudo-label is evaluated by calculating its confidence. The confidence calculation method may vary depending on the application scenario and the model used. The model's output probability (for classification tasks) or the distance near the decision boundary (for regression tasks) can be used as a confidence metric.
[0075] S402: Construct the extended training set according to the confidence of the first pseudo label of each target data.
[0076] Specifically, a confidence threshold is set in advance, and only target data with confidence higher than the threshold and their first pseudo labels are selected to construct the extended training set. The purpose is to ensure that only relatively reliable data points are included to reduce the impact of noise and incorrect labels on model training.
[0077] In an alternative embodiment, see Figure 5 As shown, Figure 5 A flowchart of a specific method for constructing an extended training set provided by the first embodiment of the present invention is shown, wherein the method for constructing the extended training set according to the confidence of the first pseudo label of each target data includes steps S501 to S504: S501: Determine whether the confidence level of the first pseudo label of each target data exceeds a preset threshold.
[0078] Specifically, a confidence threshold is preset, which reflects the minimum requirement for the reliability of the prediction results.
[0079] S502: If there is valid target data whose confidence level of the first pseudo label exceeds the preset threshold, construct the extended training set based on each valid target data.
[0080] Specifically, if there is target data whose first pseudo-label confidence exceeds a preset threshold, these target data are regarded as valid target data. The confidence of these data is high, so their contribution to model training is relatively reliable.
[0081] S503: If the valid target data does not exist, the model parameters of the first initial model are updated to obtain a second initial model, and the second initial model is trained using the labeled data set to obtain a second target model.
[0082] Specifically, if there is no valid target data with a confidence level exceeding a preset threshold, this means that the current first initial model may not be sufficient to accurately predict the label of the target data. The model parameters of the first initial model are then updated to attempt to improve the model's performance. This update involves adjusting the model's hyperparameters, changing the model structure, or employing other optimization strategies. The updated model is called the second initial model, and it is trained using the original labeled dataset to obtain the second target model.
[0083] S504: Determine a second pseudo label for each target data in the target data set by using the second target model, and construct the extended training set according to the second pseudo label for each target data.
[0084] Specifically, after obtaining the second target model, prediction is performed again for each target data in the target dataset to obtain a second pseudo label. Since the second target model may have been improved and optimized, it may more accurately predict the true label of the target data. Therefore, an extended training set is constructed based on the second pseudo label of the target data.
[0085] In an alternative embodiment, see Figure 6 As shown, Figure 6 The flowchart of the second specific method for constructing an extended training set provided by the first embodiment of the present invention is shown, wherein the method for constructing the extended training set based on each valid target data includes steps S601 to S602: S601: Construct an auxiliary training set based on each valid target data.
[0086] Specifically, only valid target data whose first pseudo-label confidence exceeds a preset threshold are used. The valid target data and their corresponding high-confidence pseudo-labels are collected to form an auxiliary training set.
[0087] S602: Construct the extended training set based on the auxiliary training set and the labeled data set.
[0088] Specifically, the auxiliary training set is combined with the original labeled dataset to construct an extended training set. The labeled dataset contains real data with known labels, which provides supervision during model training. The auxiliary training set contains high-confidence pseudo-labeled data predicted by the model. This data serves as additional training samples, helping the model learn more features and information. By combining these two datasets, the model's generalization ability can be enhanced, while also leveraging the useful information in the unlabeled data to improve model performance.
[0089] The grinding process of a copper mine is studied using the model training method for overflow concentration of the grinding process provided by this application. Grinding mainly adopts the method of wet grinding, which helps grinding by water or other media. First, the raw ore is placed on No. 1 iron plate, No. 2 iron plate and No. 3 iron plate respectively, and the iron plate is placed on No. 1 belt conveyor to transport the raw ore to the semi-autogenous mill via No. 3 belt conveyor, and at the same time, the water pipeline adds water to the semi-autogenous mill for grinding. After the ground ore is vibrated and screened by a vibrating screen, the hard-to-grind stubborn stone is conveyed to the stubborn stone silo via No. 5 belt conveyor, and then transported to the cone crusher for crushing. After crushing, it can be sent to the semi-autogenous mill for grinding again. In addition, the ore pulp is tested for concentration using an overflow concentration tester, and the ore pulp is transported to No. 1 cyclone and No. 2 cyclone for classification. The ore pulp that does not meet the particle size requirements is discharged through the sand settling port and enters the ball mill for regrinding. After regrinding, the slurry flows back to the pump sump for subsequent circulation. Meanwhile, minerals meeting the required particle size are discharged as overflow and transported to the pump sump by a mortar pump, forming overflow slurry. The overflow concentration reflects the system's classification efficiency and operating status and also affects ore recovery during flotation. This slurry is then transported to the flotation process for further processing.
[0090] In the research process, the data set consists of 808 labeled samples, 3232 unlabeled samples, 22 input variables and the output variable overflow concentration. The input variables are shown in Table 1. Table 1 shows the sequence number, name and unit of each input variable of the grinding process.
[0091]
[0092] Table 1 Traditional soft sensing models suffer from the following issues: they ignore valuable information in unlabeled data. In this process, the target variable is measured at 5-minute intervals using an overflow concentration analyzer; other process variables, such as feed rate, ball mill power, and slag pool level, are measured at 1-minute intervals. The large number of unlabeled samples in historical data may contain valuable information. Therefore, it is necessary to develop a semi-supervised soft sensing sensor to address this issue.
[0093] In this application, the first 80% of the dataset is used for model training, and the remaining 20% is used for testing. This application compares its performance with that of the Stochastic Configuration Network (SCN), Multilayer Perceptron (MLP), Locality Preserving Stochastic Configuration Network (LPSCN), Locality Preserving Multilayer Perceptron (LPMLP), and Incremental Self-Training Semi-Supervised Multilayer Perceptron (ISTSS-MLP). Algorithm performance is evaluated using mean squared error (MSE), mean absolute error (MAE), and coefficient of determination (R-squared, R²).
[0094]
[0095] in, Indicates the number of samples in the test set; and 、 Respectively represent The true value, predicted value and mean of the true value of the samples.
[0096] The table shows the average results of 10 simulation experiments, and gives the mean, standard deviation, and best performance of the developed and compared algorithms in terms of MSE, MAE, and R². Table 2 shows the simulation results of overflow concentration prediction using different algorithms.
[0097]
[0098] Table 2 Example 2 The second embodiment of the present invention provides a model training device for overflow concentration in a grinding process, see Figure 7 As shown, Figure 7 A schematic structural diagram of a model training device for overflow concentration in a grinding process provided by a second embodiment of the present invention is shown, wherein the device comprises: A data set construction module 701 is used to construct a labeled data set and an unlabeled data set based on the input variables and output variables of each labeled data and the input variables of each unlabeled data; A candidate list building module 702 is configured to build a candidate list based on the labeled dataset and the unlabeled dataset; A data set expansion module 703 is configured to train a first initial model using the labeled data set to obtain a first target model, and to expand the unlabeled data set using the candidate list to obtain a target data set; A training set construction module 704 is configured to determine a first pseudo label for each target data in the target data set using the first target model, and construct an extended training set based on the first pseudo label for each target data; The model training module 705 is used to train the second initial model using the extended training set to obtain a second target model.
[0099] In an optional embodiment, the device further comprises: The overflow concentration prediction module is used to predict the overflow concentration during the grinding process using the second target model.
[0100] In an optional embodiment, constructing the labeled data set and the unlabeled data set based on the input variables and output variables of each labeled data and the input variables of each unlabeled data includes: Determine the weight value of each input variable according to the variable characteristics of the input variables and output variables of each labeled data; Perform weighted processing on each labeled data and each unlabeled data based on the weight value of each input variable; A labeled dataset and an unlabeled dataset are constructed based on the weighted labeled data and unlabeled data respectively.
[0101] In an optional embodiment, constructing a candidate list based on the labeled dataset and the unlabeled dataset includes: Performing clustering processing on the labeled data set to obtain a number of clusters; Calculating the distance between the centroid of each cluster and the unlabeled data set in the feature space; The candidate list is constructed based on the minimum distance corresponding to the centroid of each cluster.
[0102] In an optional embodiment, constructing an extended training set based on the first pseudo label of each target data includes: Calculate the confidence of the first pseudo label of each target data; The extended training set is constructed according to the confidence of the first pseudo label of each target data.
[0103] In an optional embodiment, constructing the extended training set according to the confidence of the first pseudo label of each target data includes: Determine whether the confidence of the first pseudo label of each target data exceeds a preset threshold; If there is valid target data whose confidence level of the first pseudo label exceeds the preset threshold, constructing the extended training set based on each valid target data; If the valid target data does not exist, updating the model parameters of the first initial model to obtain a second initial model, and training the second initial model using the labeled data set to obtain a second target model; The second target model is used to determine a second pseudo label for each target data in the target data set, and the extended training set is constructed according to the second pseudo label for each target data.
[0104] In an optional embodiment, constructing the extended training set based on each valid target data includes: Construct auxiliary training sets based on each valid target data; The extended training set is constructed based on the auxiliary training set and the labeled dataset.
[0105] Example 3 Based on the same application concept, see Figure 8 As shown, Figure 8 FIG. 1 shows a schematic diagram of the structure of a computer device provided by the third embodiment of the present invention, wherein Figure 8 As shown, a computer device 800 provided in the third embodiment of the present application includes: A processor 801, a memory 802 and a bus 803, wherein the memory 802 stores machine-readable instructions executable by the processor 801. When the computer device 800 is running, the processor 801 communicates with the memory 802 via the bus 803. When the processor 801 is running, the machine-readable instructions execute the steps of the model training method for overflow concentration of the grinding process shown in the above-mentioned embodiment 1.
[0106] Example 4 Based on the same application concept, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the model training method for overflow concentration of the grinding process described in any one of the above embodiments are executed.
[0107] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0108] The computer program product for model training for overflow concentration in a grinding process provided in an embodiment of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the previous method embodiment. For specific implementation, please refer to the method embodiment and will not be repeated here.
[0109] The model training device for overflow concentration of the grinding process provided by the embodiment of the present invention can be specific hardware on the device or software or firmware installed on the device. The implementation principle and technical effects of the device provided by the embodiment of the present invention are the same as those of the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices and units described above can all refer to the corresponding processes in the aforementioned method embodiment, and will not be repeated here.
[0110] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some communication interface, the indirect coupling or communication connection of the device or unit may be electrical, mechanical or other forms.
[0111] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0112] In addition, each functional unit in the embodiment provided by the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0113] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0114] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance.
[0115] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. However, such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. They should all be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A model training method for overflow concentration in a grinding process, characterized in that: The method comprises: Constructing a labeled data set and an unlabeled data set based on the input variables and output variables of each labeled data and the input variables of each unlabeled data; constructing a candidate list based on the labeled dataset and the unlabeled dataset; Using the labeled data set to train a first initial model to obtain a first target model, and using the candidate list to expand the unlabeled data set to obtain a target data set; Determining a first pseudo label for each target data in the target data set by using the first target model, and constructing an extended training set according to the first pseudo label for each target data; The second initial model is trained using the extended training set to obtain a second target model.
2. The method according to claim 1, characterized in that The method further comprises: The overflow concentration during the grinding process is predicted using the second target model.
3. The method according to claim 1, characterized in that The step of constructing a labeled dataset and an unlabeled dataset based on the input variables and output variables of each labeled data and the input variables of each unlabeled data includes: Determine the weight value of each input variable according to the variable characteristics of the input variables and output variables of each labeled data; Perform weighted processing on each labeled data and each unlabeled data based on the weight value of each input variable; A labeled dataset and an unlabeled dataset are constructed based on the weighted labeled data and unlabeled data respectively.
4. The method according to claim 1, wherein The constructing a candidate list based on the labeled data set and the unlabeled data set includes: Performing clustering processing on the labeled data set to obtain a number of clusters; Calculating the distance between the centroid of each cluster and the unlabeled data set in the feature space; The candidate list is constructed based on the minimum distance corresponding to the centroid of each cluster.
5. The method according to claim 1, wherein The step of constructing an extended training set according to the first pseudo labels of each target data includes: Calculate the confidence of the first pseudo label of each target data; The extended training set is constructed according to the confidence of the first pseudo label of each target data.
6. The method according to claim 5, characterized in that The step of constructing the extended training set according to the confidence of the first pseudo label of each target data includes: Determine whether the confidence of the first pseudo label of each target data exceeds a preset threshold; If there is valid target data whose confidence level of the first pseudo label exceeds the preset threshold, constructing the extended training set based on each valid target data; If the valid target data does not exist, updating the model parameters of the first initial model to obtain a second initial model, and training the second initial model using the labeled data set to obtain a second target model; The second target model is used to determine a second pseudo label for each target data in the target data set, and the extended training set is constructed according to the second pseudo label for each target data.
7. The method according to claim 6, characterized in that The step of constructing the extended training set based on each valid target data includes: Construct auxiliary training sets based on each valid target data; The extended training set is constructed based on the auxiliary training set and the labeled dataset.
8. A model training device for overflow concentration in a grinding process, characterized in that: The device comprises: A data set construction module, configured to construct a labeled data set and an unlabeled data set based on the input variables and output variables of each labeled data and the input variables of each unlabeled data; a candidate list building module, configured to build a candidate list based on the labeled data set and the unlabeled data set; a data set expansion module, configured to train a first initial model using the labeled data set to obtain a first target model, and to expand the unlabeled data set using the candidate list to obtain a target data set; a training set construction module, configured to determine a first pseudo label for each target data in the target data set by using the first target model, and construct an extended training set according to the first pseudo label for each target data; The model training module is used to train the second initial model using the extended training set to obtain a second target model.
9. A computer device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor and the memory communicate via the bus. When the machine-readable instructions are executed by the processor, the steps of the model training method for overflow concentration of a grinding process as described in any one of claims 1 to 7 are performed.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the model training method for overflow concentration in a grinding process as claimed in any one of claims 1 to 7.
Citation Information
Patent Citations
Classification model training method and device
CN112488237A
Neural network training method and device
CN113705769A
Incremental self-training framework and semi-supervised width learning classification method
CN114722908A
Optical device parameter determination method and device based on machine learning
CN117933411A
Model adaptive updating method and device during model reasoning
CN119360125A