A loom abnormal data processing method based on probability distribution and XGBoost decision algorithm

By using probability distribution and the XGBoost decision algorithm to process abnormal data from looms, the problem of missing data caused by abnormal data in textile production was solved. Data cleaning and missing data repair were achieved, improving the accuracy of loom production information and the quality of fabrics, and increasing enterprise efficiency.

CN116467653BActive Publication Date: 2026-06-02ZHEJIANG SCI-TECH UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG SCI-TECH UNIV
Filing Date
2023-04-03
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

The textile production environment is complex, and factors such as network communication failures and damage to measurement sensor hardware result in a large amount of abnormal and outlier data mixed in with textile big data, causing data loss, low availability, inaccurate production information, and affecting the quality of woven fabrics.

Method used

A method for processing abnormal loom data based on probability distribution and XGBoost decision algorithm is adopted. The confidence interval is determined by adaptive regression benchmark threshold, a Bayesian network abnormal data identification model is constructed to remove abnormal data points, and a missing data repair model based on XGBoost decision method is constructed for repair.

Benefits of technology

It improved the accuracy of loom production information and the quality of woven fabrics, enhanced enterprise production efficiency, and increased data availability and repair success rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116467653B_ABST
    Figure CN116467653B_ABST
Patent Text Reader

Abstract

The application discloses a loom abnormal data processing method based on a probability distribution and an XGBoost decision algorithm, and belongs to the field of intelligent manufacturing of weaving workshops in the textile industry, and the method comprises the following steps: collecting loom original data at regular time intervals; calculating the change difference between each adjacent data point; calculating a self-adaptive regression reference threshold under different time windows; updating the confidence interval at different time according to the self-adaptive regression reference threshold; determining the data points outside the confidence interval range as abnormal data; constructing a Bayesian network abnormal data identification model based on a probability distribution; determining the loom parameters causing the abnormal data points to generate data abnormalities through the Bayesian network abnormal data identification model based on the probability distribution; constructing a loom missing data repair model based on an XGBoost decision method; training the loom missing data repair model based on the XGBoost decision method; and repairing the abnormal data points through the trained loom missing data repair model based on the XGBoost decision method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent manufacturing in textile weaving workshops, and specifically relates to a method for processing abnormal data of weaving machines based on probability distribution and XGBoost decision algorithm. Background Technology

[0002] Significant progress has been made in intelligent manufacturing within my country's textile industry. Based on the Internet of Things (IoT), achieving high-quality data interconnection in the textile sector is a prerequisite for data-driven analysis, intelligent production scheduling, and other intelligent manufacturing processes. Therefore, it is necessary to analyze the causes and forms of abnormal textile data, and then develop abnormal data processing methods tailored to specific textile production scenarios. Among textile equipment, the loom's process is complex and generates a vast amount of operational information. As the final processing equipment in fabric formation, its production status directly affects the final quality of the woven fabric. Therefore, improving the accuracy of weaving data and processing abnormal loom data is crucial for achieving high-quality development in intelligent textile manufacturing.

[0003] Currently, the processing of abnormal data mainly includes cleaning and repairing the data. This involves analyzing the characteristics of abnormal data in the industry to identify and clean the data. However, due to the complex textile production environment, network communication failures, and damage to measurement sensor hardware, a large amount of abnormal and outlier data is mixed in with textile big data. This results in data gaps, low usability, inaccurate production information, and ultimately affects the quality of woven fabrics, for which there is no effective method to address the problem. Summary of the Invention

[0004] The purpose of this invention is to provide a method for processing abnormal data of looms based on probability distribution and XGBoost decision algorithm. This method can solve the technical problem that a large number of abnormal and outlier data are mixed in with textile big data due to the complex textile production environment, network communication failures and damage to measurement sensor hardware, resulting in data loss, low availability, inaccurate production information, and ultimately affecting the quality of woven fabrics.

[0005] To solve the above-mentioned technical problems, the present invention is implemented as follows:

[0006] This invention provides a method for processing loom anomaly data based on probability distribution and the XGBoost decision algorithm, including:

[0007] S101: Timed acquisition of raw data from the loom;

[0008] S102: Calculate the difference in change between adjacent data points;

[0009] S103: Calculate the adaptive regression benchmark threshold under different time windows;

[0010] S104: Update the confidence interval at different times based on the adaptive regression baseline threshold;

[0011] S105: Data points whose values ​​are outside the confidence interval are identified as;

[0012] S106: Construct a Bayesian network-based model for identifying anomalous data based on probability distribution;

[0013] S107: Using a Bayesian network anomaly data identification model based on probability distribution, determine the loom parameters that cause data anomalies in anomalous data points;

[0014] S108: Construct a loom missing data repair model based on XGBoost decision method;

[0015] S109: Training a loom missing data repair model based on XGBoost decision method;

[0016] S110: Abnormal data points are repaired using a loom missing data repair model based on the trained XGBoost decision method.

[0017] In this embodiment of the invention, firstly, an adaptive regression threshold is set to determine the confidence interval of changes in each original data of the loom, narrowing the scope for locating abnormal data. Then, an abnormal data identification model based on probability distribution is constructed to further accurately locate abnormal data. This Bayesian network abnormal data identification model based on probability distribution determines the loom parameters that cause the abnormal data points to produce data anomalies. Finally, a loom missing data repair model based on the XGBoost decision method is constructed. The XGBoost decision method is used to repair data missing data caused by abnormal data, simultaneously cleaning and filling in missing data. This addresses the technical problem of low usability of missing data in sampled data, improves the accuracy of loom production information and the quality of the final woven fabric, and significantly enhances enterprise production efficiency. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a method for processing abnormal loom data based on probability distribution and XGBoost decision algorithm provided in an embodiment of the present invention.

[0019] Figure 2 This is a network relationship diagram for identifying abnormal data of a loom provided in an embodiment of the present invention.

[0020] The realization of the objective, functional characteristics and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0022] The following description, in conjunction with the accompanying drawings, details the loom anomaly data processing method based on probability distribution and XGBoost decision algorithm provided by the present invention through specific embodiments and application scenarios.

[0023] Reference Figure 1 The diagram shows a flowchart of a loom anomaly data processing method based on probability distribution and XGBoost decision algorithm provided by an embodiment of the present invention.

[0024] This invention provides a method for processing loom anomaly data based on probability distribution and XGBoost decision algorithm, comprising:

[0025] S101: Timed acquisition of raw data from the loom.

[0026] It should be noted that during the operation of the loom and the network data collection process, issues such as communication abnormalities, sensor malfunctions, and unstable network signals can lead to the presence of abnormal interference data in the raw collected information, reducing the reliability of the collected data. Therefore, the collected raw data from the loom contains both abnormal and normal data, and this raw data provides the data foundation for subsequent processing of abnormal data. Those skilled in the art can set their own timing methods for collecting the raw data from the loom; this solution does not limit the timing method.

[0027] Table 1

[0028]

[0029]

[0030] It should be noted that the function of a loom is to interweave the warp yarns of the warp beam and the weft yarns of the weft bobbin to form fabric. Due to the complexity of the loom's processing and the long processing time for each warp beam, the amount of data generated during loom operation is enormous. Loom data can be divided into two categories: static data and dynamic data. Static data refers to data that can be set and predicted in advance and will not change within a short period, such as equipment attributes and product process information. Dynamic data refers to data that changes in real time according to the equipment's working status and processing requirements, such as loom production output, weft insertion count, loom status, running time, running speed, running efficiency, and warp and weft stop counts. Of these two types of data, dynamic data changes more frequently and reflects the overall production status of the loom. During the actual operation of the loom, the magnitude of change among the various dynamic data varies, but they are all temporally correlated, and the changes among the data are interconnected. The correlation between various loom parameters and time is shown in Table 1. The fabric production of the loom during normal operation will show a continuous growth trend over time. Due to the short transition time of the loom's state change, the collected values ​​of the loom speed will show short-term discrete low-frequency jumps.

[0031] In actual operation, the changes in various data are interconnected. The fabric output per unit time of the loom maintains a synchronous and positive correlation with the loom speed. When the loom is running, the output value should also increase synchronously, and the loom speed should be greater than 0. Similarly, when the loom is stopped, the output value should remain unchanged, and the loom speed should be equal to 0. The theoretical relationship between fabric output, running time, running speed, and running status is as follows:

[0032]

[0033]

[0034] Where M represents the weaving length of the loom in cm, T represents the weaving time of the loom in min, W represents the weft density parameter value of the fabric, that is, the number of weft yarns per 1 cm of fabric, S represents the running speed of the loom when weaving, that is, the number of weft beats per minute of the loom, and K represents the current working status of the loom, where 1 is running and 0 is stopped.

[0035] S102: Calculate the difference in change between adjacent data points.

[0036] It should be noted that due to the influence of various factors during the operation of the loom, the data values ​​between any two adjacent data points in the collected raw data cannot be exactly the same. By calculating the difference in change between adjacent data points, theoretically, if the working conditions are the same during continuous operation of the loom, the difference in change between adjacent data will remain within a certain range. Therefore, by comparing the difference in change between adjacent data, the location of abnormal data can be initially located, laying the foundation for subsequent accurate location of abnormal data.

[0037] S103: Calculate the adaptive regression benchmark threshold under different time windows.

[0038] It should be noted that, compared with existing technologies, the adaptive regression threshold can automatically correct the baseline threshold at different time periods as the original data is continuously updated. This avoids the situation where normal data fluctuations are misjudged as abnormal data due to intermittent communication failures between the acquisition end and the underlying device terminal, thereby improving the accuracy of abnormal data location.

[0039] Table 2

[0040]

[0041] As shown in Table 2, when calculating the adaptive regression benchmark threshold, the update time of each data value is taken into account in the calculation, resulting in a data change difference sequence within a unit time. Then, a moving average is applied to the change difference sequence of time-series features under an n-dimensional time window to finally obtain the adaptive regression benchmark threshold under different time windows.

[0042] In one possible implementation, S103 specifically includes:

[0043] S1031: Obtain the data change difference sequence within a unit of time;

[0044] S1032: Calculate the moving average of the data change differences in the current time window in chronological order, and use it as the regression benchmark threshold in the current time window. The current time window is an n-dimensional time window, which includes n units of time.

[0045] The regression benchmark threshold is calculated as follows:

[0046]

[0047] Where F(i) represents the baseline threshold in the i-th time window, x k t represents the data value of the parameter at time k. k This represents the data update time of the parameter at time k, and n represents the length of the time window.

[0048] It should be noted that n units of time constitute a time window. The first time window refers to the first to n units of time, and the second time window refers to the (n+1)th to 2nth units of time.

[0049] S104: Update the confidence interval at different times based on the adaptive regression benchmark threshold.

[0050] In one possible implementation, the upper and lower bounds of the confidence interval are respectively:

[0051]

[0052]

[0053] Among them, H t L represents the upper limit of the confidence interval at time t. t y represents the lower bound of the confidence interval at time t. t-a Let t represent the valid data value before time t, and C represent the amplitude coefficient.

[0054] It should be noted that, based on the adaptive regression baseline threshold, the degree to which different data values ​​deviate from the baseline threshold can be determined. By calculating the upper and lower limits of the confidence interval, it can be determined that data values ​​within the range of the upper and lower limits of the confidence interval are normal data values, which is equivalent to determining normal data values. Using this method, not only can the probability of normal data being misjudged as abnormal data be avoided, but the collected raw data can also be initially distinguished into normal data and abnormal data.

[0055] S105: Data points whose values ​​are outside the confidence interval are identified as anomalous data.

[0056] It is understandable that the upper and lower limits of the confidence interval calculated in S104 can be used to determine whether the data value is normal by whether the data value is within the confidence interval. In other words, data values ​​that are not within the confidence interval are abnormal data values, and the data points corresponding to the abnormal data values ​​can be identified as abnormal data, thus initially locating the location of the abnormal data.

[0057] S106: Construct a Bayesian network-based model for identifying anomalous data based on probability distribution.

[0058] In one possible implementation, S106 specifically includes:

[0059] S1061: Establish a sample dataset D, which includes fabric production (x1), weft insertion count (x2), running time (x3), running efficiency (x4), running speed (x5), loom status (x6), and data type (x7). The composition relationship of the sample dataset D is as follows:

[0060]

[0061] Where D has a dimension of i×7, and i represents the number of time-series data acquisition points for the loom data.

[0062] S1062: The sample dataset is discretized using the equal-frequency binning method;

[0063] It should be noted that, due to the superior processing performance of Bayesian networks on discrete data, the sample elements of the sample dataset D need to be discretized. The loom status is a binary element reflecting the loom's operation and shutdown, while the data type is a ternary element representing whether the loom data at each time point belongs to normal data points, outliers, or inactive outliers; these do not require discretization. The equal-frequency binning discretization method is used to discretize the five types of data—fabric production, weft insertion count, running time, running efficiency, and loom speed—to facilitate precise further location of outliers.

[0064] Specifically, the equal-frequency bin discretization method is used to discretize the fabric production (x1), weft insertion times (x2), running time (x3), running efficiency (x4), and running speed (x5).

[0065] In one possible implementation, the process after S1062 further includes:

[0066] S1062A: Introduces correlation coefficients between parameters in a Bayesian network. The correlation coefficients are calculated as follows:

[0067]

[0068] Where a and b represent the elements whose correlation needs to be determined, cov represents the covariance between the elements, σ ​​represents the standard deviation between the elements, and ρ(a,b) represents the correlation coefficient between elements a and b.

[0069] It should be noted that the correlation coefficients between various parameters of the loom are the basis for determining the relationship structure of the Bayesian network. This paper uses the Pearson correlation function to perform correlation analysis on each network node in order to determine the relationship structure of the probability distribution network.

[0070] S1063: Determine the network structure of the Bayesian network anomaly data identification model based on probability distribution;

[0071] The network structure is a Bayesian network, and the Bayesian network relation is:

[0072] B = (J, T)

[0073] Where J represents the network structure graph describing the relationships between elements, including element nodes and relationship pointing lines, and T represents the relationship dataset describing the probability distribution between element nodes in the network.

[0074] S1064: Based on the prior probability P(x) under each child node i |D), i=1,2,···,6, train the probability distribution among each child node to determine the conditional probability distribution P(x7=m) between evidence nodes and child nodes in the loom anomaly identification network. j |x i ).

[0075] S1065: Calculate the total probability for each node:

[0076]

[0077] Where, P(x7=m) j ) represents the joint total probability of the conditions for each data type being true, that is, the probability of the outcome under the combined influence of six types of data: fabric production, number of weft insertions, running time, running efficiency, running speed, and loom status.

[0078] S1066: Determine the type of loom data based on the full probability distribution of each node.

[0079] S1067: When the output of a child node is abnormal data, infer the posterior probability of each evidence node based on the probability distribution of abnormal data, and finally locate the loom parameter that caused the data anomaly. The posterior probability is calculated as follows:

[0080]

[0081] Wherein, P(x i |x7) indicates that the data type of the child node is known, and the parent node is x. i The probability that the condition is true is called the posterior probability.

[0082] S1068: Input the sample dataset into the Bayesian network anomaly data identification model based on probability distribution for training.

[0083] S107: Using a Bayesian network anomaly data identification model based on probability distribution, determine the loom parameters that cause the data anomalies in the anomalous data points.

[0084] In one possible implementation, the process after S107 includes:

[0085] S111: Clear loom parameters that cause abnormal data points to generate data anomalies.

[0086] It should be noted that, after locating the abnormal data and the loom parameters, the exact location of the abnormal data was finally determined. The identified abnormal data was then cleared to avoid its impact on the subsequent repair process. After that, the missing data was repaired to improve the overall quality of the data.

[0087] S108: Construct a loom missing data repair model based on XGBoost decision method.

[0088] XGBoost, a decision tree ensemble algorithm based on extreme gradient boosting, uses gradient descent to integrate multiple base learners to gradually reduce the residual between the model's repair results and the actual loom values. Since loom data has multi-dimensional characteristics, compared to existing technologies, this paper utilizes XGBoost to fully leverage the correlations between various dimensions of loom data to repair missing data. While avoiding insufficient consideration during the repair process, it sets up multiple base learners to progressively repair missing data to approximate the actual loom values, thus improving the repair effect.

[0089] In one possible implementation, S108 specifically includes:

[0090] S1081: Construct a loom missing data repair model based on the XGBoost decision method using a regression tree as the base learner. The expression for the loom missing data repair model based on the XGBoost decision method is:

[0091]

[0092] in, x represents the loom data repair result at time i. i Let f represent the known associated input samples of the loom parameters to be repaired at time i, N represent the number of base learners, and f k This represents the k-th base learner.

[0093] S109: Training a loom missing data repair model based on the XGBoost decision method.

[0094] In one possible implementation, S109 specifically includes:

[0095] S1091: The normal data points of each parameter of the loom obtained after anomaly identification are used as feature samples of the loom missing data repair model based on XGBoost decision method. The feature sample set is constructed and divided into training set and test set according to a preset ratio.

[0096] It should be noted that using the normal data points of each parameter of the loom obtained after anomaly identification as the feature sample set of the driving model can eliminate the overfitting of the model caused by the high-dimensional sparse features of the original data and improve the data repair capability of the model.

[0097] Optionally, the default ratio of the training set to the test set is 7:3.

[0098] S1092: Select the regression tree as the base learner for the loom missing data repair model based on XGBoost decision method, and set the objective loss function of the loom missing data repair model based on XGBoost decision method as mean squared loss regression.

[0099] S1093: Adjust the learning rate, number of base learners, and regression tree depth of the loom missing data repair model based on the XGBoost decision method according to the principle of minimizing error.

[0100] It should be noted that the data restoration problem in this paper is a regression problem. When selecting the base learner (booster) and objective loss function of the model, a regression tree is chosen as the base learner for the loom missing data restoration model, and the objective loss function of the loom missing data restoration model is set to mean squared loss regression. The learning rate, the number of base learners (n_estimator), and the regression tree depth (max_depth) of the model are adjusted according to the principle of minimizing error. In order to prevent the model from overfitting the sample data during training, the proportion of random sampling (subsample) and the proportion of random features (colsample_bytree) when constructing the base learner are set to improve the restoration quality and reduce the restoration error.

[0101] S1094: Train the loom missing data repair model based on XGBoost decision method using the training set.

[0102] S1095: When the number of iterations of the loom missing data repair model based on XGBoost decision method reaches the number of base learners, output the output result of the loom missing data repair model based on XGBoost decision method.

[0103] S1096: Determine whether the output result is between the upper and lower limits of the confidence interval of each corresponding parameter of the loom. If the output result is between the upper and lower limits of the confidence interval, proceed to S1097; otherwise, proceed to S1098.

[0104] S1097: Perform outlier validation on the output result. Determine whether the output result passes the outlier validation. If the output result passes the outlier validation, proceed to S1099; otherwise, proceed to S1098.

[0105] S1098: Adjust the learning rate, number of base learners, and regression tree depth, return to S1094, and retrain the loom missing data repair model based on the XGBoost decision method.

[0106] It should be noted that the model is trained and constructed according to the initial parameters. When the number of training iterations reaches the maximum number of base learners, the model training results and errors are output. It is then determined whether the training results are within the confidence interval of the corresponding parameters of the loom. The outlier detection method is used again to check the results for outliers. If the check fails, the parameter values ​​of the model learning rate and the number of base learners are adjusted, and the model is retrained until the output results meet the condition of being within the confidence interval and can pass the outlier check. At this point, the parameter adjustment is completed, and the training of the loom missing data repair model based on the XGBoost decision method is finished.

[0107] S1099: Randomly input sample information associated with the loom data to be repaired in the test set, compare the deviation between the target output result and the corresponding real parameters of the loom, and end the training of the loom missing data repair model based on XGBoost decision method if the deviation between the target output result and the corresponding real parameters of the loom is within the preset range.

[0108] Understandably, the output of the previously trained loom missing data repair model based on the XGBoost decision method has indeed been validated. However, to avoid unexpected situations, the model will be further validated using a test set. If the validation results are successful, it indicates that the loom missing data repair model based on the XGBoost decision method has general applicability. Further validation will improve the reliability and credibility of the trained model.

[0109] S110: Abnormal data points are repaired using a loom missing data repair model based on the trained XGBoost decision method.

[0110] It is understandable that a well-trained loom missing data repair model based on the XGBoost decision method can be used to repair the located abnormal data points, fill in the missing positions, improve the availability of loom sampling data, and increase the success rate of abnormal data repair.

[0111] In one possible implementation, the process after S110 includes:

[0112] S112: Verify the reliability of the repaired data.

[0113] It should be noted that, in order to avoid inaccurate repair or failure to meet preset conditions after the abnormal data is located and repaired, the reliability of the repaired data is verified to verify the reliability of the repair effect.

[0114] In one possible implementation, S112 specifically includes:

[0115] S1121: Use the mean absolute error index (MAE), root mean square error index (RMSE), and fit coefficient index (R²). 2 The reliability of the repaired data was verified:

[0116]

[0117]

[0118]

[0119] Among them, y i This represents the actual value of the test sample in the test set. This represents the predicted value of the test sample. The mean of the test sample is represented by MAE, RMSE, and R. 2 The values ​​range from [0,1]. The smaller the values ​​of MAE and RMSE, the closer the repair result is to the true value. 2 The closer the value is to 1, the higher the accuracy of the loom missing data repair model based on XGBoost decision method in repairing loom missing data.

[0120] In practical use, most existing looms serving textile enterprises have information transmission interfaces for external data acquisition and communication, such as those from Toyota, Tsudakoma, and Picanol. This paper takes the data collection from a Picanol OMNIPLUS-340 air-jet loom in a textile enterprise in Shijiazhuang as an example to verify the effectiveness of the proposed abnormal data processing method. Existing data acquisition equipment mainly uses external data acquisition terminals to collect various production data from the loom. To align with the actual shift-based production schedule of the weaving workshop, this paper collects complete shift data of the loom, with a data acquisition frequency of once per minute. The collected dataset contains a total of 720 samples, each including seven categories of data: fabric production, weft insertion count, running time, operating efficiency, running speed, loom status, and abnormal conditions.

[0121] An adaptive regression threshold method was used to define the confidence intervals for the changes in various loom parameters. This paper uses four parameters—weaving output, weft insertion count, operating efficiency, and loom speed—which have a wide range of variable data and are easily affected by environmental interference, as validation data for the processing effect of this method. Analysis of the data variation characteristics shows that the instantaneous changes in weaving output and weft insertion count under normal conditions are relatively small, while the values ​​of loom speed and efficiency can fluctuate significantly in a short period due to the rapid switching between running and stopped states. Therefore, when processing loom speed and efficiency, the time window length needs to be longer than that for weaving output and weft insertion count. However, an excessively long time window will reduce the real-time update rate of the base threshold; an excessively short window will include too little historical data, making it difficult to reflect the data's changing trend. The figure shows the anomaly identification of various loom parameters under different time window lengths.

[0122] Experiments showed that the time window lengths for loom fabric production, weft insertion frequency, operating efficiency, and operating speed were set to 30, 23, 48, and 57, respectively. After determining the confidence intervals for the changes in each parameter data, 25 anomalies in fabric production were effectively identified, with a recognition rate of 32.05%; 17 anomalies in weft insertion frequency were effectively identified, with a recognition rate of 21.15%; 10 anomalies in operating efficiency were effectively identified, with a recognition rate of 13.33%; and 17 anomalies in operating speed were effectively identified, with a recognition rate of 26.98%. Furthermore, by using an adaptive regression threshold to define the confidence interval for data changes, the overall trend of loom parameters can be estimated, narrowing the scope of locating abnormal loom data and enabling preliminary identification of abnormal data. Among the anomalous data points identified solely by defining a trusted region, most are global deviation anomalies. Local deviation anomalies and inactive anomalous data points, due to their insignificant fluctuation amplitude, are included within the trusted region and thus cannot be effectively identified. This demonstrates the effectiveness of introducing an adaptive regression benchmark threshold to define a trusted interval and significantly improves the ability to identify anomalous data.

[0123] Table 3 Correlation coefficients among parameters

[0124] <![CDATA[x1]]> <![CDATA[x2]]> <![CDATA[x3]]> <![CDATA[x4]]> <![CDATA[x5]]> <![CDATA[x6]]> <![CDATA[x1]]> 1.00 0.97 0.96 0.34 0.09 0.08 <![CDATA[x2]]> 0.97 1.00 0.99 0.35 0.09 0.08 <![CDATA[x3]]> 0.96 0.99 1.00 0.34 0.09 0.08 <![CDATA[x4]]> 0.34 0.35 0.34 1.00 0.11 0.12 <![CDATA[x5]]> 0.09 0.09 0.09 0.11 1.00 0.93 <![CDATA[x6]]> 0.08 0.08 0.08 0.12 0.93 1.00

[0125] While adaptive regression thresholding can effectively locate global deviation anomalies, it is less effective at identifying local deviations and inactive anomalies with insignificant fluctuations. This paper constructs a Bayesian network anomaly identification model based on parameter correlation to further pinpoint the timing and type of anomalies. First, the correlations of six parameters—weaving output, weft insertion count, running time, running efficiency, running speed, and weaving machine status—are analyzed to determine the degree of interrelationship among these parameters, serving as the basis for constructing the anomaly identification network structure. The correlation coefficients among the parameters are shown in Table 3.

[0126] The correlation coefficients in Table 3 show that there is a correlation between loom running time, weft insertion frequency, and fabric output. The correlation between running time and weft insertion frequency is the strongest, with a correlation coefficient of 0.99. There is also a correlation between loom status and loom speed, with a correlation coefficient of 0.93.

[0127] Reference Figure 2 The diagram illustrates a network relationship diagram for identifying abnormal data of a loom provided by an embodiment of the present invention.

[0128] Depend on Figure 2 It can be seen that the abnormalities in the loom data are jointly determined by six types of evidence nodes: fabric production, number of weft insertions, running time, running efficiency, running speed, and loom status.

[0129] Table 4. Abnormal data identification results of each method

[0130]

[0131] To verify the effectiveness of the proposed identification network in identifying anomalous data from looms, the method presented in this paper is compared with that of decision trees and the K-nearest neighbor algorithm. Table 4 shows the anomaly identification results of each method for four types of parameters: fabric production, weft insertion frequency, operating efficiency, and loom speed. As shown in Table 4, the probability distribution-based anomaly identification method in this paper achieves anomaly identification rates of 98.71%, 95.00%, 98.67%, and 98.41% for the four types of loom parameters, respectively, with an average identification rate of 97.70%, which is higher than the other two comparison methods. This is because the proposed method fully utilizes the correlation between the data.

[0132] Table 5 XGBoost Model Parameters

[0133]

[0134] The XGBoost decision method was used to repair missing data points on the loom. The model parameters are shown in Table 5. To quantitatively analyze the reliability of the model in repairing missing data on the loom, the feature sample set was randomly split into a 7:3 ratio to form a training set and a test set. Hermite interpolation, cubic spline interpolation, and the method proposed in this paper were used to repair the missing data points on the loom. Five observation points at different times were randomly selected from each of the four types of parameters: loom production, weft insertion count, operating efficiency, and operating speed for data repair.

[0135] The results show that, comparing the fit between the imputed data values ​​and the actual data, the XGBoost decision method proposed in this paper provides a closer match to the actual values ​​of the loom parameters in repairing missing data. Furthermore, the two comparison methods show significant errors between the repaired values ​​and the actual values ​​at the observed points. This is because, compared to obtaining feature information from a single dimension, which relies solely on the trend of data changes, the XGBoost decision method considered the interrelationships between multiple loom parameters when repairing missing loom values. By utilizing existing correlation information of these parameters, the missing data can be repaired more accurately.

[0136] Table 6 Evaluation results of each parameter repair

[0137]

[0138] Table 6 shows the evaluation results of Hermite interpolation, cubic spline interpolation, and the method proposed in this paper for repairing various parameters of the loom. According to the evaluation indicators in the table, the method proposed in this paper outperforms the other two methods in repairing missing loom data, with the lowest MAE and RMSE values. The R values ​​corresponding to loom production, weft insertion frequency, operating efficiency, and operating speed are also significantly higher. 2 The values ​​were 0.9649, 0.9563, 0.9832, and 0.9736, respectively. After verifying the accuracy of the XGBoost decision-making method for repairing missing data of the loom, the trained repair model was used to repair the missing data of each parameter of the loom, thereby improving the overall quality of the loom data.

[0139] In this embodiment of the invention, firstly, an adaptive regression threshold is set to determine the confidence interval of changes in each original data of the loom, narrowing the scope for locating abnormal data. Then, an abnormal data identification model based on probability distribution is constructed to further accurately locate abnormal data. This Bayesian network abnormal data identification model based on probability distribution determines the loom parameters that cause the abnormal data points to produce data anomalies. Finally, a loom missing data repair model based on the XGBoost decision method is constructed. The XGBoost decision method is used to repair data missing data caused by abnormal data, simultaneously cleaning and filling in missing data. This addresses the technical problem of low usability of missing data in sampled data, improves the accuracy of loom production information and the quality of the final woven fabric, and significantly enhances enterprise production efficiency.

[0140] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A method for processing loom anomaly data based on probability distribution and XGBoost decision algorithm, characterized in that, include: S101: Timed acquisition of raw data from the loom; S102: Calculate the difference in change between adjacent data points; S103: Calculate the adaptive regression benchmark threshold under different time windows; S104: Update the confidence interval at different times according to the adaptive regression benchmark threshold; S105: Data points whose values ​​are outside the confidence interval are identified as anomalous data; S106: Construct a Bayesian network-based model for identifying anomalous data based on probability distribution; S107: Using the Bayesian network anomaly data identification model based on probability distribution, determine the loom parameters that cause the anomalies in the anomaly data points; S108: Construct a loom missing data repair model based on XGBoost decision method; S109: Train the loom missing data repair model based on the XGBoost decision method; S110: The abnormal data points are repaired using a loom missing data repair model based on the trained XGBoost decision method; Specifically, S103 includes: S1031: Obtain the data change difference sequence within a unit of time; S1032: Calculate the moving average of the data change differences in the current time window according to the time order, and use it as the regression benchmark threshold in the current time window, wherein the current time window is an n-dimensional time window and the current window includes n units of time; The regression benchmark threshold is calculated as follows: in, Indicates the first i The baseline threshold under each time window Indicates the parameter in k Data values ​​at any given time Indicates the parameter in k The data update time at any given moment n Indicates the length of the time window; Specifically, S104 is as follows: S1041: Based on the aforementioned adaptive regression benchmark threshold, calculate the upper and lower limits of the confidence interval for the loom parameters at time t: in, express t The upper limit of the confidence interval at time t, express t The lower bound of the confidence interval at time t, express t The last valid data value before the specified time. C This represents the amplitude coefficient.

2. The loom anomaly data processing method based on probability distribution and XGBoost decision algorithm according to claim 1, characterized in that, S106 specifically includes: S1061: Establish a sample dataset D The sample dataset D Including fabric production ( x 1) Number of weft insertions ( x 2) Running time ( x 3) Operating efficiency ( x 4) Operating speed ( x 5) Loom status x 6) and data types ( x 7), the sample dataset D The compositional relationship is as follows: in, D The dimension is , i This indicates the number of time-series data acquisition points for the loom; S1062: The sample dataset is discretized using the equal-frequency binning discretization method; S1063: Determine the network structure of the Bayesian network anomaly data identification model based on probability distribution; The network structure is a Bayesian network, and the Bayesian network relation is as follows: in, J A network structure diagram representing the relationships between elements, including element nodes and relationship lines. T This represents a dataset that describes the probability distribution among nodes in a network. S1064: Based on the prior probabilities of each child node The probability distributions between each child node are trained to determine the conditional probability distributions between evidence nodes and child nodes in the loom anomaly detection network. ; S1065: Calculate the total probability for each node: in, This represents the joint total probability of the conditions for each data type being true, namely the probability of the outcome under the combined influence of six types of data: fabric production, number of weft insertions, running time, running efficiency, running speed, and loom status. S1066: Determine the type of loom data based on the full probability distribution of each node; S1067: When the output of a child node is abnormal data, the posterior probability of each evidence node is inferred based on the probability distribution of the abnormal data, and the loom parameter that caused the data abnormality is finally located. The posterior probability is calculated as follows: in, Parent node represents the data type of the known result child node. The probability that the condition is true, i.e., the posterior probability; S1068: Input the sample dataset into the probability distribution-based Bayesian network anomaly data identification model for training.

3. The loom anomaly data processing method based on probability distribution and XGBoost decision algorithm according to claim 2, characterized in that, Following S1062, the following is also included: S1062A: In the Bayesian network, a correlation coefficient is introduced between the parameters. The correlation coefficient is calculated as follows: in, a , b This indicates that we need to find the elements with a certain degree of relevance. cov This represents the covariance value between elements. This represents the standard deviation between elements. Represents element a , b The correlation coefficient between them.

4. The loom anomaly data processing method based on probability distribution and XGBoost decision algorithm according to claim 1, characterized in that, Following S107, the following is also included: S111: Clear the loom parameters that cause the abnormal data points to generate data abnormalities.

5. The loom anomaly data processing method based on probability distribution and XGBoost decision algorithm according to claim 1, characterized in that, Specifically, S108 is: S1081: Construct the loom missing data repair model based on the XGBoost decision method using a regression tree as the base learner. The expression of the loom missing data repair model based on the XGBoost decision method is: in, express i The current loom data repair results express i Known associated input samples of the parameters of the loom to be repaired at the current moment. N This indicates the number of base learners. Indicates the first k Each base learner.

6. The loom anomaly data processing method based on probability distribution and XGBoost decision algorithm according to claim 5, characterized in that, S109 specifically includes: S1091: The normal data points of each parameter of the loom obtained after anomaly identification are used as feature samples of the loom missing data repair model based on XGBoost decision method. A feature sample set is constructed, and the feature sample set is divided into a training set and a test set according to a preset ratio. S1092: Select a regression tree as the base learner of the loom missing data repair model based on XGBoost decision method, and set the objective loss function of the loom missing data repair model based on XGBoost decision method as mean squared loss regression. S1093: Adjust the learning rate, number of base learners, and regression tree depth of the loom missing data repair model based on the XGBoost decision method according to the principle of minimizing error; S1094: Train the loom missing data repair model based on the XGBoost decision method using the training set; S1095: When the number of iterations of the loom missing data repair model based on XGBoost decision method reaches the number of base learners, output the output result of the loom missing data repair model based on XGBoost decision method. S1096: Determine whether the output result is between the upper and lower limits of the confidence interval of each corresponding parameter of the loom. If the output result is between the upper and lower limits of the confidence interval, proceed to S1097; otherwise, proceed to S1098. S1097: Perform outlier verification on the output result, and determine whether the output result passes the outlier verification. If the output result passes the outlier verification, proceed to S1099; otherwise, proceed to S1098. S1098: Adjust the learning rate, the number of base learners, and the depth of the regression tree, return to S1094, and retrain the loom missing data repair model based on the XGBoost decision method; S1099: Randomly input sample information associated with the loom data to be repaired in the test set, compare the deviation between the target output result and the corresponding real parameters of the loom, and end the training of the loom missing data repair model based on XGBoost decision method if the deviation between the target output result and the corresponding real parameters of the loom is within a preset range.

7. The loom anomaly data processing method based on probability distribution and XGBoost decision algorithm according to claim 1, characterized in that, Following S110, the following is also included: S112: Verify the reliability of the repaired data.

8. The loom anomaly data processing method based on probability distribution and XGBoost decision algorithm according to claim 7, characterized in that, Specifically, S112 is as follows: S1121: Use the mean absolute error index (MAE), root mean square error index (RMSE), and fitting coefficient index. The reliability of the repaired data was verified: in, This represents the actual value of the test sample in the test set. This represents the predicted value of the test sample. The mean, MAE, RMSE, and RMSE of the test samples represent the mean of the test samples. The values ​​range from [0,1]. The smaller the MAE and RMSE values, the closer the repair result is to the true value. The closer the value is to 1, the higher the accuracy of the loom missing data repair model based on XGBoost decision method in repairing loom missing data.