A data grading and classification method based on the entire supply chain process

By dynamically adjusting the optimal proximity number and weighted Euclidean distance, the problem of inaccurate LOF values in the prior art is solved, and the accurate grading of supply chain data is achieved, ensuring the reliability and accuracy of data grading.

CN120123953BActive Publication Date: 2025-07-18INSPUR SMART SUPPLY CHAIN TECH (SHANDONG) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510600787.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-18
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

In the data analysis of supply chain processes in the prior art, the problem of inaccurate LOF values by fixedly setting the nearest number will affect the accuracy of data grading.

Method used

By dynamically adjusting the optimal proximity number, combining the anomaly degree and correlation of data in each dimension, weighted Euclidean distance is calculated, and the optimal proximity number is determined to calculate the LOF value, thereby achieving accurate grading of product data.

Benefits of technology

It improves the calculation accuracy of LOF values, realizes the accurate division of abnormal levels of product data, avoids density estimation fluctuations caused by improper selection of k values, and ensures the reliability of data grading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123953B_ABST
    Figure CN120123953B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of supply chain data management, and specifically relates to a data grading and classification method based on the entire supply chain process. The method includes: obtaining the standardized product data of multiple products in the same batch in the supply chain process; the product data includes multiple dimension data; calculating the LOF value of each product data; performing data grading on the LOF values of all product data to obtain the graded product data. That is, the solution of the present invention can accurately grade the supply chain process data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of supply chain data management. More specifically, the present invention relates to a method for classifying and grading data based on the entire supply chain process. Background Art

[0002] In the entire supply chain process, data classification and grading is an important management method for better organizing, managing, and utilizing data to support decision-making, optimizing processes, and improving efficiency. Data grading in the product supply chain can convert massive amounts of data into actionable priority management bases by quantifying the risk value of data. For example, data with a low risk level can be directly archived without review, while data with a high risk level triggers an alarm, enabling managers to prioritize the processing of high-risk level data.

[0003] Among them, the data in the entire supply chain process includes financial data, production data, quality data, procurement data, and logistics data, etc.

[0004] In the calculation process of quantifying data risk, the LOF value calculated by using the LOF anomaly detection algorithm can be used as a parameter for measuring data risk. Among them, the LOF (Local Outlier Factor) anomaly detection algorithm is a density-based unsupervised anomaly detection algorithm. By comparing the density of data points with that of their neighboring data points, it can effectively identify outliers in a local area. Its specific steps include: (1) Selecting the number of nearest neighbors k; (2) Calculating the distance between each data point and all its neighbors; (3) Determining the reachable distance; (4) Calculating the local reachable density; (5) Calculating the LOF value.

[0005] Among them, the selection of the number of nearest neighbors in the LOF anomaly detection algorithm directly affects the reliability and accuracy of the calculation results. However, in the prior art, the number of nearest neighbors of all data points is generally set to a fixed value. Since the different-dimensional parameters in product data are different, the number of nearest neighbors may also be different. Therefore, the fixed setting of the number of nearest neighbors in the prior art will lead to inaccurate LOF values for different product data. Summary of the Invention

[0006] The object of the present invention is to propose a method for classifying and grading data based on the entire supply chain process to solve the problem of inaccuracy in the analysis of supply chain process data in the prior art. To this end, the present invention provides a solution in the following aspect.

[0007] A method for classifying and grading data based on the entire supply chain process provided by the present invention includes:

[0008] Obtain the standardized product data of multiple products in the same batch in the supply chain process; the product data includes multiple dimension data, and the data of all products in the same dimension constitutes a dimension sequence;

[0009] Calculate the LOF value of each product data;

[0010] Perform data grading on the LOF values of all products to obtain the graded product data;

[0011] Among them, when calculating the LOF value of each product data, it includes the step of determining the optimal neighbor number; the optimal neighbor number is the abscissa corresponding to the inflection point in the constructed curve; the curve is obtained by sorting the weighted Euclidean distances between the i-th product data and the corresponding set number of neighboring product data from small to large, where the abscissa of the curve is the number of the neighboring product data of the i-th product data;

[0012] The weight of each dimension data in the weighted Euclidean distance is the normalized value of the sum of the product of the standard deviation of all abnormal degrees and the maximum correlation in the corresponding dimension sequence and the initial weight; the maximum correlation is the maximum value of the correlation between any dimension sequence and the remaining other dimension sequences; the abnormal degree characterizes the abnormal conditions of each data in the corresponding dimension sequence.

[0013] In the above solution, through the preliminary abnormal analysis of different dimension data in each product data, the abnormal conditions of each dimension data are determined. Then, according to the abnormal conditions and the correlation between different dimension data, the importance of the corresponding dimension data is determined, so as to determine the optimal neighbor number for calculating the LOF value of each product data. That is, the solution of the present invention can dynamically adjust the optimal neighbor number to meet the needs of different products, ultimately improving the accuracy of calculating the LOF value, and then realizing the accurate classification of the abnormal level of product data.

[0014] Optionally, the multiple dimension data are production work order data, quality inspection data, and logistics tracking data; the production work order data is the operation status data of production equipment, the quality inspection data is the product qualification rate, and the logistics tracking data is the transportation time during transportation.

[0015] Optionally, the weighted Euclidean distance is the first distance is:

[0016] ;

[0017] Among them, is the weight of the n-th dimension data, N is the number of dimensions, is the n-th dimension data in the standardized product data corresponding to the i-th product, is the n-th dimension data in the standardized product data corresponding to the j-th product.

[0018] In the above solution, by determining the weights of the data in each dimension, the weighted Euclidean distance between the product data and the neighboring products can be accurately obtained.

[0019] Optionally, the weighted Euclidean distance is the second distance which is:

[0020] ;

[0021] where is the degree of abnormality of the n-th dimensional data in the standardized product data corresponding to the i-th product, is the weight of the n-th dimensional data, N is the number of dimensions, is the n-th dimensional data in the standardized product data corresponding to the i-th product, is the n-th dimensional data in the standardized product data corresponding to the j-th product.

[0022] In the above solution, by introducing the weights and the degree of abnormality of the data in each dimension, the weighted Euclidean distance between the product data and the neighboring products can be accurately obtained.

[0023] Optionally, the degree of abnormality is:

[0024] ;

[0025] In the formula, is the n-th dimensional data in the standardized product data corresponding to the i-th product, is the mean of the n-th dimensional data of all products, is the coefficient of variation of the n-th dimensional data of all products; the coefficient of variation is the ratio of the standard deviation of any dimensional sequence in all products to the mean of the data in the corresponding dimensional sequence.

[0026] In the above solution, by analyzing the fluctuation of the data in different dimensions of each product relative to all products, the abnormal conditions of the corresponding dimensional data are determined.

[0027] Optionally, the correlation is the Pearson correlation coefficient between any dimensional sequence and the remaining dimensional sequences.

[0028] Optionally, the standardized product data is obtained by performing standardization processing on the product data of each product obtained by using the normalization method; the normalization method is the maximum-minimum normalization method.

[0029] Through the above normalization method, not only can the dimension be eliminated, but also the calculation can be simplified.

[0030] Optionally, when the abscissa corresponding to the inflection point in the constructed curve is not an integer, perform rounding on the abscissa corresponding to the inflection point.

[0031] Optionally, it further includes: when there is no inflection point in the curve, set the optimal proximity number to the set number.

[0032] Optionally, the data grading of the LOF values of all products to obtain the graded product data includes: using the ostu algorithm to grade the LOF values of all products to obtain product data with high-risk levels and product data with low-risk levels.

[0033] The beneficial effects of the present invention are:

[0034] The solution of the present invention quantifies the degree of abnormality of each dimension data, combines the distribution of the degree of abnormality of all product data in a single dimension and the correlation between the dimension sequence and other dimension sequences to obtain weights, and uses the weights to quantify the importance of different dimension data when calculating the Euclidean distance, amplifying the abnormality of a single dimension and avoiding the problem that the traditional LOF algorithm ignores the abnormality of a single dimension; determines the optimal proximity number k using the change curve of the weighted Euclidean distance between any product data and its neighboring product data, filtering the density estimation fluctuations caused by improper selection of the k value, and making the calculation of the LOF values of each product data obtained more reliable. Description of the Drawings

[0035] Figure 1 Schematically shows a step flowchart of a data grading and classification method based on the entire supply chain process in this embodiment. Detailed Embodiments

[0036] Specifically, as Figure 1 shown, a data grading and classification method based on the entire supply chain process in this embodiment includes the following steps:

[0037] Step S1, obtain product data of multiple products in the same batch in the supply chain process. Among them, the data of different products in the same dimension constitute a dimension sequence, which is obtained in the form of time series data and retains its corresponding data label, such as the serial number of the product.

[0038] In this embodiment, multiple dimension data of product data of each link in the supply chain process are obtained in real time, including production work order data, quality inspection data, logistics tracking data, etc.

[0039] Specifically, production work order data: operating status data of production equipment (equipment startup time, shutdown time, number of faults, etc.), production progress data (number of completed products, number of uncompleted products, deviation between production progress and plan, etc.). These data are used for daily production scheduling and equipment maintenance.

[0040] Quality inspection data: raw material inspection results, quality inspection data during the production process (qualified rate, defective rate, defect types, etc.), finished product quality inspection results, etc. These data are used for quality control and improvement.

[0041] Logistics tracking data: transportation mode, transportation time, transportation cost, carrier information, goods tracking information (transportation trajectory, estimated arrival time, etc.). These data are used for optimizing logistics distribution and cost control.

[0042] In this embodiment, taking the production work order data as the operating status data of the production equipment, the quality inspection data as the product qualified rate, and the logistics tracking data as the transportation time as examples, the product data of multiple products are classified. Exemplarily, taking a certain electronic product as an example, the operating status data, qualified rate, and transportation cost of the production equipment of the electronic product can be obtained.

[0043] In this embodiment, the product data of multiple collected products are also subjected to data cleaning and interpolation processing. Among them, the interpolation processing can fill the missing values in the product data by using the LSTM network model or the interpolation processing method.

[0044] For the product data of the i-th product, in order to avoid the problem of dimensionality and affect the subsequent result of quantifying the data risk level, the above different-dimensional data are respectively normalized. In this embodiment, the normalization method is used to standardize the product data of each obtained product; the normalization method is the maximum-minimum normalization method.

[0045] Step S2, calculate the LOF value of each product data.

[0046] It should be noted that in the process of calculating the LOF value of each product data, when the k value is too small, the neighborhood range is limited to very close neighbors, and the local density estimation is greatly disturbed by noise, resulting in relatively violent fluctuations in the local reachable density. When the k value is too large, the model is too sensitive to local data, easily leading to overfitting and falling into the problem of local minimum. Therefore, in this embodiment, an improved method for selecting an appropriate k value is given to obtain the optimal number of neighbors.

[0047] The process of obtaining the optimal number of neighbors includes steps S21 - S25, specifically:

[0048] Step S21, obtain the degree of abnormality of each dimension data in the standardized product data corresponding to each product.

[0049] In this embodiment, considering that when calculating the LOF value, the local reachability density is the reciprocal of the reachability distance, and the reachability distance is the Euclidean distance between each data point and its neighboring data points. When there are multi-dimensional data in product data and the anomalies of different dimensional data are different (if only the data in a certain dimension of the product data numbered i is abnormal), if the Euclidean distance is directly calculated without considering the anomalies of dimensional data, then when calculating the local reachability density and the subsequent LOF value, it may not be possible to detect whether the product data is a truly abnormal data point. Therefore, in this embodiment, the anomalies of each dimensional data are introduced to provide accurate data support for the subsequent anomaly detection of product data.

[0050] In one embodiment, the degree of anomaly is:

[0051] ;

[0052] In the formula, is the nth-dimensional data in the standardized product data corresponding to the i-th product, is the mean of the nth-dimensional data of all products, is the coefficient of variation of the nth-dimensional data of all products.

[0053] Among them, the coefficient of variation is the ratio of the standard deviation of any dimensional sequence in all products to the mean of the data in the corresponding dimensional sequence.

[0054] The above The closer the value is to 1, the smaller the difference between the nth-dimensional data of the i-th product and the nth-dimensional data of its same-batch products, and the smaller the possibility that the nth-dimensional data of the i-th product is abnormal. The coefficient of variation reflects the degree of dispersion of the data of all products in the nth dimension, and it can measure the confidence level of the anomaly degree of the nth-dimensional data of the i-th product. If the degree of dispersion of the nth-dimensional data of all products is relatively high, then deviates from the mean , it cannot be determined that the nth-dimensional data of the i-th product is abnormal, that is, the confidence level is relatively low; if the nth-dimensional data of all products is relatively concentrated, when deviates from , it can be judged that the nth-dimensional data of the i-th product is abnormal, that is, the confidence level is relatively high. The value range of is a positive number less than 1. The larger the value, the greater the possibility that the nth-dimensional data of the i-th product is abnormal compared with its same-batch products.

[0055] In another embodiment, the degree of anomaly can also be obtained by using algorithms such as random forest algorithm and clustering method. Since the random forest algorithm and clustering method are both existing technologies, they will not be elaborated here.

[0056] Step S22: Calculate the correlation between any dimension sequence and all the remaining dimension sequences respectively to obtain the maximum correlation of any dimension sequence.

[0057] Among them, the Pearson correlation coefficient is selected as the correlation calculation method.

[0058] In this embodiment, taking the maximum value of the correlation between the nth dimension sequence and other dimension sequences can measure the importance of each dimension when calculating the Euclidean distance according to the maximization of the correlation between different dimension sequences.

[0059] Step S23: Determine the weight of the corresponding dimension data according to the maximum correlation and the degree of abnormality.

[0060] Since the traditional Euclidean distance calculation method may make it difficult to detect the abnormality of single-dimensional data, therefore, when calculating the Euclidean distance, the weights of each dimension data can be used to weight the squared difference of the corresponding dimension data to amplify the importance of the abnormal dimension.

[0061] Among them, in addition to considering the degree of abnormality of the data in each dimension when obtaining the weight, the correlation between the data of each dimension also needs to be considered. Highly correlated dimensions may represent the same phenomenon. If only a single abnormal dimension is weighted, the collaborative abnormality of the associated dimensions may be ignored.

[0062] Specifically, the weight is:

[0063] ; where is the standard deviation of the degree of abnormality of the nth-dimensional data of all products, is the maximum value of the correlation between the nth dimension sequence of all products and the remaining dimension sequences, and N is the number of dimensions.

[0064] Among them, the initial weight of each dimension data is 1. In this embodiment, the purpose of setting the initial weight is to ensure that each dimension data in the product data participates in the calculation to obtain the quasi-importance of different dimension data.

[0065] The above standard deviation can reflect the concentration degree of the degree of abnormality of the nth-dimensional data of all products. If the standard deviation is small, it indicates that the degree of abnormality of the nth-dimensional data of all products is relatively concentrated; on the contrary, if the standard deviation is large, it indicates that the degree of abnormality of the nth-dimensional data of all products is more discrete, and the probability of abnormal data appearing is large. At this time, the weight of the nth-dimensional data should be increased.

[0066] Step S24: Calculate the weighted Euclidean distance between each product and its neighboring products based on the weights of each dimension data.

[0067] In one embodiment, the weighted Euclidean distance is the first distance is:

[0068] ; where is the weight of the n-th dimensional data, and N is the number of dimensions, is the n-th dimensional data in the standardized product data corresponding to the i-th product, is the n-th dimensional data in the standardized product data corresponding to the j-th product.

[0069] In another embodiment, the weighted Euclidean distance is the second distance is:

[0070] ;

[0071] where is the degree of abnormality of the n-th dimensional data in the standardized product data corresponding to the i-th product, is the weight of the n-th dimensional data, and N is the number of dimensions, is the n-th dimensional data in the standardized product data corresponding to the i-th product, is the n-th dimensional data in the standardized product data corresponding to the j-th product.

[0072] In this embodiment, the weighted Euclidean distance can increase the distance of the dimension where the abnormal data is located to a greater extent, thereby improving the detection accuracy of the abnormality of the single-dimensional data.

[0073] Step S25: Draw a curve of the weighted Euclidean distance and the number of different neighboring products, and determine the optimal number of neighbors.

[0074] In this embodiment, after obtaining the weighted Euclidean distances from each product to its nearest neighboring products, the weighted Euclidean distances are sorted from small to large, and the neighboring products corresponding to the sorted weighted Euclidean distances are renumbered. A curve of the number of neighboring products corresponding to each product and the weighted Euclidean distance is drawn according to the re-sorted weighted Euclidean distances and numbers. Where the abscissa of this curve is the number of the neighboring product of the i-th product, and the ordinate is the weighted Euclidean distance between the i-th product and the product data of the corresponding set number of neighboring products.

[0075] Exemplarily, set the set number to the maximum number of neighbors k max ; and obtain the Euclidean distance between the i-th product and its k max -th neighboring product, where the Euclidean distance between the x-th neighboring product and the i-th product is D ix , where the number of the neighboring product closest to the i-th product is 1, and the number of the neighboring product farthest from the i-th product is k max , that is, a curve sorted from small to large according to the weighted Euclidean distance can be obtained.

[0076] It should be noted that the curve drawn in this embodiment is a monotonically increasing curve.

[0077] In this embodiment, after determining the curve, calculate the second derivative of each point on the curve, and select the abscissa (the number of the adjacent product) corresponding to the point where the second derivative is 0 (i.e., the inflection point on the curve) as the optimal adjacent number.

[0078] When the abscissa corresponding to the inflection point in the constructed curve is not an integer, perform a rounding operation on the abscissa corresponding to the inflection point.

[0079] Furthermore, it also includes: when there is no inflection point in the curve, take the set number as the optimal adjacent number. The value of the maximum adjacent number in this embodiment can be 50; of course, this maximum adjacent number can also be selected according to the actual situation.

[0080] In this embodiment, when the k value increases to cover the locally dense uniform area, the local reachability density of the corresponding product data changes slowly and tends to be stable. Therefore, by taking the k value corresponding to the inflection point in the process of the change of the local reachability density caused by the change of the k value as the optimal adjacent number when calculating the LOF value of this point, the density fluctuations caused by improper selection of the k value can be filtered out, making the LOF value more reliable.

[0081] After obtaining the optimal adjacent number of each product, in this embodiment, the local outlier factor algorithm is used to calculate the LOF value of each product data.

[0082] Step S3, perform data grading on the LOF values of all products to obtain the graded product data.

[0083] In this embodiment, the ostu algorithm is used to perform data grading on the LOF values of all products.

[0084] Specifically, use the ostu algorithm to calculate the threshold of the LOF values of these product data, divide them into product data of high-risk level and low-risk level, directly classify and store the product data of the low-risk level with lower LOF values, and for the product data of the high-risk level with higher LOF values, it is necessary to remind the management personnel to conduct a review to complete the grading process of the supply chain product data.

[0085] Among them, the ostu algorithm is a prior art and will not be elaborated here.

[0086] The solution of the present invention can perform risk level grading on the product data in the supply chain, overcoming the problem that the existing data grading method is limited by static grading rules, resulting in incorrect real-time data grading.

[0087] In the description of this specification, the meaning of "a plurality of" is at least two, such as two, three or more, etc., unless otherwise specifically defined.

[0088] Although this specification has shown and described multiple embodiments of the present invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art will think of many changes, alterations and alternative ways without departing from the spirit and idea of the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed in the practice of the present invention.

Claims

1. A data grading and classification method based on the entire supply chain process, characterized in that, Including: Obtaining the standardized product data of multiple products in the same batch in the supply chain process; the product data includes multiple dimension data, and the data of all products in the same dimension constitutes a dimension sequence; The data of multiple dimensions are production work order data, quality inspection data, and logistics tracking data; the production work order data is the operation status data of production equipment, the quality inspection data is the product qualification rate, and the logistics tracking data is the transportation time during transportation; Calculating the LOF value of each product data; Performing data grading on the LOF values of all products to obtain the graded product data; Among them, when calculating the LOF value of each product data, it includes the step of determining the optimal number of neighbors; the optimal number of neighbors is the abscissa corresponding to the inflection point in the constructed curve; the curve is obtained by sorting the weighted Euclidean distances between the i-th product data and the corresponding set number of neighboring product data from small to large, where the abscissa of the curve is the number of the neighboring product data of the i-th product data; The weight of each dimension data in the weighted Euclidean distance is the normalized value of the sum of the product of the standard deviation of all abnormal degrees and the maximum correlation in the corresponding dimension sequence and the initial weight; the maximum correlation is the maximum value of the correlation between any dimension sequence and the other remaining dimension sequences; the abnormal degree characterizes the abnormal conditions of each data in the corresponding dimension sequence; The weighted Euclidean distance is the first distance which is ; Among them, is the weight of the n-th dimensional data, and N is the number of dimensions. is the n-th dimensional data in the standardized product data corresponding to the i-th product. is the n-th dimensional data in the standardized product data corresponding to the j-th product. Degree of abnormality is as follows: ; wherein, is the mean value of the n-th dimensional data of all products, is the coefficient of variation of the n-th dimensional data of all products; the coefficient of variation is the ratio of the standard deviation of any dimensional sequence among all products to the mean value of the data in the corresponding dimensional sequence.

2. A data classification method based on the entire supply chain process according to claim 1, characterized in that The weighted Euclidean distance is the second distance It is as follows: ; Among them, is the degree of abnormality of the n-th dimensional data in the standardized product data corresponding to the i-th product, is the weight of the n-th dimensional data, N is the number of dimensions, is the n-th dimensional data in the standardized product data corresponding to the i-th product, is the n-th dimensional data in the standardized product data corresponding to the j-th product.

3. A data classification method based on the entire supply chain process according to claim 1, characterized in that The correlation is the Pearson correlation coefficient between any dimension sequence and the other remaining dimension sequences.

4. A data classification method based on the entire supply chain process according to claim 1, characterized in that, The standardized product data is obtained by performing standardized processing on the product data of each product obtained by using the normalization method; the normalization method is the maximum-minimum normalization method.

5. A data classification method based on the entire supply chain process according to claim 1, characterized in that When the abscissa corresponding to the inflection point in the constructed curve is not an integer, perform a rounding operation on the abscissa corresponding to the inflection point.

6. A data classification method based on the entire supply chain process according to claim 1, characterized in that, It also includes: When there is no inflection point in the curve, take the optimal number of neighbors as the set number.

7. A data classification method based on the entire supply chain process according to claim 1, characterized in that, The performing data grading on the LOF values of all products to obtain the graded product data includes: using the ostu algorithm to perform data grading on the LOF values of all products to obtain the product data of high-risk levels and the product data of low-risk levels.

Citation Information

Patent Citations

  • Time series data anomaly monitoring system and method based on LOF and isolated forest

    CN115577275A

  • Dust monitoring method of micron-sized powder dust-free charging system

    CN116702082A