Supply chain emergency management key data anomaly sensing method, device and equipment
By applying multi-scale feature subset extraction and multi-dimensional random hyperplanar interface division technology in supply chain data, combined with a hierarchical integrated learning decision-making mechanism, the problem of insufficient accuracy in the processing of high-dimensional multi-source data is solved, and more efficient and robust anomaly detection is achieved.
Patent Information
- Application Number
- CN202510675825.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-20
AI Technical Summary
Existing anomaly detection technologies often face problems such as insufficient accuracy when processing high-dimensional and multi-source data, and it is difficult to adapt to the real-time monitoring needs of complex dynamic supply chains.
Multi-scale feature subset extraction technology, multi-dimensional random hyperplanar interface division strategy, and hierarchical integrated learning decision-making mechanism are adopted to improve the accuracy and robustness of anomaly detection.
By extracting multi-scale feature subsets and random hyperplanar interface divisions, complex anomalies in the data can be captured more accurately, the ability to detect multi-dimensional anomalies and the adaptability and robustness of the model in complex data environments can be enhanced.
Smart Images

Figure CN120180290A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data anomaly detection, and in particular, to a method, device, and equipment for anomaly perception of key data in supply chain emergency management. Background Art
[0002] The rise of big data technology has brought new opportunities for the improvement of supply chain emergency management. Through multi-dimensional anomaly detection technology, supply chain managers can accurately identify potential risks in complex data environments, providing timely data support for emergency response and decision-making, thereby reducing economic losses caused by supply chain disruptions. Therefore, efficient and accurate anomaly detection has become an important research issue for enhancing supply chain resilience and ensuring economic security.
[0003] Currently, significant progress has been made in anomaly detection research in supply chain emergency management both at home and abroad. At the international level, the research focus is mainly on building a responsive and sustainable supply chain system. In terms of algorithm theory, some have introduced an improved Transformer into supply chain time series anomaly detection, and the proposed adaptive attention mechanism has increased the detection accuracy to 94.2%; some have deeply studied the key scientific issues in supply chain resilience and security, providing important theoretical guidance for the development of anomaly detection technology; in domestic research, some have proposed an improved isolation forest algorithm based on node evaluation and maximum inter-class variance, significantly improving the accuracy of anomaly detection, and some have developed an improved isolation forest algorithm for anomaly behavior detection, showing excellent performance in dealing with high-dimensional data.
[0004] However, existing anomaly detection technologies often face problems such as insufficient accuracy when dealing with high-dimensional and multi-source data, and it is difficult to meet the real-time monitoring requirements of complex dynamic supply chains. Currently, there is no technical solution that can solve the above technical problems, nor is there a method, device, and equipment for anomaly perception of key data in supply chain emergency management. Summary of the Invention
[0005] The present invention provides a method, device, and equipment for anomaly perception of key data in supply chain emergency management, which integrates multi-scale feature subset extraction technology, multi-dimensional random hyperplane boundary division strategy, and hierarchical integrated learning decision-making mechanism to improve the accuracy and robustness of anomaly detection.
[0006] In a first aspect, the present invention provides a method for anomaly perception of key data in supply chain emergency management, including: Processing the corresponding data set, feature dimension set, preset subset size, and preset step length of supply chain emergency management by using a multi-scale feature subset extraction mechanism to obtain a target set containing multiple trained forest data sets; For each trained forest dataset in the target set, randomly select two different data points from the trained forest dataset, calculate the vertical vector between the two data points, determine a random intercept based on the two different data points and the vertical vector, and use the random intercept to divide the trained forest dataset to obtain a first subset and a second subset. For any subset, select two different data points from the subset to continue the division until the preset sample number is reached, obtain all the divided subsets, and determine the isolation tree corresponding to the trained forest dataset according to all the divided subsets; Use the isolation tree to calculate the anomaly score of each data point in the trained forest dataset, obtain the anomaly score corresponding to the trained forest dataset, traverse all the trained forest datasets, obtain the anomaly score corresponding to each trained forest dataset, fuse the anomaly scores corresponding to all the trained forest datasets, obtain a preliminary fusion score, and process the preliminary fusion score to obtain a final anomaly detection result.
[0007] According to the key data anomaly perception method for supply chain emergency management provided by the present invention, the use of the multi-scale feature subset extraction mechanism to process the dataset, feature dimension set, preset subset size, and preset step corresponding to supply chain emergency management to obtain a target set including multiple trained forest datasets includes: When the total number of feature dimensions in the feature dimension set is greater than the preset subset size, initialize the feature subset set, set the initial index, and repeatedly execute the following steps: When the initial index is less than the total number of feature dimensions, select a feature subset of the preset subset size from the feature dimension set of the dataset corresponding to supply chain emergency management to construct a feature dataset, and use a preset isolation forest algorithm to train the feature dataset to obtain a trained forest dataset; Add the trained forest dataset to the initialized feature subset set to obtain an updated set corresponding to the updated index, where the updated index is determined by adding the preset step to the initial index; Until the updated index is greater than or equal to the total number of feature dimensions, obtain a target set including multiple trained forest datasets.
[0008] According to the key data anomaly perception method for supply chain emergency management provided by the present invention, the use of the random intercept to divide the trained forest dataset to obtain a first subset and a second subset includes: =filter(X, X·w + b ≤ 0); = filter(X, X·w + b > 0); Where, is the first subset, is the second subset, X is the trained forest dataset, w is the vertical vector between two data points, and b is the random intercept.
[0009] According to the key data anomaly perception method for supply chain emergency management provided by the present invention, calculating the anomaly score of each data point in the trained forest dataset by using the isolation tree to obtain the anomaly score corresponding to the trained forest dataset includes: For any data point, obtain the current path length of the data point in the isolation tree and obtain the expected path length of the data point in the isolation tree; Determine the anomaly score of the data point according to the current path length and the expected path length, traverse all data points, and determine each anomaly score corresponding to all data points; Determine the anomaly score corresponding to the trained forest dataset according to each anomaly score corresponding to all data points.
[0010] According to the key data anomaly perception method for supply chain emergency management provided by the present invention, fusing the anomaly scores corresponding to all trained forest datasets to obtain a preliminary fusion score includes:
[0011] wherein, is the preliminary fusion score, is the anomaly score corresponding to any trained forest dataset, is the weight coefficient corresponding to any trained forest dataset, is the number of all trained forest datasets.
[0012] According to the key data anomaly perception method for supply chain emergency management provided by the present invention, processing the preliminary fusion score to obtain the final anomaly detection result includes: In the case that the preliminary fusion score is greater than the global anomaly threshold, determine that there are anomaly points in the corresponding dataset of the supply chain emergency management; Determine all data points with anomaly scores above the first quantile threshold and below the second quantile threshold as slightly abnormal data points, and determine all data points with anomaly scores above the second quantile threshold as significantly abnormal data points; The first quantile threshold is less than the second quantile threshold.
[0013] According to the method for abnormal perception of key data in supply chain emergency management provided by the present invention, the corresponding data set for supply chain emergency management includes unique identifier features, product ID features, temperature features, process temperature features, rotational speed features, torque features, tool wear features, machine failure label features, tool wear failure features, heat dissipation failure features, power failure features, overstrain failure features, and random failure features.
[0014] According to the method for abnormal perception of key data in supply chain emergency management provided by the present invention, before using the multi-scale feature subset extraction mechanism to process the corresponding data set, feature dimension set, preset subset size, and preset step size of supply chain emergency management, the method further includes: Performing data cleaning and denoising on the original supply chain emergency management data to obtain a first data set; Performing feature selection and extraction on the first data set to obtain a second data set; Performing data standardization and normalization processing on the second data set to obtain the corresponding data set for supply chain emergency management.
[0015] In a second aspect, a device for abnormal perception of key data in supply chain emergency management is provided, including: A processing unit configured to use a multi-scale feature subset extraction mechanism to process the corresponding data set, feature dimension set, preset subset size, and preset step size of supply chain emergency management to obtain a target set including multiple trained forest data sets; An acquisition unit configured to, for each trained forest data set in the target set, randomly select two different data points from the trained forest data set, calculate the vertical vector between the two data points, determine a random intercept according to the two different data points and the vertical vector, use the random intercept to divide the trained forest data set to obtain a first subset and a second subset, for any subset, select two different data points from the subset to continue the division until the preset sample number is reached, obtain all divided subsets, and determine the isolation tree corresponding to the trained forest data set according to all divided subsets; A calculation unit configured to use the isolation tree to calculate the anomaly score of each data point in the trained forest data set to obtain the anomaly score corresponding to the trained forest data set, traverse all trained forest data sets to obtain the anomaly score corresponding to each trained forest data set, fuse the anomaly scores corresponding to all trained forest data sets to obtain a preliminary fusion score, and process the preliminary fusion score to obtain a final anomaly detection result.
[0016] In a third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method for abnormal perception of key data in supply chain emergency management is implemented.
[0017] By extracting feature subsets at different scales, the present invention can accurately identify various abnormal patterns in supply chain data. This mechanism enhances the sensitivity of the algorithm to the recognition of local and global anomalies, and effectively addresses the heterogeneity and dynamic changes of supply chain data. In response to the challenge of high-dimensional data processing, the present invention proposes a mechanism for simultaneously performing hyperplane segmentation in multiple dimensions, enabling the model to more precisely capture complex abnormal patterns in the data. Especially in a data environment with multiple levels and intertwined factors, it demonstrates stronger detection capabilities and higher adaptability. By integrating multiple isolation forest models and analyzing and detecting anomalies in the data hierarchically, the adaptability and robustness of the model in a complex data environment are enhanced. The hierarchical structure enables the model to capture subtle changes in the data from different levels, improving the detection ability for multi-dimensional abnormal patterns. The present invention takes improving the efficiency of supply chain emergency management as the core and building an efficient anomaly detection mechanism as the focus, constructs an anomaly detection framework suitable for the modern complex supply chain environment. By integrating multi-dimensional data analysis techniques, it realizes the accurate identification and rapid response to supply chain abnormal states, thereby enhancing the resilience and reliability of the supply chain. A feature extraction method applicable to multi-source heterogeneous data in the supply chain is established. By effectively identifying and extracting features at different scales, the accuracy and comprehensiveness of anomaly detection are improved. It focuses on solving technical problems such as high dimensionality and complex features of supply chain data, laying a foundation for subsequent anomaly detection. Based on the anomaly boundary division algorithm of multi-dimensional random hyperplanes, by optimizing the boundary division method, the robustness and generalization ability of anomaly detection are improved. It focuses on breaking through the computational efficiency bottleneck of traditional methods when dealing with high-dimensional data and realizing the rapid identification of abnormal patterns. A hierarchical integrated learning framework is constructed. Based on multi-scale feature extraction and multi-dimensional random hyperplane boundary division methods, a systematic anomaly identification criterion is established. It focuses on breaking through the problems of consistency and reliability of anomaly detection results and enhancing the robustness and adaptive ability of the detection system in dynamic complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1It is a schematic flowchart of the method for abnormal perception of key data in supply chain emergency management provided by the present invention; Figure 2 It is a scatter plot of the isolation forest algorithm provided by the present invention; Figure 3 It is a two-dimensional graph of the isolation forest algorithm provided by the present invention; Figure 4 It is an example graph of the mean and standard deviation of supply chain data features provided by the present invention; Figure 5 It is a graph for abnormal detection of supply chain data provided by the present invention; Figure 6 It is a visualization graph for abnormal detection of grid segmentation supply chain provided by the present invention; Figure 7 It is a graphic effect diagram after magnifying 50 times provided by the present invention; Figure 8 It is a graphic effect diagram after magnifying 100 times provided by the present invention; Figure 9 It is a schematic structural diagram of the device for abnormal perception of key data in supply chain emergency management provided by the present invention; Figure 10 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments
[0020] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0021] Isolation Forest is an unsupervised anomaly detection algorithm based on trees. It "isolates" anomaly points by randomly selecting features and partitioning the data space. The advantage of this method is that anomaly points can usually be identified in fewer partitioning steps. Its high efficiency, lack of need for supervised learning, and good scalability make it particularly suitable for processing large-scale datasets. However, as the data dimension increases, the effects of feature selection and partitioning weaken, making it difficult to capture complex anomaly patterns. At the same time, the algorithm may also be affected by the characteristics of the data distribution when identifying local anomalies, reducing the detection accuracy. In response to the deficiencies of the Isolation Forest algorithm in aspects such as the recognition sensitivity of multi-dimensional data parameters and local anomaly detection, related research has proposed a variety of optimization strategies. The following are several main optimization methods: The adaptive weight adjustment strategy enhances the algorithm's adaptability to data features by dynamically adjusting the weights of the trees. Especially in the case of high data heterogeneity, it improves the detection effect; The Deep Isolation Forest combines deep learning techniques to automatically extract the features of the trees, enhancing the model's ability to express features and more effectively processing high-dimensional and complex datasets; The Graph Isolation Forest integrates the graph model into the Isolation Forest framework to analyze the mutual relationships between data, which is particularly suitable for processing complex datasets with intrinsic structures, such as social network or supply chain data; The Density-Enhanced Isolation Forest introduces data density information, which can more accurately identify anomaly points in sparse regions and enhances the recognition sensitivity to anomaly patterns; The Robust Isolation Forest strengthens the ability to handle noise and missing data, improving the robustness of the model and enabling it to maintain a stable performance when facing incomplete or noisy data.
[0022] Although these optimization methods have significantly improved the performance of anomaly detection in various application scenarios, in practical applications, choosing the most suitable optimization method still requires comprehensive consideration of factors such as data scale, real-time requirements, and data characteristics. Therefore, the algorithm optimization for specific application scenarios remains the forefront direction of technical research. Based on the multi-level anomaly detection theory of the improved Isolation Forest algorithm, the present invention innovatively constructs a detection and analysis framework that includes multi-scale feature extraction and random hyperplane isolation technology, expands the supply chain risk identification theory, and proposes a hierarchical integrated learning decision-making mechanism, providing a new direction for the supply chain decision-making theory. At the practical level, the anomaly detection framework constructed by the present invention enhances the enterprise's prevention and control of supply chain risks, reduces interruption risks through early identification and warning; at the same time, the present invention can be directly used for supply chain emergency management, optimizing production plans and resource allocation, and improving emergency efficiency; again, it provides a technical path for the digital transformation of the enterprise's supply chain management, enhancing operational efficiency and market competitiveness.
[0023] The present invention will deeply explore the research methods and implementation frameworks adopted, with a focus on introducing the step-by-step optimization process of the Isolation Forest algorithm, and completing three core aspects: data preprocessing, algorithm optimization design, and experimental verification. In the data preprocessing stage, multi-dimensional data related to the supply chain is cleaned, integrated, and feature selected to ensure the accuracy of the data and its applicability in anomaly detection. In terms of algorithm design, a step-by-step fusion mechanism of multi-scale feature subset extraction, multi-dimensional random hyperplane boundary division, and hierarchical ensemble learning is adopted, which effectively overcomes the limitations of traditional Isolation Forest algorithms in dealing with high-dimensional data and complex anomaly patterns. Finally, in the experimental verification section, the experimental design and evaluation metrics are introduced in detail, and by comparing with existing algorithms, the practical application effect of the proposed method in detecting key data anomalies in supply chain emergency management is evaluated, providing a theoretical and practical basis for the performance and visualization application of the algorithm.
[0024] Figure 1 It is a schematic flowchart of the method for anomaly perception of key data in supply chain emergency management provided by the present invention. The method for anomaly perception of key data in supply chain emergency management includes: Step 101: Use the multi-scale feature subset extraction mechanism to process the dataset, feature dimension set, preset subset size, and preset step size corresponding to supply chain emergency management, and obtain a target set containing multiple trained forest datasets; Step 102: For each trained forest dataset in the target set, randomly select two different data points from the trained forest dataset, calculate the vertical vector between the two data points, determine the random intercept according to the two different data points and the vertical vector, use the random intercept to divide the trained forest dataset to obtain a first subset and a second subset. For any subset, select two different data points from the subset to continue the division until the preset sample number is reached, obtain all the divided subsets, and determine the isolation tree corresponding to the trained forest dataset according to all the divided subsets; Step 103: Use the isolation tree to calculate the anomaly score of each data point in the trained forest dataset, obtain the anomaly score corresponding to the trained forest dataset, traverse all the trained forest datasets, obtain the anomaly score corresponding to each trained forest dataset, fuse the anomaly scores corresponding to all the trained forest datasets to obtain a preliminary fusion score, and process the preliminary fusion score to obtain the final anomaly detection result.
[0025] In step 101, the multi-scale feature subset extraction mechanism uses the sliding window technique to construct feature subsets of different sizes. By decomposing the multi-dimensional features in the original data, it dynamically generates subsets containing different feature combinations to cover a wide range of feature combinations. The advantage of this is that it not only improves the model's adaptability to diverse data but also effectively avoids problems such as insufficient information or bias that may arise from using a single feature subset. Generally speaking, the multi-scale feature subset extraction mechanism comprehensively captures potential abnormal features in the data by deeply exploring feature subsets at different scales, laying a data foundation for subsequent algorithm processing.
[0026] Optionally, the input features in the algorithm steps of step 101 are the dataset X, the feature dimension set , the total number of feature dimensions u, the subset size L, and the step size step. The output is a set New_p of multiple isolation forests. First, perform the initialization operation. If u ≤ L, return the set {X} containing the original dataset; otherwise, initialize the feature subset set New_p = {} and set the index i = 1. Then generate the feature subset and train the iForest. When i < u, select a feature subset of size L from the feature dimension set Dims of the dataset X: "Subset = SelectSubset(Dims, i, L)"; construct a new dataset through the selected feature subset: "New_Data = GenerateData(X, Subset, step)", train the isolation forest iForest on the new dataset New_Data: "Forest = iForestTraining(New_Data)"; add the trained forest to the set New_p: "New_p = New_p ∪ {Forest}"; update the index: "i = i + step"; return: output the set New_p containing multiple isolation forests.
[0027] In step 102, the multi-dimensional random hyperplane boundary division mechanism randomly selects the direction of the hyperplane each time it divides the data, enabling the hyperplane to more flexibly and dynamically cross the data space. This random division method reduces the dependence on specific dimensions, allowing the model to more comprehensively capture the complex relationships between features. At the same time, the multi-dimensional random hyperplane boundary division mechanism integrates multi-dimensional feature combinations and effectively enhances the ability to identify complex abnormal patterns during the division process by dynamically adjusting the position and direction of the hyperplane.
[0028] Specifically, the input of the multi-dimensional random hyperplane boundary division mechanism is the data set X, the current tree height h, and the tree construction threshold, such as the maximum height or the minimum number of samples. The output is the node structure of the multi-dimensional hyperplane isolation tree. For the determination of leaf nodes, if the current tree height h reaches the preset maximum height threshold, or the number of samples in the data set |X| ≤ 1, then this node is considered a leaf node, and the leaf node is returned and the number of samples is recorded: "Return LeafNode{Size=|X|}". This judgment is used to terminate the recursive division and ensure the rationality of the tree depth and division.
[0029] In step 102, first randomly select two different data points from the data set, denoted as and ." , =RandomGenerateData(X)". The selected data points will be used to determine the direction vector of the hyperplane; then calculate the perpendicular vector, calculate the perpendicular vector between the two data points to construct a random hyperplane: "w=GetNormal( , )", where w is the direction vector, representing the direction of the generated hyperplane, so that the hyperplane can divide the data set in different dimensions; then generate a random intercept, and generate a random intercept b according to the selected two data points and the direction vector w to determine the specific position of the hyperplane: "b=RandomGenerateIntercept( , )". This intercept is used to offset the hyperplane, so that different division methods produce different data subsets, improving the randomness and diversity of the model; then, perform data division, and use the generated hyperplane to divide the data set X into two subsets and . The method of dividing the trained forest data set by using the random intercept to obtain the first subset and the second subset includes: =filter(X,X·w+b≤0); = filter(X,X·w+b>0); Among them, is the first subset, is the second subset, X is the trained forest data set, w is the perpendicular vector between the two data points, and b is the random intercept. This step divides the data space through the hyperplane to ensure that samples in different regions are gradually isolated, improving the accuracy of anomaly detection.
[0030] Finally, generate internal nodes and construct subtrees, recursively build the left and right subtrees, and return the structure of the current node. Each node contains the generated hyperplane information and the structure of its subtree "Return Node{Left=RMHIiTree( ,h+1,Threshold) Right=RMHIiTree( ,h+1,Threshold) Scope=w,Intercept=b}" Recursively call RMHIiTree to continue splitting the subset until the set maximum depth or sample number limit is reached.
[0031] In step 103, the hierarchical integrated learning decision mechanism is a comprehensive method for the abnormal detection results of multi-dimensional data. This mechanism hierarchically weights and integrates the detection results of the isolated forest algorithms of each subset to form a global anomaly score. Different from the direct output results of a single model, the hierarchical integrated learning decision mechanism adopts a hierarchical fusion strategy, from the preliminary subset anomaly scores to the final comprehensive score, gradually optimizing the abnormal detection results. This hierarchical structure can not only balance the contributions of different feature subsets to abnormal detection, but also enhance the stability and robustness of the model in complex data scenarios. Through hierarchical weighted fusion, the hierarchical integrated learning decision mechanism can accurately identify different types and degrees of anomalies, providing more reliable support for the abnormal perception of supply chain data.
[0032] In an optional embodiment, the present invention first takes the original data set X, the feature dimension set , the subset size L, the step size step, and the maximum depth of the RMHI tree as inputs, and outputs the anomaly scores for identifying abnormal data points. Specifically, in the generation of feature subsets in the initial stage, call the multi-scale feature subset extraction mechanism to generate a set of feature subsets { , ,… }, where each subset contains a set of feature combinations for the abnormal detection of data from multiple perspectives. Among them, the inputs are X, , L, step, and the output is the set of feature subsets S. Apply the multi-dimensional hyperplane isolation mechanism in the first-level integration, the random hyperplane isolation mechanism on each feature subset: for each feature subset ∈S, perform the following steps: construct an isolation tree on the feature subset through the multi-dimensional random hyperplane boundary division mechanism (RMHI), where each tree divides data points through a random hyperplane. Among them, the inputs are the subset , and the maximum depth of the RMHI tree ; The output is an isolation tree .
[0033] Calculating the anomaly score of each data point in the trained forest dataset by using the isolation tree to obtain the anomaly score corresponding to the trained forest dataset includes: for any data point, obtaining the current path length of the data point in the isolation tree, and obtaining the expected path length of the data point in the isolation tree; determining the anomaly score of the data point according to the current path length and the expected path length, traversing all data points, and determining each anomaly score corresponding to all data points; determining the anomaly score corresponding to the trained forest dataset according to each anomaly score corresponding to all data points.
[0034] Optionally, Isolation Forest is a clustering tree structure algorithm mainly used for unsupervised outlier detection. The core idea of this method is to gradually "isolate" outliers by randomly selecting features and partitioning the data space. The basis for identifying outliers is the isolation of data points rather than similarity. Isolation Forest is favored for its high efficiency, unsupervised learning ability, and excellent scalability, especially suitable for processing large-scale datasets. In the construction of an isolation tree (Isolation Tree, iTree), the core of Isolation Forest consists of isolation trees. Each isolation tree recursively selects a feature and randomly determines a split point within the range of the feature's values to partition the dataset. This process continues until the leaf node contains only a single data point or reaches the set maximum tree depth limit. In PathLength, given a data point x, the path length ℎ(x) refers to the number of edges passed from the root node to this point in the tree. For outliers, due to their large differences from other data points, they are usually isolated on shorter paths; while normal points usually require deeper paths to be isolated. The expected path length (Expected Path Length) of Isolation Forest is the average path length of a given dataset, which is determined based on the dataset size and the distribution of data points. The formula for the expected path length E(h(x)) is: E(h(x)) = 2·(In(n - 1) + 0.5772) where n is the size of the dataset, and 0.5772 is the Euler's constant, which is related to the average depth of the tree.
[0035] Optionally, anomaly scoring is an indicator to measure the degree of anomaly of a data point compared with other points. The anomaly score of each data point is based on the comparison of its path length and the expected path length, and the formula is as follows: S(x)=
[0036] where is the path length of the data point in an isolated tree, is the expected path length of the data point.
[0037] Optionally, a normalization constant is used to adjust the relationship between the path length and the dataset size.
[0038] Optionally, for the final anomaly score, for T trees in the isolation forest, the final anomaly score of the data point x is the average of the scores in all trees: S(x)=
[0039] where (x) is the anomaly score of the data point X in the t-th tree, and T is the number of trees in the isolation forest.
[0040] Optionally, calculate the anomaly score for each training forest dataset: Use each isolation tree for each record in to calculate the anomaly score. The anomaly score is based on the isolation depth of the data point, and a higher depth indicates a more anomalous point. Finally, obtain the initial set of anomaly scores F={ , , …, }.
[0041] Optionally, include hierarchical anomaly score fusion in the second-level integration. In the feature subset anomaly score fusion, use weighted average or voting to fuse the anomaly scores of each subset to obtain the initial fusion anomaly score .
[0042] Optionally, fusing the anomaly scores corresponding to all training forest datasets to obtain the initial fusion score, including:
[0043] where is the initial fusion score, is the anomaly score corresponding to any training forest dataset, is the weight coefficient corresponding to any training forest dataset, which can be set based on the detection accuracy or experience of each subset, is the number of all training forest datasets. If the voting method is used, count the frequency of anomalies in different subsets.
[0044] Further hierarchical fusion: The initial fusion score Transferred to the second layer, a higher-level weighted voting mechanism is used to combine the abnormal distribution of the overall data points to obtain the final anomaly detection decision 。
[0045] Optionally, processing the preliminary fusion score to obtain the final anomaly detection result includes: when the preliminary fusion score is greater than the global anomaly threshold, determining that there are anomaly points in the corresponding dataset of the supply chain emergency management; determining all data points with anomaly scores above the first quantile threshold and below the second quantile threshold as slightly abnormal data points, and determining all data points with anomaly scores above the second quantile threshold as significantly abnormal data points; the first quantile threshold is less than the second quantile threshold.
[0046] Decision criterion: If exceeds the global anomaly threshold, it is marked as an anomaly point; if using quantile thresholds, data points above the median of the anomaly scores can be identified as slightly abnormal, and points above 75% can be identified as significantly abnormal. Output the final anomaly scores and decision results, output: the set of anomaly scores after final fusion ; Anomaly point identification based on the decision-making mechanism.
[0047] Optionally, the corresponding dataset of the supply chain emergency management includes unique identifier features, product ID features, temperature features, process temperature features, rotation speed features, torque features, tool wear features, machine failure label features, tool wear failure features, heat dissipation failure features, power failure features, overstrain failure features, and random failure features.
[0048] To verify the effectiveness of the improved isolation forest algorithm in the anomaly detection of key data in supply chain emergency management, the present invention selects the traditional isolation forest algorithm (Isolation Forest) as a control model for comprehensive comparative experiments. The main purpose of the experiment is to systematically evaluate and demonstrate the superiority of the improved algorithm in multiple dimensions such as anomaly detection accuracy, robustness, response efficiency, and visualization effect. Specifically, it is carried out from the following two aspects: Visualization comparison, the present invention will use data visualization technology to intuitively present the anomaly recognition effects of different algorithms through two-dimensional space mapping, so as to more clearly understand how the improved algorithm effectively identifies anomaly points in a complex multi-dimensional data environment, highlighting the advantages and limitations of different algorithms in processing high-dimensional and multi-source supply chain data; Performance index evaluation, the present invention will use 10 standardized quantitative evaluation indicators to analyze the performance of the two algorithms in terms of detection ability, stability, and response speed. The main evaluation indicators include detection accuracy, recall rate, F1 value, algorithm running time, and the sensitivity of the algorithm to noise data, etc.
[0049] The present invention evaluates the improved algorithm through the following steps: Dataset selection: In the present invention, an industrial equipment predictive maintenance dataset is used as the experimental object. This dataset contains 10,000 sample records and 14 feature dimensions, featuring high dimensionality, dynamic changes, etc., and can effectively simulate complex scenarios in the actual production environment; Data preprocessing: Before the experiment, the original data is subjected to standardization and normalization. Standardization adjusts the feature data to a standard normal distribution with a mean of 0 and a variance of 1; normalization compresses the data to the interval [0, 1] through linear mapping. By eliminating the dimensionality differences, it ensures that the features of each dimension have the same weight in model training, thereby improving the reliability of the algorithm performance; Model training and testing: Two algorithms are respectively used for model training based on the same training set and test set. The hyperparameter settings are kept consistent during the training process to eliminate the model differences caused by parameter selection, so as to ensure the fairness and scientific nature of the experiment; Visualization display: To visually display the performance differences between the two algorithms in anomaly detection, PCA (Principal Component Analysis) is used in the present invention for data dimensionality reduction, mapping the high-dimensional data to a two-dimensional space. By comparing the visualization images, the differences in anomaly recognition of the two algorithms at different data points are observed, and the ability of the improved algorithm to improve the accuracy of anomaly detection in a complex multi-dimensional data environment is understood more clearly; Performance evaluation: The present invention uses multi-dimensional evaluation indicators to evaluate the algorithm performance, mainly including key indicators such as Precision, Recall, Accuracy, AUC-ROC curve, and F1-score.
[0050] To comprehensively evaluate the performance of anomaly detection of the two algorithms, the following evaluation indicators are selected in the present invention. These indicators can help analyze the performance of the model in different dimensions, especially its performance in multi-dimensional data and complex environments, as shown in Table 1 below: Table 1 Evaluation indicators
[0051] Taking the relatively complex production and manufacturing link in the supply chain as an example, this invention selects the AI4I 2020 predictive maintenance dataset as the experimental dataset, aiming to verify the effectiveness of the proposed method for abnormal perception of key data in multi-dimensional supply chain emergency management based on the improved isolation forest algorithm. This dataset is from the UCI Machine Learning Repository and is provided under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. The AI4I 2020 predictive maintenance dataset has multi-dimensional data characteristics and contains various failure modes, which is suitable for evaluating the application of anomaly detection methods in complex supply chain environments.
[0052] Table 2 lists the main variables in the AI4I 2020 dataset and their descriptions: Table 2 AI4I 2020 Dataset Characteristics
[0053] Optionally, before using the multi-scale feature subset extraction mechanism to process the dataset, feature dimension set, preset subset size, and preset step size corresponding to supply chain emergency management, the method further includes: performing data cleaning and denoising on the original supply chain emergency management data to obtain a first dataset; performing feature selection and extraction on the first dataset to obtain a second dataset; and performing data standardization and normalization processing on the second dataset to obtain the dataset corresponding to supply chain emergency management.
[0054] Before conducting the experiment, for the AI4I 2020 predictive maintenance dataset, this invention performed the following data preprocessing operations. By cleaning the data, performing feature selection, standardization, and label processing and other steps, it ensured that the data could be effectively input into the anomaly detection model for training and evaluation.
[0055] Table 3 Data Preprocessing Steps
[0056] Optionally, to verify the effectiveness of the proposed method for abnormal perception of key data in multi-dimensional supply chain emergency management based on the improved isolation forest algorithm, this invention designed a comparative experiment scheme. By comparing with the traditional isolation forest algorithm anomaly detection method, it evaluated the performance of the improved model in the same task and verified its advantages in complex data environments.
[0057] To present the experimental results more intuitively, Python code was written, which includes data loading, preprocessing (processing categorical data and extracting numerical features), training an Isolation Forest model to detect outliers, and finally plotting a scatter plot. After running the Python code, a two-dimensional graph as shown in Figure 2 is generated, showing the Isolation Forest outlier detection and severity levels. As can be observed from Figure 2 , when processing 10,000 data points and their multi-dimensional features, the visualization effect is not ideal. To improve the visualization, the dimensionality reduction of the standardized data was performed, compressing it from a high-dimensional space to a two-dimensional space, so that it can be more intuitively visualized on a two-dimensional plane. In addition, PCA dimensionality reduction code has been added to the original code to enhance the data visualization effect. Running the Python code again generates a two-dimensional graph as shown in Figure 3 , showing the Isolation Forest outlier detection and severity classification (after dimensionality reduction). Through the two-dimensional space after PCA dimensionality reduction, the visualization effect of the outlier level and distribution of each data point has been significantly improved.
[0058] To implement the Multi-Scale Feature Subset Extraction Algorithm (MSFSE), corresponding Python code was designed and written, mainly used to read and process the AI4I 2020 dataset, select features using the MSFSE algorithm to generate feature subsets; and calculate the mean and standard deviation of each feature subset and plot them in two subgraphs respectively, facilitating the comparison of the feature statistical characteristics of different subsets, and through clear visualization, making data analysis and result interpretation more intuitive. After running the relevant Python code, instance graphs of the mean and standard deviation of the data features were generated, as shown in Figure 4 . Figure 4 Figure (a) in Figure 4 shows the mean change of each subset feature, and Figure 4 Figure (b) in
[0059] shows the standard deviation change of each subset feature. The mean and standard deviation of each subset are represented by different lines respectively, facilitating the comparison of the feature distribution differences between different subsets.
[0059] Optionally, an improved Isolation Forest RMHI algorithm proposed by the present invention intuitively reflects the abnormal distribution of data through multi-level score division and spatial display. The AI4I 2020 dataset is also used, the data is abnormally scored, and classified according to the abnormal score, dividing the data into three categories: normal, slightly abnormal, and severely abnormal. Subsequently, principal component analysis (PCA) is used to reduce the high-dimensional data to a two-dimensional space to simplify the data display and highlight the abnormal relationships between the data. Finally, by overlaying a refined quadrilateral grid area on the two-dimensional space and color-coding the abnormal states of each area, a more refined abnormal pattern distribution is presented, and Python code for implementing this algorithm is designed. After running the relevant code, the generated two-dimensional scatter plot is as shown inFigure 3 As shown, where red represents severe anomaly points, yellow represents mild anomaly points, and green represents normal points.
[0060] Figure 5 This is the supply chain data anomaly detection graph provided by the present invention, showing the anomaly detection and grid area demarcation. From Figure 5 it can be seen that this improved RMHI algorithm adds a hierarchical anomaly scoring mechanism, visualization processing of multi-dimensional data, and space segmentation on the basis of the traditional algorithm. It classifies and displays anomalies in smaller grid areas, making the distribution of anomaly points clearer and facilitating subsequent analysis and decision-making. The hierarchical ensemble learning decision-making mechanism (HELDM) is based on the MSFSE and RMHI algorithms. By weighted fusion of the anomaly scores of multiple feature subsets, it improves the accuracy and robustness of anomaly detection, and processes the heterogeneity of feature subsets through a multi-level decision-making structure to reduce bias. At the same time, it combines the global anomaly distribution to avoid local misjudgment. The Python code for implementing the HELDM algorithm is designed. After running the relevant Python code of the HELDM algorithm, the visualization effect of 10,000 data points in the multi-dimensional space is as Figure 6 shown, showing the anomaly detection of industrial equipment maintenance data based on grid division. From Figure 6 it can be seen that when dealing with complex and high-dimensional data, the HELDM mechanism algorithm can still effectively perform three-dimensional anomaly detection. Using the PCA dimensionality reduction technology, the algorithm effectively transforms the high-dimensional data into a two-dimensional space representation, so that different types of anomalies (normal, moderately abnormal, severely abnormal) are clearly distinguished in the visualization, and the distribution of anomaly points is also clearly presented in the filling of the quadrilateral grid area. The color coding of green, yellow, and red not only makes the anomaly level clear at a glance, but also enhances the interpretability of the graph.
[0061] To further verify the accuracy and robustness of the algorithm, the effect of anomaly detection can be further observed through Figure 6 the locally enlarged graph area. In the locally enlarged area as Figure 7 , Figure 8 shown, the distribution of different anomaly levels can be seen more clearly, verifying that the HELDM mechanism can accurately capture the anomaly patterns in the data at the fine-grained level and still perform stably in data-intensive or complex scenarios, ensuring the efficiency and operability of anomaly detection.
[0062] Based on the graphs generated from the experiments, the present invention conducts a comparative analysis from the following aspects: the accuracy of anomaly detection. In the traditional Isolation Forest algorithm, the distribution of anomaly points is relatively difficult to visually identify, and the distinction between minor and significant anomalies is not obvious. After introducing the multi-scale feature subset extraction and multi-dimensional random hyperplane boundary division mechanism, the improved algorithm can more accurately label anomalies of different severity levels; the clarity of the anomaly point distribution. In the visualization graph of the traditional Isolation Forest algorithm, although dimensionality reduction processing is adopted, the anomaly points are also scattered and irregularly distributed, lacking obvious structural and trend characteristics. The improved model optimizes the feature extraction and isolation mechanism, making the distribution of anomaly points in the visualization graph more concentrated and regular, and intuitively differentiating normal points and anomaly points through color or shape markings, improving the recognition and clarity of the visualization results; the real-time monitoring and early warning capabilities. In the monitoring graph of the traditional Isolation Forest algorithm, due to the unclear boundary between normal points and anomaly points, anomaly points are easily overlooked or misjudged, affecting the triggering of timely early warnings. The improved model's hierarchical anomaly scoring mechanism, visualization processing of multi-dimensional data, and spatial segmentation classify and display anomalies in smaller grid areas, making the distribution of anomaly points clearer.
[0063] In the performance test experiment, the experimental scheme of the present invention is carried out through the following steps: Feature engineering and data optimization. The present invention adopts a multi-level data preprocessing strategy, including data standardization processing, missing value filling, and outlier identification, and combines the information gain criterion for feature screening. By constructing a high-quality feature space, it lays a foundation for subsequent model training and improves the generalization ability of the algorithm; Parameter optimization design. The core parameters of the two algorithms are systematically tuned, including key variables such as the proportion of abnormal samples, the scale of decision trees, and the sampling capacity. The optimal parameter combination is determined through the grid search method, and a parameter sensitivity analysis framework is established to deeply analyze the influence mechanism of each parameter on the model performance; Model structure optimization. To enhance the accuracy and robustness of anomaly detection, the key structural parameters of the algorithm, including the dimension of the feature subspace, the iteration step size, and the depth of the isolation tree, are optimized to improve the model's representation ability for complex patterns and enhance the detection efficiency of the algorithm in high-dimensional data scenarios; Robustness verification. Different types of data interference factors, including dimensional redundancy and noise pollution, are introduced in the experimental design to construct a multi-level verification system to verify the adaptability and stability of the improved algorithm in complex environments; Performance evaluation system. A systematic performance evaluation framework is constructed using multi-dimensional evaluation indicators, focusing on core indicators such as the classification accuracy, operation efficiency, and stability of the model. Based on the test program implemented in Python, through multiple rounds of cross-validation, the repeatability and reliability of the experimental results are ensured, providing a scientific basis for model performance evaluation, as shown in Table 4 below: Table 4 Experimental Results
[0064] Analysis of Performance Test Results: By running the performance comparison test code, a quantitative analysis was conducted on the performance of the two algorithms in various performance metrics. The main results are as follows: In the traditional Isolation Forest algorithm, Accuracy: decreased from 0.9922 to 0.9865, with a decrease of 0.58%, indicating sensitivity to noise; Precision: decreased from 0.9980 to 0.9960, with a relatively small decrease (0.20%), suggesting that noise has a limited impact on the ability to identify outliers; Recall: decreased from 0.9792 to 0.9700, with a decrease of 0.94%, showing that the model is sensitive to false negatives; AUC-ROC: decreased from 0.9895 to 0.9870, with a decrease of 0.25%, reflecting a slight impact of noise on the model's discrimination ability; Robustness: the accuracy decreased by 0.58%, indicating that the traditional algorithm performs poorly in the presence of noisy data; Runtime: increased from 0.35 seconds to 0.37 seconds, showing that noise has a relatively small impact on the running time.
[0065] In the improved Isolation Forest algorithm, Accuracy: decreased from 0.9950 to 0.9920, with a decrease of 0.30%, showing more stability compared to the traditional algorithm; Precision: decreased from 0.9990 to 0.9975, with a decrease of 0.15%, indicating that the model has a stronger ability to identify outliers in noisy data; Recall: decreased from 0.9850 to 0.9800, with a decrease of 0.51%, demonstrating relatively high robustness; AUC-ROC: decreased from 0.9930 to 0.9910, with a decrease of 0.20%, which is better than the traditional algorithm; Robustness: the accuracy decreased by 0.30%, indicating that the model's performance is more stable in the presence of noisy data; Runtime: increased from 0.45 seconds to 0.48 seconds. Although there is an increase, the impact is relatively small and within an acceptable range.
[0066] The improved isolation forest algorithm outperforms the traditional algorithm in core evaluation metrics such as precision, recall, and AUC-ROC. Especially in a data environment with noise interference, the improved algorithm shows stronger robustness and environmental adaptability; the improved model optimizes the spatial distribution characteristics of abnormal samples, effectively distinguishing minor anomalies from significant anomalies. Through the optimized visualization scheme, it provides a more intuitive decision-making basis for the real-time monitoring and early warning mechanism of supply chain anomalies; in a noise-interfered environment, the performance decay of the improved algorithm is significantly lower than that of the traditional algorithm, and the accuracy decline rate drops from 0.58% to 0.30%, demonstrating stronger anti-noise ability and model stability; although the calculation time of the improved algorithm increases, its performance is significantly improved. This trade-off between performance and efficiency better meets the dual requirements of precision and real-time in supply chain emergency management.
[0067] In today's supply chain emergency management practice, efficient and accurate anomaly detection methods are crucial for enhancing the resilience and emergency response capabilities of the supply chain. The present invention proposes a method for anomaly perception of key data in supply chain emergency management based on an improved isolation forest algorithm, which integrates multi-scale feature subset extraction technology, multi-dimensional random hyperplane boundary division strategy, and hierarchical integrated learning decision mechanism to construct a systematic multi-level anomaly detection framework and thereby improve the accuracy and robustness of anomaly detection. This method can effectively and quickly identify potential anomalies in a multi-dimensional and dynamically changing data environment and visually display the anomaly warning information. Experimental results show that the method proposed in the present invention significantly improves the anomaly perception ability of key data in supply chain emergency management, can provide technical support for an intelligent and automated supply chain management system, and can be efficiently applied in supply chain risk monitoring, early warning, and emergency response.
[0068] Figure 9 It is a schematic structural diagram of a device for anomaly perception of key data in supply chain emergency management provided by the present invention. The device for anomaly perception of key data in supply chain emergency management includes: Processing unit 1, which is used to process the corresponding data set, feature dimension set, preset subset size, and preset step length of supply chain emergency management by using a multi-scale feature subset extraction mechanism to obtain a target set containing multiple trained forest data sets; An obtaining unit 2, where the obtaining unit 2 is configured to, for each trained forest data set in the target set, randomly select two different data points from the trained forest data set, calculate a vertical vector between the two data points, determine a random intercept according to the two different data points and the vertical vector, divide the trained forest data set by using the random intercept to obtain a first subset and a second subset, for any subset, select two different data points from the subset to continue the division until a preset sample number is reached, obtain all divided subsets, and determine an isolation tree corresponding to the trained forest data set according to all the divided subsets; A calculating unit 3, where the calculating unit 3 is configured to calculate an anomaly score of each data point in the trained forest data set by using the isolation tree to obtain an anomaly score corresponding to the trained forest data set, traverse all the trained forest data sets to obtain an anomaly score corresponding to each trained forest data set, fuse the anomaly scores corresponding to all the trained forest data sets to obtain a preliminary fusion score, and process the preliminary fusion score to obtain a final anomaly detection result.
[0069] Figure 10 It is a schematic structural diagram of an electronic device provided by the present invention. As Figure 10As shown in the figure, the electronic device may include: a processor 110, a communications interface 120, a memory 130, and a communication bus 140. Among them, the processor 110, the communications interface 120, and the memory 130 complete communication with each other through the communication bus 140. The processor 110 may call the logical instructions in the memory 130 to execute the method for abnormal perception of key data in supply chain emergency management. The method includes: processing the corresponding data set, feature dimension set, preset subset size, and preset step size of supply chain emergency management by using a multi-scale feature subset extraction mechanism to obtain a target set including multiple trained forest data sets; for each trained forest data set in the target set, randomly select two different data points from the trained forest data set, calculate the vertical vector between the two data points, determine the random intercept according to the two different data points and the vertical vector, and use the random intercept to divide the trained forest data set to obtain a first subset and a second subset. For any subset, select two different data points from the subset to continue the division until the preset sample quantity is reached, obtain all the divided subsets, and determine the isolation tree corresponding to the trained forest data set according to all the divided subsets; use the isolation tree to calculate the anomaly score of each data point in the trained forest data set to obtain the anomaly score corresponding to the trained forest data set, traverse all the trained forest data sets to obtain the anomaly score corresponding to each trained forest data set, fuse the anomaly scores corresponding to all the trained forest data sets to obtain a preliminary fusion score, and process the preliminary fusion score to obtain the final anomaly detection result.
[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for abnormal perception of key data in supply chain emergency management, characterized in that, Including: Processing the dataset, feature dimension set, preset subset size, and preset step size corresponding to supply chain emergency management by using a multi-scale feature subset extraction mechanism to obtain a target set including multiple trained forest datasets; For each trained forest dataset in the target set, randomly select two different data points from the trained forest dataset, calculate the vertical vector between the two data points, determine a random intercept according to the two different data points and the vertical vector, and use the random intercept to divide the trained forest dataset to obtain a first subset and a second subset. For any subset, select two different data points from the subset to continue the division until the preset sample quantity is reached, obtain all divided subsets, and determine the isolation tree corresponding to the trained forest dataset according to all divided subsets; Calculating the anomaly score of each data point in the trained forest dataset by using the isolation tree to obtain the anomaly score corresponding to the trained forest dataset, traversing all trained forest datasets to obtain the anomaly score corresponding to each trained forest dataset, fusing the anomaly scores corresponding to all trained forest datasets to obtain a preliminary fusion score, and processing the preliminary fusion score to obtain a final anomaly detection result.
2. The method for abnormal perception of key data in supply chain emergency management according to claim 1, characterized in that, The processing the dataset, feature dimension set, preset subset size, and preset step size corresponding to supply chain emergency management by using a multi-scale feature subset extraction mechanism to obtain a target set including multiple trained forest datasets includes: When the total number of feature dimensions in the feature dimension set is greater than the preset subset size, initialize the feature subset set, set the initial index, and repeatedly execute the following steps: When the initial index is less than the total number of feature dimensions, select a feature subset of the preset subset size from the feature dimension set of the dataset corresponding to supply chain emergency management to construct a feature dataset, and train the feature dataset by using a preset isolation forest algorithm to obtain a trained forest dataset; Add the trained forest dataset to the initialized feature subset set to obtain an updated set corresponding to the updated index, where the updated index is determined by adding the preset step size to the initial index; Until the updated index is greater than or equal to the total number of feature dimensions, obtain a target set including multiple trained forest datasets.
3. The method for abnormal perception of key data in supply chain emergency management according to claim 1, characterized in that, The dividing the trained forest dataset by using the random intercept to obtain a first subset and a second subset includes: =filter(X, X·w + b ≤ 0); = filter(X, X·w + b > 0); Among them, is the first subset, is the second subset, X is the trained forest dataset, w is the vertical vector between two data points, and b is the random intercept.
4. The method for abnormal perception of key data in supply chain emergency management according to claim 1, characterized in that, The calculating the anomaly score of each data point in the trained forest dataset by using the isolation tree to obtain the anomaly score corresponding to the trained forest dataset includes: For any data point, obtain the current path length of the data point in the isolation tree and the expected path length of the data point in the isolation tree; Determine the anomaly score of the data point according to the current path length and the expected path length, traverse all data points, and determine each anomaly score corresponding to all data points; Determine the anomaly score corresponding to the trained forest dataset according to each anomaly score corresponding to all data points.
5. The method for abnormal perception of key data in supply chain emergency management according to claim 1, characterized in that, Fusing the anomaly scores corresponding to all the trained forest data sets to obtain a preliminary fusion score, including: ; Among them, is the preliminary fusion score, is the anomaly score corresponding to any trained forest dataset, is the weight coefficient corresponding to any trained forest dataset, is the number of all trained forest datasets.
6. The method for abnormal perception of key data in supply chain emergency management according to claim 1, characterized in that, Processing the preliminary fusion score to obtain a final anomaly detection result, including: When the preliminary fusion score is greater than the global anomaly threshold, determining that there are anomaly points in the corresponding data set of supply chain emergency management; Determining all data points with anomaly scores above the first quantile threshold and below the second quantile threshold as slightly anomalous data points, and determining all data points with anomaly scores above the second quantile threshold as significantly anomalous data points; The first quantile threshold is less than the second quantile threshold.
7. The method for abnormal perception of key data in supply chain emergency management according to claim 1, characterized in that, The corresponding data set of supply chain emergency management includes unique identifier features, product ID features, temperature features, process temperature features, rotational speed features, torque features, tool wear features, machine failure label features, tool wear failure features, heat dissipation failure features, power failure features, overstrain failure features, and random failure features.
8. The method for abnormal perception of key data in supply chain emergency management according to claim 7, characterized in that, Before processing the corresponding data set, feature dimension set, preset subset size, and preset step size of supply chain emergency management using the multi-scale feature subset extraction mechanism, the method further includes: Performing data cleaning and denoising on the original supply chain emergency management data to obtain a first data set; Performing feature selection and extraction on the first data set to obtain a second data set; Performing data standardization and normalization processing on the second data set to obtain the corresponding data set of supply chain emergency management.
9. A key data anomaly perception device for supply chain emergency management, characterized in that, Including: A processing unit for processing the corresponding data set, feature dimension set, preset subset size, and preset step size of supply chain emergency management using the multi-scale feature subset extraction mechanism to obtain a target set including multiple trained forest data sets; An acquisition unit for, for each trained forest data set in the target set, randomly selecting two different data points from the trained forest data set, calculating the vertical vector between the two data points, determining a random intercept according to the two different data points and the vertical vector, dividing the trained forest data set using the random intercept to obtain a first subset and a second subset, for any subset, selecting two different data points from the subset to continue splitting until the preset sample number is reached, obtaining all the split subsets, and determining the isolation tree corresponding to the trained forest data set according to all the split subsets; A calculation unit for calculating the anomaly score of each data point in the trained forest data set using the isolation tree to obtain the anomaly score corresponding to the trained forest data set, traversing all the trained forest data sets to obtain the anomaly score corresponding to each trained forest data set, fusing the anomaly scores corresponding to all the trained forest data sets to obtain a preliminary fusion score, and processing the preliminary fusion score to obtain a final anomaly detection result.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the key data anomaly perception method for supply chain emergency management according to any one of claims 1 to 8.
Citation Information
Patent Citations
Hair restorer production process management method and system
CN117391641A
Equipment inspection method and inspection system
CN117585554A
Abnormality determination method and device for wind turbine generator, medium, equipment and program product
CN119244458A