Data anomaly detection method for vehicle part production
By combining the isolation forest algorithm and the maximum inter-class variance algorithm, the anomaly detection problem of high-dimensional vehicle parts production data was solved, and efficient identification of local and global anomalies was achieved, which optimized the production process and improved production efficiency and product quality.
Patent Information
- Application Number
- CN202510858475.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing abnormal data detection methods are inefficient when processing high-dimensional and large-volume vehicle parts production data, making it difficult to meet production needs, especially when identifying local anomalies.
The isolation forest algorithm is used to build a model to detect anomalies in vehicle parts production data. The node evaluation strategy and the maximum inter-class variance algorithm are combined to adaptively set the threshold to improve the ability to identify abnormal data.
It realizes comprehensive monitoring of vehicle parts production data, can identify global and local anomalies, reduce misjudgment rate, optimize production parameters, and improve production efficiency and product quality.
Smart Images

Figure BDA0005466746640000071 
Figure BDA0005466746640000081 
Figure BDA0005466746640000082
Abstract
Description
Technical Field
[0001] The present application relates to the field of industrial data processing, and in particular to a data anomaly detection method for vehicle parts production. Background Art
[0002] Industrial production data anomaly detection involves analyzing various production data to identify data points or sequences that do not conform to normal production patterns. In the vehicle parts production process, real-time monitoring of production data and the timely identification of anomalies can optimize production processes and adjust production parameters in a timely manner, avoiding production halts and product quality degradation caused by abnormalities, thereby improving overall production efficiency.
[0003] Existing methods for detecting abnormal data include statistics, distance metrics, density methods, and clustering. Statistical methods primarily detect abnormal data based on the distribution of data and excel at detecting univariate data. Distance metrics detect outliers by calculating the distance between data points, such as the K-nearest neighbor detection method, which has the ability to detect global outliers. Density methods detect outliers by determining whether sample points are in areas with low density and are capable of identifying local anomalies. Clustering methods identify outliers based on the relationships between data domain classes. With the increase in data dimensionality and volume, existing methods for detecting abnormal data are struggling to meet expectations in industrial information and data processing. Therefore, a more efficient data anomaly detection method is urgently needed to process large-volume and high-dimensional vehicle parts production data. Summary of the Invention
[0004] To address the technical issues of existing technologies, this application provides a data anomaly detection method for vehicle parts production. This method uses an isolation forest algorithm to detect anomalies in production data during the vehicle parts production process, and introduces a node evaluation strategy to comprehensively evaluate production data samples. Furthermore, within the isolation forest algorithm, a maximum inter-class variance algorithm is used to adaptively select the anomaly score threshold for the isolation forest, reducing errors. This invention addresses the issue of incomplete sample evaluation and improves the ability to detect anomaly data.
[0005] This application provides a data anomaly detection method for vehicle parts production, comprising:
[0006] (1) Collect relevant historical production data from the production process of vehicle parts and pre-process the collected production data;
[0007] (2) Constructing an isolation forest model based on the preprocessed historical production data of vehicle parts;
[0008] (3) Inputting the real-time collected vehicle parts production data into the isolation forest model, calculating the average path length, and further calculating the outlier score using the node evaluation strategy;
[0009] (4) The maximum inter-class variance algorithm is used to adaptively determine the outlier score threshold and to identify abnormal data in the vehicle parts production data based on the threshold;
[0010] (5) According to the abnormal data detection results of the isolation forest model, the relevant parameter settings in the vehicle parts production process are adjusted to optimize the production process.
[0011] Furthermore, the collected vehicle parts production data includes production equipment data and quality inspection data. Production equipment data is collected through sensors installed on production equipment and includes machine tool spindle speed, hydraulic system pressure, motor current, and ambient temperature. Quality inspection data is obtained through parts quality inspection equipment and includes part dimensional data, hardness data, and surface roughness data.
[0012] Furthermore, an industrial automation system is used to automatically collect data. The system is connected to sensors and testing equipment via a communication protocol, reading data in real time and transmitting it to a database. The database uses a relational database to store structured data based on the production process. It records production data at each timestamp within the sampling period, with each row representing a record and each column representing the type of production data.
[0013] Furthermore, the pre-processing method of the collected historical production data of vehicle parts includes data cleaning and data standardization.
[0014] First, check whether there are exactly the same duplicate records in the data, and delete the duplicate records using the deduplication function in the data processing tool, retaining only one valid data; second, check the error values in the data, that is, data that exceeds the normal measurement range of sensors and detection equipment, and mark them as missing data; then, for the missing data in the data, use the linear interpolation method to fill the missing data by calculating the average value of the data in the neighborhood of the missing data; finally, use Z-score standardization to convert the data into a distribution with a mean of 0 and a standard deviation of 1, so that different types of production data are on the same scale, which is convenient for subsequent modeling.
[0015] Furthermore, based on the pre-processed historical production data of vehicle parts, the detailed steps of constructing the isolation forest model include:
[0016] Determine the isolation forest parameters, i.e., the number of isolated trees and subsample size; the number of isolated trees is determined based on the size of the historical production dataset, and the subsample size is 25%-30% of the dataset size;
[0017] Randomly select n samples from the historical production data set without replacement as the root nodes of the isolation tree;
[0018] Randomly select a type of production data, randomly select a split value between the maximum and minimum values of the data, and divide the data set into two child nodes on the left and right;
[0019] Repeat the above steps to divide the left and right child nodes until the number of data in the child nodes is less than or equal to 1, thus establishing an isolated tree;
[0020] According to the above method, t isolated trees are generated to form an isolation forest model.
[0021] Furthermore, each data point in the real-time collected vehicle parts production data is input into each isolated tree in the isolation forest model, and its path length h(x) in each tree is calculated, that is, the number of edges passed from the root node to the leaf node plus the average path length from the root node to the leaf node in the tree.
[0022] A node evaluation strategy is used to calculate outlier scores based on the average path length of the isolation tree. Traditional isolation forest scoring mechanisms only analyze the global anomaly level of data points, lacking consideration of local anomalies. This limits their ability to detect abnormal data. However, in the production of automotive parts, subtle changes in data are crucial for adjusting production parameters and are fundamental to ensuring part quality.
[0023] The node evaluation mechanism utilizes the number and degree of separation of data points by the isolation tree, introducing the concepts of node depth and relative quality into the scoring process. Node depth represents the number of data point divisions and is a measure of global anomalies. Relative quality is defined as the ratio of the number of data points in a leaf node to the number of data points in its root node. It represents the degree of data point dispersion and is a measure of local anomalies.
[0024] Furthermore, the maximum inter-class variance algorithm is applied to adaptively determine the outlier score threshold for vehicle parts production data. The maximum inter-class variance algorithm is a method for determining the threshold for image binarization segmentation. Therefore, the outlier scores for the vehicle parts production data are first converted into a scatter plot. This scatter plot is then converted into a grayscale image containing both foreground and background pixels. The ratio of foreground and background pixels to the total image is calculated. Based on this ratio, the average grayscale of the foreground and background is calculated. Based on this ratio, the inter-class variance is calculated. The score threshold that maximizes the variance is determined and used as the optimal outlier score threshold.
[0025] The present invention discloses the following technical effects:
[0026] The present invention proposes a method for detecting abnormal data in vehicle parts production. By establishing an isolation forest model for the collected vehicle parts production data, anomaly detection is performed on the real-time collected production data. The present invention uses an industrial automation data acquisition system to automatically collect multiple types of production data in the vehicle parts production process, overcoming the single data source limitation of data anomaly detection, enriching data diversity, and achieving comprehensive monitoring of production data. In addition, the isolation forest algorithm adopted by the present invention introduces a node evaluation mechanism to improve the calculation method of the anomaly score, so that it has the ability to identify both global and local anomalies, and captures minor anomalies in vehicle parts production data, which is crucial for improving production parameters. When judging abnormal data points in the production of multiple vehicle parts, the maximum inter-class variance algorithm is introduced to adaptively determine the outlier score threshold, reduce interference between different types of abnormal data points, and reduce the error rate. The isolation forest algorithm has advantages in industrial information and data processing, especially its ability to process abnormal data. The present invention applies the algorithm to the production of vehicle parts, monitors abnormal situations in various production data, conducts in-depth analysis of the abnormal situations in a timely manner, adjusts the parameter settings of production equipment, and further optimizes the production process; the algorithm has low complexity and strong real-time performance, and while improving the production efficiency of vehicle parts, it also ensures the quality of the parts products and reduces the product failure rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention are briefly introduced below. Flowcharts are used in this application to illustrate the operations performed according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in precise order. Instead, various steps may be processed in reverse order or simultaneously as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0028] Figure 1 A flowchart of a data anomaly detection method for vehicle parts production provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below.
[0030] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0031] In the following description, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules that are not explicitly listed or are inherent to these processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application only.
[0032] The present application embodiment provides a data anomaly detection method for vehicle parts production, such as Figure 1 As shown, the method includes:
[0033] Step S10: collecting relevant historical production data from the vehicle parts production process and preprocessing the collected production data.
[0034] In this embodiment, the collected vehicle parts production data includes production equipment data and quality inspection data. Production equipment data is collected through sensors installed on the production equipment and includes machine tool spindle speed, hydraulic system pressure, motor current, and ambient temperature. Quality inspection data is obtained through parts quality inspection equipment and includes part dimensional data, hardness data, and surface roughness data.
[0035] Sensors installed on production equipment include vibration sensors, pressure sensors, voltage and current sensors, and temperature sensors, which monitor equipment operating status in real time. Quality inspection data is collected using coordinate measuring machines, hardness testers, and surface roughness testers. During the data collection process, the industrial parameter settings of production equipment are recorded to provide a reference for parameter adjustment.
[0036] An industrial automation system is used to automatically collect data. The system sets the collection time and cycle, connects the system to sensors and detection equipment via a communication protocol, and reads and transmits data to a database in real time. The system uses a programmable logic controller (PLC) to control the data collection process through a logic control program. The Modbus industrial communication protocol is used for data transmission.
[0037] The database uses a relational database based on MySQL to store structured data based on the production process, recording the production data of each timestamp within the sampling time. Each row represents a record, and each column represents the production data type.
[0038] The preprocessing methods for the collected historical production data of vehicle parts include data cleaning and data standardization. First, the data is checked for identical duplicate records. The duplicate records are deleted using the deduplication function in the data processing tool, retaining only one valid data. Second, the data is checked for error values, that is, data outside the normal measurement range of sensors and detection equipment, and marked as missing data. Then, for missing data in the data, linear interpolation is used to calculate the average value of the data in the neighborhood of the missing data to fill the gap. Finally, Z-score standardization is used to transform the data into a distribution with a mean of 0 and a standard deviation of 1. The standardization process is expressed as follows:
[0039]
[0040] Among them, x is the original data point, x std is the standardized data point, μ and σ are the mean and standard deviation of this type of production data in the sampling period, respectively; data standardization processing makes different types of production data on the same scale, which is convenient for subsequent modeling.
[0041] Step S20: constructing an isolation forest model based on the preprocessed historical production data of vehicle parts.
[0042] In this embodiment, the detailed steps of constructing an isolation forest model based on the preprocessed historical production data of vehicle parts include:
[0043] Determine the isolation forest parameters, i.e., the number of isolated trees and subsample size; the number of isolated trees is determined based on the size of the historical production dataset, and the subsample size is 25%-30% of the dataset size;
[0044] Randomly select n samples from the historical production data set without replacement as the root nodes of the isolation tree;
[0045] Randomly select a type of production data, randomly select a split value between the maximum and minimum values of the data, and divide the data set into two child nodes on the left and right;
[0046] Repeat the above steps to divide the left and right child nodes until the number of data in the child nodes is less than or equal to 1, thus establishing an isolated tree;
[0047] According to the above method, t isolated trees are generated to form an isolation forest model.
[0048] In step S30 , the real-time collected vehicle parts production data is input into the isolation forest model, the average path length is calculated, and the outlier score is further calculated using a node evaluation strategy.
[0049] In this embodiment, each data point in the real-time collected vehicle parts production data is input into each isolated tree in the isolation forest model, and its path length h(x) in each tree is calculated. This is the sum of the number of edges passed from the root node to the leaf node plus the average path length from the root node to the leaf node in the tree. The calculation formula is as follows:
[0050] h(x)=e+c(T.size)
[0051] Where e is the number of edges that data point x experiences from the root node to the leaf node of the tree, that is, the number of times the node is partitioned; T.size represents the number of data points that share the same leaf node as data point x, and c(T.size) is a correction value representing the average path length of an isolated tree constructed from T.size data points. The calculation formula is as follows:
[0052]
[0053] Among them, c(n) represents the average path length of n data points to construct an isolation tree.
[0054] Based on the average path length of the isolation tree, a node evaluation strategy is used to calculate the outlier score. The node evaluation mechanism uses the number and degree of separation of data points by the isolation tree to introduce the concepts of node depth and relative quality into the scoring process. Node depth represents the number of data points divided and is a consideration of global anomalies; relative quality is defined as the ratio of the number of data points in a leaf node to the number of data points in its root node, representing the degree of dispersion of the data points and a consideration of local anomalies. The calculation formula for the outlier score S(x) is expressed as:
[0055]
[0056] Where u(·) represents the number of data points in the node; l(x) represents the leaf node where data point x is located; and f(x) represents the root node where data point x is located. S(x) is normalized to the range [0, 1] to facilitate setting the outlier score threshold. The average of the outlier scores of t isolated trees is taken as the final outlier score for data point x.
[0057] Step S40 , applying the maximum inter-class variance algorithm to adaptively determine an outlier score threshold, and demarcate abnormal data in the vehicle parts production data according to the threshold.
[0058] In this embodiment, the maximum inter-class variance algorithm is used to adaptively determine the outlier score threshold to improve the threshold accuracy and reduce the misjudgment of abnormal data points. The steps for adaptively setting the outlier score threshold are as follows:
[0059] The final outlier score is plotted in the form of a scatter plot, which is achieved by calling the scatter statement in MATLAB software;
[0060] Call the itshow statement to convert the scatter plot into a grayscale image, including the foreground and background, with a grayscale range of [0,255], and calculate the proportion of the foreground and background to the entire image. The formulas are:
[0061]
[0062] Among them, k is the threshold for distinguishing foreground and background; p i is the proportion of pixels with gray level i in the entire image; ω0 and ω1 are the proportions of foreground and background pixels in the entire image respectively. The average grayscale calculation formula for foreground and background is:
[0063]
[0064] Among them, u0 and u1 are the average grayscale of foreground and background respectively. Based on the above, let the global average grayscale u=ω0u0+ω1u1, and the inter-class variance σ 2 The calculation formula is:
[0065] σ 2 =ω0(u0-u) 2 +ω1(u1-u) 2
[0066] The score threshold k is traversed from the 256 grayscale levels in ascending order. When the inter-class variance reaches its maximum, the difference between normal and abnormal data points is the greatest, and the k value at this time is the optimal threshold. Data points above the threshold are identified as abnormal data, and data points below the threshold are identified as normal data, thus filtering out abnormalities in the vehicle parts production data.
[0067] Step S50 , adjusting relevant parameter settings in the vehicle parts production process according to the abnormal data detection results of the isolation forest model to optimize the production process.
[0068] The above specific embodiments do not constitute a limitation to the scope of protection of this application. It should be understood by those skilled in the art that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of this application should be included in the scope of protection of this application. In some cases, the actions or steps recorded in this application can be performed in an order different from that in the embodiments and can still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A data anomaly detection method for vehicle parts production, characterized in that: The method comprises: (1) Collect relevant historical production data from the production process of vehicle parts and pre-process the collected production data; (2) Constructing an isolation forest model based on the preprocessed historical production data of vehicle parts; (3) Inputting the real-time collected vehicle parts production data into the isolation forest model, calculating the average path length, and further calculating the outlier score using the node evaluation strategy; (4) The maximum inter-class variance algorithm is used to adaptively determine the outlier score threshold and to identify abnormal data in the vehicle parts production data based on the threshold; (5) According to the abnormal data detection results of the isolation forest model, the relevant parameter settings in the vehicle parts production process are adjusted to optimize the production process.
2. The data anomaly detection method for vehicle parts production according to claim 1, characterized in that: In step (1), the vehicle parts production data includes production equipment data and quality inspection data; the production equipment data is collected through sensors installed on the production equipment, including machine tool spindle speed, hydraulic system pressure, motor current and ambient temperature; The quality inspection data is obtained through component quality inspection equipment, including component size data, hardness data and surface roughness data.
3. The data anomaly detection method for vehicle parts production according to claim 2, characterized in that: Use industrial automation systems to automatically collect data, set collection time and collection cycle, connect the system with sensors and detection equipment through communication protocols, read data in real time and transmit it to the database; The database uses a relational database to store structured data based on the production process, recording the production data of each timestamp within the sampling time. Each row represents a record, and each column represents the type of production data.
4. The data anomaly detection method for vehicle parts production according to claim 1, characterized in that: In step (1), the method of preprocessing the collected historical production data of vehicle parts includes data cleaning and data standardization.
5. The data anomaly detection method for vehicle parts production according to claim 4, characterized in that: The detailed steps of data preprocessing include: Check whether there are identical duplicate records in the data, and use the deduplication function in the data processing tool to delete the duplicate records, leaving only one valid data; Check the data for erroneous values, i.e., data outside the normal measurement range of sensors and detection equipment, and mark them as missing data; For missing data in the data, linear interpolation method is used to fill it by calculating the average value of the data in the neighborhood of the missing data; Z-score standardization is used to transform the data into a distribution with a mean of 0 and a standard deviation of 1, so that different types of production data are on the same scale.
6. The data anomaly detection method for vehicle parts production according to claim 1, characterized in that: In step (2), the detailed steps of constructing the isolation forest model include: Determine the isolation forest parameters, i.e., the number of isolated trees and subsample size; the number of isolated trees is determined based on the size of the historical production dataset, and the subsample size is 25%-30% of the dataset size; Randomly select n samples from the historical production data set without replacement as the root nodes of the isolation tree; Randomly select a type of production data, randomly select a split value between the maximum and minimum values of the data, and divide the data set into two child nodes on the left and right; Repeat the above steps to divide the left and right child nodes until the number of data in the child nodes is less than or equal to 1, thus establishing an isolated tree; According to the above method, t isolated trees are generated to form an isolation forest model.
7. The data anomaly detection method for vehicle parts production according to claim 1, characterized in that: In step (3), each data point in the real-time collected vehicle parts production data is input into each isolated tree in the isolation forest model, and its path length h(x) in each tree is calculated, that is, the number of edges passed from the root node to the leaf node plus the average path length from the root node to the leaf node in the tree. The calculation formula is as follows: h(x)=e+c(T.size) Where e is the number of edges that data point x experiences from the root node to the leaf node of the tree, that is, the number of times the node is partitioned; T.size represents the number of data points that share the same leaf node as data point x, and c(T.size) is a correction value representing the average path length of an isolated tree constructed from T.size data points. The calculation formula is as follows: Among them, c(n) represents the average path length of n data points to construct an isolation tree.
8. The data anomaly detection method for vehicle parts production according to claim 1, characterized in that: In step (3), based on the average path length of the isolation tree, the node evaluation strategy is used to calculate the outlier score to detect local and global anomalies in the vehicle parts production data. The calculation formula of the outlier score S(x) is expressed as: Where u(·) represents the number of data points in the node; l(x) represents the leaf node where the data point x is located; f(x) represents the root node where the data point x is located; S(x) is normalized to the range of [0,1] to facilitate the subsequent setting of the outlier score threshold; the average of the outlier scores of t isolated trees is taken as the final outlier score of the data point x.
9. The data anomaly detection method for vehicle parts production according to claim 1, characterized in that: In step (4), the step of applying the maximum inter-class variance algorithm to adaptively determine the outlier score threshold includes: The final outlier score is plotted in the form of a scatter plot, converted into a grayscale image, including the foreground and background, with a grayscale range of [0,255], and the proportion of the foreground and background in the entire image is calculated. The formulas are: Among them, k is the threshold for distinguishing foreground and background; p i is the proportion of pixels with gray level i in the entire image; ω0 and ω1 are the proportions of foreground and background pixels in the entire image respectively; the average grayscale calculation formula for foreground and background is: Among them, u0 and u1 are the average grayscale of the foreground and background respectively; on this basis, let the global average grayscale u=ω0u0+ω1u1, and the inter-class variance σ - The calculation formula is: s 2 =ω0(u0-u) 2 +ω1(u1-u) 2 The score threshold k is traversed from the 256 gray levels in ascending order. When the inter-class variance reaches the maximum, the difference between normal data points and abnormal data points is the largest, and the k value at this time is the optimal threshold.