Digital factory production data abnormity self-checking method and system and medium

By classifying and resampling production data from digital factories, constructing training datasets, and fine-tuning anomaly detection models, the problem of insufficient learning of new data in digital factories is solved, and the accuracy and timeliness of production data anomaly self-inspection are improved.

CN121211284AActive Publication Date: 2025-12-26CHANGCHUN INST OF TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511746412.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2025-12-26
Estimated Expiration
2045-11-26

AI Technical Summary

Technical Problem

When faced with massive amounts of production data, existing digital factories struggle to quickly update their self-inspection thresholds based on new production data, resulting in poor accuracy in self-inspection of abnormal production data. Furthermore, newly emerging data that differs from previous data is offset by a large amount of already learned data, reducing learning efficiency.

Method used

By classifying current production data and historical production data to form production data clusters, and resampling the second production data based on data redundancy indicators, a training dataset is constructed, and the anomaly detection model is fine-tuned to improve self-inspection accuracy.

Benefits of technology

It effectively filters out information valuable for learning from new data, avoids interference from duplicate data, and improves the accuracy and timeliness of self-checking for anomalies in production data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121211284A_ABST
    Figure CN121211284A_ABST
Patent Text Reader

Abstract

The invention discloses a digital factory production data abnormity self-checking method and system and a medium, and relates to the technical field of data processing. The method comprises the following steps: classifying first production data in current production data and second production data in historical production data to obtain a plurality of production data clusters; comparing the first production data with the second production data based on the production data cluster, and determining a data redundancy index between the first production data and the second production data; the data redundancy index is used for representing data rule consistency between the first production data and the second production data; resampling the second production data based on the data redundancy index, and constructing a training data set of the first production data; and based on the training data set, performing fine tuning on the anomaly detection model, so that data anomaly self-inspection is performed on the first production data based on the fine-tuned anomaly detection model. According to the invention, the accuracy of production data abnormity self-inspection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to a method, system, and medium for self-inspection of anomalies in production data in a digital factory. Background Technology

[0002] Digital factories, as an important development direction of modern manufacturing, leverage digital technologies such as the Internet of Things and big data to uniformly manage enterprise factory production-related information, achieving efficient, automated, and adaptive advanced manufacturing. While digital factories can effectively improve production efficiency, they are highly dependent on production data. The accuracy and reliability of this data are crucial for ensuring production efficiency, product quality, and operating costs. Therefore, timely monitoring of production data is essential to avoid the negative impact of abnormal data.

[0003] Existing technologies mainly predict and analyze production data in digital factories by pre-setting data logic to determine the normal range of production data. When production data exceeds the normal range, anomalies can be detected, thereby achieving self-checking of production data anomalies.

[0004] However, digital factories cover the entire product manufacturing process, generating massive amounts of production data every moment at an extremely high speed. Existing methods for determining the normal range of production data mostly rely on learning from production data to adaptively adjust. A large amount of this data overlaps with previously learned data, and new production data that differs from previous data is often offset by a large amount of already learned data. This reduces the efficiency of learning from new data and makes it difficult to quickly update self-inspection thresholds based on new production data, resulting in poor accuracy in self-inspection of production data anomalies. Summary of the Invention

[0005] This invention provides a method, system, and medium for self-inspection of production data anomalies in a digital factory, which can improve the accuracy of self-inspection of production data anomalies.

[0006] A first aspect of this invention provides a method for self-detecting anomalies in production data of a digital factory, comprising: classifying first production data in current production data and second production data in historical production data to obtain several production data clusters; the first production data being production data with the same equipment source and processing attributes in the current production data, and the second production data being production data with the same equipment source and processing attributes as the first production data in historical production data, wherein each production data cluster includes first production data and second production data in the same processing state; comparing the first production data and the second production data based on the production data clusters to determine a data redundancy index between the first production data and the second production data; the data redundancy index being used to characterize the consistency of data patterns between the first production data and the second production data; resampling the second production data based on the data redundancy index to construct a training dataset for the first production data; and fine-tuning anomaly detection model based on the training dataset to enable the fine-tuned anomaly detection model to perform data anomaly self-detection on the first production data.

[0007] Furthermore, the present invention proposes to classify the first production data in the current production data and the second production data in the historical production data to obtain several production data clusters, including: for each production data in the first production data and the second production data, constructing a target processing vector of the production data with the processing attribute of the production data as the dimension; for each target processing vector, determining the degree of change of the processing state of the target processing vector based on the similarity between the target processing vector and the reference processing vector at the previous sampling time; and clustering each target processing vector based on the degree of change of the processing state of each target processing vector to obtain several production data clusters.

[0008] Furthermore, this invention proposes a method based on a production data cluster to compare first production data with second production data and determine a data redundancy index between the first and second production data. This includes: obtaining the current production stage corresponding to the current processing vector sequence; the current processing vector sequence is formed by arranging each target processing vector in the first production data in chronological order; determining the product consistency probability between the current production stage and historical production stages based on the material property differences between the current production stage and historical production stages; historical production stages are the production stages corresponding to historical processing vector sequences, which are formed by arranging each target processing vector in the second production data in chronological order; filtering each historical production stage based on the product consistency probability between the current production stage and each historical production stage to determine the reference production stage corresponding to the current production stage; determining the production data consistency degree between adjacent reference production stages based on the difference in the degree of change in processing status between adjacent reference production stages and the production data cluster; identifying adjacent reference production stages with a production data consistency degree less than a preset consistency degree threshold as an optimization process group; and determining the data redundancy index between the first and second production data based on the differences between the production stages before and after optimization in each optimization process group.

[0009] Furthermore, the present invention also proposes to obtain the current production stage corresponding to the current processing vector sequence, including: arranging each target processing vector in the first production data in chronological order to construct the current processing vector sequence; constructing a processing state change sequence based on the absolute value of the difference in the degree of change of processing state between two adjacent target processing vectors in the current processing vector sequence; and dividing the current processing vector sequence into at least one current production stage based on the processing state change sequence.

[0010] Furthermore, this invention proposes determining the likelihood of product consistency between the current production stage and historical production stages based on the differences in material properties between the current and historical production stages. This includes: standardizing the difference between various current material property values ​​in the current production stage and their corresponding historical material property values ​​in the historical production stages to obtain several standard material property difference values; dividing the sum of the several standard material property difference values ​​by the total number of material properties to obtain the average contribution value of a single standard material property difference value; and determining the likelihood of product consistency between the current and historical production stages based on the average contribution value of a single standard material property difference value.

[0011] Furthermore, the present invention proposes to determine the consistency of production data between adjacent reference production stages based on the difference in the degree of change of processing state between adjacent reference production stages and the production data clusters, including: averaging the differences in the corresponding degrees of change of processing state between adjacent reference production stages to obtain the comprehensive degree of difference in processing state between adjacent reference production stages; obtaining the number of first processing vectors located in the same production data cluster between adjacent reference production stages; dividing the number of first processing vectors by the total number of processing vectors in adjacent reference production stages to obtain the proportion of processing vectors; and determining the consistency of production data between adjacent reference production stages based on the comprehensive degree of difference in processing state and the proportion of processing vectors.

[0012] Furthermore, the present invention proposes to determine a data redundancy index between the first production data and the second production data based on the differences between the production stages before and after optimization in each optimization process group. This includes: determining the processing optimization vector of each optimization process group based on the difference between the corresponding target processing vectors of the production stages before and after optimization in each optimization process group; determining the comprehensive optimization consistency between the current production stage and historical production stages based on the processing optimization vectors of each optimization process group; and determining the data redundancy index between the first production data and the second production data based on the comprehensive optimization consistency between the current production stage and historical production stages.

[0013] Furthermore, the present invention proposes that the data redundancy index includes the sub-redundancy index of the first production data in each current production stage; based on the data redundancy index, the second production data is resampled to construct a training dataset of the first production data, including: using the sub-redundancy index of each current production stage and the original sampling frequency of the reference production stage corresponding to the current production stage in the second production data to determine the resampling frequency of each reference production stage; based on the resampling frequency of each reference production stage, the second production data of each reference production stage is resampled to construct a training dataset of the first production data.

[0014] A second aspect of this invention provides a digital factory production data anomaly self-inspection system, comprising: a data classification module, used to classify first production data in current production data and second production data in historical production data to obtain several production data clusters; the first production data refers to production data in current production data that have the same equipment source and the same processing attributes, and the second production data refers to production data in historical production data that have the same equipment source and the same processing attributes as the first production data, and the production data clusters include first production data and second production data in the same processing state; a data comparison module, used to compare the first production data and the second production data based on the production data clusters to determine a data redundancy index between the first production data and the second production data; the data redundancy index is used to characterize the consistency of data patterns between the first production data and the second production data; a data resampling module, used to resample the second production data based on the data redundancy index to construct a training dataset for the first production data; and a model training module, used to fine-tune anomaly detection model based on the training dataset so that the fine-tuned anomaly detection model can perform data anomaly self-inspection on the first production data.

[0015] A third aspect of the present invention provides a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the above-described method for self-checking anomalies in digital factory production data.

[0016] The present invention has the following beneficial effects: The digital factory production data anomaly self-detection method provided in this invention first categorizes first and second production data with the same equipment source and processing attributes from the current and historical production data into production data clusters, ensuring that each cluster covers production data in the same processing state. Next, a data redundancy index is determined by comparing the first and second production data, confirming the consistency of data patterns between them. Based on this data redundancy index, a training dataset is constructed by resampling the second production data, effectively filtering out data valuable for learning from new data and avoiding interference from duplicate data. Finally, the anomaly detection model is fine-tuned using this training dataset, enabling it to more accurately perform anomaly self-detection on the first production data within the current production data, overcoming the problem of insufficient learning from new data in existing methods, thereby improving the accuracy of production data anomaly self-detection. Attached Figure Description

[0017] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating a method for self-checking anomalies in production data in a digital factory, as provided in one embodiment of the present invention. Figure 2 This is a schematic flowchart of S200 provided in one embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a digital factory production data anomaly self-inspection system provided in one embodiment of the present invention. Detailed Implementation

[0019] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a digital factory production data anomaly self-inspection method, system, and medium proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0020] In specific embodiments of this invention, various mathematical calculations and formulas will be involved. Those skilled in the art should understand that certain specific boundary conditions may arise when processing actual production data. For example, when performing operations involving division, if the result of the calculation of the variable or expression used as the denominator is zero, this invention can employ numerical processing methods known in the art to ensure the robustness and numerical stability of the calculation.

[0021] Since the numerical range of the denominator in the various calculation formulas involved in this invention is greater than 0, and may be equal to 0 in extreme cases, an exemplary processing method is to add a very small positive number to the denominator (for example, a preset non-zero constant, such as 0.001, the specific value of which can be set by the implementer according to the actual situation, and this application embodiment does not make a specific limitation), so that the denominator is not 0 after being added to a very small positive number, so as to avoid the error of dividing by zero.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0023] In traditional production data anomaly detection methods, adaptive models trained on historical data face the dual challenges of data redundancy and feature dilution. When production equipment continuously generates data with the same source and attributes under the same processing conditions, historical datasets and real-time data streams exhibit highly repetitive features, leading to ineffective iterations during model training. Simultaneously, newly emerging non-repetitive data features are submerged in a large number of historical samples, making it difficult for the model to quickly identify subtle shifts in data patterns. This results in anomaly detection threshold updates lagging behind actual changes in production conditions.

[0024] Faced with the aforementioned problems, this invention first considers how to effectively distinguish repetitive features in real-time data streams and historical data. Traditional methods directly mix new data with all historical data for training, causing the model to fail to focus on effective shifts in data patterns. To address this, this invention attempts to approach the issue from a data classification perspective, clustering production data with similar origins and attributes according to processing status to form data clusters with consistent patterns. Further analysis reveals that the degree of data redundancy varies across different processing statuses; applying a uniform sampling strategy to historical data may dilute the features of new data. Based on this, this invention proposes quantifying the consistency of patterns between real-time and historical data under the same processing status and dynamically adjusting the sampling weights of historical data, thereby enhancing the learning of new data patterns while preserving effective historical features.

[0025] In this regard, such as Figure 1 As shown, this invention provides a flowchart of a method for self-checking anomalies in production data of a digital factory. This method can be applied to electronic devices and may include the following steps S100 to S400: S100, classify the first production data in the current production data and the second production data in the historical production data to obtain several production data clusters; the first production data is the production data in the current production data that has the same equipment source and the same processing attributes, and the second production data is the production data in the historical production data that has the same equipment source and the same processing attributes as the first production data. The production data cluster includes the first production data and the second production data that are in the same processing state.

[0026] In this step, current production data refers to the data generated by production activities in the current time period, reflecting real-time information such as the operating status of production equipment, processing parameters, and production progress in the current period; historical production data refers to the data generated by production activities in the past period. This data records various situations in the past production process and can be used to analyze production patterns and compare the current production status.

[0027] First-line production data refers to production data within the current production data that meets the criteria of having the same equipment source and the same processing attributes. Same equipment source means the data was generated by the same piece of equipment or a group of equipment with the same model and function; same processing attributes indicate that the processing technology, processing parameters, etc., corresponding to these data are consistent.

[0028] Secondary production data refers to historical production data that shares the same equipment source and processing attributes as the primary production data. In other words, it is historical data generated by the same specific equipment and processed in the same way as the primary production data.

[0029] A production data cluster refers to several data sets obtained by classifying first production data and second production data. The data in each production data cluster are in the same processing state and include first production data from current production and second production data from historical production.

[0030] Specifically, the process begins by acquiring current production data, which reflects real-time information such as equipment operation, processing parameters, and production progress generated during the current production period; and historical production data, which records production activities over a past period. Next, data from the current production data that shares the same equipment source and processing attributes is selected as the first production data. Then, historical production data that shares both the same equipment source and processing attributes with the first production data is identified as the second production data. Finally, the first and second production data are categorized based on their processing status, grouping data in the same processing state together to obtain several production data clusters.

[0031] By filtering data based on two key dimensions—equipment origin and processing attributes—it is ensured that data included in the same data cluster share similar generation environments and processing characteristics. Categorizing data according to the same processing state is crucial because data under the same processing state exhibits more similar patterns and characteristics. This categorization facilitates subsequent analysis and comparison of data patterns, laying the foundation for accurately assessing data redundancy and constructing effective training datasets.

[0032] S200, based on the production data cluster, compare the first production data with the second production data to determine the data redundancy index between the first production data and the second production data; the data redundancy index is used to characterize the consistency of data patterns between the first production data and the second production data.

[0033] In this step, the data redundancy index refers to an indicator used to characterize the consistency of data patterns between the first and second production data. By comparing the first and second production data, the similarity between them in terms of data distribution, trends, and correlations is analyzed to derive the data redundancy index. This data redundancy index quantifies the proximity of data patterns between the two; a higher value indicates a more consistent data pattern and potentially more redundant information, while a lower value indicates a greater difference in data patterns.

[0034] Specifically, a detailed comparison is first made between the first and second production data. In terms of data distribution, the similarity in central tendency and dispersion of the data values ​​is examined; trends are analyzed to observe whether the data changes consistently over time or due to other factors; correlations are studied to determine whether a relationship exists between the two and the strength of that relationship. Based on the combined results of these analyses and the production data clusters, a data redundancy index is calculated to characterize the consistency of the data patterns between the two datasets.

[0035] Among these factors, data distribution, trends, and correlations are crucial aspects reflecting data patterns. Similar data distributions indicate that the data are similar in value range and concentration; consistent trends mean that the data exhibit the same patterns of change over time or with other variables; and strong correlations indicate a stable relationship between the two. By comprehensively considering these factors, the data redundancy index can accurately quantify the similarity in data patterns between the first and second production data, providing a key basis for subsequent resampling operations.

[0036] S300, based on the data redundancy index, resamples the second production data to construct the training dataset of the first production data.

[0037] In this step, resampling refers to the operation of resampling the second production data. Based on the data redundancy index, a portion of data is selectively extracted from the second production data to construct a training dataset suitable for training models related to the first production data. The purpose of resampling may be to balance the amount of data, remove redundant data, or enhance the representativeness of the data, so that the constructed training dataset can better reflect the characteristics and patterns related to the first production data.

[0038] The training dataset refers to the data set used to train the anomaly detection model after resampling and other processing. This training dataset includes filtered and organized second-generation production data, designed to provide learning samples for the anomaly detection model, enabling it to learn the characteristics and patterns of normal production data, so as to accurately perform data anomaly self-checks on the first-generation production data subsequently.

[0039] Specifically, based on the calculated data redundancy index, the second production data is resampled. If the data redundancy index is high, it indicates that the second production data is highly consistent with the first production data and may contain a lot of redundant information. In this case, the number of samples drawn from this part of the second production data can be appropriately reduced. If the data redundancy index is low, it indicates that the data patterns of the two are significantly different. To enhance the representativeness of the data, the proportion of samples drawn from this type of second production data needs to be increased. Through this selective extraction, a training dataset suitable for training models related to the first production data is constructed.

[0040] The data redundancy index reflects the value of the second production data for training an anomaly detection model targeting the first production data. Data with high redundancy may not provide new and effective information for the anomaly detection model, and may even interfere with the model's ability to learn the true characteristics of the first production data; while data with low redundancy may contain important patterns not reflected in the first production data. By resampling based on the data redundancy index, the amount of data can be balanced, redundant data can be removed, and the representativeness of the data can be enhanced, making the training dataset more accurately reflect the features and patterns related to the first production data, thereby improving the training effect of the anomaly detection model.

[0041] S400, based on the training dataset, fine-tunes the anomaly detection model so that the fine-tuned anomaly detection model can perform data anomaly self-checks on the first production data.

[0042] In this step, the anomaly detection model refers to a model used to detect anomalies in production data. This anomaly detection model learns the characteristics and patterns of normal production data, establishes corresponding judgment criteria, and when new production data (such as the first production data) is input, it can determine whether it deviates from the normal range, thereby discovering possible anomalies, such as equipment failure or abnormal processing parameters.

[0043] Fine-tuning refers to the process of further adjusting and optimizing an anomaly detection model that has already been initially trained. Based on the constructed training dataset, by adjusting the parameters, structure, or training strategy of the anomaly detection model, the model can be better adapted to the characteristics and needs of the first production data, thereby improving the accuracy and reliability of the anomaly detection model in performing data anomaly self-checks on the first production data.

[0044] Data anomaly self-checking refers to the process of automatically analyzing and judging primary production data using a finely tuned anomaly detection model to detect the presence of abnormal data. Without manual intervention, the anomaly detection model can quickly and accurately identify anomalies in the primary production data based on learned normal data patterns, providing timely and effective information for production process monitoring and optimization.

[0045] Specifically, the pre-trained anomaly detection model is fine-tuned using a pre-constructed training dataset. During fine-tuning, the parameters of the anomaly detection model are adjusted based on the characteristics of the training dataset, such as the weights and thresholds in the neural network; the model structure is optimized, for example, by adding or removing certain layers; and the training strategy is adjusted, such as changing the learning rate and the number of iterations. After a series of adjustments and optimizations, a fine-tuned anomaly detection model is obtained. This model is then used to perform a self-check for data anomalies on the first batch of production data. The anomaly detection model automatically analyzes and determines whether any anomalies exist in the first batch of production data.

[0046] The initially trained anomaly detection model may not fully adapt to the characteristics and requirements of the first production data. The training dataset contains filtered data related to the first production data. By fine-tuning the anomaly detection model based on this dataset, it can learn patterns that better match the characteristics and regularities of the first production data, thereby establishing more accurate judgment criteria. In this way, when the first production data is input, the anomaly detection model can quickly and accurately identify abnormal data that deviates from the normal range based on the learned normal data patterns, achieving effective anomaly self-checking of the first production data and providing timely and reliable information for production process monitoring and optimization.

[0047] This invention optimizes the adaptability of the anomaly detection model to the current production scenario by dynamically quantifying the consistency of patterns between current production data and historical production data, filtering and reconstructing the training dataset, solving the problem of low model update efficiency caused by redundant data in massive datasets, and improving the accuracy and timeliness of anomaly self-detection.

[0048] As an example, on an automotive parts production line, a press generates multi-dimensional sensor data every minute, including tonnage, stroke, and temperature. First, real-time temperature data collected over the past eight hours is used as the first production data, and historical temperature data from the past two years is used as the second production data. This data is then categorized into multiple production data clusters, each representing a processing state.

[0049] Furthermore, by combining the comparison results between the first and second production data and the production data clusters, a data redundancy index is calculated to characterize the consistency of the data patterns between the two.

[0050] Therefore, the second production data is resampled based on the data redundancy index. For highly repetitive processing states, the sampling frequency is reduced; for newly emerging processing states, the sampling frequency is increased. This method constructs a training dataset that retains valuable information from historical data while highlighting the characteristics of new data.

[0051] Finally, the anomaly detection model was fine-tuned using this resampled training dataset. The fine-tuned anomaly detection model was then used to perform real-time anomaly detection on the first production data, enabling it to quickly identify minor anomalies in the production process, such as abnormal increases in mold temperature.

[0052] In this embodiment, production data with the same equipment source and processing attributes, such as first and second production data, are first classified into production data clusters based on current and historical production data. This ensures that each production data cluster includes production data with the same processing status. Next, a data redundancy index is determined by comparing the first and second production data, confirming the consistency of data patterns between them. Based on this data redundancy index, the second production data is resampled to construct a training dataset. This effectively filters out data valuable for learning from new data and avoids interference from duplicate data. Finally, the anomaly detection model is fine-tuned using this training dataset, enabling it to more accurately perform anomaly self-checks on the first production data within the current production data. This overcomes the problem of insufficient learning on new data in existing models, thereby improving the accuracy of production data anomaly self-checks.

[0053] In some of the solutions described above in this invention, when classifying current production data and historical production data, a simple division is insufficient to distinguish data under the same equipment and processing attributes but different processing states. Because changes in processing states are not effectively identified, the same production data cluster may contain data under different processing states, leading to inaccurate calculations of subsequent data redundancy indicators and affecting the quality of training dataset construction.

[0054] In this regard, the present invention further proposes that S100 includes: For each production data in the first and second production data, a target processing vector for the production data is constructed using the processing attributes of the production data as dimensions. For each target processing vector, the degree of change in the processing state of the target processing vector is determined based on the similarity between the target processing vector and the reference processing vector at the previous sampling time. Based on the degree of change in processing status of each target processing vector, the target processing vectors are clustered to obtain several production data clusters.

[0055] In this embodiment, the target processing vector is constructed by extracting processing attribute dimensions. The processing state change degree is determined by calculating the cosine similarity between the target processing vector and the reference processing vector at the previous sampling time, with the value controlled between 0 and 1. Specifically, for each target processing vector, firstly, the target processing vector corresponding to the production data at the previous sampling time is determined as its corresponding reference processing vector; then, the cosine similarity between the target processing vector and the reference processing vector at the previous sampling time is calculated, and this cosine similarity is normalized to obtain a value controlled between 0 and 1; finally, the processing state change degree of the target processing vector is obtained by subtracting the normalized cosine similarity from 1.

[0056] The clustering process employs a density-based clustering algorithm, with a set cluster radius threshold of 0.3 and a minimum sample size of 5. Vectors with a difference in processing state variability less than 0.2 are grouped into the same cluster. Specifically, during density-based clustering, processing state variability is used as the clustering feature. For each target processing vector, the difference in processing state variability between it and surrounding processing vectors (within the set cluster radius) is calculated. When the difference is less than 0.2, these processing vectors are grouped into the same potential cluster. Finally, the production data clusters are determined based on constraints such as the minimum sample size.

[0057] As an example, we first classify the first production data in the current production data and the second production data in the historical production data to obtain several production data clusters. Specifically, for each production data in the first and second production data, we construct a target processing vector for the production data, using the processing attributes of the production data as dimensions. For example, processing attributes can be processing temperature, processing pressure, processing time, etc. Further, for each target processing vector, we determine the degree of change in the processing state of the target processing vector based on the similarity between the target processing vector and the reference processing vector corresponding to its previous sampling time. The similarity can be measured by calculating the cosine similarity between the two vectors. Therefore, based on the degree of change in the processing state of each target processing vector, we cluster the target processing vectors to obtain several production data clusters. Clustering can use common clustering methods such as the K-means algorithm.

[0058] This embodiment constructs a target processing vector and calculates the degree of change in processing state, effectively capturing the dynamic characteristics of production data. Furthermore, clustering based on the degree of change in processing state allows data with similar processing states to be grouped together, thus achieving refined classification of production data. This provides a more accurate data foundation for subsequent data comparison and anomaly detection, helping to improve the accuracy and efficiency of self-inspection of production data anomalies.

[0059] In some of the solutions described above in this invention, when determining data redundancy indicators by comparing current production data with historical production data, if the differences in material properties in different production stages are not considered, it may lead to a deviation in the judgment of data consistency, thereby affecting the accuracy of the calculation of redundancy indicators.

[0060] In this regard, such as Figure 2 As shown, the present invention further proposes that S200 includes the following S210 to S260: S210, Obtain the current production stage corresponding to the current processing vector sequence; the current processing vector sequence is formed by arranging the target processing vectors in the first production data in chronological order. S220, based on the material property differences between the current production stage and the historical production stage, determine the probability of product consistency between the current production stage and the historical production stage; the historical production stage is the production stage corresponding to the historical processing vector sequence, which is formed by arranging the target processing vectors in the second production data in chronological order; S230, based on the probability of product consistency between the current production stage and each historical production stage, screen each historical production stage and determine the reference production stage corresponding to the current production stage. S240, Based on the difference in the degree of change in processing status between adjacent reference production stages and the production data cluster, determine the consistency of production data between adjacent reference production stages; S250, adjacent reference production links whose production data consistency is less than the preset consistency threshold are identified as a group of optimization processes; S260, based on the differences between the production links before and after optimization in each optimization process group, and combined with the overall optimization consistency between the current production links and historical production links, determine the data redundancy index between the first production data and the second production data.

[0061] In this embodiment, when obtaining the current production stage corresponding to the current processing vector sequence, the target processing vectors in the first production data need to be arranged in chronological order to construct the current processing vector sequence; a processing state change sequence is constructed based on the difference in processing state change degree between adjacent target processing vectors, and then the current production stage is divided based on the processing state change sequence.

[0062] When determining the probability of product consistency, the following logic can be used for calculation: First, calculate the difference between the corresponding material attribute values ​​of the current production stage and the historical production stages. Summate these differences to obtain the cumulative material attribute difference value, then divide it by the total number of material attributes to obtain the mean of the attribute differences. Finally, determine the probability of product consistency based on the mean of the attribute differences. It is important to note that the mean of the attribute differences has an inverse relationship with the probability of product consistency; that is, the smaller the mean of the attribute differences, the closer the material attributes of the current production stage are to those of the historical production stages, and the higher the probability of product consistency; conversely, the larger the mean of the attribute differences, the lower the probability of product consistency. Therefore, the probability of product consistency can be simply obtained by taking the reciprocal of the mean of the attribute differences.

[0063] It's important to note that the current method of calculating the reciprocal of the mean of attribute differences in determining the probability of product consistency shares the same underlying logic as potentially more complex calculations used later. Both revolve around the core element of attribute difference degree to measure the probability of product consistency. Essentially, both rely on the magnitude of material attribute differences between production stages to determine product consistency. However, the current approach of using the reciprocal of the mean of attribute differences is more direct and basic, while other potentially more complex calculation methods will be more refined and diversified in their handling of attribute differences, differing in specific calculation steps and forms.

[0064] When selecting reference production stages, historical production stages with minimal differences in material properties from the current production stage should be retained based on the likelihood of product consistency. Specifically, a threshold can be pre-set, and historical production stages with a greater-than-threshold likelihood of product consistency with the current production stage can be identified as reference production stages corresponding to the current production stage. The threshold value can be 0.7, and implementers can set it according to the specific implementation scenario.

[0065] When determining production data consistency, it is necessary to calculate the comprehensive processing status difference between adjacent reference production stages to assess the consistency. When selecting optimization process groups, the following methods are used: A preset consistency threshold is set, and all combinations of adjacent reference production stages are iterated through. The production data consistency between each combination is calculated, and combinations of adjacent reference production stages with a consistency lower than the preset threshold are marked. Each such combination of adjacent reference production stages is identified as an optimization process group. The method for obtaining the pre-optimization and post-optimization production stages is as follows: In the determined optimization process groups, the reference production stages that occur earlier in time are the pre-optimization stages, and those that occur later are the post-optimization stages. The consistency threshold can be 0.3, and implementers can set it according to the specific implementation scenario.

[0066] Specifically, after constructing the current processing vector sequence, production stages are divided by abrupt changes in the processing state change sequence to ensure smooth changes in processing state within each stage. When selecting reference production stages, if the probability of product consistency is higher than a preset threshold, the historical production stage is retained. The determination of the data redundancy index is first achieved by comparing the processing optimization vectors of production stages before and after optimization (e.g., the average processing vector of the stage before optimization is [1.2, 3.5], and after optimization it is [1.0, 3.3], with a difference vector of [0.2, 0.2]), further calculating the comprehensive optimization consistency between the current production stage and the historical production stages, and finally mapping this comprehensive optimization consistency to the data redundancy index.

[0067] As an example, the first step is to obtain the current production stage corresponding to the current processing vector sequence. For instance, the target processing vectors in the first production data can be arranged in chronological order to construct the current processing vector sequence. Then, based on the difference in processing state change between adjacent target processing vectors, a processing state change sequence is constructed. Finally, the current processing vector sequence is divided into multiple current production stages according to the processing state change sequence.

[0068] Next, the differences in material properties between the current production stage and historical production stages are calculated to determine the likelihood of product consistency. For example, the differences in the values ​​of each material property can be calculated, standardized separately, and then summed and divided by the total number of material properties to obtain the average contribution value of the difference in a single standard material property, thereby determining the likelihood of product consistency.

[0069] Then, historical production stages are screened based on the probability of product consistency to determine reference production stages. For adjacent reference production stages, the mean of the difference in processing status variability is calculated to obtain the comprehensive processing status difference. Simultaneously, the proportion of processing vectors located in the same production data cluster within adjacent reference production stages is statistically analyzed. Production data consistency is determined based on these two indicators.

[0070] Then, adjacent reference production data with a consistency of less than a preset threshold are selected. The production process is identified as an optimization process group. For each optimization process group, the difference between the target processing vectors of the production process before and after optimization is calculated to obtain the optimized processing vector. Based on the optimized processing vectors of each optimization process group, the overall optimization consistency between the current production process and historical production processes is determined, thereby determining the data redundancy index.

[0071] This embodiment enables the comparison of first and second production data based on production data clusters to determine data redundancy indicators. By meticulously comparing and filtering current and historical production processes, production data with similar processing characteristics can be accurately identified. Simultaneously, by calculating production data consistency and optimizing the process, data patterns and trends during production can be effectively captured. This method avoids simply treating all historical data equally, instead employing intelligent filtering and weight allocation based on actual production conditions. Therefore, this invention can more accurately assess the correlation between current and historical production data, improving the accuracy and representativeness of data redundancy indicators. This lays a solid foundation for subsequent data resampling and anomaly detection model optimization, contributing to improved performance and reliability of the entire production data anomaly self-inspection system.

[0072] In some of the solutions described above, a method is proposed to determine the likelihood of product consistency based on the differences in material properties between the current production stage and historical production stages, and to select reference production stages to determine the consistency of production data. However, in this process, if the current processing vector sequence cannot accurately reflect the continuous changes in actual production stages, it may lead to a bias in the selection of reference production stages, thereby affecting the accuracy of the calculation of the consistency of production data.

[0073] In this regard, the present invention further proposes that S210 includes: Arrange the target processing vectors in the first production data in chronological order to construct the current processing vector sequence; Based on the absolute value of the difference in the degree of change of processing state between two adjacent target processing vectors in the current processing vector sequence, a processing state change sequence is constructed; Based on the processing state change sequence, the current processing vector sequence is divided into at least one current production stage.

[0074] In this embodiment, the target processing vectors are arranged in chronological order to form the current processing vector sequence, ensuring the temporal continuity of the data. The absolute value of the difference in the degree of change of processing state is calculated by subtracting the degree of change of processing state between two adjacent target processing vectors and taking the absolute value, reflecting the abrupt change points of state transitions during production. The processing state change sequence is constructed by continuously calculating the absolute value of the difference in the degree of change of adjacent processing states, forming a continuous signal characterizing state changes. The production process is divided into stable intervals based on the distribution location of the abrupt change points.

[0075] Specifically, the target processing vectors are arranged in ascending order by timestamp, forming a current processing vector sequence with temporal correlation. For each pair of adjacent vectors in the current processing vector sequence, the absolute value of the difference in their processing state changes is calculated, generating a processing state change sequence containing the absolute values ​​of all adjacent differences. By setting a mutation threshold, the locations of mutation points in the processing state change sequence where the absolute value of the difference exceeds the threshold are identified. The vector intervals between adjacent mutation points are divided into independent production stages, ensuring that the processing state changes within each stage are stable. For example, when the absolute value of the difference in the processing state changes between adjacent vectors exceeds a preset threshold of 0.15, that location is determined to be a production stage segmentation point. Thus, the current processing vector sequence is divided into multiple production stages with stable processing states, providing accurate stage boundary conditions for subsequent selection of reference production stages, avoiding misjudgments of data consistency due to stage segmentation deviations, and improving the reliability of data redundancy index calculations.

[0076] As an example, when obtaining the current production stage corresponding to the current processing vector sequence, the target processing vectors in the first production data are first arranged in chronological order to construct the current processing vector sequence. For example, assuming the first production data contains 10 target processing vectors, denoted as V1, V2, ..., V10, the current processing vector sequence obtained after chronological order is [V1, V2, ..., V10].

[0077] Next, a processing state change sequence is constructed based on the absolute difference of the processing state change degree between two adjacent target processing vectors in the current processing vector sequence. Specifically, the absolute difference of the processing state change degree between two adjacent target processing vectors is calculated to obtain a processing state change sequence of length 9. For example, assume that the absolute difference of the processing state change degree between two adjacent target processing vectors is [0.1, 0.2, 0.5, 0.1, 0.3, 0.4, 0.2, 0.1, 0.3].

[0078] Then, based on the processing state change sequence, the current processing vector sequence is divided into at least one current production stage. Specifically, a threshold can be set, for example, 0.4. When the absolute value of the difference in the degree of processing state change is greater than this threshold, it is considered the start of a new production stage. According to the above example, the current processing vector sequence can be divided into 3 current production stages: [V1,V2,V3], [V4,V5,V6], and [V7,V8,V9,V10].

[0079] This embodiment enables accurate identification and segmentation of the current production stage, providing a reliable foundation for subsequent data comparison and anomaly detection. This improves the accuracy and efficiency of production data anomaly self-checking, helps to promptly detect and handle abnormal situations in the production process, and ensures production quality and efficiency.

[0080] In some of the solutions described above in this invention, when determining the likelihood of product consistency by the difference in material properties between the current production stage and the historical production stage, the method of measuring the difference in material properties is rather vague, making it difficult to accurately quantify the degree of influence of different material properties on product consistency. This results in insufficient accuracy in selecting reference production stages and affects the calculation accuracy of subsequent data redundancy indicators.

[0081] In this regard, the present invention further proposes that S220 includes: The current material attribute values ​​in the current production process are compared with the corresponding historical material attribute values ​​in the historical production process, and then standardized to obtain several standard material attribute difference values. The sum of several standard material property differences is divided by the total number of material properties to obtain the average contribution value of a single standard material property difference. Based on the average contribution value of the difference in individual standard material properties, the likelihood of product consistency between the current production stage and the historical production stage is determined.

[0082] In this embodiment, the standard material property difference is calculated by standardizing the material property values ​​of the same type in the current production stage and historical production stages, such as temperature, pressure, or raw material composition ratio. The total number of material properties refers to the total number of material property types participating in the comparison in the current production stage. The average contribution value is obtained by summing the absolute values ​​of all standard material property differences and dividing by the total number, reflecting the average impact of individual attribute differences on overall consistency. When determining the probability of product consistency based on the average contribution value, the smaller the contribution value, the smaller the impact of material property differences on product consistency, and the higher the probability of product consistency.

[0083] Specifically, when calculating the probability of product consistency, the process first extracts material attribute values ​​of the same type from the current and historical production stages. For example, if the temperature value in the current stage is T1 and the temperature value in a historical stage is T2, the temperature attribute difference is |T1-T2|, which is then standardized. This process is repeated for all material attributes to obtain the difference between each standard material attribute. The absolute values ​​of all standard material attribute differences are summed and divided by the number of material attribute types to obtain the average contribution value. The smaller this average contribution value, the smaller the difference in material attributes between the current and historical production stages, and the higher the probability of product consistency. This method quantifies material attribute differences into specific numerical values, avoiding errors caused by subjective judgment, and thus more accurately selecting historical reference stages that match the current production stage, improving the reliability of data redundancy indicators.

[0084] The probability of product consistency can be determined using the following formula 1: Formula 1 In formula 1, Used to characterize the probability that the product from the a-th production stage is identical to that from the k-th production stage. Used to characterize the nth material property value in the a-th production stage. The parameter `norm` is used to characterize the value of the nth material property in the kth production stage, and `norm` is used to characterize the standardization process. Used to characterize the total number of material properties Used to characterize exponential function operations; The proportion of the same processing state time between the a-th production stage and the k-th production stage can be specifically calculated by dividing the amount of data in the same production data cluster of the a-th and k-th production stages by the total amount of data in the a-th and k-th production stages.

[0085] As an example, we first calculate the differences between the current material property values ​​in the current production stage and the corresponding historical material property values ​​in previous production stages, and then standardize these differences to obtain several standard material property difference values. For instance, for a certain production stage, the current material's hardness, density, and thermal conductivity are 10, 5, and 2, respectively, while the corresponding material property values ​​in previous production stages are 9, 6, and 3. The resulting material property difference values ​​are 1, -1, and -1. These differences are then standardized to obtain several standard material property difference values.

[0086] Next, the sum of several material property differences is divided by the total number of material properties to obtain the average contribution value of a single material property difference. Finally, based on the average contribution value of a single material property difference, the probability of product consistency between the current production stage and the historical production stage is determined using Formula 1 above.

[0087] This embodiment enables accurate assessment of product consistency between the current and historical production stages based on differences in material properties. This effectively filters out historical production data that matches the current stage, improving the accuracy of subsequent data analysis and anomaly detection. Furthermore, by introducing the average contribution value of individual material property differences, the influence of multiple material properties can be balanced, avoiding misjudgments caused by excessive differences in a single attribute, thereby improving the reliability and stability of product consistency assessment.

[0088] In some of the solutions described above in this invention, when determining the consistency of production data between adjacent reference production stages, if the evaluation is based solely on the difference in the degree of change in processing status, the calculated consistency of production data may deviate from the actual situation. This single-dimensional evaluation method reduces the accuracy of production data consistency, thereby affecting the reliability of subsequent data redundancy indicators, leading to inaccurate construction of the training dataset, and ultimately affecting the fine-tuning effect of the anomaly detection model.

[0089] In this regard, the present invention further proposes that S240 includes: The average value of the difference in the degree of change in the corresponding processing state between adjacent reference production stages is used to obtain the comprehensive degree of difference in the processing state between adjacent reference production stages. Obtain the number of first processing vectors located in the same production data cluster between adjacent reference production stages; Divide the number of the first processing vector by the total number of processing vectors in the adjacent reference production stage to obtain the proportion of the number of processing vectors. Based on the overall processing status difference and the proportion of processing vectors, the consistency of production data between adjacent reference production stages is determined.

[0090] In this embodiment, the overall processing state difference is obtained by calculating the average of the differences in processing state changes between all corresponding adjacent reference production stages. For example, if adjacent stages contain three processing state change differences of 0.3, 0.5, and 0.4 respectively, the overall difference is 0.4. The proportion of processing vectors is calculated by counting the number of processing vectors that exist simultaneously in two reference production stages within the same production data cluster, and then calculating their proportion to the total number of processing vectors in both stages. For example, if there are 100 processing vectors in adjacent stages, and 60 of them are located in the same production data cluster, then the proportion is 60%. The production data consistency is calculated by weighting the overall processing state difference with the proportion of processing vectors.

[0091] Specifically, the mean value of the difference in processing status changes between adjacent reference production stages is first processed to eliminate interference from single-point fluctuations, obtaining a comprehensive index characterizing the overall processing status differences. Further, the proportion of processing vectors belonging to the same production data cluster in both stages is statistically analyzed to reflect the degree of overlap in data distribution. The comprehensive processing status difference is combined with the proportion of processing vectors; for example, the lower the comprehensive processing status difference and the higher the proportion of processing vectors, the higher the consistency of production data. By fusing the two indices through weighted or multiplicative methods, both dynamic changes in processing status and static characteristics of data distribution can be considered simultaneously, avoiding the limitations of a single index. This method can more comprehensively evaluate the consistency of production data between adjacent reference production stages, providing a reliable basis for subsequent optimization of process group division and calculation of data redundancy indices, thereby improving the accuracy of training dataset construction.

[0092] The consistency of production data between adjacent reference production stages can be determined using the following formula 2: Formula 2 In formula 2, Used to characterize the consistency of production data between the a-th and k-th production stages. Used to characterize the overall processing status difference between the a-th and k-th production stages. This is used to characterize the number of first processing vectors that are located in the same production data cluster between the a-th production stage and the k-th production stage. Used to characterize the number of processing vectors in the a-th production stage. Used to characterize the number of processing vectors in the k-th production stage. Used to characterize positive correlation normalization.

[0093] As an example, we first average the differences in the degree of change in processing status between adjacent reference production stages to obtain the overall degree of difference in processing status between adjacent reference production stages. For example, by calculating the average of the differences in the degree of change in processing status between adjacent reference production stages A and B, we obtain an overall degree of difference in processing status of 0.35.

[0094] Next, obtain the number of first processing vectors belonging to the same production data cluster between adjacent reference production stages. Specifically, count the number of processing vectors in reference production stages A and B that simultaneously belong to production data cluster X, assuming it is 50.

[0095] The number of the first processing vector is then divided by the total number of processing vectors in the adjacent reference production stages to obtain the percentage of processing vectors. Further, if the total number of processing vectors in reference production stages A and B is 200, then the percentage of processing vectors is 50 / 200 = 25%.

[0096] Finally, based on the overall processing status difference and the proportion of processing vectors, the consistency of production data between adjacent reference production stages is determined. Therefore, the overall processing status difference of 0.35 and the proportion of processing vectors of 25% can be substituted into Formula 2 above to obtain the consistency of production data between reference production stage A and reference production stage B.

[0097] This embodiment comprehensively considers the changes in processing status and data distribution characteristics between adjacent reference production stages, accurately quantifying the consistency of production data between them. This method avoids the one-sidedness that may result from relying on a single indicator, improving the comprehensiveness and accuracy of production data consistency assessment. Furthermore, by introducing the concept of production data clusters, this scheme can better capture data characteristics under similar processing states, thereby enhancing the relevance and reliability of consistency assessment.

[0098] In some of the above-mentioned solutions of the present invention, the process group is selected and optimized based on the consistency of production data between adjacent reference production links. However, it is difficult to effectively measure the impact of the differences between production links before and after optimization on data redundancy in the optimization process group, resulting in insufficient accuracy of data redundancy indicators and affecting the subsequent resampling and model fine-tuning effects.

[0099] In this regard, the present invention further proposes that S260 includes: Based on the difference between the corresponding target processing vectors between the production steps before and after optimization in each optimization process group, the processing optimization vector of each optimization process group is determined. Based on the processing optimization vectors of each optimization process group, the overall optimization consistency between the current production process and the historical production process is determined; Based on the overall optimization consistency between the current production process and the historical production process, a data redundancy index is determined between the first production data and the second production data.

[0100] In this embodiment, the processing optimization vector of the optimization process group can be determined by the following formula 3: Formula 3 In formula 3, The processing optimization vector used to characterize the i-th optimization process group Used to characterize the t-th target processing vector of the production stage before optimization in the i-th optimization process group. The term T is used to characterize the t-th target processing vector in the optimized production process group of the i-th optimization process group, and T is used to characterize the number of target processing vectors in the production process.

[0101] The degree of consistency between the current production process and the historical production process can be determined using the following formula 4: Formula 4 In formula 4, Used to characterize the overall optimization consistency between the current production step and historical production steps. The cosine similarity is used to characterize the processing optimization vector between the u-th current production stage and the most recent reference production stage, and the processing optimization vector of the i-th optimization process group. The cosine similarity between the processing optimization vector of the i-th optimization process group and the mean of the processing optimization vectors of all optimization process groups is used, and I is used to represent the number of optimization process groups. The processing optimization vector between the current production stage and the most recent reference production stage can be determined using the same calculation method as in Formula 3 above. The current production stage is the optimized production stage in Formula 3 above, and the most recent reference production stage is the unoptimized production stage in Formula 3 above.

[0102] The data redundancy index can be determined using the following formula 5: Formula 5 In formula 5, This indicator is used to characterize the data redundancy between the current production stage (u-th) and historical production stages. Used to characterize the consistency of production data between the current production stage (u-th stage) and the most recent reference production stage. The mean value used to characterize the consistency of production data between the current production stage (u-th stage) and each reference production stage. Used to characterize the overall optimization consistency between the current production stage (u-th) and historical production stages.

[0103] As an example, the processing optimization vector for each optimization process group is first determined based on the difference between the corresponding target processing vectors between the production stages before and after optimization. Specifically, the difference between the corresponding target processing vectors between the production stages before and after optimization is substituted into Formula 3 above to obtain the processing optimization vector for the optimization process group.

[0104] Furthermore, based on the processing optimization vectors of each optimization process group, the overall optimization consistency between the current production stage and the historical production stage is determined. For example, by substituting the processing optimization vectors of each optimization process group into Formula 4 above, the overall optimization consistency between the current production stage and the historical production stage can be obtained.

[0105] Therefore, based on the overall optimization consistency between the current production stage and the historical production stage, a data redundancy index is determined between the first production data and the second production data. Specifically, the data redundancy index is calculated using Formula 5 above. Where there is only one current production stage, Formula 5 calculates the data redundancy index between the first and second production data; where there are multiple current production stages, Formula 5 calculates the sub-redundancy index between the current and historical production stages. The combination of these sub-redundancy indices yields the data redundancy index between the first and second production data.

[0106] This embodiment enables accurate quantification of data redundancy between the first and second production data based on changes in production stages within an optimized process group. By introducing the concepts of processing optimization vectors and overall optimization consistency, the changes in production stages before and after optimization can be comprehensively considered, avoiding the bias caused by a single indicator. Simultaneously, the use of vector calculation and cosine similarity methods effectively captures the multidimensional features of production data, improving the accuracy and reliability of data redundancy indicators. This method more accurately reflects the changing patterns of production data, providing a reliable basis for subsequent data resampling and fine-tuning of the anomaly detection model, thereby improving the performance and efficiency of the entire digital factory production data anomaly self-inspection system.

[0107] In some of the solutions described above in this invention, when resampling historical production data based on data redundancy indicators, if a globally uniform resampling strategy is adopted, it is impossible to distinguish the differences in data patterns across different production stages. The redundancy levels of historical production data and current production data differ across different production stages. Uniform sampling would result in an excessively high proportion of historical data with high redundancy in the training set, while the proportion of historical data with low redundancy would be insufficient, affecting the model's fine-tuning effect.

[0108] To address this, the present invention further proposes a data redundancy index including sub-redundancy indices of the first production data in each current production stage; S300 includes: The resampling frequency of each reference production stage is determined by using the sub-redundancy index of each current production stage and the original sampling frequency of the reference production stage corresponding to the current production stage in the second production data. Based on the resampling frequency of each reference production stage, the second production data of each reference production stage are resampled to construct the training dataset of the first production data.

[0109] In this embodiment, the sub-redundancy index is obtained by calculating the overall optimization consistency between the current production process and historical production processes, reflecting the degree of matching between their data patterns. The original sampling frequency is the actual sampling frequency of the reference production process in historical data, used to measure the distribution characteristics of historical data. The resampling frequency is generated by inversely proportionalizing the sub-redundancy index and the original sampling frequency; for example, the sub-redundancy index can be used as a divisor to adjust the original sampling frequency. For reference production processes with high sub-redundancy indices, it indicates that their data patterns are highly consistent with the current data patterns, containing a lot of redundant information, and their resampling frequency should be reduced to reduce redundant data interference. For reference production processes with low sub-redundancy indices, it indicates that their data patterns differ significantly, containing more valuable learning information, and their resampling frequency should be maintained or appropriately increased to enhance the model's learning of new patterns.

[0110] Specifically, when determining the resampling frequency, the sub-redundancy index of the current production stage is first extracted. This sub-redundancy index is calculated by the difference in processing optimization vectors between the production stages before and after optimization in the optimization process group. Simultaneously, the original sampling frequency of the reference production stage is obtained. This original sampling frequency is determined by the equipment parameters or process requirements at the time of historical data acquisition. The sub-redundancy index is used as an adjustment factor and inversely proportional to the original sampling frequency to obtain the adjusted resampling frequency. For example, if the sub-redundancy index of a reference production stage is 1.2 and the original sampling frequency is 10Hz, the resampling frequency can be adjusted to 8.3Hz (e.g., 10Hz / 1.2). Subsequently, the second production data of the reference production stage is resampled according to the adjusted frequency, for example, using oversampling or undersampling methods, so that the data distribution of each production stage in the training dataset dynamically matches the redundancy index of the current production stage. In this way, the training dataset can more accurately reflect the data patterns of the current production stage, thereby improving the efficiency and accuracy of fine-tuning the anomaly detection model.

[0111] The resampling frequency can be determined using the following formula 6: Formula 6 In formula 6, The resampling frequency is used to characterize the reference production stage corresponding to the u-th current production stage. This indicator is used to characterize the data redundancy between the current production stage (u-th) and historical production stages. The original sampling frequency used to characterize the reference production stage corresponding to the u-th current production stage.

[0112] As an example, in a semiconductor wafer fabrication scenario, current production data is divided into three consecutive production stages. The sub-redundancy indices for each stage are calculated to be 0.5, 1.0, and 2.0, respectively. The original sampling frequency for the corresponding reference production stages in historical data is 20Hz. According to Formula 6, the resampling frequencies are adjusted to 40Hz, 20Hz, and 10Hz, respectively. Linear interpolation is used to downsample the historical production data. When the resampling frequency is higher than the original frequency, supplementary data points are generated using a time series prediction model. After aligning the resampled historical data with the current production data by timestamp, a training dataset containing 12,000 samples is constructed.

[0113] This embodiment effectively solves the problem of low model learning efficiency caused by the inconsistency between historical and current data distribution. By dynamically adjusting the sampling density of historical data, it not only preserves the effective features of historical data but also strengthens the data distribution characteristics that match current production patterns. This allows the training dataset to accurately reflect the actual state of the current production process, thereby improving the sensitivity of the anomaly detection model in identifying new types of production data.

[0114] This invention provides a method for self-checking anomalies in production data of a digital factory. Accordingly, this invention also provides a specific embodiment of a self-checking system for anomalies in production data of a digital factory.

[0115] like Figure 3 As shown, a schematic diagram of a self-inspection system for abnormal production data in a digital factory is provided. The self-inspection system 300 for abnormal production data in a digital factory includes a data classification module 310, a data comparison module 320, a data resampling module 330, and a model training module 340.

[0116] The data classification module 310 is used to classify the first production data in the current production data and the second production data in the historical production data to obtain several production data clusters; the first production data is the production data in the current production data that has the same equipment source and the same processing attributes, and the second production data is the production data in the historical production data that has the same equipment source and the same processing attributes as the first production data. The production data cluster includes the first production data and the second production data that are in the same processing state. The data comparison module 320 is used to compare the first production data with the second production data based on the production data cluster, and determine the data redundancy index between the first production data and the second production data; the data redundancy index is used to characterize the consistency of data patterns between the first production data and the second production data. The data resampling module 330 is used to resample the second production data based on the data redundancy index to construct a training dataset for the first production data. The model training module 340 is used to fine-tune the anomaly detection model based on the training dataset, so that the fine-tuned anomaly detection model can perform data anomaly self-checking on the first production data.

[0117] In the digital factory production data anomaly self-inspection system provided in this invention embodiment, firstly, firstly, production data and secondly, which share the same equipment source and processing attributes in the current and historical production data, are classified into production data clusters to ensure that each cluster includes production data with the same processing state. Next, a data redundancy index is determined by comparing the first and second production data to confirm the consistency of data patterns between them. Based on this data redundancy index, a training dataset is constructed by resampling the second production data, effectively filtering out data valuable for learning from new data and avoiding interference from duplicate data. Finally, the anomaly detection model is fine-tuned using this training dataset, enabling the model to more accurately perform anomaly self-inspection on the first production data within the current production data, overcoming the problem of insufficient learning from new data in existing systems, thereby improving the accuracy of production data anomaly self-inspection.

[0118] Furthermore, in conjunction with the digital factory production data anomaly self-checking method in the above embodiments, this invention can be implemented using a computer storage medium. This computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the digital factory production data anomaly self-checking methods in the above embodiments.

[0119] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0120] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A digital plant production data anomaly self-checking method, characterized by, The method comprises: classifying first production data in current production data and second production data in historical production data to obtain a plurality of production data clusters; the first production data is production data with the same equipment source and the same processing attribute in the current production data, the second production data is production data with the same equipment source and the same processing attribute as the first production data in the historical production data, and the production data cluster includes the first production data and the second production data in the same processing state; comparing the first production data and the second production data based on the production data cluster to determine a data redundancy index between the first production data and the second production data; the data redundancy index is used to represent the consistency of data rules between the first production data and the second production data; based on the data redundancy index, resampling the second production data to construct a training data set of the first production data; based on the training data set, fine-tuning the anomaly detection model to enable the anomaly detection model after fine-tuning to perform data anomaly self-checking on the first production data.

2. The digital plant production data anomaly self-checking method according to claim 1, characterized in that, The classification of the first production data in the current production data and the second production data in the historical production data to obtain a plurality of production data clusters comprises: for each production data in the first production data and the second production data, constructing a target processing vector of the production data with the processing attribute of the production data as the dimension; for each target processing vector, determining the processing state change degree of the target processing vector based on the similarity between the target processing vector and the reference processing vector at the previous sampling time; based on the processing state change degree of each target processing vector, clustering each target processing vector to obtain a plurality of production data clusters.

3. The digital plant production data anomaly self-checking method of claim 1, wherein, The comparison of the first production data and the second production data based on the production data cluster to determine the data redundancy index between the first production data and the second production data comprises: obtaining a current production link corresponding to a current processing vector sequence; the current processing vector sequence is formed by arranging each target processing vector in the first production data in time sequence; based on the material attribute difference between the current production link and the historical production link, determining the product consistency possibility between the current production link and the historical production link; the historical production link is a production link corresponding to a historical processing vector sequence, and the historical processing vector sequence is formed by arranging each target processing vector in the second production data in time sequence; based on the product consistency possibility between the current production link and each historical production link, screening each historical production link to determine a reference production link corresponding to the current production link; based on the difference of the processing state change degree between adjacent reference production links and the production data cluster, determining the production data consistency degree between adjacent reference production links; The adjacent reference production links with the production data consistency less than the preset consistency threshold are determined as a group of optimization process groups; Based on the difference between the pre-optimization production link and the post-optimization production link in each optimization process group, a data redundancy index between the first production data and the second production data is determined.

4. The digital plant production data anomaly self-checking method according to claim 3, characterized in that, The current production link corresponding to the current machining vector sequence is obtained, including: Arranging each target machining vector in the first production data in time sequence to obtain a current machining vector sequence; Based on the absolute value of the difference between the machining state change degrees of two adjacent target machining vectors in the current machining vector sequence, a machining state change sequence is constructed; Based on the machining state change sequence, the current machining vector sequence is divided into at least one current production link.

5. The digital plant production data anomaly self-checking method of claim 3, wherein, Based on the material attribute difference between the current production link and the historical production link, the product consistency possibility between the current production link and the historical production link is determined, including: The current material attribute values in the current production link are subtracted from the corresponding historical material attribute values in the historical production link to obtain a plurality of standard material attribute difference values; The cumulative value of the plurality of standard material attribute difference values is divided by the total number of material attributes to obtain an average contribution value of a single standard material attribute difference value; Based on the average contribution value of the single standard material attribute difference value, the product consistency possibility between the current production link and the historical production link is determined.

6. The digital plant production data anomaly self-checking method of claim 3, wherein, Based on the difference between the machining state change degrees of the adjacent reference production links and the production data cluster, the production data consistency between the adjacent reference production links is determined, including: The difference between the machining state change degrees corresponding to the adjacent reference production links is processed by mean value to obtain a comprehensive machining state difference degree between the adjacent reference production links; The number of first machining vectors between the adjacent reference production links in the same production data cluster is obtained; The number of machining vectors is divided by the total number of machining vectors in the adjacent reference production links to obtain a machining vector quantity proportion; Based on the comprehensive machining state difference degree and the machining vector quantity proportion, the production data consistency between the adjacent reference production links is determined.

7. The digital plant production data anomaly self-checking method of claim 3, wherein, Based on the difference between the pre-optimization production link and the post-optimization production link in each optimization process group, a data redundancy index between the first production data and the second production data is determined, including: Based on the difference between the target machining vectors corresponding to the pre-optimization production link and the post-optimization production link in each optimization process group, a machining optimization vector of each optimization process group is determined; Based on the machining optimization vectors of each optimization process group, a comprehensive optimization consistency degree of the current production link and the historical production link is determined; Based on the comprehensive optimization consistency degree of the current production link and the historical production link, a data redundancy index between the first production data and the second production data is determined.

8. The digital plant production data anomaly self-checking method of claim 1, wherein, The data redundancy index comprises a sub-redundancy index of the first production data at each current production link; The resampling of the second production data based on the data redundancy index comprises: determining a resampling frequency of each reference production link based on the sub-redundancy index of each current production link and an original sampling frequency of a reference production link corresponding to the current production link in the second production data; and resampling the second production data of each reference production link based on the resampling frequency of each reference production link to construct the training data set of the first production data.

9. A digital factory production data anomaly self-checking system characterized by, The system comprises: a data classification module configured to classify first production data in current production data and second production data in historical production data to obtain a plurality of production data clusters; the first production data is production data with the same equipment source and the same processing attribute in the current production data, the second production data is production data with the same equipment source and the same processing attribute as the first production data in the historical production data, and the production data clusters comprise the first production data and the second production data in the same processing state; a data comparison module configured to compare the first production data and the second production data based on the production data clusters to determine a data redundancy index between the first production data and the second production data; the data redundancy index is used to represent the consistency of data rules between the first production data and the second production data; a data resampling module configured to resample the second production data based on the data redundancy index to construct a training data set of the first production data; a model training module configured to fine-tune an anomaly detection model based on the training data set to enable the anomaly detection model after fine-tuning to perform data anomaly self-checking on the first production data.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer program instructions, and the computer program instructions are executed by the processor to implement the digital factory production data anomaly self-checking method of any one of claims 1-8.

Citation Information

Patent Citations

  • Intelligent factory data optimization acquisition method based on digital twinning

    CN116578890A

  • Fault diagnosis method and system of turboset, computer storage medium and equipment

    CN117574322A

  • Abnormal electricity utilization detection method and device based on semi-fixed integrated data and medium

    CN119179990A

  • A data quality monitoring method based on time series large model

    CN119782965A

  • Environmental monitoring data anomaly detection method, medium and system

    CN119807728A