A digital factory production data anomaly self-checking method, system and medium

By classifying and resampling production data in a digital factory, constructing a training dataset, and fine-tuning the model, the problem of insufficient learning from new data is solved, and the accuracy and efficiency of production data anomaly self-detection are improved.

CN121211284BActive Publication Date: 2026-02-27CHANGCHUN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511746412.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-02-27
Estimated Expiration
2045-11-26

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly update self-inspection thresholds based on new production data in digital factories, resulting in poor accuracy in self-inspection of production data anomalies. This is especially true in environments with massive amounts of data, where the learning efficiency of new data is low and it is difficult to identify subtle deviations.

Method used

By classifying current production data and historical production data to form production data clusters, and resampling the second production data based on data redundancy indicators, a training dataset is constructed, and the anomaly detection model is fine-tuned to improve self-inspection accuracy.

Benefits of technology

It effectively filters out information valuable for learning from new data, avoids interference from duplicate data, improves the accuracy and timeliness of self-checking for anomalies in production data, and adapts to the characteristics of current production data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121211284B_ABST
    Figure CN121211284B_ABST
Patent Text Reader

Abstract

The application discloses a kind of digital chemical plant production data anomaly self-checking method, system and medium, it is related to data processing technical field.The method comprises: the first production data in current production data and the second production data in historical production data are classified, obtain several production data clusters;Based on production data cluster, the first production data is compared with the second production data, determine the data redundancy index between the first production data and the second production data;Data redundancy index is used to characterize the data rule consistency between the first production data and the second production data;Based on data redundancy index, the second production data is resampled, and the training data set of the first production data is constructed;Based on training data set, fine-tuning is carried out to abnormal detection model, so that the first production data is carried out data anomaly self-checking based on fine-tuned abnormal detection model.The application can improve the accuracy of production data anomaly self-checking.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a digital factory production data anomaly self-checking method, system and medium. BACKGROUND

[0002] As an important development direction of modern manufacturing industry, the digital factory unifies the management of enterprise factory production related information by means of digital technologies such as Internet of Things and big data, and realizes efficient, automatic and adaptive advanced manufacturing production. The digital factory can effectively improve production efficiency, but it highly depends on production data, and the accuracy and reliability of production data are crucial to guarantee production efficiency, product quality and operating cost, so it is necessary to monitor production data in time to avoid negative effects of abnormal data.

[0003] The prior art mainly determines the normal range of production data by presetting data logic to predict and analyze the production data of the digital factory, and detects the abnormality of production data when the production data exceeds the normal range, so as to realize the anomaly self-checking of production data.

[0004] However, the digital factory covers the whole production process of products, and generates a large amount of production data every moment, and the data generation speed is extremely high. The existing method of determining the normal range of production data mainly adjusts adaptively by learning production data, a large amount of data is repeated with the past learning data, and the newly appeared production data different from the past is offset by a large amount of learned data, which leads to low learning efficiency of new data, and it is difficult to quickly update the self-checking threshold according to new production data, resulting in poor accuracy of production data anomaly self-checking. SUMMARY

[0005] The embodiments of the present application provide a digital factory production data anomaly self-checking method, system and medium, which can improve the accuracy of production data anomaly self-checking.

[0006] In a first aspect, the present application provides a digital factory production data anomaly self-checking method, comprising: classifying first production data in current production data and second production data in historical production data to obtain a plurality of production data clusters; the first production data is production data with the same equipment source and the same processing attribute in the current production data, the second production data is production data with the same equipment source and the same processing attribute as the first production data in the historical production data, and the production data cluster includes the first production data and the second production data in the same processing state; comparing the first production data and the second production data based on the production data cluster to determine a data redundancy index between the first production data and the second production data; the data redundancy index is used to represent the consistency of the data rule between the first production data and the second production data; resampling the second production data based on the data redundancy index to construct a training data set of the first production data; and fine-tuning an anomaly detection model based on the training data set to enable the anomaly detection model after fine-tuning to perform data anomaly self-checking on the first production data.

[0007] Further, the present application also provides that the first production data in the current production data and the second production data in the historical production data are classified to obtain a plurality of production data clusters, comprising: for each production data in the first production data and the second production data, taking the processing attribute of the production data as the dimension to construct a target processing vector of the production data; for each target processing vector, determining the processing state change degree of the target processing vector based on the similarity between the target processing vector and the reference processing vector at the previous sampling time; and clustering each target processing vector based on the processing state change degree of each target processing vector to obtain a plurality of production data clusters.

[0008] Further, the application also proposes that, based on the production data cluster, the first production data is compared with the second production data to determine the data redundancy index between the first production data and the second production data, including: obtaining a current production link corresponding to a current machining vector sequence; the current machining vector sequence is formed by arranging each target machining vector in the first production data in time sequence; based on the material attribute difference between the current production link and the historical production link, the product consistency possibility between the current production link and the historical production link is determined; the historical production link is the production link corresponding to a historical machining vector sequence, and the historical machining vector sequence is formed by arranging each target machining vector in the second production data in time sequence; based on the product consistency possibility between the current production link and each historical production link, each historical production link is screened to determine the reference production link corresponding to the current production link; based on the difference value of the machining state change degree between adjacent reference production links and the production data cluster, the production data consistency degree between adjacent reference production links is determined; the adjacent reference production links with the production data consistency degree less than the preset consistency degree threshold are determined as a group of optimization process groups; based on the difference between the optimization before production link and the optimization after production link in each optimization process group, the data redundancy index between the first production data and the second production data is determined.

[0009] Further, the application also proposes that the current production link corresponding to the current machining vector sequence is obtained, including: arranging each target machining vector in the first production data in time sequence to construct the current machining vector sequence; based on the absolute value of the difference value of the machining state change degree between adjacent two target machining vectors in the current machining vector sequence, a machining state change sequence is constructed; based on the machining state change sequence, the current machining vector sequence is divided into at least one current production link.

[0010] Further, the application also proposes that, based on the material attribute difference between the current production link and the historical production link, the product consistency possibility between the current production link and the historical production link is determined, including: the current material attribute values in the current production link are subtracted from the corresponding historical material attribute values in the historical production link and then standardized to obtain a plurality of standard material attribute difference values; the cumulative value of the plurality of standard material attribute difference values is divided by the total number of material attributes to obtain the average contribution value of a single standard material attribute difference value; based on the average contribution value of a single standard material attribute difference value, the product consistency possibility between the current production link and the historical production link is determined.

[0011] Further, the application also proposes that, based on the difference value of the processing state change degree between adjacent reference production links and the production data cluster, the production data consistency degree between adjacent reference production links is determined, including: the difference value of the corresponding processing state change degree between adjacent reference production links is processed by mean value to obtain the comprehensive processing state difference degree between adjacent reference production links; the number of first processing vectors in the same production data cluster between adjacent reference production links is obtained; the number of first processing vectors is divided by the total number of processing vectors in adjacent reference production links to obtain the processing vector quantity proportion; based on the comprehensive processing state difference degree and the processing vector quantity proportion, the production data consistency degree between adjacent reference production links is determined.

[0012] Further, the application also proposes that, based on the difference between the production data before optimization and the production data after optimization in each optimization process group, the data redundancy index between the first production data and the second production data is determined, including: based on the difference between the corresponding target processing vectors between the production data before optimization and the production data after optimization in each optimization process group, the processing optimization vector of each optimization process group is determined; based on the processing optimization vector of each optimization process group, the comprehensive optimization consistency degree of the current production link and the historical production link is determined; based on the comprehensive optimization consistency degree of the current production link and the historical production link, the data redundancy index between the first production data and the second production data is determined.

[0013] Further, the application also proposes that the data redundancy index includes the sub-redundancy index of the first production data in each current production link; based on the data redundancy index, the second production data is resampled to construct the training data set of the first production data, including: the resampling frequency of each reference production link is determined by using the sub-redundancy index of each current production link and the original sampling frequency of the reference production link corresponding to the current production link in the second production data; based on the resampling frequency of each reference production link, the second production data of each reference production link is resampled to construct the training data set of the first production data.

[0014] In a second aspect, the embodiment of the present application provides a digital factory production data anomaly self-checking system, comprising: a data classification module, configured to classify first production data in current production data and second production data in historical production data to obtain a plurality of production data clusters; the first production data is production data with the same equipment source and the same processing attribute in the current production data, the second production data is production data with the same equipment source and the same processing attribute as the first production data in the historical production data, and the production data cluster includes the first production data and the second production data in the same processing state; a data comparison module, configured to compare the first production data and the second production data based on the production data cluster to determine a data redundancy index between the first production data and the second production data; the data redundancy index is used to represent the consistency of the data rule between the first production data and the second production data; a data resampling module, configured to resample the second production data based on the data redundancy index to construct a training data set of the first production data; and a model training module, configured to fine-tune an anomaly detection model based on the training data set to enable the anomaly detection model after fine-tuning to perform data anomaly self-checking on the first production data.

[0015] In a third aspect, the embodiment of the present application provides a computer readable storage medium, and the computer readable storage medium stores computer program instructions, and the computer program instructions are executed by a processor to implement the digital factory production data anomaly self-checking method.

[0016] The present application has the following advantages:

[0017] In the digital factory production data anomaly self-checking method provided by the embodiment of the present application, the first production data and the second production data with the same equipment source and the same processing attribute in the current production data and the historical production data are classified into production data clusters, so that the production data clusters cover production data in the same processing state. Then, the data redundancy index is determined by comparing the first production data and the second production data, and the consistency of the data rule between the first production data and the second production data is determined. The training data set is constructed by resampling the second production data based on the data redundancy index, which can effectively filter out valuable data for learning new data and avoid repeated data interference. Finally, the anomaly detection model is fine-tuned using the training data set, so that the anomaly detection model can more accurately perform anomaly self-checking on the first production data in the current production data, overcoming the problem of insufficient learning of new data, thereby improving the accuracy of production data anomaly self-checking. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, and the advantages thereof, a brief introduction will be given to the drawings that need to be used in the embodiments or prior art description. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without any creative effort.

[0019] Figure 1 A flowchart of a digital factory production data anomaly self-checking method provided by an embodiment of the present application;

[0020] Figure 2 A flowchart of S200 provided by an embodiment of the present application;

[0021] Figure 3 A structural diagram of a digital factory production data anomaly self-checking system provided by an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined purposes, the following describes a digital factory production data anomaly self-checking method, system and medium according to the present application, the specific implementation, structure, features and effects thereof in detail, with reference to the drawings and preferred embodiments. Different "one embodiment" or "another embodiment" in the following description do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0023] In the detailed description of the present application, various mathematical calculations and formulas will be involved. Those skilled in the art should understand that when processing actual production data, certain specific boundary conditions may occur. For example, when performing an operation involving division, if the calculation result of the variable or expression as the denominator is zero, in order to ensure the robustness and numerical stability of the calculation, the present application can use the numerical processing means known in the art.

[0024] Since the numerical range of the denominator in each calculation formula involved in the present application is greater than 0, and in extreme cases it can be equal to 0, an exemplary processing method is to add a very small positive number (for example, a preset non-zero constant such as 0.001, the specific value of which can be set by the implementer according to the actual situation, and the present embodiment does not make specific limitations) to the denominator, so as to ensure that the denominator is not zero after adding a very small positive number, so as to avoid division by zero error.

[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.

[0026] In traditional production data anomaly detection methods, adaptive models trained on historical data face the dual challenges of data redundancy and feature dilution. When production equipment continuously generates data with the same source and attributes under the same processing conditions, historical datasets and real-time data streams exhibit highly repetitive features, leading to ineffective iterations during model training. Simultaneously, newly emerging non-repetitive data features are submerged in a large number of historical samples, making it difficult for the model to quickly identify subtle shifts in data patterns. This results in anomaly detection threshold updates lagging behind actual changes in production conditions.

[0027] Faced with the aforementioned problems, this invention first considers how to effectively distinguish repetitive features in real-time data streams and historical data. Traditional methods directly mix new data with all historical data for training, causing the model to fail to focus on effective shifts in data patterns. To address this, this invention attempts to approach the issue from a data classification perspective, clustering production data with similar origins and attributes according to processing status to form data clusters with consistent patterns. Further analysis reveals that the degree of data redundancy varies across different processing statuses; applying a uniform sampling strategy to historical data may dilute the features of new data. Based on this, this invention proposes quantifying the consistency of patterns between real-time and historical data under the same processing status and dynamically adjusting the sampling weights of historical data, thereby enhancing the learning of new data patterns while preserving effective historical features.

[0028] In this regard, such as Figure 1 As shown, this invention provides a flowchart of a method for self-checking anomalies in production data of a digital factory. This method can be applied to electronic devices and may include the following steps S100 to S400:

[0029] S100, classify the first production data in the current production data and the second production data in the historical production data to obtain several production data clusters; the first production data is the production data in the current production data that has the same equipment source and the same processing attributes, and the second production data is the production data in the historical production data that has the same equipment source and the same processing attributes as the first production data. The production data cluster includes the first production data and the second production data that are in the same processing state.

[0030] In this step, current production data refers to the data generated by production activities in the current time period, reflecting real-time information such as the operating status of production equipment, processing parameters, and production progress in the current period; historical production data refers to the data generated by production activities in the past period. This data records various situations in the past production process and can be used to analyze production patterns and compare the current production status.

[0031] The first production data refers to production data in the current production data that meets the same equipment source and the same processing attribute. The same equipment source means that the data is generated by the same equipment or the same group of equipment with the same model and function; the same processing attribute means that the processing technology, processing parameters, etc. corresponding to the data have consistency.

[0032] The second production data refers to production data in the historical production data that meets the same equipment source and the same processing attribute as the first production data. That is, the historical data is also generated by a specific equipment and has the same processing method as the first production data.

[0033] The production data cluster refers to a plurality of data sets obtained by classifying the first production data and the second production data. The data in each production data cluster is in the same processing state, and includes the first production data from the current production and the second production data from the historical production.

[0034] Specifically, first, the current production data, i.e. the data reflecting the real-time information of equipment operation, processing parameters, production progress, etc. generated by the production activities in the current time period, and the historical production data, i.e. the data recorded by the production activities in the past period of time, are obtained. Then, the data with the same equipment source and the same processing attribute are selected from the current production data as the first production data, and the data with the same equipment source and the same processing attribute as the first production data are found from the historical production data as the second production data. Then, the first production data and the second production data are classified according to the processing state, and the two types of data in the same processing state are grouped into a group to obtain a plurality of production data clusters.

[0035] Among them, the data is screened through the two key dimensions of equipment source and processing attribute, which can ensure that the data included in the same data cluster has similar generation environment and processing characteristics. According to the same processing state classification, because the change rule and characteristics of the data in the same processing state are more similar, such classification is helpful for subsequent analysis and comparison of data rules, and lays a foundation for accurate evaluation of data redundancy and construction of effective training data set.

[0036] In S200, based on the production data cluster, the first production data and the second production data are compared to determine the data redundancy index between the first production data and the second production data; the data redundancy index is used to represent the consistency of the data rules between the first production data and the second production data.

[0037] In this step, the data redundancy index refers to an index for characterizing the consistency of data regularity between the first production data and the second production data. By comparing the first production data and the second production data, the similarity of their data distribution, change trend, correlation, etc. is analyzed, and the data redundancy index is obtained. The data redundancy index can quantify the closeness of the data regularity between the two, and the higher the value, the more consistent the data regularity, and there may be more redundant information; the lower the value, the greater the difference in data regularity.

[0038] Specifically, the first production data and the second production data are compared in detail first. From the aspect of data distribution, whether the concentration trend and dispersion degree of the data values of the two are similar is viewed; the change trend is analyzed to observe whether the change trend of the data with time or other factors is consistent; the correlation is studied to determine whether there is an association between the two and the closeness of the association. The analysis results of these aspects and the production data cluster are combined to calculate a data redundancy index for characterizing the consistency of the data regularity of the two.

[0039] Among them, data distribution, change trend and correlation are important aspects reflecting data regularity. Similar data distribution means that the data is similar in value range and concentration degree; consistent change trend means that the data has the same change pattern over time or other variables; and close correlation means that there is a stable relationship between the two. The data redundancy index calculated by comprehensively considering these factors can accurately quantify the closeness of the data regularity between the first production data and the second production data, providing a key basis for subsequent resampling operations.

[0040] S300, based on the data redundancy index, resampling the second production data to construct a training data set of the first production data.

[0041] In this step, resampling refers to the operation of resampling the second production data. According to the data redundancy index, some data is selectively extracted from the second production data to construct a training data set suitable for training a model related to the first production data. The purpose of resampling may be to balance the data volume, remove redundant data, or enhance the representativeness of the data, so that the training data set constructed can better reflect the characteristics and regularity related to the first production data.

[0042] The training data set refers to a data set for training an anomaly detection model after resampling and other processing. The training data set contains the second production data that has been filtered and sorted, and is intended to provide learning samples for the anomaly detection model, so that the anomaly detection model can learn the characteristics and patterns of normal production data to accurately perform data anomaly self-checking on the first production data.

[0043] Specifically, the second production data is resampled according to the calculated data redundancy index. If the data redundancy index is high, it indicates that the second production data has high consistency with the first production data, and there may be more redundant information. At this time, the number of samples extracted from this part of the second production data can be appropriately reduced. If the data redundancy index is low, it indicates that the data rules of the two are quite different. To enhance the representativeness of the data, the proportion of samples extracted from this type of second production data needs to be increased. After such selective extraction, a training data set suitable for training the first production data related model is constructed.

[0044] The data redundancy index reflects the value of the second production data for training the anomaly detection model for the first production data. Data with a high redundancy index may not provide new effective information for the anomaly detection model, and may even interfere with the anomaly detection model learning the true characteristics of the first production data. Data with a low redundancy index may contain important rules not reflected in the first production data. By resampling according to the data redundancy index, the data volume can be balanced, redundant data can be removed, and the representativeness of the data can be enhanced, so that the training data set can more accurately reflect the characteristics and rules related to the first production data, thereby improving the training effect of the anomaly detection model.

[0045] S400, based on the training data set, fine-tuning the anomaly detection model to enable the anomaly detection model based on the fine-tuning to perform data anomaly self-checking on the first production data.

[0046] In this step, the anomaly detection model refers to a model used to detect abnormal conditions in production data. The anomaly detection model learns the characteristics and rules of normal production data and establishes corresponding judgment criteria. When new production data (such as the first production data) is input, it can determine whether it deviates from the normal range, thereby discovering possible abnormal conditions such as equipment failure, abnormal processing parameters, etc.

[0047] Fine-tuning refers to the process of further adjusting and optimizing the anomaly detection model that has been preliminarily trained. Based on the constructed training data set, the parameters, structure or training strategy of the anomaly detection model are adjusted to make the anomaly detection model better adapt to the characteristics and needs of the first production data, and improve the accuracy and reliability of the anomaly detection model in performing data anomaly self-checking on the first production data.

[0048] Data anomaly self-checking refers to the process of automatically analyzing and judging the first production data using the fine-tuned anomaly detection model to detect whether there is abnormal data. Without human intervention, the anomaly detection model can quickly and accurately identify abnormal conditions in the first production data according to the learned normal data patterns, providing timely and effective information for the monitoring and optimization of the production process.

[0049] Specifically, the fine-tuning is performed on the preliminarily trained anomaly detection model using the constructed training data set. During the fine-tuning, the parameters of the anomaly detection model, such as the weights in the neural network, the threshold, etc., are adjusted according to the characteristics of the training data set; the structure of the model is optimized, for example, some layers are added or removed; the training strategy is adjusted, such as changing the learning rate, the number of iterations, etc. After a series of adjustments and optimizations, the fine-tuned anomaly detection model is obtained, and then the fine-tuned anomaly detection model is used for data anomaly self-checking of the first production data. The anomaly detection model automatically analyzes and judges whether there is an abnormal situation in the first production data.

[0050] The preliminarily trained anomaly detection model may not fully adapt to the characteristics and requirements of the first production data. The training data set contains data related to the first production data after screening. By fine-tuning the anomaly detection model based on this data set, the anomaly detection model can learn patterns that are more consistent with the characteristics and rules of the first production data, thereby establishing more accurate judgment criteria. In this way, when the first production data is input, the anomaly detection model can quickly and accurately identify abnormal data that deviates from the normal range based on the learned normal data patterns, achieving effective anomaly self-checking of the first production data and providing timely and reliable information for the monitoring and optimization of the production process.

[0051] The present application dynamically quantifies the consistency of the current production data and the historical production data, screens and reconstructs the training data set, thereby optimizing the adaptability of the anomaly detection model to the current production scene, solving the problem of low model update efficiency caused by redundant data in massive data, and improving the accuracy and timeliness of anomaly self-checking.

[0052] As an example, in an automobile parts production line, a press machine generates multi-dimensional sensor data including tonnage, stroke, temperature, etc. every minute. First, the real-time temperature data collected in the last 8 hours of the automobile parts production line is taken as the first production data, and the historical temperature data of the automobile parts production line in the past two years is taken as the second production data. These data are classified to obtain multiple production data clusters, each cluster representing a processing state.

[0053] Further, a data redundancy index is calculated to represent the consistency of the data rules of the first production data and the second production data based on the comparison results between the first production data and the second production data and the production data clusters.

[0054] Therefore, the second production data is resampled according to the data redundancy index. For highly repetitive processing states, the sampling frequency is reduced; for newly appearing processing states, the sampling frequency is increased. The training data set constructed in this way not only retains the effective information of the historical data, but also highlights the features of the new data.

[0055] Finally, the abnormality detection model is fine-tuned using the resampled training data set. The fine-tuned abnormality detection model is used for real-time abnormality detection of the first production data, and can quickly identify small abnormalities in the production process, such as abnormal temperature rise of the mold.

[0056] Through the embodiment, the first production data and the second production data with the same equipment source and processing attribute in the current production data and the historical production data are classified into production data clusters first, so that the production data clusters cover the production data in the same processing state. Then, the data redundancy index is determined by comparing the first production data and the second production data, and the data regularity consistency between the first production data and the second production data is determined. Based on the data redundancy index, the second production data is resampled to construct a training data set, which can effectively filter out valuable data for learning new data and avoid repeated data interference. Finally, the abnormality detection model is fine-tuned using the training data set, so that the abnormality detection model can more accurately perform abnormality self-detection on the first production data in the current production data, overcoming the problem of insufficient learning of new data, thereby improving the accuracy of production data abnormality self-detection.

[0057] In some schemes of the present application, when classifying the current production data and the historical production data, if only simple division is performed, it is difficult to distinguish data in different processing states under the same equipment and processing attribute. Since the processing state change is not effectively identified, the same production data cluster may contain data in different processing states, which leads to inaccurate calculation of the data redundancy index and affects the quality of the training data set construction.

[0058] To this end, the present application further provides that S100 comprises:

[0059] For each production data in the first production data and the second production data, the processing attribute of the production data is taken as a dimension to construct a target processing vector of the production data;

[0060] For each target processing vector, based on the similarity between the target processing vector and a reference processing vector at a previous sampling time, a processing state change degree of the target processing vector is determined;

[0061] Based on the processing state change degree of each target processing vector, each target processing vector is clustered to obtain a plurality of production data clusters.

[0062] In this embodiment, the target processing vector is constructed by extracting the processing attribute dimension. The processing state change degree is determined by calculating the cosine similarity between the target processing vector and the reference processing vector at the previous sampling time, and the numerical range is controlled between 0 and 1. Specifically, for each target processing vector, first determine the target processing vector corresponding to the production data at the previous sampling time as the corresponding reference processing vector; then, calculate the cosine similarity between the target processing vector and the reference processing vector at the previous sampling time, normalize the cosine similarity to obtain a value with a numerical range controlled between 0 and 1; finally, subtract the normalized cosine similarity from 1 to obtain the processing state change degree of the target processing vector.

[0063] The clustering process adopts a density clustering algorithm, and the clustering radius threshold can be set to 0.3, the minimum sample number is 5, and the vectors with a processing state change degree difference less than 0.2 are classified into the same cluster. The specific operation is: in the density clustering process, the processing state change degree is used as the clustering feature, for each target processing vector, the difference value of the processing state change degree with the surrounding processing vectors (within the set clustering radius range) is calculated, when the difference value is less than 0.2, these processing vectors are classified into the same potential cluster, and finally the production data cluster is determined according to the minimum sample number and other limitation conditions.

[0064] As an example, the first production data in the current production data and the second production data in the historical production data are classified to obtain a plurality of production data clusters. Specifically, for each production data in the first production data and the second production data, the target processing vector of the production data is constructed by taking the processing attribute of the production data as the dimension. For example, the processing attribute can be processing temperature, processing pressure, processing time, etc. Further, for each target processing vector, the processing state change degree of the target processing vector is determined based on the similarity between the target processing vector and the reference processing vector corresponding to the previous sampling time. The similarity can be measured by calculating the cosine similarity between the two vectors. Thus, based on the processing state change degree of each target processing vector, the target processing vectors are clustered to obtain a plurality of production data clusters. The clustering can use common clustering methods such as K-means algorithm.

[0065] Through this embodiment, the target processing vector is constructed and the processing state change degree is calculated, which can effectively capture the dynamic change characteristics of the production data. Further, based on the processing state change degree, similar processing states can be classified into one category, thereby realizing the fine classification of the production data. Thus, a more accurate data basis is provided for subsequent data comparison and anomaly detection, which helps to improve the accuracy and efficiency of production data anomaly detection.

[0066] In the foregoing schemes of the present application, when determining the data redundancy index by comparing the current production data with the historical production data, if the difference in material properties in different production links is not considered, the consistency of data rules may be judged with deviation, and thus the calculation accuracy of the redundancy index is affected.

[0067] To this end, as shown in Figure 2 the present application further provides that S200 comprises the following S210 to S260:

[0068] S210, a current production link corresponding to a current machining vector sequence is acquired; the current machining vector sequence is formed by arranging each target machining vector in the first production data in time sequence;

[0069] S220, based on the difference in material properties between the current production link and a historical production link, a product consistency possibility between the current production link and the historical production link is determined; the historical production link is a production link corresponding to a historical machining vector sequence formed by arranging each target machining vector in the second production data in time sequence;

[0070] S230, based on the product consistency possibility between the current production link and each historical production link, each historical production link is screened to determine a reference production link corresponding to the current production link;

[0071] S240, based on the difference in machining state change degree between adjacent reference production links and the production data cluster, a production data consistency degree between the adjacent reference production links is determined;

[0072] S250, the adjacent reference production links with the production data consistency degree less than a preset consistency threshold are determined as a group of optimization process groups;

[0073] S260, based on the difference between the production link before optimization and the production link after optimization in each optimization process group, and in combination with the comprehensive optimization consistency degree between the current production link and the historical production link, a data redundancy index between the first production data and the second production data is determined.

[0074] In the present embodiment, when acquiring the current production link corresponding to the current machining vector sequence, the target machining vectors in the first production data need to be arranged in time sequence to construct the current machining vector sequence; a machining state change sequence is constructed based on the difference in machining state change degree between adjacent target machining vectors, and then the current production link is divided based on the machining state change sequence.

[0075] In determining the product consistency possibility, the following logic can be used for calculation: first, calculate the difference between the current production link and the historical production link corresponding to each material attribute value, accumulate the material attribute difference value, divide the total number of material attribute values, and obtain the average value of the attribute difference value, and then determine the product consistency possibility based on the average value of the attribute difference value. It needs to be clear that the average value of the attribute difference value and the product consistency possibility are inversely related, that is, the smaller the average value of the attribute difference value, the closer the material attributes of the current production link and the historical production link, and the higher the product consistency possibility; on the contrary, the larger the average value of the attribute difference value, the lower the product consistency possibility. Therefore, the average value of the attribute difference value can be simply inverted to obtain the product consistency possibility.

[0076] It should be noted that in the calculation logic of determining the product consistency possibility, the current method of inverting the average value of the attribute difference value is consistent with the internal idea of determining the product consistency possibility through more complex operations in the future. Both of them measure the product consistency possibility around the core element of attribute difference degree, and the core essence is to judge the consistency degree of the product according to the difference between the material attributes of the production links. Only the current method of inverting the average value of the attribute difference value is a more direct and basic way, while other complex operation methods may be more detailed and diversified in the processing of attribute difference, and there are differences in specific calculation steps and operation forms.

[0077] When screening the reference production link, the historical production link with smaller material attribute difference from the current production link needs to be retained according to the product consistency possibility. Specifically, a threshold value can be set in advance, and the historical production link with product consistency possibility greater than the threshold value from the current production link is determined as the reference production link corresponding to the current production link. The threshold value can be 0.7, and the implementer can set it according to the specific implementation scene.

[0078] When determining the production data consistency degree, the comprehensive processing state difference between adjacent reference production links is calculated to evaluate the production data consistency degree. In the screening and optimization process group, the following methods are used: a preset consistency degree threshold is set in advance, all combinations of adjacent reference production links are traversed, their production data consistency degrees are calculated one by one, and the adjacent reference production link combinations with production data consistency degree less than the preset consistency degree threshold are marked out. Each such adjacent reference production link combination is determined as an optimization process group. The acquisition method of the optimization before production link and the optimization after production link is: in the determined optimization process group, the reference production link in time sequence is the optimization before production link, and the reference production link in time sequence is the optimization after production link. The consistency degree threshold can be 0.3, and the implementer can set it according to the specific implementation scene.

[0079] Specifically, after constructing the current machining vector sequence, the production links are divided by the mutation points in the machining state change sequence to ensure that the machining state changes within each link are smooth. When screening the reference production link, if the product consistency probability is higher than the preset threshold, the historical production link is retained. The determination of the data redundancy index is first to compare the machining optimization vectors of the production links before and after optimization (for example, the machining vector mean of the link before optimization is [1.2, 3.5], and the machining vector mean of the link after optimization is [1.0, 3.3], and the difference vector is [0.2, 0.2]), and then the comprehensive optimization consistency between the current production link and the historical production link is calculated, and the data redundancy index is finally mapped based on the comprehensive optimization consistency.

[0080] As an example, first, the current production link corresponding to the current machining vector sequence is obtained. For example, the target machining vectors in the first production data can be arranged in chronological order to construct the current machining vector sequence. Then, based on the machining state change degree difference between adjacent target machining vectors, the machining state change sequence is constructed. Finally, the current machining vector sequence is divided into multiple current production links according to the machining state change sequence.

[0081] Next, the material attribute difference between the current production link and the historical production link is calculated to determine the product consistency probability. For example, the difference of each material attribute value can be calculated and then standardized, and then the average contribution value of a single standard material attribute difference is obtained by accumulating and dividing by the total number of material attributes, and then the product consistency probability is determined.

[0082] Then, based on the product consistency probability, the historical production link is screened to determine the reference production link. And for adjacent reference production links, the mean of the machining state change degree difference is calculated to obtain the comprehensive machining state difference. At the same time, the proportion of the number of machining vectors in the same production data cluster in adjacent reference production links is counted. Based on these two indicators, the production data consistency is determined.

[0083] Again, the adjacent reference production links with production data consistency less than the preset threshold are determined as the optimization process group.

[0084] The production links corresponding to the target machining vectors before and after optimization are calculated to obtain the machining optimization vector. Based on the machining optimization vectors of each optimization process group, the comprehensive optimization consistency between the current production link and the historical production link is determined, and then the data redundancy index is determined.

[0085] Through the embodiment, the first production data and the second production data can be compared based on the production data cluster to determine the data redundancy index. By carefully comparing and screening the current production link and the historical production link, production data with similar processing characteristics can be accurately identified. At the same time, by calculating the production data consistency and optimizing the process, the data rules and change trends in the production process can be effectively captured. This method avoids simply treating all historical data equally, but intelligently screens and allocates weights according to actual production conditions. Therefore, the present application can more accurately evaluate the relevance between the current production data and the historical production data, improve the accuracy and representativeness of the data redundancy index. This lays a solid foundation for subsequent data resampling and abnormal detection model optimization, and helps to improve the performance and reliability of the entire production data abnormal self-checking system.

[0086] In some schemes of the present application, the product consistency possibility is determined based on the material attribute difference between the current production link and the historical production link, and the reference production link is screened to determine the production data consistency. However, in this process, if the division method of the current processing vector sequence cannot accurately reflect the continuity change of the actual production link, it may cause deviation in screening the reference production link, and thus affect the calculation accuracy of the production data consistency.

[0087] To this end, the present application further proposes that S210 comprises:

[0088] The target processing vectors in the first production data are arranged in time sequence to construct the current processing vector sequence;

[0089] Based on the absolute value of the difference between the processing state change degrees of two adjacent target processing vectors in the current processing vector sequence, a processing state change sequence is constructed;

[0090] Based on the processing state change sequence, the current processing vector sequence is divided into at least one current production link.

[0091] In the present embodiment, the target processing vectors are arranged in time sequence to form the current processing vector sequence, ensuring the time continuity of the data. The absolute value of the difference between the processing state change degrees of two adjacent target processing vectors is calculated by subtracting the processing state change degree of the adjacent target processing vector and taking the absolute value, reflecting the mutation point of state transition in the production process. The processing state change sequence is constructed by continuously calculating the absolute value of the difference between the adjacent processing state change degrees, forming a continuous signal representing state change. The production link is divided based on the distribution position of the mutation point, cutting the sequence into interval segments with stable state.

[0092] Specifically, the target processing vectors are arranged in ascending order of timestamps to form a current processing vector sequence with time sequence correlation. For each pair of adjacent vectors in the current processing vector sequence, the absolute value of the difference between the processing state change degrees of the two vectors is calculated to generate a processing state change sequence containing all adjacent difference absolute values. By setting a mutation threshold, the mutation point positions in the processing state change sequence where the difference absolute values exceed the threshold are identified. The vector interval between adjacent mutation points is divided into independent production links to ensure that the processing state change within each link is smooth. For example, when the absolute value of the difference between the processing state change degrees of adjacent vectors exceeds a preset threshold of 0.15, it is determined that the position is a production link segmentation point. In this way, the current processing vector sequence is divided into multiple production links with stable processing states, providing accurate link boundary conditions for subsequent screening of reference production links, avoiding misjudgment of data consistency due to link division deviation, and improving the calculation reliability of the data redundancy index.

[0093] As an example, when obtaining the current production link corresponding to the current processing vector sequence, first, the target processing vectors in the first production data are arranged in chronological order to construct the current processing vector sequence. For example, assuming that the first production data contains 10 target processing vectors, denoted as V1, V2,..., V10, the current processing vector sequence obtained after arranging in chronological order is [V1, V2,..., V10].

[0094] Next, based on the absolute values of the differences between the processing state change degrees of adjacent two target processing vectors in the current processing vector sequence, a processing state change sequence is constructed. Specifically, the absolute values of the differences between the processing state change degrees of adjacent two target processing vectors are calculated to obtain a processing state change sequence with a length of 9. For example, assuming that the absolute values of the differences between the processing state change degrees of adjacent two target processing vectors are [0.1, 0.2, 0.5, 0.1, 0.3, 0.4, 0.2, 0.1, 0.3].

[0095] Then, based on the processing state change sequence, the current processing vector sequence is divided into at least one current production link. Specifically, a threshold can be set, for example, 0.4, and when the absolute value of the difference between the processing state change degrees is greater than the threshold, it is considered as the start of a new production link. According to the above example, the current processing vector sequence can be divided into three current production links: [V1, V2, V3], [V4, V5, V6], and [V7, V8, V9, V10].

[0096] Through this embodiment, the current production link can be accurately identified and divided, providing a reliable basis for subsequent data comparison and anomaly detection. In this way, the accuracy and efficiency of production data anomaly self-checking are improved, which helps to timely discover and handle abnormal situations in the production process, ensuring production quality and efficiency.

[0097] In some schemes of the present application, when determining the product consistency possibility based on the material attribute difference between the current production link and the historical production link, the measurement of the material attribute difference is relatively vague, and it is difficult to accurately quantify the influence degree of different material attributes on the product consistency, resulting in insufficient accuracy of the screening reference production link and affecting the calculation accuracy of the subsequent data redundancy index.

[0098] To this end, the present application further proposes that S220 comprises:

[0099] The current material attribute values in the current production link are subtracted from the corresponding historical material attribute values in the historical production link, and then standardized to obtain a plurality of standard material attribute difference values;

[0100] The cumulative value of the plurality of standard material attribute difference values is divided by the total number of material attributes to obtain an average contribution value of a single standard material attribute difference value;

[0101] Based on the average contribution value of a single standard material attribute difference value, the product consistency possibility between the current production link and the historical production link is determined.

[0102] In this embodiment, the standard material attribute difference value is calculated by standardizing the material attribute values of the same type in the current production link and the historical production link one by one, such as temperature, pressure or raw material composition ratio. The total number of material attributes refers to the total number of material attributes participating in comparison in the current production link. The average contribution value is obtained by summing the absolute values of all standard material attribute difference values and dividing by the total number, which reflects the average influence of a single attribute difference on the overall consistency. When determining the product consistency possibility based on the average contribution value, the smaller the contribution value, the smaller the influence of the material attribute difference on the product consistency, and the higher the product consistency possibility.

[0103] Specifically, when calculating the product consistency possibility, first, extract the material attribute values of the same type in the current production link and the historical production link, for example, the temperature value of the current link is T1, and the temperature value of the historical link is T2, then the temperature attribute difference is |T1-T2|, and then it is standardized. Then, repeat this operation for all material attributes to obtain each standard material attribute difference value. Sum the absolute values of all standard material attribute difference values and divide by the number of material attribute types to obtain the average contribution value. The smaller the average contribution value, the smaller the material attribute difference between the current production link and the historical production link, and the higher the product consistency possibility. In this way, the material attribute difference is quantified as a specific value, avoiding errors caused by subjective judgment, so as to more accurately screen out the historical reference link matching the current production link and improve the reliability of the data redundancy index.

[0104] The product consistency possibility can be determined by the following formula 1:

[0105] Formula 1

[0106] In formula 1, is used to represent the product consistency possibility between the a th production link and the k th production link, is used to represent the a th material attribute value of the a th production link, is used to represent the k th material attribute value of the k th production link, and norm is used to represent the standardization processing. is used to represent the total number of material attributes, is used to represent the exponential function operation. is used to represent the same processing state moment proportion of the a th production link and the k th production link. Specifically, the same processing state moment proportion of the a th production link and the k th production link can be calculated by dividing the data amount of the a th production link and the k th production link in the same production data cluster by the total data amount of the a th production link and the k th production link.

[0107] As an example, the current material attribute values in the current production link are first subtracted from the corresponding historical material attribute values in the historical production link to obtain a plurality of standard material attribute difference values after standardization processing. For example, for a production link, the hardness, density, and thermal conductivity of the current material are 10, 5, and 2, respectively, and the corresponding material attribute values in the historical production link are 9, 6, and 3, respectively, to obtain material attribute difference values of 1, -1, and -1. Then, the standardization processing is performed on the material attribute difference values to obtain a plurality of standard material attribute difference values.

[0108] The accumulated value of the plurality of material attribute difference values is then divided by the total number of material attributes to obtain the average contribution value of a single material attribute difference value. Finally, based on the average contribution value of a single material attribute difference value, the product consistency possibility between the current production link and the historical production link is determined by the above formula 1.

[0109] Through the embodiment, the product consistency between the current production link and the historical production link can be accurately evaluated based on the material attribute difference. Thus, the historical production data matching the current production link can be effectively screened, and the accuracy of subsequent data analysis and anomaly detection can be improved. Further, by introducing the average contribution value of a single material attribute difference value, the influence of multiple material attributes can be considered, and the misjudgment caused by excessive difference in a single attribute can be avoided, thereby improving the reliability and stability of the product consistency evaluation.

[0110] In some schemes of the present application, when determining the production data consistency degree between adjacent reference production links, if only a single dimension evaluation is made based on the difference of the processing state variation degree, the calculation result of the production data consistency degree may deviate from the true situation. Such single dimension evaluation method may reduce the accuracy of the production data consistency degree, and further affect the reliability of the subsequent data redundancy index, leading to inaccurate construction of the training data set, and finally affecting the fine-tuning effect of the anomaly detection model.

[0111] To this end, the present application further provides that S240 comprises:

[0112] The difference of the corresponding processing state variation degrees between adjacent reference production links is subjected to mean value processing to obtain a comprehensive processing state difference degree between adjacent reference production links;

[0113] The number of first processing vectors located in the same production data cluster between adjacent reference production links is obtained;

[0114] The number of first processing vectors is divided by the total number of processing vectors in the adjacent reference production links to obtain a processing vector quantity proportion;

[0115] Based on the comprehensive processing state difference degree and the processing vector quantity proportion, the production data consistency degree between adjacent reference production links is determined.

[0116] In the present embodiment, the comprehensive processing state difference degree is obtained by calculating the average value of all corresponding processing state variation degree differences between adjacent reference production links. For example, the adjacent links contain three processing state variation degree differences of 0.3, 0.5 and 0.4, and the comprehensive difference degree is 0.4. The processing vector quantity proportion is obtained by counting the number of processing vectors existing in both reference production links in the same production data cluster, and calculating the proportion of the total number of processing vectors in the two links. For example, the adjacent links have a total of 100 processing vectors, of which 60 are located in the same production data cluster, and the proportion is 60%. The production data consistency degree is obtained by weighting the comprehensive processing state difference degree and the processing vector quantity proportion.

[0117] Specifically, first, the difference value of the processing state change degree of the adjacent reference production link is processed by mean value to eliminate single-point fluctuation interference, and a comprehensive index representing the overall processing state difference is obtained. Further, the proportion of the number of processing vectors belonging to the same production data cluster in the two links is counted to reflect the overlap degree of data distribution. The comprehensive processing state difference degree and the proportion of the number of processing vectors are combined, for example, the lower the comprehensive processing state difference degree and the higher the proportion of the number of processing vectors, the higher the consistency of production data. By weighting or multiplying the two indexes, the dynamic change of the processing state and the static characteristics of the data distribution can be considered at the same time, and the limitations of a single index can be avoided. This method can more comprehensively evaluate the consistency of production data between adjacent reference production links, provide a reliable basis for subsequent optimization process group division and data redundancy index calculation, and thus improve the accuracy of training data set construction.

[0118] The consistency of production data between adjacent reference production links can be determined by the following formula 2:

[0119] Formula 2

[0120] In formula 2, is used to represent the consistency of production data between the a th production link and the k th production link, is used to represent the comprehensive processing state difference degree between the a th production link and the k th production link, is used to represent the number of first processing vectors between the a th production link and the k th production link in the same production data cluster, is used to represent the number of processing vectors of the a th production link, is used to represent the number of processing vectors of the k th production link, is used to represent the positive correlation normalization.

[0121] As an example, first, the difference value of the corresponding processing state change degree between adjacent reference production links is processed by mean value to obtain the comprehensive processing state difference degree between adjacent reference production links. For example, the average value of the processing state change degree difference between adjacent reference production link A and reference production link B is calculated to obtain a comprehensive processing state difference degree of 0.35.

[0122] Then, the number of first processing vectors between adjacent reference production links in the same production data cluster is obtained. Specifically, the number of processing vectors in reference production link A and reference production link B that belong to production data cluster X at the same time is counted, and is assumed to be 50.

[0123] The first processing vector quantity is divided by the total number of processing vectors in the adjacent reference production link to obtain a processing vector quantity proportion. Further, if the total number of processing vectors of the reference production links A and B is 200, the processing vector quantity proportion is 50 / 200=25%.

[0124] Finally, based on the comprehensive processing state difference degree and the processing vector quantity proportion, the production data consistency degree between the adjacent reference production links is determined. Thus, the comprehensive processing state difference degree 0.35 and the processing vector quantity proportion 25% can be substituted into the above formula 2 to obtain the production data consistency degree between the reference production link A and the reference production link B.

[0125] Through the embodiment, the processing state change and the data distribution characteristics between the adjacent reference production links can be comprehensively considered, and the production data consistency between the adjacent reference production links can be accurately quantified. This method avoids the one-sidedness caused by relying on a single indicator, and improves the comprehensiveness and accuracy of the production data consistency evaluation. Further, by introducing the concept of production data cluster, the present scheme can better capture the data characteristics under similar processing states, thereby improving the pertinence and reliability of the consistency evaluation.

[0126] In some schemes of the present application, the production data consistency degree between the adjacent reference production links is used to screen the optimization process groups, which is difficult to effectively measure the influence of the difference between the production links before and after optimization in the optimization process group on the data redundancy, resulting in insufficient accuracy of the data redundancy index and affecting the subsequent resampling and model fine-tuning effect.

[0127] To this end, the present application further proposes that S260 comprises:

[0128] Based on the difference between the corresponding target processing vectors between the production links before optimization and the production links after optimization in each optimization process group, a processing optimization vector of each optimization process group is determined;

[0129] Based on the processing optimization vector of each optimization process group, a comprehensive optimization consistency degree of the current production link and the historical production link is determined;

[0130] Based on the comprehensive optimization consistency degree of the current production link and the historical production link, a data redundancy index between the first production data and the second production data is determined.

[0131] In the present embodiment, the processing optimization vector of the optimization process group can be determined by the following formula 3:

[0132] Formula 3

[0133] In formula 3, is used to represent the processing optimization vector of the i-th optimization process group, a tth target processing vector for representing a production link before optimization in the iith optimization process group, a tth target processing vector for representing a production link after optimization in the iith optimization process group, T represents a number of target processing vectors in a production link.

[0134] The comprehensive optimization consistency between the current production link and the historical production link can be determined by the following formula 4:

[0135] Formula 4

[0136] In formula 4, a comprehensive optimization consistency between the uuth current production link and the historical production link, a cosine similarity between a processing optimization vector between the uuth current production link and the last reference production link and a processing optimization vector of the iith optimization process group, a cosine similarity between a processing optimization vector of the iith optimization process group and an average of processing optimization vectors of each optimization process group, I represents a number of optimization process groups. The processing optimization vector between the current production link and the last reference production link can be determined in the same way as formula 3. The current production link is the production link after optimization in formula 3, and the last reference production link is the production link before optimization in formula 3.

[0137] The data redundancy index can be determined by the following formula 5:

[0138] Formula 5

[0139] In formula 5, a data redundancy index between the uuth current production link and the historical production link, a production data consistency between the uuth current production link and the last reference production link, an average of production data consistencies between the uuth current production link and each reference production link, a comprehensive optimization consistency between the uuth current production link and the historical production link.

[0140] As an example, the processing optimization vector of each optimization process group is determined based on the difference between the corresponding target processing vectors between the production link before optimization and the production link after optimization in each optimization process group. Specifically, the difference between the corresponding target processing vectors between the production link before optimization and the production link after optimization is substituted into formula 3 to obtain the processing optimization vector of the optimization process group.

[0141] Further, based on the processing optimization vectors of each optimization process group, the comprehensive optimization consistency degree between the current production link and the historical production link is determined. For example, the processing optimization vectors of each optimization process group are substituted into the above formula 4 to obtain the comprehensive optimization consistency degree between the current production link and the historical production link.

[0142] Therefore, based on the comprehensive optimization consistency degree between the current production link and the historical production link, the data redundancy index between the first production data and the second production data is determined. Specifically, the data redundancy index is calculated by the above formula 5. Wherein, in the case of only one current production link, the data redundancy index between the first production data and the second production data is calculated by the above formula 5; in the case of several current production links, the sub-redundancy index between the current production link and the historical production link is calculated by the above formula 5, and the combination of each sub-redundancy index obtains the data redundancy index between the first production data and the second production data.

[0143] Through the embodiment, the data redundancy degree between the first production data and the second production data can be accurately quantified based on the change of the production link in the optimization process group. By introducing the concepts of processing optimization vector and comprehensive optimization consistency degree, the changes before and after the optimization of the production link can be considered comprehensively, and the one-sidedness caused by a single index can be avoided. At the same time, by using vector calculation and cosine similarity method, the multi-dimensional characteristics of production data can be effectively captured, and the accuracy and reliability of the data redundancy index can be improved. This method can more accurately reflect the change rule of production data, provide reliable basis for subsequent data resampling and abnormal detection model fine-tuning, and improve the performance and efficiency of the entire digital factory production data abnormal self-checking system.

[0144] In some of the above schemes of the application, when resampling the historical production data based on the data redundancy index, if a global unified resampling strategy is used, the data rule differences of different production links cannot be distinguished. The redundancy degree of the historical production data corresponding to different production links and the current production data is different, and unified sampling will cause the proportion of high-redundancy historical data in the training set to be too high, and the proportion of low-redundancy historical data to be insufficient, affecting the model fine-tuning effect.

[0145] To this end, the application further provides that the data redundancy index comprises sub-redundancy indexes of the first production data in each current production link; S300 comprises:

[0146] The sub-redundancy indexes of each current production link and the original sampling frequency of the reference production link corresponding to the current production link in the second production data are used to determine the resampling frequency of each reference production link;

[0147] The second production data of each reference production link is resampled based on a resampling frequency of each reference production link, and a training data set of the first production data is constructed.

[0148] In the embodiment, the sub-redundancy index is obtained by calculating the comprehensive optimization consistency between the current production link and the historical production link, and reflects the matching degree of the data rules of the two. The original sampling frequency is the actual sampling frequency of the reference production link in the historical data, which is used to measure the distribution characteristics of the historical data. The resampling frequency is generated by inverse operation of the sub-redundancy index and the original sampling frequency, for example, the sub-redundancy index is used as the divisor to adjust the original sampling frequency. For the reference production link with a high sub-redundancy index, it indicates that the consistency of the data rules is high, and there is more redundant information, so the resampling frequency should be reduced to reduce the interference of redundant data; for the reference production link with a low sub-redundancy index, it indicates that the difference of the data rules is large, and contains more valuable learning information, so the resampling frequency should be maintained or appropriately increased to enhance the learning of the model to the new rules.

[0149] Specifically, when determining the resampling frequency, first, the sub-redundancy index of the current production link is extracted, which is calculated by the difference between the processing optimization vectors of the production link before and after optimization in the optimization process group. At the same time, the original sampling frequency of the reference production link is obtained, which is determined by the equipment parameters or process requirements during historical data collection. The sub-redundancy index is used as an adjustment factor to perform inverse operation on the original sampling frequency to obtain the adjusted resampling frequency. For example, if the sub-redundancy index of a certain reference production link is 1.2 and the original sampling frequency is 10Hz, the resampling frequency can be adjusted to 8.3Hz (for example, 10Hz / 1.2). Subsequently, the second production data of the reference production link is resampled according to the adjusted frequency, for example, using oversampling or undersampling methods, so that the data distribution of each production link in the training data set dynamically matches the redundancy index of the current production link. In this way, the training data set can more accurately reflect the data rules of the current production link, thereby improving the efficiency and accuracy of the fine-tuning of the anomaly detection model.

[0150] The resampling frequency can be determined by the following formula 6:

[0151] Formula 6

[0152] In formula 6, is used to represent the resampling frequency of the reference production link corresponding to the u-th current production link, is used to represent the data redundancy index between the u-th current production link and the historical production link, is used to represent the original sampling frequency of the reference production link corresponding to the u-th current production link.

[0153] As an example, in a semiconductor wafer processing scene, the current production data is divided into three continuous production links. The corresponding sub-redundancy indicators of each link are calculated as 0.5, 1.0, and 2.0, respectively. The original sampling frequency of the corresponding reference production link in the historical data is 20Hz. According to formula 6, the resampling frequency is adjusted to 40Hz, 20Hz, and 10Hz, respectively. The linear interpolation method is used to downsample the historical production data. When the resampling frequency is higher than the original frequency, supplementary data points are generated by the time series prediction model. After the historical data and the current production data are aligned by timestamp, a training data set containing 12000 samples is constructed.

[0154] Through the embodiment, the problem of low learning efficiency of the model caused by inconsistent distribution of historical data and current data is effectively solved. By dynamically adjusting the sampling density of the historical data, the effective features in the historical data are retained, and the data distribution characteristics matching the current production law are strengthened, so that the training data set can accurately reflect the actual state of the current production link, and the recognition sensitivity of the abnormal detection model to new production data is improved.

[0155] Based on the digital factory production data anomaly self-checking method provided by the embodiment of the application. Correspondingly, the application also provides a specific embodiment of a digital factory production data anomaly self-checking system.

[0156] As shown in Figure 3 The structure diagram of the digital factory production data anomaly self-checking system 300 is provided, which includes a data classification module 310, a data comparison module 320, a data resampling module 330, and a model training module 340.

[0157] The data classification module 310 is used for classifying the first production data in the current production data and the second production data in the historical production data to obtain a plurality of production data clusters; the first production data is the production data with the same equipment source and the same processing attribute in the current production data, the second production data is the production data with the same equipment source and the same processing attribute as the first production data in the historical production data, and the production data cluster includes the first production data and the second production data in the same processing state;

[0158] The data comparison module 320 is used for comparing the first production data and the second production data based on the production data cluster to determine the data redundancy indicator between the first production data and the second production data; the data redundancy indicator is used to represent the data law consistency between the first production data and the second production data;

[0159] The data resampling module 330 is configured to resample the second production data based on the data redundancy index, and construct a training data set of the first production data.

[0160] The model training module 340 is configured to fine-tune the anomaly detection model based on the training data set, so that the anomaly detection model after fine-tuning performs data anomaly self-checking on the first production data.

[0161] In the digital factory production data anomaly self-checking system provided by the embodiment of the present application, the first production data and the second production data with the same equipment source and processing attribute in the current production data and the historical production data are classified into a production data cluster, so that the production data cluster covers the production data in the same processing state. Then, the data redundancy index is determined by comparing the first production data and the second production data, and the data regularity consistency between the first production data and the second production data is determined. The training data set is constructed by resampling the second production data based on the data redundancy index, which can effectively filter out the data valuable for learning new data and avoid repeated data interference. Finally, the anomaly detection model is fine-tuned using the training data set, so that the anomaly detection model can more accurately perform anomaly self-checking on the first production data in the current production data, overcoming the problem of insufficient learning of new data, thereby improving the accuracy of production data anomaly self-checking.

[0162] In addition, in combination with the digital factory production data anomaly self-checking method in the above-mentioned embodiments, the embodiment of the present application can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; the computer program instructions are executed by the processor to implement any one of the digital factory production data anomaly self-checking methods in the above-mentioned embodiments.

[0163] It should be noted that the above-mentioned embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous.

[0164] Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the difference from other embodiments.

Claims

1. A digital plant production data anomaly self-checking method, characterized by, The method comprises: classifying first production data in current production data and second production data in historical production data to obtain a plurality of production data clusters; the first production data is production data with the same equipment source and the same processing attribute in the current production data, the second production data is production data with the same equipment source and the same processing attribute as the first production data in the historical production data, and the production data cluster includes the first production data and the second production data in the same processing state; comparing the first production data and the second production data based on the production data cluster to determine a data redundancy index between the first production data and the second production data; the data redundancy index is used to represent the consistency of data rules between the first production data and the second production data; based on the data redundancy index, resampling the second production data to construct a training data set of the first production data; based on the training data set, fine-tuning the anomaly detection model to enable the anomaly detection model after fine-tuning to perform data anomaly self-checking on the first production data.

2. The digital plant production data anomaly self-checking method according to claim 1, characterized in that, The classification of the first production data in the current production data and the second production data in the historical production data to obtain a plurality of production data clusters comprises: for each production data in the first production data and the second production data, constructing a target processing vector of the production data with the processing attribute of the production data as the dimension; for each target processing vector, determining the processing state change degree of the target processing vector based on the similarity between the target processing vector and the reference processing vector at the previous sampling time; based on the processing state change degree of each target processing vector, clustering each target processing vector to obtain a plurality of production data clusters.

3. The digital plant production data anomaly self-checking method of claim 2, wherein, The comparison of the first production data and the second production data based on the production data cluster to determine the data redundancy index between the first production data and the second production data comprises: obtaining a current production link corresponding to a current processing vector sequence; the current processing vector sequence is formed by arranging each target processing vector in the first production data in time sequence; based on the material attribute difference between the current production link and the historical production link, determining the product consistency possibility between the current production link and the historical production link; the historical production link is a production link corresponding to a historical processing vector sequence, and the historical processing vector sequence is formed by arranging each target processing vector in the second production data in time sequence; based on the product consistency possibility between the current production link and each historical production link, screening each historical production link to determine a reference production link corresponding to the current production link; based on the difference of the processing state change degree between adjacent reference production links and the production data cluster, determining the production data consistency degree between adjacent reference production links; The adjacent reference production links with the production data consistency less than the preset consistency threshold are determined as a group of optimization process groups; Based on the difference between the pre-optimization production link and the post-optimization production link in each optimization process group, a data redundancy index between the first production data and the second production data is determined.

4. The digital plant production data anomaly self-checking method according to claim 3, characterized in that, The current production link corresponding to the current machining vector sequence is obtained, including: Arranging each target machining vector in the first production data in time sequence to obtain a current machining vector sequence; Based on the absolute value of the difference between the machining state change degrees of two adjacent target machining vectors in the current machining vector sequence, a machining state change sequence is constructed; Based on the machining state change sequence, the current machining vector sequence is divided into at least one current production link.

5. The digital plant production data anomaly self-checking method of claim 3, wherein, Based on the material attribute difference between the current production link and the historical production link, the product consistency possibility between the current production link and the historical production link is determined, including: The current material attribute values in the current production link are subtracted from the corresponding historical material attribute values in the historical production link to obtain a plurality of standard material attribute difference values; The cumulative value of the plurality of standard material attribute difference values is divided by the total number of material attributes to obtain an average contribution value of a single standard material attribute difference value; Based on the average contribution value of the single standard material attribute difference value, the product consistency possibility between the current production link and the historical production link is determined.

6. The digital plant production data anomaly self-checking method of claim 3, wherein, Based on the difference between the machining state change degrees of the adjacent reference production links and the production data cluster, the production data consistency between the adjacent reference production links is determined, including: The difference between the machining state change degrees corresponding to the adjacent reference production links is processed by mean value to obtain a comprehensive machining state difference degree between the adjacent reference production links; The number of first machining vectors between the adjacent reference production links in the same production data cluster is obtained; The number of machining vectors is divided by the total number of machining vectors in the adjacent reference production links to obtain a machining vector quantity proportion; Based on the comprehensive machining state difference degree and the machining vector quantity proportion, the production data consistency between the adjacent reference production links is determined.

7. The digital plant production data anomaly self-checking method of claim 3, wherein, Based on the difference between the pre-optimization production link and the post-optimization production link in each optimization process group, a data redundancy index between the first production data and the second production data is determined, including: Based on the difference between the target machining vectors corresponding to the pre-optimization production link and the post-optimization production link in each optimization process group, a machining optimization vector of each optimization process group is determined; Based on the machining optimization vectors of each optimization process group, a comprehensive optimization consistency degree of the current production link and the historical production link is determined; Based on the comprehensive optimization consistency degree of the current production link and the historical production link, a data redundancy index between the first production data and the second production data is determined.

8. The digital plant production data anomaly self-checking method of claim 3, wherein, The data redundancy index comprises a sub-redundancy index of the first production data at each current production link; The resampling of the second production data based on the data redundancy index comprises: determining a resampling frequency of each reference production link based on the sub-redundancy index of each current production link and an original sampling frequency of a reference production link corresponding to the current production link in the second production data; and resampling the second production data of each reference production link based on the resampling frequency of each reference production link to construct the training data set of the first production data.

9. A digital factory production data anomaly self-checking system characterized by, The system comprises: a data classification module configured to classify first production data in current production data and second production data in historical production data to obtain a plurality of production data clusters; the first production data is production data with the same equipment source and the same processing attribute in the current production data, the second production data is production data with the same equipment source and the same processing attribute as the first production data in the historical production data, and the production data clusters comprise the first production data and the second production data in the same processing state; a data comparison module configured to compare the first production data and the second production data based on the production data clusters to determine a data redundancy index between the first production data and the second production data; the data redundancy index is used to represent the consistency of data rules between the first production data and the second production data; a data resampling module configured to resample the second production data based on the data redundancy index to construct a training data set of the first production data; a model training module configured to fine-tune an anomaly detection model based on the training data set to enable the anomaly detection model after fine-tuning to perform data anomaly self-checking on the first production data.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer program instructions, and the computer program instructions are executed by the processor to implement the digital factory production data anomaly self-checking method of any one of claims 1-8.

Citation Information

Patent Citations

  • Intelligent factory data optimization acquisition method based on digital twinning

    CN116578890A

  • Environmental monitoring data anomaly detection method, medium and system

    CN119807728A