Steel defect mode mining method and system based on time sequence curve form clustering

By using a time-series curve morphology clustering method, the problem of dynamic feature extraction from high-frequency time-series data in steel manufacturing was solved. This method enables automatic clustering and visualization of defect patterns, improves the accuracy and interpretability of defect identification, and supports early warning and process optimization in the production process.

CN121479346APending Publication Date: 2026-02-06SHANGHAI BAOSIGHT SOFTWARE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511583012.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing technologies struggle to extract dynamic feature patterns closely related to defect formation from high-dimensional, high-sampling-rate high-frequency time-series data of steel manufacturing. Furthermore, they lack visualization and pattern summarization of different types of defect time-series morphologies, resulting in inaccurate defect identification, delayed early warnings, and a high false alarm rate, making it difficult to support quality closed-loop control under intelligent manufacturing systems.

Method used

By acquiring and preprocessing time-series curves of high-frequency feature parameters and defect information, a standardized sample set is constructed, a time-series morphological distance matrix is ​​calculated, hierarchical clustering analysis is performed, high-defect-risk clusters are identified, and typical defect pattern templates are visualized, thus achieving unsupervised automatic clustering and visualization of defect patterns.

Benefits of technology

It achieves accurate mapping between high-frequency process timing data in steel manufacturing and downstream quality inspection results, improving the accuracy and interpretability of defect identification, supporting early warning and process optimization, and enhancing the robustness and universality of defect pattern recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121479346A_ABST
    Figure CN121479346A_ABST
Patent Text Reader

Abstract

The invention provides a steel defect mode mining method and system based on time sequence curve form clustering. The method comprises the steps that time sequence curves and defect information corresponding to high-frequency characteristic parameters are obtained and preprocessed; constructing a standardized sample set and dividing the standardized sample set into a defect sample set and a non-defect sample set; calculating a time sequence form distance matrix; performing hierarchical clustering analysis based on the matrix and dividing clusters; counting the proportion of defect samples in each cluster to identify a high-defect risk cluster; and visualizing the time sequence curve in the high risk deficiency cluster and extracting a typical defect mode template. The technical problem that the defect dynamic mode is difficult to identify and cluster from the high-frequency time sequence data in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of nondestructive testing technology for metallic materials, specifically to a method and system for mining steel defect patterns by time-series curve morphology clustering. Background Technology

[0002] In the steel manufacturing process, product quality is affected by the dynamic changes of various complex process parameters, such as temperature, pressure, flow rate, and casting speed. With the development of production equipment automation and data acquisition technology, a large amount of high-frequency time-series data can be acquired in real time at the production site, providing a data foundation for process monitoring and quality analysis. However, traditional defect detection and early warning methods mainly rely on manual experience or low-frequency statistical indicators, making it difficult to extract dynamic feature patterns closely related to defect formation from high-dimensional, high-sampling-rate process data.

[0003] In existing technologies, some methods attempt to utilize machine learning algorithms for defect identification and classification. For example, Chinese patent application number CN202510202246.5 discloses a method for classifying internal defects in steel billets based on particle swarm optimization and support vector machine (PSO-SVM). This method trains a model using ultrasonic signal data to achieve defect identification. However, such methods mainly target the static feature classification of single-type defects, relying on manually constructed features and making it difficult to capture dynamic features of morphological changes in time series. Furthermore, these methods typically employ static metrics such as Euclidean distance, which are not robust to phenomena such as time axis misalignment and signal drift, resulting in inaccurate defect feature extraction, delayed warnings, and a high false alarm rate.

[0004] Meanwhile, existing technologies still have shortcomings in terms of the interpretability of defect patterns and automatic clustering capabilities. Because traditional methods are mainly based on black-box classification, they lack visualization and pattern summarization of the temporal morphology of different types of defects, making it difficult to support subsequent root cause analysis and process optimization, thus restricting the quality closed-loop control under the intelligent manufacturing system.

[0005] Therefore, there is an urgent need for a technical solution that can achieve time-shift robust similarity modeling, unsupervised automatic clustering of defect patterns, generation and visualization of typical defect templates for high-frequency time-series data in steel manufacturing processes, so as to improve the accuracy, timeliness and interpretability of defect identification and support early warning and process optimization for high-risk production conditions. Summary of the Invention

[0006] To address the shortcomings of existing technologies, the purpose of this invention is to provide a method and system for mining steel defect patterns using time-series curve morphology clustering.

[0007] According to one aspect of the present invention, a method for mining steel defect patterns by time-series curve morphology clustering is provided, comprising: Step 1: acquiring and preprocessing time-series curves and defect information corresponding to high-frequency feature parameters; Step 2: constructing a standardized sample set of time-series curves, identifying defect samples and non-defect samples based on defect information, and dividing the sample set into a defect sample set and a non-defect sample set; Step 3: calculating the time-series morphological distance matrix between time-series curves based on the sample set; Step 4: performing hierarchical clustering analysis on the sample set based on the time-series morphological distance matrix, and dividing the sample set into clusters according to the analysis results; Step 5: statistically analyzing the proportion of defect samples in each cluster based on the defect sample set and the non-defect sample set, and identifying high-risk defect clusters; Step 6: visualizing the time-series curves in the high-risk defect clusters, and extracting the time-series curves in the high-risk defect clusters as typical defect pattern templates.

[0008] Preferably, in step one, acquiring and preprocessing the time-series curves and defect information corresponding to the high-frequency characteristic parameters includes: acquiring the time-series curves of the high-frequency characteristic parameters during the steel manufacturing process, wherein the high-frequency characteristic parameters include temperature, pressure, flow rate, and drawing speed; acquiring quality inspection data of downstream processes, wherein the quality inspection data includes product number and defect information, wherein the defect information includes defect type, defect size, and defect location; acquiring production material tracking data, wherein the production material tracking data includes product number, material number, product length, material length, and quality judgment information; performing noise suppression and normalization processing on the high-frequency time-series data, and accurately mapping the defect information to the time-series curves based on the time-series curves, quality inspection data, and production material tracking data through position mapping and traceability matching.

[0009] Preferably, the step of performing noise suppression and normalization processing on high-frequency time-series data, and accurately mapping defect information to the time-series curve based on the time-series curve, quality inspection data, and production material tracking data through location mapping and traceability matching, includes: Based on the material number, the defect information in the quality inspection data is associated with the production material tracking data to obtain the physical location of the defect in the length direction of the material; Based on the actual length of the material in the production line and the cumulative casting length collected by the sensors, the physical location of the defect is linearly mapped to the corresponding length coordinates of the high-frequency time-series curve. The mapping relationship is as follows:

[0010] in, It is the length from the starting point of the defect to the head of the material during the continuous casting stage. It is the distance from the defect termination point to the head of the material during the continuous casting stage. It is the actual length of the material during the hot rolling stage. It is the length from the defect initiation point to the head of the material during the hot rolling stage. It is the length from the defect termination end to the material head during the hot rolling stage. The actual length of the material during the continuous casting stage; By aligning the mapped defect locations with the time series curves according to the material number, defect labels or non-defect labels are assigned to the time series curve segments.

[0011] Preferably, in step two, constructing a standardized sample set of time-series curves, identifying defective and non-defective samples based on defect information, and dividing the sample set into a defective sample set and a non-defective sample set includes: for each process characteristic parameter, organizing the corresponding sample set based on the time-series curve, identifying defective and non-defective samples based on defect information, and dividing the sample set into a defective sample set and a non-defective sample set; applying different strategies to window the defective and non-defective samples to form a balanced sample dataset containing defective and non-defective samples; and using a stratified sampling strategy to divide the balanced sample dataset into a training set and a test set.

[0012] Preferably, in step three, the temporal morphological distance matrix between time series curves is calculated based on the sample set, including: calculating the temporal morphological distance between each pair of time series curves in the training set; organizing the calculation results into a symmetric distance matrix and storing it in a specified path.

[0013] Preferably, in step four, hierarchical clustering analysis is performed on the sample set based on the temporal morphological distance matrix, and clusters are divided according to the analysis results. This includes: performing clustering analysis using a hierarchical clustering algorithm based on the temporal morphological distance matrix; merging similar samples using the Ward linking method, average linking method, or full linking method; generating corresponding clustering dendrograms; visualizing the hierarchical relationship between samples through the clustering dendrograms; determining the optimal number of clusters K based on maximizing the silhouette coefficient or process experience; dividing the sample set into K clusters; and recording the cluster number to which each sample belongs.

[0014] Preferably, in step five, based on the defect sample set and the non-defect sample set, the proportion of defect samples in each cluster is statistically analyzed to identify high-risk clusters, including: statistically analyzing the proportion of defect samples within each cluster; screening out high-risk defect clusters, wherein the screening method is one or more of setting an absolute threshold, setting a relative threshold, and statistical testing; and outputting the time-series curves of all process characteristic parameters contained in the high-risk defect clusters and their corresponding original production information.

[0015] Preferably, in step six, visualizing the time-series curves in the high-defect-risk cluster and extracting the time-series curves in the high-defect-risk cluster as typical defect pattern templates includes: performing time dimension alignment processing on all defect samples in the high-defect-rate cluster by resampling each time-series curve on the normalized time axis to make them have a uniform length; calculating the mean curve of the defect samples after the above processing and using the mean curve as the typical defect pattern template of the cluster; or, extracting the median curve of the defect samples after the above processing as the typical defect pattern template of the cluster; or, extracting the principal component curve through principal component analysis as the typical defect pattern template; visually displaying the superimposed spectrum of all defect samples in the high-defect-risk cluster and highlighting its average trend. The mean curve is calculated using the following formula:

[0016] : This indicates the output of the mean curve at the nth time point; M represents the number of defective samples within the cluster; : Represents the feature value of the i-th aligned defect sample at the n-th time point; : Represents the unified time dimension index after resampling, and N is the length of the aligned sequence.

[0017] The median curve is calculated using the following formula:

[0018] : This indicates the output of the median curve at the nth time point; : is the feature value of the i-th aligned defect sample at the n-th time point; : This represents the median of a set of values ​​and has a strong ability to resist outliers.

[0019] The principal component curve is used to extract the first principal component through principal component analysis. To obtain, to be satisfied:

[0020] : Represents the direction vector of the first principal component, that is, the most representative common trend of change; : This is a data matrix consisting of all aligned defect samples, with each row corresponding to one sample and each column corresponding to one time point; : This is the sample covariance matrix (before centering); V: is a unit direction vector that maximizes its projection onto the direction of the data variance; : Represents the optimal vector V that maximizes the objective function; Time-series curve images of defect samples in the remaining clusters are plotted and classified according to the clusters to form a defect pattern atlas.

[0021] Preferably, the step of aligning all defect samples within the high defect rate cluster in terms of time dimension, by resampling each time series curve on a normalized time axis to achieve a uniform length, includes: for each original time series curve, linearly mapping its original sampling time point to the interval [0, 1] to form a normalized time series reference axis; establishing an interpolation model based on this axis, and resampling at N uniformly distributed time positions corresponding to the target length to obtain a new curve with uniform length and time alignment, wherein the resampling adopts the following formula:

[0022] Where x is the original time series curve, The aligned, equal-length timing curves , which are preset, equally spaced, normalized time points. This represents the interpolation function, which can be one of linear interpolation, cubic spline interpolation, or nearest neighbor interpolation.

[0023] According to another aspect of the present invention, a steel defect pattern mining system based on time-series curve morphology clustering is provided, characterized by comprising: module M1: acquiring and preprocessing time-series curves and defect information corresponding to high-frequency feature parameters; module M2: constructing a standardized sample set of time-series curves, identifying defect samples and non-defect samples based on defect information, and dividing the sample set into a defect sample set and a non-defect sample set; module M3: calculating the time-series morphological distance matrix between time-series curves based on the sample set; module M4: performing hierarchical clustering analysis on the sample set based on the time-series morphological distance matrix, and dividing the sample set into clusters according to the analysis results; module M5: statistically analyzing the proportion of defect samples in each cluster based on the defect sample set and the non-defect sample set, and identifying high-risk clusters; module M6: visualizing the time-series curves in the high-risk clusters, and extracting the time-series curves in the high-risk clusters as typical defect pattern templates.

[0024] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention achieves precise mapping and time alignment between high-frequency process timing data in steel manufacturing and downstream quality inspection results. By introducing material tracking data to establish a spatial-temporal-product number mapping relationship, it effectively solves the synchronization problem between upstream continuous sampling data and downstream discrete inspection results, achieving a precise correspondence between defect features and process parameters, and providing a reliable data foundation for subsequent pattern mining.

[0025] 2. This invention constructs a two-way comparison system for defective and non-defective samples. By organizing, windowing, and standardizing defective and normal samples separately, not only is data balance and sample comparability guaranteed, but a reference benchmark for normal operating conditions is also provided for defect pattern analysis, thereby improving the accuracy of model identification and the interpretability of results.

[0026] 3. This invention employs Dynamic Time Warping (DTW) and its improved algorithm in temporal similarity calculation, overcoming the limitations of traditional Euclidean distance, which is sensitive to time shifts and struggles to capture temporal morphological differences. This method can identify morphologically similar anomalous curves even when time scales are not entirely consistent, effectively improving the robustness and universality of defect pattern recognition.

[0027] 4. This invention achieves unsupervised morphological clustering of multidimensional time-series curves through hierarchical clustering analysis based on the DTW distance matrix and dendrogram visualization. This process automatically discovers potential abnormal patterns without requiring manual definition of defect types or thresholds, and intuitively presents the hierarchical relationships between samples through a tree structure, providing a clear structural basis for summarizing defect patterns.

[0028] 5. The high-defect-risk cluster identification mechanism proposed in this invention achieves automatic identification of high-risk operating conditions by statistically analyzing the defect ratio within clusters and setting a significance threshold. This method combines clustering results with quality information, enabling the clustering results to have quantitative discrimination capabilities and engineering applicability, and can be directly used for risk warning and process optimization decisions in the production process.

[0029] 6. The typical defect pattern template and normal baseline template generated by this invention provide technicians with a visual and interpretable basis for analysis. By comparing the differences in the time-series curve morphology of the two types of templates, the dynamic laws of defect formation and changes in key process characteristics can be intuitively revealed, thereby supporting defect root cause tracing, process parameter optimization, and production process improvement. Attached Figure Description

[0030] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 A simplified flowchart of a steel defect pattern mining method based on time-series curve morphology clustering; Figure 2 A detailed flowchart of a steel defect pattern mining method based on time-series curve morphology clustering. Detailed Implementation

[0031] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0032] For ease of understanding, the terms and concepts used in this application are explained below: (1) High-frequency characteristic parameters refer to key process parameters collected at high sampling frequency during the steel manufacturing process, such as temperature, pressure, flow rate, and pulling speed, which are used to reflect the operating status of equipment and the characteristics of the production process.

[0033] (2) The time series curve represents a continuous numerical sequence of high-frequency characteristic parameters changing over time, and is the basic data form for process status analysis and defect correlation identification.

[0034] (3) Defect information refers to the product defect data identified in downstream quality inspection, including defect type, defect size, defect location, etc., which is used to mark the quality results of the upstream time series curve.

[0035] (4) Production material tracking data refers to data that reflects the flow and related information of products in each link of the production line, such as product number, material number, length and quality judgment information, which are used to realize data traceability and defect location.

[0036] (5) Temporal morphological distance is an indicator that measures the degree of similarity between two time series curves in terms of morphology. It takes into account the time axis misalignment and numerical differences and is used to measure the dynamic correspondence between curves.

[0037] (6) Dynamic Time Warping Algorithm: A commonly used temporal similarity measurement algorithm. It uses dynamic programming to find the optimal matching path on the time axis, thereby eliminating the impact of temporal differences.

[0038] (7) Hierarchical clustering is an unsupervised clustering method that merges or splits samples layer by layer based on the distance between samples. It visualizes the hierarchical relationship between samples by generating a clustering tree diagram.

[0039] (8) Clustering clusters refer to the sample set formed by grouping time series curves with similar shapes into the same class through clustering algorithms, which are used to analyze curve characteristics and defect distribution patterns.

[0040] This application provides a method for mining steel defect patterns based on time-series curve morphology clustering. This method is based on high-frequency time-series characteristic parameters and corresponding quality inspection data from the steel manufacturing process. Through data preprocessing, it achieves accurate mapping between defect information and time-series curves. Subsequently, a standardized sample set is constructed, distinguishing between defective and non-defective samples. A dynamic time warping algorithm is used to calculate the time-series morphological distance between curves, forming a distance matrix. Based on this matrix, hierarchical clustering analysis is performed to divide the data into clusters, and high-defect-risk clusters are identified by statistically analyzing the defect proportion within each cluster. Finally, the time-series curves within high-risk clusters are aligned and interpolated, and the mean or principal component curves are extracted to form typical defect pattern templates, enabling pattern recognition and visualization analysis of steel defects.

[0041] The following is combined with Figure 1 , Figure 2 The steps involved in this method will be explained in detail: Step 1: Acquire and preprocess the time-series curves and defect information corresponding to high-frequency characteristic parameters. This step focuses on key process stages in steel manufacturing, collecting time-series curves of high-frequency characteristic parameters (temperature, pressure, flow rate, casting speed); and simultaneously retrieving downstream process quality inspection data (product number, defect type, defect size, defect location) and production material tracking data (product number, material number, product length, material length, quality judgment information). Noise suppression and normalization are applied to the high-frequency time-series data to ensure comparability of curves across different time periods; subsequently, based on the alignment strategy of "location mapping and traceability matching," the defect information identified downstream is accurately mapped to the upstream high-frequency curves, achieving a one-to-one correspondence and traceable association between defect labels and process curves.

[0042] Based on the above scheme, it can be understood that noise suppression and normalization are performed on high-frequency time series data, and defect information is accurately mapped to the time series curve through location mapping and traceability matching, based on the time series curve, quality inspection data, and production material tracking data. Based on the material number, the defect information in the quality inspection data is associated with the production material tracking data to obtain the physical location of the defect in the length direction of the material; Based on the actual length of the material in the production line and the cumulative casting length collected by the sensors, the physical location of the defect is linearly mapped to the corresponding length coordinates of the high-frequency time-series curve. The mapping relationship is as follows:

[0043] in, It is the length from the starting point of the defect to the head of the material during the continuous casting stage. It is the distance from the defect termination point to the head of the material during the continuous casting stage. It is the actual length of the material during the hot rolling stage. It is the length from the defect initiation point to the head of the material during the hot rolling stage. It is the length from the defect termination end to the material head during the hot rolling stage. The actual length of the material during the continuous casting stage; By aligning the mapped defect locations with the time series curves according to the material number, defect labels or non-defect labels are assigned to the time series curve segments.

[0044] Step two involves constructing a standardized sample set of time-series curves and dividing them into defective and non-defective samples. For each process characteristic parameter, a corresponding sample set is organized, and "defective samples" and "non-defective samples" are labeled based on the defect information obtained in Step one, thus obtaining defective and non-defective sample sets. To mitigate class imbalance, different strategies are used for windowing the two types of samples, generating a balanced sample dataset containing both defective and non-defective samples. Then, a stratified sampling strategy is used to divide this dataset into training and testing sets, ensuring consistency in class distribution across the training and testing ends and facilitating the reliability of subsequent evaluations.

[0045] Step 3: Calculate the temporal morphological distance matrix. In the training set, calculate the "temporal morphological distance" for each pair of time series curves. Dynamic Time Warping (DTW) is preferably used as the preset algorithm model to measure the similarity of unaligned time series. Optional embodiments may support ERP, LCSS, or ShapeDTW to adapt to different process fluctuation patterns. The pairwise distances are organized into a symmetric distance matrix and stored in a specified path for reuse and experiment reproduction. In large-scale data scenarios, FastDTW or constrained DTW can be used for approximate acceleration to improve computational efficiency.

[0046] Step four: Perform hierarchical clustering and divide the sample set into clusters based on the temporal morphological distance matrix. A hierarchical clustering algorithm is used to perform cluster analysis on the sample set. The Ward linking method is preferred for its agglomerative hierarchical clustering to minimize intra-cluster variance. In alternative implementations, the average linking method or the fully linking method can also be used to accommodate different inter-cluster distance metrics. A cluster dendrogram is generated to visualize the hierarchical relationships of the samples. The optimal number of clusters K is determined by maximizing the silhouette coefficient and / or through process experience. The dendrogram is then cut to obtain the final cluster division, and the cluster number of each sample is recorded for subsequent statistics and visualization.

[0047] Step 5: Calculate the proportion of defective samples in each cluster and identify high-risk clusters. For each cluster, calculate its internal "defective sample proportion" and use this to screen for high-risk clusters. Absolute thresholds (e.g., defect rate > 30%), relative thresholds (e.g., more than twice the overall average defect rate), or statistical tests (e.g., chi-square test / hypothesis test) can be used to determine "significantly higher than the average level." For the identified high-risk clusters, output the time-series curves of all process characteristic parameters within the cluster and their corresponding original production information, providing a basis for subsequent process cause analysis and closed-loop optimization.

[0048] Step 6: Visualize high-risk defect clusters and extract typical defect pattern templates. First, align all defect samples within the high-risk defect clusters along the time dimension. Then, use interpolation methods (linear interpolation, cubic spline interpolation, or nearest neighbor interpolation) to resample the curves to a uniform length for morphological comparison and template extraction. Based on the above scheme, for each original time series curve, its original sampling time point is linearly mapped to the interval [0,1] to form a normalized time series reference axis; An interpolation model is established based on this axis. Resampling is performed at N uniformly distributed time positions corresponding to the target length to obtain a new curve with uniform length and time alignment. The resampling uses the following formula:

[0049] Where x is the original time series curve, The aligned, equal-length timing curves , which are preset, equally spaced, normalized time points. This represents the interpolation function, which can be one of linear interpolation, cubic spline interpolation, or nearest neighbor interpolation.

[0050] Subsequently, the mean curve of the aligned curves is calculated as a template for the typical defect pattern of the cluster, or the median curve is selected, or the principal component curve is extracted by principal component analysis (PCA) as a template. For visualization, a superimposed map of defect samples for the cluster is generated and its average trend is highlighted; at the same time, defect sample curves in the other clusters are plotted and classified by cluster to form a defect pattern atlas to support the comparison, archiving and subsequent deployment of multiple types of patterns.

[0051] The mean curve is calculated using the following formula:

[0052] : Represents the mean curve output at the nth time point; M is the number of defective samples within this cluster; : Represents the feature value of the i-th aligned defect sample at the n-th time point; : Represents the unified time dimension index after resampling, and N is the length of the aligned sequence.

[0053] The median curve is calculated using the following formula:

[0054] : This indicates the output of the median curve at the nth time point; : is the feature value of the i-th aligned defect sample at the n-th time point; : This represents the median of a set of values ​​and has a strong ability to resist outliers.

[0055] The principal component curve is used to extract the first principal component through principal component analysis. To obtain, to be satisfied:

[0056] : Represents the direction vector of the first principal component, that is, the most representative common trend of change; : This is a data matrix consisting of all aligned defect samples, with each row corresponding to one sample and each column corresponding to one time point; : is the sample covariance matrix (before centering); V: is the unit direction vector, which maximizes its projection onto the data variance direction; : Represents the optimal vector V that maximizes the objective function.

[0057] This invention also provides a method system for mining steel defect patterns using time-series curve morphology clustering. This method system can be implemented by executing the process steps of the time-series curve morphology clustering method for mining steel defect patterns. That is, those skilled in the art can understand the time-series curve morphology clustering method for mining steel defect patterns as a preferred embodiment of the time-series curve morphology clustering method system for mining steel defect patterns.

[0058] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0059] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A method for mining steel defect patterns using time-series curve morphology clustering, characterized in that, include: Step 1: Obtain and preprocess the time-series curves and defect information corresponding to the high-frequency characteristic parameters; Step 2: Construct a standardized sample set of time series curves. Based on defect information, identify defective samples and non-defective samples, and divide the sample set into a defective sample set and a non-defective sample set. Step 3: Based on the sample set, calculate the temporal morphological distance matrix between time series curves; Step 4: Based on the temporal morphological distance matrix, perform hierarchical cluster analysis on the sample set, and divide the sample into clusters according to the analysis results; Step 5: Based on the defective sample set and the non-defective sample set, calculate the proportion of defective samples in each cluster and identify high-risk clusters; Step 6: Visualize the time series curves in the high-risk defect clusters and extract the time series curves in the high-risk defect clusters as templates for typical defect patterns.

2. The method according to claim 1, characterized in that, In step one, acquiring and preprocessing the time-series curves and defect information corresponding to the high-frequency feature parameters includes: The time-series curves of high-frequency characteristic parameters during the steel manufacturing process are obtained, including temperature, pressure, flow rate, and casting speed. Acquire quality inspection data from downstream processes. The quality inspection data includes product number and defect information. The defect information includes defect type, defect size, and defect location. Acquire production material tracking data, which includes product number, material number, product length, material length, and quality judgment information; Noise suppression and normalization are performed on high-frequency time-series data. Through location mapping and traceability matching, defect information is accurately mapped to the time-series curve based on the time-series curve, quality inspection data, and production material tracking data.

3. The method according to claim 1, characterized in that, The process of noise suppression and normalization of high-frequency time-series data, and the accurate mapping of defect information to the time-series curve based on the time-series curve, quality inspection data, and production material tracking data through location mapping and traceability matching, includes: Based on the material number, the defect information in the quality inspection data is associated with the production material tracking data to obtain the physical location of the defect in the length direction of the material; Based on the actual length of the material in the production line and the cumulative casting length collected by the sensors, the physical location of the defect is linearly mapped to the corresponding length coordinates of the high-frequency time-series curve. The mapping relationship is as follows: in, It is the length from the starting point of the defect to the head of the material during the continuous casting stage. It is the distance from the defect termination point to the head of the material during the continuous casting stage. It is the actual length of the material during the hot rolling stage. It is the length from the defect initiation point to the head of the material during the hot rolling stage. It is the length from the defect termination end to the material head during the hot rolling stage. The actual length of the material during the continuous casting stage; By aligning the mapped defect locations with the time series curves according to the material number, defect labels or non-defect labels are assigned to the time series curve segments.

4. The method according to claim 2, characterized in that, In step two, a standardized sample set of time-series curves is constructed. Based on defect information, defective samples and non-defective samples are identified, and the sample set is divided into a defective sample set and a non-defective sample set, including: For each process characteristic parameter, based on the time series curve, a corresponding sample set is organized, and based on the defect information, defect samples and non-defect samples are identified, and the sample set is divided into a defect sample set and a non-defect sample set. Different strategies are used to window defective and non-defective samples to form a balanced sample dataset containing defective and non-defective samples. A stratified sampling strategy is used to divide the balanced sample dataset into a training set and a test set.

5. The method according to claim 1, characterized in that, In step three, based on the sample set, the temporal morphological distance matrix between time series curves is calculated, including: Calculate the temporal morphological distances between all pairs of time series curves in the training set; The calculation results are organized into a symmetric distance matrix and stored in the specified path.

6. The method according to claim 1, characterized in that, In step four, hierarchical clustering analysis is performed on the sample set based on the temporal morphological distance matrix, and clusters are divided according to the analysis results, including: Based on the temporal morphological distance matrix, a hierarchical clustering algorithm is used to perform cluster analysis. Similar samples are merged using the Ward linking method, average linking method, or full linking method, and corresponding cluster dendrograms are generated. The hierarchical relationship between samples is visualized through the cluster dendrograms. The optimal number of clusters K is determined based on maximizing the profile coefficient or process experience, and the sample set is divided into K clusters. The cluster number to which each sample belongs is recorded.

7. The method according to claim 6, characterized in that, In step five, based on the defective sample set and the non-defective sample set, the proportion of defective samples in each cluster is statistically analyzed to identify high-risk clusters, including: Calculate the proportion of defective samples within each cluster; High-risk defect clusters are screened out, and the screening method is one or more of the following: setting an absolute threshold, setting a relative threshold, and statistical testing. Output the timing curves of all process characteristic parameters contained in the high-risk defect cluster and their corresponding raw production information.

8. The method according to claim 1, characterized in that, In step six, the time-series curves in the high-risk defect clusters are visualized, and the time-series curves in the high-risk defect clusters are extracted as typical defect pattern templates, including: All defect samples within the high defect rate cluster are aligned in the time dimension by resampling each time series curve on the normalized time axis to make them have a uniform length. Calculate the mean curve of the defect samples after the above processing, and use the mean curve as the template of the typical defect pattern of the cluster. Alternatively, extract the median curve of the defect samples after the above processing as the template of the typical defect pattern of the cluster. Alternatively, extract the principal component curve through principal component analysis as the template of the typical defect pattern. Visualize the superimposed spectrum of all defect samples in the high defect risk cluster and highlight its average trend. The mean curve is calculated using the following formula: : This indicates the output of the mean curve at the nth time point; M represents the number of defective samples within the cluster; : Represents the feature value of the i-th aligned defect sample at the n-th time point; : Represents the unified time dimension index after resampling, and N is the length of the aligned sequence. The median curve is calculated using the following formula: : This indicates the output of the median curve at the nth time point; : is the feature value of the i-th aligned defect sample at the n-th time point; : This represents the median of a set of values ​​and has a strong ability to resist outliers. The principal component curve is used to extract the first principal component through principal component analysis. To obtain, to be satisfied: : Represents the direction vector of the first principal component, that is, the most representative common trend of change; : This is a data matrix consisting of all aligned defect samples, with each row corresponding to one sample and each column corresponding to one time point; : This is the sample covariance matrix (before centering); V: is a unit direction vector that maximizes its projection onto the direction of the data variance; : Represents the optimal vector V that maximizes the objective function; Time-series curve images of defect samples in the remaining clusters are plotted and classified according to the clusters to form a defect pattern atlas.

9. The method according to claim 8, characterized in that, The process of aligning all defect samples within a high defect rate cluster along the time dimension involves resampling each time series curve on a normalized time axis to ensure a uniform length. This includes: For each original time series curve, its original sampling time point is linearly mapped to the interval [0, 1] to form a normalized time series reference axis; An interpolation model is established based on this axis. Resampling is performed at N uniformly distributed time positions corresponding to the target length to obtain a new curve with uniform length and time alignment. The resampling uses the following formula: Where x is the original time series curve, The aligned, equal-length timing curves , which are preset, equally spaced, normalized time points. This represents the interpolation function, which can be one of linear interpolation, cubic spline interpolation, or nearest neighbor interpolation.

10. A steel defect pattern mining system based on time-series curve morphology clustering, characterized in that, include: Module M1: Acquires and preprocesses the time-series curves and defect information corresponding to high-frequency characteristic parameters; Module M2: Constructs a standardized sample set of time series curves, identifies defective and non-defective samples based on defect information, and divides the sample set into a defective sample set and a non-defective sample set; Module M3: Based on the sample set, calculates the temporal morphological distance matrix between time series curves; Module M4: Based on the temporal morphological distance matrix, it performs hierarchical cluster analysis on the sample set and divides the data into clusters according to the analysis results; Module M5: Based on the defective sample set and the non-defective sample set, it calculates the proportion of defective samples in each cluster and identifies high-risk clusters. Module M6: Visualizes the time-series curves in high-risk defect clusters and extracts them as templates for typical defect patterns.

Citation Information

Patent Citations

  • Steel billet internal defect classification method, system and equipment based on PSO-SVM (Particle Swarm Optimization-Support Vector Machine)

    CN120046019A