A crystallization anomaly identification method based on unsupervised learning

By using unsupervised learning-based methods, multi-camera video stream data, an improved SAM2 visual segmentation algorithm, and a DINOv2 feature extraction model, the comprehensiveness and robustness of crystallization anomaly identification technology under dynamic process parameter changes were solved, enabling the identification of unknown anomalies and high-precision identification of the main crystallization region.

CN120543498BActive Publication Date: 2026-04-21融域智慧(西安)智能科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
融域智慧(西安)智能科技有限公司
Filing Date
2025-05-15
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing crystallization anomaly identification technologies cannot adapt to dynamic changes in process parameters and require a large amount of labeled data, resulting in poor comprehensiveness and robustness of identification.

Method used

An unsupervised learning-based approach was adopted to acquire crystallization video stream data from multiple cameras. The SAM2 visual segmentation algorithm and DINOv2 feature extraction model were used to extract features and compare similarities of the main crystallization region. An adaptive threshold was then used to determine crystallization anomalies.

Benefits of technology

It enables the identification of unknown anomaly types, improves the comprehensiveness and robustness of crystallization anomaly identification, can adapt to dynamic process parameter changes, reduces dependence on labeled data, and improves the identification accuracy of the main crystallization region.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543498B_ABST
    Figure CN120543498B_ABST
Patent Text Reader

Abstract

The application discloses a crystallization abnormality recognition method based on unsupervised learning, relates to the technical field of crystallization abnormality recognition, and comprises the following steps: setting a plurality of positions, acquiring crystallization video stream data, frame extraction and analysis to obtain multi-position crystallization frame images; each image is cut into a plurality of subblocks; an unsupervised learning algorithm is used to recognize the crystallization main body area of each subblock; feature extraction is performed on the crystallization main body area to obtain crystallization main body features; the similarity between the crystallization main body features of each subblock is compared, and the minimum similarity is taken as the overall similarity; the overall similarity is compared with an adaptive threshold value to determine whether there is crystallization abnormality. The application realizes all-around abnormality recognition by acquiring crystallization video stream data through multiple positions, and combines the unsupervised learning algorithm, feature extraction, similarity comparison and adaptive threshold value, so that the comprehensiveness and robustness of crystallization abnormality recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of crystallization anomaly recognition technology, and in particular to a crystallization anomaly recognition method based on unsupervised learning. Background Technology

[0002] The main phenomena of crystallization anomalies are irregular crystal forms (e.g., flocculent matter appearing instead of crystals under normal conditions) or uneven grain size distribution (in high-purity scenarios). The impact of crystallization anomalies on chemical experiments mainly includes abnormal yields (decreased product purity) and safety hazards (toxic gases from side reactions, pressure changes). Therefore, research is needed on methods for identifying crystallization anomalies.

[0003] In the prior art, Chinese patent CN117036261A discloses an adipic acid reaction chamber crystallization anomaly monitoring system and method, including: acquiring video footage; sampling single-frame images; processing the single-frame images using a reaction chamber crystallization monitoring method to determine the crystallization status within the reaction chamber; performing frequency domain image conversion on the single-frame grayscale image using Fast Fourier Transform (FFT), filtering the frequency domain image, inputting the filtered frequency domain image into a Long Short-Term Memory (LSTM) network for prediction, and obtaining a predicted image in the time domain using Inverse Fast Fourier Transform (IFFT); inputting the frequency domain image into a denoising autoencoder network, and obtaining a denoised image using IFFT; calculating the displacement rate matrix of crystal boundary pixels in the predicted image and the denoised image; fusing the displacement rate matrix and the denoised image and inputting them into a binary classifier to determine whether there is a crystallization anomaly, and issuing an early warning if an anomaly is detected.

[0004] However, the aforementioned existing technologies cannot adapt to dynamic changes in process parameters, require a large amount of labeled data for training, and are difficult to identify unseen crystallization anomalies, resulting in poor comprehensiveness and robustness in crystallization anomaly identification. Summary of the Invention

[0005] This application provides a crystallization anomaly identification method based on unsupervised learning to address the problem of poor comprehensiveness and robustness in existing crystallization anomaly identification technologies.

[0006] On the one hand, this application provides a method for identifying crystallization anomalies based on unsupervised learning, including the following steps:

[0007] Step 1: Set up several camera positions, acquire the crystal video stream data of each camera position, and extract and parse the frames to obtain multi-camera crystal frame images.

[0008] Step 2: Cut each image of the multi-camera crystal frame image into several sub-blocks.

[0009] Step 3: Use an unsupervised learning algorithm to identify the main crystallization region of each sub-block in Step 2.

[0010] Step four: Extract features from the main crystallization region to obtain the main crystallization features of each sub-block.

[0011] Step 5: Compare the similarity between the main crystalline features of each sub-block, and take the minimum similarity as the overall similarity.

[0012] Step six: Compare the overall similarity with the adaptive threshold to determine whether there is any crystallization anomaly.

[0013] In one possible implementation, in step three, the unsupervised learning algorithm employs the SAM2 visual segmentation algorithm.

[0014] In one possible implementation, in step three, the network modules of the SAM2 visual segmentation algorithm include: an image encoder, a memory attention module, a cue encoder, a mask decoder, a memory encoder, and a memory bank.

[0015] In one possible implementation, step three involves improving the network of the SAM2 visual segmentation algorithm, including: parameter freezing, low-rank adaptation, mixed-precision training, loss function improvement, evaluation metric improvement, and dataset improvement.

[0016] In one possible implementation, in step four, the DINOv2 feature extraction model is used to extract features from the main crystal region.

[0017] In one possible implementation, in step four, the self-supervised learning framework of the DINOv2 feature extraction model adopts a dual-network architecture based on the DINO network and the iBOT network.

[0018] In one possible implementation, in step four, the loss function of the DINOv2 feature extraction model adopts a multi-objective loss function.

[0019] In one possible implementation, in step five, the similarity between the crystalline main features of each sub-block is calculated using cosine similarity.

[0020] In one possible implementation, in step six, the adaptive threshold is calculated using an adaptive thresholding method based on local statistics.

[0021] The crystallization anomaly identification method based on unsupervised learning in this application has the following advantages:

[0022] By acquiring crystallization video stream data from multiple cameras, comprehensive anomaly identification is achieved. Combining unsupervised learning algorithms, feature extraction, similarity comparison, and adaptive thresholds improves the comprehensiveness and robustness of crystallization anomaly identification. The unsupervised learning algorithm eliminates dependence on labeled data, enabling the identification of unknown anomaly types, while the adaptive thresholds can adapt to dynamic changes in process parameters.

[0023] The proposed improvements to the SAM2 visual segmentation algorithm include parameter freezing, low-rank adaptation, mixed-precision training, improved loss function, improved evaluation metrics, and improved dataset, which enhance the segmentation performance of the SAM2 visual segmentation algorithm and thus improve the recognition accuracy of the crystallized main body region. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart illustrating a crystallization anomaly identification method based on unsupervised learning, provided in an embodiment of this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] like Figure 1 As shown in the figure, this application provides a method for identifying crystallization anomalies based on unsupervised learning, including the following steps:

[0028] Step 1: Set up several camera positions, acquire the crystal video stream data of each camera position, and extract and parse the frames to obtain multi-camera crystal frame images.

[0029] Step 2: Cut each image of the multi-camera crystal frame image into several sub-blocks.

[0030] Step 3: Use an unsupervised learning algorithm to identify the main crystallization region of each sub-block in Step 2.

[0031] Step four: Extract features from the main crystallization region to obtain the main crystallization features of each sub-block.

[0032] Step 5: Compare the similarity between the main crystalline features of each sub-block, and take the minimum similarity as the overall similarity.

[0033] Step six: Compare the overall similarity with the adaptive threshold to determine whether there is any crystallization anomaly.

[0034] Specifically, in this embodiment, in step one, three camera positions are set up to acquire crystal video stream data from different angles through RTSP video stream or SDK video stream, and frame extraction and parsing are used to obtain multi-camera crystal frame images, with three crystal frame images existing at the same time.

[0035] In step two, each of the three crystal frame images is cut into a preset number of sub-blocks.

[0036] For example, in step three, the unsupervised learning algorithm uses the SAM2 visual segmentation algorithm.

[0037] Specifically, the SAM2 visual segmentation algorithm can avoid the crystal diversity problem caused by the complexity of chemical experiments, enabling the algorithm to achieve good accuracy even when encountering untrained samples.

[0038] For example, in step three, the network modules of the SAM2 visual segmentation algorithm include: an image encoder, a memory attention module, a cue encoder, a mask decoder, a memory encoder, and a memory bank.

[0039] Specifically, the image encoder is responsible for converting video frames into a series of feature embeddings, which can then be used by other parts of the model for video frame segmentation and target tracking.

[0040] The role of memory attention is to adjust the feature representation of the current frame, making it dependent on features from past frames, predictions, and any new cues. It takes into account the temporal continuity in the video, thus better understanding the context of the current frame. The model consists of L stacked transformer blocks. Each transformer block performs a specific attention operation. The first transformer block receives the image encoding of the current frame as input.

[0041] The cue encoder is the same as the encoder in the SAM model, capable of handling different types of cues, including positive / negative clicks, bounding boxes, or masks, to define the extent of an object in a given frame. For sparse cues (such as clicks), they are represented by adding a positional encoding to the learned embedding for each cue type. Positional encoding helps the model understand the specific location of the cue in the image. For masked cues, the mask is embedded using a convolutional operation and then added to the frame embedding. In this way, the mask information is integrated into the feature representation of the frame.

[0042] The mask decoder is designed to largely follow the SAM model. It uses bidirectional transformer blocks to update cues and frame embeddings. For ambiguous cues (such as a single click) that may produce multiple compatible target masks, the model predicts multiple masks. This design ensures that the model outputs a valid mask. In video, ambiguity can span multiple frames. Therefore, the model predicts multiple masks per frame. If no subsequent cues resolve this ambiguity, the model only propagates the mask with the highest predicted IoU (Intersection over Union) for the current frame.

[0043] The memory encoder generates a memory by downsampling the output mask using convolutional modules and adding it element-wise to the unconditional frame embedding provided by the image encoder. This process combines mask information with features from the original frame. Lightweight convolutional layers are then used to fuse this information, creating a memory representation that includes object segmentation information.

[0044] The memory bank retains past prediction information of target objects in the video by maintaining a first-in, first-out (FIFO) queue, which contains memories of at most N most recent frames. Similarly, the memory bank also stores information from cues, which is stored in a FIFO queue containing at most M cue frames.

[0045] For example, in step three, the network of the SAM2 visual segmentation algorithm is improved, including: parameter freezing, low-rank adaptation, mixed precision training, loss function improvement, evaluation metric improvement, and dataset improvement.

[0046] Specifically, in this embodiment, parameter freezing includes: retaining the SAM2 image encoder (ViT or Hiera architecture), fine-tuning only the decoder and cue encoder, which is suitable for scenarios with small data volumes, and unfreezing the last few layers of the image encoder and the decoder to balance computational cost and performance.

[0047] Low-rank adaptation involves inserting a low-rank matrix into the decoder, reducing the number of trainable parameters (by only 0.1%).

[0048] Mixed-precision training includes using FP16 (16-bit Boolean parameters) to accelerate training and reduce memory usage.

[0049] The loss function improvements include: using the Dice loss function + BCE loss function.

[0050] Improvements to the evaluation metrics include: adopting Dice coefficient, IoU (Intersection over Union), and recall rate in occluded scenes as evaluation metrics.

[0051] Dataset improvements include: using the crystal body as the segmentation object and performing mask annotation (COCO format or binary mask) to obtain the crystal body segmentation dataset used for fine-tuning. A strategy of alternating training with video data (multi-frame) and still images (single frame) was adopted.

[0052] The improved SAM2 visual segmentation algorithm network achieved an accuracy increase from 85% to 93%, with a 6% increase in the Dice coefficient, demonstrating improved segmentation performance and thus enhanced recognition accuracy of the main crystal region.

[0053] For example, in step four, the DINOv2 feature extraction model is used to extract features from the main crystal region.

[0054] Specifically, the DINOv2 feature extraction model generates general visual features through training on large-scale data. Its core idea borrows from the unsupervised pre-training paradigm in NLP, extracting semantic information directly from images through self-supervised learning, rather than relying on text supervision or manual annotation. This greatly reduces annotation work and improves the model's generalization ability.

[0055] For example, in step four, the self-supervised learning framework of the DINOv2 feature extraction model adopts a dual-network architecture based on the DINO network and the iBOT network.

[0056] Specifically, in this embodiment, a dual-network architecture based on the DINO network and the iBOT network is used, combined with a teacher-student collaborative distillation mechanism. The teacher network is updated through the parameter exponential moving average (EMA) of the student network, and both input different enhanced views of the same image to maximize feature consistency.

[0057] For example, in step four, the loss function of the DINOv2 feature extraction model adopts a multi-objective loss function.

[0058] Specifically, in this embodiment, the multi-objective loss function includes image-level loss (DINO), block-level loss (iBOT), and SwAV's Sinkhorn-Knopp centering method, while KoLeo regularization is introduced to prevent feature space collapse.

[0059] For example, in step five, the similarity between the main crystalline features of each sub-block is calculated using cosine similarity.

[0060] Specifically, in this embodiment, the formula for calculating cosine similarity is as follows:

[0061]

[0062] Where m(s,r) represents the cosine similarity between the main crystalline features of sub-block s and sub-block r, cosine s The similarity function is represented by cosine similarity, f(s) and f(r) represent the main crystalline features of sub-blocks s and r, respectively, and ||f(s)||2 and ||f(r)||2 represent the L2 norms of the main crystalline features of sub-blocks s and r, respectively.

[0063] By calculating the cosine similarity between every two main crystal features, the minimum similarity is taken as the overall similarity.

[0064] For example, in step six, the adaptive threshold is calculated using an adaptive thresholding method based on local statistics.

[0065] Specifically, in this embodiment, the adaptive threshold for the subsequent N+1 frames is predicted based on the overall similarity results of the first N frames of the video. The calculation formula is as follows:

[0066] T(x,y)=μ local (x,y)-C·σ local (x,y).

[0067] Where T(x,y) represents the adaptive threshold corresponding to the sliding window centered at point (x,y), and μ local (x,y) represents the mean of the overall similarity within a sliding window centered at point (x,y), reflecting the local average level. σ local (x,y) represents the standard deviation of the overall similarity within the sliding window centered at point (x,y), reflecting the degree of local fluctuation. C represents the adjustment factor, which is between 0.5 and 2, controlling the sensitivity of the threshold to fluctuation.

[0068] In this embodiment, when the overall similarity is greater than or equal to the adaptive threshold, it indicates that the main features of the crystal are basically similar, and the crystal is judged to be normal, so the process returns to step two to continue the loop; when the overall similarity is less than the adaptive threshold, it indicates that the main features of the crystal are significantly different, and the crystal is judged to be abnormal, so that the abnormal crystal situation can be handled accordingly in the future.

[0069] This application's embodiments achieve comprehensive anomaly identification by acquiring crystallization video stream data from multiple cameras. Combining unsupervised learning algorithms, feature extraction, similarity comparison, and adaptive thresholds improves the comprehensiveness and robustness of crystallization anomaly identification. The unsupervised learning algorithm eliminates dependence on labeled data, enabling the identification of unknown anomaly types, while the adaptive threshold can adapt to dynamic changes in process parameters.

[0070] The proposed improvements to the SAM2 visual segmentation algorithm include parameter freezing, low-rank adaptation, mixed-precision training, improved loss function, improved evaluation metrics, and improved dataset, which enhance the segmentation performance of the SAM2 visual segmentation algorithm and thus improve the recognition accuracy of the crystallized main body region.

[0071] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0072] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A crystallization anomaly recognition method based on unsupervised learning, characterized in that, Includes the following steps: Step 1: Set up several camera positions, acquire the crystal video stream data of each camera position, and extract and parse the frames to obtain multi-camera crystal frame images; Step 2: Cut each image of the multi-camera crystal frame image into several sub-blocks; Step 3: Use an unsupervised learning algorithm to identify the main crystallization region of each sub-block in Step 2; Step four: Extract features from the main crystallization region to obtain the main crystallization features of each sub-block; Step 5: Compare the similarity between the main crystal features of each sub-block, and take the minimum similarity as the overall similarity. Step six: Compare the overall similarity with the adaptive threshold to determine whether there is any crystallization anomaly; In step three, the unsupervised learning algorithm uses the SAM2 visual segmentation algorithm; In step three, the network modules of the SAM2 visual segmentation algorithm include: image encoder, memory attention, cue encoder, mask decoder, memory encoder, and memory bank; In step three, the network of the SAM2 visual segmentation algorithm is improved, including: parameter freezing, low-rank adaptation, mixed precision training, loss function improvement, evaluation metric improvement, and dataset improvement. In step four, the DINOv2 feature extraction model is used to extract features from the main crystallization region; Parameter freezing includes: retaining the SAM2 image encoder, fine-tuning only the decoder and cue encoder, and unfreezing the image encoder and the last few layers of the decoder; Low-rank adaptation includes inserting a low-rank matrix into the decoder to reduce the number of trainable parameters; Mixed-precision training includes: using FP16 to accelerate training and reduce GPU memory usage; The loss function improvement includes: using the Dice loss function + BCE loss function; Improvements to the evaluation metrics include: using Dice coefficient, IoU (Intersection over Union) ratio, and recall rate in occluded scenes as evaluation metrics; The dataset improvements include: using the crystal body as the segmentation object and performing mask annotation to obtain the crystal body segmentation dataset for fine-tuning; and adopting a strategy of alternating training with video data and static images. In step six, the adaptive threshold is calculated using an adaptive thresholding method based on local statistics; The adaptive threshold for the subsequent N+1 frames is predicted based on the overall similarity results of the first N frames of the video. The calculation formula is as follows: , wherein, denotes the adaptive threshold value corresponding to the sliding window centered at point , denotes the mean of the overall similarity within the sliding window centered at point , reflecting the local average level, denotes the standard deviation of the overall similarity within the sliding window centered at point , reflecting the local fluctuation degree, and C denotes an adjustment factor.

2. The crystallization anomaly recognition method based on unsupervised learning according to claim 1, characterized in that, In step four, the self-supervised learning framework of the DINOv2 feature extraction model adopts a dual-network architecture based on the DINO network and the iBOT network. 3.The crystallization anomaly recognition method based on unsupervised learning according to claim 1, characterized in that, In step four, the loss function of the DINOv2 feature extraction model adopts a multi-objective loss function.

4. The crystallization anomaly identification method based on unsupervised learning according to claim 1, characterized in that, In step five, the similarity between the main crystalline features of each sub-block is calculated using cosine similarity.

Citation Information

Patent Citations

  • Adipic acid reaction chamber crystallization abnormity monitoring system and method thereof

    CN117036261A

  • Abnormal area detection method, automatic optical detection equipment and storage medium

    CN114820561A

  • Crystallization detection method and device and electronic equipment

    CN117058075A