Drug crystal target detection tracking method based on YOLO-SAM2
By improving the YOLO-SAM2 model, the automatic detection and tracking of drug crystals is used by using the feature pyramid network and the all-around convolution module, the problems of uneven light, noise interference and small object detection are solved, and efficient drug crystal recognition and segmentation are achieved.
Patent Information
- Application Number
- CN202510467168.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has problems such as uneven light, noise interference, crystal light transmittance, insufficient training data, and small targets in the early stage of crystal growth in drug crystal recognition and segmentation, resulting in poor recognition effect and low degree of automation.
Using the improved YOLO-SAM2 model, by optimizing the YOLO11 model, feature integration is performed using the feature pyramid network SOEP and the all-around convolution module CSP-OKM, combined with data preprocessing and annotation, automatic detection and tracking of drug crystals is realized.
It improves the recognition and segmentation effect of drug crystals, reduces manual intervention, improves the degree of automation and accuracy of data analysis, and overcomes the problems of low efficiency and poor accuracy of traditional methods.
Smart Images

Figure CN120339584A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of crystal recognition, and particularly to a method for detecting and tracking drug crystal targets based on YOLO-SAM2. Background Art
[0002] Image recognition technology is an important field of artificial intelligence, which refers to the technology of object recognition for images to identify various different patterns of targets and objects. Image segmentation technology goes a step further on the basis of image recognition technology. It can not only identify the targets and objects of different patterns in the picture, but also segment the targets and objects in the picture from the picture. In the field of biopharmaceuticals, the type of drug crystals can be identified by using image recognition technology, and crystal particles can be segmented and screened, the crystallization result can be predicted, and the crystal size can be controlled to improve the crystallization quality. However, at present, the research on the recognition and segmentation of the drug crystallization process in industrial production is still in its infancy. For example, the patent titled "An Image Segmentation Method and System for Crystal Particles" (application number 201810110876.3) uses traditional machine vision algorithms to segment crystals. The two-step Otsu double-threshold segmentation algorithm is used to segment particles, and then the morphological algorithm is used to connect and fill the crystal particles. Its essence is a traditional threshold algorithm, which has high requirements for picture quality and is easily interfered by picture noise; for example, the patent titled "Method, Device, Equipment and Medium for Segmenting and Positioning Crystal Bodies Based on U-Net Structure" (application number 202011623485.1) uses the classical U-Net network structure to segment crystal bodies, without achieving the recognition of crystal categories. At the same time, due to the shortcomings of the U-Net structure itself, the network depth is not enough, resulting in mediocre performance in multi-classification. At the same time, the convolutional neural network (CNN) also has the problem of a small receptive field; for example, the patent titled "A Method for Identifying and Segmenting Drug Crystal Types Based on Vision Transformer" (application number CN202210659832), Transformer may perform poorly when dealing with small targets because the self-attention mechanism tends to model global features and ignores local details.
[0003] When using the traditional machine vision method to identify and segment drug crystals, the following problems will occur when processing crystal particle information:
[0004] 1. The brightness of the photo taken is uneven, resulting in bright and dark areas on the picture, which affects the recognition effect of the algorithm.
[0005] 2. Drug crystal particles have light transmissivity, which is manifested as the brightness of the light-transmitting part of the crystal in the crystal picture being not much different from the background light, resulting in the inability to achieve complete segmentation of the crystal.
[0006] 3. The noise of crystal images is large, which is mostly related to the imaging quality of microscopic equipment. Turbid solutions will also have a certain impact, making the crystal boundaries unclear.
[0007] When using traditional artificial intelligence methods to identify and segment drug crystals, the following problems will occur when processing crystal particle information:
[0008] 1. It is difficult to obtain training pictures of drug crystals, and insufficient training data leads to inaccurate prediction results.
[0009] 2. The target is small in the initial stage of crystal growth, and it is impossible to detect and track it in time. It can only be compensated by manual intervention.
[0010] 3. It is impossible to automatically analyze crystal growth data. The type, quantity, and change process may make the data too crowded, affecting the user's analysis efficiency. Summary of the Invention
[0011] The purpose of the present invention is to solve the problems in the prior art, and propose a drug crystal target detection and tracking method based on YOLO-SAM2. This method can not only identify the types of drug crystals, but also has remarkable segmentation effect, reduces the need for manual intervention, improves the automation degree and accuracy of data analysis, and has high practical value.
[0012] To achieve the above purpose, the present invention proposes a drug crystal target detection and tracking method based on YOLO-SAM2, including the following steps:
[0013] S1. Collect the original samples of drug crystals and preprocess them, and manually annotate to make a dataset as the training data of the model.
[0014] S2. Model optimization and training. Optimize the YOLO11 model, use the improved feature pyramid network (SOEP) to improve the small target detection performance of the model. Input the training data in step S1 into the improved YOLO11-SOEP model for training, and save the weight file obtained from the training.
[0015] S3. Crystal target detection and tracking. Input the video to be detected and recognized into the program with the encapsulated model, use the improved YOLO11-SOEP to obtain the bounding box prediction for crystal target detection as the prompt input of the SAM2 model to track crystal growth, and automatically obtain and save crystal growth data.
[0016] S4. Data analysis. Analyze the crystal growth data according to the results of detection and tracking.
[0017] Preferably, the crystal preprocessing and dataset production in step S1 specifically include the following steps:
[0018] S1.1 Collect the original samples of drug crystallization. Record the information of drug crystals and solvents, conduct heating dissolution and cooling crystallization experiments on the drug, use a microscope digital photography system to photograph the drug crystallization process, and obtain images and videos of crystal growth. The images can be automatically photographed at regular intervals in the crystallization experiment, and can also be intercepted by extracting frames from the photographed crystallization video. Select the data with obvious crystallization process and clear boundaries from the obtained crystallization data.
[0019] The data preprocessing in S1.2 includes processes of solving uneven illumination, image denoising, and enhancing clarity. The multi-scale Retinex algorithm combined with the color restoration mechanism is used to solve the problem of uneven illumination distribution. The denoising method uses the BilateralFilter algorithm, i.e., the bilateral filtering algorithm, to remove the noise in the image; the contrast-limited adaptive histogram equalization method (CLAHE) is used to enhance the clarity of the image and can also suppress noise.
[0020] S1.3 Make a dataset. Label the processed images. Here, ISAT_with_segment_anything is used as the image annotation tool. This tool supports the segmentation of target masks by the SAM series of models. Here, the SAM2 model is selected for loading and annotation. Add the crystal category before annotation, click on the target for segmentation according to the crystal category, and then use this software to export the dataset format required for YOLO model training to generate the corresponding TXT file.
[0021] Preferably, the optimization of the YOLO11-SOEP model in step S2 specifically includes the following steps:
[0022] S2.1 The P2 feature layer of the YOLO11 model is passed through the spatial-to-depth convolution SPD-Conv to obtain features rich in small target information and given to P3 for fusion.
[0023] S2.2 Use the cross-stage local CSP idea and the all-round convolutional module (omni-kernel module) based on AAAI-24 for improvement to obtain CSP-OKM for feature integration. The CSP-OKM module consists of three branches, including a global branch, a large-scale branch, and a local branch, to effectively learn the feature representation from global to local.
[0024] Preferably, the specific steps for crystal target detection, classification, and tracking in step S3 are as follows:
[0025] S3.1 The program loads the weight file trained by the improved model YOLO11-SOEP in step S2 and selects a SAM2 model with a suitable size according to the system environment.
[0026] S3.2 Input the crystallization video to be detected and recognized into the program. First, YOLO11-SOEP detects and classifies the crystals in the first frame of the video, and uses the obtained bounding box prediction as the prompt input for the SAM2 model. The sparse prompt is represented by adding the positional encoding to the learned embedding of each prompt type, while the mask uses convolutional embedding and adds it to the frame embedding to reduce the need for manual selection. As a result, a segmentation mask of the growth process of the crystals recognized in the entire video is obtained, and the obtained classification results and crystal growth data are recorded and saved.
[0027] Preferably, the data analysis in step S4 specifically includes the following steps:
[0028] S4.1 Perform data analysis on the segmentation mask obtained in step S3. Use the find-Contours function in OpenCV to determine the boundary of the target, then use the arclength function to calculate the perimeter (in pixels), and finally multiply the pixel perimeter by the scale size to obtain the true perimeter of the target.
[0029] S4.2 Calculate the number of pixel values of the target mask and multiply it by the actual area of each pixel to obtain the area value of the target.
[0030] Advantages of the present invention: Improving the YOLO11 model enhances the small target detection performance. Compared with traditional methods, adding the P2 detection layer causes problems such as increased computational complexity and more time-consuming post-processing. The P2 feature layer is obtained through spatial-to-depth convolution SPD-Conv to get features rich in small target information and fused with P3, and then the CSP-OKM improved by using the cross-stage local CSP idea and the omnipotent convolution module OKM is used for feature integration, which improves the problem that small targets in the crystal growth process cannot be detected. By using the bounding box prediction of the crystal by YOLO as the prompt input for the SAM2 model, the need for manual intervention is reduced, and the goal of automatic detection, classification, tracking, and data analysis of crystals is achieved. This method overcomes the problems of low efficiency and poor accuracy of traditional methods and has high practical value.
[0031] The features and advantages of the present invention will be described in detail through embodiments in conjunction with the accompanying drawings. Description of the Drawings
[0032] Figure 1 is a schematic flow chart of a method for detecting and tracking drug crystal targets based on YOLO-SAM2 of the present invention;
[0033] Figure 2 is the improved YOLO11-SOEP model in the method for detecting and tracking drug crystal targets based on YOLO-SAM2 of the present invention.
[0034] Figure 3It is the model structure diagram of a drug crystal target detection and tracking method based on YOLO-SAM2 of the present invention. Specific embodiments
[0035] The present invention proposes a drug crystal target detection and tracking method based on YOLO-SAM2, optimizes the YOLO11 model, uses an improved feature pyramid network (SOEP), passes the P2 feature layer of the model through spatial-to-depth convolution SPD-Conv to obtain features rich in small target information and gives them to P3 for fusion, and then uses the cross-stage local CSP idea and the omnipotent convolution module OKM to improve and obtain CSP-OKM for feature integration, so as to improve the small target detection performance of the model; at the same time, uses the improved YOLO11-SOEP to detect crystal targets, and the obtained bounding box predictions are used as the prompt input of the SAM2 model to track crystal growth, and the segmentation mask of the entire crystallization process is obtained, so as to analyze crystal growth, which is specifically realized in the following ways.
[0036] Refer to Figure 1 , according to a drug crystal target detection and tracking method based on YOLO-SAM2 of the present invention, specifically implement the identification and segmentation of crystal particle types during the crystallization process, which specifically includes the following steps:
[0037] Step 1.1: Collect the original samples of drug crystallization. Record the information of drug crystals and solvents, and use a microscope digital photography system to photograph the drug crystallization process to obtain images and videos of crystal growth.
[0038] Step 1.2: The data preprocessing process includes solving uneven illumination, image denoising, and enhancing clarity. The multi-scale Retinex algorithm combined with a color restoration mechanism is used to solve the problem of uneven illumination distribution, and the denoising method uses the BilateralFilter, that is, the bilateral filtering algorithm to remove the noise in the image; the contrast-limited adaptive histogram equalization method (CLAHE) is used to enhance the clarity of the image.
[0039] Step 1.3: Make a dataset. Use ISAT_with_segment_anything as an image annotation tool, select to load the SAM2 model for annotation. Add crystal categories before annotation, click on the target for segmentation according to the crystal categories, and then use this software to export the dataset format required for YOLO model training to generate the corresponding TXT file.
[0040] Step 2.1: Model optimization and training. Optimize the YOLO11 model, use an improved feature pyramid network (SOEP) to improve the small target detection performance of the model, and its specific implementation steps are as follows:
[0041] Step 2.2: The P2 feature layer of the YOLO11 model is passed through a spatial-to-depth convolution SPD-Conv to obtain features rich in small object information and given to P3 for fusion.
[0042] Step 2.3: Use the cross-stage partial CSP idea and the all-round convolution module (omni-kernel module) based on AAAI-24 for improvement to obtain CSP-OKM for feature integration. The CSP-OKM module consists of three branches, including a global branch, a large-scale branch, and a local branch, for effectively learning feature representations from global to local.
[0043] Step 2.4: Input the training data in Step S1.3 into the improved YOLO11-SOEP model for training, and save the obtained weight file.
[0044] Step 3.1: Input the video to be detected and recognized into the automated crystal detection, classification, and tracking program encapsulated with the model. The specific steps of using this program are as follows:
[0045] Step 3.2: The program loads the weight file obtained by training the improved model YOLO11-SOEP in Step S2 and selects a SAM2 model of appropriate size according to the system environment.
[0046] Step 3.3: Input the crystallization video to be detected and recognized into this program. First, YOLO11-SOEP detects and classifies the crystals in the first frame of the video, and uses the obtained bounding box prediction as the prompt input for the SAM2 model. The sparse prompt is represented by adding the positional encoding to the learned embedding of each prompt type, while the mask uses convolutional embedding and adds it to the frame embedding, reducing the need for manual selection. As a result, a segmentation mask of the growth process of the crystals recognized in the entire video is obtained, and the obtained classification results and crystal growth data are recorded and saved.
[0047] Step 4: Perform data analysis on the segmentation mask obtained in Step S3.3. Use the find-Contours function in OpenCV to determine the boundary of the target, then use the arclength function to calculate the perimeter (in pixels), and finally multiply the pixel perimeter by the scale size to obtain the true perimeter of the target. Calculate the number of pixel values of the target mask and multiply it by the actual area of each pixel to obtain the area value of the target.
[0048] The above embodiments are illustrative of the present invention, not limiting of the present invention. Any solution obtained by simply transforming the present invention falls within the protection scope of the present invention.
Claims
1. A drug crystal target detection and tracking method based on YOLO-SAM2, characterized in that: It includes the following steps: S1. Collect the original samples of drug crystals and preprocess them, and manually label and make a dataset as the training data for the model; S2. Model optimization and training; Optimize the YOLO11 model, use the improved feature pyramid network (SOEP) to enhance the small object detection performance of the model; Input the training data in step S1 into the improved YOLO11-SOEP model for training, and save the weight file obtained from the training; S3. Crystal object detection and tracking; Input the video to be detected and recognized into the program with the encapsulated model, use the improved YOLO11-SOEP for crystal object detection, and use the obtained bounding box prediction as the prompt input for the SAM2 model to track crystal growth, automatically obtain crystal growth data and save it; S4. Data analysis; Analyze the crystal growth data according to the results of detection and tracking.
2. The method for detecting and tracking drug crystal targets based on YOLO-SAM2 according to claim 1, characterized in that: The crystal preprocessing and dataset making in step S1 specifically include the following steps: S1.1 Collect the original samples of drug crystallization; Record the information of drug crystals and solvents, conduct a heating and dissolving and cooling crystallization experiment on the drug, use a microscope digital photography system to photograph the drug crystallization process, and obtain the images and videos of crystal growth; The images can be automatically photographed at regular intervals in the crystallization experiment, and can also be intercepted from the captured crystallization videos; Select the data with obvious crystallization process and clear boundaries from the obtained crystallization data; S1.2 The data preprocessing process includes solving uneven illumination, image denoising, and enhancing clarity; Use the multi-scale Retinex algorithm combined with the color restoration mechanism to solve the problem of uneven illumination distribution, and use the BilateralFilter (bilateral filtering algorithm) as the denoising method to remove the noise in the image; Use the contrast-limited adaptive histogram equalization method (CLAHE) to enhance the clarity of the image and suppress noise at the same time; S1.3 Make a dataset; Label the processed images. Here, use ISAT_with_segment_anything as the image annotation tool. This tool supports the segmentation target mask of the SAM series models. Here, select to load the SAM2 model for annotation; Add the crystal category before annotation, click on the target for segmentation according to the crystal category, and then use this software to export the dataset format required for YOLO model training to generate the corresponding TXT file.
3. The drug crystal target detection and tracking method based on YOLO-SAM2 according to claim 1, characterized in that: The optimization of the YOLO11-SOEP model in step S2 specifically includes the following steps: S2.1 Pass the P2 feature layer of the YOLO11 model through the spatial-to-depth convolution SPD-Conv to obtain the features rich in small object information and give them to P3 for fusion; S2.2 Use the cross-stage local CSP idea and the all-round convolutional module (omni-kernel module) based on AAAI-24 for improvement to obtain CSP-OKM for feature integration. The CSP-OKM module consists of three branches, including the global branch, the large-scale branch, and the local branch, which are used to effectively learn the feature representations from global to local.
4. A drug crystal target detection and tracking method based on YOLO-SAM2 according to claim 1, characterized in that: The crystal object detection, classification and tracking in step S3 specifically include the following steps: S3.1 The program loads the weight file trained by the improved model YOLO11-SOEP in step S2, and selects the appropriate-sized SAM2 model according to the system environment; S3.2 Input the crystallization video to be detected and recognized into this program. First, YOLO11-SOEP detects and classifies the crystals in the first frame of the video, and uses the obtained bounding box prediction as the prompt input for the SAM2 model. The sparse prompt is represented by adding the positional encoding to the learned embedding of each prompt type, while the mask uses the convolutional embedding and adds it to the frame embedding to reduce the need for manual selection; as a result, the segmentation mask of the growth process of the crystals recognized in the entire video is obtained, and the obtained classification results and crystal growth data are recorded and saved.
5. A drug crystal target detection and tracking method based on YOLO-SAM2 according to claim 1, characterized in that: The data analysis in step S4 specifically includes the following steps: S4.1 Perform data analysis on the segmentation mask obtained in step S3. Use the find-Contours function in OpenCV to determine the boundary of the target, then use the arclength function to calculate the perimeter (in pixels), and finally multiply the pixel perimeter by the scale size to obtain the true perimeter of the target; S4.2 Calculate the number of pixel values of the target mask and multiply it by the actual area of each pixel to obtain the area value of the target.
Citation Information
Patent Citations
Microwave cooking equipment, microwave heating control method and storage medium
CN108337758A
Crystal segmentation and positioning method and device based on U-net structure, equipment and medium
CN112766313A
Drug crystal type identification and segmentation method based on visual Transform
CN115187778A
Cited By
Clam body size character measuring method and system based on improved YOLO and SAM2
CN122368797A