Method, device and system for classifying interventional pneumology cytopathology images
Patent Information
- Application Number
- CN202610955329.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
AI Technical Summary
然而,现有的图像分类方法大多依赖于单一的深度学习模型进行病理图像分类,难以充分利用病理图像中所蕴含的多维度信息,导致特征提取片面,从而影响分类精度和泛化能力
[0011]上述介入呼吸病学细胞病理学图像的分类方法、装置及系统,通过对每张待检测图像的统计特征进行特征提取,得到第一特征子集,将第一特征子集输入训练好的第一专家模型进行初级分类,得到待检测图像的第一分类概率;对每张待检测图像的纹理特征进行特征提取,得到第二特征子集,将第二特征子集输入训练好的第二专家模型进行初级分类,得到待检测图像的第二分类概率;对每张待检测图像的局部特征进行特征提取,得到第三特征子集,并将第三特征子集输入训练好的第三专家模型进行初级分类,得到待检测图像的第三分类概率;结合每张待检测图像的各分类概率以及各专家模型对应的预设融合权重,得到各待检测图像分别对应的各初级分类概率;各初级分类概率用于确定待筛查对象隶属于各初级类别的预测概率。本申请通过分别提取待检测图像的统计特征、纹理特征和局部特征,并分别输入对应的专家模型进行独立分类,得到三个不同特征维度的分类概率,再结合各专家模型的预设融合权重进行加权融合,实现了对不同类型特征的有效整合与协同决策,避免了单一模型仅依赖单一类型特征进行决策的局限性,使分类结果能够综合反映图像在不同特征空间下的判别信息,从而提升了病理医学图像分类的精度和泛化能力。
Smart Images

Figure CN122821548A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus and system for classifying interventional respiratory cytopathology images. Background Technology
[0002] Respiratory diseases are among the most common diseases worldwide, posing a serious threat to human health. Rapid and accurate diagnosis is crucial for developing individualized treatment plans. With the rapid development of digital imaging technology and machine learning, AI-based automated evaluation technologies for pathological medical images are emerging. Rapid On-Site Evaluation (ROSE) is an auxiliary assessment method used during interventional respiratory procedures. It involves rapid slide preparation, staining, and microscopic evaluation of acquired specimens, providing real-time feedback on specimen quality and offering preliminary diagnostic opinions. The results can assist clinicians in quickly determining the nature of lesions and optimizing the diagnostic and treatment process.
[0003] With the rapid development of digital imaging technology and machine learning, artificial intelligence-based pathological medical image classification methods have gradually emerged, providing a feasible technical path for the intelligentization of ROSE. However, most existing image classification methods rely on a single deep learning model for pathological image classification, making it difficult to fully utilize the multi-dimensional information contained in pathological images, resulting in one-sided feature extraction and thus affecting classification accuracy and generalization ability. Summary of the Invention
[0004] Therefore, it is necessary to provide a classification method, device, and system for interventional respiratory cytopathology images that can improve the accuracy and generalization ability of medical image classification, addressing the aforementioned technical problems.
[0005] Firstly, this application provides a method for classifying interventional respiratory cytopathology images, including: In response to an auxiliary classification request for the object to be screened, multiple images of the object to be screened are acquired. The statistical features of each image to be detected are extracted to obtain a first feature subset. The first feature subset is then input into a trained first expert model for primary classification to obtain the first classification probability of the image to be detected. Texture features are extracted from each image to be detected to obtain a second feature subset. The second feature subset is then input into a trained second expert model for primary classification to obtain the second classification probability of the image to be detected. Feature extraction is performed on the local features of each image to be detected to obtain a third feature subset, and the third feature subset is input into a trained third expert model for primary classification to obtain the third classification probability of the image to be detected. By combining the classification probabilities of each image to be detected with the preset fusion weights corresponding to each expert model, the primary classification probabilities corresponding to each image to be detected are obtained; each primary classification probability is used to determine the predicted probability that the object to be screened belongs to each primary category.
[0006] Secondly, this application also provides a classification device for interventional respiratory cytopathology images, comprising: The acquisition module is used to acquire multiple images of the object to be screened in response to an auxiliary classification request for the object to be screened. The first module is used to extract statistical features from each of the images to be detected to obtain a first feature subset, and input the first feature subset into a trained first expert model for primary classification to obtain the first classification probability of the image to be detected. The second module is used to extract texture features from each of the images to be detected, obtain a second feature subset, and input the second feature subset into a trained second expert model for primary classification to obtain the second classification probability of the image to be detected. The third module is used to extract features from the local features of each image to be detected, obtain a third feature subset, and input the third feature subset into the trained third expert model for primary classification to obtain the third classification probability of the image to be detected. The prediction module is used to combine the classification probabilities of each image to be detected with the preset fusion weights corresponding to each expert model to obtain the primary classification probabilities corresponding to each image to be detected; the primary classification probabilities are used to determine the predicted probability that the object to be screened belongs to each primary category.
[0007] Thirdly, this application also provides a classification system for interventional respiratory cytopathology images, including: An image acquisition device is used to acquire multiple images of the object to be screened. A processor, connected to the image acquisition device, is used to perform the steps of the classification method for interventional respiratory cytopathology images provided in the first aspect of this application.
[0008] Fourthly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the classification method for interventional respiratory cytopathology images provided in the first aspect of this application.
[0009] Fifthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the classification method for interventional respiratory cytopathology images provided in the first aspect of this application.
[0010] Sixthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the classification method for interventional respiratory cytopathology images provided in the first aspect of this application.
[0011] The aforementioned classification method, device, and system for interventional respiratory cytopathology images involves: extracting features from the statistical features of each image to be tested to obtain a first feature subset; inputting the first feature subset into a trained first expert model for primary classification to obtain a first classification probability of the image to be tested; extracting features from the texture features of each image to be tested to obtain a second feature subset; inputting the second feature subset into a trained second expert model for primary classification to obtain a second classification probability of the image to be tested; extracting features from the local features of each image to be tested to obtain a third feature subset; inputting the third feature subset into a trained third expert model for primary classification to obtain a third classification probability of the image to be tested; and combining the classification probabilities of each image to be tested with the preset fusion weights corresponding to each expert model to obtain the primary classification probabilities corresponding to each image to be tested. Each primary classification probability is used to determine the predicted probability of the subject to be screened belonging to each primary category. This application extracts statistical features, texture features, and local features of the image to be detected separately, and inputs them into the corresponding expert models for independent classification to obtain classification probabilities of three different feature dimensions. Then, it combines the preset fusion weights of each expert model for weighted fusion, realizing the effective integration and collaborative decision-making of different types of features. This avoids the limitation of a single model relying on only a single type of feature for decision-making, and enables the classification results to comprehensively reflect the discriminative information of the image in different feature spaces, thereby improving the accuracy and generalization ability of pathological and medical image classification.
[0012] Meanwhile, this application can comprehensively determine the predicted probability of the screening object belonging to each primary category by integrating the primary classification probabilities of each image to be detected, realizing a comprehensive evaluation from single image level classification to object level. It can make full use of multiple image information of the same screening object, and assist in judging the overall status of the object by integrating the classification results of multiple images. It effectively avoids classification fluctuations caused by sampling bias or local abnormalities in individual images, improves the stability and reliability of auxiliary evaluation results, and better meets the actual needs of clinical auxiliary decision-making. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a diagram illustrating the application environment of a classification method for interventional respiratory cytopathology images in one embodiment. Figure 2 This is a flowchart illustrating a classification method for interventional respiratory cytopathology images in one embodiment; Figure 3 This is a schematic diagram of the process for creating observation samples in one embodiment; Figure 4 This is a flowchart illustrating a classification method for interventional respiratory cytopathology images in another embodiment; Figure 5 This is a flowchart illustrating a classification method for interventional respiratory cytopathology images in another embodiment; Figure 6 This is a flowchart illustrating a classification method for interventional respiratory cytopathology images in another embodiment; Figure 7 This is a schematic diagram of the region of interest extraction results in one embodiment; Figure 8 This is a schematic diagram showing the before-and-after comparison of the super-resolution reconstruction effect in one embodiment; Figure 9 This is a structural block diagram of a classification device for interventional respiratory cytopathology images in one embodiment; Figure 10 This is a schematic diagram of the homepage of a classification system for interventional respiratory cytopathology images in one embodiment; Figure 11 This is a schematic diagram of the page showing the operation of the classification module in a classification system for interventional respiratory cytopathology images in one embodiment; Figure 12 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0016] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0017] The classification method for interventional respiratory cytopathology images provided in this application embodiment can be applied to, for example... Figure 1 The application environment shown. The image acquisition device 102 can acquire multiple images of the object to be screened, which can be medical pathological images of the object, especially interventional respiratory cytopathological images. After receiving an auxiliary classification request for an object to be screened, computer device 104 can acquire multiple images to be detected through image acquisition device 102. It then extracts features from the statistical features of each image to obtain a first feature subset. Using the first feature subset and a trained first expert model, it performs primary classification to obtain the first classification probability of the image to be detected. Next, it extracts features from the texture features of each image to obtain a second feature subset. Using the second feature subset and a trained second expert model, it performs primary classification to obtain the second classification probability of the image to be detected. Finally, it extracts features from the local features of each image to obtain a third feature subset. Using the third feature subset as input to a trained third expert model, it performs primary classification to obtain the third classification probability of the image to be detected. Finally, by combining the classification probabilities of each image with the preset fusion weights corresponding to each expert model, it obtains the primary classification probabilities of each image to be detected, thereby determining the predicted probability of the object to be screened belonging to each primary category.
[0018] The computer device 104 can be a terminal or a server. A terminal can be, but is not limited to, various personal computers, laptops, smartphones, and tablets. A server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0019] The image acquisition device 102 can be a microscope system equipped with an image acquisition module. This microscope system may include an optical microscope body and a digital camera connected to the optical microscope body, used to acquire images of cell smear samples at a preset magnification. Optionally, the image acquisition device 102 can be a digital microscope with integrated camera functionality, or a conventional optical microscope that acquires images via an external camera. The images to be tested acquired by the image acquisition device 102 can be stored on the local storage medium of the computer device 104 or in cloud storage for subsequent retrieval and processing.
[0020] In one exemplary embodiment, such as Figure 2 As shown, a classification method for interventional respiratory cytopathology images is provided, which can be applied to... Figure 1 The following steps are used as an example of computer equipment, including steps 201 to 205. Wherein: Step 201: In response to the auxiliary classification request for the object to be screened, acquire multiple images of the object to be screened.
[0021] The individuals to be screened can refer to those requiring medical image-assisted classification, such as patients undergoing interventional pulmonology examinations. The images to be examined can refer to medical images of the individuals to be screened acquired through an image acquisition device, such as cytological smear images taken under a microscope.
[0022] Optionally, the image to be detected can be an interventional respiratory cytopathology image corresponding to the subject to be screened. Alternatively, the image to be detected can be a preprocessed image of the cytopathology image. The computer device can load the color cytopathology image, apply Gaussian filtering to denoise the cytopathology image to suppress noise interference introduced during image acquisition, and convert the denoised color image into a grayscale image, which is then identified as the image to be detected.
[0023] Optionally, an auxiliary classification request may refer to an instruction initiated by a user (such as a doctor or operator) through a computer device to classify the cytopathological images of the object to be screened. The request carries the identification information of the object to be screened and the corresponding storage path or image data of the image to be detected.
[0024] Optionally, the computer device can receive an auxiliary classification request triggered by a user through an interactive interface, which includes identification information of the object to be screened. Based on this identification information, the computer device reads multiple images associated with the object to be screened from local storage or cloud storage.
[0025] Optionally, the computer equipment can communicate directly with the image acquisition device (such as a microscope system equipped with a digital camera) and, after the user triggers an auxiliary classification request, control the image acquisition device in real time to take pictures of the cell smears of the object to be screened and receive multiple images of the object to be tested.
[0026] Step 202: Extract statistical features from each image to be detected to obtain a first feature subset. Input the first feature subset into the trained first expert model for primary classification to obtain the first classification probability of the image to be detected.
[0027] Statistical features refer to mathematical characteristics used to describe the distribution of image pixel values, such as the mean, standard deviation, maximum, and minimum values of image pixels, which reflect the grayscale distribution characteristics of the image. The first classification probability represents the probability that the image to be detected belongs to each primary category.
[0028] Optionally, the primary category for interventional respiratory cytopathology images may include a tumor category and a non-tumor category. The first classification probability includes the probability that the image to be detected, output by the first expert model, belongs to the tumor category and the probability that it belongs to the non-tumor category.
[0029] Optionally, the first expert model can refer to a classification model trained on statistical features, whose input is the statistical features of the image, and whose output is the classification probability that the image belongs to a certain category (such as tumor category or non-tumor category). The first expert model can be a classification model based on support vector machine (SVM), an ensemble learning model based on random forest, or a classification model based on logistic regression.
[0030] Step 203: Extract texture features from each image to be detected to obtain a second feature subset. Input the second feature subset into the trained second expert model for primary classification to obtain the second classification probability of the image to be detected.
[0031] Texture features refer to features used to describe the spatial variation patterns of pixel grayscale values in an image, such as local binary pattern features or texture features based on the gray-level co-occurrence matrix. Texture features can reflect microscopic structural information such as cell morphology and arrangement in the image. The second classification probability includes the probability that the image to be detected belongs to each primary category, as output by the second expert model.
[0032] Optionally, the second expert model can refer to a classification model trained on texture features, whose input is the texture feature vector of the image, and whose output is the classification probability that the image belongs to a certain category. The second expert model can be a classification model based on support vector machine (SVM), an ensemble learning model based on random forest, or a classification model based on logistic regression.
[0033] Step 204: Extract features from the local features of each image to be detected to obtain a third feature subset, and input the third feature subset into the trained third expert model for primary classification to obtain the third classification probability of the image to be detected.
[0034] Local features can refer to informational features used to describe key points or salient structures within local regions of an image, such as Scale Invariant Feature Transform (SIFT) features or Speed-Up Robust Features (SURF). Local features can characterize discriminative local structural information in an image. The third classification probability includes the probability that the image to be detected belongs to each primary category, as output by the third expert model.
[0035] Optionally, the third expert model can be a classification model trained on local features, whose input is the local feature vector of the image, and whose output is the classification probability that the image belongs to a certain category. This third expert model can be a classification model based on Support Vector Machine (SVM), an ensemble learning model based on Random Forest, or a classification model based on Logistic Regression.
[0036] Step 205: Combine the classification probabilities of each image to be detected with the preset fusion weights of each expert model to obtain the primary classification probabilities corresponding to each image to be detected.
[0037] The primary classification probabilities are used to determine the predicted probability that the object to be screened belongs to each primary category. For a given image to be detected, the primary classification probability of the image to be detected includes the probability that the image belongs to each primary category after aggregating the classification probabilities output by the three expert models.
[0038] Optionally, the preset fusion weights corresponding to each expert model can refer to the weight coefficients pre-assigned to each expert model, which are used to reflect the degree of contribution of different expert models in the primary classification decision.
[0039] For example, the computer device acquires the first classification probability, the second classification probability, and the third classification probability corresponding to each image to be detected, as well as the preset fusion weights corresponding to each expert model. For each image to be detected, the computer device multiplies the probability of belonging to a certain primary category in each classification probability by the corresponding preset fusion weight, and then sums the weighted probability values to obtain the probability that the image to be detected belongs to a certain primary category, so as to obtain the primary classification probability that the image to be detected belongs to each primary category.
[0040] For example, the primary categories can include tumor and non-tumor. The first classification probability for a given image is [0.85, 0.15], meaning the first expert model outputs a probability of 0.85 for tumor and 0.15 for non-tumor. Similarly, the second classification probability is [0.75, 0.25], and the third classification probability is [0.80, 0.20]. The preset fusion weights for the first, second, and third expert models are 0.4, 0.3, and 0.3, respectively. Therefore, in the weighted fusion of the primary classification probabilities for this image, the primary classification probability for tumor is 0.805 (0.85×0.4 + 0.75×0.3 + 0.80×0.3); and the primary classification probability for non-tumor is 0.195 (0.15×0.4 + 0.25×0.3 + 0.20×0.3). The computer device performs the aforementioned fusion operation on each image to be detected, obtaining the primary classification probabilities corresponding to each image. Further, the computer device can combine the primary classification probabilities of all images to be detected for the object to be screened, for example, by taking the average of the primary classification probabilities of all images belonging to a certain primary category, thereby determining the predicted probability of the object to be screened belonging to each primary category, and outputting this as an auxiliary classification result.
[0041] In the aforementioned classification method for interventional respiratory cytopathology images, the statistical features of each image to be tested are extracted to obtain a first feature subset. This first feature subset is then input into a trained first expert model for primary classification to obtain the first classification probability of the image to be tested. The texture features of each image to be tested are extracted to obtain a second feature subset. This second feature subset is then input into a trained second expert model for primary classification to obtain the second classification probability of the image to be tested. The local features of each image to be tested are extracted to obtain a third feature subset. This third feature subset is then input into a trained third expert model for primary classification to obtain the third classification probability of the image to be tested. Combining the classification probabilities of each image to be tested with the preset fusion weights corresponding to each expert model, the primary classification probabilities corresponding to each image to be tested are obtained. Each primary classification probability is used to determine the predicted probability that the object to be screened belongs to each primary category. This application embodiment extracts statistical features, texture features, and local features of the image to be detected, and inputs them into the corresponding expert models for independent classification, obtaining classification probabilities of three different feature dimensions. Then, it combines the preset fusion weights of each expert model for weighted fusion, realizing the effective integration and collaborative decision-making of different types of features. This avoids the limitation of a single model relying on only a single type of feature for decision-making, and enables the classification results to comprehensively reflect the discriminative information of the image in different feature spaces, thereby improving the accuracy and generalization ability of cytopathological image classification.
[0042] Meanwhile, the embodiments of this application can comprehensively determine the predicted probability of the screening object belonging to each primary category by integrating the primary classification probabilities of each image to be detected, realizing a comprehensive evaluation from the classification at the single image level to the object level. It can make full use of the image information of multiple images of the same screening object, and assist in judging the overall status of the object by integrating the classification results of multiple images. It effectively avoids the classification fluctuation caused by sampling deviation or local abnormalities in individual images, improves the stability and reliability of the auxiliary evaluation results, and better meets the actual needs of clinical auxiliary decision-making.
[0043] Understandably, before classifying interventional respiratory cytopathology images, it is necessary to acquire the images to be tested or training sample images using an image acquisition device. In practical applications, such as... Figure 3 As shown, after sample preparation processes such as material collection, slide preparation, and staining, cell smears are examined at a preset magnification, and representative fields of view are photographed to obtain image data for subsequent classification processing. This image acquisition process can be applied to the acquisition of sample images during the training phase, as well as the acquisition of images to be detected during the actual classification phase. In some embodiments, both training sample images and images to be detected are acquired under the same imaging conditions to ensure consistency in data distribution between the model training and inference phases. During the above image acquisition process, a low-magnification objective lens (such as a 10× objective lens) is typically used to scan the slides and acquire images of the lesion area, in order to reduce the difficulty of data acquisition and storage costs while ensuring the field of view.
[0044] Optionally, during the sampling process, the data collection techniques used in the embodiments of this application include transbronchial lung biopsy (TBLB), transbronchial mucosal biopsy (EBB), and transbronchial needle aspiration (TBNA). Transbronchial lung biopsy is suitable for sampling parenchymal or peripheral lung lesions; transbronchial mucosal biopsy is for superficial airway mucosal lesions (such as inflammation or tumors); and transbronchial needle aspiration is mainly used for puncture sampling of mediastinal or hilar lymphadenopathy and submucosal lesions.
[0045] Optionally, the film preparation process may include pretreatment and coating, as well as fixation.
[0046] During pretreatment and smear preparation, for specimens obtained by clamping or puncture, the tissue specimen should be pressed against the slide using the tip of a No. 7 syringe needle for concentric smear preparation. For needle aspiration specimens, the specimen should first be pushed out using the core of the puncture needle, then a 20mL empty syringe should be connected, and air should be quickly injected to expel most of the specimen into the cell preservation solution. Then, the puncture needle should be moved to the slide, and air should be injected again to expel the trace amount of liquid specimen remaining in the puncture needle tube. The droplet should be rolled horizontally using the axis of the syringe needle. If the specimen is a viscous droplet, a separate slide can be used, placed horizontally on top of the supporting slide, and the specimen gently flattened before being pulled horizontally apart. The smearing process should be completed in one continuous motion in the same direction to avoid back-and-forth compression.
[0047] Fixation is a cell treatment technique used to prevent cell autolysis or degradation by bacteria or fungi. The fixation process employs air-drying fixation; after the specimen smear is prepared, it is placed in a well-ventilated area to accelerate drying.
[0048] Optionally, the materials used in the staining process employ the Diff-Quick staining method. The Diff-Quick staining kit contains two bottles of staining solutions, A and B. Solution A mainly contains eosin and methanol, while solution B mainly contains methylene blue. During staining, after air-drying the cell slides, immerse them in solution A for approximately 30 seconds, then rinse with phosphate-buffered saline solution and blot dry with a paper towel. Next, immerse them in solution B for approximately 40 seconds, rinse with phosphate-buffered saline solution, and blot dry with a paper towel. Staining is then complete.
[0049] Optionally, during slide review and imaging, first quickly scan the slides under a 10× low-power objective lens, following a wave-like route, scanning the entire specimen area from left to right and from bottom to top to avoid omissions. If a suspicious area is found during the scan, center the target area in the field of view, then switch to a 40× medium-power objective lens, fine-tuning the focus until the field of view is clear. Use a digital camera with an image acquisition system to record and save what is seen under low and medium power. Images are acquired according to patient level, with approximately 3 to 6 images acquired for each patient, saved in a folder named after the patient's information, and then the folder is stored in the corresponding patient category folder.
[0050] In practical applications, before using the trained expert models to classify the images to be detected, the embodiments of this application can first construct and train each expert model and determine the preset fusion weights corresponding to each expert model.
[0051] In one exemplary embodiment, such as Figure 4 As shown, another method for classifying interventional respiratory cytopathology images is provided, comprising steps 401 to 408. Wherein: Step 401: Obtain the training sample set, which includes multiple training sample images of multiple training sample objects and the primary category labels corresponding to each training sample object.
[0052] The primary category labels include tumor labels and non-tumor labels.
[0053] Optionally, the training sample objects can refer to patients with known diagnoses, and the corresponding multiple training sample images are cell smear images acquired during pathological examinations in interventional pulmonology. These training sample images can be lesion area images acquired at the first magnification (e.g., the low magnification corresponding to a 10× objective lens). The primary category label refers to the pathological category corresponding to the patient, such as a tumor label or a non-tumor label. The training sample set is used to train each expert model, enabling the model to learn the feature distribution patterns of images of different categories.
[0054] For example, each patient with a known diagnosis has a known primary category label (tumor or non-tumor), and all training sample images for that patient inherit this label. During data organization, granulomatous inflammation and non-granulomatous inflammation samples are merged into the non-tumor category, while tumor samples remain in the tumor category. Based on the total number of images, the tumor and non-tumor datasets are divided into training, validation, and test sets for each patient, ensuring that images from the same patient do not cross sets. For example, if a patient has 5 training sample images and their primary category label is tumor, then all 5 images are labeled tumor and belong to the same subset (training, validation, or test set), without being scattered across different subsets.
[0055] Optionally, the training sample images can be preprocessed before subsequent feature extraction. The computer loads the color training sample images, applies Gaussian filtering to denoise the images to suppress noise interference introduced during image acquisition, and converts the denoised color images into grayscale images. The converted grayscale images are then used as the training sample images.
[0056] Step 402: Extract features from each training sample image to obtain the first sample feature subset, the second sample feature subset, and the third sample feature subset.
[0057] The first sample feature subset includes statistical features, the second sample feature subset includes texture features, and the third sample feature subset includes local features.
[0058] For example, statistical features may include at least one of the following: the mean value reflecting the overall brightness of the image, the maximum value representing the brightest pixel value, the minimum value representing the darkest pixel value, and the standard deviation measuring the image contrast.
[0059] For example, texture features may include at least one of Local Binary Pattern (LBP) features and Haralick texture features based on a Gray-Level Co-occurrence Matrix (GLCM). A computer device can describe LBP features by generating binary codes by comparing the gray values of each pixel with those of its neighboring pixels. Alternatively, a computer device can describe Haralick texture features based on a GLCM by statistically analyzing the co-occurrence probabilities of gray values between pixels at specific distances and orientations in an image, extracting six attributes: contrast, dissimilarity, homogeneity, energy, correlation, and angular second moment (ASM).
[0060] For example, local features may include KAZE features, which use a KAZE detector to extract key points in an image and calculate corresponding descriptors. The local structural information of the image is characterized by statistically summarizing the descriptors (such as calculating the mean and standard deviation).
[0061] Optionally, the computer device can extract features from the statistical features, texture features, and local features of each training sample image to obtain an initial first sample feature subset, an initial second sample feature subset, and an initial third sample feature subset. Then, all the extracted features are standardized to make each feature have the same dimension, so as to improve the stability of subsequent model training, and thus obtain the first sample feature subset, the second sample feature subset, and the third sample feature subset.
[0062] Step 403: Train the first expert model based on the first sample feature subset and primary category label of each training sample image to obtain the trained first expert model.
[0063] The first expert model can be used to learn the mapping relationship between statistical features and primary categories. Optionally, the first expert model can use a support vector machine as the base classifier, which maps low-dimensional features to a high-dimensional space through a kernel function to handle linearly inseparable problems. During training, probability prediction is enabled to output classification confidence, and the class weights are set to a balanced mode to handle class imbalance problems. Optionally, the first expert model can also use an ensemble learning model based on random forest or a classification model based on logistic regression.
[0064] For example, the computer device acquires a first subset of sample features and corresponding primary class labels for each training sample image. Before training, all first subsets of sample features are standardized to ensure that all features are on the same scale, thereby improving the stability of model training. The computer device trains the first expert model using a Support Vector Machine (SVM) algorithm, employing a Radial Basis Function (RBF) kernel. A regularization parameter C is set to control the penalty for misclassification, the kernel function type is specified as RBF, and a random seed is set to ensure the repeatability of training results. The regularization parameter C and the kernel function parameter gamma are optimized using grid search, and the model performance is evaluated based on the validation set accuracy. During training, the probability prediction function is enabled to allow the model to output classification confidence, and the class weights are set to a balanced mode to alleviate class imbalance. After parameter optimization, the trained first expert model is obtained.
[0065] Step 404: Train the second expert model based on the second sample feature subset and primary category label of each training sample image to obtain the trained second expert model.
[0066] The second expert model is used to learn the mapping relationship between texture features and primary categories. The second expert model can adopt the same model structure as the first expert model, such as using a support vector machine as the base classifier and enabling probability prediction and class weight balancing.
[0067] For example, the computer device acquires a second subset of sample features and corresponding primary category labels for each training sample image. Using this second subset of features as input, the second expert model is trained using the same support vector machine algorithm and training parameters (including radial basis function kernel, regularization parameter C, kernel function parameter gamma, random seed, etc.) as the first expert model. During training, probability prediction is enabled, and the category weights are set to balanced mode. After training is complete, the trained second expert model is obtained.
[0068] Step 405: Train a third expert model based on the third sample feature subset and primary category labels of each training sample image to obtain a trained third expert model.
[0069] The third expert model is used to learn the mapping relationship between local features and primary categories. The third expert model can adopt the same model structure as the first expert model, such as using a support vector machine as the base classifier and enabling probability prediction and class weight balancing.
[0070] For example, the computer device acquires a third sample feature subset and its corresponding primary category label from each training sample image. Using this third sample feature subset as input, a third expert model is trained using the same support vector machine algorithm and training parameters as the first expert model. During training, probability prediction is enabled, and the category weights are set to a balanced mode. After training is complete, the trained third expert model is obtained.
[0071] Step 406: Use the validation sample set to determine the classification accuracy of each trained expert model.
[0072] Among them, classification accuracy refers to the proportion of the number of validation samples correctly classified by each expert model on the validation sample set to the total number of validation samples, which is used to reflect the classification ability of each expert model for different types of features.
[0073] For example, the computer device acquires the first sample feature subset, the second sample feature subset, and the third sample feature subset corresponding to each validation sample image in the validation sample set, and inputs them into the trained first expert model, second expert model, and third expert model, respectively. Each expert model outputs the predicted probability of each validation sample image belonging to each primary category. The computer device compares the primary category corresponding to the maximum predicted probability with the true primary category label of the validation sample image, counts the number of validation samples that match, and divides each by the total number of validation samples to obtain the classification accuracy corresponding to the first expert model, the second expert model, and the third expert model.
[0074] Step 407: Determine the preset fusion weights corresponding to each trained expert model based on the classification accuracy.
[0075] The classification accuracy is directly proportional to the corresponding preset fusion weights.
[0076] Step 408: Use the trained first expert model, second expert model and third expert model to perform primary classification on multiple images to be detected, and obtain the primary classification probabilities corresponding to each image to be detected.
[0077] Understandably, the specific implementation method for performing primary classification on the image to be detected can be referred to the aforementioned methods. Figure 2 This is an example of a computer device that can also generate a classification report and confusion matrix for each image to be detected, used to evaluate the model's classification performance at the image level.
[0078] In this embodiment, a training sample set containing training sample objects and corresponding primary category labels is obtained. Statistical features, texture features, and local features are extracted from each training sample image. Three independent expert models are trained based on each feature subset and label, enabling each model to focus on learning category discrimination information of different feature dimensions. The classification accuracy of each expert model is then determined using a validation sample set, and the preset fusion weight of each expert model is dynamically determined accordingly. This allows the expert model with stronger classification ability to occupy a larger proportion in subsequent fusion decisions, thereby improving the accuracy and robustness of classification decisions.
[0079] Understandably, primary categories include tumor and non-tumor categories. In some cases, when the predicted probability of a candidate belonging to the non-tumor category after primary classification is high, a more granular classification can be performed to further distinguish different subcategories within the non-tumor category. For example, the non-tumor category can be further subdivided into granulomatous inflammation and non-granulomatous inflammation categories.
[0080] In one exemplary embodiment, such as Figure 5 As shown, another method for classifying interventional respiratory cytopathology images is provided, comprising steps 501 to 504. Wherein: Step 501: Aggregate the primary classification probabilities of all images to be detected of the object to be screened to obtain the predicted probability of the object to be screened belonging to each primary category.
[0081] For example, the computer device acquires the primary classification probabilities of all images of the object to be screened. For each primary category, it calculates the average of the primary classification probabilities of all images of the object belonging to that primary category, and uses the average as the predicted probability that the object belongs to that primary category. For instance, if there are 5 images of the object to be screened, and the primary classification probabilities of each image belonging to the non-tumor category are 0.80, 0.85, 0.78, 0.82, and 0.90, respectively, then the average value of 0.83 is taken as the predicted probability that the object belongs to the non-tumor category.
[0082] Step 502: If the predicted probability that the subject to be screened belongs to a non-tumor category is greater than the first preset probability threshold, the region of interest of the lesion in multiple images to be detected is extracted to obtain multiple images of the lesion center.
[0083] Please refer to Figure 6 and Figure 7When the predicted probability that the subject to be screened belongs to a non-tumor category is greater than a first preset probability threshold (e.g., 0.7), the computer device converts each image to be detected to a preset color space to obtain an image after color space conversion. The brightness channel and saturation channel are extracted from the image after color space conversion. The saturation channel is segmented using the Otsu's method to generate a saturation mask. The brightness channel is segmented using a fixed threshold to generate a brightness mask. The saturation mask and the brightness mask are bitwise ANDed to obtain a candidate lesion region mask. The candidate lesion region mask is subjected to morphological opening operation to obtain a denoised mask. Connected component analysis is performed on the denoised mask, and the connected component with the largest area is selected as the target lesion region. The geometric center of the target lesion region is calculated, and the image to be detected is cropped according to a preset size based on the geometric center to obtain the lesion center image.
[0084] For example, the image to be detected can be a lesion region image acquired at a first magnification (e.g., the low magnification corresponding to a 10× objective lens). The computer equipment can convert each image to the HLS color space (Hue, Lightness, Saturation) and extract the luminance and saturation channels respectively. For the saturation channel, the Otsu's method (OTSU) is used to automatically calculate the threshold and segment it, generating a saturation mask to suppress low-saturation background areas. For the luminance channel, a fixed threshold range (e.g., 80 to 220) is set to generate a luminance mask to exclude necrotic dark areas (L<80) and blank backgrounds (L>220). A bitwise AND operation is performed between the saturation mask and the luminance mask to obtain a candidate lesion region mask. A morphological opening operation (using a 5×5 structuring element) is performed on the candidate lesion region mask to remove isolated noise and retain clustered cells. All external connected components are obtained through contour detection, sorted in descending order of area, and the connected component with the largest area is selected as the target lesion region. The geometric center coordinates of the target lesion region are calculated using image moments and used as the lesion center seed point. Based on the lesion center seed point, the image is cropped at a preset ratio according to the original image size (e.g., width and height are each one-quarter of the original image), and the cropping boundary is restricted to ensure that the cropped area is always within the effective range of the image. This cropped image is then used to obtain the lesion center image. This process is repeated for each image to be detected, resulting in multiple lesion center images. The cropped lesion center images can be saved according to the naming convention "original filename_roi.jpg" and stored in the corresponding result subfolder according to the folder structure corresponding to the object to be screened, forming low-resolution lesion center image samples for subsequent super-resolution reconstruction.
[0085] Step 503: Perform super-resolution reconstruction on multiple lesion center images to obtain multiple lesion reconstructed images.
[0086] Please refer to Figure 6 and Figure 8 Computer equipment can pre-build a blind super-resolution model and obtain a reconstructed sample set. The blind super-resolution model includes a noise correction module and a resolution restoration module. The reconstructed sample set includes high-resolution sample images, real low-resolution sample images, and clean low-resolution sample images.
[0087] The image to be detected is a lesion area image acquired at a first magnification, the high-resolution sample image is a sample image of the lesion area acquired at a second magnification, the second magnification being greater than the first magnification; the real low-resolution sample image is a cropped image of the lesion center from the sample image of the lesion area acquired at the first magnification, and the clean low-resolution sample image is obtained by downsampling the high-resolution sample image.
[0088] The computer device can use the reconstructed sample set to jointly train the blind super-resolution model, enabling the noise correction module to learn the mapping relationship between the real low-resolution sample image and the clean low-resolution sample image, and enabling the resolution restoration module to learn the mapping relationship between the clean low-resolution sample image and the high-resolution sample image.
[0089] The computer equipment can input multiple images of the lesion center into a trained blind super-resolution model, and then perform super-resolution reconstruction through the noise correction module and the resolution restoration module in sequence to obtain multiple reconstructed images of the lesion.
[0090] For example, the image to be detected is an image of the lesion area acquired at a low magnification corresponding to a 10× objective lens, the high-resolution sample image is a sample image of the lesion area acquired at a medium magnification corresponding to a 40× objective lens, the real low-resolution sample image is a cropped image of the lesion center from a sample image of the lesion area acquired at a low magnification corresponding to a 10× objective lens, and the clean low-resolution sample image is obtained from the high-resolution sample image through bicubic downsampling.
[0091] Computer devices can pre-build classification-driven unpaired blind super-resolution (CUBSR) models. Based on the CinCGAN framework, the blind super-resolution model introduces an intermediate clean low-resolution domain, decomposing the blind super-resolution problem into two sub-problems: noise correction and resolution restoration.
[0092] Optionally, the blind super-resolution model may include a noise correction module and a resolution recovery module.
[0093] The noise correction module includes a first generator. Second generator and the first discriminator This forms the first loop structure. The first generator converts a noisy, real low-resolution image (real LR) into a clean low-resolution image (clean LR), the second generator converts the clean low-resolution image (clean LR) back into a noisy, real low-resolution image (real LR), and the first discriminator... It is used to distinguish clean low-resolution images from real low-resolution images after noise correction by the second generator, and generates clean images that are closer to the real distribution through adversarial training.
[0094] The resolution restoration module includes a third generator. Fourth generator Second discriminator This forms the second loop structure. The third generator... This is an upsampling generator that employs an improved version of the Residual Feature Distillation Network (RFDN) with intermediate residual removal and addition operations. A third generator is used to convert a clean low-resolution image (clean LR) into a high-resolution reconstructed image (SR), a fourth generator is used to convert the high-resolution reconstructed image (SR) into a true low-resolution image (true LR), and a second discriminator... It is used to distinguish between real high-resolution images and high-resolution reconstructed images generated by a third generator, and generates high-resolution images with higher visual quality and more realistic details through adversarial training.
[0095] The blind super-resolution model also includes a task-driven module, which can employ a classifier (such as MobileNetV2) to output global semantic features and class logistic values to compute the classification-driven loss.
[0096] During training, the computer equipment jointly trains the blind super-resolution model using the reconstructed sample set. Optionally, the computer equipment constructs a data loader to read reconstructed samples from the patient folder, including high-resolution sample images (images under medium magnification), real low-resolution sample images (ROIs at the center of lesions under low magnification), and clean low-resolution sample images (obtained by bicubic downsampling of medium magnification images). Label files are automatically generated, and the training and test sets are divided into training and test sets at a 6:4 ratio based on the patient. A progressive training strategy is adopted in the early stages of training, gradually increasing the classification-driven loss weights through a gating function to avoid early task supervision damaging the image structure.
[0097] Specifically, the computer device inputs real low-resolution sample images into the first generator. Get the first generator The output is a pseudo-clean low-resolution image; the pseudo-clean low-resolution image is then input into the second generator. Get the second generator The output is a first reconstructed low-resolution image; the computer device also inputs a clean low-resolution sample image into a second generator. Get the second generator The first pseudo-realistic low-resolution image output.
[0098] The computer device calculates a first cycle consistency loss based on the difference between the real low-resolution sample image and the first reconstructed low-resolution image, making... and During the conversion process, the content information of the original image is preserved as much as possible; based on the difference between the real low-resolution sample image and the pseudo-clean low-resolution image, the first identity mapping loss (forward identity loss) is calculated; based on the difference between the clean low-resolution sample image and the first pseudo-real low-resolution image, the second identity mapping loss (reverse identity loss) is calculated to constrain... and Preserve image identity features during the conversion process.
[0099] The computer device inputs the real low-resolution sample image and the first reconstructed low-resolution image into the first discriminator. Through the first discriminator Distinguish between real low-resolution sample images and the first reconstructed low-resolution image to drive the second generator. Generate outputs that more closely resemble real low-resolution sample images, and enable the first discriminator Continuously improve the discrimination ability. Based on the first discriminator The output results are used to calculate the first adversarial loss (e.g., using an adversarial loss function such as hinge loss or least squares loss), thereby optimizing the performance of the noise correction module in the adversarial game between the generator and the discriminator.
[0100] For the resolution restoration module, the computer device inputs a clean, low-resolution sample image into the third generator. Obtain the third generator Output a pseudo-high-resolution image; input the pseudo-high-resolution image into the fourth generator. Obtain the fourth generator The output is a second reconstructed low-resolution image. The computer calculates a second cycle consistency loss based on the difference between the clean low-resolution sample image and the second reconstructed low-resolution image, to constrain the third and fourth generators to preserve as much of the original image's content information as possible during the transformation process.
[0101] The computer equipment also inputs clean, low-resolution sample images into the fourth generator. Obtain the fourth generator The output second pseudo-real low-resolution image is used to calculate the third identity mapping loss based on the difference between the clean low-resolution sample image and the second pseudo-real low-resolution image, so as to constrain the fourth generator to maintain the identity features of the image during the transformation process.
[0102] The computer device inputs the pseudo-high-resolution image and the high-resolution sample image into the second discriminator. Through the second discriminator Distinguish between high-resolution sample images and pseudo-high-resolution images to drive the third generator. Generate output that is closer to the high-resolution sample image, and enable the second discriminator Continuously improve discrimination capabilities. Based on the second discriminator. The output results are used to calculate the second adversarial loss, thereby optimizing the performance of the resolution restoration module in the adversarial game between the generator and the discriminator.
[0103] The computer device inputs a pseudo-high-resolution image into the task-driven module to extract its global semantic features and predicted logistic values; it also inputs a high-resolution sample image into the task-driven module to extract its global semantic features. The computer device calculates the semantic feature alignment loss based on the mean square error of the feature centers between the global semantic features of the pseudo-high-resolution image and the global semantic features of the high-resolution sample image; and it calculates the classification consistency loss based on the cross-entropy between the predicted logistic values of the pseudo-high-resolution image and the class labels of the real low-resolution sample image.
[0104] Finally, the computer device weights and sums the first cycle consistency loss, first identity mapping loss, second identity mapping loss, first adversarial loss, second cycle consistency loss, third identity mapping loss, second adversarial loss, and classification-driven loss (including semantic feature alignment loss and classification consistency loss) according to preset weights to obtain the overall loss. The optimizer then updates the parameters of each generator and discriminator to reduce the loss. After each training round, image-level and object-level classification accuracy, macro-average F1 score, and domain-level reconstruction quality metrics such as Learned Perceptual Image Patch Similarity (LPIPS) and Kernel Inception Distance (KID) are evaluated on the test set. The optimal weight model is saved based on the evaluation metrics. After training, a trained blind super-resolution model is obtained.
[0105] In the model application phase, the computer loads a pre-trained blind super-resolution model and inputs multiple images of lesion centers (i.e., low-resolution lesion center images). For each lesion center image, the computer converts it into a tensor and normalizes it, then inputs it into the pre-trained blind super-resolution model. The model then sequentially passes through a noise correction module (which converts noisy real low-resolution images into clean low-resolution images using a first generator) and a resolution restoration module (which upsamples the clean low-resolution images using a third generator to generate high-resolution images) for super-resolution reconstruction, generating high-resolution reconstructed lesion images. The computer saves the reconstructed lesion images to the corresponding subfolder of the objects to be screened, retaining their original filenames, thus creating high-quality images that can be directly used for subsequent subdivision and classification, completing the super-resolution reconstruction task.
[0106] Step 504: Input multiple lesion reconstruction images into a pre-trained subdivision classification model and output the subdivision classification probability of each lesion reconstruction image belonging to the granulomatous inflammation category and the non-granulomatous inflammation category.
[0107] Among them, the sub-classification probabilities of each lesion reconstruction image are used to determine the predicted probability that the subject to be screened belongs to each sub-class.
[0108] The subdivision classification model is a lightweight convolutional neural network based on depthwise separable convolutions. The subdivision classification model includes a backbone network, a global average pooling layer, and a fully connected layer.
[0109] For example, the computer device can input the reconstructed lesion image into the backbone network to extract deep features, obtaining a deep feature map; the deep feature map is then input into a global average pooling layer for global average pooling, obtaining a compressed feature vector; the compressed feature vector is then input into a fully connected layer, which maps the feature vector to sub-classification probabilities belonging to the granulomatous inflammation category and the non-granulomatous inflammation category, respectively. The computer device can then aggregate the sub-classification probabilities of all reconstructed lesion images of the subject to be screened to obtain the predicted probability of the subject belonging to each sub-class.
[0110] For example, a subdivision classification model can use a lightweight MobileNetV2 network structure as the backbone encoder, extracting semantic features of the image through a series of depthwise separable convolutional layers. The backbone output is then connected to a global average pooling layer for dimensionality reduction, compressing the spatial dimension of each feature channel into a single value. A classification head (such as a fully connected layer) is then constructed to map the extracted image features to a category distribution, outputting the final predicted probability of belonging to each sub-category.
[0111] Before training the subdivision classification model, the computer device can acquire a second training sample set. This second training sample set includes multiple images of various training sample objects and subdivision category labels corresponding to each training sample object. The subdivision category labels include granulomatous inflammation labels and non-granulomatous inflammation labels. The second training sample images are sample images of lesion regions acquired at a second magnification (e.g., a medium magnification corresponding to a 40× objective lens). The computer device performs data augmentation on the second training sample images, including random cropping and random horizontal flipping, to improve the model's generalization ability; it also performs standardization on the augmented images to unify the input distribution; and it converts the images and labels into tensor form and encapsulates them into batch data for subsequent training.
[0112] During training, the computer system divides the granulomatous inflammation and non-granulomatous inflammation datasets into training, validation, and test sets according to preset ratios. A stochastic gradient descent (SGD) optimizer is used, with hyperparameters such as learning rate, momentum, and weight decay set, and a learning rate scheduling strategy (e.g., StepLrUpdater, with a decay factor of 0.98 per round) is constructed. The batch size is configured to 32, the number of data loading threads to 4, and the number of training rounds to 100. The accuracy and loss curves for both the training and validation sets are recorded. During training, the weights with the minimum loss during training and the weights with the highest accuracy during validation are saved for subsequent testing. After training, a well-trained, finely segmented classification model is obtained.
[0113] In the model application phase, the computer device inputs the multiple lesion reconstruction images obtained in step 503 into the trained sub-classification model. The backbone network extracts features layer by layer from each lesion reconstruction image through a series of depthwise separable convolutional layers, obtaining the depth feature map of the image. The computer device inputs the depth feature map into a global average pooling layer, performing global average pooling on the depth feature map in the spatial dimension, that is, averaging the feature values at all spatial locations in each feature channel to obtain a compressed feature vector. The computer device inputs the compressed feature vector into a fully connected layer, which maps the feature vector to the sub-classification probabilities of granulomatous inflammation and non-granulomatous inflammation, outputting the predicted probability of the lesion reconstruction image belonging to each sub-class. The computer device aggregates (e.g., averages) the sub-classification probabilities of all lesion reconstruction images of the subject to be screened to obtain the predicted probability of the subject belonging to each sub-class.
[0114] In this embodiment, a subsequent subdivision process is triggered when the predicted probability of the subject belonging to a non-tumor category exceeds a preset threshold. This process extracts and reconstructs the central region of the lesion using super-resolution, and then inputs the reconstructed high-resolution image into the subdivision classification model. This achieves a progressive classification process from coarse-grained primary classification to fine-grained subclass classification. Based on the primary classification results, this method performs super-resolution reconstruction on the central image of the lesion acquired under low magnification, restoring the low-resolution image to a high-resolution image. This enhances details such as cell morphology and texture, avoiding the data acquisition difficulties and computational overhead associated with acquiring high-resolution images for the target image, while still obtaining high-quality detail information for samples requiring further differentiation. Furthermore, this embodiment utilizes an automatic region of interest extraction method based on traditional image processing, eliminating the need for doctors to manually label the lesion location beforehand. This effectively reduces data labeling costs and improves the model's scalability and applicability. Furthermore, by aggregating (e.g., averaging) the subdivision classification probabilities of multiple lesion reconstruction images of the subject to be screened, the predicted probability of the subject belonging to each subdivision category is determined. This achieves comprehensive auxiliary assessment from single image level classification to object level, avoiding classification fluctuations caused by single image sampling bias or local anomalies, and improving the stability and reliability of the auxiliary assessment results.
[0115] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0116] Based on the same inventive concept, this application also provides a classification device for interventional respiratory cytopathology images to implement the classification method for interventional respiratory cytopathology images described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations of one or more embodiments of the classification device for interventional respiratory cytopathology images provided below can be found in the limitations of the classification method for interventional respiratory cytopathology images described above, and will not be repeated here.
[0117] In one exemplary embodiment, such as Figure 9 As shown, a classification device for interventional respiratory cytopathology images is provided, comprising: an acquisition module 901, a first module 902, a second module 903, a third module 904, and a prediction module 905, wherein: The acquisition module 901 is used to acquire multiple images of the object to be screened in response to an auxiliary classification request for the object to be screened.
[0118] The first module 902 is used to extract features from the statistical features of each image to be detected, obtain a first feature subset, input the first feature subset into a trained first expert model for primary classification, and obtain the first classification probability of the image to be detected.
[0119] The second module 903 is used to extract texture features from each of the images to be detected, obtain a second feature subset, and input the second feature subset into a trained second expert model for primary classification to obtain the second classification probability of the image to be detected.
[0120] The third module 904 is used to extract features from the local features of each image to be detected, obtain a third feature subset, and input the third feature subset into the trained third expert model for primary classification to obtain the third classification probability of the image to be detected.
[0121] The prediction module 905 is used to combine the classification probabilities of each of the images to be detected with the preset fusion weights corresponding to each expert model to obtain the primary classification probabilities corresponding to each of the images to be detected; the primary classification probabilities are used to determine the predicted probability that the object to be screened belongs to each primary category.
[0122] Each module in the aforementioned classification device for interventional respiratory cytopathology images can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0123] In one exemplary embodiment, a classification system for interventional respiratory cytopathology images is provided. The system includes an image acquisition device and a processor. The image acquisition device is used to acquire multiple images of an object to be screened. The processor is connected to the image acquisition device and is used to perform the steps of the classification method for interventional respiratory cytopathology images provided in the first aspect of the present application.
[0124] Optionally, the classification system for interventional respiratory cytopathology images also includes a display device for displaying the interactive interface of the classification system.
[0125] For example, the interactive interface can be a visual interactive interface built on the PyQt framework, integrating a login and registration module, an image import module, a classification processing module, and a result display module. The login and registration module verifies user identity, supports new user account registration and login, ensuring data security and user management. The image import module allows users to select image folders of a single or multiple objects to be screened through the interface, batch importing lesion area images under low magnification, and automatically identifying and loading multiple images corresponding to each object to be screened. The classification processing module responds to user-triggered classification commands, calling the classification method of interventional respiratory cytopathology images from any of the aforementioned embodiments to process the imported images, and displays the current object to be screened, processing progress, processing time, and processing status in real time. The result display module, after classification, displays the category prediction results (such as the confidence level corresponding to granulomatous inflammation, non-granulomatous inflammation, or tumor) of the current object to be screened in the display area of the interactive interface, and displays the detection records of all tested objects in tabular form. Each record includes a serial number, patient information, confidence level, model used, and processing time. Optionally, the results display module also allows users to export test records to an Excel file with one click, making it easier for doctors to review and archive them later.
[0126] For a specific implementation method, please refer to Figure 10 The interactive interface displays the first page. The homepage navigation bar includes options such as "Homepage" and "Classification Based on Low Magnification." The top of the first page displays the system title, such as "Interventional Pulmonology Intelligent Rapid On-Site Assessment (ROSE) System." The middle of the first page displays a button for accessing the low-magnification classification module. Please refer to [link / reference]. Figure 11 In response to a click on the option button, the interactive interface displays a second page. The left side of the second page is a navigation bar for switching between different functional modules; the middle section is the main operation and display area, including a selection area for the objects to be tested, a thumbnail display area for the images to be tested, a results display area, and a test record table area; the right side of the second page is the operation function area, including a file import button (for selecting one or more objects to be screened), a results display sub-area, and operation buttons for testing, saving, and clearing. Through this visual interactive interface, doctors do not need programming or deep learning expertise; they can complete the entire process from image import to classification result acquisition with simple clicks, effectively lowering the threshold for clinical application and improving the automation and efficiency of rapid on-site evaluation in interventional pulmonology.
[0127] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 12 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and databases. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a classification method for interventional respiratory cytopathology images.
[0128] Those skilled in the art will understand that Figure 12 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0129] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the classification method for interventional respiratory cytopathology images provided in the first aspect of the present application.
[0130] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the classification method for interventional respiratory cytopathology images provided in the first aspect of the present application.
[0131] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the classification method for interventional respiratory cytopathology images provided in the first aspect of the present application.
[0132] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0133] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0134] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0135] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A classification method for interventional respiratory cytopathology images, characterized in that, The method includes: In response to an auxiliary classification request for the object to be screened, multiple images of the object to be screened are acquired. The statistical features of each image to be detected are extracted to obtain a first feature subset. The first feature subset is then input into a trained first expert model for primary classification to obtain the first classification probability of the image to be detected. Texture features are extracted from each image to be detected to obtain a second feature subset. The second feature subset is then input into a trained second expert model for primary classification to obtain the second classification probability of the image to be detected. Feature extraction is performed on the local features of each image to be detected to obtain a third feature subset, and the third feature subset is input into a trained third expert model for primary classification to obtain the third classification probability of the image to be detected. By combining the classification probabilities of each image to be detected with the preset fusion weights corresponding to each expert model, the primary classification probabilities corresponding to each image to be detected are obtained; each primary classification probability is used to determine the predicted probability that the object to be screened belongs to each primary category.
2. The method according to claim 1, characterized in that, Before acquiring multiple images of the object to be screened in response to an auxiliary classification request for the object to be screened, the process includes: Obtain a training sample set, which includes multiple training sample images of multiple training sample objects and primary category labels corresponding to each training sample object; the primary category labels include tumor labels and non-tumor labels. For each of the training sample images, feature extraction is performed to obtain a first sample feature subset, a second sample feature subset, and a third sample feature subset; the first sample feature subset includes statistical features, the second sample feature subset includes texture features, and the third sample feature subset includes local features; A first expert model is trained based on the first sample feature subset and the primary category label of each training sample image to obtain the trained first expert model; The second expert model is trained based on the second sample feature subset of each training sample image and the primary category label to obtain the trained second expert model; The third expert model is trained based on the third sample feature subset and the primary category label of each training sample image to obtain the trained third expert model.
3. The method according to claim 2, characterized in that, After obtaining the trained third expert model, the process includes: The classification accuracy of each trained expert model is determined using the validation sample set. The preset fusion weights corresponding to each trained expert model are determined based on the classification accuracy; the classification accuracy is proportional to the corresponding preset fusion weight.
4. The method according to any one of claims 1 to 3, characterized in that, The primary categories include tumor categories and non-tumor categories, and the method further includes: The primary classification probabilities of all images to be detected for the object to be screened are aggregated to obtain the predicted probability of the object to be screened belonging to each primary category; If the predicted probability that the object to be screened belongs to the non-tumor category is greater than a first preset probability threshold, the region of interest of the lesion in multiple images to be detected is extracted to obtain multiple images of the lesion center. Super-resolution reconstruction was performed on multiple images of the lesion center to obtain multiple reconstructed lesion images; Multiple reconstructed images of the lesions are input into a pre-trained subdivision classification model, which outputs the subdivision classification probability of each reconstructed image of the lesions belonging to the granulomatous inflammation category and the non-granulomatous inflammation category. The subdivision classification probability of each reconstructed image of the lesions is used to determine the predicted probability of the object to be screened belonging to each subdivision category.
5. The method according to claim 4, characterized in that, The step involves extracting regions of interest (ROIs) from multiple images to be detected to obtain multiple images of the lesion centers, including: Each image to be detected is converted to a preset color space to obtain an image after color space conversion; Extract the brightness and saturation channels from the color space converted image. The saturation channel is thresholded using the Otsu's method to generate a saturation mask; The brightness channel is segmented using a fixed threshold to generate a brightness mask; Perform a bitwise AND operation between the saturation mask and the brightness mask to obtain the candidate lesion region mask; A morphological opening operation is performed on the candidate lesion region mask to obtain the denoised mask; Perform connected component analysis on the denoised mask and select the connected component with the largest area as the target lesion region; Calculate the geometric center of the target lesion region, and crop the image to be detected according to a preset size based on the geometric center to obtain the lesion center image.
6. The method according to claim 4, characterized in that, The image to be detected is an image of the lesion area acquired at a first magnification. The super-resolution reconstruction of multiple lesion center images yields multiple reconstructed lesion images, including: A blind super-resolution model is constructed, which includes a noise correction module and a resolution restoration module. A reconstruction sample set is obtained, comprising high-resolution sample images, true low-resolution sample images, and clean low-resolution sample images; wherein, the high-resolution sample images are sample images of the lesion region acquired at a second magnification factor, which is greater than the first magnification factor; the true low-resolution sample images are cropped images of the lesion center from the sample images of the lesion region acquired at the first magnification factor; and the clean low-resolution sample images are obtained by downsampling the high-resolution sample images. The blind super-resolution model is jointly trained using the reconstructed sample set, enabling the noise correction module to learn the mapping relationship between the real low-resolution sample image and the clean low-resolution sample image, and enabling the resolution restoration module to learn the mapping relationship between the clean low-resolution sample image and the high-resolution sample image. Multiple images of the lesion centers are input into a trained blind super-resolution model, and then processed sequentially by the noise correction module and the resolution restoration module to obtain multiple reconstructed images of the lesions.
7. The method according to claim 4, characterized in that, The subdivision classification model is a lightweight convolutional neural network based on depthwise separable convolutions. The subdivision classification model includes a backbone network, global average pooling layers, and fully connected layers. The step of inputting multiple reconstructed lesion images into the pre-trained subdivision classification model and outputting the subdivision classification probability of each reconstructed lesion image belonging to the granulomatous inflammation category and the non-granulomatous inflammation category includes: The reconstructed image of the lesion is input into the backbone network for depth feature extraction to obtain a depth feature map; The deep feature map is input into the global average pooling layer for global average pooling to obtain the compressed feature vector. The compressed feature vector is input into the fully connected layer, which maps the feature vector to subdivision probabilities belonging to the granulomatous inflammation category and the non-granulomatous inflammation category, respectively.
8. A classification device for interventional respiratory cytopathology images, characterized in that, The device includes: The acquisition module is used to acquire multiple images of the object to be screened in response to an auxiliary classification request for the object to be screened. The first module is used to extract statistical features from each of the images to be detected to obtain a first feature subset, and input the first feature subset into a trained first expert model for primary classification to obtain the first classification probability of the image to be detected. The second module is used to extract texture features from each of the images to be detected, obtain a second feature subset, and input the second feature subset into a trained second expert model for primary classification to obtain the second classification probability of the image to be detected. The third module is used to extract features from the local features of each image to be detected, obtain a third feature subset, and input the third feature subset into the trained third expert model for primary classification to obtain the third classification probability of the image to be detected. The prediction module is used to combine the classification probabilities of each of the images to be detected with the preset fusion weights corresponding to each expert model to obtain the primary classification probabilities corresponding to each of the images to be detected; each primary classification probability is used to determine the predicted probability that the object to be screened belongs to each primary category.
9. A classification system for interventional respiratory cytopathology images, characterized in that, The system includes: An image acquisition device is used to acquire multiple images of the object to be screened. A processor, connected to the image acquisition device, is used to perform the steps of the method according to any one of claims 1 to 7.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.