Focus category division method and system, electronic equipment and storage medium

By combining the feature extraction methods of CNN and Transformer, multimodal data is used to divide lesion categories, which solves the problem of inaccurate classification of lesion categories in the prior art, and achieves higher accuracy and generalization ability of lesion category identification.

CN120339177APending Publication Date: 2025-07-18SHENYANG NEUSOFT INTELLIGENT MEDICAL TECH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510295256.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing lesion category classification method mainly relies on CT image data, making it difficult to accurately learn differential characteristics between classes. In addition, existing mixed models are prone to missing global semantic features or local detailed characteristics when dividing lesion categories, resulting in inaccurate lesion category classification.

Method used

The lesion category division model based on multimodal data is adopted, combining the local characteristics of CNN and the global representation of Transformer. By extracting the global characteristics, local characteristics and modal features of other modal data of the CT sequence, and performing feature interaction and fusion processing, the precise division of lesion categories is achieved.

Benefits of technology

By strengthening the closeness of the characteristics of different receptive fields, the accuracy of lesion classification is improved, and the recognition ability and generalization performance of lesion categories are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339177A_ABST
    Figure CN120339177A_ABST
Patent Text Reader

Abstract

The invention provides a lesion category division method and system, electronic equipment and a storage medium, and relates to the technical field of image processing, and the method comprises the steps: obtaining CT sequence data and other modal data corresponding to a lesion region; preprocessing the CT sequence data and the other modal data to obtain a target CT sequence corresponding to the CT sequence data and target modal features corresponding to the other modal data; the target CT sequence and the target modal features are input into a pre-trained lesion category division model, a category division result of the lesion area is obtained, the lesion category division model extracts global features and local features based on the target CT sequence, and the lesion category division result of the lesion area is obtained; and carrying out feature interaction fusion processing on the global feature, the local feature and the target modal feature, and carrying out lesion category division based on the fused multi-modal feature to obtain a category division result of the lesion region. The accuracy of focus category division can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and particularly to a method, a system, an electronic device and a storage medium for lesion category classification. Background Art

[0002] Lesion category classification is of great significance for medical diagnosis and treatment, which can help doctors more accurately judge the condition, formulate treatment plans, evaluate the prognosis and guide follow-up. At the same time, it also helps to improve the treatment effect and quality of life of patients.

[0003] Recently, many studies have been based on deep learning image classification models to improve the accuracy of lesion category classification. However, in existing image classification models, only a small number of CT images are used for judgment, and the model data only includes CT image data. For different types of lesions with similar characteristics, the inter-class differences are small, and it is difficult to accurately learn the inter-class difference features only relying on CT image data. In addition, when existing models perform lesion category classification, they often miss global semantic features or local detail features, and thus it is difficult to ensure the accurate classification of lesion categories. Summary of the Invention

[0004] The present invention provides a method, a system, an electronic device and a storage medium for lesion category classification, which can improve the accuracy of lesion category classification.

[0005] In a first aspect, a method for lesion category classification is provided, including: Obtaining CT sequence data and other modality data corresponding to a lesion area; Preprocessing the CT sequence data and other modality data to obtain a target CT sequence corresponding to the CT sequence data and target modality features corresponding to the other modality data; Inputting the target CT sequence and the target modality features into a pre-trained lesion category classification model to obtain a category classification result of the lesion area, wherein the lesion category classification model extracts global features and local features based on the target CT sequence, performs feature interaction and fusion processing on the global features, local features and target modality features, and performs lesion category classification based on the fused multi-modal features to obtain the category classification result of the lesion area.

[0006] In a second aspect, a system for lesion category classification is provided, including: An obtaining module, configured to obtain CT sequence data and other modality data corresponding to a lesion area; A processing module, configured to preprocess the CT sequence data and other modality data to obtain a target CT sequence corresponding to the CT sequence data and target modality features corresponding to the other modality data; A partitioning module for inputting a target CT sequence and target modality features into a pre-trained lesion category partitioning model to obtain a category partitioning result of the lesion region, where the lesion category partitioning model extracts global features and local features based on the target CT sequence, performs feature interaction and fusion processing on the global features, local features, and target modality features, and performs lesion category partitioning based on the fused multi-modal features to obtain the category partitioning result of the lesion region.

[0007] In a third aspect, an electronic device is provided, including: a processor and a memory, where the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in the first aspect or its various implementation manners.

[0008] In a fourth aspect, a computer-readable storage medium is provided for storing a computer program, and the computer program causes a computer to execute the method in the first aspect or its various implementation manners.

[0009] Through the technical solution provided by the present invention, after obtaining the CT sequence data and other modality data corresponding to the lesion region, the CT sequence data and other modality data can be preprocessed to obtain the target CT sequence corresponding to the CT sequence data and the target modality features corresponding to the other modality data; then the target CT sequence and the target modality features can be input into a pre-trained lesion category partitioning model. In the lesion category partitioning model, global features and local features are extracted based on the target CT sequence, and feature interaction and fusion processing are performed on the global features, local features, and target modality features, and lesion category partitioning is performed based on the fused multi-modal features to obtain the category partitioning result of the lesion region. In the technical solution of this application, by simultaneously extracting the global features, local features in the CT sequence, and the modality features of other modality data, and performing feature interaction and fusion processing on the global features, local features, and target modality features, the feature fusion closeness of different receptive fields can be strengthened, and the accuracy of the features can be improved. Accurate partitioning of the lesion category can be achieved based on the fused multi-modal features.

[0010] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. Other features and advantages of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 It is an application scenario diagram provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of a method for classifying lesion categories provided by an embodiment of the present application; Figure 3 It is a schematic diagram of the model framework of a lesion category classification model provided by an embodiment of the present application; Figure 4 It is a schematic flowchart of a method for classifying lesion categories provided by another embodiment of the present application; Figure 5 It is a schematic diagram of the model structure of a lesion category classification model provided by an embodiment of the present application; Figure 6 It is a schematic diagram of patch splitting provided by an embodiment of the present application; Figure 7 It is a schematic diagram of the convolutional block structure in a local feature extraction module provided by an embodiment of the present application; Figure 8 It is a schematic diagram of the structure of a dual-branch feature fusion module provided by an embodiment of the present application; Figure 9 It is a schematic diagram of the structure of a multi-modal feature fusion module provided by an embodiment of the present application; Figure 10 It is a schematic diagram of the structure of a lesion category classification system provided by an embodiment of the present invention; Figure 11 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0014] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0015] The classification of lesion categories is of great significance for medical diagnosis and treatment. It can help doctors more accurately judge the condition, formulate treatment plans, evaluate the prognosis and guide follow-up. At the same time, it also helps to improve the treatment effect and quality of life of patients.

[0016] The development of artificial intelligence (AI) technology has provided new ideas and methods for the research of lesion category classification. Among them, deep learning (DL) can automatically extract implicit disease diagnosis features from large-scale medical image datasets and has quickly become a research hotspot in medical image analysis. Convolutional neural networks (CNNs) have achieved good performance in medical image classification tasks. However, there are still deficiencies in lesion category classification.

[0017] I. Data limitations 1. Data modality limitations. In existing lesion category classification methods, only CT images are used for judgment, and the data of the model is only image data. For diseases such as pneumonia, different types of pneumonia have similar characteristics, and the characteristic manifestations in the images overlap, such as co-occurring features like multiple nodules, ground-glass opacities, and crescent signs. The differences between classes are small, and it is difficult to accurately learn the inter-class difference features only relying on image data, and it is difficult to accurately learn the heterogeneity of different types of pneumonia from CT images, and thus it is difficult to accurately identify the types of pneumonia.

[0018] 2. Data quantity limitations. Deep learning methods need to learn the feature distribution of data through a large amount of raw data to achieve functions such as classification or segmentation. However, medical images are affected by multiple factors, and it is difficult to expand the data scale. On the one hand, it is difficult for doctors to annotate images, and the workload is large. On the other hand, medical data involves privacy, ethics and other issues, making it difficult to achieve large-scale public access.

[0019] II. Model limitations Most models are supervised and overly reliant on image data. The development of Artificial Intelligence (AI) technology has provided new ideas and methods for medical image research. Among them, Deep Learning (DL) can automatically extract implicit image features from large-scale medical image datasets and has quickly become a research hotspot in medical image analysis. Convolutional Neural Networks (CNN) have achieved good performance in medical image segmentation tasks, such as brain Magnetic Resonance Imaging (MRI) and liver lesion segmentation. However, due to the inherent inductive bias characteristics of CNN, it lacks the ability to perceive global information in limited perceptual fields, resulting in local invariance and translational invariance. Although some methods, such as expanding convolutional kernels, dilated convolutions, and pyramid pooling, have addressed this issue, it has not been fully resolved. Vision Transformer (ViT) is good at using self-attention mechanisms to extract global features, which are superior to local receptive fields, but it lacks the inductive bias and shift invariance of CNN. In addition, fully training a ViT model requires too much labeled data to complete. Due to the persistent lack of high-quality labeled medical images, how to train a ViT model on a small-scale dataset has become an urgent problem to be solved. Therefore, combining the features extracted by the CNN representation model (local features) with the features extracted by the Transformer representation model (global features) is beneficial for lesion detection. However, existing hybrid methods usually use CNN convolutional blocks in the low-level stage and Transformer blocks in the high-level stage, which may result in missing some global semantic information in the low-level stage and missing local detail features in the high-level stage. That is to say, these hybrid strategies cannot meet the requirements of segmentation and detection-intensive visual tasks and are difficult to ensure the accurate classification of lesion categories.

[0020] To solve the above technical problems, the concept of the present invention is to propose a lesion category classification model based on multi-modal data, which combines the local features of CNN with the global representation of Transformer for lesion category classification in CT images. Specifically, by simultaneously extracting the global features, local features, and modal features of other modal data in the CT sequence and performing feature interaction and fusion processing on the global features, local features, and target modal features, the feature fusion closeness of different receptive fields can be strengthened, and the accuracy of the features can be improved. Based on the fused multi-modal features, lesion category classification can be performed, which can greatly improve the accuracy of lesion category classification.

[0021] It should be understood that the technical solution of the present invention can be applied to the following scenarios, but is not limited to: In some implementable ways, Figure 1A scenario diagram provided by an embodiment of the present invention is as follows Figure 1 As shown, this application scenario may include an electronic device 110 and a network device 120. The electronic device 110 may establish a connection with the network device 120 through a wired network or a wireless network.

[0022] Exemplarily, the electronic device 110 may be a desktop computer, a laptop computer, a tablet computer, etc., but is not limited thereto. The network device 120 may be a terminal device or a server, but is not limited thereto. In an embodiment of the present invention, the electronic device 110 may send a request message to the network device 120, and the request message may be used to request to obtain the classification result of the lesion area. Further, the electronic device 110 may receive a response message sent by the network device 120, and the response message includes the classification result of the lesion area.

[0023] It should be noted that the technical solution in this application can be applied to the classification of lesion categories for non-therapeutic purposes. The types of lesions classified can vary according to the actual application scenario, and the types of lesions classified include but are not limited to cancer, pneumonia, gastritis, etc. In the following embodiment steps of this application, the classification of pneumonia lesions is taken as an example to illustrate the technical solution in this application, but it does not constitute a specific limitation.

[0024] After introducing the application scenario of the embodiment of the present invention, the technical solution of the present invention will be elaborated in detail below: Figure 2 A flowchart of a method for classifying lesion categories provided by an embodiment of the present invention, and this method may be executed by an electronic device 110 as shown Figure 1 As shown, but is not limited thereto. As Figure 2 shown, this method may include the following steps: Step 210: Obtain CT sequence data and other modality data corresponding to the lesion area.

[0025] Among them, CT sequence data refers to a series of two-dimensional image data obtained by continuous X-ray slice scanning during CT scanning of a lesion area, and these data can be used to reconstruct the internal structure of the three-dimensional lesion area; other modality data are auxiliary modality data for classifying lesion categories, used to increase the feature dimension, and are used together with the CT sequence data as multi-modal data for classifying lesion categories. The other modality data may include at least one of the following: medical data provided by various medical imaging techniques such as clinical data, Magnetic Resonance Imaging (MRI), Positron Emission Tomography (PET), Ultrasound (US), Single-Photon Emission Computed Tomography (SPECT), and X-ray. It should be noted that the data types included in the other modality data can be determined according to the lesion types that actually need to be classified in a specific application scenario. In the following embodiment steps of the present disclosure, taking the lesion type to be classified as pneumonia and the other modality data including clinical data as an example, the technical solution in the present application is described, but it does not constitute a specific limitation.

[0026] Step 220: Preprocess the CT sequence data and other modality data to obtain a target CT sequence corresponding to the CT sequence data and target modality features corresponding to the other modality data.

[0027] As a possible implementation manner, when preprocessing the CT sequence data, the preprocessing process may include data verification processing and first data preprocessing. By successively performing data verification processing and first data preprocessing on the CT sequence data, a target CT sequence can be obtained.

[0028] Data verification processing may include: validity verification, repeatability verification, integrity verification, etc. Through data verification processing, the correctness and consistency of CT sequence data can be ensured. Validity verification is used to confirm whether the CT sequence data exists and can be accessed and used, and to check whether the CT sequence data is damaged, such as data loss caused by transmission errors, storage medium failures, or other reasons. When performing validity verification on CT sequence data, it can be determined whether the number of sequence layers of the CT sequence data is less than a certain threshold (such as 60), or the slice thickness is greater than a certain threshold (such as 2 mm), etc. Such CT sequence data is usually considered an invalid sample because the size of the lesion is not fixed and may be very small, and the CT sequence data in the above situation will not be able to see the lesion; Integrity verification is used to verify whether the CT sequence data is complete and whether there are any missing or tampered parts. When performing integrity verification on CT sequence data, it can be verified whether there are missing layers in the CT sequence data. The method is to first sort based on the unique sequence code (Instance Number) of each slice, and then check the slice spacing (SpacingBetween Slices) between adjacent slices based on the relative coordinate position (Slice Location) of each slice. If the slice spacing between two adjacent slices is 2 times or more that of other adjacent slice spacings, it means that there is a missing slice between these two slices; Repeatability verification is used to ensure that each frame in the CT sequence data is unique, avoiding repeated calculation or analysis of the same content during processing. When performing repeatability verification on CT sequence data, it can first be sorted based on the unique sequence code (Instance Number) of each slice, and then traverse the sequence codes of all slices. If the sequence code of the current slice is the same as the previous one, there is a repeated slice sequence code, that is, there are repeated slices.

[0029] The first data preprocessing may include: foreground cropping processing and normalization processing. Through image preprocessing, the high availability of data can be ensured, and the generalization ability and robustness of the method can be improved. Through foreground cropping processing, the locking of the range where the features are located, the positioning of specific features, and information extraction can be accelerated, improving the efficiency of image classification. The purpose of normalizing the CT sequence data is to make the pixel values of each pixel point in the CT sequence data be normalized and distributed between 0 and 1, so as to improve the contrast of the lesion area, that is, to make the lesion area more obvious relative to other areas, facilitating the subsequent extraction of image features. Specifically, a pixel value range can be set first, and the pixel value range can be set differently according to different lesion areas, and no specific limitation is made here.

[0030] As a possible implementation, when preprocessing other modality data, the preprocessing process may include second data preprocessing. By performing second data preprocessing on other modality data, the target modality features corresponding to the other modality data can be obtained. The second data preprocessing may include, for example, foreground cropping processing, normalization processing, and feature encoding processing. Through foreground cropping processing, the locking of the feature range, specific feature localization, and information extraction can be accelerated, and the efficiency of lesion category classification can be improved. The purpose of normalizing other modality data is to make the other modality data uniformly distributed between 0 and 1 to improve the contrast of the region of interest, that is, to make the region of interest more obvious relative to other regions, facilitating subsequent extraction of image features. By performing feature encoding processing using a preset feature encoding method, the target modality features corresponding to the other modality data can be obtained. The preset feature encoding method may include, but is not limited to, One-Hot Encoding, Target Encoding, Mean Encoding, etc.

[0031] Step 230: Input the target CT sequence and the target modality features into the pre-trained lesion category classification model to obtain the category classification result of the lesion area.

[0032] Among them, as Figure 3 shown, the lesion category classification model includes a dual-branch feature extraction module 31, a multi-modal feature fusion module 32, and a classifier 33. The dual-branch feature extraction module 31 serves as the input end of the lesion category classification model. The output end of the dual-branch feature extraction module 31 is connected to the input end of the multi-modal feature fusion module 32. The classifier 33 serves as the output end of the lesion category classification model. The input end of the classifier 33 is connected to the output end of the multi-modal feature fusion module 32. After inputting the target CT sequence and the target modality features into the pre-trained lesion category classification model, the dual-branch feature extraction module 31 can extract global features and local features based on the target CT sequence, and fuse the global features and local features to obtain dual-branch fusion features; the multi-modal feature fusion module 32 performs feature interaction and fusion processing on the global features, local features, dual-branch fusion features, and target modality features to obtain multi-modal features; the classifier 33 performs lesion category classification based on the fused multi-modal features to obtain the category classification result of the lesion area.

[0033] When pre-training the lesion category classification model, iterative training is performed on the lesion category classification model using the sample CT sequences with labeled lesion categories and the sample modality features until the loss function of the lesion category classification model reaches a convergent state, and it is determined that the pre-training of the lesion category classification model is completed. When performing iterative training on the lesion category classification model, data verification processing and first data preprocessing can be performed on the sample CT sequences first. Among them, the data verification processing can include availability verification, integrity verification, repeatability verification, and consistency verification. The implementation processes of the availability verification, integrity verification, and repeatability verification are the same as those in step 220 of the embodiment and will not be elaborated here. The consistency verification is used to verify whether the labeled data of the sample CT sequence is consistent with the actual situation of the sample CT sequence. Through the consistency verification, the training accuracy of the lesion category classification model can be ensured, and interference from mislabeled data can be avoided. The first data preprocessing can include foreground cropping processing, normalization processing, building a data dictionary, data loading, annotation file processing, resampling, random cropping, and data augmentation processing, etc. The implementation processes of the foreground cropping processing and normalization processing are the same as those in step 220 of the embodiment and will not be elaborated here. When building the data dictionary, for each sample CT sequence, a mapping relationship between it and the corresponding labeled training label can be created and encapsulated as a list of dictionaries in the Python language dictionary format, which is convenient for verifying the consistency and accuracy of the annotation through the list of dictionaries to ensure that the annotation of each sample CT sequence is correct; when performing data loading, for each sample CT sequence and its corresponding annotation file, a channel dimension can be added in front of its spatial dimension. Adding the channel dimension can be used to represent whether each pixel point or region belongs to multiple categories; when the sample CT sequence is a CT image, resampling is used to resample the CT image and the annotation to unify the anatomical coordinates and voxel spacing so that the lesion category classification model can learn consistent target features of interest. The anatomical coordinate system is a continuous three-dimensional space composed of three planes (transverse plane, coronal plane, sagittal plane), and the anatomical coordinates are used to describe the position of the standard human body anatomically; random cropping is used to randomly crop one or more input blocks (crops) of a fixed size from the corresponding positions of the sample CT sequence and the annotation. On the one hand, it is for batch training of the sample CT sequence. On the other hand, through this random cropping, the lesion area is located at different positions in space, weakening the sensitivity of the lesion category classification model to the target position to improve the generalization ability of the training model. In each iteration cycle during the model training process, the spatial positions of the random cropping are almost different; data augmentation is used to perform image augmentation on the sample CT sequence to ensure the data volume of the training samples. When performing data augmentation, random affine transformations (rotation, scaling, etc.) can be performed on the input blocks after random cropping, also to improve the generalization ability of the training model.

[0034] After performing the above image preprocessing on the sample CT sequence, 20% of the sample CT sequence can be randomly divided into a validation set, and 80% into a training set. The training set is used to continuously optimize and train the lesion category classification model, and the validation set is used to verify the training accuracy of the lesion category classification model during the training process until the lesion category classification model is trained. After verifying that the lesion category classification model reaches a certain accuracy on the validation set, the lesion category classification model can be used to classify the lesion categories based on the test set (i.e., the target CT sequence and the target modal features).

[0035] Correspondingly, when pre-training the lesion category classification model, the steps of the embodiment may include: determining the preprocessed sample CT sequence and sample modal features, as well as the preset training labels corresponding to the sample images, where the preset training labels at least include the classification labels corresponding to the lesion regions; inputting the sample images, sample modal features, and preset training labels into the lesion category classification model to perform lesion category classification training on the lesion category classification model; wherein, during the lesion category classification training process, the sample images and sample modal features are used as input features, and the preset training labels are used as training labels to iteratively update the model parameters in the lesion category classification model until the loss function of the lesion category classification model is less than a preset threshold, and it is determined that the lesion category classification model is trained. The preset threshold is a value between 0 and 1, and the specific value can be set according to the actual application scenario and will not be specifically limited here.

[0036] In summary, according to the lesion category classification method provided by the present invention, after obtaining the CT sequence data and other modal data corresponding to the lesion region, the CT sequence data and other modal data can be preprocessed to obtain the target CT sequence corresponding to the CT sequence data and the target modal features corresponding to the other modal data; then the target CT sequence and the target modal features can be input into the pre-trained lesion category classification model. In the lesion category classification model, global features and local features are extracted based on the target CT sequence, and feature interaction and fusion processing are performed on the global features, local features, and target modal features, and the lesion categories are classified based on the fused multi-modal features to obtain the classification result of the lesion region. In the technical solution of this application, by simultaneously extracting the global features, local features in the CT sequence, and the modal features of other modal data, and performing feature interaction and fusion processing on the global features, local features, and target modal features, the feature fusion closeness of different receptive fields can be strengthened, and the accuracy of the features can be improved. Based on the fused multi-modal features, accurate classification of the lesion categories can be achieved.

[0037] Based on Figure 2 the embodiments shown, as a refinement and extension of the above embodiments, in order to fully illustrate the specific implementation process of the method of this embodiment, this embodiment provides as Figure 4The specific method shown. Figure 4 Based on Figure 2 The embodiment shown. As shown in FIG. 4, the method includes the following steps: Step 410: Obtain the CT sequence data corresponding to the lesion area and other modality data.

[0038] Step 420: Preprocess the CT sequence data and other modality data to obtain the target CT sequence corresponding to the CT sequence data and the target modality features corresponding to the other modality data.

[0039] For the embodiments of the present disclosure, the specific implementation process can refer to the relevant description in step 220 of the embodiment, which will not be elaborated here.

[0040] Step 430: Input the target CT sequence into the dual-branch feature extraction module, extract the target global feature and the target local feature corresponding to the target CT sequence, and perform feature mapping interaction fusion and feature splicing processing on the target global feature and the target local feature to obtain the target dual-branch fusion feature.

[0041] Among them, as Figure 5 shown, the dual-branch feature extraction module 31 includes multiple feature extraction stages. Each feature extraction stage includes a global feature extraction module 311, a local feature extraction module 312, and a dual-branch feature fusion module 313 (i.e., Dual-branch feature fusion module). The number of multiple feature extraction stages can be set according to the actual application scenario. Each feature extraction stage is configured with different numbers of channels, and the number of channels increases layer by layer in multiple feature extraction stages for gradually extracting implicit features at a deeper level. In the following embodiment steps of the present disclosure, taking the dual-branch feature extraction module 31 including 3 feature extraction stages (Stage 1, Stage 2, and Stage 3) as an example, the technical solution in the present application will be described, but it does not constitute a specific limitation.

[0042] Correspondingly, for the embodiments of the present disclosure, step 430 of inputting the target CT sequence into the dual-branch feature extraction module, extracting the target global feature and the target local feature corresponding to the target CT sequence, and performing feature mapping interaction fusion and feature splicing processing on the target global feature and the target local feature to obtain the target dual-branch fusion feature may include the following steps: Step 430-1: For any one of the multiple feature extraction stages, use the globally configured feature extraction module therein to perform global feature extraction on the first input feature, use the locally configured feature extraction module therein to perform local feature extraction on the second input feature, and use the dual-branch feature fusion module configured therein to perform feature mapping interaction fusion on the third input feature, so as to obtain the dual-branch fusion feature of the current feature extraction stage.

[0043] Among them, when the current feature extraction stage is the first feature extraction stage among the multiple feature extraction stages, the first input feature and the second input feature are the target CT sequences, and the third input feature includes the first global feature and the first local feature extracted in the current feature extraction stage; when the current feature extraction stage is any one of the multiple feature extraction stages except the first feature extraction stage, the first input feature is the first global feature extracted in the previous feature extraction stage corresponding to the current feature extraction stage, the second input feature is the first local feature extracted in the previous feature extraction stage corresponding to the current feature extraction stage, the third input feature includes the first global feature and the first local feature extracted in the current feature extraction stage, and the first dual-branch fusion feature output in the previous feature extraction stage corresponding to the current feature extraction stage.

[0044] For example, as Figure 5 shown, when the dual-branch feature extraction module 31 includes 3 feature extraction stages (Stage 1, Stage 2, Stage 3), for the first feature extraction stage Stage 1, the first input feature and the second input feature are the target CT sequences, and the third input feature includes the first global feature extracted by the global feature extraction module 311 and the first local feature extracted by the local feature extraction module 312 in the current feature extraction stage; for the second feature extraction stage Stage 2, the first input feature is the first global feature extracted in the first feature extraction stage Stage 1, the second input feature is the first local feature extracted in the first feature extraction stage Stage 1, and the third input feature includes the first global feature extracted by the global feature extraction module 311, the first local feature extracted by the local feature extraction module 312, and the first dual-branch fusion feature output in the first feature extraction stage Stage 1 in the current feature extraction stage; for the third feature extraction stage Stage 3, the first input feature is the first global feature extracted in the second feature extraction stage Stage 2, the second input feature is the first local feature extracted in the second feature extraction stage Stage 2, and the third input feature includes the first global feature extracted by the global feature extraction module 311, the first local feature extracted by the local feature extraction module 312, and the first dual-branch fusion feature output in the second feature extraction stage Stage 2 in the current feature extraction stage.

[0045] Correspondingly, as Figure 5 shown, the global feature extraction module 311 includes a Swin transformer block (i.e., a Swin transformer block) and a patch merge block (i.e., Patch Merge). Different from the standard multi-head self-attention (Multi-Head Self-Attention, MSA) mechanism of ordinary transformers, the Swin transformer block is based on a self-attention mechanism with a moving window (W-MSA) and a self-attention mechanism based on a shifted window (SW-MSA) using the current token. Compared with ordinary transformers, it greatly reduces the computational complexity while enhancing the ability to model strong correlations between different tokens. The input of the Swin transformer block is a sequence of image patches (a sequence of patches). Before inputting the target CT sequence into the Swin transformer block, patch splitting will be performed, that is, the input image is split into non-overlapping patches, and each is regarded as a patch token (abbreviated as token).

[0046] As Figure 6 shown in the patch splitting schematic diagram, for the schematic diagram, the image is split into patches of 16×(10×10×1), and the size of each token is 10×10×1. In other words, there are tokens. In the schematic diagram, both H and W are equal to 40, and each image 40×40×1 ( ) is processed into (i.e., ) image patches, and each path is flattened into a 100 (10×10×1 = )-dimensional token vector, and the overall flattening is 16×100 = ( )×100-dimensional 2D patch sequence.

[0047] The patch merge block is to generate hierarchical global features with different receptive fields. As the network deepens, the number of tokens is gradually reduced by Patch Merge. The first Patch Merge concatenates each group of 2×2 adjacent patches, and the number of patch tokens becomes 1 / 4 of the original. For example, before the Stage1 processing, it is H / 2×W / 2, and after the Patch Merge in Stage1, it is H / 4×W / 4.

[0048] Correspondingly, for any one of the multiple feature extraction stages, when using the global feature extraction module configured therein to perform global feature extraction on the first input feature, the steps of the embodiment may include: for any one of the multiple feature extraction stages, using the Swin transformer block to extract the self-attention features of each image patch in the first input feature, and using the patch merging block to merge the self-attention features of the image patches to obtain the first global feature extracted by the global feature extraction module in the current feature extraction stage. Then, the first global feature is input into the global feature extraction module 311 of the next Stage and the dual-branch feature fusion module 313 of the current Stage.

[0049] In a specific application scenario, since the regions and sizes of pneumonia lesions vary greatly, if only global feature extraction is used and local features are ignored, it may be difficult to learn the relevant information of small lesions, resulting in ignoring small lesions and causing classification errors. As Figure 5 shown, the local feature extraction module 312 includes a convolutional block and a downsampling module (i.e., Average Pooling). The structure of the convolutional block in the local feature extraction module 312 is as Figure 7 shown. The convolutional block includes a convolutional layer, a DropBlock layer, and a batch normalization and activation layer. The input of the local feature extraction module 312 first passes through a 3×3 convolution, and random block dropout, that is, DropBlock, is performed on the convolved structure to improve the generalization ability and prevent overfitting. Subsequently, batch normalization + ReLU processing is performed. Then the above process is repeated once. The processed result is subjected to an average pooling operation through the downsampling module to achieve downsampling and obtain the first local feature. Then, the first local feature is input into the local feature extraction module 312 of the next Stage and the dual-branch feature fusion module 313 of the current Stage.

[0050] Correspondingly, for any one of the multiple feature extraction stages, when using the local feature extraction module configured therein to perform local feature extraction on the second input feature, the steps of the embodiment may include: inputting the second input feature into the convolutional block and iteratively performing convolutional feature processing until a preset number of iterations is reached to obtain convolutional features, where the convolutional feature processing sequentially includes convolutional processing, random feature block dropout processing, batch normalization processing, and activation processing; using the downsampling module to perform downsampling processing on the convolutional features to obtain the first local feature extracted by the local feature extraction module in the current feature extraction stage.

[0051] After obtaining the output features of the dual-scale encoder (the first global feature and the first local feature), the remaining problem is how to effectively fuse them to form the learning of multi-scale feature representation. A direct method is to simply concatenate the multi-scale features and then perform convolution operations. However, this direct method cannot capture the long-term dependencies and global context connections between features of different scales.

[0052] Therefore, we propose a new dual-branch feature fusion method based on mapping interaction, as Figure 8 shown. The dual-branch feature fusion module 313 utilizes the efficient interaction between features. For the first global feature F1, whose dimension is , it first undergoes convolution and then obtains the same size as the first local feature F2 after convolution . Then, the corresponding mapping feature weights are obtained through pooling operations respectively, both with the size of (1, 1, c). Subsequently, the two mapping feature weights are cross-multiplied with the first local feature and the first global feature respectively to obtain the interacted features, namely the second local feature mapping global information and the second global feature mapping local information. Then, the second local feature and the second global feature are concatenated and convolved to obtain the finally fused dual-branch interaction information. During the concatenation operation, the first feature extraction stage Stage 1 performs the operation in the dotted box a, that is, the fused feature obtained by fusing the second local feature and the second global feature is used as the first dual-branch fusion feature of the first feature extraction stage Stage 1; the second feature extraction stage Stage 2 and the third feature extraction stage Stage 3 perform the operation in the dotted box b, that is, after fusing the second local feature and the second global feature to obtain the fused feature, it is also necessary to concatenate the first dual-branch fusion feature of the previous feature extraction stage corresponding to the current feature extraction stage on the basis of the fused feature to obtain the first dual-branch fusion feature of the current feature extraction stage.

[0053] Correspondingly, for any one of the multiple feature extraction stages, when using the dual-branch feature fusion module configured therein to perform feature mapping interaction fusion on the third input feature to obtain the dual-branch fusion feature of the current feature extraction stage, the steps of the embodiment may include: determining the first mapping feature weight of the first global feature extracted in the current feature extraction stage and the second mapping feature weight of the first local feature extracted in the current feature extraction stage; performing feature cross-multiplication on the first global feature and the first local feature respectively based on the first mapping feature weight and the second mapping feature weight to obtain the second local feature mapping global information and the second global feature mapping local information; performing feature fusion on the second local feature and the second global feature to obtain the second fusion feature; when the current feature extraction stage is the first feature extraction stage among the multiple feature extraction stages, using the second fusion feature as the first dual-branch fusion feature of the current feature extraction stage; or, when the current feature extraction stage is any one of the multiple feature extraction stages other than the first feature extraction stage, splicing the second fusion feature and the first dual-branch fusion feature corresponding to the previous feature extraction stage of the current feature extraction stage to obtain the first dual-branch fusion feature of the current feature extraction stage.

[0054] Step 430-2: Determine the target global feature and the target local feature as the first global feature and the first local feature extracted in the last feature extraction stage among the multiple feature extraction stages, and determine the first dual-branch fusion feature output by the dual-branch feature fusion module configured in the last feature extraction stage as the target dual-branch fusion feature.

[0055] Step 440: In the multi-modal feature fusion module, perform feature fusion processing on the target dual-branch fusion feature, the target global feature, and the target local feature to obtain the first fusion feature, and splice the target modal feature on the first fusion feature to obtain the multi-modal feature.

[0056] As Figure 9 shown, there are multiple inputs to the multi-modal feature fusion module 32, including the target dual-branch fusion feature, the target global feature, the target local feature, and other modal features. In this module, first, perform feature mapping interaction fusion on the target dual-branch fusion feature, the target global feature, and the target local feature input to the module, then perform a flatten operation on the fusion feature to process the fusion feature into a (1×N) feature vector. Secondly, connect the feature vector of the target modal feature (1×n), such as the clinical information feature encoding, etc., with the flattened feature vector, and further strengthen the connection between the multi-modal and the image features through a linear layer to obtain the multi-modal feature output by the multi-modal feature fusion module.

[0057] Correspondingly, for the embodiments of the present disclosure, the embodiment steps may include: performing fusion processing on the target double-branch fusion feature, the target global feature, and the target local feature to obtain a third fusion feature; determining a first feature vector of the third fusion feature and a second feature vector of the target modality feature; concatenating the first feature vector and the second feature vector, and performing linear processing on the concatenated feature vector to obtain the multi-modal feature output by the multi-modal feature fusion module.

[0058] Step 450: Use a classifier to classify the lesion area based on the multi-modal feature to obtain a classification result.

[0059] For the embodiments of the present disclosure, the classifier may output prediction scores for each preset lesion category based on the multi-modal feature, and determine the preset lesion category with the highest corresponding prediction score as the classification result of the lesion area. Among them, the preset lesion category may be determined according to the classified lesion type. For example, when the classified lesion type is pneumonia, the preset lesion category may include fungal pneumonia, bacterial pneumonia, viral pneumonia, etc.

[0060] In summary, for the technical solution in this application, after obtaining the CT sequence data and other modality data corresponding to the lesion area, the CT sequence data and other modality data can be preprocessed to obtain the target CT sequence corresponding to the CT sequence data and the target modality feature corresponding to the other modality data; then the target CT sequence and the target modality feature can be input into the pre-trained lesion category classification model. In the lesion category classification model, global features and local features are extracted based on the target CT sequence, and feature interaction and fusion processing are performed on the global features, local features, and target modality features. The lesion category is classified based on the fused multi-modal feature to obtain the classification result of the lesion area. For the technical solution in this application, by simultaneously extracting the global features, local features in the CT sequence, and the modality features of other modality data, and performing feature interaction and fusion processing on the global features, local features, and target modality features, the feature fusion closeness of different receptive fields can be strengthened, and the accuracy of the features can be improved. Classifying the lesion category based on the fused multi-modal feature can greatly improve the accuracy of lesion category classification.

[0061] Based on the above Figure 2 、 Figure 4 specific description of the lesion category classification method provided, as Figure 10 shown, Figure 10 is a block diagram of a lesion category classification system shown according to an exemplary embodiment. As Figure 10 shown, the system includes: An acquisition module 1010, which can be used to acquire CT sequence data and other modality data corresponding to the lesion area; The processing module 1020 can be used to preprocess the CT sequence data and other modality data to obtain the target CT sequence corresponding to the CT sequence data and the target modality features corresponding to the other modality data; The partitioning module 1030 can be used to input the target CT sequence and the target modality features into a pre-trained lesion category partitioning model to obtain the category partitioning result of the lesion region. Among them, the lesion category partitioning model extracts global features and local features based on the target CT sequence, and performs feature interaction and fusion processing on the global features, local features, and target modality features, and performs lesion category partitioning based on the fused multi-modal features to obtain the category partitioning result of the lesion region.

[0062] In some embodiments of the present application, the lesion category partitioning model includes a dual-branch feature extraction module, a multi-modal feature fusion module, and a classifier; the partitioning module 1030 is specifically configured to input the target CT sequence into the dual-branch feature extraction module, extract the target global feature and the target local feature corresponding to the target CT sequence, and perform feature mapping interaction fusion and feature splicing processing on the target global feature and the target local feature to obtain the target dual-branch fusion feature; in the multi-modal feature fusion module, perform feature fusion processing on the target dual-branch fusion feature, the target global feature, and the target local feature to obtain the first fusion feature, and splice the target modality features on the first fusion feature to obtain the multi-modal feature; use the classifier to perform category partitioning on the lesion region based on the multi-modal feature to obtain the category partitioning result.

[0063] In some embodiments of the present application, the dual-branch feature extraction module includes multiple feature extraction stages, and each feature extraction stage includes a global feature extraction module, a local feature extraction module, and a dual-branch feature fusion module; when inputting the target CT sequence into the dual-branch feature extraction module, extracting the target global feature and the target local feature corresponding to the target CT sequence, and performing feature mapping interaction fusion and feature splicing processing on the target global feature and the target local feature to obtain the target dual-branch fusion feature, the partitioning module 1030 is specifically configured to, for any one of the multiple feature extraction stages, use the global feature extraction module configured therein to extract the global feature of the first input feature, use the local feature extraction module configured therein to extract the local feature of the second input feature, and use the dual-branch feature fusion module configured therein to perform feature mapping interaction fusion on the third input feature to obtain the first dual-branch fusion feature of the current feature extraction stage; determine the first global feature and the first local feature extracted in the last feature extraction stage among the multiple feature extraction stages as the target global feature and the target local feature, and determine the first dual-branch fusion feature output by the dual-branch feature fusion module configured in the last feature extraction stage as the target dual-branch fusion feature; Among them, when the current feature extraction stage is the first feature extraction stage among multiple feature extraction stages, the first input feature and the second input feature are the target CT sequences, and the third input feature includes the first global feature and the first local feature extracted in the current feature extraction stage; when the current feature extraction stage is any feature extraction stage other than the first feature extraction stage among multiple feature extraction stages, the first input feature is the first global feature extracted in the previous feature extraction stage corresponding to the current feature extraction stage, the second input feature is the first local feature extracted in the previous feature extraction stage corresponding to the current feature extraction stage, the third input feature includes the first global feature and the first local feature extracted in the current feature extraction stage, and the first dual-branch fusion feature output in the previous feature extraction stage corresponding to the current feature extraction stage.

[0064] In some embodiments of the present application, the global feature extraction module includes a Swin transformer block and a patch merging block; for any feature extraction stage among multiple feature extraction stages, when using the global feature extraction module configured therein to extract the global feature of the first input feature, the partitioning module 1030 can specifically be used to, for any feature extraction stage among multiple feature extraction stages, extract the self-attention feature of each image patch in the first input feature by using the Swin transformer block, and merge the self-attention features of the image patches by using the patch merging block to obtain the first global feature extracted by the global feature extraction module in the current feature extraction stage.

[0065] In some embodiments of the present application, the local feature extraction module includes a convolutional block and a downsampling module, and the convolutional block includes a convolutional layer, a DropBlock layer, and a batch normalization and activation layer; for any feature extraction stage among multiple feature extraction stages, when using the local feature extraction module configured therein to extract the local feature of the second input feature, the partitioning module 1030 can specifically be used to input the second input feature into the convolutional block and iteratively perform convolutional feature processing until a preset number of iterations is reached to obtain convolutional features, where the convolutional feature processing sequentially includes convolutional processing, random feature block discarding processing, batch normalization processing, and activation processing; use the downsampling module to perform downsampling processing on the convolutional features to obtain the first local feature extracted by the local feature extraction module in the current feature extraction stage.

[0066] In some embodiments of the present application, for any one of the multiple feature extraction stages, when using the dual-branch feature fusion module configured therein to perform feature mapping interaction fusion on the third input feature to obtain the dual-branch fusion feature of the current feature extraction stage, the partitioning module 1030 can specifically be used to determine the first mapping feature weight of the first global feature extracted in the current feature extraction stage, and the second mapping feature weight of the first local feature extracted in the current feature extraction stage; perform feature cross-multiplication on the first global feature and the first local feature respectively based on the first mapping feature weight and the second mapping feature weight to obtain the second local feature mapping global information and the second global feature mapping local information; perform feature fusion on the second local feature and the second global feature to obtain a second fusion feature; when the current feature extraction stage is the first feature extraction stage among the multiple feature extraction stages, use the second fusion feature as the first dual-branch fusion feature of the current feature extraction stage; or, when the current feature extraction stage is any one of the multiple feature extraction stages other than the first feature extraction stage, splice the second fusion feature and the first dual-branch fusion feature of the previous feature extraction stage corresponding to the current feature extraction stage to obtain the first dual-branch fusion feature of the current feature extraction stage.

[0067] In some embodiments of the present application, in the multi-modal feature fusion module, when performing feature fusion processing on the target dual-branch fusion feature, the target global feature, and the target local feature to obtain a first fusion feature, and splicing the target modal feature on the first fusion feature to obtain the multi-modal feature, the partitioning module 1030 can specifically be used to perform fusion processing on the target dual-branch fusion feature, the target global feature, and the target local feature to obtain a third fusion feature; determine the first feature vector of the third fusion feature and the second feature vector of the target modal feature; splice the first feature vector and the second feature vector, and perform linear processing on the spliced feature vector to obtain the multi-modal feature output by the multi-modal feature fusion module.

[0068] Regarding the system in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0069] In the embodiments of the present application, by learning and integrating multi-modal data such as clinical information to assist in image diagnosis, the ability and accuracy of the model to distinguish multiple types of pneumonia are greatly improved. In addition, the dual-branch feature extraction method proposed by the present invention can better adaptively learn the lesion features of different scales, different receptive fields, and different shapes, and has good generalization performance in dealing with lesion changes. At the same time, the present invention proposes a dual-branch feature fusion module based on feature mapping interaction, which uses mapping weights to realize the interaction between local features and global features, strengthens the closeness of feature fusion in different receptive fields, improves the accuracy of features, and further improves the performance of the model and the recognition ability of different lesions.

[0070] In the above, the lesion category division system of the embodiments of the present invention has been described from the perspective of functional modules in combination with the accompanying drawings. It should be understood that the functional module can be implemented in the form of hardware, can also be implemented by instructions in the form of software, or can be implemented by a combination of hardware and software modules. Specifically, each step of the method embodiment of the lesion category division in the embodiments of the present invention can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in the form of software. The steps of the lesion category division method applied in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or can be executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the method embodiment of the above-mentioned lesion category division.

[0071] Figure 11 It is a schematic block diagram of an electronic device 1100 according to an embodiment provided by the present invention.

[0072] As Figure 11 shown, the electronic device 1100 may include: A memory 1110 and a processor 1120. The memory 1110 is used to store a computer program and transmit the program code to the processor 1120. In other words, the processor 1120 can call and run the computer program from the memory 1110 to implement the method in the embodiments of the present invention.

[0073] For example, the processor 1120 can be used to execute the above method embodiment according to the instructions in the computer program.

[0074] In some embodiments of the present invention, the processor 1120 may include, but is not limited to: General-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.

[0075] In some embodiments of the present invention, the memory 1110 includes, but is not limited to: Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0076] In some embodiments of the present invention, the computer program can be divided into one or more modules, and the one or more modules are stored in the memory 1110 and executed by the processor 1120 to complete the method provided by the present invention. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the controller.

[0077] As Figure 11 shown, the electronic device 1100 may further include: A transceiver 1130, which can be connected to the processor 1120 or the memory 1110.

[0078] Among them, the processor 1120 can control the transceiver 1130 to communicate with other devices. Specifically, it can send data to other devices or receive data sent by other devices. The transceiver 1130 can include a transmitter and a receiver. The transceiver 1130 can further include an antenna, and the number of antennas can be one or more.

[0079] It should be understood that the various components in the electronic device are connected through a bus system. Among them, the bus system includes, in addition to the data bus, a power bus, a control bus, and a status signal bus.

[0080] The present invention also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by the computer, the computer can execute the methods in the above method embodiments. Or rather, an embodiment of the present invention also provides a computer program product containing instructions. When the instructions are executed by the computer, the computer executes the methods in the above method embodiments.

[0081] When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center containing one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a Digital Video Disc (DVD)), or a semiconductor medium (such as a Solid State Disk (SSD)), etc.

[0082] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments claimed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0083] In several embodiments provided by the present invention, it should be understood that the disclosed systems, systems, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the systems or modules can be in electrical, mechanical, or other forms.

[0084] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, the various functional modules can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0085] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for classifying lesion categories, characterized in that, Including: Obtaining CT sequence data and other modality data corresponding to a lesion region; Preprocessing the CT sequence data and the other modality data to obtain a target CT sequence corresponding to the CT sequence data and target modality features corresponding to the other modality data; Inputting the target CT sequence and the target modality features into a pre-trained lesion category classification model to obtain a category classification result of the lesion region, where the lesion category classification model extracts global features and local features based on the target CT sequence, and performs feature interaction fusion processing on the global features, the local features, and the target modality features, and performs lesion category classification based on the fused multi-modal features to obtain the category classification result of the lesion region.

2. The method according to claim 1, wherein The lesion category classification model includes a dual-branch feature extraction module, a multi-modal feature fusion module, and a classifier; The step of inputting the target CT sequence and the target modality features into a pre-trained lesion category classification model to obtain a category classification result of the lesion region includes: Inputting the target CT sequence into the dual-branch feature extraction module, extracting a target global feature and a target local feature corresponding to the target CT sequence, and performing feature mapping interaction fusion and feature splicing processing on the target global feature and the target local feature to obtain a target dual-branch fusion feature; In the multi-modal feature fusion module, performing feature fusion processing on the target dual-branch fusion feature, the target global feature, and the target local feature to obtain a first fusion feature, and splicing the target modality features on the first fusion feature to obtain multi-modal features; Using the classifier to perform category classification on the lesion region based on the multi-modal features to obtain a category classification result.

3. The method according to claim 2, wherein The dual-branch feature extraction module includes multiple feature extraction stages, and each feature extraction stage respectively includes a global feature extraction module, a local feature extraction module, and a dual-branch feature fusion module; The step of inputting the target CT sequence into the dual-branch feature extraction module, extracting a target global feature and a target local feature corresponding to the target CT sequence, and performing feature mapping interaction fusion and feature splicing processing on the target global feature and the target local feature to obtain a target dual-branch fusion feature includes: For any one of the multiple feature extraction stages, using the global feature extraction module configured therein to extract global features from a first input feature, using the local feature extraction module configured therein to extract local features from a second input feature, and using the dual-branch feature fusion module configured therein to perform feature mapping interaction fusion on a third input feature to obtain a first dual-branch fusion feature of the current feature extraction stage; Determine the first global feature and the first local feature extracted in the last feature extraction stage among the multiple feature extraction stages as the target global feature and the target local feature, and determine the first dual-branch fusion feature output by the dual-branch feature fusion module configured in the last feature extraction stage as the target dual-branch fusion feature; Among them, when the current feature extraction stage is the first feature extraction stage among the multiple feature extraction stages, the first input feature and the second input feature are the target CT sequence, and the third input feature includes the first global feature and the first local feature extracted in the current feature extraction stage; when the current feature extraction stage is any feature extraction stage among the multiple feature extraction stages except the first feature extraction stage, the first input feature is the first global feature extracted in the previous feature extraction stage corresponding to the current feature extraction stage, the second input feature is the first local feature extracted in the previous feature extraction stage corresponding to the current feature extraction stage, the third input feature includes the first global feature and the first local feature extracted in the current feature extraction stage, and the first dual-branch fusion feature output in the previous feature extraction stage corresponding to the current feature extraction stage.

4. The method according to claim 3, wherein The global feature extraction module includes a Swin transformer block and a patch merging block; For any feature extraction stage among the multiple feature extraction stages, using the global feature extraction module configured therein to perform global feature extraction on the first input feature includes: For any feature extraction stage among the multiple feature extraction stages, use the Swin transformer block to extract the self-attention features of each image patch in the first input feature, and use the patch merging block to merge the self-attention features of the image patches to obtain the first global feature extracted by the global feature extraction module in the current feature extraction stage.

5. The method according to claim 3, characterized in that, The local feature extraction module includes a convolutional block and a downsampling module, and the convolutional block includes a convolutional layer, a DropBlock layer, and a batch normalization and activation layer; The use of the local feature extraction module configured therein to perform local feature extraction on the second input feature includes: Input the second input feature into the convolutional block, and perform convolutional feature processing iteratively until the preset number of iterations is reached to obtain convolutional features, where the convolutional feature processing sequentially includes convolutional processing, random feature block discarding processing, batch normalization processing, and activation processing; Use the downsampling module to perform downsampling processing on the convolutional features to obtain the first local feature extracted by the local feature extraction module in the current feature extraction stage.

6. The method according to claim 3, wherein The use of the dual-branch feature fusion module configured therein to perform feature mapping interaction fusion on the third input feature to obtain the dual-branch fusion feature in the current feature extraction stage includes: Determine the first mapping feature weight of the first global feature extracted in the current feature extraction stage, and the second mapping feature weight of the first local feature extracted in the current feature extraction stage; Perform feature cross - multiplication on the first global feature and the first local feature respectively based on the first mapping feature weight and the second mapping feature weight to obtain a second local feature of the mapped global information and a second global feature of the mapped local information; Perform feature fusion on the second local feature and the second global feature to obtain a second fusion feature; When the current feature extraction stage is the first feature extraction stage among the multiple feature extraction stages, use the second fusion feature as the first dual - branch fusion feature of the current feature extraction stage; or, When the current feature extraction stage is any feature extraction stage other than the first feature extraction stage among the multiple feature extraction stages, concatenate the second fusion feature and the first dual - branch fusion feature of the previous feature extraction stage corresponding to the current feature extraction stage to obtain the first dual - branch fusion feature of the current feature extraction stage.

7. The method according to claim 2, characterized in that In the multi - modal feature fusion module, perform feature fusion processing on the target dual - branch fusion feature, the target global feature, and the target local feature to obtain a first fusion feature, and concatenate the target modal feature on the first fusion feature to obtain a multi - modal feature, including; Perform fusion processing on the target dual - branch fusion feature, the target global feature, and the target local feature to obtain a third fusion feature; Determine the first feature vector of the third fusion feature and the second feature vector of the target modal feature; Concatenate the first feature vector and the second feature vector, and perform linear processing on the concatenated feature vectors to obtain the multi - modal feature output by the multi - modal feature fusion module.

8. A lesion category classification system, characterized in that, Including: An acquisition module for acquiring CT sequence data and other modal data corresponding to a lesion area; A processing module for pre - processing the CT sequence data and the other modal data to obtain a target CT sequence corresponding to the CT sequence data and a target modal feature corresponding to the other modal data; A division module for inputting the target CT sequence and the target modal feature into a pre - trained lesion category division model to obtain a category division result of the lesion area, where the lesion category division model extracts global features and local features based on the target CT sequence, performs feature interaction and fusion processing on the global features, the local features, and the target modal feature, and performs lesion category division based on the fused multi - modal features to obtain the category division result of the lesion area.

9. An electronic device, characterized in that, Including: A processor and a memory, the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 - 7.

10. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program causes a computer to execute the method according to any one of claims 1 - 7.