A method, device, equipment and storage medium for training an image classification model

CN117372796BActive Publication Date: 2026-09-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210748948.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-28
Publication Date
2026-09-25
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

[0012]本申请实施例提供了一种训练图像分类模型的方法、装置、计算机设备及存储介质,用于解决训练得到的目标图像分类模型的分类准确性和分类可靠性较低的问题

Benefits of technology

[0066]本申请实施例中,可以通过样本图像模态转换出不同图像模态的样本转换图像,从而图像分类模型可以基于一种图像模态的样本图像,学习到多种图像模态的图像特征,使得训练得到的目标图像分类模型,仅需要针对一个目标获取一种图像模态的待分类图像,就可以基于待分类图像进行准确且可靠的分类任务。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117372796B_ABST
    Figure CN117372796B_ABST
Patent Text Reader

Abstract

The application provides a method and device for training an image classification model, equipment and a storage medium, which can be applied to the field of artificial intelligence or the field of intelligent medical treatment, and is used to solve the problem of low classification accuracy and reliability of a target image classification model. The method comprises the following steps: for a selected sample image, using a modal processing model, converting the sample image modal into a corresponding sample converted image according to the image modal of at least one associated reference image; using an image classification model, extracting common features of the sample image and at least one obtained sample converted image, and determining the predicted classification of the sample image based on the common features; and adjusting the model parameters based on the at least one sample converted image, the at least one reference image and the predicted classification. The image features of multiple image modalities can be learned based on the sample image of one image modal, and the classification accuracy of the target image classification model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device, and storage medium for training an image classification model. Background Technology

[0002] With the continuous development of technology, more and more devices can provide image classification services, which can be used to determine the category to which a target in a target image belongs.

[0003] Since the same target can be associated with target images of multiple image modalities, and each image modal target image can represent different feature information of the target, when training an image classification model, multiple sample image sets of different image modalities can be combined to train the image classification model to obtain a more accurate target image classification model.

[0004] For example, for the same person's brain, there are associated magnetic resonance imaging (MRI) and positron emission tomography (PET) scans. MRI primarily represents the structural features of the brain, while PET primarily represents its functional features. Therefore, a more accurate target image classification model can be obtained by combining sample image sets containing MRI images and sample image sets containing PET images.

[0005] Traditional methods for training image classification models by combining multiple sample image sets from different image modalities include two types:

[0006] The first method is to first use image synthesis technology to combine sample images of multiple image modalities of the same target into a synthetic sample image, and then train the image classification model based on the set of synthetic sample images.

[0007] The second approach involves first extracting features from sample images of multiple image modalities of the same target, then using feature synthesis technology to synthesize the obtained multiple image features into synthetic features, and finally training the image classification model based on the synthetic features.

[0008] However, using the above two methods for model training has the following drawbacks:

[0009] First, the target image classification model trained using image synthesis technology still requires obtaining multiple image modalities of the same target for classification during its use. However, under certain technologies, it is difficult to obtain multiple image modalities of the same target in some image classification scenarios, which limits the application of the target image classification model.

[0010] Secondly, in sample images of multiple image modalities of the same target, not all pixels, nor all extracted image features, are helpful in determining the category of the target. This introduces a lot of noise information during the process of merging sample images or merging image features, resulting in a large training error.

[0011] It is evident that none of the training methods employed under the relevant technologies can guarantee the classification accuracy and reliability of the target image classification model obtained through training. Summary of the Invention

[0012] This application provides a method, apparatus, computer device, and storage medium for training an image classification model, which addresses the problem of low classification accuracy and reliability of the trained target image classification model.

[0013] Firstly, a method for training an image classification model is provided, including:

[0014] Acquire training data comprising a set of sample images and at least one set of reference images, wherein different sets of reference images have different image modalities, and each sample image is associated with a reference image containing the same target in each set of reference images;

[0015] Based on the obtained training data, and combined with the modality processing model, the image classification model to be trained is subjected to multiple rounds of iterative training to obtain the target image classification model. Each round of iteration performs the following operations:

[0016] For the selected sample image, the modality processing model is used to convert the sample image modality into a corresponding sample transformed image according to the image modality of at least one associated reference image;

[0017] Using the image classification model, common features are extracted from the sample image and at least one converted sample image, and the predicted classification of the sample image is determined based on the common features.

[0018] Based on the at least one sample transformed image, the at least one reference image, and the predicted classification, the model parameters of the modality processing model and the image classification model are adjusted.

[0019] Secondly, an apparatus for training an image classification model is provided, comprising:

[0020] Acquisition module: used to acquire training data containing a set of sample images and at least one set of reference images, wherein different sets of reference images have different image modalities, and each sample image in each set of reference images is associated with a reference image containing the same target;

[0021] Processing module: Based on the obtained training data and combined with the modality processing model, this module performs multiple rounds of iterative training on the image classification model to be trained, thereby obtaining the target image classification model. Each iteration performs the following operations:

[0022] The processing module is specifically used to: for the selected sample image, use the modality processing model to convert the sample image modality into a corresponding sample transformed image according to the image modality of at least one associated reference image;

[0023] The processing module is specifically used to: use the image classification model to extract common features of the sample image and at least one obtained sample transformed image, and determine the predicted classification of the sample image based on the common features;

[0024] The processing module is specifically used to: adjust the model parameters of the modality processing model and the image classification model based on the at least one sample converted image, the at least one reference image, and the predicted classification.

[0025] Optionally, the modality processing model includes at least one sample transformation network, which outputs images of different modalities from the sample image set, with different sample transformation networks outputting images corresponding to different image modalities; the processing module is specifically used for:

[0026] For each sample transformation network, perform the following operations:

[0027] The sample transformation network is used to perform multi-scale downsampling on the selected sample images to obtain multiple intermediate feature maps, each of which has a different resolution.

[0028] Based on the resolution of the sample image, the multiple intermediate feature maps are upsampled and fused to obtain the sample transformed image.

[0029] Optionally, the processing module is specifically used for:

[0030] Data distillation is performed on the sample image and the at least one transformed sample image to determine the common features between the sample image and the at least one transformed sample image;

[0031] The image classification model is used to extract the common features from the sample images, and the predicted classification of the sample images is determined based on the common features.

[0032] Optionally, the processing module is specifically used for:

[0033] Based on the at least one sample transformed image and the at least one reference image, the adversarial loss, cyclic loss, correlation loss, and reinforcement loss of the modality processing model are determined, wherein the adversarial loss characterizes the realism of the image after modality transformation, the cyclic loss characterizes the reproducibility of the image after modality transformation, the correlation loss characterizes the consistency of the image before and after modality transformation, and the reinforcement loss characterizes the accuracy of modality transformation when the image before and after modality transformation is the same image modality;

[0034] Based on the predicted classification, the classification loss of the image classification model is determined;

[0035] The model parameters of the modality processing model and the image classification model are adjusted based on the adversarial loss, the recurrence loss, the correlation loss, the reinforcement loss, and the classification loss.

[0036] Optionally, the modality processing model includes a discriminant network, which is used to: predict the discriminant probability of the input image obtained through modality transformation; the processing module is specifically used to:

[0037] The sample image, the at least one transformed sample image, and the at least one reference image are used as input images for the discrimination network to predict the corresponding discrimination probabilities.

[0038] Based on the error between each obtained discrimination probability and the reference probability associated with its corresponding input image, the adversarial loss is determined, and based on the at least one sample transformed image and the at least one reference image, the cyclic loss, correlation loss and reinforcement loss of the modality processing model are determined.

[0039] Optionally, the modality processing model includes a sample transformation network and a reference transformation network. The sample transformation network is used to obtain images of image modalities different from the sample image set, and the reference transformation network is used to obtain images of image modalities identical to the sample image set. The processing module is specifically used for:

[0040] Using the reference transformation network, the at least one reference image mode is transformed into a corresponding reference transformation image according to the image mode of the sample image;

[0041] Using the sample conversion network, at least one reference converted image modality is converted into a corresponding verification converted image according to the image modality of the at least one reference image;

[0042] Based on the error between each obtained verification transformed image and the sample image, the cyclic loss is determined, and based on the at least one sample transformed image and the at least one reference image, the adversarial loss, correlation loss, and reinforcement loss of the modality processing model are determined.

[0043] Optionally, the at least one reference image is associated with a corresponding reference transformed image, which is obtained by performing a mode transformation on the corresponding reference image according to the image mode of the sample image; the processing module is specifically used for:

[0044] Extract the common sample structure features between the sample image and the at least one sample transformed image, and extract the common reference structure features between the at least one reference image and the corresponding reference transformed image;

[0045] Based on the sample structural features and the reference structural features, the correlation loss is determined, and based on the at least one sample transformed image and the at least one reference image, the adversarial loss, cyclic loss, and reinforcement loss of the modality processing model are determined.

[0046] Optionally, the modality processing model includes a sample transformation network and a reference transformation network. The sample transformation network is used to obtain images of image modalities different from the sample image set, and the reference transformation network is used to obtain images of image modalities identical to the sample image set. The processing module is specifically used for:

[0047] Each sample image is used as the at least one reference image, and modality transformation is performed using the reference transformation network to obtain the corresponding sample enhanced image;

[0048] Each of the at least one reference image is used as the sample image, and modality conversion is performed using the sample conversion network to obtain the corresponding reference enhanced image;

[0049] The enhancement loss is determined based on the error between the obtained at least one sample enhanced image and the sample image, and the error between the obtained at least one reference enhanced image and its corresponding reference image. The adversarial loss, cyclic loss and correlation loss of the modality processing model are determined based on the at least one sample transformed image and the at least one reference image.

[0050] Optionally, the processing module is specifically used for:

[0051] If, among the obtained training losses, there is a training loss that does not meet its corresponding training objective, the model parameters of the modality processing model and the image classification model are adjusted based on the training losses, and the next round of iterative training is initiated until it is determined that all the obtained training losses meet their corresponding training objectives.

[0052] The training losses include the adversarial loss, the recurrence loss, the correlation loss, the reinforcement loss, and the classification loss.

[0053] Optionally, the processing module is further configured to:

[0054] Based on the obtained training data, and combined with the modality processing model, the image classification model to be trained is iteratively trained in multiple rounds to obtain the target image classification model, and then the image to be classified is obtained. The image modality of the image to be classified is the same as the image modality of the sample image set, and the image to be classified contains multiple image slices.

[0055] The target image classification model is used to extract features from the multiple image slices to obtain corresponding slice feature maps, and the obtained multiple slice feature maps are fused into an image feature map.

[0056] An attention mechanism is used to evaluate the importance of the multiple image slices based on the multiple slice feature maps and the image feature map, and obtain the corresponding importance evaluation values.

[0057] The multiple slice feature maps, the image feature map, and the obtained multiple importance evaluation values ​​are fused and stitched together to obtain the comprehensive features of the image to be classified, and the target classification of the image to be classified is determined based on the comprehensive features.

[0058] Optionally, the processing module is further configured to:

[0059] After determining the target classification of the image to be classified based on the comprehensive features, at least one target image slice with an importance evaluation value greater than a preset evaluation threshold is selected from the plurality of image slices;

[0060] The image classification interface presents at least one target image slice and its corresponding importance evaluation value.

[0061] Thirdly, a computer program product is provided, including a computer program that, when executed by a processor, implements the method described in the first aspect.

[0062] Fourthly, a computer device is provided, comprising:

[0063] Memory, used to store program instructions;

[0064] A processor is configured to invoke program instructions stored in the memory and execute the method described in the first aspect according to the obtained program instructions.

[0065] Fifthly, a computer-readable storage medium is provided, the computer-readable storage medium storing computer-executable instructions for causing a computer to perform the method as described in the first aspect.

[0066] In this embodiment, sample image transformation images of different image modalities can be generated by sample image modal transformation. Thus, the image classification model can learn the image features of multiple image modalities based on sample images of one image modality. This allows the trained target image classification model to perform accurate and reliable classification tasks based on the image to be classified by only obtaining an image of one image modality for one target.

[0067] When training an image classification model, the sample image modality is converted into the image modality of each corresponding reference image to obtain the corresponding sample-transformed images. The image classification model determines the predicted classification of the sample image through the common features between the sample image and the sample-transformed images, thereby achieving the purpose of training the image classification model. This allows the image classification model to learn the feature information of the sample image under multiple image modalities through the sample image and each sample-transformed image. Furthermore, through the common information, effective information is extracted from various feature information for the target. This effective information includes the features of the target in the sample image under multiple image modalities, eliminating the influence of noisy pixels. Thus, through multiple rounds of iterative training, the classification accuracy and reliability of the trained target image classification model are improved.

[0068] The training of the modality processing model is used to assist in the training of the image classification model. However, the modality transformation accuracy of the trained modality processing model is not used as a primary criterion for evaluating the classification accuracy and reliability of the target image classification model. This avoids a situation where, even with high modality transformation accuracy, the classification accuracy and reliability of the target image classification model may be low if the realism of the transformed sample images cannot be guaranteed. Simultaneously, in the process of improving the classification accuracy and reliability of the target image classification model, the semantic representation ability of the various image modalities obtained after modality transformation by the modality processing model is also improved, rather than simply the similarity between pixels. Attached Figure Description

[0069] Figure 1A A schematic diagram of MRI and PET images provided in the embodiments of this application;

[0070] Figure 1B This is one application scenario of the method for training an image classification model provided in the embodiments of this application;

[0071] Figure 2 A flowchart illustrating a method for training an image classification model provided in this application embodiment;

[0072] Figure 3 A schematic diagram illustrating the principle of a method for training an image classification model provided in this application embodiment;

[0073] Figure 4A A schematic diagram of the principle of the method for training an image classification model provided in the embodiments of this application. Figure 2 ;

[0074] Figure 4B A schematic diagram of the principle of the method for training an image classification model provided in the embodiments of this application. Figure 3 ;

[0075] Figure 4C A schematic diagram four illustrating the principle of a method for training an image classification model provided in this application embodiment;

[0076] Figure 4D A schematic diagram five illustrating the principle of a method for training an image classification model provided in this application embodiment;

[0077] Figure 5A A schematic diagram six illustrating the principle of a method for training an image classification model provided in this application embodiment;

[0078] Figure 5B A schematic diagram seven illustrating the principle of a method for training an image classification model provided in this application embodiment;

[0079] Figure 5C A schematic diagram of the principle of the method for training an image classification model provided in the embodiments of this application. Figure 8 ;

[0080] Figure 5D A schematic diagram of the principle of the method for training an image classification model provided in the embodiments of this application. Figure 9 ;

[0081] Figure 6A A flowchart illustrating a method for training an image classification model provided in an embodiment of this application. Figure 2 ;

[0082] Figure 6B A schematic diagram ten illustrating the principle of a method for training an image classification model provided in this application embodiment;

[0083] Figure 6C11. A schematic diagram illustrating the principle of a method for training an image classification model provided in this application embodiment;

[0084] Figure 6D A schematic diagram twelve illustrating the principle of a method for training an image classification model provided in this application embodiment;

[0085] Figure 6E A schematic diagram thirteen illustrating the principle of a method for training an image classification model provided in this application embodiment;

[0086] Figure 6F Fourteen is a schematic diagram illustrating the principle of a method for training an image classification model provided in this application embodiment;

[0087] Figure 6G A schematic diagram fifteen illustrating the principle of a method for training an image classification model provided in this application embodiment;

[0088] Figure 7A A schematic diagram sixteen illustrating the principle of a method for training an image classification model provided in this application embodiment;

[0089] Figure 7B A schematic diagram seventeen illustrating the principle of a method for training an image classification model provided in this application embodiment;

[0090] Figure 7C Schematic diagram eighteen illustrating the principle of a method for training an image classification model provided in this application embodiment;

[0091] Figure 7D A schematic diagram nineteen illustrating the principle of a method for training an image classification model provided in an embodiment of this application;

[0092] Figure 8 A schematic diagram of a device for training an image classification model provided in an embodiment of this application;

[0093] Figure 9 A schematic diagram of the structure of an apparatus for training an image classification model provided in the embodiments of this application. Figure 2 . Detailed Implementation

[0094] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0095] The following explanations of some terms used in the embodiments of this application are provided to facilitate understanding by those skilled in the art.

[0096] (1) Magnetic Resonance Image (MRI):

[0097] Magnetic resonance imaging utilizes the principle of nuclear magnetic resonance. Based on the different attenuations of the released energy in different structural environments within a substance, and by detecting the emitted electromagnetic waves through an external gradient magnetic field, the location and type of atomic nuclei that make up the object can be determined, and an image of the object's internal structure can be drawn.

[0098] The magnetic resonance imaging (MRI) scan consists of two sequences: T1 and T2. T1 highlights the difference in T1 relaxation (longitudinal relaxation) of tissues and is used to examine whether there are abnormalities in anatomical structures; T2 highlights the difference in T2 relaxation (transverse relaxation) of tissues and is used to examine whether there are abnormalities in the signal of tissue lesions.

[0099] (2) Positron emission tomography (PET) and fluorodeoxyglucose (FDG):

[0100] Positron emission tomography (PET) is a non-invasive imaging technique that generates three-dimensional medical images by detecting the gamma rays emitted when certain radioactive tracers are injected into the human body.

[0101] Fluorodeoxyglucose is a fluorinated derivative of 2-deoxyglucose. FDG is most commonly used in positron emission tomography (PET) medical imaging equipment. The fluorine in the FDG molecule can be fluorine-18, a positron-emitting radioactive isotope. After FDG is injected into a patient, a PET scanner can construct an image reflecting the distribution of FDG in the body.

[0102] (3) Alzheimer's Disease (AD):

[0103] Alzheimer's disease is a slowly progressing neurodegenerative disease that worsens over time. Clinically, it is characterized by comprehensive dementia manifestations, including memory impairment, aphasia, apraxia, agnosia, visuospatial skill impairment, executive dysfunction, and personality and behavioral changes.

[0104] (4) ADNI Database:

[0105] The ADNI database is currently the world's leading data center for Alzheimer's disease research. The ADNI dataset originates from the Alzheimer's Neuroimaging Project, initiated in 2013. This project aims to collect and organize data on Alzheimer's patients to further explore the changes and mechanisms in the disease process in order to find a cure. The project has employed over 1,500 adults aged approximately 55-90 years as participants, including healthy elderly individuals, those with early or late-stage mild cognitive impairment (MCI), and early-stage Alzheimer's patients.

[0106] The data in the ADNI database can currently be divided into four phases: ADNI-1, ADNI-GO, ADNI-2, and ADNI-3. All research-related agreements are reviewed by the local review committee of each participant and signed by the participant after obtaining their consent.

[0107] (5) Mini-Mental State Examination (MMSE), Functional Activities Questionnaire (FAQ), and apolipoprotein E4 (APOe4) allele:

[0108] The Mini-Mental State Examination (MMSE) is obtained through the Mini-Mental State Examination. The MMSE can comprehensively, accurately, and quickly reflect the intellectual state and degree of cognitive impairment of the subjects, and can provide a scientific basis for clinical psychological diagnosis and treatment as well as neuropsychological research.

[0109] The Functional Activity Scale is used to identify and evaluate elderly patients with less severe functional impairment, i.e., those with early-onset or mild dementia.

[0110] (6) Super-resolution test sequence (Visual Geometry Group, VGG) 19 network:

[0111] The VGG19 network contains 19 hidden layers, including 16 convolutional layers and 3 fully connected layers. The structure of the VGG network is very consistent, using 3x3 convolutional kernels and 2x2 max pooling throughout.

[0112] This application relates to the field of Artificial Intelligence (AI), and is designed based on Computer Vision (CV) and Machine Learning (ML) technologies. It can be applied to fields such as cloud computing, smart transportation, smart healthcare, and mapping.

[0113] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that studies the design principles and implementation methods of various machines, attempting to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence, enabling machines to have perception, reasoning, and decision-making functions.

[0114] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, interactive operating systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, machine learning / deep learning, autonomous driving, and intelligent transportation. With the development and progress of AI, it has been researched and applied in numerous fields, such as smart homes, intelligent customer service, virtual assistants, smart speakers, intelligent marketing, wearable devices, autonomous driving, drones, robots, smart healthcare, vehicle networking, and intelligent transportation. It is believed that with further technological advancements, AI will be applied in even more fields, playing an increasingly important role. The solutions provided in this application's embodiments relate to deep learning and augmented reality technologies in AI, which are further illustrated by the following examples.

[0115] Computer vision is the science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in identifying, tracking, and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), autonomous driving, intelligent transportation, and other technologies, as well as common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0116] Machine learning is a multidisciplinary field that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and many other disciplines. It specifically studies how computers acquire new knowledge or skills by simulating human learning behavior, reorganize existing knowledge structures, and continuously improve their performance.

[0117] Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence. Its applications span all areas of artificial intelligence. Deep learning, the core of machine learning, is a technology that enables machine learning. Machine learning typically includes techniques such as deep learning, reinforcement learning, transfer learning, inductive learning, artificial neural networks, and instructional learning. Deep learning includes techniques such as convolutional neural networks (CNNs), deep belief networks, recurrent neural networks, autoencoders, and generative adversarial networks.

[0118] It should be noted that in the embodiments of this application, data related to sample image sets, reference image sets, or images to be classified are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0119] The following is a brief introduction to the application areas of the method for training image classification models provided in the embodiments of this application.

[0120] With the continuous development of technology, more and more devices can provide image classification services, which can be used to determine the category to which a target in a target image belongs.

[0121] Since the same target can be associated with target images of multiple image modalities, and each image modal target image can represent different feature information of the target, when training an image classification model, multiple sample image sets of different image modalities can be combined to train the image classification model to obtain a more accurate target image classification model.

[0122] For example, for the same person's brain, there are associated images such as magnetic resonance imaging (MRI) and positron emission tomography (PET). MRI primarily represents the structural features of the brain, while PET primarily represents its functional features. Therefore, a more accurate target image classification model can be obtained by combining sample image sets containing MRI images and sample image sets containing PET images.

[0123] Traditional methods for training image classification models by combining multiple sample image sets from different image modalities include two types:

[0124] The first method is to first use image synthesis technology to combine sample images of multiple image modalities of the same target into a synthetic sample image, and then train the image classification model based on the set of synthetic sample images.

[0125] The second approach involves first extracting features from sample images of multiple image modalities of the same target, then using feature synthesis technology to synthesize the obtained multiple image features into synthetic features, and finally training the image classification model based on the synthetic features.

[0126] However, using the above two methods for model training has the following drawbacks:

[0127] First, the target image classification model trained using image synthesis technology still requires obtaining multiple image modalities of the same target for classification during its use. However, under certain technologies, it is difficult to obtain multiple image modalities of the same target in some image classification scenarios, which limits the application of the target image classification model.

[0128] For example, in the context of diagnosing Alzheimer's Disease (AD), incomplete medical data from multiple image modalities is often unavoidable due to subjective factors of the subjects or objective factors such as data quality. These challenges include excessively high costs of multiple examinations, incomplete equipment availability in different hospitals, and difficulties in data collection. Therefore, when using a target image classification model to perform image classification tasks on either magnetic resonance imaging (MRI) or positron emission tomography (PET) images of a subject, it will be difficult to obtain accurate classification results.

[0129] Even with modality transformation tasks that incorporate related technologies, the image data generated by modality transformation tasks is highly controversial due to the inherent differences between medical images and natural images. These images still differ significantly from real images, rendering modality transformation tasks lacking truly practical clinical value and consequently causing image classification tasks to lack corresponding diagnostic persuasiveness.

[0130] Secondly, in sample images of multiple image modalities of the same target, not all pixels, nor all extracted image features, are helpful in determining the category of the target. This introduces a lot of noise information during the process of merging sample images or merging image features, resulting in a large training error.

[0131] For example, regarding magnetic resonance images, please refer to... Figure 1A (1) This is a magnetic resonance imaging (MRI) image of the human brain. The classification objective of this image is to determine the lesions in the brain. Besides features related to lesions, this image also includes various other features such as high contrast between pixels and clear visibility of each ventricle. For positron emission tomography (PET) imaging, please refer to... Figure 1A (2) This is a positron emission tomography (PET) image of the human brain. The classification objective in this image is still to determine the lesion condition of the human brain. However, in addition to the features related to the lesion, the image also includes various other features such as low contrast between pixels and indistinct edges of each ventricle. Compared to the classification objective of determining the lesion condition of the human brain, the features other than those related to the lesion are all noise information, which can easily cause errors in the training process of the image classification model, resulting in low classification accuracy and reliability of the trained target image classification model.

[0132] It is evident that none of the training methods employed under the relevant technologies can guarantee the classification accuracy and reliability of the target image classification model obtained through training.

[0133] To address the issue of low classification accuracy and reliability in trained target image classification models, this application proposes a method for training image classification models. In this method, after acquiring training data comprising a set of sample images and at least one set of reference images, the image classification model to be trained is iteratively trained multiple times based on the acquired training data and in conjunction with a modality processing model to obtain the target image classification model. Different reference image sets have different image modalities, and each sample image in each reference image set is associated with a reference image containing the same target.

[0134] Each iteration performs at least the following operations: For the selected sample image, using a modality processing model, the sample image modality is converted into a corresponding transformed sample image according to the image modality of at least one associated reference image. Using an image classification model, common features are extracted from the sample image and the obtained at least one transformed sample image, and the predicted classification of the sample image is determined based on these common features. Based on at least one transformed sample image, at least one reference image, and the predicted classification, the model parameters of the modality processing model and the image classification model are adjusted.

[0135] In this embodiment, sample image transformation images of different image modalities can be generated by sample image modal transformation. Thus, the image classification model can learn the image features of multiple image modalities based on sample images of one image modality. This allows the trained target image classification model to perform accurate and reliable classification tasks based on the image to be classified by only obtaining an image of one image modality for one target.

[0136] When training an image classification model, the sample image modality is converted into the image modality of each corresponding reference image to obtain the corresponding sample-transformed image. The image classification model determines the predicted classification of the sample image through the common features between the sample image and the sample-transformed image, thereby achieving the purpose of training the image classification model. This allows the image classification model to learn the feature information of the sample image under multiple image modalities through the sample image and each sample-transformed image. Furthermore, through the common information, effective information is extracted from various feature information for the target. This effective information includes the features exhibited by the target in the sample image under multiple image modalities. Thus, through multiple rounds of iterative training, the classification accuracy and reliability of the trained target image classification model are improved.

[0137] The training of the modality processing model is used to assist in the training of the image classification model. However, the modality transformation accuracy of the trained modality processing model is not used as a primary criterion for evaluating the classification accuracy and reliability of the target image classification model. This avoids a situation where, even with high modality transformation accuracy, the presence of numerous noisy pixels in the images across multiple modalities could lead to lower classification accuracy and reliability in the target image classification model. Simultaneously, in improving the classification accuracy and reliability of the target image classification model, the semantic representation capabilities of the images obtained after modality transformation by the modality processing model are also enhanced, rather than simply improving the similarity between pixels.

[0138] The following describes the application scenarios of the method for training the image classification model provided in this application.

[0139] Please refer to Figure 1B This diagram illustrates an application scenario of the image classification model training method provided in this application. The application scenario includes a client 101 and a server 102. The client 101 and the server 102 can communicate with each other. The communication method can be wired, such as through a network cable or serial cable; or wireless, such as through Bluetooth or Wi-Fi. No specific limitation is imposed.

[0140] Client 101 generally refers to a device that can provide sample image sets to server 102 or use a trained target image classification model, such as a terminal device, a third-party application accessible to the terminal device, or a webpage accessible to the terminal device. Terminal devices include, but are not limited to, mobile phones, computers, smart medical devices, smart home appliances, vehicle terminals, or aircraft. Server 102 generally refers to a device that can train an image classification model, such as a terminal device or a server. Servers include, but are not limited to, cloud servers, local servers, or associated third-party servers. Both client 101 and server 102 can use cloud computing to reduce the consumption of local computing resources; similarly, they can also use cloud storage to reduce the consumption of local storage resources.

[0141] As one embodiment, the client 101 and the server 102 can be the same device, and there is no specific limitation. In this embodiment, the client 101 and the server 102 are described as different devices.

[0142] The following is based on Figure 1B Using server 102 as the main component, this paper provides a detailed description of the method for training an image classification model provided in this embodiment. Please refer to [link / reference]. Figure 2 This is a flowchart illustrating a method for training an image classification model provided in an embodiment of this application.

[0143] S201, Obtain training data containing a sample image set and at least one reference image set.

[0144] Before training an image classification model, training data containing a set of sample images and at least one set of reference images can be obtained. The sample image set and each reference image set have different image modalities, and different reference image sets have different image modalities. Each sample image in each reference image set is associated with a reference image containing the same target.

[0145] For example, taking a magnetic resonance imaging (MRI) image set as the sample image set and a positron emission tomography (PET) image set as the reference image set, for a magnetic resonance image in the MRI image set, there exists a PET image in the PET image set. The magnetic resonance image and the PET image are brain images of the same subject.

[0146] As one example, in the context of diagnosing Alzheimer's disease, the sample image and the reference image can be one of MRI, PET, computed tomography (CT) images, and ultrasound images, respectively.

[0147] Image classification tasks in image classification models can be tasks that classify the entire image, or tasks that classify pixels corresponding to targets within the image, etc., with no specific limitations. Image classification tasks can be used for diagnosing Alzheimer's disease, for target registration, for radiotherapy such as localized treatments, for face synthesis, and for target identification, etc., with no specific limitations. In this embodiment, the image classification task of classifying the entire image is used as an example for description. The training process for other classification tasks is similar and will not be repeated here.

[0148] As one embodiment, training data can be received from other devices or read from a storage unit. Training data can also be obtained by downloading from the ADNI database in online resources. For example, using magnetic resonance imaging (MRI) as sample images and positron emission tomography (PET) as reference images, a total of 564 T1 MRI images and 549 FDG PET images were obtained from the ADNI1 and ADNI2 stages of the ADNI database, representing both control normal (CN) and Alzheimer's disease (AD) patients. Please refer to Table 1. Note that the MRI and PET data may not all come from the same subject; some subjects may only have image data for a single modality.

[0149] Table 1

[0150]

[0151] Wherein, the symbol * indicates that no significant differences in gender or age were found among different subject groups in the corresponding image modality; This indicates that significant differences were found in MMSE, FAQ, and APOe4 among different subject groups in the corresponding image modality. Chi-square tests can be used to assess gender and APOe4; t-tests can be used to assess age, MMSE, and FAQ scores.

[0152] As one embodiment, when receiving images of multiple image modalities sent by other devices, reading images of multiple image modalities from a storage unit, or downloading images of multiple image modalities from network resources, each image can be preprocessed to remove unnecessary content from the image and to achieve quantitative analysis for each image.

[0153] For example, please refer to human brain MRI and PET images used to diagnose Alzheimer's disease. Figure 3When acquiring images of multiple image modalities from the ADNI database, the initial MRI and PET images can be preprocessed using the ADNI preprocessing pipeline. Following the ADNI preprocessing pipeline, software for processing 3D brain data (such as Freesurfer) can be used to remove skull fragments from the MRI and PET images. Then, medical image registration software (such as Ants) can be used to register the MRI and PET images to the MNI152 standard space. Finally, outlier handling and standardization are performed on the MRI and PET images to complete the preprocessing operations, resulting in preprocessed MRI and PET images. Based on these preprocessed images, a sample image set and at least one reference image set can be obtained.

[0154] S202, based on the obtained training data, combined with the modality processing model, the image classification model to be trained is subjected to multiple rounds of iterative training to obtain the target image classification model.

[0155] After obtaining training data containing at least one set of reference images from a set of sample images, the image classification model to be trained can be iteratively trained multiple times based on the obtained training data and in conjunction with a modality processing model to obtain the target image classification model.

[0156] The modality processing model can be a model pre-trained on other datasets. Therefore, the image classification model can be trained based on the model parameters of the modality processing model obtained after pre-training. This avoids the situation where the model parameters of the modality processing model and the image classification model are highly random, which leads to low efficiency in training the image classification model.

[0157] During multiple rounds of iterative training of the image classification model, if the training losses of the modality processing model and the image classification model meet the training objective, then the trained target image classification model is output. This target image classification model can be used to classify images to be classified that have the same image modality as the sample image set. If the training losses of the modality processing model and the image classification model do not meet the training objective, then the model parameters of the modality processing model and the image classification model are adjusted based on the training loss, and the next round of iterative training begins.

[0158] The following sections S203 to S205 will use one round of iterative training as an example to introduce the process. Each round of iterative training is similar and will not be described in detail here.

[0159] S203, for the selected sample image, a modality processing model is used to convert the sample image modality into the corresponding sample transformed image according to the image modality of at least one associated reference image.

[0160] During a certain round of iterative training, a sample image can be selected from the sample image set, and at least one corresponding reference image can be obtained from at least one reference image set for use in this round of iterative training. Each reference image contains the same target as the sample image.

[0161] For example, if the sample image set is an MRI image set and at least one reference image set is a PET image set, then an MRI image can be selected from the MRI image set, and a PET image belonging to the same subject as the MRI image can be obtained from the PET image set.

[0162] After obtaining the sample image and at least one corresponding reference image, a modality processing model can be used to convert the sample image modality into a corresponding sample transformed image according to the image modality of the at least one reference image. The image modality of different sample transformed images corresponds to the image modality of different reference images.

[0163] For example, if the sample image is an MRI image and at least one reference image is a PET image, then please refer to... Figure 4A A modality processing model can be used to convert the MRI image modality into a pseudo-PET image, i.e., a sample-converted image, according to the image modality of the PET image.

[0164] As one embodiment, the modality processing model may include at least one sample transformation network. This network outputs images of different modalities than the sample image set, with each network corresponding to a different modality. Therefore, when there are multiple reference image sets, a separate sample transformation network can be used to transform the sample image modality into the image modality corresponding to each reference image set. Thus, multiple sample transformation networks can transform the sample image modality into the image modality corresponding to each of the multiple reference image sets.

[0165] For example, please refer to Figure 4B If the sample image is an MRI image and at least one reference image is a PET image and a CT image, then the modality processing model can contain two sample transformation networks: one for converting the sample image modality into a pseudo-PET image and the other for converting the sample image modality into a pseudo-CT image. This avoids the situation where a single sample transformation network is used, which would result in a complex sample transformation network structure and a long training process when multiple image modalities need to be converted. This simplifies the design difficulty of the sample transformation network and improves the training efficiency.

[0166] The following section uses a sample transformation network to illustrate the modality transformation process. The modality transformation process of each sample transformation network is similar and will not be elaborated further here.

[0167] A sample transformation network is used to perform multi-scale downsampling on the sample image, obtaining multiple intermediate feature maps, each with a different resolution. The resolution of each intermediate feature map can be less than or equal to the resolution of the sample image. Based on the resolution of the sample image, the sample transformation network then performs upsampling and fusion on these multiple intermediate feature maps to obtain the sample-transformed image.

[0168] Please refer to Figure 4C Taking an MRI image as the sample image and a PET image as the reference image as an example, the resolution of the sample image is H×W, where H is the height of the sample image and W is the width of the sample image. Since both MRI and PET images are 3D images, a 3D convolutional kernel can be used to perform four multi-scale downsampling operations on the sample image sequentially. After the first downsampling operation, a first intermediate feature map is obtained, with a resolution of [missing information]. A second downsampling operation is then performed on the first intermediate feature map. After the second downsampling operation, a second intermediate feature map is obtained, with a resolution of [resolution missing]. The second intermediate feature map is then subjected to a third downsampling operation. After the third downsampling operation, a third intermediate feature map is obtained, with a resolution of [resolution missing]. The third intermediate feature map is then subjected to a fourth downsampling operation. After the fourth downsampling operation, a fourth intermediate feature map is obtained, with a resolution of [resolution missing].

[0169] For the fourth intermediate feature map, a convolution operation is performed using the same padding method to output a fifth intermediate feature map, which has the same resolution as the fourth intermediate feature map. Then, based on the skip connection method, corresponding elements of the fourth and fifth intermediate feature maps are added to obtain a sixth intermediate feature map, which has a resolution of [missing information].

[0170] Please refer to Figure 4D Based on the resolution of the sample images, upsampling and fusion processing is performed on the first, second, third, fourth, and sixth intermediate feature maps. This can be achieved using 3D transposed convolution. A first concatenation operation is performed on the fourth and sixth intermediate feature maps, followed by a first upsampling operation on the concatenated feature map to obtain the first fused feature map. The resolution of this first fused feature map is [resolution missing]. The third intermediate feature map and the first fused feature map are then concatenated a second time, and the concatenated feature map is upsampled a second time to obtain the second fused feature map. The resolution of this second fused feature map is [resolution missing]. The second intermediate feature map and the second fused feature map are then concatenated a third time, and the concatenated feature map is upsampled a third time to obtain the third fused feature map. The resolution of this third fused feature map is [resolution missing]. The first intermediate feature map and the third fused feature map are then stitched together for the fourth time, and the stitched feature map is then upsampled for the fourth time to obtain a sample-transformed image. The resolution of the sample-transformed image is H×W, where H is the height of the sample-transformed image and W is the width of the sample-transformed image.

[0171] As an example, to reduce information loss during multi-scale downsampling and upsampling fusion, mirror padding and sample normalization are used for 3D convolution and 3D transposed convolution, respectively, and LeakyReLU activation function is used for all activation functions except for the last layer which uses Tanh activation function.

[0172] As one example, the process of generating 3D sample-converted images differs from that of generating 2D sample-converted images, requiring significantly more storage space. Furthermore, to adapt to sample images of different scales for texture feature learning, each original image can be randomly cropped into multiple sample images of uniform size during the training phase. For example, if the original image resolution is 256×256, each original image can be randomly cropped into four sample images with a resolution of 128×128. This not only expands the scale of the training data, facilitating the acquisition of a target image classification model with higher accuracy, but also accelerates the model's achievement of a more stable performance state.

[0173] S204. Using an image classification model, extract the common features of the sample image and at least one transformed sample image, and determine the predicted classification of the sample image based on the common features.

[0174] After obtaining at least one sample transformed image, an image classification model can be used to extract the common features of the sample image and the obtained at least one sample transformed image. The common features are used to characterize the feature information contained in both the sample image and the at least one sample transformed image. Thus, through the common features, noise information in the sample image and the at least one sample transformed image that does not help with image classification can be masked, making the image classification model more focused on learning the semantic expression of the sample image rather than shallow features.

[0175] An image classification model is used to determine the predicted classification of sample images based on common features. Since common features can characterize the semantic expression of sample images, the model parameters of the image classification model are corrected through predicted classification. Through iterative training, the trained target image classification model can have a more accurate classification ability. The process of determining the predicted classification of sample images using an image classification model is similar to the process of determining the target classification of an image to be classified using a target image classification model, which will be described in detail later and will not be repeated here.

[0176] As one embodiment, data distillation can be used to extract underlying shared features between a sample image and at least one transformed sample image. This shared feature information can then be used to train an image classification model. For example, data distillation is performed on a sample image and at least one transformed sample image to determine the shared features between them. An image classification model is then used to extract these shared features from the sample images, and based on these features, the predicted classification of the sample images is determined.

[0177] Alternatively, the image classification model can be trained based on the underlying shared features extracted from the sample images and the underlying shared features extracted from the transformed sample images. For example, a first shared feature is extracted from the sample images, and at least one second shared feature is extracted from at least one transformed sample image. Based on the first shared feature, a first predicted classification of the sample image is determined, and based on at least one second shared feature, at least one second predicted classification of the sample image is determined. The first predicted classification and at least one second predicted classification are then used as predicted classifications to train the image classification model.

[0178] S205, based on at least one sample transformed image, at least one reference image, and predicted classification, adjust the model parameters of the modality processing model and the image classification model.

[0179] After obtaining at least one sample transformed image, at least one reference image, and a predicted classification, the training loss of the modality processing model and the image classification model can be determined based on these elements. When the training loss satisfies the training objective, the trained target image classification model is output. If the training loss does not meet the training objective, the model parameters of the modality processing model and the image classification model are adjusted, and the next round of iteration training begins.

[0180] As one example, based on at least one sample transformed image and at least one reference image, various training losses for the modality processing model can be determined, including adversarial loss, cyclic loss, correlation loss, and reinforcement loss. Adversarial loss characterizes the realism of the image after modality transformation, cyclic loss characterizes the reproducibility of the image after modality transformation, correlation loss characterizes the consistency of the image before and after modality transformation, and reinforcement loss characterizes the accuracy of modality transformation when the images before and after transformation belong to the same image modality.

[0181] Based on predictive classification, the classification loss of the image classification model can be determined, and the model parameters of the modality processing model and the image classification model can be adjusted based on adversarial loss, cyclic loss, correlation loss, reinforcement loss and classification loss.

[0182] As one example, when determining the classification loss of an image classification model based on predictive classification, each sample image is associated with a corresponding classification label. Taking the scenario of diagnosing Alzheimer's disease as an example, the classification label can characterize whether the human brain in the corresponding sample image is the brain of an Alzheimer's disease patient (AD) or a normal person (Control Normal, CN). Predictive classification can predict whether the human brain in the sample image is the brain of an Alzheimer's patient or a normal person. Therefore, the classification loss of the image classification model can be determined based on the error between the predicted classification and the classification label.

[0183] When determining the classification loss of an image classification model based on the error between the predicted classification and the classification label, the cross-entropy function can be used for calculation. The predicted classification can be the probability distribution formed by predicting whether the human brain in the sample image belongs to the brain of a person with Alzheimer's disease and the probability that it belongs to the brain of a normal person. Similarly, the classification label is also the probability distribution formed by predicting whether the human brain in the sample image belongs to the brain of a person with Alzheimer's disease and the probability that it belongs to the brain of a normal person. Therefore, based on the error between the predicted classification and the classification label, the classification loss L of the image classification model can be determined. Classify Please refer to formula (1).

[0184]

[0185] Where, p i_AD Let y represent the probability that the predicted classification of the i-th sample image belongs to the brain of a person with Alzheimer's disease. i_AD p represents the probability that the classification label of the i-th sample image belongs to the brain of a person with Alzheimer's disease; i_CN Let y represent the probability that the predicted classification of the i-th sample image belongs to a normal human brain. i_CNThe probability that the i-th sample image belongs to the normal human brain in the classification label; N represents the number of sample images in the sample image set.

[0186] By adjusting the model parameters, the image classification model can learn to extract the semantic information of the target represented by multiple image modalities from the sample image. The trained target image classification model can determine the semantic information of the target represented by multiple image modalities in the image to be classified based on the image to be classified in one image modality, thus obtaining an accurate target classification.

[0187] There are several methods for determining the training loss of a modal processing model. Below, we will introduce one method for determining the adversarial loss, cyclic loss, correlation loss, and reinforcement loss of a modal processing model. Other methods will not be listed here.

[0188] Combat losses:

[0189] In addition to a sample transformation network for modality transformation, a modality processing model may also include a discriminant network. The discriminant network is used to predict the discrimination probability of an input image obtained through modality transformation. Therefore, after obtaining at least one sample-transformed image, the sample image, at least one sample-transformed image, and at least one reference image can be used as input images to the discriminant network to predict the corresponding discrimination probabilities.

[0190] The adversarial loss is determined based on the error between each obtained discrimination probability and the reference probability associated with its corresponding input image. The reference probability associated with the input image can be preset. For example, since the sample transformed image is obtained through modality transformation, the reference probability can be set to 1; since the sample image and reference image are original images and have not undergone modality transformation, the reference probability can be set to 0, etc. There are no specific restrictions.

[0191] As one embodiment, the modality processing model may include a reference transformation network, in addition to a sample transformation network for obtaining images of image modalities different from the sample image set, for obtaining images of the same image modalities as the sample image set.

[0192] By using a reference transformation network, at least one reference image mode can be transformed into a corresponding reference transformation image according to the image mode of the sample image. The image mode of the reference transformation image is the same as that of the sample image.

[0193] The modality processing model may further include a sample discriminant network and a reference discriminant network. The sample discriminant network is used to predict the discriminant probability of the input image obtained through modality transformation when the input image is either a sample image or a sample-transformed image. The reference discriminant network is used to predict the discriminant probability of the input image obtained through modality transformation when the input image is either a reference image or a reference-transformed image.

[0194] After obtaining at least one sample transformed image and at least one reference transformed image, the sample image and at least one sample transformed image can be used as input images for the sample discrimination network to predict the corresponding sample discrimination probability; at least one reference image and at least one reference transformed image can be used as input images for the reference discrimination network to predict the corresponding reference discrimination probability.

[0195] The adversarial loss is determined based on the error between the obtained discrimination probabilities of each sample and each reference discrimination probability and the reference probability associated with their corresponding input image. The reference probabilities associated with the input image are similar to the aforementioned reference probabilities and will not be repeated here.

[0196] For example, if the sample image is an MRI image and at least one reference image is a PET image, then please refer to... Figure 5A The modality processing model includes a sample transformation network and a reference transformation network, a sample discrimination network and a reference discrimination network.

[0197] MRI images are modally transformed using a sample transformation network to obtain pseudo-PET images. The MRI images and pseudo-PET images are then input into a sample discrimination network to obtain the corresponding sample discrimination probabilities.

[0198] PET images are modally transformed using a reference transformation network to obtain pseudo-MRI images. The PET images and pseudo-MRI images are then input into a reference discrimination network to obtain their respective reference discrimination probabilities.

[0199] Based on the errors between the sample discrimination probability and the reference probability of MRI images, the errors between the sample discrimination probability and the reference probability of pseudo-PET images, the errors between the reference discrimination probability and the reference probability of PET images, and the errors between the sample discrimination probability and the reference probability of pseudo-MRI images, the adversarial loss L is determined. GAN Please refer to formula (2).

[0200] L GAN =-E PET [(D PET (PET)-1) 2 ]-E MRI [(D PET (GPET (MRI) 2 (2)

[0201] Among them, -E PET [(D PET (PET)-1) 2 The error between the sample discrimination probability of an MRI image and the reference probability of an MRI image, and the error between the sample discrimination probability of a pseudo-PET image and the reference probability of a pseudo-PET image; -E MRI [(D PET (G PET (MRI) 2 The error between the reference discrimination probability of a PET image and the reference probability of a PET image, and the error between the sample discrimination probability of a spoofed MRI image and the reference probability of a spoofed MRI image.

[0202] Cyclic loss:

[0203] In a modality processing model that includes a sample transformation network and a reference transformation network, where the sample transformation network is used to obtain images of different modalities from the sample image set, and the reference transformation network is used to obtain images of the same modalities as the sample image set, the reference transformation network is used to convert at least one reference image modality into a corresponding reference transformed image according to the image modality of the sample image. Then, the sample transformation network is used to convert at least one obtained reference transformed image modality into a corresponding verification transformed image according to the image modality of at least one reference image. The cyclic loss is determined based on the error between each obtained verification transformed image and the sample image.

[0204] As one embodiment, after using a reference transformation network to convert at least one reference image mode into a corresponding reference transformation image according to the image mode of the sample image, a sample transformation network can be used to convert the obtained at least one reference transformation image mode into a corresponding sample verification transformation image according to the image mode of at least one reference image.

[0205] After using a sample transformation network to convert the sample image modality into a corresponding sample transformation image according to the image modality of at least one reference image, a reference transformation network can be used to convert the obtained at least one sample transformation image modality into a corresponding reference verification transformation image according to the image modality of the sample image.

[0206] Based on the errors between each sample verification transformed image and the sample image, and the errors between each reference verification transformed image and the corresponding reference image, the cyclic loss is determined.

[0207] For example, if the sample image is an MRI image and at least one reference image is a PET image, then the modality processing model includes a sample transformation network and a reference transformation network. Please refer to [reference needed]. Figure 5B (1) A reference conversion network is used to convert the PET image modality into a pseudo-MRI image according to the image modality of the MRI image. Then, a sample conversion network is used to convert the pseudo-MRI image modality into the corresponding verification PET image according to the image modality of the PET image.

[0208] Please refer to Figure 5B (2) A sample conversion network is used to convert the MRI image modality into a pseudo-PET image according to the image modality of the PET image. Then, a reference conversion network is used to convert the pseudo-PET image modality into the corresponding verification MRI image according to the image modality of the MRI image.

[0209] The cyclic loss is determined based on the errors between validating MRI images and the errors between validating PET images. For example, the weighted sum of the L1 norms of the errors between validating MRI images and the errors between validating PET images is used as the cyclic loss.

[0210] As one embodiment, when determining the cyclic loss, the L1 norm of the error between MRI images and the L1 norm of the error between PET images can be combined, and the determination can be based on the cyclic reinforcement loss network.

[0211] The recurrent reinforcement loss network consists of two layers: the 7th and 12th layers of a pre-trained VGG19 network, denoted as F. i Where i = 1, 2. The 7th and 12th layers of the network can be stacked in a certain proportion, denoted by the proportion coefficient λ. i , where i = 1, 2.

[0212] Taking an MRI image as the sample image and at least one PET image as the reference image as an example, the image classification model focuses more on the process of generating pseudo-PET images from MRI images. Therefore, λ can be further set. CyclePET and λ CycleMRI This is used to control the weight of the two generation paths: generating pseudo-PET images from MRI images and generating pseudo-MRI images from PET images.

[0213] Therefore, based on verifying the errors between MRI images and the errors between PET images, the cyclic loss L is determined. Cycle Please refer to formula (3).

[0214]

[0215] Among them, G MRI (PET) characterization of pseudo-MRI images generated from PET images, G PET (G MRI (PET) characterization was performed by generating verification PET images from pseudo-MRI images; G PET (MRI) characterization of pseudo-PET images generated from MRI images, G MRI (G PET (MRI) characterization was performed by generating verification MRI images from pseudo-PET images; ∑ i λ i ‖F i (G MRI (G PET (MRI)))-F i (MRI)‖1 Characterizes and verifies the sum of errors of MRI images when they pass through each layer of the cyclic enhancement loss network; ∑ i λ i ‖F i (G PET (G MRI (PET)))-F i (PET)‖1 characterizes and verifies the sum of errors of PET images and PET images when passed through each layer of a recurrent reinforcement loss network.

[0216] Related loss:

[0217] When at least one reference image is associated with a corresponding reference transformed image, which is obtained by modality transformation of the corresponding reference image according to the image modality of the sample image, the common sample structural features between the sample image and at least one sample transformed image can be extracted, as well as the common reference structural features between the structural features of at least one reference image and the corresponding reference transformed image can be extracted. Based on the sample structural features and the reference structural features, the association loss is determined.

[0218] As one embodiment, the sample structural features shared between the sample image and at least one sample transformed image can be represented by the covariance between the sample image and each sample transformed image; the reference structural features shared between the structural features of at least one reference image and the corresponding reference transformed image can be represented by the covariance between each reference image and the corresponding reference transformed image.

[0219] For example, please refer to Figure 5CTaking an MRI image as the sample image and at least one PET image as the reference image as an example, the MRI image modality is converted into a pseudo-PET image according to the image modality of the PET image; similarly, the PET image modality is converted into a pseudo-MRI image according to the image modality of the MRI image. Therefore, based on the shared sample structural features (i.e., their covariance) between the pseudo-PET image and the MRI image, and the shared reference structural features (i.e., their covariance) between the pseudo-MRI image and the PET image, and further combining the variances of the sample image, the converted sample image, the reference image, and the converted reference image, the correlation loss is determined. This correlation loss L is calculated using the variances of the MRI image, the pseudo-PET image, the PET image, and the pseudo-MRI image. Cor Please refer to formula (4).

[0220]

[0221] Among them, Cov(G PET (MRI), MRI) represents the covariance between pseudo-PET images and MRI images, Cov(G) MRI (PET), PET) represents the covariance between the pseudo-MRI image and the PET image, σ MRI Characterizing the variance of MRI images, The variance, σ, characterizes the variance of a pseudo-PET image. PET Characterizing the variance of PET images, The variance of pseudo-MRI images is characterized.

[0222] Enhanced loss:

[0223] In a modality processing model that includes a sample transformation network and a reference transformation network, where the sample transformation network is used to obtain images of different modalities from the sample image set, and the reference transformation network is used to obtain images of the same modalities as the sample image set, each sample image can be used as at least one reference image. Modality transformation is then performed using the reference transformation network to obtain the corresponding sample-enhanced image. Simultaneously, each at least one reference image can be used as a sample image, and modality transformation is performed using the sample transformation network to obtain the corresponding reference-enhanced image. The enhancement loss is determined based on the error between the obtained at least one sample-enhanced image and the sample image, and the error between the obtained at least one reference-enhanced image and its corresponding reference image.

[0224] For example, please refer to Figure 5DTaking an MRI image as the sample image and at least one PET image as the reference image as an example, the modality processing model includes a sample transformation network and a reference transformation network. Using the MRI image as the PET image, the reference transformation network performs modality transformation on the MRI image according to its image modality to obtain an enhanced MRI image. Similarly, using the PET image as the MRI image, the sample transformation network performs modality transformation on the PET image according to its image modality to obtain an enhanced PET image.

[0225] The enhancement loss is determined based on the errors between enhanced MRI images and between enhanced PET images. For example, the weighted sum of the L1 norms of the errors between enhanced MRI images and between enhanced PET images is used as the enhancement loss.

[0226] As one embodiment, when determining enhancement loss, the L1 norm of the error between enhanced MRI images and the L1 norm of the error between enhanced PET images can also be combined, and the determination can be based on a cyclic enhancement loss network, which can be referred to the above description and will not be repeated here.

[0227] Continuing with the example of using MRI images as sample images and at least one PET image as a reference image, the image classification model focuses more on the sample transformation network that generates pseudo-PET images from MRI images. Therefore, λ can be further adjusted. idenPET and λ idenMRI This is used to adjust the weight of the sample transformation network and the reference transformation network, respectively, in the two generation paths.

[0228] Therefore, based on the errors between enhanced MRI images and between enhanced PET images, the enhancement loss L is determined. identity Please refer to formula (5).

[0229]

[0230] Among them, represents the error between enhanced PET images, and represents the error between enhanced MRI images; ∑ i λ i ‖F i (G MRI (MRI)-F i (MRI)‖1 Characterization test: the sum of errors between enhanced MRI images and MRI images passing through each layer of the cyclic enhancement loss network; ∑ i λ i ‖Fi (G PET (PET))-F i (PET)‖1 characterizes and verifies the sum of errors of PET images and PET images when passed through each layer of a recurrent reinforcement loss network.

[0231] As one example, after obtaining various training losses such as adversarial loss, recurrence loss, correlation loss, reinforcement loss, and classification loss, if it is determined that there is a training loss among the obtained training losses that does not meet its corresponding training objective, the model parameters of the modality processing model and the image classification model can be adjusted based on each training loss, and the next round of iterative training can be entered until it is determined that each obtained training loss meets its corresponding training objective.

[0232] Multiple training losses can be weighted and summed to obtain a comprehensive training loss. If the comprehensive training loss does not meet the comprehensive training objective, the model parameters of the modality processing model and the image classification model are adjusted based on the comprehensive training loss, and the next round of iterative training is started until the comprehensive training loss meets the comprehensive training objective.

[0233] Taking various training losses, including adversarial loss, recurrence loss, correlation loss, reinforcement loss, and classification loss, as an example, since the recurrence loss and reinforcement loss have already been weighted according to different training paths in the aforementioned calculation process, we can simply set the corresponding weight coefficients for adversarial loss, correlation loss, and classification loss respectively, thereby comprehensively calculating the training loss L. total Please refer to formula (6).

[0234] L total =λ GAN L GAN +λ cor L cor +L Cycle +L identity +λ classify L classify (6)

[0235] Where, λ GAN Characterizing adversarial loss L GAN The weighting coefficient, λ cor Characterizing correlation loss L cor The weighting coefficient, λ classify Characterization classification loss L classify The weighting coefficients.

[0236] As one example, after obtaining a trained target image classification model, the target image classification model can be used to perform image classification tasks. In the process of using the target image classification model, since the target image classification model has learned how to extract the common features of multiple image modalities in the image to be classified, it can accurately identify the semantic expression of the image to be classified. Therefore, it is not necessary to combine it with a modality processing model to achieve accurate image classification tasks.

[0237] The following describes the process of using a target image classification model. Please refer to [link / reference]. Figure 6A .

[0238] S601, Obtain the image to be classified.

[0239] Since the target image classification model is trained by classifying sample images, when using the target image classification model to classify images, the image modality of the image to be classified is the same as the image modality of the sample image set.

[0240] Taking the scenario of diagnosing Alzheimer's disease as an example, the images to be classified, such as MRI images or PET images, are all 3D images. Therefore, the images to be classified can be regarded as containing multiple image slices, and the image slices are 2D images.

[0241] When diagnosing Alzheimer's disease, not every image slice in the images to be classified is helpful for diagnosis. In a subject with Alzheimer's, only one or a few image slices in the images to be classified are highly correlated with the disease. Due to individual differences, the image slices highly correlated with the disease vary from person to person in the images to be classified for each Alzheimer's patient. If medical professionals were to review each image slice individually, not only would the diagnostic efficiency be low, but under high-intensity work conditions, misdiagnosis would be easily made, such as diagnosing an Alzheimer's patient as a normal person, or vice versa.

[0242] Therefore, the target image classification model obtained by using the training method described in the embodiments of this application can accurately classify each image slice, thereby locking one or several image slices that are highly related to the disease among many image slices. This not only improves diagnostic efficiency but also reduces the workload of medical workers and avoids misjudgment.

[0243] S602 uses a target image classification model to extract features from multiple image slices to obtain corresponding slice feature maps, and then merges the obtained multiple slice feature maps into an image feature map.

[0244] After obtaining the image slices contained in the image to be classified, a target image classification model can be used to extract features from each image slice to obtain the corresponding slice feature map. The target image classification model can include a feature extraction network that treats each image slice as a group of image slices and uses grouped convolution to extract features.

[0245] For example, please refer to Figure 6B The feature extraction network consists of a first deep convolutional layer, a second inverted network layer, a third inverted network layer, a fourth inverted network layer, and a fifth inverted network layer. The activation function can be the ReLU6 activation function.

[0246] The first depthwise convolutional layer consists of multiple convolutional layers formed by 7×7 depthwise convolutional kernels, and multiple max pooling layers with a span of 2. The first depthwise convolutional layer performs the first round of feature extraction on this set of image slices, obtaining the first set of slice feature maps.

[0247] For example, a set of image slices has dimensions D×1×W×H, such as 256×1×256×256. The first depthwise convolutional layer contains 256 convolutional layers formed by 7×7 depthwise convolutional kernels, and multiple max-pooling layers with a span of 2. The feature map is quickly compressed through the first depthwise convolutional layer, resulting in the first set of slice feature maps with dimensions of... That is, 256×1×64×64.

[0248] The second inverted network layer contains multiple groups of unpadded convolutional layers with 1×1 kernels and a span of 1, multiple groups of convolutional layers with 1 padding, 3×3 kernels, and a span of 1, and multiple groups of unpadded convolutional layers with 1×1 kernels and a span of 1. The second inverted network layer performs a second round of feature extraction on the first group of slice feature maps, and the second group of slice feature maps is obtained based on the sum of the extracted feature maps and the first group of slice feature maps.

[0249] For example, the size of the first set of slice feature maps is For example, 256×1×64×64. Taking the second inverted network layer as an example to illustrate the structure of an inverted network layer, please refer to [reference needed]. Figure 6C The structure of other inverted network layers is similar and will not be described in detail here. The second inverted network layer contains 256 groups of unpadded convolutional layers with 1×1 kernels and a span of 1, which increases the D of the first group of slice feature maps to 8 times its original value; then it connects to 256 groups of grouped convolutional layers with 1 padding, 3×3 kernels, and a span of 1, without changing the spatial resolution, i.e., without changing the size of the feature maps. Feature extraction is performed under the premise of [missing information]. Then, 256 groups of unpadded convolutional layers with 1×1 kernels and a span of 1 are connected. D is restored to the same number of slice feature maps as the first group. Based on the sum of the extracted feature maps and the first group of slice feature maps, the second group of slice feature maps is obtained, with a size of [missing information]. That is, 256×1×64×64.

[0250] The third inverted network layer is the same as the second inverted network layer, so it will not be described again here. The third set of slice feature maps can be obtained through the third inverted network layer.

[0251] For example, the size of the second set of slice feature maps is For example, if the size is 256×1×64×64, then the third group of slice feature maps has the following dimensions: That is, 256×1×64×64.

[0252] The fourth inverted network layer contains multiple groups of unpadded convolutional layers with 1×1 kernels and a span of 2, multiple groups of convolutional layers with 1 padding, 3×3 kernels, and a span of 2, and multiple groups of unpadded convolutional layers with 1×1 kernels and a span of 2. This fourth inverted network layer performs a fourth round of feature extraction on the third group of slice feature maps, and the fourth group of slice feature maps is obtained based on the sum of the extracted feature maps and the third group of slice feature maps.

[0253] For example, the size of the third set of slice feature maps is For example, in a 256×1×64×64 network, the fourth inverted network layer contains 256 groups of unpadded, 1×1 kernel, and 2-span grouped convolutional layers, which increases the D of the third group of slice feature maps to 8 times its original value. Then, it connects to 256 groups of 1-padded, 3×3 kernel, and 2-span grouped convolutional layers, reducing both the height and width of the spatial resolution by a factor of 2, thus becoming... Feature extraction is performed under the premise of [missing information]. Then, 256 groups of unpadded convolutional layers with 1×1 kernels and a span of 2 are connected. D is restored to the same number of slice feature maps as the third group. Based on the sum of the extracted feature maps and the third group slice feature maps, a fourth group slice feature map is obtained, with a size of [missing information]. That is, 256×1×32×32.

[0254] The fifth inverted network layer is the same as the fourth inverted network layer, so it will not be described again here. Through the fifth inverted network layer, the fifth set of slice feature maps can be obtained.

[0255] For example, the size of the fourth set of slice feature maps is For example, if the size is 256×1×32×32, then the fifth group of slice feature maps has the following dimensions: That is, 256×1×16×16.

[0256] After obtaining multiple slice feature maps, a target image classification model can be used to fuse these slice feature maps into a single image feature map. These multiple slice feature maps can be the fifth set of slice feature maps mentioned earlier. The target image classification model can also include a feature fusion network to fuse the multiple slice feature maps into a single image feature map.

[0257] Please refer to Figure 6D The feature fusion network can contain pointwise convolutional layers with 1×1 kernels, and the activation function can be the ReLU6 activation function. Multiple slice feature maps are fused into an image feature map. For example, multiple slice feature maps can be combined into a group of slice feature maps, i.e., the fifth group of slice feature maps mentioned above, with a size of... For example, an image with dimensions of 256×1×16×16 can be obtained through a feature fusion network, resulting in an image feature map with dimensions of [missing information]. That is, 1×1×16×16.

[0258] S603 employs an attention mechanism, which evaluates the importance of multiple image slices based on multiple slice feature maps and image feature maps, and obtains the corresponding importance evaluation values.

[0259] After obtaining multiple slice feature maps and image feature maps, an attention mechanism can be used to evaluate the importance of multiple image slices based on the multiple slice feature maps and image feature maps, and obtain the corresponding importance evaluation values.

[0260] Please refer to Figure 6E Multiple slice feature maps are grouped into an ordered set of slice feature maps, such as the fifth set of slice feature maps mentioned above, with a size of [size missing]. For example, a set of slice feature maps with dimensions 256×1×16×16 can be reshaped to obtain a transformed set of slice feature maps with dimensions D×1×W×1, such as 256×1×256×1. The image feature map dimensions mentioned above are... For example, if the image feature map is 1×1×16×16, its shape can be transformed to obtain a transformed image feature map with a size of 1×1×W×1, such as 1×1×256×1.

[0261] The target image classification model also includes a softmax activation network. The transformed set of slice feature maps and the transformed image feature map can be input into the softmax activation network for dot product operation to obtain the importance evaluation vector output by the softmax activation network. Each element in the importance evaluation vector represents the importance evaluation value of the corresponding image slice. For example, the first element is the importance evaluation value of the first image slice in a set of ordered image slices.

[0262] For example, in the context of diagnosing Alzheimer's disease, the importance assessment value can characterize whether the image slice is highly correlated with the disease. For instance, a higher importance assessment value indicates that the image slice is highly correlated with the disease, while a lower importance assessment value indicates that the image slice is less correlated with the disease.

[0263] As one embodiment, a target image classification model can be used to classify the aforementioned sample image set, i.e., the image data in the ADNI database, to obtain the importance assessment value of each image slice contained in each sample image. Based on the average importance assessment value of the image slices in each order of all sample images, it is possible to determine which image slices are generally highly related to the disease. Thus, medical workers can directly view the corresponding image slices in the images of newly added cases and make diagnoses based on the image slices that are generally highly related to the disease, which can also improve diagnostic efficiency.

[0264] Each sample image contains image slices arranged in the order they make up the sample image. For example, if each sample image contains 256 image slices, please refer to [reference needed]. Figure 6F This represents the average importance assessment value for each image slice within the ADNI database, ranked according to its position. The figure shows that for all sample images, the average importance assessment value for slice 69 ranges from 0.2 to 0.4, while the average importance assessment value for slice 120 ranges from 0 to 0.2. The average importance assessment value is primarily concentrated between slices 69 and 171. Therefore, medical professionals can directly examine slices 69 to 171 of newly diagnosed cases for diagnosis, thus aiding in medical diagnostics.

[0265] As one embodiment, after obtaining the images to be classified for a new case, a target image classification model can be used to classify the images for that case. This process obtains image slices highly relevant to the disease from the images to be classified for that case. From multiple image slices, at least one target image slice with an importance assessment value greater than a preset assessment threshold is selected. The image classification interface displays at least one target image slice and its corresponding importance assessment value. This allows medical professionals to make targeted diagnoses for the case, avoiding situations where, due to individual differences or other reasons, the image slices highly relevant to the disease in certain cases differ from those generally considered highly relevant to the disease, thus improving diagnostic accuracy.

[0266] S604: Multiple slice feature maps, image feature maps, and multiple importance evaluation values ​​are fused and stitched together to obtain the comprehensive features of the image to be classified, and the target classification of the image to be classified is determined based on the comprehensive features.

[0267] After obtaining the importance assessment values, a target image classification model can be used to fuse and stitch together multiple slice feature maps, image feature maps, and the obtained importance assessment values ​​to obtain the comprehensive features of the image to be classified. Based on the comprehensive features, the target classification of the image to be classified can be determined. The fusion and stitching can be performed by converting the multiple slice feature maps, image feature maps, and importance assessment values ​​into a unified format before stitching; alternatively, after converting to a unified format, a weighted sum can be performed, and the sum can be stitched together with the converted slice feature maps, image feature maps, and importance assessment values, etc., with no specific restrictions.

[0268] As one embodiment, the target image classification model may also include a fusion and stitching network. For details on the fusion and stitching network, please refer to... Figure 6G The previously mentioned multiple slice feature maps can be combined into a set of slice feature maps, namely the fifth set of slice feature maps. The transformed set of slice feature maps is then multiplied with the importance evaluation vector composed of each importance evaluation value to obtain the fused feature vector.

[0269] By fusing and stitching networks, the fused feature vectors can be stitched together with the transformed image feature maps obtained by transforming their shapes to obtain the comprehensive features of the image to be classified.

[0270] The target image classification model can also include a classification network, which can determine the target category of the image based on comprehensive features. In the scenario of diagnosing Alzheimer's disease, the target classification can include the probability that the human brain in the image has Alzheimer's disease, and the probability that the human brain in the image belongs to a normal person.

[0271] The following example, using the scenario of diagnosing Alzheimer's disease, illustrates the method for training an image classification model in this application.

[0272] The sample image set consists of all MRI images, and at least one reference image set consists of all PET images. The modality processing model includes a sample transformation network and a reference transformation network, a sample discrimination network and a reference discrimination network.

[0273] This section uses the process of iteratively training a modality processing model and an image classification model on a single sample image, namely an MRI image, as an example. Please refer to [link / reference needed]. Figure 7A .

[0274] MRI images are input into a sample transformation network. This network, based on the image modality of PET images, converts the MRI image modality into a sample-transformed image, i.e., a pseudo-PET image. PET images are then input into a reference transformation network. This network, based on the image modality of MRI images, converts the PET image modality into a reference-transformed image, i.e., a pseudo-MRI image.

[0275] After obtaining pseudo-PET and pseudo-MRI images, the MRI and pseudo-PET images are input into a sample discriminant network to predict the discrimination probability of whether the corresponding input image was obtained through modality transformation. Similarly, the PET and pseudo-MRI images are input into a reference discriminant network to predict the discrimination probability of whether the corresponding input image was obtained through modality transformation. Therefore, based on the discrimination probabilities of the MRI image, pseudo-PET image, PET image, and pseudo-MRI image, the adversarial loss L of the modality processing model can be determined. GAN .

[0276] After obtaining pseudo-PET and pseudo-MRI images, the pseudo-PET images are input into a reference transformation network. This network, based on the image modality of the MRI images, transforms the pseudo-PET image modality into a reference verification image, i.e., a verification MRI image. Similarly, the pseudo-MRI images are input into a sample transformation network. Again, using the reference transformation network and based on the image modality of the PET images, the pseudo-MRI image modality is transformed into a sample verification image, i.e., a verification PET image. Therefore, the cyclic loss L can be determined based on the errors between the verification MRI images and the MRI images, as well as the errors between the verification PET images. Cycle .

[0277] After obtaining pseudo-PET and pseudo-MRI images, common sample structural features between the MRI and pseudo-PET images are extracted, as well as common reference structural features between the PET and pseudo-MRI images. Based on these sample and reference structural features, the correlation loss L can be determined. Cor .

[0278] MRI images are used as PET images and input into a reference transformation network. This network performs mode transformation on the MRI images according to their image modalities to obtain enhanced MRI images. Similarly, PET images are used as MRI images and input into a sample transformation network. This network performs mode transformation on the PET images according to their image modalities to obtain enhanced PET images. Therefore, the enhancement loss L can be determined based on the errors between enhanced MRI images and between enhanced PET images. identity .

[0279] After obtaining the pseudo-PET image, data distillation can be performed on the MRI image and the pseudo-PET image to obtain the common features between the MRI image and the pseudo-PET image. The common features are extracted from the MRI image and input into the image classification model. Based on the common features, the image classification model determines the predicted classification of the MRI image. The predicted classification includes the predicted probability that the human brain in the MRI image has Alzheimer's disease and the predicted probability that the human brain in the MRI image belongs to a normal person.

[0280] Taking the training of an image classification model using a sample image set obtained from the ADNI database as an example, in order to quantitatively analyze the training process, the training data can be used to train the image classification model, and the training target image classification model can be tested using test data. The test data can be obtained from the ADNI database, or it can be obtained from real cases in a certain region within a certain time period, etc., without any specific restrictions.

[0281] In the quantitative analysis of the training process, three standard performance indicators can be used to quantitatively analyze the quality effect of the modality processing model during modality conversion: Mean Absolute Error (MAE), Peak Signal-to-Noise Ratio (PSNR), and Structural Similarity Index (SSIM). Among them, MAE and PSNR represent the pixel-level matching degree between the modality-converted image and the real image, such as the pixel-level matching degree between the pseudo PET image and the real PET image; SSIM represents the structural similarity between the modality-converted image and the real image, such as the structural similarity between the pseudo PET image and the real PET image.

[0282] For the calculation method of MAE, please refer to formula (7); for the calculation method of PSNR, please refer to formula (8); for the calculation method of SSIM, please refer to formula (9).

[0283]

[0284]

[0285]

[0286] Among them, y j Representing the j-th reference image or the j-th sample image, y j ′ The sample transformed image represents the j-th sample image after mode transformation, or the reference transformed image represents the j-th reference image after mode transformation. K represents the number of images in the reference image set or the number of images in the sample image set.

[0287] MSE represents the mean square error between the original image and the modality-transformed image, and n represents the number of pixels.

[0288] μ x μ represents the average pixel value of image x. y This represents the average pixel value of image y. The pixel variance characterizing image x. The pixel variance σ of image y is a characteristic of the image. xy The covariance of image x and image y is represented by c1 and c2, respectively, which are pixel-related parameters.

[0289] When quantitatively analyzing the training process, two performance metrics, accuracy (ACC) and F1 score, can be used to measure the classification accuracy of the image classification model. ACC reflects the ability of the image classification model to accurately identify both Alzheimer's disease and normal individuals, while the F1 score reflects the average ability of the image classification model in terms of precision and recall.

[0290] For the calculation method of ACC, please refer to formula (10), and for the calculation method of F1 score, please refer to formula (11).

[0291]

[0292]

[0293] Among them, TP represents the probability of predicting Alzheimer's disease from sample images of Alzheimer's patients; TN represents the probability of predicting normal people from sample images of normal people; FP represents the probability of predicting Alzheimer's disease from sample images of normal people; and FN represents the probability of predicting normal people from sample images of Alzheimer's patients.

[0294] Taking the sample image set obtained from the ADNI database as an example, the following quantitative analysis and comparison are performed on the image classification model A, trained based on adversarial loss, cyclic loss, and reinforcement loss provided in the embodiments of this application, wherein the cyclic loss and reinforcement loss are calculated using only the L1 norm without combining the cyclic reinforcement loss network calculation; the image classification model B, trained based on adversarial loss, cyclic loss, correlation loss, and reinforcement loss, wherein the cyclic loss and reinforcement loss are calculated using only the L1 norm without combining the cyclic reinforcement loss network calculation; the image classification model C, trained based on adversarial loss, cyclic loss, correlation loss, reinforcement loss, and reinforcement loss, wherein the cyclic loss and reinforcement loss are calculated using both the L1 norm and the cyclic reinforcement loss network calculation; and the image classification model D, trained based on adversarial loss, cyclic loss, correlation loss, reinforcement loss, and classification loss, wherein the cyclic loss and reinforcement loss are calculated using both the L1 norm and the cyclic reinforcement loss network calculation. Please refer to Table 2 for details.

[0295] Table 2

[0296]

[0297] Among them, because the image classification model focuses more on the process of generating pseudo-PET images from MRI images, therefore in L Cycle Set λ in CyclePET and λ CycleMRI The values ​​are 15 and 10 respectively, while λ idenPET and λ idenMRI All values ​​are set to 2. Since the three performance metrics of image classification model B are roughly the same as those of image classification model A, λ is set to 2. cor The result is 0.1. For image classification model C, all three metrics show significant improvements, with λ1 = 0.1 and λ2 = 0.05 in the recurrent reinforcement loss network. For image classification model D, generating pseudo-PET images yields the best results, and the introduction of the classification task also improves the accuracy of pseudo-PET image generation.

[0298] Please refer to Figure 7BThis is a qualitative comparative analysis of image classification models A, B, C, and D. The first column contains real MRI images, showing human brain images from different perspectives. The second column contains pseudo-PET images A obtained by modal processing model A (trained together with image classification model A) after modal transformation of real MRI images. The third column contains pseudo-PET images B obtained by modal processing model B (trained together with image classification model B) after modal transformation of real MRI images. The fourth column contains pseudo-PET images C obtained by modal processing model C (trained together with image classification model C) after modal transformation of real MRI images. The fifth column contains pseudo-PET images D obtained by modal processing model D (trained together with image classification model D) after modal transformation of real MRI images. The sixth column contains real PET images, showing human brain images from different perspectives.

[0299] For modal processing model B, compared to modal processing model A, the texture features of the generated pseudo-PET images are more concentrated in specific areas, such as near the ventricles, which are highly correlated with the disease. For modal processing model C, the global features of the generated pseudo-PET images are more prominent. For modal processing model D, the generated pseudo-PET images are the most realistic and have the best effect.

[0300] Although the generated pseudo-PET images still differ significantly from real PET images, due to the different generation principles of real MRI and real PET images, pseudo-PET images generated by any method cannot reflect the functional characteristics of real PET images. Therefore, by combining a modality processing model and an image classification model for training, the focus is more on enabling the image classification model to extract shared low-level feature information from the generated pseudo-PET images and real MRI images. This allows the image classification model to achieve relatively good image classification results even when only real MRI images are input.

[0301] Taking the sample image set obtained from the ADNI database as an example, a quantitative analysis and comparison are performed on the ResNet18 model in related technologies; the image classification model E trained based on adversarial loss, cyclic loss, correlation loss and reinforcement loss provided in the embodiments of this application; and the image classification model F trained based on adversarial loss, cyclic loss, correlation loss, reinforcement loss and classification loss. Please refer to Table 3.

[0302] Table 3

[0303]

[0304] Among them, the two performance metrics of image classification model E are slightly lower than those of the traditional ResNet18 model. Due to the introduction of depthwise separable convolution, the number of parameters in image classification model F is significantly smaller than that of the ResNet18 model. Image classification model F further mines the underlying brain imaging information shared by MRI images and pseudo-PET images, thus its two performance metrics surpass those of other models.

[0305] The following example, using the scenario of diagnosing Alzheimer's disease, illustrates the method of using a trained target image classification model in this application.

[0306] The target image classification model consists of a feature extraction network, a feature fusion network, a fusion and stitching network, and a classification network. Please refer to [reference needed]. Figure 7C This paper will take the process of classifying an image using a target image classification model, specifically an MRI image, as an example.

[0307] MRI images are 3D images; therefore, multiple 2D image slices can be extracted from them. Taking the extraction of 256 image slices from an MRI image, each with a resolution of 256×256, as an example, since MRI images are black and white, each image slice is also black and white, meaning each slice has 1 color channel. The slices can be arranged according to their position within the MRI image. Thus, the size of each image slice can be represented as (256, 1, 256, 256), that is, 256 image slices with 1 color channel and a resolution of 256×256.

[0308] The MRI images are input into the feature extraction network of the target image classification model. The feature extraction network can be referred to in the previous introduction and will not be repeated here. The feature extraction network extracts features from each image slice to obtain the fifth set of slice feature maps. The size of the fifth set of slice feature maps is (256×1×16×16), that is, 256 color channels with a resolution of 16×16.

[0309] The fifth set of slice feature maps is input into the feature fusion network of the target image classification model. The feature fusion network can be referred to in the previous introduction and will not be repeated here. The feature fusion network performs feature fusion on the fifth set of slice feature maps to obtain an image feature map. The size of the image feature map is (1×1×16×16), that is, an image feature map with 1 color channel and a resolution of 16×16.

[0310] The fifth set of slice feature maps and the image feature map are input into a fusion and stitching network. The fusion and stitching network can be referred to in the previous description and will not be repeated here. The fusion and stitching network first converts the size of the fifth set of slice feature maps to (256, 1, 256), that is, a 256×256 matrix with 1 color channel. Then, it converts the size of the image feature map to (1, 256), that is, a vector containing 256 elements with 1 color channel. Next, a dot product is performed on the transformed fifth set of slice feature maps and the transformed image feature map to obtain the importance evaluation value of each image slice. This importance evaluation value forms an importance evaluation vector according to the order of the image slices.

[0311] Finally, the product of the deformed fifth group of slice feature maps and the importance evaluation vector is concatenated with the deformed image feature map to obtain the comprehensive features of the MRI image.

[0312] The comprehensive features are input into a classification network, which then determines the target classification of the MRI image based on these features. The target classification represents the probability that the MRI image contains a human brain with Alzheimer's disease, and the probability that the MRI image contains a human brain that is normal.

[0313] Please refer to Figure 7D (1) This is a visualization of the slice feature maps corresponding to several image slices of an Alzheimer's patient, including the visualization results of the slice feature maps for image slices 69, 77, 85, 87, 89, 123, 167, 171, 181, and 187. The highlighted areas in the figure represent the regions that the target image classification model focuses on during image classification or feature extraction. It can be seen that because Alzheimer's disease is often accompanied by varying degrees of ventricular atrophy, the target image classification model, through learning, pays more attention to some ventricular regions related to the disease. Please refer to the visualization results of the slice feature map for image slice 85. Figure 7D (2) is an enlarged image of the 85th image. Referring to the two elliptical areas marked in the image, it can be seen that the target image classification model pays more attention to the vicinity of the hippocampus, which is the area highly related to the disease.

[0314] The target image classification model in this embodiment differs from the traditional integration of generation and classification tasks. This embodiment utilizes a modal processing model to assist in training the target image classification model. This model helps the target image classification model mine low-level semantic features from MRI images. When performing image classification, only the target image classification model needs to be deployed to achieve good results; the modal processing model is not required. Furthermore, the target image classification model borrows the idea of ​​depthwise separable convolution, achieving good classification accuracy while having a relatively small number of parameters, making it suitable for real-time edge computing deployment on mobile devices. The obtained importance assessment values ​​can simulate the doctor's consultation process, avoiding misjudgments or inefficiencies that can occur when doctors subjectively select disease-related image slices from numerous image slices. In addition to providing classification results, the importance assessment values, which have interpretable magnitudes, can provide more auxiliary diagnostic functions, helping doctors make quick and accurate diagnoses.

[0315] Based on the same inventive concept, embodiments of this application provide an apparatus for training an image classification model, capable of achieving the functions corresponding to the aforementioned method for training an image classification model. Please refer to... Figure 8 The device includes an acquisition module 801 and a processing module 802, wherein:

[0316] Acquisition module 801: used to acquire training data containing a set of sample images and at least one set of reference images, wherein different sets of reference images have different image modalities, and each sample image in each set of reference images is associated with a reference image containing the same target;

[0317] Processing module 802: Based on the obtained training data and combined with the modality processing model, it performs multiple rounds of iterative training on the image classification model to be trained to obtain the target image classification model. Each iteration performs the following operations:

[0318] The processing module 802 is specifically used to: for the selected sample image, adopt a modality processing model, and convert the sample image modality into the corresponding sample transformed image according to the image modality of at least one associated reference image;

[0319] The processing module 802 is specifically used to: use an image classification model to extract common features of the sample image and at least one sample transformed image obtained, and determine the predicted classification of the sample image based on the common features;

[0320] The processing module 802 is specifically used to: adjust the model parameters of the modality processing model and the image classification model based on at least one sample converted image, at least one reference image, and prediction classification.

[0321] In one possible embodiment, the modality processing model includes at least one sample transformation network, which outputs images of different image modalities than the sample image set, with different sample transformation networks outputting images corresponding to different image modalities; the processing module 802 is specifically used for:

[0322] For each sample transformation network, perform the following operations:

[0323] A sample transformation network is used to perform multi-scale downsampling on the selected sample images to obtain multiple intermediate feature maps, each with a different resolution.

[0324] Based on the resolution of the sample image, multiple intermediate feature maps are upsampled and fused to obtain the sample-transformed image.

[0325] In one possible embodiment, the processing module 802 is specifically used for:

[0326] Data distillation is performed on the sample image and at least one transformed sample image to determine the common features between the sample image and at least one transformed sample image;

[0327] An image classification model is used to extract common features from sample images, and the predicted classification of the sample images is determined based on these common features.

[0328] In one possible embodiment, the processing module 802 is specifically used for:

[0329] Based on at least one sample transformed image and at least one reference image, the adversarial loss, cyclic loss, correlation loss and reinforcement loss of the modality processing model are determined. The adversarial loss characterizes the realism of the image after modality transformation, the cyclic loss characterizes the reproducibility of the image after modality transformation, the correlation loss characterizes the consistency of the image before and after modality transformation, and the reinforcement loss characterizes the accuracy of modality transformation when the image before and after modality transformation is the same image modality.

[0330] Based on the predicted classification, determine the classification loss of the image classification model;

[0331] The model parameters of the modality processing model and the image classification model are adjusted based on adversarial loss, cyclic loss, correlation loss, reinforcement loss, and classification loss.

[0332] In one possible embodiment, the modality processing model includes a discriminant network, which is used to: predict the discriminant probability of the input image obtained through modality transformation; the processing module 802 is specifically used to:

[0333] The sample image, at least one transformed sample image, and at least one reference image are used as input images for the discrimination network to predict the corresponding discrimination probabilities.

[0334] Based on the error between each obtained discrimination probability and the reference probability associated with its corresponding input image, the adversarial loss is determined, and based on at least one sample transformed image and at least one reference image, the recurrence loss, correlation loss and reinforcement loss of the modality processing model are determined.

[0335] In one possible embodiment, the modality processing model includes a sample transformation network and a reference transformation network. The sample transformation network is used to obtain images of image modalities different from the sample image set, and the reference transformation network is used to obtain images of image modalities identical to the sample image set. The processing module 802 is specifically used for:

[0336] A reference transformation network is used to transform at least one reference image mode into a corresponding reference transformation image according to the image mode of the sample image;

[0337] A sample transformation network is used to transform at least one reference transformed image modality into a corresponding verification transformed image according to the image modality of at least one reference image.

[0338] Based on the error between each verified transformed image and the sample image, the cyclic loss is determined, and based on at least one sample transformed image and at least one reference image, the adversarial loss, correlation loss and reinforcement loss of the modal processing model are determined.

[0339] In one possible embodiment, at least one reference image is associated with a corresponding reference transformed image, which is obtained by performing mode transformation on the corresponding reference image according to the image mode of the sample image; the processing module 802 is specifically used for:

[0340] Extract the common sample structure features between the sample image and at least one sample transformed image, and extract the common reference structure features between at least one reference image and the corresponding reference transformed image;

[0341] Based on sample structural features and reference structural features, the correlation loss is determined, and based on at least one sample transformed image and at least one reference image, the adversarial loss, cyclic loss and reinforcement loss of the modal processing model are determined.

[0342] In one possible embodiment, the modality processing model includes a sample transformation network and a reference transformation network. The sample transformation network is used to obtain images of image modalities different from the sample image set, and the reference transformation network is used to obtain images of image modalities identical to the sample image set. The processing module 802 is specifically used for:

[0343] Each sample image is used as at least one reference image, and a reference transformation network is used to perform modality transformation to obtain the corresponding sample enhanced image.

[0344] At least one reference image is used as a sample image, and a sample transformation network is used to perform mode transformation to obtain the corresponding reference enhancement image;

[0345] Based on the error between at least one sample augmented image and the sample image, and the error between at least one reference augmented image and its corresponding reference image, the augmentation loss is determined, and based on at least one sample transformed image and at least one reference image, the adversarial loss, cyclic loss and correlation loss of the modal processing model are determined.

[0346] In one possible embodiment, the processing module 802 is specifically used for:

[0347] If any of the obtained training losses does not meet its corresponding training objective, the model parameters of the modality processing model and the image classification model are adjusted based on each training loss, and the next round of iterative training is started until it is determined that each obtained training loss meets its corresponding training objective.

[0348] The training losses include adversarial loss, cyclic loss, correlation loss, reinforcement loss, and classification loss.

[0349] In one possible embodiment, the processing module 802 is further configured to:

[0350] Based on the obtained training data, combined with the modality processing model, the image classification model to be trained is iterated and trained in multiple rounds to obtain the target image classification model. After obtaining the target image classification model, the image to be classified is obtained. The image modality of the image to be classified is the same as the image modality of the sample image set, and the image to be classified contains multiple image slices.

[0351] A target image classification model is used to extract features from multiple image slices to obtain corresponding slice feature maps, and the obtained multiple slice feature maps are then fused into an image feature map.

[0352] An attention mechanism is employed to evaluate the importance of multiple image slices based on multiple slice feature maps and image feature maps, thereby obtaining corresponding importance evaluation values.

[0353] Multiple slice feature maps, image feature maps, and multiple importance evaluation values ​​are fused and stitched together to obtain the comprehensive features of the image to be classified, and the target classification of the image to be classified is determined based on the comprehensive features.

[0354] In one possible embodiment, the processing module 802 is further configured to:

[0355] After determining the target classification of the image to be classified based on comprehensive features, at least one target image slice with an importance evaluation value greater than a preset evaluation threshold is selected from multiple image slices.

[0356] The image classification interface presents at least one target image slice and its corresponding importance assessment value.

[0357] Please refer to Figure 9 The aforementioned apparatus for training the image classification model can run on a computer device 900. The current and historical versions of the data storage program, as well as the application software corresponding to the data storage program, can be installed on the computer device 900, which includes a processor 980 and a memory 920. In some embodiments, the computer device 900 may include a display unit 940, which includes a display panel 941 for displaying a user-interactive interface, etc.

[0358] In one possible embodiment, the display panel 941 may be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED).

[0359] The processor 980 is used to read a computer program and then execute the methods defined by the computer program. For example, the processor 980 reads a data storage program or file, thereby running the data storage program on the computer device 900 and displaying the corresponding interface on the display unit 940. The processor 980 may include one or more general-purpose processors, and may also include one or more DSPs (Digital Signal Processors) for performing related operations to implement the technical solutions provided in the embodiments of this application.

[0360] The memory 920 generally includes main memory and secondary storage. Main memory can be random access memory (RAM), read-only memory (ROM), and cache, etc. Secondary storage can be a hard disk, optical disk, USB flash drive, floppy disk, or magnetic tape drive, etc. The memory 920 is used to store computer programs and other data. The computer programs include applications corresponding to each client, and other data may include data generated after the operating system or applications are run, including system data (e.g., operating system configuration parameters) and user data. In this embodiment, program instructions are stored in the memory 920, and the processor 980 executes the program instructions in the memory 920 to implement any of the methods described in the preceding figures.

[0361] The aforementioned display unit 940 is used to receive input digital information, character information, or contact touch operations / non-contact gestures, and to generate signal inputs related to user settings and function control of the computer device 900. Specifically, in this embodiment, the display unit 940 may include a display panel 941. The display panel 941, for example, is a touch screen, which can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or on the display panel 941), and drive corresponding connection devices according to a pre-set program.

[0362] In one possible embodiment, the display panel 941 may include two parts: a touch detection device and a touch controller. The touch detection device detects the player's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 980. It can also receive and execute commands from the processor 980.

[0363] The display panel 941 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the display unit 940, in some embodiments, the computer device 900 may also include an input unit 930. The input unit 930 may include an image input device 931 and other input devices 932, wherein the other input devices may include, but are not limited to, one or more of the following: a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick.

[0364] In addition to the above, the computer device 900 may also include a power supply 990 for powering other modules, an audio circuit 960, a near-field communication module 970, and an RF circuit 910. The computer device 900 may also include one or more sensors 950, such as an accelerometer, a light sensor, and a pressure sensor. The audio circuit 960 specifically includes a speaker 961 and a microphone 962, for example, the computer device 900 can use the microphone 962 to collect the user's voice and perform corresponding operations.

[0365] As one embodiment, the number of processors 980 can be one or more, and the processors 980 and the memory 920 can be coupled together or relatively independent.

[0366] As one example, Figure 9 The processor 980 in the middle can be used to implement, for example Figure 8 The functions of the acquisition module 801 and the processing module 802 in the process.

[0367] As one example, Figure 9The processor 980 in the text can be used to implement the functions of the server or terminal devices discussed above.

[0368] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0369] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of software products, for example, through a computer program product. This computer program product is stored in a storage medium and includes several instructions to cause a computer device to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0370] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for training an image classification model, characterized in that, include: Acquire training data comprising a set of sample images and at least one set of reference images, wherein different sets of reference images have different image modalities, and each sample image is associated with a reference image containing the same target in each set of reference images; Based on the obtained training data, and combined with the modality processing model, the image classification model to be trained is subjected to multiple rounds of iterative training to obtain the target image classification model. Each round of iteration performs the following operations: For the selected sample image, the modality processing model is used to convert the sample image modality into a corresponding sample transformed image according to the image modality of at least one associated reference image; Using the image classification model, common features are extracted from the sample image and at least one obtained sample transformed image, and the predicted classification of the sample image is determined based on the common features. Using the reference transformation network in the modality processing model, the at least one reference image modality is converted into a corresponding reference transformed image according to the image modality of the sample image; the reference transformation network is used to obtain an image with the same image modality as the sample image set; Using the sample conversion network in the modality processing model, at least one reference converted image modality is converted into a corresponding verification converted image according to the image modality of the at least one reference image; the sample conversion network is used to obtain images of image modalities different from the sample image set; Based on the error between each verified transformed image and the sample image, a cyclic loss is determined; the cyclic loss characterizes the reversibility of the image after mode transformation. Based on the cyclic loss, the model parameters of the modality processing model and the image classification model are adjusted by combining the at least one sample transformed image, the at least one reference image, and the predicted classification.

2. The method according to claim 1, characterized in that, The modality processing model includes at least one sample conversion network, and the images output by different sample conversion networks correspond to different image modalities. The step of using the modality processing model to convert the selected sample image modality into a corresponding sample transformed image according to the image modality of at least one associated reference image includes: For each sample transformation network, perform the following operations: The sample transformation network is used to perform multi-scale downsampling on the selected sample images to obtain multiple intermediate feature maps, each of which has a different resolution. Based on the resolution of the sample image, the multiple intermediate feature maps are upsampled and fused to obtain the sample transformed image.

3. The method according to claim 1, characterized in that, The step of using the image classification model to extract common features from the sample image and at least one transformed sample image, and determining the predicted classification of the sample image based on the common features, includes: Data distillation is performed on the sample image and the at least one transformed sample image to determine the common features between the sample image and the at least one transformed sample image; The image classification model is used to extract the common features from the sample images, and the predicted classification of the sample images is determined based on the common features.

4. The method according to claim 1, characterized in that, The step of adjusting the model parameters of the modality processing model and the image classification model based on the cyclic loss, combined with the at least one sample transformed image, the at least one reference image, and the predicted classification, includes: Based on the at least one sample transformed image and the at least one reference image, the adversarial loss, correlation loss, and reinforcement loss of the modality processing model are determined, wherein the adversarial loss characterizes the realism of the image after modality transformation, the correlation loss characterizes the consistency of the image before and after modality transformation, and the reinforcement loss characterizes the accuracy of modality transformation when the image before and after modality transformation is the same image modality; Based on the predicted classification, the classification loss of the image classification model is determined; The model parameters of the modality processing model and the image classification model are adjusted based on the adversarial loss, the recurrence loss, the correlation loss, the reinforcement loss, and the classification loss.

5. The method according to claim 4, characterized in that, The modality processing model includes a discriminant network, which is used to: predict the discriminant probability of the input image being obtained through modality transformation; The step of determining the adversarial loss, correlation loss, and reinforcement loss of the modality processing model based on the at least one sample transformed image and the at least one reference image includes: The sample image, the at least one transformed sample image, and the at least one reference image are used as input images for the discrimination network to predict the corresponding discrimination probabilities. Based on the error between each obtained discrimination probability and the reference probability associated with its corresponding input image, the adversarial loss is determined, and based on the at least one sample transformed image and the at least one reference image, the correlation loss and reinforcement loss of the modality processing model are determined.

6. The method according to claim 4, characterized in that, The at least one reference image is associated with a corresponding reference transformation image, which is obtained by performing mode transformation on the corresponding reference image according to the image mode of the sample image; The determination of the adversarial loss, correlation loss, and reinforcement loss of the modality processing model based on the at least one sample transformed image and the at least one reference image includes: Extract the common sample structure features between the sample image and the at least one sample transformed image, and extract the common reference structure features between the at least one reference image and the corresponding reference transformed image; Based on the sample structural features and the reference structural features, the correlation loss is determined, and based on the at least one sample transformed image and the at least one reference image, the adversarial loss and reinforcement loss of the modality processing model are determined.

7. The method according to claim 4, characterized in that, The determination of the adversarial loss, correlation loss, and reinforcement loss of the modality processing model based on the at least one sample transformed image and the at least one reference image includes: Each sample image is used as the at least one reference image, and modality transformation is performed using the reference transformation network to obtain the corresponding sample enhanced image; Each of the at least one reference image is used as the sample image, and modality conversion is performed using the sample conversion network to obtain the corresponding reference enhanced image; The enhancement loss is determined based on the error between the obtained at least one sample enhanced image and the sample image, and the error between the obtained at least one reference enhanced image and its corresponding reference image. The adversarial loss and correlation loss of the modality processing model are determined based on the at least one sample transformed image and the at least one reference image.

8. The method according to claim 4, characterized in that, The step of adjusting the model parameters of the modality processing model and the image classification model based on the adversarial loss, the recurrence loss, the correlation loss, the reinforcement loss, and the classification loss includes: If, among the obtained training losses, there is a training loss that does not meet its corresponding training objective, the model parameters of the modality processing model and the image classification model are adjusted based on the training losses, and the next round of iterative training is initiated until it is determined that all the obtained training losses meet their corresponding training objectives. The training losses include the adversarial loss, the recurrence loss, the correlation loss, the reinforcement loss, and the classification loss.

9. The method according to any one of claims 1 to 8, characterized in that, After obtaining the target image classification model by performing multiple rounds of iterative training on the image classification model to be trained based on the obtained training data and combining it with the modality processing model, the process further includes: Obtain an image to be classified, wherein the image modality of the image to be classified is the same as the image modality of the sample image set, and the image to be classified contains multiple image slices; The target image classification model is used to extract features from the multiple image slices to obtain corresponding slice feature maps, and the obtained multiple slice feature maps are fused into an image feature map. An attention mechanism is used to evaluate the importance of the multiple image slices based on the multiple slice feature maps and the image feature map, and obtain the corresponding importance evaluation values. The multiple slice feature maps, the image feature map, and the obtained multiple importance evaluation values ​​are fused and stitched together to obtain the comprehensive features of the image to be classified, and the target classification of the image to be classified is determined based on the comprehensive features.

10. The method according to claim 9, characterized in that, After determining the target classification of the image to be classified based on the comprehensive features, the method further includes: From the plurality of image slices, at least one target image slice with an importance evaluation value greater than a preset evaluation threshold is selected; The image classification interface presents at least one target image slice and its corresponding importance evaluation value.

11. An apparatus for training an image classification model, characterized in that, include: Acquisition module: used to acquire training data containing a set of sample images and at least one set of reference images, wherein different sets of reference images have different image modalities, and each sample image in each set of reference images is associated with a reference image containing the same target; Processing module: Based on the obtained training data and combined with the modality processing model, this module performs multiple rounds of iterative training on the image classification model to be trained, thereby obtaining the target image classification model. Each iteration performs the following operations: The processing module is specifically used to: for the selected sample image, use the modality processing model to convert the sample image modality into a corresponding sample transformed image according to the image modality of at least one associated reference image; The processing module is specifically used to: use the image classification model to extract common features of the sample image and at least one obtained sample transformed image, and determine the predicted classification of the sample image based on the common features; The processing module is specifically used to: adjust the model parameters of the modality processing model and the image classification model based on the at least one sample converted image, the at least one reference image, and the predicted classification; The processing module is specifically used to: employ the reference transformation network in the modality processing model to convert the at least one reference image modality into a corresponding reference transformation image according to the image modality of the sample image; the reference transformation network is used to obtain an image with the same image modality as the sample image set; Using the sample conversion network in the modality processing model, at least one reference converted image modality is converted into a corresponding verification converted image according to the image modality of the at least one reference image; the sample conversion network is used to obtain images of image modalities different from the sample image set; Based on the error between each verified transformed image and the sample image, a cyclic loss is determined; the cyclic loss characterizes the reversibility of the image after mode transformation. Based on the cyclic loss, the model parameters of the modality processing model and the image classification model are adjusted by combining the at least one sample transformed image, the at least one reference image, and the predicted classification.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method as described in any one of claims 1 to 10.

13. A computer device, characterized in that, include: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the method as described in any one of claims 1 to 10 according to the obtained program instructions.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the method as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Multi-mode three-dimensional medical image fusion method and system and electronic equipment

    CN110580695A

  • COVID-19 early screening and severity degree evaluation method and system based on attention guidance

    CN111755131A

  • Visible light-infrared cross-modal pedestrian re-recognition method based on knowledge distillation

    CN112597866A

  • Generator training method and device, storage medium and electronic device

    CN113822976A