Parkinson's disease multi-modal brain template construction method and device based on semantic weak supervision guidance

By constructing a multimodal brain template based on semantic weak supervision, and through iterative optimization of the template generation subnetwork and the registration network, the problems of low efficiency and limited accuracy of brain template registration in the prior art are solved. The generated template improves the accuracy and consistency of Parkinson's disease diagnosis.

CN122391307APending Publication Date: 2026-07-14RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
Filing Date
2026-03-23
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency and limited accuracy in brain template registration, especially when multimodal image fusion results in severe information loss, making it difficult to achieve efficient and accurate diagnosis of Parkinson's disease.

Method used

A multimodal brain template construction method based on semantic weak supervision is adopted. The template generation sub-network and the registration network are used to construct the brain template through an iterative optimization process. The semantic weak supervision loss and deformation field regularization are combined to ensure that the template generation has high-quality spatial alignment and semantic information.

Benefits of technology

The generated brain templates not only possess high-quality spatial alignment properties, but also naturally embed semantic information, significantly improving the performance and usability of downstream tasks, preserving key information, and enhancing the accuracy and consistency of Parkinson's disease diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391307A_ABST
    Figure CN122391307A_ABST
Patent Text Reader

Abstract

The application relates to a Parkinson's disease multi-modal brain template construction method and equipment based on semantic weak supervision guidance, which utilizes a template generation subnetwork and a registration network to iteratively construct a brain template: a sample is randomly selected, and after pretreatment, the sample is rigidly registered to MNI52 space to initialize an initialization template; a learnable parameter image is constructed, and after being processed through the template generation subnetwork, the learnable parameter image is added to the initialization template to obtain a learning template; a double-channel image is extracted, and the initialization template and the learning template are aligned through the registration network to obtain an alignment template; a loss is calculated based on the alignment template and the sample, a discriminator is fixed, parameters of a generation network and the registration network are optimized, and a parameter image is updated; a new sample is selected, and the previous steps are repeated, a discriminator loss is calculated, other networks are fixed, and the discriminator is updated; iteration is performed to a maximum number of times, and a brain template is output. Compared with the prior art, the application realizes more accurate brain template construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer-aided diagnosis technology for medical images, and in particular to a method and device for constructing multimodal brain templates for Parkinson's disease based on semantic weak supervision guidance. Background Technology

[0002] Magnetic resonance imaging (MRI), as a radiation-free imaging technique with high soft tissue resolution, can clearly display the fine anatomical structures within the brain, particularly morphological changes in deep nuclei such as the substantia nigra and striatum. It has become an indispensable imaging tool for the diagnosis and differential diagnosis of Parkinson's disease. In clinical practice, MRI is used not only to rule out other structural lesions that may lead to Parkinson's syndrome (such as cerebrovascular disease and tumors), but also to capture characteristic changes related to neurodegenerative processes through specific imaging sequences (such as iron-sensitive sequences and neuromelanin imaging). However, the interpretation of MRI images currently still relies primarily on the visual assessment of physicians. This method is not only time-consuming but also has limited ability to identify early, subtle structural changes, is heavily influenced by the interpreter's subjective experience, and makes it difficult to guarantee diagnostic consistency among different physicians. To improve the accuracy and reproducibility of MRI-based Parkinson's disease diagnosis, constructing a standardized brain template is a crucial prerequisite. A brain template provides a unified reference space; by registering brain images from different individuals to this template, differences in brain morphology between individuals can be eliminated, thereby enabling statistical analysis at the population level and precise lesion localization. For diseases like Parkinson's disease, which involve dysfunction in specific brain regions (such as the basal ganglia and substantia nigra), brain templates can place individual image data in a standard coordinate system, enabling voxel-based statistical analysis and thus more sensitively capturing early or subtle metabolic abnormalities. By combining the brain's natural symmetry and constructing symmetrical brain templates, the differences between lesion areas and their contralateral mirror images can be systematically compared, simulating the diagnostic logic of clinicians and effectively improving the detection of small lesions or asymmetrical lesions. For example, Chinese patent application CN114334130A uses a traditional iterative optimization method based on single-channel images to achieve brain template registration. While it successfully constructs brain templates, it suffers from slow registration speed and limited accuracy. Furthermore, when multimodal or multi-contrast images are involved, the traditional approach typically uses linear addition to fuse multi-channel information into a single-channel image before sending it into the registration process. Although this approach can integrate boundary information from different images, the selection of linear coefficients lacks objective standards, easily leading to information loss.

[0003] Therefore, the purpose of this invention is to solve the problems of low brain template registration efficiency and accuracy limitations due to image fusion strategies in existing technologies. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a method and device for constructing multimodal brain templates for Parkinson's disease based on semantic weak supervision. This ensures that brain templates that can characterize the average brain state of the research population, visualize brain differences between different groups, and be used for various downstream tasks such as brain region of interest segmentation and key point localization can be generated, thereby achieving more accurate Parkinson's disease identification.

[0005] The objective of this invention can be achieved through the following technical solutions: According to a first aspect of the present invention, a method for constructing a multimodal brain template for Parkinson's disease based on semantically weakly supervised guidance is provided. The method utilizes a template generation subnetwork and a registration network to iteratively execute the following steps to construct the brain template: S1. Randomly select any multimodal image sample, perform preprocessing, rigidly register it onto the MNI52 template image, and initialize it to obtain the initialized template image; S2. Construct a learnable parameter image of a preset size, process the parameter image using a template generation sub-network, and add it to the initial template image to obtain a learned template image; S3. Extract dual-channel images from the learning template image and the initialization template image respectively. Based on the dual-channel images, use a registration network to spatially align the initialization template image and the learning template image to obtain an aligned template image. S4. Randomly select a sample image, calculate the loss function based on the alignment template image and the sample image, fix the constructed discriminator parameters, optimize the model parameters of the template generation sub-network and the registration network based on the loss function, and update the parameter image. S5. Randomly select another multimodal image sample and repeat S1~S3 in combination with the updated parameter image. Calculate the discriminator loss function, fix the template to generate sub-network and registration network parameters, and update the discriminator parameters based on the discriminator loss function. S6. Determine if the maximum number of iterations has been reached. If not, return to execute S1; otherwise, based on the parameter image updated in the current iteration, generate the brain template output from the network using the template of the current iteration.

[0006] As a preferred technical solution, the multimodal image samples include five modalities: T1, QSM, NM-MRI, tSWI, and FW, and their preprocessing is as follows: For the T1 modality, Fastsurfer was used to remove the skull from the T1 images and the four regions of white matter, gray matter, ventricles and cerebellum were labeled. For the QSM modality, the deep brain nuclei are labeled, including the caudate nucleus, putamen, globus pallidus, substantia nigra, red nucleus, and dentate nucleus; For the NM-MRI modality, substantia nigra and reference background are extracted from NM-MRI images and annotated, and then the annotated NM-MRI images are standardized. For the tSWI modality, the swallowtail feature in the tSWI image is labeled, specifically by placing a center point in the central region of the swallowtail feature in the tSWI image.

[0007] As a preferred technical solution, the rigid registration method is as follows: using the preprocessed T1 image as a reference, all modal sample images are registered onto the MNI52 template image.

[0008] As a preferred technical solution, the spatial alignment method is as follows: The dual-channel image is used as input to the registration network to generate a velocity field; The deformation field is obtained by applying an approximate integration operation to the velocity field. The learning template image is transformed into the space of the initial template image based on the deformation field, as follows: ,in, Represents the learning template image, This represents a spatial transformation operation. It represents the deformation field.

[0009] As a preferred technical solution, in S4, if the current iteration number is less than the preset iteration number, then the image similarity loss and deformation field regularization loss are calculated as loss functions based on the aligned template image and the sample image; otherwise, the image similarity loss, deformation field regularization loss and semantic weak supervision loss are calculated as loss functions.

[0010] As a preferred technical solution, the image similarity loss includes a local normalization coefficient term and a generator loss term, wherein the local normalization coefficient term... Represented as: ; Represents the set of spatial locations in an image; Represents any point in the set of spatial locations of an image; Indicated by Center point Points within the range ; In the neighborhood Pixel intensity value; In the neighborhood The average pixel value, and ; Indicates spatial transformation operations; Represents the learning template image; Represents the deformation field; This indicates that the neighborhood mean is calculated for the learning template image; Using the template generation subnetwork as the generator, the generator loss term is expressed as follows: ; This represents the discriminator introduced during the training process; Indicates spatial transformation operations; Represents the learning template image, and , Represents a parametric image; Represents the deformation field; This indicates that the template generates a subnetwork; Represents a parametric image; Indicates the registration network; This represents a sample image.

[0011] As a preferred technical solution, the deformation field regularization loss is expressed as: ; in, , and All represent weighting coefficients; Represents the deformation field; This represents the average of the deformation field, and ; Represents the set of spatial locations in an image; Represents any point in the set of spatial locations of an image; Represents the deformation field At a point in space Spatial gradient on; Represents the calculation of the L2 norm; Represents the deformation field In a specific spatial pixel The value at that location.

[0012] As a preferred technical solution, the semantic weak supervision loss is expressed as: ; ; in, The Dice coefficient represents the brain region labeling; This represents the spatial distance between key points, and it is the average spatial distance between the swallowtail features of the template image and the sample image. Represents a point in space; Brain region annotation map; Brain region annotation map representing the learning template image; Indicates spatial transformation operations; It represents the deformation field.

[0013] As a preferred technical solution, the discriminator loss function is: ; in, Indicates the discriminator; The real image represents a real image obtained by random sampling from the training data; Indicates spatial transformation operations; Represents the learning template image; Represents the deformation field; Indicates weight; This indicates that when the input is a real image At that time, the output of the discriminator network is related to its parameters. The gradient.

[0014] According to a second aspect of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described thereon.

[0015] Compared with the prior art, the present invention has the following beneficial effects: 1) To address the problem that traditional template construction only focuses on image registration itself and ignores subsequent application requirements, this invention incorporates brain region segmentation and key point annotation in the initialized template sample image into the model training process in the form of semantic weak supervision loss during the training of the template generation sub-network and the registration network. This makes the generated brain template not only have high-quality spatial alignment characteristics, but also naturally embed semantic information, which greatly improves the performance and practicality of the template when used for downstream tasks.

[0016] 2) To address the problem of blurred brain region boundaries and information loss caused by the linear fusion of multimodal images into a single-channel image in traditional methods in the background technology, the template generation subnetwork constructed in this invention directly organizes and splices multimodal images in a multi-channel manner and then extracts features, avoiding the difficulty of selecting linear fusion coefficients, preserving key information in each modality image, and ensuring that the constructed brain template accurately represents the abnormal brain changes related to Parkinson's disease.

[0017] 3) To address the problem that directly using all multimodal images as input to the registration network may lead to non-convergence or misleading learning during network training, this invention uses only T1 and tSWI to construct dual-channel images as input during the registration process. This not only preserves the necessary brain regions but also avoids the negative impacts of QSM and NM-MRI, ensuring stable training of the registration network and optimal image alignment performance. Attached Figure Description

[0018] Figure 1The following is the flowchart of the network training method of the present invention; Figure 2 This is a schematic diagram of the framework of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0020] A brain template is a standardized, representative digital model of the "average brain" or "reference brain." It can be used to describe the average brain state of a research population, visualize brain differences between different groups, segment brain regions of interest, and locate key points, among other downstream tasks. Based on this, brain templates can serve as an auxiliary diagnostic tool for Parkinson's disease. However, existing brain template construction methods suffer from low registration efficiency and accuracy limitations imposed by image fusion strategies. To address these technical problems, this invention provides a method for constructing multimodal brain templates for Parkinson's disease based on semantically weakly supervised guidance. The method utilizes a trained template generation sub-network and a registration network to construct the brain template. Specifically, the method provided by this invention uses the template generation sub-network and the registration network to iteratively execute... Figure 1 The process shown constructs a brain template, and the steps include: S1. Randomly select any multimodal image sample, perform preprocessing, rigidly register it onto the MNI52 template image, and initialize it to obtain the initialized template image.

[0021] S11. Data Acquisition and Preprocessing.

[0022] This invention collected data from 90 patients with progressive disease (PD) and 90 age- and sex-matched healthy controls. Each sample included five imaging modalities: T1-weighted imaging, QSM, NM-MRI, tSWI, and free water (FW). The preprocessing for each sample was as follows: 1) For the T1 modality, Fastsurfer was used to remove the skull from the T1 images and the four regions of white matter, gray matter, ventricles and cerebellum were labeled.

[0023] 2) For the QSM modality, the deep brain nuclei are labeled, including the caudate nucleus, putamen, globus pallidus, substantia nigra, red nucleus, and dentate nucleus.

[0024] 3) For the NM-MRI modality, substantia nigra and reference background are extracted from NM-MRI images and labeled. Standardization is then performed on the labeled NM-MRI images. The standardized images can be used as a semi-quantitative image. Standardization can be expressed as: ; in, Represents the original NM-MRI image; This represents the image mean of the reference background region in NM-MRI images; This represents the relative contrast ratio image obtained after standardization.

[0025] 4) For the tSWI modality, the swallowtail feature in the tSWI image is annotated. Specifically, the center point is placed in the central area of ​​the swallowtail feature in the tSWI image.

[0026] S12, rigid registration.

[0027] Using the preprocessed T1 image as a reference, all modal sample images are registered onto the MNI52 template image to achieve spatial alignment of all sample images.

[0028] S13, Initialization.

[0029] Linear averaging is performed on the registered multimodal sample images to obtain an initial template image. In this embodiment, the initial template image is organized in a multi-channel format using five modalities, including T1, tSWI, QSM, NM, and FW, with an image size of [missing information]. (Length * Width * Height * Channel).

[0030] The main goal of constructing a brain template is to obtain an average image that is as close as possible to all images in the dataset. This invention achieves this by, as shown in the example... Figure 2 The framework shown achieves this goal, and includes a template generation subnetwork and a registration network. The specific steps are as follows: S2. Construct a pre-sized and learnable parameter image. After processing the parameter image using a template generation sub-network, add it to the initial template image to obtain the learned template image.

[0031] S21, Parametric Image Construction.

[0032] The parameter image size is set to 1 / 8 of the initial template image size, meaning its length, width, and height can be expressed as: ( ); Furthermore, each position in the parametric image is a learnable parameter.

[0033] S22, Parametric image magnification processing.

[0034] The parameter image is passed through multiple transposed convolutional modules, gradually enlarged to the size of the initial template image, and then added to the initial template image to obtain the learned template image, which can be represented as: The subscripts correspond to the five different modalities of MRI.

[0035] S3. Extract dual-channel images from the learning template image and the initialization template image respectively. Based on the dual-channel images, use a registration network to spatially align the initialization template image and the learning template image to obtain an aligned template image.

[0036] Specifically, the purpose of the registration network provided by this invention is to align two input images. The input to the registration network is set as a first dual-channel image and a second dual-channel image, and a spatial deformation field is output. Then, the deformation field is used to perform spatial transformation on the learning template image to obtain an aligned template image. The aligned template image and the initial template image are compared to calculate their similarity. The obtained similarity is used to optimize the parameters in the registration network through a backpropagation algorithm. After multiple iterations, the registration network can achieve the best registration of the two input images, that is, the aligned template image and the initial template image have the highest image similarity.

[0037] S31. Among the above five image modalities, T1 and tSWI provide the clearest descriptions of brain structures and deep nuclei, making them most beneficial for the registration task. Therefore, the T1 and tSWI channels of the learning template image are extracted as the first dual-channel image, represented as: The T1 and tSWI images of the initial template image are used as one input to the registration network; correspondingly, the T1 and tSWI images of the initial template image are combined to form a second dual-channel image, which is used as the other input to the registration network. .

[0038] S32. Using the first dual-channel image and the second dual-channel image as input to the registration network, generate the velocity field. .

[0039] S33. Apply an approximate integration operation to the velocity field to obtain the deformation field. This process is common knowledge to those skilled in the art and will not be elaborated here.

[0040] S34. Based on the deformation field, the learned template image is transformed into the space of the initialized template image, represented as: ,in, Represents the learning template image, and , Represents a parametric image; This represents a spatial transformation operation. It represents the deformation field.

[0041] It should also be noted that, in practice, the sign of the velocity field can also be reversed. ), to obtain the spatial deformation from the initial template image to the learned template image ( ).

[0042] S4. Randomly select a sample image, calculate the loss function based on the alignment template image and the sample image, fix the constructed discriminator parameters, optimize the model parameters of the template generation sub-network and the registration network based on the loss function, and update the parameter image.

[0043] If the current iteration is less than 50, then the image similarity loss and deformation field regularization loss are calculated based on the aligned template image and the sample image as the loss function; otherwise, the image similarity loss, deformation field regularization loss and semantic weak supervision loss are calculated as the loss function.

[0044] 1) Image similarity loss.

[0045] Image similarity loss includes a local normalization coefficient term and a generator loss term.

[0046] Among them, the local normalization coefficient term Supervised training of the network, using image similarity as a metric, is represented as: ; Represents the set of spatial locations in an image; Represents any point in the set of spatial locations of an image; Indicated by Center point Points within the range ; In the neighborhood Pixel intensity value; In the neighborhood The average pixel value, and ; Indicates spatial transformation operations; Represents the learning template image; Represents the deformation field; This indicates that the neighborhood mean is calculated for the learning template image; The template generation subnetwork is used as the generator. The purpose of the generator is to generate template images that conform to the representation of real images. Therefore, the deformed template images should be as close as possible to the representation of real images. To achieve this goal, the generator loss term is constructed as follows: ; This represents the discriminator introduced during the training process; Indicates spatial transformation operations; Represents the learning template image; Represents the deformation field; This indicates that the template generates a subnetwork; Represents a parametric image; Indicates the registration network; This represents a sample image.

[0047] 2) Deformation field regularization loss.

[0048] The deformation field regularization loss is used to constrain the smoothness and spatial transformation amplitude of the deformation field, and the supervised registration network outputs a reasonable deformation field.

[0049] ; in, , and All represent weighting coefficients; This represents the average of the deformation field, and ; Represents the deformation field; indicates; Represents the set of spatial locations in an image; Represents any point in the set of spatial locations of an image; Represents the deformation field At a point in space Spatial gradient on; Represents the calculation of the L2 norm; Represents the deformation field In a specific spatial pixel The value at that location.

[0050] Regarding the first item, This represents the average deformation field from the template image to all sample images. This is achieved by minimizing... The first term can be used to ensure that the template image is located at the center of the sample space and is as close as possible to all sample images; the second term represents the gradient of the deformation field. Minimizing the gradient of the deformation field can constrain the deformation field to be smooth and avoid the deformation field from flipping as much as possible; the third term represents the overall deformation amplitude of the deformation field, which requires the registration network to achieve spatial alignment between the template image and the sample image with the smallest possible spatial transformation.

[0051] It should be noted that in this embodiment, image similarity loss and deformation field regularization loss are used as registration losses.

[0052] 3) Semantic weak supervision loss.

[0053] To achieve more accurate image alignment, this invention incorporates semantically weakly supervised loss, including the Dice coefficient and keypoint spatial distance. For brain region annotation, the Dice coefficient... During training, through deformation fields The brain region annotations of the template image are mapped to the sample space, and the degree of alignment of the brain region of interest is measured by calculating the Dice coefficient between the template image and the brain region annotations of the sample image.

[0054] For keypoint spatial distance, since the swallowtail sign in the substantia nigra is an important potential imaging marker in Parkinson's disease research, the deformation field of the registration network is used to transform the swallowtail sign annotation into the template space. Then, the average spatial distance between the swallowtail sign in the template image and the swallowtail sign in the sample image is calculated. The least squares of the spatial distance is used to calculate the semantic weak supervision loss function. The supervised registration network achieves accurate swallowtail sign localization through spatial transformation.

[0055] Specifically, the expression for this loss function is: ; ; in, The Dice coefficient represents the brain region labeling; This represents the spatial distance between key points, and it is the average spatial distance between the swallowtail features of the template image and the sample image. Represents a point in space; Brain region annotation map; Brain region annotation map representing the learning template image; Indicates spatial transformation operations; It represents the deformation field.

[0056] It should be noted that the brain region annotations mentioned above were obtained through automated algorithms or manual annotation during the preprocessing stage.

[0057] Brain region annotations for the learning template image are obtained through multi-map fusion. Specifically, after the two networks have been trained for a period of time, the spatial positions and boundaries of multiple brain regions in the learning template image have become relatively fixed and will not change significantly during later training. At this point, the registration network can be used to obtain the deformation field from the sample image to the learning template image, mapping the brain region annotations of the sample image to the learning template image space. Then, based on the image similarity between the sample image and the learning template image, the brain region annotations of the sample image are weighted and averaged. This process can be implemented in engineering using the ants.joint_label_fusion function in the ANTspy library.

[0058] S5. Randomly select another multimodal image sample and repeat S1~S3 in combination with the updated parameter image. Then, randomly select a real image, calculate the discriminator loss function based on the real image and the learning template, fix the template generation sub-network and registration network parameters, and update the discriminator parameters based on the discriminator loss function.

[0059] To improve the detail representation of the template image, adversarial generative training is used to supervise the training of both sub-networks. Specifically, a Unet network is constructed as the discriminator. The algorithm takes an alignment template image and a real image as input images. The discriminator performs pixel-by-pixel regression on the input image to determine whether it is a real sample. The parameters in the generator and discriminator networks are updated alternately using the training method of adversarial generative network. The real image is obtained by randomly sampling from the training data.

[0060] In this invention, the discriminator needs to have the ability to distinguish between input samples, specifically by distinguishing between real images ( ) and alignment template image ( The discriminant loss function can be constructed by regressing to 1 and 0 respectively, thus regressing to 1 and 0 respectively. ; in, Indicates the discriminator; Represents a real image; Indicates spatial transformation operations; Represents the learning template image; Represents the deformation field; Indicates weight; This indicates that when the input is a real image At that time, the output of the discriminator network is related to its parameters. The gradient.

[0061] S6. Determine if the maximum number of iterations has been reached. If not, return to execute S1; otherwise, based on the parameter image updated in the current iteration, generate the brain template output from the network using the template of the current iteration.

[0062] The above method can achieve more accurate brain template construction. Specifically, the traditional deep learning-based brain template construction method does not involve the construction of multi-contrast brain templates. If we simply transplant the previous deep learning algorithm, we will directly input the multi-contrast image as a multi-channel image into the registration network.

[0063] However, this approach does not take into account some special cases, such as QSM images, which depict the distribution of iron in the brain. In the cerebral cortex, the image appears as noise. In practice, such images can easily cause non-convergence during the training of the registration network, specifically, the network cannot achieve the purpose of registration. Also, NM-MRI (neuromelanin-sensitive imaging) is problematic because Parkinson's patients cannot remain still during MRI scans. Therefore, NM-MRI only scans the part below the striatum and cannot cover the whole brain. This means that the brain covered by NM-MRI is not consistent between different samples, which can also lead to incorrect guidance for the learning of the registration network.

[0064] The approach of using only T1 and tSWI dual-channel images as the input to the registration network in this invention can avoid the negative effects of QSM and NM-MRI on the one hand, and preserve the clear boundaries of necessary brain regions, especially deep nuclei that are important to PD, thereby enabling more accurate brain template matching.

[0065] The present invention also provides an electronic device including a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) or loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0066] Multiple components in the device are connected to the I / O interface, including: input units such as keyboards and mice; output units such as various types of displays and speakers; storage units such as disks and optical discs; and communication units such as network interface cards (NICs), modems, and wireless transceivers. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0067] The processing unit executes the various methods and processes described above, such as methods S1 to S6. For example, in some embodiments, methods S1 to S6 may be implemented as computer software programs tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via ROM and / or a communication unit. When the computer program is loaded into RAM and executed by the CPU, one or more steps of methods S1 to S6 described above may be performed. Alternatively, in other embodiments, the CPU may be configured to execute methods S1 to S6 by any other suitable means (e.g., by means of firmware).

[0068] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0069] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0070] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for constructing multimodal brain templates for Parkinson's disease based on semantically weakly supervised guidance, characterized in that, The method described above utilizes templates to generate subnetworks and registration networks, and iteratively executes the following steps to construct a brain template: S1. Randomly select any multimodal image sample, perform preprocessing, rigidly register it onto the MNI52 template image, and initialize it to obtain the initialized template image; S2. Construct a learnable parameter image of a preset size, process the parameter image using a template generation sub-network, and add it to the initial template image to obtain a learned template image; S3. Extract dual-channel images from the learning template image and the initialization template image respectively. Based on the dual-channel images, use a registration network to spatially align the initialization template image and the learning template image to obtain an aligned template image. S4. Randomly select a sample image, calculate the loss function based on the alignment template image and the sample image, fix the constructed discriminator parameters, optimize the model parameters of the template generation sub-network and the registration network based on the loss function, and update the parameter image. S5. Randomly select another multimodal image sample and repeat S1~S3 in combination with the updated parameter image. Calculate the discriminator loss function, fix the template to generate sub-network and registration network parameters, and update the discriminator parameters based on the discriminator loss function. S6. Determine if the maximum number of iterations has been reached. If not, return to execute S1. Conversely, based on the parameter image updated in the current iteration, the brain template output from the network is generated using the template of the current iteration.

2. The method for constructing a multimodal brain template for Parkinson's disease based on semantically weak supervision as described in claim 1, characterized in that, The multimodal image samples include five modalities: T1, QSM, NM-MRI, tSWI, and FW. Their preprocessing is as follows: For the T1 modality, Fastsurfer was used to remove the skull from the T1 images and the four regions of white matter, gray matter, ventricles and cerebellum were labeled. For the QSM modality, the deep brain nuclei are labeled, including the caudate nucleus, putamen, globus pallidus, substantia nigra, red nucleus, and dentate nucleus; For the NM-MRI modality, substantia nigra and reference background are extracted from NM-MRI images and annotated, and then the annotated NM-MRI images are standardized. For the tSWI modality, the swallowtail feature in the tSWI image is labeled, specifically by placing a center point in the central region of the swallowtail feature in the tSWI image.

3. The method for constructing a multimodal brain template for Parkinson's disease based on semantically weak supervision as described in claim 2, characterized in that, The rigid registration method is as follows: using the preprocessed T1 image as a reference, all modal sample images are registered onto the MNI52 template image.

4. The method for constructing a multimodal brain template for Parkinson's disease based on semantically weak supervision as described in claim 1, characterized in that, The spatial alignment method is as follows: The dual-channel image is used as input to the registration network to generate a velocity field; The deformation field is obtained by applying an approximate integration operation to the velocity field. The learning template image is transformed into the space of the initial template image based on the deformation field, as follows: ,in, Represents the learning template image, This represents a spatial transformation operation. It represents the deformation field.

5. The method for constructing a multimodal brain template for Parkinson's disease based on semantically weak supervision as described in claim 1, characterized in that, In S4, if the current iteration number is less than the preset iteration number, then the image similarity loss and deformation field regularization loss are calculated as loss functions based on the aligned template image and the sample image; otherwise, the image similarity loss, deformation field regularization loss and semantic weak supervision loss are calculated as loss functions.

6. The method for constructing a multimodal brain template for Parkinson's disease based on semantically weak supervision as described in claim 5, characterized in that, The image similarity loss includes a local normalization coefficient term and a generator loss term, wherein the local normalization coefficient term... Represented as: ; Represents the set of spatial locations in an image; Represents any point in the set of spatial locations of an image; Indicates Center point Points within the range ; In the neighborhood Pixel intensity value; In the neighborhood The average pixel value, and ; Indicates spatial transformation operations; Represents the learning template image; Represents the deformation field; This indicates that the neighborhood mean is calculated for the learning template image; Using the template generation subnetwork as the generator, the generator loss term is expressed as follows: ; This represents the discriminator introduced during the training process; Indicates spatial transformation operations; This represents the learning template image, and , Represents a parametric image; Represents the deformation field; This indicates that the template generates a subnetwork; Represents a parametric image; Indicates the registration network; This represents a sample image.

7. A method for constructing a multimodal brain template for Parkinson's disease based on semantically weak supervision, as described in claim 5, is characterized in that... The deformation field regularization loss is expressed as: ; in, , and All represent weighting coefficients; Represents the deformation field; This represents the average of the deformation field, and ; Represents the set of spatial locations in an image; Represents any point in the set of spatial locations of an image; Represents the deformation field At a point in space Spatial gradient on; Represents the calculation of the L2 norm; Represents the deformation field In a specific spatial pixel point The value at that location.

8. A method for constructing a multimodal brain template for Parkinson's disease based on semantically weak supervision, as described in claim 5, is characterized in that... The semantic weak supervision loss is expressed as: ; ; in, The Dice coefficient represents the brain region labeling; This represents the spatial distance between key points, and it is the average spatial distance between the swallowtail features of the template image and the sample image. Represents a point in space; Brain region annotation map; Brain region annotation map representing the learning template image; Indicates spatial transformation operations; It represents the deformation field.

9. A method for constructing a multimodal brain template for Parkinson's disease based on semantically weak supervision, as described in claim 1, is characterized in that... The discriminator loss function is: ; in, Indicates the discriminator; The real image represents a real image obtained by random sampling from the training data; Indicates spatial transformation operations; Represents the learning template image; Represents the deformation field; Indicates weight; This indicates that when the input is a real image At that time, the output of the discriminator network is related to its parameters The gradient.

10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 9.