Functional lung partitioning method and device, equipment and medium
By extracting common features from functional lung imaging using a cross-modal partitioning model, the functional lung regions are automatically divided, solving the subjectivity and accuracy problems of functional lung partitioning results in existing technologies. This achieves highly accurate functional lung partitioning, adapting to diverse clinical application scenarios and reducing the risk of radiation-induced lung injury.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- OUR INNOBEAM MEDICAL CO LTD
- Filing Date
- 2026-01-04
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies for functional lung zoning results are highly subjective, have poor repeatability, and low accuracy, failing to meet the needs of radiotherapy planning. In particular, when single-modal or multi-modal imaging is incomplete, it is difficult to fully reflect the lung function status.
A cross-modal partitioning model is used to extract cross-modal common features from functional lung imaging in at least one modality. The pre-trained cross-modal partitioning model is used to partition the functional lungs, automatically dividing multiple functional regions of the lungs and generating highly accurate partitioning results.
It improves the accuracy and repeatability of functional lung zoning, adapts to diverse clinical application scenarios, reduces the radiation dose to high-functioning areas of the lungs, lowers the risk of radiation-induced lung injury, and enhances the protective effect of radiotherapy plans.
Smart Images

Figure CN121837296A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of medical image processing technology, and in particular to methods, devices, equipment and media for functional lung zoning. Background Technology
[0002] With the development of medical technology, radiotherapy for lung tumors according to a pre-established radiotherapy plan has become one of the main methods of treating lung tumors.
[0003] When developing a radiotherapy plan for lung tumors, it is usually combined with functional lung zoning (that is, dividing the lung into multiple functional areas) to develop the radiotherapy plan. This is to reduce the radiation dose to high-functional areas of the lung while meeting the required dose for the target area and the maximum radiation dose that organs at risk can tolerate. This can effectively reduce the risk of radiation-induced lung injury and improve the protection of the lungs.
[0004] Currently, functional lung zoning is mainly based on functional lung imaging, using manual or threshold-based methods. However, both of these methods are highly dependent on experience, resulting in subjective and poorly reproducible functional lung zoning results. This leads to low accuracy in functional lung zoning results, which cannot meet the needs of radiotherapy planning. Summary of the Invention
[0005] This disclosure provides a method, apparatus, device, and medium for functional lung zoning; it can output highly accurate functional lung zoning results without relying on experience, thus ensuring the objectivity and repeatability of the functional lung zoning results.
[0006] The technical solution disclosed herein is implemented as follows: In a first aspect, this disclosure provides a functional lung partitioning method, comprising: acquiring at least one modality of functional lung imaging of a target object; generating functional lung partitioning results based on the at least one modality of functional lung imaging using a pre-trained cross-modal partitioning model; the functional lung partitioning results being used to indicate the spatial distribution of multiple functional regions of the lung divided according to functional status.
[0007] In a second aspect, this disclosure provides a functional lung partitioning device, comprising: an acquisition module for acquiring at least one modality of functional lung imaging of a target object; and a partitioning module for generating functional lung partitioning results based on the at least one modality of functional lung imaging using a pre-trained cross-modal partitioning model; wherein the functional lung partitioning results are used to indicate the spatial distribution of multiple functional regions of the lung divided according to functional state.
[0008] Thirdly, this disclosure provides an electronic device including a memory and a processor, the memory storing executable instructions; when executed by the processor, the instructions implement the steps of the functional lung partitioning method of the first aspect.
[0009] Fourthly, this disclosure provides a computer storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the functional lung partitioning method of the first aspect.
[0010] Fifthly, this disclosure provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the functional lung partitioning method of the first aspect.
[0011] By applying the embodiments of this disclosure, functional lung zoning is performed through a cross-modal zoning model. Whether based on single-modal or multi-modal functional lung imaging, highly accurate functional lung zoning results can be output without relying on experience, ensuring the objectivity and repeatability of the functional lung zoning results. Attached Figure Description
[0012] Figure 1 This is an architectural diagram of a functional lung partitioning system provided in an embodiment of the present disclosure.
[0013] Figure 2 A structural block diagram of an electronic device provided in an embodiment of this disclosure.
[0014] Figure 3 A flowchart of a functional lung partitioning method provided in an embodiment of this disclosure.
[0015] Figure 4 This is a flowchart illustrating a process for generating functional lung partitioning results, as provided in an embodiment of this disclosure.
[0016] Figure 5 This is a schematic diagram illustrating the composition of a cross-modal partitioning model provided in an embodiment of this disclosure.
[0017] Figure 6 This is a schematic diagram illustrating the composition of a cross-modal feature extraction model provided in an embodiment of this disclosure.
[0018] Figure 7 This is a schematic diagram illustrating the composition of a single-modal feature extraction module provided in an embodiment of this disclosure.
[0019] Figure 8 A flowchart illustrating a training method for a cross-modal feature extraction model provided in this embodiment of the disclosure.
[0020] Figure 9 This is a schematic diagram illustrating the training process of a cross-modal feature extraction model provided in an embodiment of this disclosure.
[0021] Figure 10 This is a schematic diagram illustrating the composition of a training feature extraction model provided in an embodiment of this disclosure.
[0022] Figure 11 This is a schematic diagram of a functional lung partition image provided in an embodiment of this disclosure.
[0023] Figure 12 This is a schematic diagram of another functional lung region image provided in an embodiment of this disclosure.
[0024] Figure 13 This is a flowchart illustrating a process for visualizing functional lung zoning results, as provided in an embodiment of this disclosure.
[0025] Figure 14 This is a schematic diagram of the composition of the functional lung partitioning device provided in this disclosure. Detailed Implementation
[0026] The technical solutions of this disclosure will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art should fall within the protection scope of this disclosure.
[0027] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0028] Furthermore, in the embodiments of this disclosure, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0029] To facilitate understanding of the technical solutions of the embodiments of this disclosure, the related technologies of the embodiments of this disclosure are described below. The following related technologies are optional solutions and can be combined with the technical solutions of the embodiments of this disclosure in any way, and they all fall within the protection scope of the embodiments of this disclosure.
[0030] In related technologies, functional lung imaging is the primary method, with functional lung zoning performed through manual or threshold-based methods. Manual zoning requires physicians or physicists to manually delineate the boundaries of each functional region based on experience, while threshold-based zoning requires manually setting thresholds, which lack standardized criteria and are highly dependent on experience. Therefore, functional lung zoning results generated using these two methods suffer from high subjectivity and poor repeatability, leading to low accuracy and failing to meet the needs of radiotherapy planning.
[0031] Furthermore, if functional lung zoning is performed based on single-modality functional lung imaging (i.e., functional lung imaging with only one modality), the single-modality features extracted can only characterize one dimension of lung function (such as ventilation or perfusion), making it difficult to comprehensively reflect the overall functional state of lung tissue, which directly limits the accuracy of functional lung zoning. Conversely, if functional lung zoning is performed based on multimodal functional lung imaging (i.e., functional lung imaging with multiple modalities), if it is impossible to obtain functional lung imaging of all modalities (e.g., some patients cannot have a certain modality of functional lung imaging due to their physical condition), it will be impossible to obtain multimodal features that can comprehensively characterize the lung function state, which will also affect the accuracy of functional lung zoning, resulting in poor clinical applicability and inability to adapt to diverse clinical application scenarios.
[0032] In view of this, the embodiments of this disclosure extract comprehensive cross-modal common features characterizing lung function from at least one modality of functional lung imaging using a cross-modal partitioning model, and divide the lung into multiple functional regions based on these cross-modal common features. This allows for the output of highly accurate functional lung partitioning results regardless of whether the functional lung imaging is based on a single modality or a multimodality, improving clinical applicability and adapting to diverse clinical application scenarios. Furthermore, it eliminates the need for manually setting thresholds or manually delineating the boundaries of functional regions, ensuring the objectivity and repeatability of the functional lung partitioning results.
[0033] The embodiments disclosed herein can also be extended to clinical diagnosis and treatment scenarios such as chest radiotherapy and routine pulmonary function analysis and testing in respiratory medicine, enabling cross-departmental applications.
[0034] See Figure 1 , Figure 1 This is an architectural diagram of a functional lung zoning system provided in an embodiment of this disclosure. The functional lung zoning system 100 is used to implement a functional lung zoning method, providing an exemplary application environment for embodiments of this method. The functional lung zoning system 100 includes an image acquisition device 110, a processing device 120, a terminal 130, a radiotherapy device 140, and a network 150. The terminal 130 includes one or more devices. The image acquisition device 110, processing device 120, radiotherapy device 140, and terminal 130 communicate via wireless connection, wired connection, or a combination thereof.
[0035] Specifically, the image acquisition device 110 generates or provides image data related to a target object by scanning the target object. The target object can include biological and / or non-biological objects. For example, the target object can include a patient, or specific parts of the patient's body, such as the head, chest, abdomen, etc., or combinations thereof. As another example, the target object can be a human-made component (such as a phantom) of living or non-living organic and / or inorganic matter. The image acquisition device 110 is a non-invasive biomedical imaging device for disease diagnosis or research purposes, including single-modal scanners and / or multi-modal scanners. Single-modal scanners can include, for example, ultrasound scanners, X-ray scanners, computed tomography (CT) scanners, magnetic resonance imaging (MRI) scanners, ultrasound examination instruments, positron emission tomography (PET) scanners, optical coherence tomography (OCT) scanners, ultrasound (US) scanners, intravascular ultrasound (IVUS) scanners, near-infrared spectroscopy (NIRS) scanners, far-infrared (FIR) scanners, etc., or any combination thereof. Multimodal scanners may include, for example, X-ray imaging-magnetic resonance imaging (X-MRI) scanners, positron emission tomography-X-ray imaging (PET-X-ray) scanners, single-photon emission computed tomography-magnetic resonance imaging (SPECT-MRI) scanners, positron emission tomography-computed tomography (PETCT) scanners, and digital subtraction angiography-magnetic resonance imaging (DSA-MRI) scanners. The image acquisition device 110 may specifically include components such as a gantry, detector, detection area, and scanning bed. The target object can be placed on the scanning bed and then moved to the detection area for scanning, thereby acquiring image data of the target object. The scanners described above are for illustrative purposes only and are not intended to limit the scope of this disclosure.
[0036] Processing device 120 can be a single server or a server cluster. The server cluster can be centralized or distributed. In some examples, processing device 120 can be local or remote to the functional lung partitioning system 100. For example, processing device 120 can access information and / or data from image acquisition device 110, radiotherapy device 140, and terminal 130 via network 150. As another example, processing device 120 can be directly connected to image acquisition device 110, terminal 130, and / or radiotherapy device 140 to access information and / or data. In some examples, processing device 120 can be implemented on a cloud platform. For example, the cloud platform can include private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, inter-cloud, multi-cloud, etc., or combinations thereof. In some examples, processing device 120 can include one or more processors (e.g., a single-chip processor or a multi-chip processor). Processing device 120 can be composed of one or more processors such as... Figure 2 The electronic device 200 implements the described components.
[0037] Terminal 130 is used to receive and display scanned images from image acquisition device 110, such as functional lung imaging of the target object; it is also used to enable interaction between the user and functional lung zoning system 100. For example, terminal 130 is used to display functional images, lung anatomy images and functional zoning results to a physicist or doctor so as to evaluate, review and adjust the functional lung zoning results.
[0038] In some examples, terminal 130 may include a mobile device, tablet computer, laptop computer, etc., or any combination thereof. For example, a mobile device may include a mobile phone, personal digital assistant (PDA), gaming device, navigation device, point-of-sale (POS) device, laptop computer, tablet computer, desktop computer, etc., or any combination thereof. In some examples, terminal 130 may include input devices, output devices, etc. In some examples, terminal 130 may be part of processing device 120.
[0039] The radiotherapy device 140 is a radiotherapy planning execution device containing dose information, used to target tumors with irradiation and protect normal functional tissues according to the dose information generated by the processing device 120. Specifically, the radiotherapy device 140 can be a linear accelerator (LINAC), a proton / heavy ion therapy device, a stereotactic radiotherapy device, etc.
[0040] Network 150 includes any suitable network that can facilitate the exchange of information and / or data between the functional lung zoning system 100. For example, processing device 120 can acquire functional lung images from image acquisition device 110 via network 150. As another example, processing device 120 can acquire user instructions from terminal 130 via network 150. Network 150 can be or includes public networks (e.g., the Internet), private networks (e.g., local area networks (LANs)), wired networks, wireless networks (e.g., 602.11 networks, Wi-Fi networks), Frame Relay networks, virtual private networks (VPNs), satellite networks, telephone networks, routers, hubs, switches, server computers, and / or any combination thereof. For example, network 150 can include cable networks, wired networks, fiber optic networks, telecommunications networks, intranets, wireless local area networks (WLANs), metropolitan area networks (MANs), public switched telephone networks (PSTNs), Bluetooth networks, ZigBee networks, near field communication (NFC) networks, etc., or any combination thereof. In some examples, network 150 can include one or more network access points. For example, network 150 may include wired and / or wireless network access points such as base stations and / or internet exchange points, through which one or more components of functional lung partition system 100 may connect to network 150 to exchange data and / or information.
[0041] The functional lung zoning system 100 of this embodiment provides the hardware and system foundation for realizing the functional lung imaging-guided functional lung zoning method, enabling subsequent method embodiments to run in a real clinical environment.
[0042] See Figure 2 , Figure 2 This is a structural block diagram of an electronic device provided in an embodiment of this disclosure. In some examples, the electronic device 200 can be at least one of a smartphone, smartwatch, desktop computer, laptop, virtual reality terminal, augmented reality terminal, wireless terminal, and laptop computer. The electronic device 200 has communication functions and can access a wired or wireless network. The electronic device 200 can refer to one of a plurality of terminals, and those skilled in the art will understand that the number of such terminals can be more or less. In some examples, the electronic device 200 can receive image data based on the accessed wired or wireless network. It is understood that the electronic device 200 undertakes the calculation and processing work of the technical solution of this disclosure, and this disclosure does not limit it in this respect.
[0043] like Figure 2 As shown, the electronic device 200 in this disclosure may include one or more of the following components: processor 210 and memory 220.
[0044] In some examples, processor 210 connects various parts of the electronic device using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 220, and by calling data stored in memory 220. Optionally, processor 210 can be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 210 can integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), Neural-network Processing Unit (NPU), and baseband chip. Among them, the CPU mainly handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required to be displayed on the touch screen; the NPU is used to implement artificial intelligence (AI) functions; and the baseband chip is used to handle wireless communication. It is understandable that the aforementioned baseband chip may not be integrated into the processor 210, but may be implemented using a separate chip.
[0045] Memory 220 may include mass storage devices, removable storage devices, volatile read-write memory, read-only memory (ROM), or any combination thereof. Exemplary mass storage devices may include disks, optical disks, solid-state drives, etc. Exemplary removable storage devices may include flash drives, floppy disks, optical disks, memory cards, compact disks, magnetic tapes, etc. Exemplary volatile read-write memory may include random access memory (RAM). Exemplary RAM may include dynamic random access memory (DRAM), double data rate synchronous dynamic access memory (DDRSDRAM), static random access memory (SRAM), thyristor random access memory (T-RAM), and zero-capacitance random access memory (Z-RAM), etc. Exemplary ROM may include mask-mode read-only memory (MROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable read-only memory (EEPROM), optical disc read-only memory (CDROM), and digital universal optical disc read-only memory, etc. Memory 220 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 220 may include a program storage area and a data storage area. The program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above method embodiments, etc. The data storage area may store data created according to the use of the electronic device, etc.
[0046] In addition, those skilled in the art will understand that the structure of the electronic device shown in the above figures does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the electronic device may also include a display screen, camera assembly, microphone, speaker, radio frequency circuit, input unit, sensors (such as accelerometer, angular velocity sensor, light sensor, etc.), audio circuit, WiFi module, power supply, Bluetooth module, etc., which will not be described in detail here.
[0047] See Figure 3 , Figure 3 This is a flowchart of a functional lung partitioning method provided in an embodiment of the present disclosure. Specifically, it includes steps S302-S304.
[0048] Step S302: Obtain at least one modality of functional lung imaging of the target object.
[0049] Specifically, in clinical settings, due to differences in examination pathways for different target subjects (e.g., some patients cannot have a certain modality of functional lung imaging acquired due to their physical condition) and the varying needs of different departments, it may be impossible to obtain functional lung imaging of all modalities for the target subject. Therefore, functional lung zoning can be performed by acquiring at least one modality of functional lung imaging that has already been acquired for the target subject.
[0050] Functional lung imaging (FLE) is a quantitative data method for obtaining lung function activity. It quantifies the functional status of different regions of the lung, dividing the lung into different functional areas to guide the dosage planning of lung radiotherapy. Lung function activity includes physiological activities (such as ventilation, blood perfusion, and respiratory movements) and biochemical activities (such as oxygen metabolism and glucose uptake), etc.
[0051] Functional lung imaging can include the target area, organs at risk, and the lungs themselves. The target area, or tumor target area, is the tumor region that needs to be covered by high-dose radiation during radiotherapy. Organs at risk are normal tissues or organs that need to be protected during radiotherapy and are sensitive to radiation, such as, but not limited to, the heart, esophagus, trachea, spinal cord, breast, and humeral head. It should be noted that in lung cancer radiotherapy, the lungs are actually considered organs at risk. Each type of functional imaging technique can acquire a modality of functional lung imaging. Functional imaging techniques can include, but are not limited to, dynamic magnetic resonance imaging (dynamic MRI), four-dimensional computed tomography (4D CT), single-photon emission computed tomography (SPECT), and positron emission tomography-computer tomography (PET-CT).
[0052] Step S304: Based on functional lung imaging of at least one modality, generate functional lung partitioning results through a pre-trained cross-modal partitioning model; the functional lung partitioning results are used to indicate the spatial distribution of multiple functional regions of the lung divided according to functional status.
[0053] Specifically, a cross-modal partitioning model can be pre-trained. This model is an end-to-end deep learning model that integrates multimodal feature learning and automated partitioning algorithms. After acquiring functional lung images, these images can be input into the cross-modal partitioning model. The model then mines and extracts cross-modal common features that comprehensively characterize lung function from single-modal or multimodal functional lung images. Based on these cross-modal common features, automated functional lung partitioning is performed to obtain the functional lung partitioning results.
[0054] The cross-modal common feature functional lung partitioning result is generated by dividing multiple functional regions of the lung according to the regional heterogeneity of lung function status. The functional lung partitioning result can indicate the spatial distribution of multiple functional regions, for example, it can be a functional lung partitioning image. The specific form of the functional lung partitioning image can be a binary mask image including high-functioning lung regions and low-functioning lung regions, or it can be implemented as a multi-class mask image including high-functioning lung regions, intermediate-functioning lung regions, and low-functioning lung regions.
[0055] Lung function refers to the state of the lungs during gas exchange, which can be assessed using at least one of the following indicators: functional strength, type of functional impairment, and pathological condition. For example, the lungs can be divided into high-functioning and low-functioning areas based on the strength of pulmonary ventilation (the ability of the alveoli to exchange gases with the external environment), blood perfusion (the ability of blood to exchange between pulmonary capillaries and alveoli), and metabolic activity. They can also be divided into ventilatory impairment and diffusion impairment areas based on the type of functional impairment. Furthermore, the lungs can be divided into normal and poorly ventilated areas based on the pathological condition. High-functioning areas of the lungs are significantly more sensitive to radiation than other areas, and these areas are more prone to radiation damage after receiving high doses of radiation. For instance, high-functioning areas are more likely to suffer radiation damage after receiving high doses of radiation than low-functioning areas, and normal areas are more likely to suffer radiation damage after receiving high doses of radiation than poorly ventilated areas.
[0056] It should be noted that cross-modal common features are actually common features of multimodal functional lung imaging. Specifically, they can take the form of a single-channel semantic feature map that can comprehensively reflect voxel-level features of multi-dimensional functional states such as lung ventilation, perfusion, and metabolism. In other words, even if only single-modal functional lung imaging is available, the cross-modal partitioning model can still output multimodal common features (i.e., cross-modal common features). Furthermore, if some modalities are missing in multimodal functional lung imaging, the cross-modal partitioning model can still output multimodal common features that can comprehensively characterize lung function (i.e., cross-modal common features).
[0057] Therefore, by applying this cross-modal partitioning model for functional lung partitioning, we can overcome the technical limitations of single-modal features extracted by single-modal functional lung imaging, which cannot fully represent lung function and improve the accuracy of functional lung partitioning. We can also solve the problem of difficulty in accurately dividing lung functional areas when certain modalities of functional lung imaging are lacking, thus improving clinical universality and flexibly adapting to diverse clinical application scenarios.
[0058] After generating the functional lung partitioning results, the results can be exchanged with the treatment planning system (TPS) via the Dicom-RT protocol. The TPS will then visualize the results for reference by physicists or physicians, thereby efficiently integrating the functional lung partitioning results into the clinical radiotherapy planning process.
[0059] This embodiment utilizes a cross-modal partitioning model to extract comprehensive cross-modal common features characterizing lung function from at least one modality of functional lung imaging. Based on these common features, the lung is divided into multiple functional regions. This allows for highly accurate functional lung partitioning results regardless of whether the functional lung imaging is based on a single modality or a multimodality, improving clinical applicability and adaptability to diverse clinical applications. Furthermore, it eliminates the need for manually setting thresholds or delineating functional region boundaries, ensuring the objectivity and reproducibility of the functional lung partitioning results. Additionally, these functional lung partitioning results can be directly integrated into the radiotherapy planning process, helping clinicians reduce the radiation dose to high-functional lung regions while meeting the required dose for the target area and the maximum acceptable radiation dose for organs at risk. This effectively reduces the risk of radiation-induced lung injury and enhances lung protection.
[0060] See Figure 4 , Figure 4 This is a flowchart illustrating a process for generating functional lung zoning results according to an embodiment of the present disclosure. Step S304 specifically includes steps S3042-S3044.
[0061] Step S3042: Based on functional lung imaging of at least one modality, extract cross-modal common features through a cross-modal feature extraction model.
[0062] Step S3044: Based on cross-modal common features, generate functional lung partitioning results through a partitioning model.
[0063] Specifically, see Figure 5The cross-modal partitioning model 400 may include a cross-modal feature extraction model 402 and a partitioning model 404. After acquiring functional lung images of at least one modality, the functional lung images of at least one modality can be spatially registered with a reference template with high precision to generate functional lung images aligned in anatomical space, so as to ensure that the spatial dimension and resolution of the data meet the input requirements of the cross-modal partitioning model 400.
[0064] Registration methods can be selected based on different modalities, or a combination of rigid and non-rigid registration can be used to ensure spatial consistency and alignment of functional lung images. Rigid registration includes correcting for rotational and translational differences in images to ensure global alignment of the overall anatomical structures. Non-rigid registration involves refining adjustments to local deformations of anatomical structures and correcting local deviations caused by respiratory, functional changes, or anatomical differences. A representative functional lung image with clear anatomical structures and high imaging quality can be selected from a dataset consisting of numerous multimodal functional lung images as a reference template to ensure that the registration benchmark has good anatomical consistency and functional representativeness.
[0065] Next, the registered functional lung images of at least one modality can be input into the cross-modal feature extraction model 402, which will then mine and extract cross-modal common features from the single-modal or multimodal functional lung images. The cross-modal feature extraction model can adopt an encoder-decoder architecture.
[0066] Then, the cross-modal common features can be input into the partitioning model 404, which will automatically perform functional lung partitioning based on the cross-modal common features using a self-supervised or unsupervised method, and output the functional lung partitioning results.
[0067] By applying this embodiment, cross-modal common features are extracted through a cross-modal feature extraction model, and then functional lung zoning is performed by a zoning model based on the cross-modal common features. This ensures that the zoning is based on common features that comprehensively characterize lung function, rather than one-sided features of a single modality, thereby improving the accuracy of functional lung zoning results.
[0068] See Figure 6 , Figure 6 This is a schematic diagram illustrating the composition of a cross-modal feature extraction model provided in an embodiment of this disclosure. Specifically, the cross-modal feature extraction model 402 includes multiple parallel single-modal feature extraction modules 4022 and single-modal feature fusion modules 4024.
[0069] Each single-modal feature extraction module 4022 is used to receive functional lung imaging of one modality and extract the single-modal common features of functional lung imaging of that modality. Each modality of functional lung imaging corresponds to one single-modal feature extraction module.
[0070] For example, the single-modal feature extraction module 4022 is a feature extraction sub-network that corresponds one-to-one with the modality of functional lung imaging. Specifically, it can adopt a single-modal encoder-decoder structure (such as an encoder-decoder branch built based on Unet3D, ResNet3D, and VNet), which can extract single-modal common features that characterize the lung function status under a specific modality.
[0071] The single-modal feature fusion module 4024 is used to fuse the single-modal common features extracted by each single-modal feature extraction module 4022 to obtain cross-modal common features.
[0072] For example, the single-modal feature fusion module 4024 can employ either an element-wise addition and averaging fusion method or an attention-weighted fusion method. The element-wise addition and averaging fusion method is implemented as follows: first, the single-modal common features output by each single-modal feature extraction module are added element-wise (the feature values at corresponding voxel positions are directly summed), and then the sum is divided by the number of modalities to obtain the average value, thus obtaining the cross-modal common features. The attention-weighted fusion method is implemented as follows: first, the attention weight of each single-modal feature is determined (for example, by obtaining a preset weight or calculating information entropy; the higher the information entropy, the higher the weight); then, each single-modal common feature is multiplied by its corresponding weight and summed to obtain the cross-modal common features.
[0073] In this embodiment, a parallel single-modal feature extraction module can extract common features from functional lung imaging of different modalities, avoiding interference between imaging features from different modalities and improving the accuracy of single-modal common features. A single-modal feature fusion module integrates these common features, further improving the quality of cross-modal common features. The cross-modal common features extracted by the cross-modal feature extraction model are applicable across departments, breaking down barriers between different diagnostic and treatment scenarios and possessing clinical universality.
[0074] For example, a schematic diagram of the module composition when the single-modal feature extraction module 4022 is specifically implemented as a single-modal encoder-decoder structure is shown below. Figure 7 As shown.
[0075] Specifically, each single-modal feature extraction module 4022 consists of a single-modal encoder and a single-modal decoder. The single-modal encoder receives functional lung imaging of a specific modality. Based on a 3D convolutional network (such as Unet3D, ResNet3D, VNet, etc.), it encodes the input functional lung imaging of that modality through hierarchical operations such as convolution and pooling, extracting local detail features and global semantic features of lung function in that modality. This maps the original voxel-level imaging data into a high-dimensional feature space representation specific to that modality, completing the transformation from raw image data to modality-specific high-dimensional features. The single-modal decoder receives the high-dimensional feature representation output by the single-modal encoder. Through operations such as deconvolution and upsampling, it gradually restores the spatial resolution of the features, performs semantic enhancement, detail completion, and spatial dimension restoration on the modality-specific features extracted in the encoding stage, and finally outputs a single-channel semantic feature map that can accurately characterize the core features of lung function in that modality, generating the single-modal common features of functional lung imaging of that modality.
[0076] See Figure 8 , Figure 8 This is a flowchart illustrating a training method for a cross-modal feature extraction model provided in this embodiment. The cross-modal feature extraction model is trained based on the following steps S502-S504.
[0077] Step S502: Construct a training feature extraction model and use the training feature extraction model as a student model.
[0078] Step S504: Use the pre-trained multimodal feature extraction model as the teacher model to guide the training of the feature extraction model and obtain the cross-modal feature extraction model.
[0079] Specifically, when training a cross-modal feature extraction model, it is first necessary to build a training feature extraction model as a student model, and then build and train a multimodal feature extraction model as a teacher model.
[0080] The training feature extraction model comprises multiple parallel training feature extraction modules and training feature fusion modules. Each training feature extraction module can employ a single-modal encoder-decoder structure to receive training functional lung imaging of one modality and extract the training single-modal common features of that modality. The training feature fusion module fuses the training single-modal common features extracted by each training feature extraction module to obtain training cross-modal common features. The multimodal feature extraction model can employ a multimodal encoder-decoder structure to receive multimodal training functional lung imaging (such as "PET-CT image + SPECT image" or "4D CT image + dynamic MRI image") and extract multimodal common features that comprehensively characterize lung function.
[0081] It should be noted that training functional lung imaging refers to functional lung imaging used for model training. This data must comply with the "Medical Imaging Data Annotation Standard" (such as the DICOM standard), include cases with different pathological types (such as non-small cell lung cancer and lung metastases) and different degrees of functional impairment (such as mild ventilatory impairment and severe perfusion defect), and be clinically annotated to ensure clinical representativeness. Training functional lung imaging can be obtained through standardized imaging examinations in clinical settings, collaborative collection in multi-center clinical research, or screening of compliant public medical datasets, while strictly adhering to medical data privacy protection and regulations.
[0082] Then, the teacher model can use knowledge distillation to distill the multimodal common features learned by the teacher model into the student model. This allows the student model to learn the teacher model's ability to extract multimodal common features, enabling it to generate common feature representations with multimodal semantic consistency (i.e., cross-modal common features) given single-modal or multimodal functional lung imaging input. After the teacher model guides the student model through training, a cross-modal feature extraction model can be obtained.
[0083] By applying this embodiment, a pre-trained multimodal feature extraction model is used as a teacher model to transmit multimodal knowledge. This allows the training feature extraction model, which acts as a student model, to output common features that comprehensively characterize lung function even when only a single modality or some modal inputs are missing. This enables the model to adapt to the differences in diagnosis and treatment pathways in different departments such as radiotherapy and respiratory medicine, making the model universally applicable across departments and improving the comprehensiveness and accuracy of lung function feature extraction.
[0084] See Figure 9 , Figure 9 This is a schematic diagram illustrating the training process of a cross-modal feature extraction model provided in an embodiment of this disclosure. Step S504 specifically includes steps S5042-S5046.
[0085] Step S5042: Fix the model weights of the teacher model, and through knowledge distillation, enable each unimodal encoder to learn the feature distribution of the multimodal encoder, and calculate the distillation loss function.
[0086] Specifically, see Figure 10 The training feature extraction model 602 includes multiple parallel training feature extraction modules 6022 and training feature fusion modules 6024. Each training feature extraction module 6022 includes a single-modal encoder and a single-modal decoder. The multimodal feature extraction model 604 includes a multimodal encoder and a multimodal decoder.
[0087] For example, when training the student model, each registered training function lung image can be input into the corresponding unimodal encoder of the student model according to its modality, and the multimodal common features output by the teacher model can be used as the label of the student model to provide the multimodal common feature distribution that the student model needs to learn. Simultaneously, all model parameters of the multimodal encoder and multimodal decoder can be fixed, and the multimodal features of the teacher model's multimodal encoder can be distilled into each unimodal encoder of the student model using knowledge distillation technology, so that each unimodal encoder of the student model can learn the feature distribution of the teacher model's multimodal encoder. L2 loss (mean squared error) and KL divergence (relative entropy) can be used as knowledge distillation loss functions (denoted as Loss1) to measure the difference between the feature distribution of the unimodal encoder and the feature distribution of the multimodal encoder.
[0088] Step S5044: Randomly discard the training single-modal common features, and calculate the reconstruction loss function based on the training cross-modal common features and the multimodal common features output by the multimodal decoder.
[0089] Specifically, during the training of the student model, a random modality dropout strategy can be adopted. In the training feature fusion module 6024, some common features of the training single modality are randomly dropped (e.g., a 50% probability dropout of common features of the training single modality for SPECT). This simulates scenarios where some modalities are missing in clinical practice, thereby enhancing the model's robustness and generalization ability when single-modality inputs are lacking or certain modalities are missing, and preventing the model from over-relying on a particular modality. The dropout probability can be set based on the clinical modality missing rate. For example, approximately 20% of patients in clinical practice cannot undergo SPECT examinations due to renal insufficiency; therefore, the dropout probability of common features of the training single modality corresponding to the SPECT modality can be set to 20%. Similarly, approximately 15% of patients cannot undergo dynamic MRI examinations due to claustrophobia; therefore, the dropout probability of common features of the training single modality corresponding to the dynamic MRI modality can be set to 15%. This makes the training process closer to real clinical scenarios and improves the practicality of the model after deployment.
[0090] Then, a reconstruction loss function (denoted as Loss2) can be calculated based on the training cross-modal common features output by the training feature fusion module 6024 and the multimodal common features output by the multimodal decoder to measure the difference between the training cross-modal common features and the multimodal common features. The reconstruction loss function can be L1 loss (mean absolute error) or cosine similarity loss, etc.
[0091] Step S5046: When the distillation loss function and the reconstruction loss function meet the preset conditions, complete the training of the training feature extraction model to obtain the cross-modal feature extraction model.
[0092] Specifically, the preset condition is the criterion for terminating student model training. For example, it could be that the total loss function (the sum of Loss1 and Loss2) is less than or equal to a preset threshold.
[0093] By applying this embodiment, the random dropout strategy and reconstruction loss constraint enable the student model to simulate real-world modality missing scenarios during training. This ensures that even with only a single modality or missing some modal inputs after deployment, it can still output cross-modal common features that comprehensively represent lung function status, thereby improving the accuracy of lung function zoning. Furthermore, the distillation loss enables the single-modal encoder to learn the feature distribution of the multimodal encoder, thereby improving the model's cross-modal and cross-scenario generality.
[0094] In some embodiments, step S3044 is specifically implemented as follows: based on the preset number of partitions and cross-modal common features, functional lung partition results are generated through a partition model.
[0095] Specifically, the partitioning model can employ a clustering model based on self-supervised learning or unsupervised learning (e.g., using a K-Means clustering model or a Gaussian mixture model).
[0096] In some examples, a preset number of partitions (i.e., the number of partition categories) can be set based on the specific needs of clinical diagnosis and treatment scenarios. Then, the preset number of partitions and cross-modal common features can be input into the partitioning model to obtain the functional lung partitioning results output by the partitioning model. Finally, the functional lung partitioning results can be post-processed. For example, morphological operations can be used to smooth the functional lung partitioning results to eliminate noise and enhance the clarity of partition boundaries.
[0097] For example, taking a lung tumor radiotherapy scenario, to meet the requirements of radiotherapy planning for lung tumors (minimizing the radiation dose to high-functioning areas of the lungs and allowing low-functioning areas of the lungs to receive appropriate radiation doses), the preset number of partitions is set to 2, corresponding to two types of functional areas: high-functioning areas and low-functioning areas. Next, the feature values corresponding to each voxel point of cross-modal common features are expanded into a feature vector, and the preset number of partitions is set again. Then, a K-Means clustering model or a Gaussian mixture model is used to classify the feature vectors to obtain clustering results, and based on the clustering results, a functional lung partition mask map is generated.
[0098] Automatically obtain the corresponding partition mask. (When the functional lung distribution of some patients is uneven, this scenario can be regarded as a case where the functional lung partition category is unknown, so density-based noisy application clustering (DBSCAN) is used. The implementation method of K-Means clustering is as follows:) For example, in special scenarios where the functional distribution is extremely uneven (such as patients with lung metastases, where the tumors are distributed in a diffuse manner, resulting in fragmentation of functional regions), it is not possible to pre-set the number of partitions. Functional lung partitioning results can be generated based on cross-modal common features through the DBSCAN clustering (Density-Based Spatial Clustering of Applications with Noise) model.
[0099] By applying this embodiment, the number of preset partitions can be matched to the needs of different scenarios or departments, which can adapt to diverse clinical scenarios. Based on the preset number of partitions and cross-modal common features, functional lung partition results are generated through the partition model. The entire process does not require manual intervention, which can improve the accuracy of functional lung partition results and the efficiency of functional area division.
[0100] In some embodiments, the functional lung partitioning result is a functional lung partitioning image; the functional lung partitioning image includes multiple functional regions.
[0101] For example, when the functional lung partition image includes four functional regions, the functional lung partition image can be as follows: Figure 11 As shown, regions 10, 21, 22, and 30 are four functional regions. Region 10 is a functional region with certain functional defects, region 21 is a functional region with complete functional defects, region 22 is a functional region with temporary functional dysfunction, and region 30 is a functional region with normal function.
[0102] In some embodiments, the functional lung partition image includes multiple functional regions, and each functional region includes multiple functional regions.
[0103] For example, when the functional lung partition image includes six functional regions, the functional lung partition image can be as follows: Figure 12 As shown, regions 11, 12, 13, 14, 15, and 16 represent six functional regions. Region 11 represents the weakest functional region, comprising five functional areas: regions 111, 112, 113, 114, and 115. Regions 12, 13, 14, 15, and 16 represent functional regions that progressively improve lung function, with each category containing multiple functional areas.
[0104] By applying this embodiment, the flexible combination of multiple functional areas and different types of functional areas can meet the needs of multiple departments such as radiotherapy, respiratory medicine, and thoracic surgery. Each type or each functional area has a clear clinical significance and can be directly used for clinical diagnosis and treatment, thereby improving the efficiency of diagnosis and treatment and the accuracy of decision-making.
[0105] See Figure 13 , Figure 13 This is a flowchart illustrating a process for visualizing functional lung zoning results, as provided in this embodiment of the disclosure. Specifically, it includes steps S702-S704.
[0106] Step S702: Acquire lung anatomical images corresponding to at least one modality of functional lung imaging.
[0107] Specifically, lung anatomy images of the corresponding modality can be selected based on the modality of functional lung imaging (such as dynamic MRI, 4D CT, SPECT, PET-CT, etc.).
[0108] Step S704: Fuse the functional lung zoning results with lung anatomical images to obtain a fused image and display the fused image.
[0109] For example, before fusing the functional lung partitioning results with lung anatomical images, preprocessing operations such as denoising and grayscale normalization, and checking whether spatial registration has been performed can be performed on the lung anatomical images.
[0110] In some examples, a pixel-level semi-transparent overlay algorithm can be used to overlay the pixel values of the functional lung region results onto the pixel values of the anatomical image in a semi-transparent manner to generate a fused image. Labeling information (such as category labels for functional regions) can also be added at fixed locations in the fused image. Then, the fused image can be transmitted to a TPS (Transmission Positioning System) using the DICOM-RT protocol for visualization. Using this embodiment, the visualization of fused images supports medical systems such as TPS, providing doctors or clinical physicists with the spatial relationships between various functional areas of the lungs, target areas, and organs at risk. The results of functional lung zoning can be directly applied to the clinical diagnosis and treatment or radiotherapy planning process to guide the setting of more reasonable dose information for various functional areas of the lungs.
[0111] Corresponding to the above-described embodiments of the functional lung zoning method, this specification also provides embodiments of the functional lung zoning device. Figure 14 This is a schematic diagram illustrating the composition of the functional lung partitioning device provided in this disclosure. The processing device 120 may internally include a functional lung partitioning device 800 composed of multiple functional modules (logically or physically), which can be implemented through software programming. For example... Figure 14 As shown, the functional lung partitioning device 800 includes: Acquisition module 802: Used to acquire at least one modality of functional lung imaging of a target object.
[0112] Partitioning module 804: Used for functional lung imaging based on at least one modality, generating functional lung partitioning results through a pre-trained cross-modal partitioning model; the functional lung partitioning results are used to indicate the spatial distribution of multiple functional regions of the lung divided according to functional status.
[0113] For example, the cross-modal partitioning model includes a cross-modal feature extraction model and a partitioning model; the partitioning module 804 is further used to: extract cross-modal common features based on functional lung imaging of at least one modality through the cross-modal feature extraction model; and generate functional lung partitioning results based on the cross-modal common features through the partitioning model.
[0114] For example, the cross-modal feature extraction model includes multiple parallel single-modal feature extraction modules and single-modal feature fusion modules; each modality of functional lung imaging corresponds to a single-modal feature extraction module; each single-modal feature extraction module is used to receive functional lung imaging of one modality and extract the single-modal common features of functional lung imaging of that modality; the single-modal feature fusion module is used to fuse the single-modal common features extracted by each single-modal feature extraction module to obtain cross-modal common features.
[0115] For example, the functional lung partitioning device 800 further includes a training module 806 for constructing a training feature extraction model and using the training feature extraction model as a student model; the training feature extraction model includes multiple parallel training feature extraction modules and a training feature fusion module; a pre-trained multimodal feature extraction model is used as a teacher model to guide the training feature extraction model to obtain a cross-modal feature extraction model; wherein, each training feature extraction module is used to receive training functional lung imaging of one modality and extract the training single-modal common features of the training functional lung imaging of that modality; the training feature fusion module is used to fuse the training single-modal common features extracted by each training feature extraction module to obtain training cross-modal common features.
[0116] For example, each training feature extraction module includes a unimodal encoder and a unimodal decoder; the multimodal feature extraction model includes a multimodal encoder and a multimodal decoder; the training module 806 is further used to: fix the model weights of the teacher model, enable each unimodal encoder to learn the feature distribution of the multimodal encoder through knowledge distillation, and calculate the distillation loss function; randomly discard the common features of the training unimodal models, and calculate the reconstruction loss function based on the common features of the training cross-modal models and the multimodal common features output by the multimodal decoder; when the distillation loss function and the reconstruction loss function meet the preset conditions, complete the training of the training feature extraction model to obtain the cross-modal feature extraction model.
[0117] For example, the partitioning module 804 is further used to: generate functional lung partitioning results based on a preset number of partitions and cross-modal common features through a partitioning model.
[0118] For example, the functional lung partitioning result is a functional lung partitioning image; the functional lung partitioning image includes multiple functional regions, or the functional lung partitioning image includes multiple types of functional regions, and each type of functional region includes multiple functional regions.
[0119] For example, the functional lung partitioning device 800 further includes a fusion module 808 for acquiring lung anatomical images corresponding to at least one modality of functional lung imaging; fusing the functional lung partitioning results with the lung anatomical images to obtain a fused image; and displaying the fused image.
[0120] The functional lung zoning device of this disclosure is used to implement the corresponding functional lung zoning method in the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here. Furthermore, the functional implementation of each module in the functional lung zoning device of this disclosure can be referred to the description of the corresponding part in the foregoing method embodiments, which will also not be repeated here.
[0121] This disclosure also provides a computer-readable storage medium storing at least one instruction that is executed by a processor to implement the functional lung partitioning method as described in the above embodiments.
[0122] This disclosure also provides a computer program product including computer instructions stored in a computer-readable storage medium; a processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the functional lung partitioning method described in the above embodiments.
[0123] Those skilled in the art will recognize that the functions described in this disclosure in one or more of the examples above can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium accessible to a general-purpose or special-purpose computer.
[0124] It should be noted that the technical solutions described in this disclosure can be combined arbitrarily as long as they do not conflict.
[0125] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for functional lung zoning, characterized in that, include: Acquire functional lung imaging of at least one modality of the target object; Based on functional lung imaging of at least one modality, functional lung partitioning results are generated through a pre-trained cross-modal partitioning model; The functional lung zoning results are used to indicate the spatial distribution of multiple functional regions of the lungs according to their functional status.
2. The functional lung zoning method according to claim 1, characterized in that, The cross-modal partitioning model includes a cross-modal feature extraction model and a partitioning model; the functional lung imaging based on the at least one modality generates functional lung partitioning results through the pre-trained cross-modal partitioning model, including: Based on functional lung imaging of at least one modality, cross-modal common features are extracted using the cross-modal feature extraction model; Based on the cross-modal common features, the functional lung partitioning results are generated through the partitioning model.
3. The functional lung zoning method according to claim 2, characterized in that, The cross-modal feature extraction model includes multiple parallel single-modal feature extraction modules and single-modal feature fusion modules; each modality of functional lung imaging corresponds to one of the single-modal feature extraction modules; Each of the single-modal feature extraction modules is used to receive a functional lung imaging of one modality and extract the single-modal common features of the functional lung imaging of that modality; The single-modal feature fusion module is used to fuse the single-modal common features extracted by each single-modal feature extraction module to obtain the cross-modal common features.
4. The functional lung zoning method according to claim 3, characterized in that, The cross-modal feature extraction model is trained based on the following steps: A training feature extraction model is constructed and used as a student model; the training feature extraction model includes multiple parallel training feature extraction modules and training feature fusion modules. The pre-trained multimodal feature extraction model is used as the teacher model to guide the training of the feature extraction model, thereby obtaining the cross-modal feature extraction model; Each of the training feature extraction modules is used to receive a training functional lung imaging of a modality and extract the common features of the training single modality of the training functional lung imaging of that modality. The training feature fusion module is used to fuse the common features of the training single modality extracted by each training feature extraction module to obtain the common features of the training cross-modality.
5. The functional lung zoning method according to claim 4, characterized in that, Each of the training feature extraction modules includes a single-modal encoder and a single-modal decoder; the multimodal feature extraction model includes a multimodal encoder and a multimodal decoder; The step of using a pre-trained multimodal feature extraction model as a teacher model to guide the training of the feature extraction model, thereby obtaining the cross-modal feature extraction model, includes: The model weights of the teacher model are fixed, and through knowledge distillation, each of the single-modal encoders learns the feature distribution of the multimodal encoder, and the distillation loss function is calculated; The training single-modal common features are randomly discarded, and the reconstruction loss function is calculated based on the training cross-modal common features and the multimodal common features output by the multimodal decoder. When the distillation loss function and the reconstruction loss function satisfy preset conditions, the training of the training feature extraction model is completed, and the cross-modal feature extraction model is obtained.
6. The functional lung zoning method according to claim 2, characterized in that, The process of generating the functional lung partitioning results based on the cross-modal common features and through the partitioning model includes: Based on the preset number of partitions and the cross-modal common features, the functional lung partitioning results are generated through the partitioning model.
7. The functional lung zoning method according to claim 1, characterized in that, The functional lung partitioning result is a functional lung partitioning image; the functional lung partitioning image includes multiple functional regions, or the functional lung partitioning image includes multiple types of functional regions, and each type of functional region includes multiple functional regions.
8. The functional lung zoning method according to claim 1, characterized in that, The method further includes: Acquire lung anatomical images corresponding to the at least one modality of functional lung imaging; The functional lung zoning results are fused with the lung anatomical images to obtain a fused image, which is then displayed.
9. A functional lung partitioning device, characterized in that, include: Acquisition module, used to acquire at least one modality of functional lung imaging of the target object; The partitioning module is used for functional lung imaging based on at least one modality, and generates functional lung partitioning results through a pre-trained cross-modal partitioning model; The functional lung zoning results are used to indicate the spatial distribution of multiple functional regions of the lungs according to their functional status.
10. An electronic device, characterized in that, include: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the functional lung partitioning method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the functional lung partitioning method as described in any one of claims 1-8.