PET / CT (positron emission tomography / computed tomography) image whole-body focus identification method and device and medium

By training and optimizing PET/CT image data using the ConvNeXt network framework, the problems of low efficiency and poor accuracy in lesion identification in existing technologies are solved, and efficient and accurate lesion identification and quantitative analysis are achieved.

CN121937359APending Publication Date: 2026-04-28SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN UNIV
Filing Date
2025-12-01
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing PSMA PET/CT image lesion identification schemes suffer from low identification efficiency, poor identification consistency, and poor identification accuracy, especially in the identification of complex lesions.

Method used

The ConvNeXt network framework was used to train PET/CT image data. Through data preprocessing and multimodal spatial alignment, shared parameters and task-specific parameters were constructed. The objective function was optimized by joint segmentation and the parameters were updated by dual weights to improve the model's recognition ability.

Benefits of technology

It improves the efficiency, accuracy, and applicability of lesion identification, enabling accurate identification of lesion locations under different imaging agents and complex lesion scenarios, and supports end-to-end quantitative analysis and visualization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937359A_ABST
    Figure CN121937359A_ABST
Patent Text Reader

Abstract

The invention discloses a PET / CT (positron emission tomography / computed tomography) image whole-body focus recognition method and device and a medium, and the method comprises the steps: carrying out the data preprocessing of obtained sample image data of a plurality of different patients, so as to obtain corresponding modal image data, taking the modal image data as a training data set, and carrying out the recognition of a whole-body focus based on the training data set; the method comprises the following steps: training an initial recognition model adopting a ConvNeXt network framework to obtain a trained target recognition model, then obtaining target image data needing to be recognized, inputting the target image data into the target recognition model to obtain a recognition result output by the target recognition model, therefore, the target image data is analyzed and recognized through the target recognition model based on the ConvNeXt network framework, and the efficiency, accuracy and applicability of focus recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image recognition, and in particular to a PET / CT image whole-body lesion recognition method, device and medium. BACKGROUND

[0002] The PET image / CT image based on prostate-specific membrane antigen (PSMA, Prostate-Specific Membrane Antigen) is an important medical imaging technology applied to the diagnosis of prostate cancer, which can clearly image the active lesions of prostate cancer at the whole-body level.

[0003] The existing lesion analysis scheme for the PET image / CT image of PSMA includes manual analysis and automatic analysis. For manual analysis, there are defects of large recognition standard difference, low recognition efficiency, and poor recognition applicability. For automatic analysis, the analysis and recognition of lesions are usually realized based on a U-net algorithm framework model, but it is difficult to model the whole and depends on long-distance semantic relationship, which leads to a significant decline in performance in complex lesion recognition, and there is a defect of poor recognition accuracy.

[0004] Therefore, the existing scheme for lesion recognition of the PET image / CT image of PSMA has defects of low recognition efficiency, poor recognition applicability, and poor recognition accuracy. SUMMARY

[0005] The technical problem to be solved by the present application is that the existing scheme for lesion recognition of the PET image / CT image of PSMA has defects of low recognition efficiency, poor recognition consistency, and poor recognition accuracy. The purpose is to provide a PET / CT image whole-body lesion recognition method, which solves the problems of low recognition efficiency, single recognition applicability, and poor recognition accuracy existing in the existing scheme.

[0006] The present application is realized by the following technical scheme:

[0007] In a first aspect, the present application provides a PET / CT image whole-body lesion recognition method, comprising:

[0008] Obtaining sample image data of a plurality of different patients, wherein the sample image data includes PET images and CT images;

[0009] Data preprocessing is performed on each of the sample image data to obtain corresponding modal image data, and a plurality of the modal image data are used as a training data set;

[0010] Based on the training data set, an initial recognition model using a ConvNeXt network framework is trained to obtain a trained target recognition model;

[0011] Acquire the target image data to be identified, wherein the target image data includes target PET images and target CT images;

[0012] The target image data is input into the target recognition model to obtain the recognition result output by the target recognition model. The recognition result includes key quantitative indicator information and lesion recognition image, wherein the lesion recognition image is used to indicate the location of the lesion.

[0013] In one possible design, the data preprocessing of each of the sample image data to obtain the corresponding modality image data includes:

[0014] Extract metadata information from the PET image and the CT image, and crop the PET image and the CT image according to the metadata information to obtain the corresponding PET display area and CT display area;

[0015] The voxel intensity values ​​of the PET display area and the CT display area are normalized to ensure that the voxel intensities of the PET display area and the CT display area are at the same scale.

[0016] The voxel intensity values ​​are normalized and unified to the preset fixed resolution of the PET display area and the CT display area as the modal image data.

[0017] In one possible design, prior to training the initial recognition model using the ConvNeXt network framework, the design further includes multimodal spatial alignment of the sample image data and the creation of dual-channel input support to enable the ConvNeXt model to accept dual-channel input from the PET data and the CT data.

[0018] In one possible design, the multimodal spatial alignment of the sample image data and the creation of dual-channel input support include:

[0019] The PET display area and the CT display area are resampled according to the information of voxel space size and fixed resolution size, respectively, to achieve spatial alignment of PET image and CT image multimodal image;

[0020] The spatially aligned display area retains the CT area channel and adds the PET area channel to achieve a ConvNeXt model architecture that can simultaneously accept input of both PET and CT data and complete the spatial area data overlay function.

[0021] In one possible design, training the initial recognition model using the ConvNeXt network framework includes:

[0022] Construct shared parameters and task-specific parameters, and establish a loss function based on the shared parameters and the task-specific parameters;

[0023] Based on the loss function, the shared parameters, and the task-specific parameters, the worst-case degradation value is calculated and obtained.

[0024] Based on the worst-case degradation value, a joint segmentation optimization objective function is constructed. Through the joint segmentation optimization objective function, sensitivity to local unstable regions is maintained during the training of the initial recognition model.

[0025] In one possible design, the calculation to obtain the worst-case degradation value includes:

[0026] The perturbation parameters are calculated and obtained based on the loss function, the shared parameters, and the task-specific parameters.

[0027] Based on the loss function, the worst-case degradation value of the corresponding task layer is calculated according to the shared parameters, the task-specific parameters, and the perturbation parameters.

[0028] In one possible design, the joint segmentation optimization objective function is:

[0029] ;

[0030] in, For the shared parameters, These are the task-specific parameters. This is the worst-case degradation value. Let be the loss function.

[0031] In another possible design, a first dual weight and a second dual weight are constructed, wherein the first dual weight is used to indicate the degree of influence of the task loss during training, and the second dual weight is used to indicate the degree of influence of the worst-case degradation value during training.

[0032] Based on the first dual weight and the second dual weight, the shared parameters and the task-specific parameters are updated.

[0033] In one possible design, the key quantitative indicator information includes the maximum standard uptake value, the average standard uptake value, and the total metabolic uptake value of the tumor.

[0034] The maximum standard uptake value is used to indicate the maximum uptake of the injected imaging agent by the lesion in the target PET image and / or the target CT image.

[0035] The average standard uptake value is used to indicate the average uptake of the injected imaging agent by each lesion in the target PET image and / or the target CT image.

[0036] Secondly, this application provides a whole-body lesion identification device for PET / CT images, comprising:

[0037] The acquisition module is used to acquire sample image data from multiple different patients, wherein the sample image data includes PET images and CT images, perform data preprocessing on each sample image data to obtain corresponding modal image data, and use the multiple modal image data as a training dataset;

[0038] The processing module is used to train the initial recognition model using the ConvNeXt network framework based on the training dataset to obtain the trained target recognition model and to obtain the target image data to be recognized, wherein the target image data includes target PET images and target CT images.

[0039] The recognition module is used to input the target image data into the target recognition model to obtain the recognition result output by the target recognition model, wherein the recognition result includes key quantitative indicator information and lesion recognition image, wherein the lesion recognition image is used to indicate the location of the lesion.

[0040] Thirdly, this application provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the above-described method for identifying whole-body lesions in PET / CT images.

[0041] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0042] This application provides a method, device, and medium for whole-body lesion identification in PET / CT images. It preprocesses sample image data from multiple different patients to obtain corresponding modal image data. These modal image data are then used as a training dataset. Based on this training dataset, an initial identification model using the ConvNeXt network framework is trained to obtain a trained target identification model. The target image data to be identified is then acquired and input into the target identification model to obtain the identification results output by the model. These results include key quantitative indicator information and lesion identification images indicating the location of the lesions. Therefore, this application improves the efficiency, accuracy, and applicability of lesion identification by analyzing and identifying target image data using a target identification model based on the ConvNeXt network framework. Attached Figure Description

[0043] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings:

[0044] Figure 1 Flowchart of the method for identifying systemic lesions provided in the embodiments of this application Figure 1 ;

[0045] Figure 2 Flowchart of the method for identifying systemic lesions provided in the embodiments of this application Figure 2 ;

[0046] Figure 3 Comparative illustrations of this application and the prior art are provided;

[0047] Figure 4 The figures illustrate application examples of different developing agents provided in the embodiments of this application.

[0048] Figure 5 This is a schematic diagram of the whole-body lesion identification device provided in the embodiments of this application. Detailed Implementation

[0049] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.

[0050] PET / CT based on prostate-specific membrane antigen (PSMA) is a nuclear medicine imaging technology that has been widely used in the diagnosis of prostate cancer in recent years. It can clearly image active lesions of prostate cancer at the whole body level with the help of specific imaging agents that target PSMA expression points. Currently, the analysis methods for PSMA PET / CT images are mainly divided into three categories: manual, semi-automatic, and automated analysis.

[0051] For manual analysis, doctors rely on visually reviewing images, observing suspicious high-metabolic lesions layer by layer in PET / CT fusion images, manually outlining lesion boundaries, and measuring metabolic indicators. However, due to differences in image details acquired by different centers and equipment, the influence of dose and post-injection time on signal intensity, and differences in the experience of the image reader, the efficiency, accuracy, and consistency of lesion identification are low.

[0052] For the semi-automatic analysis approach, a fixed standard uptake value (SUV) parameter is manually set. Then, the local tissue semi-automatic delineation function in some image post-processing devices is used to mask areas below the threshold and automatically delineate areas above the threshold, identifying them as potential lesions. Next, each delineated lesion is manually verified and quantitatively analyzed. The main limitation of this method is that there is currently no fixed threshold parameter, and different medical institutions have significant differences in the setting of this value.

[0053] For automated analysis solutions, machine learning or deep learning models are typically used to segment lesion regions. The U-net algorithm framework is used to jointly model PET / CT multimodal images to achieve automatic lesion delineation. However, the U-net algorithm framework uses local convolutions and shallow skip connections, which are difficult to model and globally rely on long-distance semantic relationships. This leads to decreased recognition performance and poor accuracy in more complex lesion identification. Furthermore, there are currently various PSMA PET imaging agents available clinically, such as... 68 Ga、 18 F et al. showed significant differences in the physiological distribution, imaging contrast, and metabolic expression of different imaging agents in the human body. Existing automated analysis models can only identify a single imaging agent, limiting their clinical applicability and resulting in poor applicability.

[0054] Therefore, this application provides a method for whole-body lesion identification in PET / CT images. It preprocesses sample image data from multiple different patients to obtain corresponding modal image data, and uses these multiple modal image data as a training dataset. Based on this training dataset, an initial identification model using the ConvNeXt network framework is trained to obtain a trained target identification model. Then, the target image data to be identified is obtained, and the target image data is input into the target identification model to obtain the identification results output by the model. The identification results include key quantitative indicator information and lesion identification images indicating the location of the lesion. Therefore, this application improves the efficiency, accuracy, and applicability of lesion identification by analyzing and identifying target image data using a target identification model based on the ConvNeXt network framework.

[0055] The technical solutions of this application and how they solve the aforementioned technical problems are described in detail below using specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0056] Example 1

[0057] Figure 1 Flowchart of the method for identifying systemic lesions provided in the embodiments of this application Figure 1 ,like Figure 1 As shown, the method includes:

[0058] S101. Acquire sample image data from multiple different patients, wherein the sample image data includes PET images and CT images.

[0059] Specifically, with the consent of each patient, PET and / or CT images of each patient are obtained and used as sample image data for each patient.

[0060] S102. Perform data preprocessing on each of the sample image data to obtain the corresponding modal image data, and use the multiple modal image data as a training dataset.

[0061] Specifically, after acquiring multiple sample image data, for each sample image data, the corresponding DICOM image file is retrieved. Based on the DICOM image file, metadata information related to image quality and spatial attributes is extracted and analyzed. According to the extracted metadata information, the corresponding image display range, that is, the corresponding PET display area and CT display area, are determined.

[0062] Furthermore, based on the determined image display range, the voxel intensity values ​​of the image display range are normalized; the determined image display range is unified to a preset fixed resolution to reduce the adverse effects of spatial differences on network training and ensure the consistency of sample image data from multiple patients in the spatial dimension.

[0063] S103. Based on the training dataset, train the initial recognition model using the ConvNeXt network framework to obtain the trained target recognition model.

[0064] Specifically, after preprocessing the image data of each sample to obtain the training dataset, the modal image data in the training dataset is used as input data and imported into the ConvNeXt network in the initial recognition model for segmentation processing. Corresponding dedicated modules are added to each training layer of the ConvNeXt network to achieve deep supervision of the model training process.

[0065] Furthermore, the ConvNeXt network adopts a five-layer symmetric encoder-decoder architecture. In the encoding stage, the improved ConvNeXt network receives spatially aligned PET and CT modal data in parallel, extracts features through depthwise separable convolutional modules, and retains key spatial location information by combining continuous skip connections. At the same time, spatial downsampling is gradually achieved in each stage through convolutional operations with a stride of 2, and the channel dimensions are successively expanded to 32, 64, 128, 256 and 320.

[0066] Furthermore, in the decoding stage, the ConvNeXt network gradually restores the spatial resolution of the image through transposed convolutional layers, achieving fine segmentation boundary reconstruction. This stage utilizes multimodal feature information deeply fused in the encoder, effectively achieving deep fusion of metabolic and anatomical structural information between the two modalities, improving the recognition model's ability to express multimodal features and the overall performance of downstream segmentation tasks.

[0067] S104. Obtain the target image data to be identified, and input the target image data into the target recognition model to obtain the recognition result output by the target recognition model.

[0068] Specifically, after training the initial recognition model and obtaining the target recognition model, the target image data to be recognized is obtained, and the target image data, including the target PET image and the target CT image, is input into the target recognition model to obtain the recognition result, which includes key quantitative index information and lesion recognition image. The lesion recognition image is an image generated by superimposing the target PET image and the target CT image. The lesion recognition image is used to indicate the location of the lesion, and the key quantitative index information is used to indicate the degree of uptake of the imaging agent by the lesion.

[0069] Among them, the target PET image is used to provide anatomical information for body structure reference, and the target CT image is used to provide functional metabolic information of the lesion. Based on the anatomical information and the functional metabolic information of the lesion, the location of the lesion in the whole body is identified and located.

[0070] This application provides a method for whole-body lesion identification in PET / CT images. It preprocesses sample image data from multiple different patients to obtain corresponding modal image data. These modal image data are then used as a training dataset. Based on this training dataset, an initial identification model using the ConvNeXt network framework is trained to obtain a trained target identification model. The target image data to be identified is then acquired and input into the target identification model to obtain the identification results output by the model. These results include key quantitative indicator information and lesion identification images indicating the location of the lesions. Therefore, this application improves the efficiency, accuracy, and applicability of lesion identification by analyzing and identifying target image data using a target identification model based on the ConvNeXt network framework.

[0071] Example 2

[0072] Figure 2 Flowchart of the method for identifying systemic lesions provided in the embodiments of this application Figure 2 , Figure 3 These are comparative example diagrams provided for embodiments of this application and the prior art. Figure 4 These are application examples of different developing agents provided in the embodiments of this application, combined with... Figures 2 to 4 As shown, the method includes:

[0073] S201. Acquire sample image data from multiple different patients, wherein the sample image data includes PET images and CT images.

[0074] Specifically, the content of this step is the same as that of step S101, and will not be repeated here.

[0075] S202. Extract the metadata information of the PET image and the CT image, and crop the PET image and the CT image according to the metadata information to obtain the corresponding PET display area and CT display area.

[0076] Specifically, after acquiring multiple sample image data, for each sample image data, metadata information related to image quality and spatial attributes is extracted and analyzed through the DICOM image files of PET and CT images in the sample image data. The metadata information includes information such as window width, window level, voxel spacing, image size and orientation.

[0077] Furthermore, based on the window width and window level information extracted from the metadata, the image display range, i.e. the corresponding PET and CT display areas, is determined, thereby enabling the truncation of PET and CT images, which in turn enhances the visibility of the region of interest and improves the contrast and recognizability of the target tissue or lesion area.

[0078] S203. The voxel intensity values ​​are normalized and unified to the preset fixed resolution of the PET display area and the CT display area as the modal image data, and multiple modal image data are used as training datasets.

[0079] Specifically, after obtaining the corresponding PET and CT display areas, the voxel intensity values ​​of the PET and CT display areas are normalized, and the voxel intensity values ​​of each display area are linearly mapped to the [0, 1] interval to improve the intensity consistency and comparability between different modal images.

[0080] Furthermore, spatial resampling is performed on the image data of the PET and CT display areas based on the voxel spacing information, unifying all images to a preset fixed resolution such as 3.125mm×3.125mm×2.68mm, in order to reduce the adverse effects of spatial differences on the subsequent training of the ConvNeXt network framework.

[0081] S204. Construct shared parameters and task-specific parameters, and establish a loss function based on the shared parameters and task-specific parameters. Based on the loss function, train the initial recognition model using the ConvNeXt network framework.

[0082] Specifically, in clinical applications, especially in whole-body tumor analysis tasks based on PSMA PET / CT images, doctors often need to simultaneously obtain spatial distribution information of multiple important structures, such as organ boundaries and tumor regions, to support accurate diagnosis and treatment planning. In this scenario, there are multiple segmentation targets that differ significantly in anatomical attributes and imaging features.

[0083] Organ segmentation tasks typically rely on clear and stable CT anatomical structures with well-defined boundaries, making training convergence easier. Tumor segmentation, however, is often more challenging due to irregular lesion morphology, blurred boundaries, and strong metabolic heterogeneity, making modeling and optimization more difficult. Therefore, when collaboratively learning different segmentation tasks within a unified framework, conflicts in optimization objectives between tasks can easily arise. That is, some tasks, due to their feature advantages, dominate the training process, causing shared network parameters to shift towards their optimization direction, thus weakening the performance of other tasks.

[0084] Furthermore, before training the initial recognition model using the ConvNeXt network framework, the process also includes multimodal spatial alignment of each sample image data and the creation of dual-channel input support to enable the ConvNeXt model to accept dual-channel input of PET and CT data.

[0085] Specifically, the PET display area and the CT display area are resampled according to the information of voxel space size and fixed resolution size, respectively, to achieve spatial alignment of the PET image and the CT image multimodal image. The spatially aligned display area retains the CT area channel and adds the PET area channel to realize the ConvNeXt model architecture that can simultaneously accept the input of the PET data and the CT data dual-channel data and complete the spatial region data overlay function.

[0086] Therefore, this application coordinates the performance degradation risk of different tasks under worst-case conditions during training and constructs shared parameters. and mission-specific parameters And based on shared parameters and mission-specific parameters , construct the first The loss function for each task is: .

[0087] S205. Calculate and obtain the worst-case degradation value based on the loss function, the shared parameters, and the task-specific parameters.

[0088] Specifically, based on the loss function, shared parameters, and task-specific parameters, the perturbation parameters are calculated and obtained using the following formula (1):

[0089] (1);

[0090] in, For disturbance parameters, To share parameters, For mission-specific parameters, The loss function is used as the basis for calculating the worst-case degradation value of the corresponding task layer based on the shared parameters, task-specific parameters, and perturbation parameters. The worst-case degradation value is obtained by the following formula (2):

[0091] (2);

[0092] in, This is the worst-case degradation value. It indicates the direction and magnitude of the maximum loss in task performance caused by system perturbation under L2 norm constraints. The worst-case degradation value can be obtained by one approximate perturbation calculation, thereby capturing the stability of the ConvNeXt model for all tasks near the current parameter point without significantly increasing the computational burden, thus improving the consistency and practicality of the segmentation criteria for different tissues, organs and tumor lesions.

[0093] S206. Based on the worst-case degradation value, construct a joint segmentation optimization objective function. Through the joint segmentation optimization objective function, maintain sensitivity to local unstable regions during the training process of the initial recognition model.

[0094] Specifically, after obtaining the worst-case degradation value, a joint segmentation optimization objective function is constructed based on the worst-case degradation value, loss function, shared parameters, and task-specific parameters, as shown in the following equation (3):

[0095] (3);

[0096] in, To share parameters, For mission-specific parameters, This is the worst-case degradation value. This is the loss function.

[0097] Furthermore, based on minimizing the conventional maximum task loss, the degradation effect caused by the worst perturbation is explicitly introduced as an auxiliary regularization term, which makes the initial recognition model sensitive to local unstable regions during training, and thus tends to a more "flat" parameter space. This helps to improve the generalization ability of the target recognition model on new samples or cross-center data after training.

[0098] In the PSMA PET / CT whole-body tumor analysis scenario, the initial recognition model is trained based on the joint segmentation optimization objective function. The resulting target recognition model can alleviate the task optimization difficulties caused by the ambiguity of lesion boundaries and the heterogeneity of metabolic expression. This allows the target recognition model to more accurately identify low-contrast and morphologically complex tumor regions while maintaining the accuracy of organ structure segmentation, thereby improving the consistency and practicality of the overall segmentation effect.

[0099] S207. Construct a first dual weight and a second dual weight, and update the shared parameters and the task-specific parameters based on the first dual weight and the second dual weight.

[0100] Specifically, when training the initial recognition model, the shared parameters and task-specific parameters are updated. The update of the shared parameters is shown in the following equation (4):

[0101] η (4);

[0102] in, As the first dual weight, For the second dual weight, To share parameters, For mission-specific parameters, For the worst-case degradation value, the update of the task-specific parameters is as shown in equation (5):

[0103] (5);

[0104] in, As the first dual weight, For the second dual weight, To share parameters, For mission-specific parameters, The worst-case degradation value is dynamically updated through the first and second dual weights, so that each task can adaptively adjust the optimization step size according to its current error and robustness during training, effectively preventing gradient shift caused by competition between tasks and improving the overall training stability.

[0105] S208. After obtaining the trained target recognition model, obtain the target image data to be recognized.

[0106] S209. Input the target image data into the target recognition model to obtain the recognition result output by the target recognition model, wherein the recognition result includes key quantitative indicator information and lesion recognition image.

[0107] Specifically, target image data, including target PET images and target CT images, are input into the target recognition model to obtain the recognition results, including key quantitative indicator information and lesion recognition images. The key quantitative indicator information includes the maximum standard uptake value SUVmax, the average standard uptake value SUVmean, and the total tumor metabolic uptake value TLU (Total Lesion Uptake).

[0108] The maximum standard uptake value (SUVmax) indicates the maximum uptake of the injected imaging agent by the lesion in the fused image formed by superimposing the target PET image and the target CT image; the mean standard uptake value (SUVmean) indicates the average uptake of the injected imaging agent by each lesion in the fused image formed by superimposing the target PET image and the target CT image.

[0109] Furthermore, after the key quantitative indicators at the lesion level are calculated, aggregate analysis can be performed on all lesion areas to output parameters such as total SUVmax, total tumor metabolic volume (MTV), and total TLU at the systemic level, thereby comprehensively reflecting the patient's overall tumor metabolic burden and activity level.

[0110] Furthermore, for the output lesion identification image, transparent pseudo-color overlay display of the original image and the segmentation mask is supported. Users can simultaneously view the original target PET and target CT images, as well as the corresponding generated lesion identification image, on slices in any direction.

[0111] Furthermore, such as Figure 3 As shown, in the imaging of a patient, the existing U-net model mistakenly identified the physiological uptake of the esophagus and ureter near the bladder as tumor lesions. The whole-body lesion identification method provided in this application successfully avoids this misjudgment.

[0112] Furthermore, such as Figure 4 As shown, for different developers such as 68 Ga-PSMA-11 and 18 Automated analysis of PET / CT images of patients with different tumor burdens using F-PSMA-1007, among which 18 F-PSMA-1007 developer and 68 The difference of Ga-PSMA-11 is that... 18 The F-PSMA-1007 imaging agent focuses more on the physiological uptake by the liver and intestines, but the whole-body lesion identification method provided in this application can still accurately identify whole-body lesions and eliminate interference from the physiological uptake by normal organs.

[0113] All quantitative analysis results can be exported as CSV tables, and this invention supports interactive modification and version control of prediction results. Users can manually edit segmentation boundaries, mark or remove abnormal lesions, and the editing record can be automatically saved and the corresponding indicators updated in real time. Ultimately, the entire "segmentation-analysis-visualization-output" process is integrated end-to-end, significantly improving the quantitative efficiency of nuclear medicine images, the degree of analysis standardization, and the convenience of physician operation, providing technical support for large-scale clinical application and remote assisted diagnosis.

[0114] This application provides a method for whole-body lesion identification in PET / CT images. It preprocesses sample image data from multiple different patients to obtain corresponding modal image data. These modal image data are then used as a training dataset. Based on this training dataset, an initial identification model using the ConvNeXt network framework is trained to obtain a trained target identification model. The target image data to be identified is then acquired and input into the target identification model to obtain the identification results output by the model. These results include key quantitative indicator information and lesion identification images indicating the location of the lesions. Therefore, this application improves the efficiency, accuracy, and applicability of lesion identification by analyzing and identifying target image data using a target identification model based on the ConvNeXt network framework.

[0115] Figure 5 This is a schematic diagram of the whole-body lesion identification device provided in the embodiments of this application, as shown below. Figure 5 As shown, the device 500 includes:

[0116] The acquisition module 501 is used to acquire sample image data from multiple different patients, wherein the sample image data includes PET images and CT images, perform data preprocessing on each sample image data to obtain corresponding modal image data, and use the multiple modal image data as a training dataset;

[0117] Processing module 502 is used to train an initial recognition model using the ConvNeXt network framework based on the training dataset to obtain a trained target recognition model and to obtain target image data to be recognized, wherein the target image data includes target PET images and target CT images.

[0118] The recognition module 503 is used to input the target image data into the target recognition model to obtain the recognition result output by the target recognition model, wherein the recognition result includes key quantitative indicator information and lesion recognition image, wherein the lesion recognition image is used to indicate the location of the lesion.

[0119] Further, the processing module 502 is specifically used to extract metadata information of the PET image and the CT image, and to crop the PET image and the CT image according to the metadata information to obtain the corresponding PET display area and CT display area;

[0120] The voxel intensity values ​​of the PET display area and the CT display area are normalized to ensure that the voxel intensities of the PET display area and the CT display area are at the same scale.

[0121] The voxel intensity values ​​are normalized and unified to the preset fixed resolution of the PET display area and the CT display area as the modal image data.

[0122] Furthermore, before training the initial recognition model using the ConvNeXt network framework, the processing module 502 is also used to perform multimodal spatial alignment on each of the sample image data and create dual-channel input support, so that the ConvNeXt model can accept dual-channel input of the PET data and the CT data.

[0123] Furthermore, the processing module 502 is specifically used to resample the PET display area and the CT display area according to the information of voxel space size and fixed resolution size, respectively, so as to achieve spatial alignment of the PET image and the CT image multimodal image.

[0124] The spatially aligned display area retains the CT area channel and adds the PET area channel to achieve a ConvNeXt model architecture that can simultaneously accept input of both PET and CT data and complete the spatial area data overlay function.

[0125] Furthermore, the processing module 502 is specifically used to construct shared parameters and task-specific parameters, and to establish a loss function based on the shared parameters and the task-specific parameters;

[0126] Based on the loss function, the shared parameters, and the task-specific parameters, the worst-case degradation value is calculated and obtained.

[0127] Based on the worst-case degradation value, a joint segmentation optimization objective function is constructed. Through the joint segmentation optimization objective function, sensitivity to local unstable regions is maintained during the training of the initial recognition model.

[0128] Further, the processing module 502 is specifically used to calculate and obtain the perturbation parameters based on the loss function, the shared parameters, and the task-specific parameters;

[0129] Based on the loss function, the worst-case degradation value of the corresponding task layer is calculated according to the shared parameters, the task-specific parameters, and the perturbation parameters.

[0130] Furthermore, the joint segmentation optimization objective function in processing module 502 is:

[0131] ;

[0132] in, For the shared parameters, These are the task-specific parameters. This is the worst-case degradation value. Let be the loss function.

[0133] Further, the processing module 502 is specifically used to construct a first dual weight and a second dual weight, wherein the first dual weight is used to indicate the degree of influence of the task loss during the training process, and the second dual weight is used to indicate the degree of influence of the worst-case degradation value during the training process.

[0134] Based on the first dual weight and the second dual weight, the shared parameters and the task-specific parameters are updated.

[0135] Furthermore, the key quantitative indicator information in the identification module 503 includes the maximum standard uptake value, the average standard uptake value, and the total metabolic uptake value of the tumor.

[0136] The maximum standard uptake value is used to indicate the maximum uptake of the injected imaging agent by the lesion in the target PET image and / or the target CT image.

[0137] The average standard uptake value is used to indicate the average uptake of the injected imaging agent by each lesion in the target PET image and / or the target CT image.

[0138] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the whole-body lesion identification method as described above.

[0139] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for identifying whole-body lesions in PET / CT images, characterized in that, The method includes: Acquire sample image data from multiple different patients, including PET images and CT images; The sample image data is preprocessed to obtain the corresponding modal image data, and the multiple modal image data are used as a training dataset. Based on the training dataset, the initial recognition model using the ConvNeXt network framework is trained to obtain the trained target recognition model. Acquire the target image data to be identified, wherein the target image data includes target PET images and target CT images; The target image data is input into the target recognition model to obtain the recognition result output by the target recognition model. The recognition result includes key quantitative indicator information and lesion recognition image, which is used to indicate the location of the lesion.

2. The method for identifying whole-body lesions in PET / CT images according to claim 1, characterized in that, The step of preprocessing the sample image data to obtain the corresponding modal image data includes: Extract metadata information from the PET image and the CT image, and crop the PET image and the CT image according to the metadata information to obtain the corresponding PET display area and CT display area; The voxel intensity values ​​of the PET display area and the CT display area are normalized to ensure that the voxel intensities of the PET display area and the CT display area are at the same scale. The voxel intensity values ​​are normalized and unified to the preset fixed resolution of the PET display area and the CT display area as the modal image data.

3. The method for identifying whole-body lesions in PET / CT images according to claim 2, characterized in that, Before training the initial recognition model using the ConvNeXt network framework, the method further includes performing multimodal spatial alignment on each of the sample image data and creating dual-channel input support so that the ConvNeXt model can accept dual-channel input from the PET data and the CT data.

4. The method for identifying whole-body lesions in PET / CT images according to claim 3, characterized in that, The process of performing multimodal spatial alignment on each of the sample image data and creating dual-channel input support includes: The PET display area and the CT display area are resampled according to the information of voxel space size and fixed resolution size, respectively, to achieve spatial alignment of PET image and CT image multimodal image; The spatially aligned display area retains the CT area channel and adds the PET area channel to achieve a ConvNeXt model architecture that can simultaneously accept input of both PET and CT data and complete the spatial area data overlay function.

5. The method for identifying whole-body lesions in PET / CT images according to claim 1, characterized in that, The training of the initial recognition model using the ConvNeXt network framework includes: Construct shared parameters and task-specific parameters, and establish a loss function based on the shared parameters and the task-specific parameters; Based on the loss function, the shared parameters, and the task-specific parameters, the worst-case degradation value is calculated and obtained. Based on the worst-case degradation value, a joint segmentation optimization objective function is constructed. Through the joint segmentation optimization objective function, sensitivity to local unstable regions is maintained during the training of the initial recognition model.

6. The method for identifying whole-body lesions in PET / CT images according to claim 5, characterized in that, The calculation to obtain the worst-case degradation value includes: The perturbation parameters are calculated and obtained based on the loss function, the shared parameters, and the task-specific parameters. Based on the loss function, the worst-case degradation value of the corresponding task layer is calculated according to the shared parameters, the task-specific parameters, and the perturbation parameters.

7. The method for identifying whole-body lesions in PET / CT images according to claim 5, characterized in that, The joint segmentation optimization objective function is: ; in, For the shared parameters, These are the task-specific parameters. This is the worst-case degradation value. Let be the loss function.

8. The method for identifying whole-body lesions in PET / CT images according to claim 5, characterized in that, It also includes constructing a first dual weight and a second dual weight, wherein the first dual weight is used to indicate the degree of influence of the task loss during the training process, and the second dual weight is used to indicate the degree of influence of the worst-case degradation value during the training process; Based on the first dual weight and the second dual weight, the shared parameters and the task-specific parameters are updated.

9. The method for identifying whole-body lesions in PET / CT images according to claim 1, characterized in that, The key quantitative indicators include the maximum standard uptake value, the average standard uptake value, and the total metabolic uptake value of the tumor. The maximum standard uptake value is used to indicate the maximum uptake of the injected imaging agent by the lesion in the target PET image and / or the target CT image. The average standard uptake value is used to indicate the average uptake of the injected imaging agent by each lesion in the target PET image and / or the target CT image.

10. A device for identifying whole-body lesions in PET / CT images, characterized in that, include: The acquisition module is used to acquire sample image data from multiple different patients, wherein the sample image data includes PET images and CT images, perform data preprocessing on each sample image data to obtain corresponding modal image data, and use the multiple modal image data as a training dataset; The processing module is used to train the initial recognition model using the ConvNeXt network framework based on the training dataset to obtain the trained target recognition model and to obtain the target image data to be recognized, wherein the target image data includes target PET images and target CT images. The recognition module is used to input the target image data into the target recognition model to obtain the recognition result output by the target recognition model, wherein the recognition result includes key quantitative indicator information and lesion recognition image, and the lesion recognition image is used to indicate the location of the lesion.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the whole-body lesion identification in PET / CT images as described in any one of claims 1 to 9.