Automatic segmentation and four-classification discrimination method and device for spinal isolated lesion image

By preprocessing multimodal image data and using a dual-domain multi-stream residual convolutional feature encoding network, the problem of insufficient specificity of MRI technology in processing image data of isolated abnormal areas of the spine is solved. This achieves efficient and reliable automated segmentation and four-class classification, supports interactive fine-tuning by doctors, and improves the standardization and consistency of image data.

CN121904072APending Publication Date: 2026-04-21PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY)
Filing Date
2025-12-31
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Conventional MRI technology lacks the ability to effectively capture microstructural and functional features when acquiring imaging data of isolated abnormal areas of the spine, resulting in overlapping and intersecting imaging data, insufficient specificity, and reliance on human experience for image interpretation, leading to poor consistency of results. Traditional detection methods are highly invasive, complex to operate, and time-consuming to interpret parameters, making them difficult to apply widely.

Method used

Multimodal image data preprocessing was employed, and the data was uniformly converted to NIfTI format and the voxel spacing was standardized. Zero-sample or interactive segmentation was performed using the MedSAM2 model. Complementary information was mined by combining a dual-domain multi-stream residual convolutional feature encoding network to generate four-class classification results for isolated spinal lesions. The results were visualized on a client-side platform to support interactive fine-tuning by doctors.

Benefits of technology

It enables standardized processing of image data, improves the efficiency and accuracy of segmentation and classification, ensures the reliability and consistency of results, and meets the needs of clinical data exchange.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121904072A_ABST
    Figure CN121904072A_ABST
Patent Text Reader

Abstract

The invention provides an automatic segmentation and four-classification discrimination method and device for a spinal isolated lesion image, and is applied to the technical field of data processing. The method comprises the following steps: acquiring a multi-modal image and clinical feature information, and preprocessing the image to generate standardized data; using a MedSAM2 model to realize zero sample or interactive segmentation, and cutting and aligning to generate an ROI input block; the method comprises the following steps of: fully mining complementary information under different imaging mechanisms through processing of a double-domain multi-stream residual convolutional feature coding network oriented to a structural phase and a functional phase, and outputting a four-classification result and confidence; the comprehensive classification probability is generated through multi-stage fusion, and finally, the client side is visualized and supports fine adjustment of a doctor, standard mask information is exported, and efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to an automated segmentation and four-class classification method and apparatus for images of isolated spinal lesions. Background Technology

[0002] Conventional magnetic resonance imaging (MRI) technology can only provide limited morphological information when acquiring data on isolated abnormal areas of the spine. For isolated abnormal areas of the spine with different biological attributes, the image data acquired by this technology often overlaps and intersects in terms of characteristics, resulting in insufficient specificity when determining attributes based on the image data. The fundamental reason for this is that conventional MRI technology lacks the ability to effectively capture the microstructural and functional characteristics of abnormal areas. Image data interpretation is highly dependent on the experience of operators. Operators with different experience levels exhibit significant differences in data processing steps such as defining abnormal area boundaries, extracting image features, and classifying data types, resulting in insufficient inter-observer consistency in data interpretation results. This problem mainly stems from the lack of objective, standardized quantitative analysis indicators. While traditional tissue sample testing methods are often used as a reference for determining the attributes of abnormal areas, they are invasive procedures with problems such as sample collection bias, postoperative complications, and low subject cooperation, making them unsuitable for widespread application and repeated data collection and evaluation. The core issue lies in the physical damage to the subject caused by the procedure itself and the limitations of the sample collection range. Multi-b-value diffusion-weighted imaging (mb-DWI) can theoretically capture microstructure and functional data in anomalous regions. However, its raw data interpretation process is complex, parameter extraction is time-consuming, and requires a high level of expertise from operators. This not only increases the workload of data processing but also limits the widespread application of this technology in routine data acquisition scenarios. The key reason for this problem lies in the high dimensionality of the data and the lack of a unified and standardized parameter interpretation process.

[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part by practice of the invention.

[0005] According to one aspect of this application, an automated segmentation and four-class classification method for images of solitary spinal lesions is provided, comprising: acquiring multimodal image data and clinical feature information; preprocessing the multimodal image data, uniformly converting it to NIfTI format and standardizing voxel spacing, completing spatial registration with T1C+ as structural image reference and b=0s / mm² sequence as diffusion image reference, evaluating registration accuracy through root mean square error, and generating standardized multimodal image data; processing the standardized multimodal image data based on the MedSAM2 model, supporting zero-shot automatic segmentation or segmentation assisted by a small amount of doctor interaction prompts, generating lesion mask files and candidate lesion lists, and then cropping the smallest cube RO. Simultaneously align the corresponding regions of the multimodal systems to generate registered multimodal ROI input blocks. Process the multimodal ROI input blocks using a dual-domain multi-stream residual convolutional feature encoding network oriented towards structural and functional phases, mining complementary information under different imaging mechanisms to generate four-class classification results and confidence probability values ​​for isolated spinal lesions. Generate comprehensive classification probabilities through a fixed-order splicing strategy in the early stage, attention-weighted feature fusion in the middle stage, or weighted voting / logistic regression fusion strategy based on modal performance weight allocation in the later stage. Process the segmentation indicators, classification indicators, and classification results using client-side visualization display functions, supporting interactive fine-tuning by doctors, and generating exportable DICOMSEG / NIfTI format mask information.

[0006] Another aspect of this application discloses an automated segmentation and four-class classification device for images of solitary spinal lesions, comprising: an acquisition module for acquiring multimodal image data and clinical feature information; a processing module for preprocessing the multimodal image data, uniformly converting it to NIfTI format and standardizing voxel spacing, completing spatial registration using T1C+ as structural image reference and b=0s / mm² sequence as diffusion image reference, evaluating registration accuracy through root mean square error, and generating standardized multimodal image data; processing the standardized multimodal image data based on the MedSAM2 model, supporting zero-sample automatic segmentation or segmentation assisted by a small amount of doctor interaction prompts, generating lesion mask files and candidate lesion lists, and then cropping out the smallest... A cubic ROI is simultaneously aligned with the corresponding multimodal regions to generate a registered multimodal ROI input block. The multimodal ROI input block is processed using a dual-domain multi-stream residual convolutional feature encoding network oriented towards structural and functional phases to mine complementary information under different imaging mechanisms, generating four-class classification results and confidence probability values ​​for isolated spinal lesions. A comprehensive classification probability is generated through a strategy of fixed-order stitching in the early stage, attention-weighted feature fusion in the middle stage, or weighted voting / logistic regression fusion based on modal performance in the later stage. The segmentation indicators, classification indicators, and classification results are processed using a client-side visualization function, supporting interactive fine-tuning by doctors, and generating exportable DICOMSEG / NIfTI format mask information.

[0007] According to another aspect of this application, an electronic device includes: a first processor; and a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute the above-described automated segmentation and four-class classification method for images of isolated spinal lesions by executing the executable instructions.

[0008] According to another aspect of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a second processor, implements the above-described method for automated segmentation and four-class classification of images of isolated spinal lesions.

[0009] This application provides an automated segmentation and four-class classification method and device for images of isolated spinal lesions. Addressing the pain points in processing images of isolated abnormal spinal regions, this application offers an automated segmentation and four-class classification data processing solution. First, multimodal images and clinical feature information are acquired. Standardized data is generated through format conversion, voxel standardization, hierarchical spatial registration, and RMSE accuracy verification. Then, zero-sample or interactive segmentation is achieved using the MedSAM2 model, and ROI input blocks are generated by cropping and alignment. The multimodal ROI input blocks are processed based on a dual-domain multi-stream residual convolutional feature encoding network oriented towards structural and functional phases, mining complementary information under different imaging mechanisms to generate four-class classification results and confidence probability values ​​for isolated spinal lesions. Finally, the results are visualized on a client-side interface, supporting interactive fine-tuning by doctors and exporting DICOMSEG / NIfTI format mask information. The technical advantages lie in standardized preprocessing ensuring data consistency, improving segmentation and classification efficiency and accuracy, multimodal fusion enhancing result reliability, and adapting to clinical data exchange and research needs.

[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0011] Figure 1 The flowchart illustrates an automated segmentation and four-class classification method for images of isolated spinal lesions provided in an embodiment of this application. Figure 2 This diagram illustrates the structure of an automated segmentation and four-class classification device for images of isolated spinal lesions provided in an embodiment of this application. Detailed Implementation

[0012] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0013] The following is combined Figure 1This application describes an automated segmentation and four-class classification method for images of solitary spinal lesions according to exemplary embodiments thereof. It should be noted that the application scenarios described below are merely illustrative for understanding the spirit and principles of this application, and the embodiments of this application are not limited in any way. Rather, the embodiments of this application are applicable to any suitable scenario.

[0014] In one implementation, Figure 1 The schematic diagram illustrates a flowchart of an automated segmentation and four-class classification method for images of isolated spinal lesions according to an embodiment of this application.

[0015] S101, acquire multimodal imaging data and clinical feature information.

[0016] In one implementation, the multimodal imaging data must include structural imaging sequences and functional imaging sequences. Structural imaging includes three sequences: T2 (T2-weighted sequence), T2-FS (T2-weighted fat-suppressed sequence), and enhanced T1C+ (enhanced T1-weighted sequence). Functional imaging uses three typical b-values ​​for diffusion-weighted imaging (b=0 s / mm², b=800 s / mm², b=1000 s / mm²). Clinical characteristic information, encompassing basic patient information, medical history, and laboratory test results, must be correlated one-to-one with the imaging data.

[0017] Specifically, raw spinal MRI data from patients were extracted from the hospital's PACS (Picture Archiving and Communication System), primarily in DICOM format (a standard clinical image storage format). The selection criteria were as follows: patients underwent spinal MRI within 4 weeks; images contained the required structural and functional sequences; lesion diameter ≥1cm (minimizing volume effects); and image quality free of severe motion artifacts. Exclusion criteria included contraindications to MRI (e.g., incompatible metal implants), multiple spinal bone lesions, incomplete image sequences, or artifacts affecting diagnosis. The selected DICOM images were uniformly converted to NIfTI format, which is easier for deep learning models to read and process, and supports dimensional alignment of multimodal data. A uniform voxel spacing of 1mm × 1mm × 1mm was set to eliminate inconsistencies in voxel size caused by differences in imaging parameters between different hospitals and equipment, ensuring consistent image spatial scale.

[0018] S102 preprocesses the multimodal image data, converts it into NIfTI format and standardizes the voxel spacing, completes spatial registration using T1C+ as the structural image reference and b=0s / mm² sequence as the diffusion image reference, evaluates the registration accuracy through root mean square error, and generates standardized multimodal image data.

[0019] In one implementation, to address the heterogeneous nature of multimodal image data, which includes structural and functional sequences, a comprehensive preprocessing mechanism is established, encompassing format conversion, NIfTI standardization, voxel spacing unification, spatial registration, and accuracy assessment to achieve data regularization. Then, a hierarchical registration strategy is employed, using T1C+ as a reference for structural images and b=0s / mm² sequences as a reference for diffuse images. This is followed by the step-by-step execution of rigid and affine registration, resulting in a multimodal image dataset covering all elements: format, spacing, coordinates, and accuracy. The comprehensive preprocessing mechanism operates around five steps: format conversion, NIfTI standardization, voxel spacing unification, spatial registration, and accuracy assessment. This addresses the heterogeneity issue of multimodal images (structural sequences + functional sequences) and ensures data regularity.

[0020] Raw multimodal image data (mostly in DICOM format, a standard clinical storage format) from different hospitals and devices were collected and uniformly converted to NIfTI format. NIfTI format supports multi-dimensional data storage, facilitating deep learning model reading and processing, and can directly correlate the spatial coordinate information of multimodal images, laying the foundation for subsequent registration. A patient's multimodal images included T1C+ (structural sequence), T2 (structural sequence), T2-FS (structural sequence), b=0 s / mm² (functional sequence), b=800 s / mm² (functional sequence), and b=1000 s / mm² (functional sequence). The raw data were all in DICOM format; after conversion, all were transformed to NIfTI format, preserving the original anatomical structures and signal characteristics of each sequence.

[0021] For the converted NIfTI format images, standardize the data structure, including pixel value range normalization, data dimension alignment (e.g., unify to a 3D matrix format: 512×512×20 layers), and file naming rules (e.g., "Patient ID_Sequence Type_Timestamp.nii"). This eliminates inconsistencies in NIfTI file structures caused by differences in imaging parameters from different devices, ensuring compatibility in subsequent processing workflows. The original pixel value range of Patient A's T1C+ sequence NIfTI file was 0-4095, T2 sequence was 0-2560, and T2-FS sequence was 0-2048. After standardization, the pixel values ​​of all three were normalized to the 0-1 range. At the same time, the 3D matrix of all sequences was unified to 512×512×20 layers, and the files were named "P2025001_T1C+_20250520.nii", "P2025001_T2_20250520.nii", and "P2025001_T2-FS_20250520.nii".

[0022] To address the voxel spacing differences across different image modalities (e.g., 0.9mm×0.9mm×2.8mm for T1C+, 1.0mm×1.0mm×3.0mm for T2, and 0.8mm×0.8mm×3mm for T2-FS, while 1.5mm×1.5mm×4mm for functional sequences), a uniform target voxel spacing (1mm×1mm×1mm in this scheme) is set, and image spatial resolution is adjusted using resampling technology. This ensures that images from different modalities of the same patient are consistent in spatial scale, avoiding measurement deviations in lesion location and size due to differences in voxel size. The original voxel spacing of patient A's b=0s / mm² sequence was 1.2mm×1.2mm×3.5mm, which was adjusted to 1mm×1mm×1mm after resampling; the original voxel spacing of the T1C+ sequence was 0.9mm×0.9mm×2.8mm, the original voxel spacing of the T2 sequence was 1.0mm×1.0mm×3.0mm, and the original voxel spacing of the T2-FS sequence was 0.8mm×0.8mm×3mm. After resampling, all were uniformly adjusted to 1mm×1mm×1mm. The proportions of the anatomical structures in the images remained unchanged, with only the spatial resolution being adapted to a uniform standard.

[0023] A "layered registration strategy + step-by-step execution" approach is adopted. First, registration layers are assigned according to modal type, and then coordinate alignment is completed through rigid registration + affine registration. Structural images use T1C+ as a reference, and functional images use b=0s / mm² as a reference, avoiding accuracy deviations caused by direct registration between different modalities. Translation and rotation corrections are performed on target images (such as T2, T2-FS, b=1000s / mm²) to eliminate global offsets caused by slight changes in patient scanning position, ensuring that the overall image orientation is consistent with the reference sequence. Based on rigid registration, the scale and shear transformation of the target images are further adjusted to eliminate global deformations caused by differences in imaging equipment parameters (such as magnetic field strength and scanning matrix), achieving more refined spatial alignment. Using patient A's T1C+ sequence as a reference, registration was performed on the T2-FS sequence: first, rigid registration corrected for a 3° rotational deviation and a 2mm translational deviation; then, affine registration adjusted for a 5% scale difference, ultimately ensuring that the spatial coordinates of each anatomical point in the T2-FS sequence (such as vertebral body margins and vascular pathways) were consistent with the corresponding points in the T1C+ sequence. Using the b=0 s / mm² sequence as a reference, the same process was performed on the b=1000 s / mm² sequence to complete the coordinate alignment within the functional sequences. Using patient A's T1C+ sequence as a reference, registration was performed on the T2 sequence, with the following specific procedure: The T1C+ sequence serves as the core reference for structural imaging, possessing clear vertebral anatomical structures and lesion enhancement features. The T2 sequence needs to be precisely aligned with it in spatial coordinates through registration, providing consistent input for the subsequent dual-domain multi-stream residual convolutional feature encoding network. First, rigid registration was performed to correct the 3° rotational deviation and 2.1mm translational deviation between the T2 and T1C+ sequences caused by the patient A's scanning position, ensuring that the overall anatomical orientation of the two sequences was consistent. Then, affine registration was performed to adjust for the 4% scale difference in the T2 sequence, eliminating global deformation caused by different imaging equipment parameters and achieving fine spatial alignment. The root mean square error (RMSE) was used to quantify the registration effect. The RMSE of the registered T2 sequence for patient A was 1.4mm, meeting the accuracy control threshold of RMSE < 1.5mm for structural sequence registration, and was therefore deemed qualified. After registration, the spatial deviation of key anatomical structures such as vertebral body edges and lesion boundaries in the T2 sequence from their corresponding positions in the T1C+ sequence was ≤ 1mm, ensuring that the dual-domain network could effectively extract lesion signal features from the T2 sequence and work synergistically with other modal features to perform classification.

[0024] The root mean square error (RMSE) was used as the core evaluation indicator to quantify the spatial coordinate difference between the registered target image and the reference sequence. An RMSE < 2 mm was set as the acceptable threshold; if this threshold was met, the process proceeded to the next step. If the threshold was exceeded, the image was marked as an abnormal registration, requiring re-optimization of the registration parameters and re-registration. After registration of patient A's T2-FS sequence, the RMSE was 1.3 mm, meeting the acceptable standard. However, after registration of the b=1000s / mm² sequence, the RMSE was 2.6 mm, exceeding the threshold. The system automatically marked this as abnormal, triggering a re-registration process.

[0025] To address the core requirements of image consistency and feature extractability in the diagnosis of isolated spinal lesions, a multimodal image registration optimization model was constructed. A reference sequence adaptation layer was added to the model to precisely analyze and refine the spatial alignment requirements and registration accuracy control standards for different modalities of images. The reference sequence adaptation layer in the registration optimization model defines registration reference standards, alignment priorities, and accuracy control thresholds for structural and functional sequences, respectively, based on their different characteristics. For structural sequences (T1C+, T2, T2-FS, etc.): T1C+ is used as the sole reference, with alignment priority in the order of "lesion boundary > vertebral body structure > surrounding soft tissue," and an accuracy control threshold of RMSE < 1.5 mm (because structural sequences need to accurately reflect the lesion morphology, higher accuracy is required).

[0026] For functional sequences (b=0s / mm², b=800s / mm², b=1000s / mm², etc.): b=0s / mm² is used as the sole reference, with alignment priority being "signal distribution in the lesion area > overall anatomical structure," and the accuracy control threshold being RMSE < 2mm (functional sequences focus on the functional characteristics of the lesion, with a moderate relaxation of overall accuracy requirements). When registering the T2 sequence (structural sequence) for patient A, the T1C+ sequence was referenced, prioritizing spatial alignment of the lesion boundary (e.g., a 1.5cm × 2cm abnormal signal area on the right side of the vertebral body), resulting in a final RMSE of 1.2mm, meeting the accuracy control threshold for structural sequences. When registering the b=800s / mm² sequence (functional sequence), the b=0s / mm² sequence was referenced, prioritizing consistency between the signal distribution in the lesion area and the reference sequence, resulting in an RMSE of 1.8mm, meeting the accuracy requirements for functional sequences.

[0027] Based on the diagnostic characteristics of isolated spinal lesions, the core alignment requirements of different imaging modalities are identified and transformed into executable registration rules. For structural sequences: precise alignment of vertebral body edges, intraspinal structures, and the boundaries between the lesion and surrounding tissues is required, enabling physicians to observe lesion morphology and enhancement characteristics through comparison of multiple structural sequences. For functional sequences: precise alignment of signal intensity distribution in the lesion area is required, supporting the analysis of lesion diffusion characteristics through multiple b-value sequences (e.g., malignant lesions often exhibit high b-value signal enhancement). Regarding the diagnostic requirement of "lesion boundary identification" for isolated spinal lesions, the registration optimization model clarifies that: during structural sequence registration, the spatial deviation of the lesion edge must be ≤1mm; during functional sequence registration, the signal overlap rate of the lesion area must be ≥90%, avoiding misjudgment of lesion diffusion characteristics due to registration deviations.

[0028] A registration strategy-accuracy index mapping model is constructed to transform the spatial alignment requirements of multimodal images into quantifiable registration execution rules. A root mean square error (RMSE) threshold is introduced to standardize and verify the registration results. An anomaly manual review mechanism is used to correct substandard data, generating standardized multimodal image data that combines data consistency with clinical diagnostic suitability. The spatial alignment requirements extracted from the registration optimization model are transformed into quantifiable registration execution rules, clarifying the parameter setting ranges for rigid registration and affine registration.

[0029] For rigid registration, the rotation angle adjustment range is -5° to 5°, and the translation distance adjustment range is -5mm to 5mm. Exceeding these ranges indicates excessive difference in scan position, requiring manual verification of image validity. For affine registration, the scale adjustment coefficient range is 0.9 to 1.1, and the shearing transformation coefficient range is -0.1 to 0.1, avoiding over-adjustment that could distort the anatomical structure of the image. During T2-FS sequence registration for patient B, the rigid registration rotation angle was adjusted to 4° (within the range of -5° to 5°), and the translation distance was adjusted to 3mm (within the range of -5mm to 5mm); the affine registration scale adjustment coefficient was 1.03 (within the range of 0.9 to 1.1), and the shearing transformation coefficient was 0.05 (within the range of -0.1 to 0.1), conforming to the quantitative execution rules.

[0030] Using RMSE < 2mm as a uniform acceptance threshold, all registered images were batch-verified to filter out data that met the accuracy requirements. The system automatically calculates the RMSE value between each registered sequence and the reference sequence. If ≤ 2mm, it is marked as "acceptable" and directly enters the subsequent dataset; if > 2mm, it is marked as "unacceptable" and triggers the exception handling process. The batch processing of multimodal image registration results from 100 patients showed that 92 patients had RMSE < 2mm for all registered sequences and were marked as acceptable; 8 patients had some sequences (e.g., b = 2000 s / mm²) with RMSE > 2mm (range 2.1~3.5mm) and were marked as unacceptable.

[0031] For data marked as "unqualified," a manual review mechanism is initiated, with professional radiologists analyzing the causes of registration abnormalities and making targeted corrections. The causes and correction methods are as follows: Cause 1: Severe motion artifacts exist in the image (e.g., patient movement during scanning), causing the registration algorithm to fail to align accurately. Correction: Remove the sequence and re-execute the registration process using the patient's original image data (artifact-free version). Cause 2: Inappropriate registration parameter settings (e.g., rigid registration rotation angle limitation is too narrow). Correction: Adjust the registration parameters (e.g., expand the rotation angle range to -8° to 8°) and re-execute the registration. Cause 3: Poor quality reference sequence (e.g., low signal-to-noise ratio in T1C+ sequence). Correction: Replace the reference sequence (e.g., use a T2WI sequence as the structural sequence reference) and re-execute the entire registration process. For example, a patient with a b=1000s / mm² sequence had a registration RMSE of 3.2mm. Upon review by a physician, it was found that the sequence contained slight motion artifacts, causing registration deviation. By re-enabling the registration process using the patient's original artifact-free b=2000s / mm² sequence, the RMSE decreased to 1.7mm, meeting the acceptable standard.

[0032] All validated sequences (structural and functional sequences) are integrated and stored according to the rule of "patient ID-modal type-registration status," forming a basic dataset covering all elements of format, spacing, coordinates, and accuracy. All images in the dataset are uniformly formatted as NIfTI, with a voxel spacing of 1mm×1mm×1mm, spatial coordinates consistent with the reference sequence, and RMSE < 2mm, achieving both data consistency and clinical diagnostic suitability. Patient A's standardized dataset contains six sequences: T1C+ (reference sequence), T2 (after registration), T2-FS (after registration), b=0s / mm² (reference sequence), b=800s / mm² (after registration), and b=1000s / mm² (after registration). All sequences meet the unified standards for format, spacing, coordinates, and accuracy, and can be directly used for subsequent segmentation processing in the MedSAM2 model.

[0033] S103 processes standardized multimodal image data based on the MedSAM2 model, supports automatic segmentation with zero samples or segmentation assisted by a small amount of doctor interaction prompts, generates lesion mask files and candidate lesion lists, then cuts out the smallest cubic ROI and simultaneously aligns it with the corresponding multimodal regions to generate the registered multimodal ROI input block.

[0034] In one implementation, model adaptation analysis and segmentation strategy transformation are performed on the lesion localization requirements, boundary recognition accuracy requirements, and multimodal collaborative segmentation conditions in standardized multimodal imaging data. This generates zero-shot automatic segmentation features, interactive auxiliary segmentation trigger parameters, and cross-modal segmentation consistency constraints, forming the basic information for lesion segmentation. The localization characteristics of isolated spinal lesions are clarified (mostly located in the vertebral body, pedicle, or spinal canal). Combined with multimodal imaging features (such as enhancement features in T1C+ sequences and signal differences in DWI), model-identifiable localization clues are extracted and transformed into zero-shot automatic segmentation features, ensuring that the model can identify lesion areas without additional training. For vertebral body lesions, "abnormal enhancement signals within the vertebral body region in T1C+ sequences" and "signal reduction regions in sequences with b=1000s / mm²" are extracted as core zero-shot segmentation features. Based on these features, the model automatically locates a 1.8cm × 2.2cm abnormal region on the left side of the L3 vertebral body of the patient, completing the initial segmentation.

[0035] Define the boundary accuracy requirements for different lesion types (e.g., malignant lesion boundaries need to be accurate to within 1mm, while non-tumor lesions can be relaxed to 2mm), set the trigger conditions for interactive segmentation, and determine the trigger parameters. When the automatic segmentation boundary accuracy does not meet the standard, activate doctor-interactive assistance. Set the boundary accuracy threshold to "the average distance between the automatic segmentation boundary and the actual lesion boundary > 1.5mm", and the trigger parameters to "the doctor selects 2-3 key points of the lesion boundary on the client" or "draws a rectangle containing the lesion". After the automatic segmentation of the patient's T5 thoracic spine lesion, the average boundary deviation was 2.1mm, triggering interactive segmentation. After the doctor selects 3 key boundary points, the model quickly corrects the segmentation boundary.

[0036] Clarify the collaborative relationship between multimodal images in segmentation (e.g., structural sequences provide anatomical boundaries, while functional sequences assist in confirming lesion extent) to avoid segmentation bias in a single modality. Establish cross-modal segmentation consistency constraints to ensure that segmentation results from different modalities match each other. The constraint rules are set as follows: "The overlap rate between the lesion extent segmented by the T1C+ sequence and the segmentation extent by the b=0s / mm² sequence is ≥90%" and "The deviation of the lesion center coordinates in multimodal segmentation is ≤1mm." After segmenting the patient's C4 lesion, the overlap rate between the T1C+ and b=0s / mm² sequences was 92%, and the center coordinate deviation was 0.8mm, satisfying the consistency constraints.

[0037] The format specifications, candidate lesion selection criteria, and boundary accuracy requirements for lesion mask generation are defined by rules and quantified by parameters. This generates NIfTI / SEG format output parameters, lesion selection threshold parameters, and boundary optimization adjustment coefficients, forming mask generation constraint information. The output format of the mask file is clearly defined (NIfTI format is commonly used clinically, and SEG format is adapted to specific systems), the file structure and coordinate system are unified, and the output parameters are determined to ensure compatibility between the mask and the original image. The NIfTI format output parameters are set to "3D matrix dimension 512×512×20 layers, pixel value 0 (background) / 1 (lesion), coordinate system consistent with T1C+ reference sequence"; the SEG format output parameters are "includes lesion location labels, modal association information, and doctor annotation record fields". The model generates NIfTI format mask files for patient lesions, where areas with pixel values ​​of 1 accurately correspond to the lesion range and can be directly read by subsequent classification models.

[0038] Define the screening criteria for effective lesions (e.g., maximum lesion diameter ≥1cm, contrast with surrounding tissue ≥30%), eliminate artifacts or small noise areas, quantify the screening threshold, and ensure the effectiveness of the candidate lesion list. Set the screening threshold parameters as "maximum diameter ≥1cm, contrast ≥30%, volume ≥0.5cm³". After model segmentation, two candidate regions were detected. One with a diameter of 0.8cm (below the threshold) was eliminated, and the other with a diameter of 2.0cm and a contrast of 35% (meeting the threshold) was included in the candidate lesion list.

[0039] For different imaging modalities (e.g., blurred boundaries in T2-FS sequences, clear boundaries in T1C+ sequences, moderate signal contrast in T2 sequences, and boundary discernibility in between), boundary optimization coefficients were set to correct boundary deviations in automatic segmentation, determine adjustment coefficients, and improve the fit between the mask boundary and the actual lesion. Boundary optimization adjustment coefficients were set as follows: 1.0 for T1C+ sequences (clear boundaries, no adjustment needed), 1.1 for T2 sequences (slightly enlarging the boundary to compensate for segmentation deviations caused by insufficient signal contrast), 1.2 for T2-FS sequences (moderately enlarging the boundary to cover the actual lesion), and 0.9 for DWI sequences (reducing the boundary to eliminate artifacts). The automatic segmentation boundary of the patient's lesion in the T2-FS sequence was too narrow; after adjustment with coefficient 1.2, the boundary expanded outward by 0.3 mm, consistent with the lesion boundary confirmed by pathological results. The automatic segmentation boundary in the T2 sequence had a 0.2 mm inward deviation; after adjustment with coefficient 1.1, the boundary expanded outward by 0.2 mm, conforming to the actual lesion area.

[0040] The objective function for ROI extraction and multimodal alignment is used, and region cropping calibration and data synchronization efficiency evaluation are performed based on the effectiveness of segmentation results. This generates minimum cube ROI cropping parameters and a multimodal region synchronization alignment benchmark, forming ROI processing optimization information. The ROI extraction objectives are clearly defined (preserving the complete lesion region, removing invalid background, and minimizing data volume). Based on the lesion size and location from the segmentation results, the cropping range is calculated, and cropping parameters are determined, including the cube's starting coordinates and side dimensions. After patient lesion segmentation, the coordinate range of the lesion region in the image is (X: 120-180 pixels, Y: 200-260 pixels, Z: 8-15 layers). Considering boundary redundancy requirements, the cropping parameters are set to "starting coordinates (115, 195, 7), side length 70 pixels × 70 pixels × 10 layers". After cropping, a minimum cube ROI is formed, which includes the complete lesion while removing 80% of the invalid background.

[0041] The alignment goals for multimodal ROIs are clearly defined (complete consistency in spatial coordinates and synchronization of data dimensions). Based on the reference sequence after prior registration, an alignment benchmark is determined, and a synchronous alignment benchmark for multimodal regions is set to ensure accurate matching of ROIs of different modalities. Using the ROI clipping parameters of the T1C+ sequence as a benchmark, the alignment rule is set as "the starting coordinates and side lengths of other modal ROIs are completely consistent with the T1C+ sequence". The ROIs of the T2-FS, b=0s / mm², and b=1000s / mm² sequences are all clipped according to the starting coordinates (115,195,7) and the layer size of 70×70×10 of T1C+. Among them, the spatial coordinate deviation between the T2 sequence ROI and the T1C+ sequence ROI is ≤0.5mm, ensuring complete spatial alignment of multimodal ROIs. This provides an accurate spatial basis for cross-modal feature fusion of the subsequent dual-domain multi-stream residual convolutional feature encoding network, ensuring complete spatial alignment of multimodal ROIs.

[0042] The process involved verifying whether the cropped ROI contained the complete lesion (effectiveness), evaluating the generation time of the multimodal ROI (synchronization efficiency), optimizing parameters to balance effectiveness and efficiency, and adjusting cropping parameters and alignment methods to ensure both effectiveness and optimal efficiency. For example, in one patient, the cropped ROI did not completely contain the lesion (missing a 0.2cm edge). Adjusting the side dimensions to 75 pixels × 75 pixels × 10 layers and re-cropping resulted in a complete lesion containment. The generation time of the multimodal ROI was evaluated; the original synchronous method took 8 seconds, but after optimizing the alignment benchmark, the time was reduced to 3 seconds, meeting clinical efficiency requirements.

[0043] This paper integrates basic information on lesion segmentation, mask generation constraints, and ROI processing optimization information to design a seamless workflow for the MedSAM2 model's segmentation inference execution, mask file generation, ROI clipping, and multimodal alignment, generating registered multimodal ROI input blocks. With the goals of "accurate segmentation, standardized masks, effective ROIs, and precise alignment," the parameters, rules, and constraints from these three types of information are correlated and matched to form a complete execution plan. "Zero-sample segmentation features" are correlated with "mask format parameters" to ensure that the segmentation results are directly output in the specified format; "boundary optimization coefficients" are correlated with "ROI clipping parameters" to adjust the clipping range based on the optimized mask boundaries.

[0044] The MedSAM2 model first segments the lesion based on basic lesion segmentation information (zero-sample or interactive), then generates a standardized mask file according to mask generation constraints, and finally completes cropping and multimodal alignment based on ROI processing optimization information. After the patient's standardized multimodal image input, the MedSAM2 model first locates and segments the T6 vertebral body lesion based on the zero-sample features of "T1C+ abnormal enhancement and increased DWI signal"; generates a mask file according to "NIfTI format, pixel value 0 / 1, and screening threshold ≥1cm"; then crops the ROI according to "starting coordinates (90,180,5) and side length 65×65×9 layers", and simultaneously aligns the ROI regions of six modalities: T1C+, T2, T2-FS, b=0s / mm², b=800s / mm², and b=1000s / mm².

[0045] The multimodal ROI input block contains cropped ROI data for each modality, featuring a unified format, spatial alignment, complete lesion preservation, and simplified background. It can be directly input into subsequent dual-domain multi-stream residual convolutional feature encoding networks targeting structural and functional phases. The final generated multimodal ROI input block contains ROI data for six modalities. Each modality's ROI is a 65×65×9-layer cube structure with perfectly aligned spatial coordinates, complete preservation of lesion regions, and an invalid background ratio of <20%, meeting the input requirements of classification models.

[0046] S104 is based on a dual-domain multi-stream residual convolutional feature encoding network oriented towards structural and functional phases to process multimodal ROI input blocks, mine complementary information under different imaging mechanisms, and generate four-class classification results and confidence probability values ​​for isolated spinal lesions.

[0047] In one implementation, a dual-domain (Structural–Functional) multi-stream residual convolutional feature encoding network was constructed, aiming to fully exploit complementary information under different imaging mechanisms and improve the classification performance of isolated spinal lesions. In the structural domain, the model input includes three sequences: T2, T2-FS, and enhanced T1C+. In the functional domain, three typical b-values ​​(b=0, b=800, b=1000) of diffusion-weighted images are used. Both domains are feature-encoded through independent multi-stream convolutional branches, where each sequence (or b-value) corresponds to a lightweight residual convolutional sub-network. This design enables the model to learn lesion feature expressions under different structural contrast mechanisms and different diffusion sensitivities.

[0048] Each branch adopts a 3D convolutional structure based on residual connections. The gradient vanishing problem of deep networks is effectively alleviated by cross-layer shortcuts, and local texture, global structure and lesion enhancement patterns are captured step by step to ensure that the expressive power of different modalities in the high-dimensional feature space is more stable and the separability is stronger.

[0049] In the inter-domain feature integration stage, the model introduces an attention-guided dual-domain fusion (AGDF) module. This module adaptively evaluates the contributions of structural and functional features in different anatomical locations and lesion regions by calculating spatial attention and channel attention weights. Channel attention is used to measure the importance of different modalities (T2, T2-FS, T1C+, b=0, b=800, b=1000) at the global semantic level; spatial attention is used to focus on the significant differences of lesion regions under different modalities; the dynamic weighted fusion mechanism can adaptively enhance the discriminative feature channels and suppress redundant or noisy features based on the differences in structural contrast and diffusion performance of lesions.

[0050] Through this attention-driven cross-domain fusion mechanism, the model not only achieves deep integration of structural and functional domain information, but also makes specific response adjustments to the enhancement patterns, tissue components and diffusion behaviors of different cases, ultimately improving the classification accuracy and generalization ability of isolated lesions.

[0051] For the classification objective function of isolated spinal lesions, parameter calibration and classification performance evaluation are performed in conjunction with clinical diagnostic criteria. This generates four-category decision thresholds, confidence calculation benchmarks, and diagnostic consistency matching parameters, forming optimized classification results. The classification objective function is calibrated using clinical diagnostic criteria (such as pathological biopsy results), setting decision thresholds for each category to ensure consistency between classification results and clinical diagnosis. The decision thresholds are determined, specifying at what threshold the network output probability should be classified as the corresponding category. Through calibration with clinical pathological data, decision thresholds are generated: non-neoplastic ≥0.7, benign ≥0.65, locally invasive ≥0.6, and malignant ≥0.75. If the network outputs a malignancy probability of 0.78, exceeding the malignancy decision threshold, it is classified as a malignant lesion. If the locally invasive probability is 0.58, below the threshold, a comprehensive judgment is made based on the probabilities of other categories.

[0052] The basis for calculating the confidence score is clearly defined (based on the probability distribution of the network output layer and the stability of feature extraction). A benchmark for calculating the confidence score is established to ensure that it reflects the reliability of the classification results. This benchmark ensures a positive correlation between the confidence score and the classification accuracy. The confidence score benchmark is set as "maximum probability value of the network output layer × feature extraction stability coefficient (0.8~1.0)". For example, the maximum probability of the network output for a benign lesion is 0.82, and the feature extraction stability coefficient is 0.95. The calculated confidence score is 0.82 × 0.95 = 0.779, reflecting a high level of reliability for this classification result.

[0053] A four-category classification of isolated spinal lesions and their confidence probability values ​​were generated. The classification results clearly identified the lesion as belonging to one of the four categories: "non-neoplastic, benign, locally invasive, or malignant." The confidence probability value reflects the reliability of the classification result (between 0 and 1, with higher reliability being closer to 1). In a test set of 100 patients, 25 were classified as non-neoplastic (mean confidence 0.78), 35 as benign (mean confidence 0.81), 20 as locally invasive (mean confidence 0.73), and 20 as malignant (mean confidence 0.85). All results met the diagnostic consistency matching parameters, with an 87% concordance rate with pathological biopsy results, satisfying clinical diagnostic requirements.

[0054] S105 generates a comprehensive classification probability through a strategy of pre-construction fixed-order splicing, mid-construction attention-weighted feature fusion, or post-construction weighted voting / logistic regression fusion based on modality performance.

[0055] In one implementation, the modal arrangement rules, channel order adaptation, and input stability requirements in the initial fixed-order splicing strategy are adapted to specific scenarios and matched with feature responses to generate acquisition sequence alignment features, modal arrangement consistency guarantee features, and input layer stability enhancement features, forming the core mechanism information for the initial fusion. The modal arrangement must adhere to the principle of "acquisition sequence + clinical priority" (structural images first, then functional images; T1C+ has the highest priority in structural images, and b=1000s / mm² has the highest priority in functional images), ensuring consistency between the arrangement and clinical diagnostic logic. This is transformed into acquisition sequence alignment features, ensuring uniform and interpretable modal arrangements. The arrangement rule is set as "T1C+→T2→T2-FS→b=1000s / mm²→b→b=800s / mm²=0s / mm²", generating "acquisition time sequence + clinical diagnostic priority" alignment features. All patients' multimodal ROI input blocks are spliced ​​in this order, ensuring a consistent modal arrangement in the network input layer and enhancing feature extraction stability.

[0056] Clearly define the channel allocation rules for different modalities (each modality occupies an independent channel group to avoid channel confusion), ensure that the channel positions of the same modality are consistent in different patient data, identify consistency guarantee features, and avoid feature extraction deviations caused by chaotic channel order.

[0057] Clearly define input stability requirements (such as pixel value range normalization and data dimension unification) to avoid the impact of input data fluctuations on network training and inference. Extract stability enhancement features to ensure the consistency and reliability of input layer data. Generate stability features that "normalize pixel values ​​to the 0-1 range + unify the 3D matrix dimension to 64×64×64". For a patient's T1C+ modality ROI, the original pixel value range is 0-255. After normalization, it is compressed to the 0-1 range, and the data dimension is adjusted to 64×64×64 to maintain consistency with other modalities, thereby improving input stability.

[0058] The effectiveness of fusion layer selection, modal weight allocation, and feature interaction logic in mid-term attention-weighted feature fusion was verified, and semantic fusion efficiency was evaluated. Feature layer adaptation parameters, dynamic weight adjustment coefficients, and shared semantic extraction scheme features were generated to form the basic information for mid-term fusion. The effects of different fusion layers were verified using a test set (after the third feature encoding module of the dual-domain multi-stream residual convolutional feature encoding network for structural and functional phases, and before the cross-domain fusion layer). A fusion layer that takes into account both low-level spatial features and high-level semantic features was selected, feature layer adaptation parameters were determined, and the specific layer for fusion execution was clarified. It was verified that fusion after the third feature encoding module of the dual-domain multi-stream residual convolutional feature encoding network for structural and functional phases can preserve low-level spatial features such as lesion edges and fuse high-level semantic features between modalities. The adaptation parameter generated is "fusion layer = third dense block output layer". At this layer, the network concatenates and fuses features from each modality to improve the feature interaction effect.

[0059] Based on the contribution of each modality to the classification task (e.g., T1C+ contributes significantly to lesion morphology assessment, and b=1000s / mm² contributes significantly to benign / malignant differentiation), a dynamic weight adjustment rule is set to enable the model to adaptively adjust modal weights and determine the dynamic weight adjustment coefficients, highlighting the feature weights of high-contribution modalities. The adjustment rule is set as "T1C+ weight coefficient 0.3, T2 weight coefficient 0.15, T2-FS weight coefficient 0.15, b=0s / mm² weight coefficient 0.2, b=1000s / mm² weight coefficient 0.2", generating "dynamic adjustment based on clinical contribution" coefficients. For patients with malignant lesions, the network automatically increases the weight coefficient of b=1000s / mm² to 0.3 through an attention mechanism to strengthen its diffusion characteristics.

[0060] The feature interaction must adhere to the principle of "redundancy removal + strong correlation" (retaining shared lesion features between modalities and eliminating duplicate and invalid information) to ensure that the features after interaction are more targeted, refine the features of the shared semantic extraction scheme, and improve the efficiency of semantic fusion. An extraction scheme of "shared lesion region signal features between modalities + weighted fusion of differentiated functional features" is generated. Through this scheme, the network extracts the "lesion boundary signal features" shared by T1C+, T2, and T2-FS, while simultaneously fusing the unique "diffusion-restricted feature" of b=1000s / mm², reducing redundant information and strengthening classification-related semantic features.

[0061] By combining performance weight allocation, result coordination rules, and probabilistic integration processes in the later-stage weighted voting / logistic regression fusion, strategic coordination planning and redundant information resolution are performed on the core mechanism information of the early-stage fusion and the basic information of the mid-stage fusion. This generates multi-strategy linkage execution coefficients and fusion result calibration parameters, forming fusion optimization information. Fusion weights are assigned based on the performance metrics (AUC, accuracy) of each modal independent model on the validation set, with higher weights for better performance. Simultaneously, the strategic characteristics of the early-stage and mid-stage fusion are considered, incorporating the performance contribution of the T2 single-modal model to determine the multi-strategy linkage execution coefficients, ensuring that the later-stage fusion works synergistically with the early-stage and mid-stage fusions.

[0062] The results showed that the AUC of the T1C+ single-modal model was 0.85, the T2 single-modal model was 0.82, the T2-FS single-modal model was 0.78, the multimodal early-stage fusion model was 0.88, and the multimodal mid-stage fusion model was 0.90. The generated weight allocation coefficients were "T1C+ (0.12), T2 (0.1), T2-FS (0.08), early-stage fusion (0.35), mid-stage fusion (0.35)", and the linkage execution coefficient was "early-stage fusion result × 0.35 + mid-stage fusion result × 0.35 + single-modal result (T1C+ × 0.12 + T2 × 0.1 + T2-FS × 0.08) × 0.3". Through this coefficient, the later-stage fusion not only integrated the structural feature advantages of the T2 modality but also retained the core value of each fusion strategy, further improving the reliability and generalization ability of the results.

[0063] The synergistic results must adhere to the principle of "clinical diagnostic priority + probability threshold screening" (the probability of malignant lesions must be strictly verified, while the threshold for non-tumor lesions can be appropriately relaxed) to avoid bias from a single strategy. Calibration parameters should be determined to correct and optimize the fusion results. The generated calibration parameters are: "Malignant lesion probability ≥ 0.75 is retained; 0.6-0.75 is further verified by combining the attention weight of mid-stage fusion; non-tumor lesion probability ≥ 0.65 is retained." For example, in a patient, the initial malignant probability of late-stage fusion was 0.72. After mid-stage fusion attention weight verification (b = 1000s / mm², weight coefficient 0.32, higher than the mean), the calibrated malignant probability was 0.76, meeting the judgment threshold.

[0064] This system integrates core fusion mechanism information from the early stage, basic fusion information from the middle stage, and fusion optimization information to generate a comprehensive classification probability that balances feature extraction stability, semantic fusion effectiveness, and result accuracy. With the goal of "clinical diagnostic logic + model performance optimization," it organically connects the early input layer fusion, the middle feature layer fusion, and the late result layer fusion, forming a complete "input-feature-result" fusion system. It links the early "modal arrangement rules" with the middle "dynamic weight adjustment coefficients" to ensure that high-priority modalities receive higher weights at the feature layer; and it links the middle "shared semantic features" with the late "performance weight allocation" to give higher fusion weights to strategies with better semantic fusion results.

[0065] First, modal splicing is completed according to the early fusion rules. After inputting into the network, attention-weighted feature fusion is performed at the mid-term fusion layer. Finally, the single-modal model results, early fusion results, and mid-term fusion results are integrated according to the late fusion weights and calibration parameters to generate a comprehensive classification probability. The patient's multimodal ROI input block is spliced ​​according to "T1C+→T2→T2-FS→b=0s / mm²→b=800s / mm²→b=1000s / mm²" (early fusion). After the third dense block, the network performs attention-weighted fusion of each modality feature (T1C+ weight 0.35, b=1000s / mm² weight 0.3) (mid-term fusion). Finally, the single-modal, early, and mid-term fusion results are integrated, and the comprehensive classification probability is calculated according to the weight coefficients: malignant 0.82, locally invasive 0.12, benign 0.05, non-tumor 0.01. After calibration, the output comprehensive classification probability is malignant 0.82.

[0066] The comprehensive classification probability clearly defines the final probability distribution of the four types of lesions. The category with the highest probability is the final classification result. The probability value reflects the reliability of the result and must meet the requirements of clinical diagnostic consistency. In 100 patients in the test set, the comprehensive classification probability and the pathological biopsy results were in agreement for 90%. Among them, the sensitivity for diagnosing malignant lesions was 92%, the specificity was 89%, and the accuracy for diagnosing non-neoplastic lesions was 88%. One patient's comprehensive classification probability was 0.78 for benign, 0.15 for locally invasive, 0.05 for malignant, and 0.02 for non-neoplastic, and the final classification was benign with a confidence level of 0.78, consistent with the physician's diagnosis.

[0067] S106 processes segmentation indicators, classification indicators, and classification results based on client-side visualization, supports interactive fine-tuning by doctors, and generates exportable DICOMSEG / NIfTI format mask information.

[0068] In one implementation, core data from the segmentation and classification processes are extracted, including segmentation indicators, classification indicators, four-class classification results, and comprehensive classification probabilities. Core indicators reflecting segmentation accuracy are selected to quantify the accuracy and completeness of lesion segmentation, ensuring reliable segmentation results. These core indicators are the Dice similarity coefficient (measuring the overlap between the segmented region and the actual lesion) and HD95 (95% Hausdorff distance, measuring the deviation of the segmentation boundary). After segmentation using MedSAM2, the extracted Dice coefficient for a patient's lumbar spine lesion was 0.89 (high overlap) and the HD95 was 1.2 mm (small boundary deviation), indicating that the segmentation accuracy met clinical requirements. However, after segmentation of another patient's lesion, the Dice coefficient was 0.72 and the HD95 was 2.3 mm, suggesting a certain deviation in the segmentation boundary, requiring further monitoring.

[0069] Core indicators reflecting classification performance were selected to quantify the reliability of the model's classification and provide data support for doctors' reference. These core indicators are AUC (Area Under the Curve), accuracy, sensitivity (recall), and specificity. For a batch of 100 patients, the classification results showed AUC = 0.91, accuracy = 88%, sensitivity for malignant lesions = 92%, and specificity = 89%, indicating stable model classification performance. However, the classification results for 12 patients were inconsistent with the pathological diagnosis, requiring further review by doctors.

[0070] The final classification result (one of four categories) and corresponding comprehensive classification probability of the model output were extracted to clarify the diagnostic conclusion and reliability. The extracted content included the probability distribution of non-tumor, benign, locally invasive, and malignant lesions, as well as the final classification category. The comprehensive classification probability of the thoracic spine lesion in the patient was "non-tumor 0.03, benign 0.12, locally invasive 0.15, malignant 0.70", and the final classification result was "malignant" with a confidence level of 0.70. The probability distribution of another patient was "non-tumor 0.68, benign 0.21, locally invasive 0.08, malignant 0.03", and the final classification result was "non-tumor" with a confidence level of 0.68.

[0071] Based on the client-side visualization function, core data is presented in an intuitive form, with an interactive fine-tuning interface that allows doctors to manually correct segmentation boundaries and classification results. The combination of "image + chart + numerical values" enables doctors to quickly grasp segmentation accuracy, classification performance, and diagnostic results. It overlays the original multimodal images, automatically segmented boundaries (marked with different colors), and classification result labels (e.g., "malignant" labeled next to the lesion). Bar charts display segmentation indicators such as Dice coefficient and HD95, while numerical lists display classification probability distributions and classification indicators such as AUC and accuracy. The left side of the client interface displays the patient's T1C+ sequence image, with red outlines marking the automatically segmented lesion boundaries, and the upper right corner labeled "malignant (confidence 0.70)". The right-side bar chart shows a Dice coefficient of 0.89 and HD95 of 1.2mm, while the list below displays the four probability distributions, along with the classification indicator AUC of 0.91, allowing doctors to clearly grasp the core information.

[0072] It provides convenient manual adjustment tools, allowing doctors to correct boundary deviations in automatic segmentation, ensuring the integrity and accuracy of the lesion area. It supports operations such as dragging boundary points with the mouse, selecting adjustment areas, and zooming in for refinement, with real-time preview of the adjusted segmentation effect. For example, if a doctor finds that the automatically segmented boundary of a patient's lesion does not completely encompass the lesion edge (HD95=2.3mm), they can use the client-side fine-tuning tool to drag the boundary point outward by 0.3mm. Real-time monitoring shows that the adjusted Dice coefficient has increased to 0.81, and HD95 has decreased to 1.5mm, meeting clinical accuracy requirements.

[0073] The system supports physicians in revising the model's classification results based on clinical experience. It also allows adjustment of classification probability thresholds, reclassification, and provides a dropdown menu for selecting the target category (non-tumor / benign / locally invasive / malignant). Manual input of the adjusted confidence level is also supported, and the system synchronously updates the comprehensive classification probability distribution. For example, if the model classifies a patient's lesion as "locally invasive (confidence 0.62)," the physician, considering the patient's clinical history (history of malignancy) and imaging characteristics (significant enhancement on T1C+, increased high b-value signal on DWI), can correct the classification to "malignant" through a fine-tuning interface, adjusting the confidence level to 0.78. The system automatically updates the four-category probability distribution to ensure consistency with clinical judgment.

[0074] Based on the corrected results, a DICOMSEG / NIfTI format mask file conforming to clinical standards is automatically generated. The corrected segmentation boundaries (spatial coordinates, range) are associated with the classification result labels, adapting to the standardized requirements of both DICOMSEG and NIfTI, two commonly used clinical formats. The DICOMSEG format must include basic patient information, image association information, segmentation region labels (e.g., "malignant lesion"), and classification result remarks; the NIfTI format must maintain consistency with the coordinate system of the original image, with pixel values ​​of 0 representing background and 1 representing the corrected lesion region. The corrected malignant lesion segmentation boundaries (coordinate range X: 110-180 pixels, Y: 190-260 pixels, Z: 6-14 layers) are integrated with the "malignant" classification label, adapting to the field requirements of the DICOMSEG format, and associating patient ID, image examination number, and other information; simultaneously, an NIfTI format file is generated, ensuring that pixel values ​​and coordinate systems are consistent with the original T1C+ sequence.

[0075] The system automatically encapsulates and integrates information according to clinical standards, generating mask files that can be directly used for clinical diagnosis, research archiving, or subsequent treatment planning. It also generates DICMSEG format (compatible with hospital PACS systems) and NIfTI format (compatible with research analysis software), allowing doctors to export as needed. After doctors complete fine-tuning, the system automatically generates a DICMSEG format mask file of the patient's lesions, including the segmented "malignant lesion" region and classification notes, which can be directly imported into the hospital PACS system and archived along with the original images. Simultaneously, it generates an NIfTI format file for researchers to use in subsequent model optimization analysis. Both file types conform to clinical data exchange standards and can be used without additional format conversion.

[0076] In one implementation, such as Figure 2 As shown, this application also provides an automated segmentation and four-class classification device for images of isolated spinal lesions, comprising: The acquisition module 201 is used to acquire multimodal image data and clinical feature information; Processing module 202 is used to preprocess multimodal image data, uniformly converting it to NIfTI format and standardizing voxel spacing. Spatial registration is completed using T1C+ as the structural image reference and b=0s / mm² sequence as the diffusion image reference. The registration accuracy is evaluated by root mean square error, generating standardized multimodal image data. Based on the MedSAM2 model, the standardized multimodal image data is processed, supporting zero-shot automatic segmentation or doctor-assisted segmentation with minimal interactive prompts. A lesion mask file and a candidate lesion list are generated. The smallest cubic ROI is then cropped and simultaneously aligned with the corresponding multimodal regions, generating the registered multimodal ROI. The input block is processed using a dual-domain, multi-stream residual convolutional feature encoding network oriented towards structural and functional phases. This process mines complementary information under different imaging mechanisms, generating four-class classification results and confidence probability values ​​for isolated spinal lesions. A comprehensive classification probability is generated through a strategy of fixed-order splicing in the early stage, attention-weighted feature fusion in the middle stage, or weighted voting / logistic regression fusion based on modal performance in the later stage. The segmentation indicators, classification indicators, and classification results are processed using client-side visualization, supporting interactive fine-tuning by doctors and generating exportable DICOMSEG / NIfTI format mask information.

[0077] The computer-readable storage medium provided in the above embodiments of this application and the automated segmentation and four-class classification method for images of isolated spinal lesions provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application stored therein.

[0078] All embodiments in this application are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of the automated segmentation and four-class classification method, electronic device, electronic device, and readable storage medium for evaluating images of isolated spinal lesions are basically similar to the embodiments of the automated segmentation and four-class classification method for images of isolated spinal lesions described above, and are therefore described simply. Relevant parts can be referred to in the description of the embodiments of the automated segmentation and four-class classification method for images of isolated spinal lesions described above.

Claims

1. An automated segmentation and four-class classification method for images of isolated spinal lesions, characterized in that, include: Acquire multimodal imaging data and clinical feature information; Multimodal image data is preprocessed, uniformly converted to NIfTI format and voxel spacing is standardized. Spatial registration is completed using T1C+ as structural image reference and b=0s / mm² sequence as diffusion image reference. Registration accuracy is evaluated by root mean square error, and standardized multimodal image data is generated. Based on the MedSAM2 model, standardized multimodal image data is processed, supporting zero-sample automatic segmentation or segmentation assisted by a small amount of doctor interaction prompts. A lesion mask file and a candidate lesion list are generated, and then the smallest cubic ROI is cropped and simultaneously aligned with the corresponding multimodal regions to generate the registered multimodal ROI input block. The dual-domain multi-stream residual convolutional feature encoding network oriented towards structural and functional phases is used to process multimodal ROI input blocks, mine complementary information under different imaging mechanisms, and generate four-class classification results and confidence probability values ​​for isolated spinal lesions. A comprehensive classification probability is generated by using a fixed-order splicing strategy in the early stage, attention-weighted feature fusion in the middle stage, or weighted voting / logistic regression fusion strategy based on modality performance in the later stage. Based on the client-side visualization display function, the segmentation indicators, classification indicators and classification results are processed, supporting interactive fine-tuning by doctors, and generating exportable DICOMSEG / NIfTI format mask information.

2. The method as described in claim 1, characterized in that, Multimodal image data is preprocessed, uniformly converted to NIfTI format, and voxel spacing is standardized. Spatial registration is completed using T1C+ as the structural image reference and b=0s / mm² sequence as the diffusion image reference. Registration accuracy is evaluated using root mean square error, generating standardized multimodal image data, including: To address the heterogeneous nature of multimodal image data, which includes structural and functional sequences, a full-process preprocessing mechanism is established, encompassing format conversion, NIfTI standardization, voxel spacing unification, spatial registration, and accuracy assessment. Then, through a hierarchical registration strategy using T1C+ as a reference for structural images and b=0s / mm² sequence as a reference for diffuse images, and the step-by-step execution of rigid and affine registration, a multimodal image base dataset covering all elements of format, spacing, coordinates, and accuracy is formed. To meet the core requirements of image consistency and feature extractability in the diagnosis of isolated spinal lesions, a multimodal image registration optimization model was constructed to accurately sort out and refine the spatial alignment requirements and registration accuracy control standards for different modal images. The spatial alignment requirements of multimodal images are transformed into quantifiable registration execution rules. Combined with a manual review mechanism for abnormal situations, substandard data is corrected to generate standardized multimodal image data that combines data consistency and clinical diagnostic suitability.

3. The method as described in claim 1, characterized in that, Based on the MedSAM2 model, standardized multimodal image data is processed, supporting automatic segmentation with zero samples or segmentation assisted by minimal doctor interaction. Lesion mask files and candidate lesion lists are generated, and then the smallest cubic Region of Interest (ROI) is cropped and simultaneously aligned with the corresponding multimodal regions, generating registered multimodal ROI input blocks, including: Model adaptation analysis and segmentation strategy transformation are performed on the lesion localization requirements, boundary recognition accuracy requirements, and multimodal collaborative segmentation conditions in standardized multimodal image data to generate zero-sample automatic segmentation features, interactive auxiliary segmentation trigger parameters, and cross-modal segmentation consistency constraints, thus forming basic information for lesion segmentation. The format specifications, candidate lesion screening criteria, and boundary accuracy requirements in the lesion mask generation requirements are defined by rules and quantified by parameters. This generates NIfTI / SEG format output parameters, lesion screening threshold parameters, and boundary optimization adjustment coefficients, forming mask generation constraint information. The objective function for ROI extraction and multimodal alignment is combined with the effectiveness of segmentation results to perform region clipping calibration and data synchronization efficiency evaluation, generate minimum cube ROI clipping parameters and multimodal region synchronization alignment benchmarks, and form ROI processing optimization information; By integrating basic information on lesion segmentation, constraint information on mask generation, and optimization information on ROI processing, a seamless design is implemented for the entire process of segmentation inference execution, mask file generation, ROI clipping, and multimodal alignment of the MedSAM2 model, generating registered multimodal ROI input blocks.

4. The method as described in claim 1, characterized in that, A dual-domain, multi-stream residual convolutional feature encoding network oriented towards structural and functional phases is used to process multimodal ROI input blocks, mining complementary information under different imaging mechanisms to generate four-class classification results and confidence probability values ​​for isolated spinal lesions, including: The network architecture of the dual-domain multi-stream residual convolutional feature coding network is analyzed and parameterized to meet the domain structure design requirements, branch feature coding requirements, and cross-domain fusion adaptation conditions. This generates structural domain / functional domain input configuration features, lightweight residual sub-network parameters, and cross-domain fusion adaptation constraints, forming the basic information of the dual-domain network. The selection of 3D convolutional structure, residual connection method and feature capture level requirements in the model feature extraction capability improvement needs are logically transformed and parameter boundaries are defined to generate residual connection 3D convolution configuration parameters, cross-layer shortcut execution parameters and multi-dimensional feature capture coefficients, forming feature extraction optimization information. For the inter-domain fusion objective function, the fusion module is calibrated and the fusion performance is evaluated by combining the modal complementarity standard. Attention weight calculation benchmark, cross-domain fusion execution threshold, and joint characterization optimization parameters are generated to form cross-domain fusion optimization information. By integrating basic information of dual-domain networks, optimization information of feature extraction, and optimization information of cross-domain fusion, a full-process collaborative design is carried out for the intra-domain feature encoding execution, cross-domain fusion strategy application, joint representation generation, and classification output of dual-domain multi-stream residual convolutional feature coding networks, generating four-class classification results and confidence probability values ​​for isolated spinal lesions.

5. The method as described in claim 1, characterized in that, A comprehensive classification probability is generated through a combination of strategies, including: initial fixed-order concatenation, mid-term attention-weighted feature fusion, or late-term weighted voting / logistic regression fusion based on modality performance. The modal arrangement rules, channel order adaptation, and input stability requirements in the early fixed-order splicing strategy are adapted to the scenario and matched with the feature response to generate collection sequence alignment features, modal arrangement consistency guarantee features, and input layer stability enhancement features, forming the core mechanism information of the early fusion. The effectiveness of the fusion layer selection, modality weight allocation, and feature interaction logic in the mid-term attention-weighted feature fusion is verified, and the semantic fusion efficiency is evaluated. Feature layer adaptation parameters, dynamic weight adjustment coefficients, and shared semantic extraction scheme features are generated to form the basic information for mid-term fusion. By combining the performance weight allocation, result coordination rules, and probability integration process in the later weighted voting / logistic regression fusion, strategic coordination planning and redundant information resolution are carried out on the core mechanism information of the early fusion and the basic information of the mid-term fusion, generating multi-strategy linkage execution coefficients and fusion result calibration parameters to form fusion optimization information. By integrating the core mechanism information from the early stage of fusion, the basic information from the mid-stage of fusion, and the optimization information from the fusion, a comprehensive classification probability is generated that takes into account the stability of feature extraction, the effectiveness of semantic fusion, and the accuracy of the results.

6. The method as described in claim 5, characterized in that, The system utilizes client-side visualization capabilities to process segmentation metrics, classification metrics, and classification results. It supports interactive fine-tuning by physicians and generates exportable DICOMSEG / NIfTI format mask information, including: Extract the core data from the segmentation and classification processes, including segmentation indicators, classification indicators, four-class classification results, and comprehensive classification probabilities; Based on the client-side visualization display function, the core data is presented in an intuitive form, and an interactive fine-tuning interface is opened at the same time to support doctors to manually correct the segmentation boundary and classification results; Based on the corrected results, a DICOMSEG / NIfTI format mask file conforming to clinical standards is automatically generated.

7. An automated segmentation and four-class classification device for images of isolated spinal lesions, characterized in that, The device is configured to perform the automated segmentation and four-class classification method for images of isolated spinal lesions as described in any one of claims 1 to 6 by executing the executable instructions.

8. An electronic device, characterized in that, include: First processor; and memory for storing executable instructions of the first processor; The first processor is configured to execute the automated segmentation and four-class classification method for images of isolated spinal lesions according to any one of claims 1 to 6 by executing the executable instructions.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the second processor, it implements the automated segmentation and four-class classification method for images of isolated spinal lesions as described in any one of claims 1 to 6.