Preoperative early warning method for sacrum tumor and related equipment

By extracting and weighted fusion multiple data of sacral tumor patients, multi-channel feature data are generated to assist in diagnosis, and the problem of high misdiagnosis rate of early diagnosis of sacral tumors in the prior art is solved, and diagnostic accuracy and safety are improved.

CN120199485AActive Publication Date: 2025-06-24PEOPLES HOSPITAL PEKING UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510270658.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-24
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

The prior art has a high misdiagnosis rate in the early diagnosis of sacral tumors, resulting in unnecessary biopsy, patient discomfort, increased costs, delayed diagnosis or death.

Method used

By collecting clinical, NCCT images, multimodal and diagnostic interval data of target users, using feature processing technology to extract features from multimodal data to obtain dynamic risk parameters, extract features from clinical data to obtain physiological attribute parameters, and generate multi-channel feature data through weighted fusion strategy, and finally input the model to generate early warning information to assist in determining the benign and malignant tumors.

Benefits of technology

It improves the accuracy of preoperative diagnosis of sacral tumors, reduces the rate of misdiagnosis, and reduces unnecessary medical costs and patient risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199485A_ABST
    Figure CN120199485A_ABST
Patent Text Reader

Abstract

The invention provides a preoperative early warning method for sacrum tumor and related equipment, and is applied to the technical field of data processing. The method comprises the following steps: processing a preset three-dimensional dense connection network model based on a target training sample set and a verification sample set to generate a target three-dimensional dense connection network model; processing the pre-operative sacrum non-contrast computed tomography image of the target user to generate a three-dimensional region-of-interest image of the target user; processing the clinical data information of the target user to generate physiological attribute parameters of the target user; and based on the target three-dimensional dense connection network model, processing the pre-operative sacrum non-contrast computed tomography image of the target user, the three-dimensional region-of-interest image of the target user, the dynamic risk parameters of the target user and the physiological attribute parameters of the target user, and generating sacrum early warning information, the sacrum early warning information is used for representing the attribute information of the pre-operative sacrum lesion of the target user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a preoperative warning method for sacral tumors and related devices. Background Art

[0002] Bone tumors seriously threaten life and health and are the third leading cause of death among cancer patients under 20 years old. It encompasses four types of tumors: primary malignant, benign, metastatic, and those caused by local invasion of visceral malignant tumors, with different biological behaviors. Benign tumors are usually stable, while sacral tumors affect the stability of the lumbosacral joint. Accurately identifying bone tumors is of great significance for clinical decision-making. However, the early symptoms of sacral tumors lack specificity and mostly manifest as neurological deficits and low back pain.

[0003] Imaging is crucial in the early warning of bone tumors. Among them, CT has high resolution, can detect lesions as small as 3 millimeters, and can provide key information such as tumor matrix mineralization to assist in diagnosis. However, the incidence of bone tumors is low, and doctors with insufficient experience are prone to misdiagnosis, leading to problems such as unnecessary biopsies, patient discomfort, increased costs, diagnostic delays, or death. In this context, developing sensitive and effective computer-aided early warning tools has become the key.

[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The purpose of the present application is to provide a preoperative warning method for sacral tumors and related devices, which at least to some extent overcomes the problems existing in the prior art. By collecting various types of data of the target user, including clinical, NCCT images, multi-modal, and diagnostic interval data, etc., in terms of feature processing, features are extracted from multi-modal data to obtain dynamic risk parameters, and features are extracted from clinical data to obtain physiological attribute parameters. After normalizing and standardizing these data and image data, a weighted fusion strategy (assigning learnable weights to the image, dynamic risk parameter, and physiological attribute parameter channels) is used to generate multi-channel feature data. Finally, the multi-channel feature data is input into the model, and warning information is generated based on the preset dynamic threshold, effectively assisting in determining the benign and malignant nature of the tumor and improving the accuracy of preoperative diagnosis.

[0006] Other features and advantages of the present application will become apparent through the following detailed description, or be learned in part through the practice of the present invention.

[0007] According to one aspect of the present application, a preoperative warning method for sacral tumors is provided, including: obtaining the clinical data information of the target user, the preoperative non-contrast computed tomography image of the sacrum of the target user, the multimodal data information of the target user, the target diagnostic interval data information of the target user, a preset three-dimensional densely connected network model, a training sample set, and a validation sample set, wherein the multimodal data information of the target user includes the T1-weighted imaging information of the target user, the T2-weighted imaging information of the target user, and positron emission tomography-computed tomography information; preprocessing the training sample set to generate a target training sample set; processing the preset three-dimensional densely connected network model based on the target training sample set and the validation sample set to generate a target three-dimensional densely connected network model; processing the preoperative non-contrast computed tomography image of the sacrum of the target user to generate a three-dimensional region of interest image of the target user; processing the multimodal data information of the target user and the target diagnostic interval data information of the target user to generate a dynamic risk parameter of the target user; processing the clinical data information of the target user to generate a physiological attribute parameter of the target user; processing the preoperative non-contrast computed tomography image of the sacrum of the target user, the three-dimensional region of interest image of the target user, the dynamic risk parameter of the target user, and the physiological attribute parameter of the target user based on the target three-dimensional densely connected network model to generate a sacral warning information, wherein the sacral warning information is used to characterize the attribute information of the preoperative sacral lesion of the target user.

[0008] Another aspect of the present application is a preoperative warning device for sacral tumors, characterized by including: an acquisition module, configured to acquire the clinical data information of a target user, the preoperative non-contrast computed tomography (NCCT) image of the sacrum of the target user, the multimodal data information of the target user, the target diagnostic interval data information of the target user, a preset three-dimensional densely connected network model, a training sample set, and a validation sample set, wherein the multimodal data information of the target user includes the T1-weighted imaging information of the target user, the T2-weighted imaging information of the target user, and positron emission tomography-computed tomography (PET-CT) information; a processing module, configured to preprocess the training sample set to generate a target training sample set; process the preset three-dimensional densely connected network model based on the target training sample set and the validation sample set to generate a target three-dimensional densely connected network model; process the preoperative non-contrast computed tomography image of the sacrum of the target user to generate a three-dimensional region of interest (ROI) image of the target user; process the multimodal data information of the target user and the target diagnostic interval data information of the target user to generate a dynamic risk parameter of the target user; process the clinical data information of the target user to generate a physiological attribute parameter of the target user; process the preoperative non-contrast computed tomography image of the sacrum of the target user, the three-dimensional ROI image of the target user, the dynamic risk parameter of the target user, and the physiological attribute parameter of the target user based on the target three-dimensional densely connected network model to generate a sacral warning information, wherein the sacral warning information is used to characterize the attribute information of the preoperative sacral lesion of the target user.

[0009] According to yet another aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a second processor, the above-mentioned preoperative warning method for sacral tumors is implemented.

[0010] For the preoperative warning method and related devices provided by the present application, the server collects various types of data of the target user, including clinical, NCCT images, multimodal, and diagnostic interval data, etc., and carefully preprocesses the training samples to ensure data quality. When constructing the model, the three-dimensional densely connected network is optimized by means of the target training and validation samples, and the final model is determined through sampling feature training and verification. At the same time, the NCCT image is processed according to the anatomical structure constraint information to generate a region of interest image. In terms of feature processing, features are extracted from multimodal data to obtain dynamic risk parameters, and features are extracted from clinical data to obtain physiological attribute parameters. After normalizing and standardizing these data and image data, a weighted fusion strategy is used to generate multi-channel feature data. Finally, the multi-channel feature data is input into the model, and warning information is generated according to a preset dynamic threshold, effectively assisting in determining the benign and malignant nature of the tumor and improving the accuracy of preoperative diagnosis.

[0011] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 The flowchart showing a preoperative warning method for a sacral tumor provided by an embodiment of the present application; Figure 2 The structural schematic diagram showing a preoperative warning device for a sacral tumor provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.

[0014] The following combines Figure 1 to describe a preoperative warning method for a sacral tumor according to an exemplary embodiment of the present application. In one embodiment, the present application also proposes a preoperative warning method for a sacral tumor and related devices. Figure 1 Schematically shown is a flowchart of a preoperative warning method for a sacral tumor according to an embodiment of the present application. As Figure 1 shown, this method is applied to a server and includes: S101, obtaining the clinical data information of the target user, the preoperative non-contrast computed tomography image of the sacrum of the target user, the multi-modal data information of the target user, the target diagnosis interval data information of the target user, a preset three-dimensional densely connected network model, a training sample set, and a validation sample set.

[0015] In one embodiment, relevant clinical data of the target user is extracted from a hospital information system (HIS), an electronic medical record (EMR), or a clinical database. It includes basic demographic information such as name, age, gender, etc.; disease history information such as previous illness conditions, surgical history, allergy history, etc.; family medical history, especially family genetic disease information related to tumors; and current symptom manifestations such as pain degree, pain location, and whether there is neurological dysfunction. Arrange for the target user to perform a preoperative non-contrast computed tomography (NCCT) of the sacrum on a qualified imaging examination device. During the scanning process, strictly operate according to the device operation procedures to ensure the acquisition of high-quality image data, including appropriate scanning angles, slice thicknesses, resolutions, etc. After the scanning is completed, store the image data in a standardized medical imaging format (such as DICOM format) and securely transmit it to a designated data analysis server or workstation to ensure that there is no damage or loss during the data transmission process.

[0016] Use a magnetic resonance imaging (MRI) device to obtain T1-weighted imaging and T2-weighted imaging data of the target user. Ensure that the settings of the MRI device, such as the magnetic field strength and scanning sequence parameters, are correct to obtain clear and accurate images. During the imaging process, guide the user to maintain the correct body position to reduce the impact of motion artifacts on the image quality. After obtaining the images, store and transmit them to the analysis system in a standard format. Conduct examinations on a device equipped with positron emission tomography (PET) and computed tomography (CT) functions to obtain PET-CT information. The acquired image data is stored and transmitted after passing quality inspections. Sort out the target diagnosis interval data from the user's previous medical records. This includes the time of first detection of suspected sacral lesions (such as the time when the user first reported symptoms, the time when relevant imaging examinations first detected abnormalities, etc.), the time of first diagnosis, the time of each follow-up visit or reexamination, and the relevant quantitative data of the lesions during each examination (such as the size and morphological changes of tumors, which can be obtained through imaging measurements; the changes in the metabolic activity of the lesions, such as the changes in SUV values in PET-CT, etc.). For time data, ensure the accuracy of the records, accurate to the specific date.

[0017] Select the 3D-DenseNet121 model as the preset three-dimensional densely connected network model. This model should have good capabilities for processing three-dimensional medical image data. The densely connected layers in its network structure can effectively extract the feature information of the images and have performance advantages in related medical image analysis tasks. The 3D-DenseNet121 model as a whole consists of multiple layers, mainly including convolutional layers (Convolutional Layers), densely connected layers (Dense Layers), transition layers (Transition Layers), and fully connected layers (Fully Connected Layers). The starting part of the model contains a series of convolutional layers used to extract low-level features of the images, such as edges and textures. These convolutional layers use three-dimensional convolutional kernels to perform convolutional operations on the images in three-dimensional space. For example, the convolutional kernel size is 3x3x3, the stride is 1, and the padding method is'same' to ensure that the size of the output feature map is the same as that of the input image (in the spatial dimension). Through the stacking of multiple convolutional layers, more complex and abstract features are gradually extracted. The densely connected layer is one of the core structures of 3D-DenseNet121. In the densely connected layer, the input of each layer comes from the concatenation of the output feature maps of all previous layers. This connection method enables information to flow directly in the network, avoiding the problem of information loss caused by the deepening of the hierarchy in traditional networks, facilitating the backpropagation of gradients, and thus making the model easier to train. There are L densely connected layers in the model. The number of input feature maps of the th layer ( = 1, 2,..., L) is (where is the number of feature maps output by the layer), and the number of output feature maps is

[0018] Each dense connection layer contains multiple convolution operations for further feature extraction and transformation of the concatenated feature maps.

[0019] To control the complexity and number of parameters of the model, transition layers are inserted between the dense connection layers. The transition layer mainly contains a 1x1x1 convolution operation to reduce the number of feature maps (playing a role in dimensionality reduction), and a 2x2x2 average pooling operation with a stride of 2 to reduce the size of the feature maps, reduce the computational amount, while maintaining a certain feature representation ability to prevent overfitting. At the end of the model is a fully connected layer for integrating and mapping the features extracted previously, and finally outputting the prediction result. The number of neurons in the fully connected layer depends on the specific classification task. For example, in the task of distinguishing benign and malignant sacral tumors, it is set to 2 neurons (corresponding to the benign and malignant categories respectively), and the output is converted into a class probability distribution through the softmax activation function.

[0020] For the collected sample data, it is divided into a training sample set and a validation sample set. The inclusion criteria include histopathologically confirmed sacral tumors, and the user has NCCT images of a single sacral tumor before surgery. The exclusion criteria include previous anti-cancer diagnoses, poor image quality, repeated users for follow-up or monitoring, and postoperative tumor recurrence. In addition, the applicant also collected data from external test cohorts in Centers 2 and 3, specifically for the final model evaluation. This cohort included 55 users with sacral tumors, and the inclusion and exclusion criteria were the same as those described above. In addition, the applicant implemented a 5-fold cross-validation scheme on the data in Cohort 1, in which the dataset was divided into five subsets. Each subset was used as an internal test set in turn, and the remaining four subsets were used for model training and validation. In the training and validation phases, the applicant randomly divided the eligible users into a training set and a validation set at a ratio of 8:2. This method ensured that each part of the dataset was used as a training set, a validation set, and an internal test set, thus maximizing the use of the data. When making the division, ensure that the training set and the validation set have similar distributions in various features (such as age distribution, gender ratio, tumor benign-malignant ratio, etc.), and avoid the impact of data bias on model training and evaluation. At the same time, data augmentation techniques (such as randomly rotating, flipping, and scaling images, etc.) are implemented on the training sample set to increase the diversity of training samples and improve the generalization ability of the model.

[0021] S102, preprocess the training sample set to generate a target training sample set.

[0022] In one implementation, data cleaning is performed on the training sample set to generate an initial training sample set after removing invalid data. Preoperative non-contrast computed tomography (NCCT) images, T1-weighted imaging, T2-weighted imaging, and positron emission tomography-computed tomography (PET-CT) images, etc. are selected from the original training sample set. Check the integrity of the images, and remove samples with missing images, severe damage (such as the image file cannot be opened normally, the image is blurred and the tumor area cannot be recognized, etc.) or obvious artifacts (such as motion artifacts caused by user movement during scanning, artifacts caused by equipment failure, etc.). Ensure that the image formats are unified. If there are images in different formats, convert them to a format suitable for subsequent processing (such as the common DICOM format or a format that can be converted to the required format for analysis).

[0023] Organize clinical data information, including the user's demographic characteristics (such as age, gender), medical history (such as past illnesses, surgical history, allergy history, etc.), family medical history (especially family genetic disease information related to tumors), and current symptom manifestations (such as pain level, pain location, presence of neurological dysfunction, etc.). Check the accuracy and consistency of the data. For obviously incorrect data (such as unreasonable age, incorrect gender record, etc.), correct it by comparing with the original medical record or communicating with the clinician. Delete duplicate user data to ensure that each user appears only once in the training sample set. Sort out the target diagnosis interval data, which involves the time of first detection of suspected sacral lesions, the time of first diagnosis, the time of each follow-up visit or reexamination, and the relevant quantitative data of the lesion during each examination (such as tumor size, morphological changes, which can be obtained by imaging measurement; changes in lesion metabolic activity, such as changes in SUV values in PET-CT, etc.). Check the logic of the time data to ensure that the time of first diagnosis is later than the time of first symptom appearance, and the order of each follow-up visit or reexamination time is reasonable. Conduct a rationality check on the quantitative lesion data and remove obviously abnormal or incorrect data (such as unreasonable large changes in tumor size within a short period of time without reasonable explanation, etc.).

[0024] Perform normalization on the initial training sample set to generate missing sample information. For the remaining various types of medical image data (NCCT, T1 / T2 weighted imaging, PET-CT), calculate the mean and standard deviation of the image pixel values. Adopt a suitable normalization method, such as zero-mean unit-variance normalization, subtract the mean from each pixel value and then divide by the standard deviation to make the image data numerically comparable and facilitate subsequent model training. For the SUV values in PET-CT images, perform a similar normalization process to ensure that they are in the same scale range as other image data. At the same time, according to the model input requirements, adjust the size of the image, such as cropping or padding the image to a unified size (such as [224, 224, 64]) to adapt to the input layer structure of the 3D-DenseNet121 model.

[0025] For numerical features (such as age) in clinical data, perform normalization and map them to a specific interval (such as [0, 1] or [-1, 1]). For example, for age data, it can be done through the formula Normalize (with the minimum age being 1 year and the maximum age being 100 years). For categorical variables (such as gender, disease types in medical history, etc.), use one-hot encoding to convert them into numerical forms. For example, gender can be encoded as male [1, 0] and female [0, 1] so that the model can process this data. Normalize the time interval data in the diagnosis interval data (such as the time interval from the first symptom discovery to the first diagnosis, the time interval between each follow-up visit, etc.) to make its numerical value match other data types. Calculate the mean and standard deviation of the time interval, and then use a normalization method similar to that of image data for processing. For numerical data such as the growth rate of lesion size, also perform normalization to ensure that its distribution is within an appropriate numerical range for convenient fusion analysis with other data. For details, see the following text, and this will not be repeated here.

[0026] Process the missing sample information to generate missing value imputation prediction information, and then process the missing value imputation prediction information to generate target variable values. Carefully check the initial training sample set after cleaning and normalization to determine the samples and features with missing values. For image data, there are cases where some image pixels are missing or some image sequences are incomplete; in clinical data, some medical history information of users is not recorded, and part of the family medical history is missing; in the diagnosis interval data, there are problems such as missing follow-up visit times or disease quantification data. Record the positions and types of missing values to prepare for subsequent imputation prediction. Select an appropriate missing value imputation method according to the data type and distribution characteristics. Specifically, for the missing pixels in image data, use interpolation methods based on image neighborhood information, such as nearest neighbor interpolation, bilinear interpolation, or spline interpolation, to estimate the values of missing pixels based on the values of surrounding pixels. For numerical features (such as age) in clinical data, if the number of missing values is small, consider using mean imputation (filling the missing values with the mean of this feature), median imputation (filling with the median), or model-based prediction imputation methods. For example, establish a simple linear regression model with other relevant features (such as gender, certain indicators in medical history, etc.) as independent variables to predict the missing age values. For categorical variables (such as gender, disease types, etc.), use mode imputation (filling the missing values with the category with the highest frequency) or model-based prediction imputation methods (such as using classification models like decision trees, naive Bayes, etc. to predict the missing categorical values based on other features). For the missing time intervals or disease quantification data in the diagnosis interval data, select an appropriate imputation method according to the time series characteristics of the data or its relationship with other relevant data, such as linear interpolation based on the data of the previous and next time points, imputation based on the average change trend of the user group, etc.

[0027] Fill in the identified missing values according to the selected interpolation method to generate missing value interpolation prediction information. When performing interpolation, attention should be paid to maintaining the rationality and consistency of the data. For example, when using mean interpolation for age, the actual range and rationality of age should be considered to avoid unreasonable interpolation results (such as the interpolated age exceeding the normal human age range). For the interpolation of categorical variables, ensure that the interpolated categories conform to the actual situation and are reasonable in the dataset. For the interpolation of image data, ensure that the interpolated image is continuous and reasonable both visually and numerically, without affecting the subsequent extraction and analysis of image features.

[0028] In the present invention, the target variable is the benign and malignant classification of sacral tumors (e.g., 0 for benign and 1 for malignant), and this information is usually obtained from the user's histopathological confirmation results. Ensure that the definition of the target variable is clear and accurate, consistent with the clinical diagnostic criteria, and avoid errors or confusion during the data annotation process. Check the integrity and accuracy of the target variable values to ensure that each sample has a corresponding target variable value and is correctly labeled. For existing mislabeling (such as mislabeling a benign tumor as malignant or vice versa), correct it by checking against the original pathology report, clinician review, etc. If the number of missing values is small, consider deleting the corresponding samples; if the number of missing values is large and there is a certain pattern or can be inferred based on other information, try to use appropriate methods to fill them (such as inferring based on other clinical features, imaging features, etc. of the user). Finally, obtain accurate and complete target variable values to provide correct supervision information for subsequent model training.

[0029] Process the target variable values to generate target variable parameter information, and generate a target training sample set based on the target variable parameter information. Conduct a statistical analysis of the distribution of the target variable (tumor benignity and malignancy), and calculate statistical indicators such as the number and proportion of benign and malignant samples. Understand whether the distribution of tumor benignity and malignancy in the dataset is balanced. If there is a serious imbalance (such as the number of malignant samples is much larger than that of benign samples or vice versa), adopt corresponding processing strategies in subsequent model training (such as oversampling, undersampling, using a weighted loss function, etc.) to avoid the model's bias towards the majority class and improve the model's prediction ability for the minority class. Generate target variable parameter information according to the distribution characteristics and analysis results of the target variable. These parameter information include the number of categories of the target variable (in the present invention, there are 2 categories: benign and malignant), the number of samples in each category, the proportion, and the category weights (if weighted processing of different categories is required in model training), etc. For example, if the proportion of benign samples is 30% and the proportion of malignant samples is 70%, in order to balance the model's attention to the two categories, set the weight of benign samples to , and the weight of malignant samples to , during the model training process, adjust the calculation of the loss function according to the weights of the samples to make the model pay more attention to the samples of the minority class. This information on the target variable parameters will play an important role in the model training process, helping the model better learn and distinguish the sample features of different classes.

[0030] Integrate the image data, clinical data, diagnostic interval data, and target variable values after data cleaning, normalization, missing value imputation, target variable value processing, and generation of target variable parameter information according to samples. Ensure that the various types of data for each sample are accurately corresponding, that is, the imaging data, clinical data, and diagnostic interval data of each user correspond one-to-one with the corresponding tumor benign / malignant target variable values to form complete sample data. Further confirm and adjust the composition of the training sample set according to the previous division results. Ensure that the training sample set is representative, covering different types of user situations (such as different ages, genders, tumor types, disease stages, etc.), and having a certain degree of diversity and balance in various features (such as imaging features, clinical features, diagnostic interval features, etc.). At the same time, perform necessary preprocessing operations on the training sample set, such as data augmentation (such as randomly rotating, flipping, scaling images, etc. to increase the diversity of training samples) to improve the generalization ability of the model. Finally, obtain the target training sample set for training the model, and make preparations for the subsequent training based on the 3D-DenseNet121 model. The specific process is described in the subsequent content. Through the above steps, comprehensively and meticulously process the training sample set to ensure reliable data quality, unified format, complete features, and accurate target variables, laying a solid data foundation for constructing an accurate and effective pre-operative warning model for sacral tumors.

[0031] S103, process the preset three-dimensional densely connected network model based on the target training sample set and the validation sample set to generate a target three-dimensional densely connected network model.

[0032] In one implementation, any number of data features in the target training sample set are obtained, and a sampling ratio is generated based on the number of each data feature in the target training sample set. Various data features are extracted from the target training sample set, including image features (such as texture, shape, signal intensity, etc. in preoperative sacral non-contrast computed tomography images, T1-weighted imaging, T2-weighted imaging, and positron emission tomography-computed tomography images), clinical data features (such as user age, gender, medical history, family medical history, etc.), and diagnostic interval data features (such as the time interval from the first symptom discovery to the first diagnosis, the growth rate of lesion size, etc.). For image features, image processing algorithms and feature extraction techniques are used (for example, the convolutional layer in a convolutional neural network can automatically extract low-level and high-level features of images); for clinical and diagnostic interval data, the corresponding numerical or categorical information is directly obtained. The number of occurrences of each data feature in the training sample set is counted. For example, the distribution number of ages in different age groups, the number of users of different genders, the occurrence frequency of various medical histories, and the number of different imaging features are counted.

[0033] According to the number of each data feature, the sampling ratio is calculated. There can be various calculation methods for the sampling ratio. For example, it can be calculated based on the proportion of the feature numbers, that is, the sampling ratio of each feature is equal to the proportion of the number of this feature to the total number of features. Or the sampling ratio can be adjusted according to the importance of the features. Higher sampling ratios are given to important features to ensure that these key information can be fully retained during the sampling process. There are n data features in the training sample set, and the number of the i-th feature is count, then its sampling ratio . In this way, the sampling ratio corresponding to each data feature is generated, providing a basis for subsequent sampling operations.

[0034] The target training sample set is sampled based on the sampling ratio to generate a preset number of sampled features. Any data feature is processed with each sampled feature to generate multiple groups of data sets, where each group of data sets contains a preset number of data samples, and at least one data sample includes identification information. According to the calculated sampling ratio, a sampling operation is performed on the target training sample set. For each data sample, according to the data features it contains and the corresponding sampling ratio, it is determined whether this sample is selected to enter the sampling set. For example, using the random sampling method, for each sample, a random number r ( ) is generated. If ( Where is the sampling ratio of a feature in the sample), the sample is selected to enter the sampling set. Repeat this process until the preset number of sampling features is reached. Through this sampling method, the number of training samples can be reduced while retaining the data feature distribution information, the training efficiency can be improved, and the overfitting problem can be avoided to a certain extent. Especially when the training sample set is large, the sampling operation can make the training process more efficient and feasible.

[0035] For the preset number of sampling features generated, any data feature is combined with each sampling feature to generate multiple data groups. The specific operation can be to use each sampling feature as the basis of a set of data, and then add other data features in sequence to form a sample group containing multiple data features. Each data group contains a preset number of data samples, and ensures that at least one data sample includes identification information (such as whether the identification sample is a risk factor affecting sacral lesions, or other identification related to the special properties of the sample). These identification information can play an important role in subsequent model training and evaluation, such as being used to distinguish different types of samples, measure the importance of samples, or serve as a reference for model prediction. By reasonably combining data features to form a data group, it is possible to fully utilize the information in the sample set, provide rich and diverse data inputs for model training, and enable the model to learn the relationship and interaction between different features.

[0036] The preset three-dimensional densely connected network model is trained based on the data samples in the multiple data sets to generate a trained three-dimensional densely connected network model. The preset three-dimensional densely connected network model (such as the 3D-DenseNet121 model) is trained using the data samples in the generated multiple data sets. Before training, the data samples are preprocessed to meet the input requirements of the model, such as adjusting the image size, normalizing the numerical data, encoding the categorical variables, etc. (the specific preprocessing steps can refer to the relevant content in the previous text). The preprocessed data samples are input into the model, and the model extracts and learns the data according to its internal network structure (including convolutional layers, densely connected layers, transition layers, and fully connected layers, etc.). During the training process, the prediction results of the model are calculated by forward propagation, and then the loss value is calculated using a loss function (such as a cross entropy loss function) based on the difference between the prediction results and the true labels (such as the benign and malignant labels of tumors). Next, the gradient of the loss function to the model parameters is calculated through the back-propagation algorithm. The model parameters are updated using an optimizer (such as the Adam optimizer) based on the gradient, and the model weights and biases are continuously adjusted to gradually make the model's prediction results approach the true label. This process is repeated until the predetermined number of training rounds is completed or other training stop conditions are met (such as the loss value no longer decreases or reaches the set minimum threshold).

[0037] Process the trained three-dimensional densely connected network model based on the validation sample set to generate a validation result. If the data samples containing identification information in the validation result are risk factors characterizing the impact of sacral lesions, then regard the trained three-dimensional densely connected network model as the target three-dimensional densely connected network model. After the model training is completed, use the validation sample set to process the trained three-dimensional densely connected network model to generate a validation result. Perform the same preprocessing operations on the data samples in the validation sample set as those on the training samples, and then input them into the trained model to obtain the prediction results of the model for the validation samples. According to the prediction results and the true labels of the validation samples, calculate various evaluation metrics, such as accuracy, recall rate, F1 score, area under the curve (AUC), etc., to comprehensively evaluate the performance of the model on unseen data. These evaluation metrics can reflect the accuracy, sensitivity, specificity, and overall discriminative ability of the model, helping to determine whether the model is overfitting or underfitting, and the effectiveness of the model in practical applications. For example, the accuracy represents the proportion of the number of samples correctly predicted by the model to the total number of samples, and the AUC measures the ability of the model to distinguish positive and negative examples by calculating the area under the receiver operating characteristic curve (ROC curve).

[0038] Check the data samples containing identification information in the validation result to determine whether these samples are risk factors characterizing the impact of sacral lesions. Determine which identification information is related to the risk of sacral lesions through data analysis. For example, certain specific imaging features, clinical features (such as older age, having a specific family history, etc.), or diagnostic interval features (such as a faster lesion growth rate) are regarded as risk factors. If the number of samples containing these risk factors in the validation result is large or the prediction performance of the model on the samples related to these risk factors is good (such as high accuracy, large AUC), it indicates that the model can better identify and process information related to risks. If the validation result shows that the trained model performs well in dealing with the risk factors affecting sacral lesions, then regard the trained three-dimensional densely connected network model as the target three-dimensional densely connected network model. This target model will be used to predict and warn of preoperative sacral lesions of the target user later, providing auxiliary support for clinical diagnosis. If the validation result is not satisfactory, it is necessary to adjust the model parameters, reselect features, increase training data, or adopt other improvement measures, and then train and validate again until a target model that meets the requirements is obtained. Through this model selection process based on the validation result, it can be ensured that the finally determined model has good performance and reliability, can accurately predict the attribute information of sacral lesions in practical applications, and provide valuable reference for clinical decision-making.

[0039] S104, Process the preoperative non-contrast computed tomography images of the sacrum of the target user to generate three-dimensional region of interest images of the target user.

[0040] In one implementation, obtain preset anatomical structure restriction information, where the preset anatomical structure restriction information is used to represent organ information that matches the sacral structure. Consult professional anatomical atlases, medical imaging anatomy textbooks, or authoritative medical databases to obtain detailed anatomical knowledge related to the sacral structure. These resources contain information such as the morphology, position of the normal human sacrum, and the anatomical relationships with surrounding organs and tissues, such as the adjacent relationships between the sacrum and the pelvis, spine, nerves, blood vessels, and surrounding soft tissues. Through in-depth study of these materials, determine the range of organs and tissues that are closely related to the sacral structure in physiological and pathological states, providing an accurate anatomical basis for subsequent processing. Define the specific content of the preset anatomical structure restriction information. Determine which organs and tissues are closely adjacent to the sacrum in space and may affect the evaluation and treatment of tumors. For example, determine the pelvic bone structure within a certain range, nearby nerve plexuses (such as the sciatic nerve, cauda equina, etc.), important blood vessels (such as branches of the internal iliac artery, etc.), and soft tissue areas that may be invaded by tumors. Describe and record this information in a standardized manner, such as by defining anatomical coordinate ranges, structure names, and relative position relationships, to form preset anatomical structure restriction information for subsequent image processing, ensuring that these key structure information can be accurately identified and utilized when processing the images of target users.

[0041] The preoperative non-contrast computed tomography (NCCT) images of the sacrum of the target user are subjected to image resampling to generate NCCT images of the sacrum with the target size. According to the input requirements of the preset three-dimensional densely connected network model and the convenience of subsequent processing, the parameters of the target size are determined. For example, referring to the optimal adaptation requirements of the model for the input image size, it is determined that the image is resampled to a size of [224, 224, 64] (width, height, depth). Such a size can not only ensure that the image contains sufficient detailed information but also achieve a good balance between computing resources and model processing efficiency. At the same time, considering the need to observe tumors and surrounding structures at different scales, the selection of the target size should also be conducive to highlighting the relationship between the lesion area and the surrounding anatomical structures, facilitating the accurate identification and analysis of tumor characteristics. According to the characteristics of the image and the requirements of the target size, a suitable resampling algorithm is selected. Considering both image quality and computational efficiency, for sacral NCCT images, the bilinear interpolation algorithm is a suitable choice. It can generate relatively smooth and accurate images of the target size without adding too much computational burden, meeting the requirements for image quality in subsequent processing. The selected resampling algorithm is used to process the preoperative non-contrast computed tomography images of the sacrum of the target user. The pixel values of the original image are recalculated and assigned according to the algorithm rules to generate NCCT images of the sacrum with the target size. During the resampling process, it is necessary to ensure that the anatomical structure information of the image is retained, especially that the relative positional relationship between the tumor area and its surrounding tissues should not be significantly distorted or distorted due to resampling. Through resampling, the original image is converted into an image with a unified size, enabling it to better meet the requirements of subsequent image analysis and model input, laying a foundation for accurately extracting the tumor area and other relevant information.

[0042] Process the non-contrast computed tomography (NCCT) images of the sacrum with the target size to generate the cropped information of the target tumor region. Use image processing algorithms and techniques to analyze the NCCT images of the sacrum with the target size and initially locate the tumor region. Adopt methods based on image gray value, texture features or morphological analysis. For example, through threshold segmentation technology, according to the difference in gray value between the tumor tissue and the surrounding normal tissue, segment the regions in the image that may belong to the tumor; or use texture feature analysis algorithms to identify the regions with specific texture patterns (such as the uneven texture often exhibited by tumor tissue) as tumor candidate regions. Combine these methods to initially determine the approximate position and scope of the tumor in the image, providing a reference for further precise cropping. According to the initially located tumor region and combined with the preset anatomical structure constraint information, perform precise cropping operations. Ensure that the cropped target tumor region not only contains the tumor tissue itself, but also includes the tissues and structures that may be affected within a certain range around it (determine the range according to the anatomical structure constraint information). During the cropping process, pay attention to avoiding cropping off the information that is of great value for tumor diagnosis and analysis, such as the relationship between the tumor and the surrounding important blood vessels and nerves. Through precise cropping, obtain a relatively small image part that focuses on the tumor and its surrounding key regions, reducing the interference of irrelevant information, highlighting the tumor features, facilitating more detailed analysis and processing of the tumor region in the follow-up, and also helping to improve the efficiency and accuracy of subsequent model processing.

[0043] Process the cropped target tumor region information based on the preset anatomical structure constraint information to generate the three-dimensional (3D) region of interest (ROI) image of the target user. Integrate the 3D anatomical structure data in the preset anatomical structure constraint information (such as the 3D models or spatial coordinate ranges of the surrounding organs and tissues) with the reconstructed 3D tumor region, so that the 3D ROI image not only contains the tumor itself, but also clearly shows its positional relationship with the surrounding important anatomical structures in the 3D space, providing more abundant information for a comprehensive assessment of the tumor. Perform post-processing and optimization operations on the generated 3D ROI image to improve the visualization effect and information expression ability of the image. For example, the contrast and brightness of the image can be adjusted to make the tumor region and the surrounding structures more clearly distinguishable; the image can be smoothed to reduce noise interference, but pay attention to avoiding excessive smoothing that may cause the loss of image details; pseudo-color coding technology can also be applied to assign different colors according to the characteristics of different tissues (such as tumor tissue, normal bone tissue, soft tissue, etc.), enhancing the readability and distinguishability of the image. Through these post-processing operations, generate a high-quality 3D ROI image of the target user, making it more suitable for subsequent clinical analysis, diagnosis, and as the input data of the preset 3D densely connected network model, providing strong support for accurately predicting the properties of sacral lesions.

[0044] S105. Process the multimodal data information of the target user and the target diagnosis interval data information of the target user to generate the dynamic risk parameter of the target user.

[0045] In one implementation, perform feature extraction processing on the T1-weighted imaging information of the target user and the T2-weighted imaging information of the target user to generate signal intensity features, texture features, shape features, and enhancement curve features. Using medical image processing professional software as a tool, load the T1-weighted imaging and T2-weighted imaging information of the target user to ensure that the image data is complete, in a unified format, and the resolution meets the analysis requirements. Apply an image segmentation algorithm to accurately identify and extract the tumor area, reducing the impact of human error on subsequent analysis. When calculating the signal intensity features, use a robust statistical method to calculate the mean, variance, maximum value, and minimum value of the signal intensity within the tumor area. At the same time, based on the statistical analysis results of large-scale clinical sample data, determine the signal intensity threshold range that is diagnostically significant for sacral tumors. For example, through the study of thousands of sacral tumor cases, it is found that when the mean signal intensity on the T1-weighted image is lower than 80 (unit) and the variance is less than 50 (unit), the possibility of a benign tumor is relatively high. Normalize the calculated signal intensity features and map them to the [0,1] interval to eliminate the numerical differences caused by different imaging devices and scanning parameters, ensuring the comparability of features in multi-source data fusion.

[0046] Based on the gray-level co-occurrence matrix (GLCM) method, determine the combination of directions (0°, 45°, 90°, 135°) and distance parameters (1 pixel, 2 pixels) applicable to the present invention according to the special research results on the texture features of sacral tumors. During the calculation process, to improve the calculation efficiency and accuracy, use parallel computing technology to accelerate the pixel pair frequency statistics in different directions and distances. After calculating texture feature parameters such as contrast, correlation, energy, and entropy, normalize each texture feature using a standardization formula (such as z-score standardization) to make its mean 0 and standard deviation 1, enhancing the stability and comparability of features among different cases. In addition, introduce the spatial frequency analysis of texture features to supplement the description of tumor texture features and further improve the ability of texture features to distinguish between benign and malignant tumors.

[0047] Use a high-precision edge detection algorithm (such as an improved Canny edge detection algorithm, combined with morphological operations to optimize the edge extraction effect) to obtain the precise contour of the tumor area. Calculate the perimeter, area, circularity of the tumor (formula: Shape features such as ( ) and eccentricity are considered. Considering the differences in tumor size and image scale among different users, a normalization method based on image moments is used to normalize the shape features, so that the shape features are not affected by the actual size of the tumor and the image resolution. At the same time, a database containing the shape features of common sacral tumors is constructed. By comparing and matching with the standard shape patterns in the database, it helps to judge the type and benign / malignant tendency of the tumor, providing an intuitive reference basis for shape features for clinical diagnosis.

[0048] For the case of T1-weighted enhanced imaging sequences, a series of T1-weighted enhanced images are acquired at clinical standard time points (such as 1 minute, 3 minutes, 5 minutes, 7 minutes, etc. after injection of the contrast agent). In the image preprocessing stage, image registration algorithms (such as registration methods based on maximizing mutual information) and motion correction techniques (such as optical flow method) are used to ensure the accurate registration of images at different time points, eliminating the interference of user movement on the calculation of enhancement curve features. For each tumor pixel, an enhancement curve of signal intensity changing with time is plotted, and its peak time, rising slope, falling slope, and area under the curve are calculated. To improve the accuracy and stability of the calculation, the spline interpolation method is used to smooth the enhancement curve before feature calculation. The enhancement curve features are normalized to make them compatible with other features in the numerical range for subsequent fusion analysis. At the same time, according to clinical case data, enhancement curve feature templates for different types of sacral tumors are established. By comparing the similarity with the templates, it helps to judge the nature and invasiveness of the tumor.

[0049] Feature extraction processing is performed on positron emission tomography-computed tomography information to generate target uptake value features. Using PET-CT image fusion technology, the PET image and the CT image are accurately registered to accurately locate the tumor area. An adaptive threshold segmentation algorithm is adopted to automatically determine the segmentation threshold according to the metabolic activity characteristics of the tumor tissue in the PET image, and the standardized uptake value (SUV) information of the tumor area is extracted. The maximum value, minimum value, average value, and standard deviation of SUV are calculated to comprehensively characterize the metabolic activity level of the tumor tissue. Considering the calibration differences of different PET-CT devices, a normalization method based on phantom calibration is used to correct and normalize the SUV features to ensure the comparability of SUV features among different devices. In addition, SUV histogram analysis is introduced to extract statistical features such as skewness and kurtosis of the SUV distribution, describing the metabolic heterogeneity of the tumor from a more comprehensive perspective and providing more valuable information for tumor diagnosis.

[0050] Perform feature quantization processing on the target diagnostic interval data information of the target user to generate access time interval information and lesion size growth rate values. Collect the target diagnostic interval data from the previous medical records of the target user, including the time of first detection of suspected sacral lesions, the time of first diagnosis, the time of each follow-up visit or follow-up, and the relevant quantitative data of the lesion during each examination (such as tumor size and morphological changes, which can be obtained by imaging measurement; changes in lesion metabolic activity, such as changes in SUV values in PET-CT, etc.). Carefully check the accuracy of the time data to ensure that the time sequence is reasonable and there are no logical errors (such as the time of first diagnosis must be later than the time of first symptom discovery). For the quantitative lesion data, adopt quality control measures to ensure the accuracy and repeatability of the measurement data, such as having experienced radiologists or professional technicians perform multiple measurements and take the average. Calculate the access time interval information, that is, the time difference from the first symptom discovery to the first diagnosis, and the time interval of each follow-up visit or follow-up, and record it in days. For the calculation of the lesion size growth rate value, according to the imaging measurement results (such as diameter or volume) of the tumor at different time points, use a linear regression model to fit the tumor growth curve, and calculate the slope of the growth curve as the lesion size growth rate value. Perform normalization processing on the access time interval information and the lesion size growth rate value, map them to a specific interval (such as [-1,1]), so that they match the numerical range of other features, facilitating subsequent fusion analysis. At the same time, analyze the correlation between the access time interval and the lesion size growth rate and the benign and malignant nature and prognosis of the tumor, providing a basis for the calculation of dynamic risk parameters.

[0051] Process the signal intensity feature, texture feature, shape feature, enhancement curve feature, and target uptake value feature to generate the target shape factor. Integrate the signal intensity feature, texture feature, shape feature, enhancement curve feature, and target uptake value feature to construct a multi-modal feature vector. Use dimensionality reduction algorithms such as principal component analysis (PCA) or independent component analysis (ICA) to perform dimensionality reduction processing on the multi-modal feature vector, and extract the main components as the target shape factor. During the dimensionality reduction process, determine the optimal number of principal components or independent components through cross-validation methods to ensure that while retaining the feature information, the data dimension is reduced and the computational complexity is reduced. Perform standardization processing on the generated target shape factor to make it follow a standard normal distribution, facilitating subsequent fusion analysis with other parameters. At the same time, through visualization techniques (such as plotting the distribution scatter plot of the target shape factor among different cases), intuitively display the relationship between the target shape factor and the benign and malignant nature of the tumor and other clinical features, providing a more intuitive diagnostic reference for clinicians.

[0052] Process the target shape factor based on the access time interval information and the lesion size growth rate value to generate the dynamic risk parameter of the target user. Based on the access time interval information and the lesion size growth rate value, construct a linear weighted model to calculate the dynamic risk parameter of the target user. For example, let the dynamic risk parameter (where represents the dynamic risk parameter, represents the target shape factor, represents the access time interval information, represents the lesion size growth rate value, is the weight coefficient obtained by training through machine learning algorithms (such as random forest regression, support vector machine regression, etc.) based on a large amount of clinical sample data. During the process of training the weight coefficient, the leave-one-out cross-validation or k-fold cross-validation (such as k = 10) method is used to evaluate the model performance, and the model parameters with the best performance are selected. According to the calculated dynamic risk parameter, establish a risk grading standard. For example, divide the dynamic risk parameter into three levels: low risk ( ), medium risk ( ), and high risk ( ), providing a clear risk assessment basis for clinical decision-making.

[0053] S106, process the clinical data information of the target user to generate the physiological attribute parameter of the target user.

[0054] In one implementation, perform feature extraction processing on the clinical data information of the target user to generate demographic characteristics, tumor location characteristics, gene characteristics of the target user, and comorbidity characteristics of the target user. Accurately extract the basic demographic information of the target user from the hospital information system (HIS) or electronic medical record (EMR), including age, gender, height, weight, race, etc. Ensure the integrity and accuracy of the data. For missing or unclear data, supplement and verify it by communicating with the user or family members, consulting other relevant records, etc. For example, if the age data is missing, it can be calculated based on information such as the user's first visit time and date of birth; if the gender information is unclear, it can be determined by asking the user or referring to relevant statements in other medical records. Organize these demographic characteristics into a standardized data format for subsequent analysis and processing.

[0055] Based on the preoperative non-contrast computed tomography (NCCT) images, magnetic resonance imaging (MRI) images, or other relevant imaging data of the target user, an experienced radiologist or a professionally trained medical imaging analyst determines the specific location of the tumor in the sacrum. The standard anatomical positioning system (such as the internationally recognized anatomical nomenclature) is used to accurately describe the tumor location, record the positional relationship of the tumor relative to each anatomical landmark of the sacrum (such as the sacral promontory, sacral hiatus, etc.), and the distribution range of the tumor in the anterior-posterior, left-right, and superior-inferior directions of the sacrum. At the same time, the adjacent relationship between the tumor and surrounding important structures (such as nerves, blood vessels, pelvic organs, etc.) is evaluated, and these location information and adjacent relationships are converted into quantifiable or classified feature data, such as the distance between the tumor and the nerve, whether the tumor invades the surrounding blood vessels, etc.

[0056] Obtain the gene detection report of the target user and extract the gene feature information related to sacral tumors, including the mutation status of specific genes (such as detecting whether there are mutations in known genes related to tumor occurrence and development, such as mutations in genes like TP53, RB1, etc.), gene expression levels (such as the mRNA expression levels of certain tumor-related genes), and gene copy number variations (such as amplification or deletion of certain gene regions). Standardize the gene feature data so that it can be analyzed in the same data framework as other clinical data. For example, perform logarithmic transformation on the gene expression levels to make the data distribution closer to a normal distribution, facilitating subsequent calculations and comparisons. Comprehensively sort out the past medical history and current health status of the target user to determine whether there are other comorbidities. The comorbidity information can be obtained from multiple aspects such as the diagnosis records, hospitalization records, outpatient visit records, and user self-report in the electronic medical record. Common comorbidities related to sacral tumors include diabetes, cardiovascular diseases (such as hypertension, coronary heart disease, etc.), pulmonary diseases (such as chronic obstructive pulmonary disease), etc. Each comorbidity is regarded as a feature and is coded in a binary classification according to its presence or absence (such as recording the presence of a comorbidity as 1 and the absence as 0), or coded in a graded manner according to the severity of the comorbidity (such as coding mild, moderate, and severe as 1, 2, 3, etc.) to reflect the impact degree of the comorbidity on the user's physiological state.

[0057] Process the demographic characteristics, tumor location characteristics, genetic characteristics of the target users, and comorbidity characteristics of the target users respectively to generate the importance score value of each clinical data information and the weight value of each clinical data information. To calculate the importance score value of each clinical data information, a Gradient Boosting Decision Tree (GBDT) model is used. The GBDT model consists of multiple decision trees, and each decision tree learns the feature relationship of the data by optimizing the objective function during the construction process. Construct a GBDT model containing M decision trees (base learners). When constructing the m-th decision tree (m = 1, 2,..., M), a certain proportion (such as 70%) of the samples are randomly drawn from the training data with replacement as the training set of this decision tree, and the remaining samples are used as the validation set.

[0058] For each decision tree, the splitting of its leaf nodes is based on information gain or other appropriate splitting criteria. For example, when calculating the splitting criterion of the m-th decision tree, the contribution of each feature to reducing the sample impurity (such as Gini impurity or information entropy) is considered. At a certain node, there are n samples, and the samples are divided into different subsets according to different values of a certain feature V, and the reduction amount of the sample impurity before and after the division is calculated. . In the calculation process, the feature vector of the sample is involved. (including demographic characteristics, tumor location characteristics, genetic characteristics, and comorbidity characteristics, etc.) and the true label. (such as the benign and malignant label of the tumor, recorded as 0 for benign and 1 for malignant). The number of leaf nodes of each decision tree will vary according to the complexity of the data and the settings of the model. For example, for complex data, more leaf nodes may be generated to better fit the data. In each decision tree, each leaf node has a sum of sample weights, which represents the total weight of the samples reaching this leaf node. This weight is adjusted according to the importance of the sample and the misclassification situation during the model training process.

[0059] The method also includes a calculation formula for obtaining the importance score value. The calculation formula is: ; where represents the importance score value of the j-th feature; M is the number of base learners (decision trees) in the gradient boosting decision tree model; is the number of leaf nodes of the m-th decision tree; is the sum of the sample weights of the t-th leaf node in the m-th decision tree; n is the number of samples; represents the feature used when splitting the t-th leaf node in the m-th decision tree; is an indicator function. When is true, , otherwise represents the reduction in impurity before and after the t-th split in the m-th decision tree for the i-th sample; is the feature vector of the i-th sample at the t-th split in the m-th decision tree, is the true label of the i-th sample (e.g., representing benign or malignant). For each of the M decision trees, traverse its leaf nodes. When at the t-th leaf node of the m-th decision tree, if the feature used for splitting at this node is exactly the age feature (i.e., ), then . At this time, calculate the reduction in impurity of sample i before and after the split at this node , where is the feature vector of sample i at the t-th split in the m-th decision tree (including the age feature and other relevant features), is the true label of sample i (tumor benign or malignant). Then, calculate the contribution value of the age feature at this node according to the formula . Sum up the contribution values of the age feature at each leaf node in all decision trees to obtain the importance score value of the age feature . The same method can be applied to calculate the importance score values of other clinical data information such as tumor location features, gene features, and comorbidity features. In this way, the importance of each feature in distinguishing tumor benign or malignant or other clinical decisions can be quantified, providing a basis for calculating weight values in the subsequent steps.

[0060] After obtaining the importance score value of each clinical data information, calculate its weight value. Normalize the importance score value of each feature so that their sum is 1 to obtain the weight value of each feature. Let the importance score value of the age feature be , the importance score value of the tumor location feature be , the importance score value of the gene feature be , and the importance score value of the comorbidity feature be , then the weight value of the age feature , and the weight values of other features can be calculated in the same way. In this way, the weight value of each feature reflects its relative importance in the overall clinical data information. The larger the weight value, the greater the impact of this feature on the physiological attribute parameters of the target user. Process the importance score value of each clinical parameter information and the weight value of each clinical data information respectively to generate the physiological attribute parameters of the target user. According to the calculated weight value of each clinical data information, perform weighted summation on the corresponding features to generate the physiological attribute parameters of the target user. The demographic features are processed to obtain the feature vector (where Represents the specific eigenvalue in demographic characteristics, such as the value after encoding or numerical conversion of age, gender, etc., and its corresponding weight vector is ; The tumor location feature vector is , and the weight vector is ; The gene feature vector is , and the weight vector is ; The comorbidity feature vector is , and the weight vector is .

[0061] Then the physiological attribute parameters of the target user can be calculated by the following formula: (where represents the transpose of vector X). By this way of weighted summation, the demographic characteristics, tumor location characteristics, gene characteristics and comorbidity characteristics are combined to generate a parameter that can comprehensively reflect the relationship between the physiological attributes of the target user and the tumor. This physiological attribute parameter can be used as part of the subsequent model input, together with other data (such as imaging data, dynamic risk parameters, etc.), to provide more comprehensive information support for accurately predicting the attribute information of sacral tumors.

[0062] S107, process the preoperative non-contrast computed tomography image of the sacrum of the target user, the three-dimensional region of interest image of the target user, the dynamic risk parameter of the target user and the physiological attribute parameter of the target user based on the target three-dimensional densely connected network model to generate sacral early warning information.

[0063] In one implementation, perform image data normalization processing on the preoperative non-contrast computed tomography image of the sacrum of the target user and the three-dimensional region of interest image of the target user to generate the mean of the image pixel values and the standard deviation of the image pixel values. We have a set of preoperative non-contrast computed tomography image datasets of the sacrum and three-dimensional region of interest image datasets . Calculate and the mean and standard deviation of each pixel value (or voxel value) in. For each pixel value in , perform normalization processing: . For each voxel value in , perform normalization processing: .

[0064] Perform parameter standardization on the dynamic risk parameters of the target user and the physiological attribute parameters of the target user to generate the mean of the dynamic risk parameters, the standard deviation of the dynamic risk parameters, the mean of the physiological attribute parameters, and the standard deviation of the physiological attribute parameters. Let the dynamic risk parameter be and the physiological attribute parameter be . Calculate the mean and standard deviation of , and the mean and standard deviation of . Standardize the dynamic risk parameters: . Standardize the physiological attribute parameters: . Perform fusion processing on the mean of the image pixel values, the standard deviation of the image pixel values, the mean of the dynamic risk parameters, the standard deviation of the dynamic risk parameters, the mean of the physiological attribute parameters, and the standard deviation of the physiological attribute parameters to generate multi-channel feature data. Let the mean and standard deviation after image data normalization be (including the fusion result of ct and 3d), and the mean and standard deviation after parameter standardization be (including the fusion result of dynamic risk and physiological attribute parameters). Fuse these data to generate multi-channel feature data , and concatenate these data together in order: .

[0065] Process the multi-channel feature data based on the target three-dimensional densely connected network model to generate the target prediction value. The method includes a calculation formula for obtaining the target prediction value, and the calculation formula is: ; where L represents the total number of layers of the three-dimensional densely connected network model (including convolutional layers, fully connected layers, etc.), represents the weight parameter of the l-th layer, and its dimension is determined according to the size of the input and output feature maps and the number of neurons in the l-th layer; represents the learning parameter vector of the l-th layer, represents the feature transformation operation performed by the l-th layer on the input multi-channel feature data F; represents the bias parameter of the l-th layer, and its dimension is the same as the size of the output feature map or the number of neurons in the l-th layer; represents the activation function of the output layer. Input the multi-channel feature data into the three-dimensional densely connected network model. The three-dimensional densely connected network model has L layers, and performs feature transformation and propagation layer by layer. At the l-th layer ( ), according to the input multi-channel feature data , combine the weight parameter , the learning parameter vector and the bias parameter for calculation.

[0066] Specifically, the three-dimensional densely connected network model has L = 3 layers. In the first layer: the weight parameter is 0.5 (a randomly initialized example value), the learning parameter vector is [0.2, 0.3] (a randomly initialized example value), and the bias parameter is 0.1 (a randomly initialized example value). The input multi-channel feature data undergoes the feature transformation of the first layer (the feature transformation function is a linear transformation plus a non-linear activation function, for example , where relu is the rectified linear unit function). The output of the first layer is . In the second layer: the weight parameter is 0.3, the learning parameter vector is [0.1, 0.4], and the bias parameter is 0.2. The input is the output of the first layer , which undergoes the feature transformation of the second layer . The output of the second layer is . In the third layer: the weight parameter is 0.2, the learning parameter vector is [0.3, 0.2], and the bias parameter is 0.3. The input is the output of the second layer , which undergoes the feature transformation of the third layer . The final target prediction value (here is the activation function of the output layer, and the sigmoid function is used to map the result between 0 and 1).

[0067] Based on a preset threshold, the target prediction value is processed to generate sacral warning information, which is used to characterize the attribute information of the target user's preoperative sacral lesion. Assume the preset threshold is T = 0.7. When the target prediction value y meets different conditions, different sacral warning information is generated: If , then high-risk sacral warning information is generated. This means that the preoperative sacral lesion of the target user has a relatively high possibility of belonging to a more serious attribute, such as the lesion may be malignant, the lesion range is large, the lesion invasiveness is strong, etc. If , then low-risk sacral warning information is generated. This indicates that the preoperative sacral lesion of the target user may have relatively mild attributes, such as the lesion may be benign, the lesion range is small, the lesion is relatively stable, etc.

[0068] First, the server widely collects various types of data of the target user, including clinical data, NCCT images, multi-modal data, and diagnostic interval data, etc., and carefully preprocesses the training samples to ensure data quality. When constructing the model, the three-dimensional densely connected network is optimized with the target training and validation samples, and the final model is determined through sampling feature training and validation. At the same time, the NCCT image is processed according to the anatomical structure constraint information to generate the region of interest image. In terms of feature processing, features are extracted from multi-modal data to obtain dynamic risk parameters, and features are extracted from clinical data to obtain physiological attribute parameters. After normalizing and standardizing these data and image data, a weighted fusion strategy (assigning learnable weights to the image, dynamic risk parameter, and physiological attribute parameter channels) is used to generate multi-channel feature data. Finally, the multi-channel feature data is input into the model, and warning information is generated according to the preset dynamic threshold, effectively assisting in determining the benign and malignant nature of the tumor and improving the accuracy of preoperative diagnosis.

[0069] In one implementation, as Figure 2 shown, the present application further provides a preoperative warning device for sacral tumors, including: An acquisition module 201, configured to acquire the clinical data information of the target user, the preoperative non-contrast computed tomography image of the sacrum of the target user, the multi-modal data information of the target user, the target diagnostic interval data information of the target user, a preset three-dimensional densely connected network model, a training sample set, and a validation sample set, wherein the multi-modal data information of the target user includes the T1-weighted imaging information of the target user, the T2-weighted imaging information of the target user, and positron emission tomography-computed tomography information; A processing module 202, configured to preprocess the training sample set to generate a target training sample set; process the preset three-dimensional densely connected network model based on the target training sample set and the validation sample set to generate a target three-dimensional densely connected network model; process the preoperative non-contrast computed tomography image of the sacrum of the target user to generate a three-dimensional region of interest image of the target user; process the multi-modal data information of the target user and the target diagnostic interval data information of the target user to generate dynamic risk parameters of the target user; process the clinical data information of the target user to generate physiological attribute parameters of the target user; process the preoperative non-contrast computed tomography image of the sacrum of the target user, the three-dimensional region of interest image of the target user, the dynamic risk parameters of the target user, and the physiological attribute parameters of the target user based on the target three-dimensional densely connected network model to generate sacral warning information, wherein the sacral warning information is used to characterize the attribute information of the preoperative sacral lesion of the target user.

[0070] Each embodiment in this application is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the preoperative warning method, electronic device, electronic equipment, and readable storage medium for evaluating sacral tumors, since they are basically similar to the embodiments of the preoperative warning method for sacral tumors described above, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the embodiments of the preoperative warning method for sacral tumors described above.

Claims

1. A preoperative early warning method for sacral tumors, characterized in that: include: Acquire clinical data information of a target user, a preoperative non-contrast computed tomography image of the sacrum of the target user, multimodal data information of the target user, target diagnostic interval data information of the target user, a preset three-dimensional densely connected network model, a training sample set, and a validation sample set, wherein the multimodal data information of the target user includes T1-weighted imaging information of the target user, T2-weighted imaging information of the target user, and positron emission tomography-computed tomography information; Preprocessing the training sample set to generate a target training sample set; Processing the preset three-dimensional densely connected network model based on the target training sample set and the verification sample set to generate a target three-dimensional densely connected network model; Processing a preoperative non-contrast computed tomography image of the sacrum of the target user to generate a three-dimensional region of interest image of the target user; Processing the multimodal data information of the target user and the target diagnostic interval data information of the target user to generate a dynamic risk parameter of the target user; Processing the clinical data information of the target user to generate physiological attribute parameters of the target user; Based on the target three-dimensional densely connected network model, the preoperative sacrum non-contrast computed tomography image of the target user, the three-dimensional region of interest image of the target user, the dynamic risk parameters of the target user and the physiological attribute parameters of the target user are processed to generate sacrum warning information, wherein the sacrum warning information is used to characterize the attribute information of the target user's preoperative sacrum lesions.

2. The method according to claim 1, characterized in that Preprocessing the training sample set to generate a target training sample set includes: Performing data cleaning processing on the training sample set to generate an initial training sample set after removing invalid data; performing normalization processing on the initial training sample set to generate missing sample information; The missing sample information is processed to generate missing value interpolation prediction information; the missing value interpolation prediction information is processed to generate a target variable value; the target variable value is processed to generate target variable parameter information; and a target training sample set is generated based on the target variable parameter information.

3. The method according to claim 2, characterized in that The preset three-dimensional densely connected network model is processed based on the target training sample set and the verification sample set to generate a target three-dimensional densely connected network model, including: Obtain any number of data features in the target training sample set; Generating a sampling ratio based on the number of each data feature in the target training sample set; Sampling the target training sample set based on the sampling ratio to generate a preset number of sampling features; Processing based on any data feature and each sampling feature to generate multiple data groups, wherein each data group includes a preset number of data samples, and at least one data sample includes identification information; Training the preset three-dimensional densely connected network model based on data samples in the multiple data groups to generate a trained three-dimensional densely connected network model; Processing the trained three-dimensional densely connected network model based on the verification sample set to generate a verification result; If the data sample containing identification information in the verification result is a risk factor for sacral lesions, the trained three-dimensional densely connected network model is used as the target three-dimensional densely connected network model.

4. The method according to claim 1, characterized in that Processing the preoperative sacrum non-contrast computed tomography image of the target user to generate a three-dimensional region of interest image of the target user includes: Acquiring preset anatomical structure restriction information, wherein the preset anatomical structure restriction information is used to characterize organ information matching the sacrum structure; performing image resampling processing on the preoperative sacrum non-contrast computed tomography image of the target user to generate a sacrum non-contrast computed tomography image of a target size; Processing the sacrum non-contrast computed tomography image of the target size to generate cropped target tumor region information; The cropped target tumor region information is processed based on the preset anatomical structure restriction information to generate a three-dimensional region-of-interest image of the target user.

5. The method according to claim 1, characterized in that Processing the multimodal data information of the target user and the target diagnostic interval data information of the target user to generate a dynamic risk parameter of the target user includes: Performing feature extraction processing on the T1-weighted imaging information of the target user and the T2-weighted imaging information of the target user to generate signal intensity features, texture features, shape features, and enhancement curve features; Performing feature extraction processing on the PET-CT information to generate target uptake value features; Performing feature quantification processing on the target diagnosis interval data information of the target user to generate access time interval information and lesion size growth rate value; Processing the signal intensity feature, the texture feature, the shape feature, the enhancement curve feature, and the target uptake value feature to generate a target shape factor; The target shape factor is processed based on the access time interval information and the lesion size growth rate value to generate a dynamic risk parameter of the target user.

6. The method according to claim 5, characterized in that The clinical data information of the target user is processed to generate physiological attribute parameters of the target user, including: Performing feature extraction processing on the clinical data information of the target user to generate demographic features, tumor location features, gene features of the target user, and comorbidity features of the target user; The demographic characteristics, tumor location characteristics, target user's genetic characteristics and target user's comorbidity characteristics are processed respectively to generate the importance score value and weight value of each clinical data information; The importance score value of each clinical parameter information and the weight value of each clinical data information are processed respectively to generate the physiological attribute parameters of the target user; The method also includes a calculation formula for obtaining an importance score value, and the calculation formula is: ; in, Represents the importance score of the jth feature; is the number of decision trees in the gradient boosted decision tree model; is the number of leaf nodes of the mth decision tree; is the sum of the sample weights of the t-th leaf node in the m-th decision tree; n is the number of samples; Represents the feature used when splitting the tth leaf node in the mth decision tree; is an indicator function, when hour, ,otherwise It represents the reduction of impurity of the i-th sample before and after the t-th split in the m-th decision tree; is the feature vector of the i-th sample at the t-th split in the m-th decision tree, is the true label of the i-th sample.

7. The method according to claim 6, characterized in that The target three-dimensional densely connected network model is used to process the preoperative sacrum non-contrast computed tomography image of the target user, the three-dimensional region of interest image of the target user, the dynamic risk parameters of the target user, and the physiological attribute parameters of the target user to generate sacrum warning information, including: Performing image data normalization processing on the preoperative sacrum non-contrast computed tomography image of the target user and the three-dimensional region of interest image of the target user to generate a mean value of image pixel values ​​and a standard deviation of image pixel values; Performing parameter standardization processing on the dynamic risk parameters of the target user and the physiological attribute parameters of the target user to generate a mean value of the dynamic risk parameters, a standard deviation of the dynamic risk parameters, a mean value of the physiological attribute parameters, and a standard deviation of the physiological attribute parameters; Performing fusion processing on the mean of the image pixel values, the standard deviation of the image pixel values, the mean of the dynamic risk parameter, the standard deviation of the dynamic risk parameter, the mean of the physiological attribute parameter, and the standard deviation of the physiological attribute parameter to generate multi-channel feature data; Processing the multi-channel feature data based on the target three-dimensional densely connected network model to generate a target prediction value; Processing the target prediction value based on a preset threshold value to generate sacral warning information; The method includes a calculation formula for obtaining a target prediction value, and the calculation formula is: ; Among them, L represents the total number of layers of the three-dimensional densely connected network model, Represents the weight parameter of the lth layer; represents the learning parameter vector of the lth layer, Represents the feature transformation operation performed by the lth layer on the input multi-channel feature data F; Represents the bias parameter of the lth layer; Represents the activation function of the output layer.

8. A preoperative warning device for sacral tumors, characterized in that: The device comprises: An acquisition module, used to acquire clinical data information of a target user, a preoperative non-contrast sacral computed tomography image of a target user, multimodal data information of a target user, target diagnostic interval data information of a target user, a preset three-dimensional densely connected network model, a training sample set, and a validation sample set, wherein the multimodal data information of a target user includes T1-weighted imaging information of a target user, T2-weighted imaging information of a target user, and positron emission tomography-computed tomography information; A processing module is used to preprocess the training sample set to generate a target training sample set; process the preset three-dimensional densely connected network model based on the target training sample set and the verification sample set to generate a target three-dimensional densely connected network model; process the preoperative sacrum non-contrast computed tomography image of the target user to generate a three-dimensional region of interest image of the target user; process the multimodal data information of the target user and the target diagnostic interval data information of the target user to generate dynamic risk parameters of the target user; process the clinical data information of the target user to generate physiological attribute parameters of the target user; based on the target three-dimensional densely connected network model, process the preoperative sacrum non-contrast computed tomography image of the target user, the three-dimensional region of interest image of the target user, the dynamic risk parameters of the target user and the physiological attribute parameters of the target user to generate sacrum warning information, wherein the sacrum warning information is used to characterize the attribute information of the preoperative sacral lesions of the target user.

9. An electronic device, characterized in that: include: a first processor; and a memory for storing executable instructions of the first processor; Wherein, the first processor is configured to execute the preoperative early warning method for sacral tumors according to any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the second processor, the preoperative early warning method for sacral tumors according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Construction method, system and equipment of bone aging degree prediction model and medium

    CN118351120A

  • Early warning method for dangerous behavior of schizophrenia patient and related equipment

    CN118588291A

  • Method and apparatus for predicting immune checkpoint inhibitor-related thyroid dysfunction

    CN118800449A

  • Lung CT image correlation analysis method

    CN119313921A

  • Autoantibody Profiles in the Early Detection and Diagnosis of Cancer

    US20150198600A1