Method and system for multi-modal bone tumor classification based on blood biochemical text guidance

By combining blood biochemical text features with bone tumor images and pathological section images, a multimodal fusion feature space is constructed, which solves the problems of bone tumor image classification in existing technologies such as reliance on manual intervention, single data modality and insufficient information utilization, and achieves high-precision and explainable bone tumor classification.

CN120599389BActive Publication Date: 2025-10-14SOUTHERN MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511106114.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-10-14
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

Existing bone tumor imaging classification technologies rely on manual intervention and have poor model scalability, over-reliance on a single data modality, insufficient model interpretability, and insufficient utilization of clinical information, resulting in insufficient classification performance and generalization capabilities.

Method used

By screening abnormal blood biochemical markers, a large language model is used to construct a descriptive text rich in clinical logic and semantics, and blood biochemical text features are extracted. Combined with the image features of bone tumors and pathological sections, a trimodal feature space is constructed. The multimodal guided fusion layer is used to perform information cross-guided fusion to generate multimodal fusion features, and finally bone tumors are classified through a linear classification layer.

Benefits of technology

It significantly improves the accuracy and interpretability of bone tumor classification, enhances the robustness and flexibility of the system, enables stable and efficient classification even when multi-source data is incomplete, and builds a more comprehensive disease portrait of patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599389B_ABST
    Figure CN120599389B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal bone tumor classification method and system based on blood biochemical text guide, the method includes the following steps: obtaining multi-modal data;Blood biochemical markers are used to construct blood biochemical text guide factor;Determine whether the bone tumor image or pathological section image in multi-modal data is missing, when not missing, the bone tumor image structure feature or pathological section image feature is extracted, when missing, the compensation feature of missing modal data is generated, and the image feature corresponding to the compensation of missing modal data is extracted;Blood biochemical text guide factor is used to cross guide fusion on bone tumor image structure feature and pathological section image feature, and multi-modal fusion feature is obtained;The probability of each bone tumor category is calculated by linear classification layer to multi-modal fusion feature, and bone tumor classification result is obtained.The application realizes the effective guidance and supplement of non-image information to image feature, and improves the accuracy and explainability of bone tumor classification result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image classification, in particular to a multi-modal bone tumor classification method and system based on blood biochemical text guidance. BACKGROUND

[0002] The existing bone tumor image classification technology has the following defects:

[0003] (1) Relies on manual intervention and has poor model expansion: Many methods require clinicians to manually crop tumor regions, which not only consumes time and effort, but also introduces subjective bias from operators; some end-to-end methods that tightly couple detection and classification tasks, although they achieve automation, limit the independent optimization and flexible expansion of classification models;

[0004] (2) Over-reliance on a single data modality: Most studies rely only on clinical images (such as X-rays, CT), ignoring other valuable clinical information. This single-modality approach is insufficient in terms of classification performance and generalization ability when faced with complex tumor types, different anatomical locations, or variable clinical scenarios;

[0005] (3) Insufficient model interpretability: Most existing deep learning models are "black boxes" with opaque decision-making processes, making it difficult for clinicians to understand and trust the model's diagnostic basis, which greatly hinders its practical application and promotion in clinical diagnosis;

[0006] (4) Inadequate use of clinical information: Even a few studies that attempt to integrate clinical data (such as patient age) usually treat them as simple numerical features and directly concatenate them with image features, failing to fully explore the biological significance and pathophysiological relationships behind these data, resulting in the value of multi-modal information not being fully utilized. SUMMARY

[0007] In order to overcome the defects and shortcomings of the existing technology, the present invention provides a multimodal bone tumor classification method and system based on blood biochemical text guidance. The present invention screens abnormal blood biochemical markers, uses a large language model to construct a descriptive text rich in clinical logic and semantics, and then extracts it as a blood biochemical text guidance factor based on a text feature extraction network, thereby enhancing the text feature expression ability and improving the text feature effectiveness by using the text feature extraction network; at the same time, combined with the bone tumor image structure features and pathological section image features accurately cropped by the target detection network, a three-modal feature space containing image, pathological and biochemical information is constructed, and the blood biochemical text guidance factor is guided and fused with the two image features through the multimodal guidance fusion layer, thereby realizing the effective guidance and supplementation of image features by non-image information, solving the limitations of traditional methods such as the single information dimension and the failure to fully utilize multi-source clinical data for collaborative analysis, thereby improving the accuracy and interpretability of bone tumor classification results.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] The present invention provides a multimodal bone tumor classification method based on blood biochemical text guidance, comprising the following steps:

[0010] Acquire multimodal data, including blood biochemical markers, bone tumor images, and one or more pathological slide images;

[0011] Screening blood biochemical markers to obtain abnormal blood biochemical markers, constructing abnormal blood biochemical marker description text based on the large language model, extracting blood biochemical text features, and normalizing them to obtain blood biochemical text guidance factors;

[0012] Determine whether bone tumor images or pathological slice images in the multimodal data are missing. If they are not missing, perform feature extraction to obtain bone tumor image structural features or pathological slice image features. If they are missing, generate compensation features for the missing modal data based on blood biochemical text guidance factors and the remaining modal data to compensate for the image features of the corresponding missing modal data.

[0013] Based on the blood biochemical text guidance factors, the structural features of bone tumor images and the features of pathological slice images are cross-guided and fused to obtain multimodal fusion features;

[0014] The multimodal fusion features are used through a linear classification layer to calculate the probability of each bone tumor category and obtain the bone tumor classification result.

[0015] As a preferred technical solution, it also includes data preprocessing and data enhancement steps for bone tumor images;

[0016] Data preprocessing includes:

[0017] The bone tumor image is converted into a standardized RGB format, and the pixel value is scaled to a predetermined range through image normalization operation, or adjusted to zero mean and unit variance, and the contrast of the bone and soft tissue is enhanced based on gray scale transformation;

[0018] The data enhancement includes:

[0019] Random horizontal or vertical flipping, random rotation, adding edge, texture or contrast enhancement information to the set channel of the image by using channel enhancement technology.

[0020] As a preferred technical solution, the feature extraction is performed to obtain the bone tumor image structure feature, specifically including:

[0021] The bone tumor target detection network is constructed, and the bone tumor region positioning information is generated based on the bone tumor target detection network. The bone tumor image is cropped based on the bone tumor region positioning information to obtain a cropped bone tumor image, and the cropped bone tumor image only contains the bone tumor.

[0022] The channel attention residual network is constructed, and the cropped bone tumor image is input into the channel attention residual network for feature extraction to obtain the bone tumor image structure feature, which is represented as:

[0023] ;

[0024] Wherein, represents the bone tumor image structure feature, 、 、 respectively represent the first, kth and nth elements of the bone tumor image structure feature .

[0025] As a preferred technical solution, the bone tumor target detection network is constructed, specifically including:

[0026] The bone tumor target detection network is constructed based on the YOLOv5 network, and the bone tumor target detection network includes a backbone network, a neck network and a detection head.

[0027] The backbone network adopts the CSPDarknet network architecture, and is used for hierarchical extraction of features from the bone tumor image;

[0028] The neck network is used for fusing the features extracted hierarchically, and detecting bone tumors of different sizes and morphologies. The neck network includes a feature pyramid network and a path aggregation network. The feature pyramid network is used for passing high-level semantic information through a top-down path, and the path aggregation network is used for enhancing bottom-level positioning information through a bottom-up path.

[0029] The detection head outputs a bounding box including a bone tumor position, a target confidence score and a category probability, and filters out redundant candidate boxes through a non-maximum suppression post-processing algorithm to generate bone tumor region positioning information.

[0030] As a preferred technical solution, feature extraction is performed to obtain pathological section image features, specifically including:

[0031] A channel attention residual network is constructed, and the pathological section image is input into the channel attention residual network for feature extraction to obtain pathological section image features, denoted as:

[0032] ;

[0033] Wherein, denotes the pathological section image features, , , denote the first, kth and nth elements of the pathological section image features .

[0034] As a preferred technical solution, the channel attention residual network includes multiple attention residual groups, a global pooling layer and a linear layer, and each attention residual group is provided with different numbers of channel attention blocks, each channel attention block performs skip concatenation on the output and the input, the global pooling layer aggregates the output feature map of the last attention residual group in the spatial dimension, converts the feature map of each feature channel into a numerical value to obtain a reduced feature vector, and the reduced feature vector is output through the linear layer.

[0035] As a preferred technical solution, the compensation features of the missing modal data are generated based on the blood biochemical text guide factor and the remaining modal data, and the image features corresponding to the missing modal data are compensated, specifically including:

[0036] When the bone tumor image is missing, an image feature compensation generator is constructed, taking the blood biochemical text guide factor and the pathological section image features as inputs to generate image compensation features, denoted as:

[0037] ;

[0038] Wherein, denotes the image compensation features, denotes the image feature compensation generator, denotes the blood biochemical text guide factor, denotes the pathological section image features;

[0039] The image compensation features are taken as the bone tumor image structure features;

[0040] When the pathological section image is missing, a pathological feature compensation generator is constructed to generate pathological compensation features by taking blood biochemical text guide factors and bone tumor image structure features as inputs, and the pathological compensation features are represented as:

[0041] ;

[0042] wherein, represents the pathological compensation features, represents the pathological feature compensation generator, represents the bone tumor image structure features;

[0043] The pathological compensation features are used as pathological section image features.

[0044] As a preferred technical solution, the bone tumor image structure features and the pathological section image features are cross-guided and fused based on the blood biochemical text guide factors to obtain multi-modal fusion features, which specifically include:

[0045] The blood biochemical text guide factors and the bone tumor image structure features are multiplied after linear transformation to calculate an image guide score, which is represented as:

[0046] ;

[0047] wherein, represents the image guide score, represents the linear transformation, represents the blood biochemical text guide factors, represents the bone tumor image structure features, represents that two vectors are multiplied element by element, represents the data length of the blood biochemical text guide factors;

[0048] The image guide weight is calculated by a softmax function :

[0049] ;

[0050] The image guide weight is multiplied with the bone tumor image structure features after linear transformation to obtain image guide features after linear transformation, which is represented as:

[0051] ;

[0052] The blood biochemical text guide factors and the pathological section image features are multiplied after linear transformation to calculate a pathological guide score, which is represented as:

[0053] ;

[0054] wherein, Represents the characteristics of pathological slice images;

[0055] Calculate pathological guidance weights through the softmax function :

[0056] ;

[0057] Multiply the pathology guidance weight with the pathology slice image feature after linear transformation, and obtain the pathology guidance feature through linear transformation. :

[0058] ;

[0059] Image-guided features and pathology-guided features are spliced ​​into multimodal fusion features.

[0060] As a preferred technical solution, the multimodal fusion features are used to calculate the probability of each bone tumor category through the linear classification layer to obtain the bone tumor classification result, which is expressed as:

[0061] ;

[0062] ;

[0063] in, represents the bone tumor classification results, represents multimodal fusion features, represents matrix multiplication, represents the connection parameter matrix of the linear classification layer, 2n represents the feature length of the multimodal fusion feature, b represents the number of classification categories, represents the activation function, Represents the mth element of the multimodal fusion feature, Represents the connection parameter matrix The element in row m and column x, represents the probability of being classified as category x.

[0064] The present invention provides a multimodal bone tumor classification system based on blood biochemical text guidance, which is used to implement the above-mentioned multimodal bone tumor classification method based on blood biochemical text guidance. The system includes: a multimodal data acquisition module, a blood biochemical text guidance factor extraction module, a modal data missing judgment module, a feature extraction module, a feature compensation generator, a multimodal fusion feature construction module, and a classification module;

[0065] The multimodal data acquisition module is used to acquire multimodal data, including blood biochemical markers, and one or more of bone tumor images and pathological section images;

[0066] The blood biochemical text guide factor extraction module is configured to extract blood biochemical text guide factors, and specifically includes:

[0067] The abnormal blood biochemical markers are screened, the abnormal blood biochemical marker description text is constructed based on the large language model, the blood biochemical text features are obtained through feature extraction, and the blood biochemical text guide factors are obtained through normalization processing;

[0068] The modal data missing judgment module is configured to judge whether the bone tumor image or the pathological section image in the multi-modal data is missing, and when the modal data is not missing, the feature extraction module extracts the bone tumor image structure feature or the pathological section image feature, and when the modal data is missing, the feature compensation generator generates compensation features for the missing modal data based on the blood biochemical text guide factors and the remaining modal data to compensate for the image features of the corresponding missing modal data.

[0069] The multi-modal fusion feature construction module is configured to cross-guide and fuse the bone tumor image structure features and the pathological section image features based on the blood biochemical text guide factors to obtain multi-modal fusion features.

[0070] The classification module is configured to calculate the probability of each bone tumor category through a linear classification layer to obtain a bone tumor classification result.

[0071] Compared with the prior art, the present application has the following advantages and beneficial effects:

[0072] (1) The present application applies abnormal blood biochemical markers to bone tumor classification, and through a large language model and its construction, a description text rich in clinical logic and semantics is obtained, solving the technical problem of insufficient information utilization in traditional methods that process clinical test data as isolated numerical values. The semantic enhances the expression ability of the text features, and the text feature extraction network efficiently converts the description text into blood biochemical text guide factors, realizing accurate encoding and effective extraction of medical text information. The blood biochemical text guide factors are guided and fused with the two image features through a multi-modal guided fusion layer, realizing collaborative learning between image, pathology and biochemical information, solving the technical problem of insufficient multi-modal feature interaction, and finally significantly improving the accuracy and interpretability of bone tumor image classification.

[0073] (2) The present application constructs an image feature compensation generator and a pathological feature compensation generator, which can automatically generate high-quality compensation features based on existing modal information when a certain modal data (such as bone tumor image or pathological section image) is missing, realize effective compensation and intelligent reasoning for missing modal data, and still realize stable and efficient bone tumor classification analysis under the condition of incomplete multi-source data, greatly improving the robustness and application flexibility of the system.

[0074] (3) The application realizes automatic, high-precision and cutting of bone tumor regions based on a bone tumor target detection network, solves the defects of traditional methods that rely on manual cutting, are low in efficiency and strong in subjectivity, and separates the detection part and the classification part for training, so that the detection model and the classification model can be optimized and iterated independently, and the flexibility and scalability of the whole system are enhanced.

[0075] (4) The application constructs a three-modal diagnosis framework containing bone tumor images, pathological section images and blood biochemical markers, which solves the limitation of existing researches that mostly rely on single image data and cannot fully reflect the macro, micro and systemic biological characteristics of tumors, and through integrating information from different aspects, a more complete patient disease portrait is constructed, so that more robust and comprehensive judgments can be made on complex bone tumor cases.

[0076] (5) The multi-modal classification method proposed by the application has superior performance, and the experimental results show that the classification accuracy reaches 0.9107, the F1 score reaches 0.8741, and the AUC value reaches 0.9764, which has significant advantages compared with the prior art, fully proving that the application can significantly improve the accuracy and effectiveness of bone tumor image classification by introducing blood biochemical text guidance and multi-modal deep fusion. BRIEF DESCRIPTION OF DRAWINGS

[0077] Figure 1 is a flowchart of the multi-modal bone tumor classification method based on blood biochemical text guidance of the application;

[0078] Figure 2 is a flowchart of the generation of blood biochemical text guidance factors of the application;

[0079] Figure 3 is a flowchart of the positioning and cutting of bone tumor regions of bone tumor images of the application;

[0080] Figure 4 is a schematic diagram of the overall network framework of the channel attention residual network of the application;

[0081] Figure 5 is a flowchart of generating bone tumor classification results of the application. DETAILED DESCRIPTION

[0082] In order to make the purpose, technical scheme and advantages of the application clearer, the application will be further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the application and not to limit the application.

[0083] Example 1

[0084] AsFigure 1 As shown, this embodiment provides a multimodal bone tumor classification method based on blood biochemical text guidance, including the following steps:

[0085] S1: Acquire multimodal data, including blood biochemical markers, bone tumor images, and one or more pathological section images;

[0086] The modal input combinations of multimodal data include: blood biochemical markers + bone tumor images + pathological slide images (three modalities), blood biochemical markers + bone tumor images (two modalities), and blood biochemical markers + pathological slide images (two modalities).

[0087] S2: Inputting the blood biochemical markers into the abnormal marker screening module to obtain abnormal blood biochemical markers, specifically including:

[0088] Assume there are n1 blood biochemical markers. The information of each blood biochemical marker includes the marker name, test result, reference range, and unit. The abnormal marker screening module detects whether the test result of each blood biochemical marker is within the reference range. Blood biochemical markers that exceed the reference range are judged as abnormal blood biochemical markers, and finally n2 abnormal blood biochemical markers are obtained.

[0089] S3: Inputting the abnormal blood biochemical markers into a large language model to construct a description text of the abnormal blood biochemical markers. The large language model of this embodiment is preferably ChatGPT4.0;

[0090] Construct a large language model prompt for the text T describing abnormal blood biochemical markers: "You are an experienced orthopedic oncologist. Please generate a clinical analysis summary of no more than 100 words based on the following patient's blood biochemical test results. Please focus on indicators that may be related to bone tumors and analyze the potential clinical significance and interrelationships of these abnormal indicators. The test results are as follows: { "[Marker Name 1]": { "Test Result": [Result 1], "Unit": [Unit 1], "Reference Range": [Reference Range 1]}, "[Marker Name 2]": { "Test Result": [Result 2], "Unit": [Unit 2], "Reference Range": [Reference Range 2], ...}", where [Marker Name], [Test Result], [Unit], and [Reference Range] are replaced with the actual values ​​corresponding to each abnormal blood biochemical marker."

[0091] S4: As Figure 2 As shown, the abnormal blood biochemical marker description text is input into the text feature extraction network for feature extraction to obtain blood biochemical text features. The blood biochemical text features are normalized to obtain the blood biochemical text guidance factor, which is expressed as:

[0092] ;

[0093] wherein, represents a blood biochemical text guide factor, , , respectively represent the first, the kth and the nth element of the blood biochemical text guide factor , n is the data length of the blood biochemical text guide factor, preferably 128, the superscript in the formula is to distinguish that it is an element of the blood biochemical text guide factor;

[0094] In this embodiment, the text feature extraction network is preferably biobert-v1.1, and the training process uses the gradient descent method;

[0095] S5: judging whether the bone tumor image modality data in the multi-modal data is missing, if not, jumping to step S6 to process the bone tumor image to obtain bone tumor image structure features, if missing, jumping to step S7 to use an image feature compensation generator to generate image compensation features;

[0096] S6: as shown in the formula Figure 3 , data preprocessing and data enhancement are performed on the bone tumor image data, automatic recognition and positioning are realized by training a bone tumor target detection network to obtain bone tumor image structure features, and the specific steps include:

[0097] S61: obtaining bone tumor image data with labeled bone tumor regions, performing data preprocessing and data enhancement on the bone tumor image data, realizing automatic recognition and positioning of the bone tumor region by training a bone tumor target detection network, generating bone tumor region positioning information, and cutting the bone tumor image data based on the bone tumor region positioning information to obtain bone tumor image containing only bone tumor;

[0098] In this embodiment, the bone tumor image data has X-ray, CT, and nuclear magnetic resonance images, and the data preprocessing includes: converting the bone tumor image data into a standardized RGB format to ensure the consistency of the data format of the input network; through image normalization operation, the pixel value is scaled to a predetermined range [0, 1], or adjusted to zero mean and unit variance, so as to eliminate the intensity difference and artifacts caused by different devices and acquisition parameters; and according to the need, apply gray scale transformation to enhance the contrast of bone and soft tissue;

[0099] In this embodiment, the training data set is expanded based on data augmentation, which includes random horizontal or vertical flipping, random rotation, and adding edge, texture or contrast enhancement information to specific channels of the image using channel enhancement technology to highlight small or low-contrast tumor lesions;

[0100] In this embodiment, the bone tumor target detection network is preferably a YOLOv5 network model, which is used to identify and locate the bone tumor region, and the loss function is CIOU loss and confidence loss;

[0101] The specific structure of the bone tumor target detection network includes:

[0102] Its backbone network is based on the CSPDarknet architecture, which is used to extract features from the input bone tumor image in a hierarchical manner;

[0103] Its neck network combines the feature pyramid network and the path aggregation network structure, the feature pyramid network passes down the high-level semantic information through the top-down path, and the path aggregation network enhances the bottom-level positioning information through the bottom-up path, both of which work together to fuse multi-scale features, thereby effectively detecting bone tumors of different sizes and shapes;

[0104] Its detection head predicts on multiple feature map scales, outputs include bone tumor position bounding box, target confidence score and class probability, and through the non-maximum suppression post-processing algorithm, redundant candidate boxes are filtered out to obtain the final accurate bone tumor region positioning information.

[0105] S62: input the image containing only the bone tumor region into the channel attention residual network for feature extraction to obtain the bone tumor image structure feature, denoted as:

[0106] ;

[0107] Wherein, represents the bone tumor image structure feature, , , respectively represent the first, kth and nth elements of the bone tumor image structure feature , The superscript in is used to distinguish that it is an element of the bone tumor image structure feature;

[0108] As Figure 4As shown, the channel attention residual network is provided with 4 attention residual groups, a global pooling layer, and a linear layer, each attention residual group contains different numbers of channel attention blocks, 3, 4, 6, and 3 channel attention blocks respectively, a total of 16 channel attention blocks. Each channel attention block contains two convolutional layers, a channel attention layer, and a skip connection inside, and the channel attention layer is used to compress and activate the input and then multiply it with the input in the skip channel;

[0109] In the first attention residual group, there are 3 channel attention blocks, and the input of the first channel attention block is the image data to be processed. For each channel attention block in this group, the output is concatenated with the input in a skip manner to serve as the input of the next channel attention block. The output after processing by the first attention residual group is used as the input of the second attention residual group, which contains 4 channel attention blocks. The same skip concatenation method is used for processing in sequence, and so on. After processing by the third attention residual group (containing 6 channel attention blocks) and the fourth attention residual group (containing 3 channel attention blocks), the obtained feature map is output to the global pooling layer.

[0110] The global pooling layer aggregates the output feature map of the last attention residual group in the spatial dimension, converting each feature channel feature map into a numerical value. Specifically, the last layer feature map [2048, 7, 7] is converted into [2048, 1, 1] after global average pooling (average of each 7x7), and then expanded into

[2048] and sent to the fully connected layer to obtain [n], thereby reducing the three-dimensional feature map to a two-dimensional feature vector. Finally, the two-dimensional feature vector is input to the linear layer to obtain the final feature.

[0111] S7: When the bone tumor image structure feature is missing due to the absence of the bone tumor image, use the image feature compensation generator to generate image compensation features:

[0112] ;

[0113] Let compensate for the missing bone tumor image structure features;

[0114] S8: Determine whether the pathological section image is missing. If not, jump to step S9 to process the pathological section image to obtain the pathological section image feature. If missing, jump to step S10 to use the pathological feature compensation generator to generate pathological compensation features;

[0115] ​​​S9: input the pathological section image into the channel attention residual network to extract features, and obtain the pathological section image features, denoted as:

[0116] ;

[0117] wherein, denotes the pathological section image features, , , denote the first, the kth and the nth element of the pathological section image features, respectively, the superscript is used to distinguish that it is an element of the pathological section image features;

[0118] S10: when the pathological section image is missing, resulting in the pathological section image features being missing, use the pathological feature compensation generator to generate the pathological compensation features, using the blood biochemical text guide factor and the bone tumor image structure features as inputs:

[0119] ;

[0120] let compensate for the missing pathological section image features;

[0121] wherein, the image feature compensation generator has the same structure as the pathological feature compensation generator , and is preferably a variational autoencoder network VAE, and the loss function used in training is preferably MSE loss;

[0122] the obtained blood biochemical text guide factor , bone tumor image structure features and pathological section image features are input into step S11 for fusion classification;

[0123] S11: as shown in Figure 5 , the blood biochemical text guide factor is used to cross-guide and fuse the bone tumor image structure features and the pathological section image features through a multi-modal guide fusion layer to obtain multi-modal fusion features, and the multi-modal fusion features are input into a linear classification layer and a softmax function to obtain the bone tumor classification results, and the specific steps include:

[0124] the blood biochemical text guide factor and the bone tumor image structure features are multiplied after linear transformation to calculate the image guide score :

[0125] ;

[0126] wherein, denotes element-wise multiplication of two vectors, denotes a linear transformation;

[0127] The image-guided weight is calculated by a softmax function :

[0128] ;

[0129] The image-guided weight is multiplied with the linearly transformed bone tumor image structure feature I, and then linearly transformed to obtain an image-guided feature:

[0130] ;

[0131] wherein, denotes element-wise multiplication of two vectors;

[0132] In this embodiment, the blood biochemical text-guided factor is multiplied with the linearly transformed pathological section image feature to calculate a pathology-guided score:

[0133] ;

[0134] wherein, denotes element-wise multiplication of two vectors;

[0135] The pathology-guided weight is calculated by a softmax function:

[0136] ;

[0137] The pathology-guided weight is multiplied with the linearly transformed pathological section image feature , and then linearly transformed to obtain a pathology-guided feature:

[0138] ;

[0139] The image-guided feature and the pathology-guided feature are spliced into a multi-modal fusion feature , which is represented as:

[0140] ];

[0141] wherein, , , Represent multimodal fusion features respectively The 1st, uth and 2nth elements of , 2n represents the feature length of the multimodal fusion feature;

[0142] Multimodal fusion features Through the linear classification layer and softmax function, the probability of each bone tumor category is calculated to obtain the bone tumor classification result. The classification result is expressed as:

[0143] ;

[0144] ;

[0145] in, represents matrix multiplication, is the mth element of the multimodal fusion feature, Represents the connection parameter matrix of the linear classification layer, which is a 1×2n matrix multiplied by a 2n×b matrix to obtain a 1×b matrix. express The element in row m and column x, represents the activation function, b represents the number of classification categories, represents the probability of classification as category x;

[0146] This example finally obtains the classification results of benign, intermediate and malignant. The classification part uses cross entropy loss and the learning rate is set at 1e -3 to 2e -3 The detection evaluation method uses the average precision, accuracy and recall rate, and the classification evaluation method uses three evaluation indicators, namely accuracy, F1 score and AUC value.

[0147] In this embodiment, the bone tumor imaging data includes a detection task dataset and a classification task dataset;

[0148] The detection task dataset consists of 1,115 bone tumor images. Experienced experts annotated the tumor lesions with bounding boxes. Data augmentation techniques such as random flipping, rotation, and channel enhancement were used to expand the dataset to 2,179 images. The dataset was randomly divided into training, validation, and test sets in a ratio of 8:1:1.

[0149] The classification task dataset consisted of 38 patient cases, each containing three types of data: bone tumor images, pathology slide images, and blood biochemical markers. By pairing these three data sets, a total of 577 data sets were constructed, all of which were pathologically confirmed and labeled as benign, intermediate, or malignant. The dataset was also divided into training, validation, and test sets in an 8:1:1 ratio, ensuring that each class was evenly distributed across the subsets.

[0150] In this embodiment, the training is divided into four stages, the first stage only trains the bone tumor target detection network, the loss function during training is CIOU loss and confidence loss, the detection stage is initialized using YOLOv5-s pre-training weight, the backbone network is frozen for the first 30 cycles, only the detection head is trained, and all parameters are unfrozen for joint training for the last 1000 cycles; the second stage only trains the channel attention residual network, the text feature extraction network, the multi-modal guided fusion layer and the linear classification layer, the input data combination used is blood biochemical markers + bone tumor images + pathological section images, and the loss function during training is cross entropy loss; the third stage only trains the image feature compensation generator , the input data combination used is blood biochemical markers + pathological section images, and the loss function during training is MSE loss; the fourth stage only trains the pathological feature compensation generator , the input data combination used is blood biochemical markers + bone tumor images, and the loss function during training is MSE loss;

[0151] In this embodiment, the performance of the proposed multi-modal bone tumor classification method is comprehensively evaluated, and the specific results are as follows:

[0152] Bone tumor detection performance: on the test set, when the IoU threshold is 0.5, the average precision of the model reaches 79.25%, and when the confidence threshold is 0.4, the model achieves the best F1 score on the detection task, reaching 0.79, at this time the accuracy is 87.86% and the recall rate is 71.01%, indicating that the model has achieved a good balance between positioning accuracy and recall rate;

[0153] Bone tumor classification performance is as follows:

[0154] (1) Ablation experiment verification: as shown in Table 1, different data modalities are used for ablation experiments in this embodiment. The results show that the performance of a single modality (only bone tumor images or pathological section images) is limited. When two modalities are fused, the performance is significantly improved. When all three modalities (bone tumor images, pathological section images, and blood biochemical markers) are fused, the model achieves the best performance, with an accuracy of 0.9107, an F1 score of 0.8741, and an AUC value of 0.9764, which proves that each modality provides indispensable supplementary information for the final classification decision, especially the introduction of blood biochemical text guide factor, which plays a key role in improving the performance of the model.

[0155] Table 1 Comparison of ablation experiment results of different modalities in this embodiment

[0156]

[0157] (2) Comparison with other methods: As shown in Table 2, the proposed method is compared with other mainstream models (such as VGG + Transformer, Inception + XGBoost), and the results show that the proposed method is superior to the comparison models in all evaluation indicators, verifying the advancement and superiority of the method in the bone tumor classification task.

[0158] Table 2 Comparison table of experimental results of classification methods of the embodiment and other methods

[0159]

[0160] (3) As shown in Table 3, under the combination of "bone tumor image + blood biochemical marker" and "pathological section image + blood biochemical marker" input modalities, after introducing the feature compensation generator proposed in the embodiment, the performance indicators are obviously improved. Among them, in the clinical image and blood biochemical marker fusion mode, the accuracy is improved from 0.8021 to 0.825, the F1 score is improved from 0.7897 to 0.812, and the AUC is improved from 0.919 to 0.936; in the pathological section and blood biochemical marker mode, the accuracy is improved from 0.825 to 0.832, the F1 score is improved from 0.8025 to 0.819, and the AUC is improved from 0.9275 to 0.941. The above results show that the feature compensation generator can effectively alleviate the influence of partial modality missing on the model performance, enhance the robustness and generalization ability of the multi-modal fusion model, and fully verify the practical value and superiority of the compensation module.

[0161] Table 3 Ablation experiment result data table of the feature compensation generator of the embodiment

[0162]

[0163] Embodiment 2

[0164] The embodiment provides a multi-modal bone tumor classification system based on blood biochemical text guidance, which is used to realize the multi-modal bone tumor classification method based on blood biochemical text guidance in the above embodiment 1. The system comprises a multi-modal data acquisition module, a blood biochemical text guidance factor extraction module, a modality data missing judgment module, a feature extraction module, a feature compensation generator, a multi-modal fusion feature construction module, and a classification module.

[0165] In the embodiment, the multi-modal data acquisition module is used to acquire multi-modal data, including blood biochemical markers, and one or more of bone tumor images and pathological section images.

[0166] In the embodiment, the blood biochemical text guide factor extraction module is configured to extract blood biochemical text guide factors, and specifically includes the following steps:

[0167] The abnormal blood biochemical markers are screened, the abnormal blood biochemical marker description text is constructed based on the large language model, the blood biochemical text features are obtained through feature extraction, and the blood biochemical text guide factors are obtained through normalization processing;

[0168] In the embodiment, the modal data missing judgment module is configured to judge whether the bone tumor image or the pathological section image in the multi-modal data is missing. If not, the feature extraction module extracts the bone tumor image structure feature or the pathological section image feature. If missing, the feature compensation generator generates the compensation feature of the missing modal data based on the blood biochemical text guide factor and the remaining modal data, and compensates the image feature of the corresponding missing modal data.

[0169] In the embodiment, the multi-modal fusion feature construction module is configured to cross-guide and fuse the bone tumor image structure feature and the pathological section image feature based on the blood biochemical text guide factor, to obtain a multi-modal fusion feature.

[0170] In the embodiment, the classification module is configured to calculate the probability of each bone tumor category through a linear classification layer, to obtain a bone tumor classification result.

[0171] The above embodiments are the preferred embodiments of the present application, but the embodiments of the present application are not limited by the above embodiments, and any changes, modifications, substitutions, combinations, simplifications made without departing from the spirit and principles of the present application are equivalent replacement methods, and are all included in the protection scope of the present application.

Claims

1. A multimodal bone tumor classification method based on blood biochemical text guidance, characterized by: The steps include: Acquire multimodal data, including blood biochemical markers, bone tumor images, and one or more pathological slide images; Screening blood biochemical markers to obtain abnormal blood biochemical markers, constructing abnormal blood biochemical marker description text based on the large language model, extracting blood biochemical text features, and normalizing them to obtain blood biochemical text guidance factors; Determine whether bone tumor images or pathological slice images in the multimodal data are missing. If they are not missing, perform feature extraction to obtain bone tumor image structural features or pathological slice image features. If they are missing, generate compensation features for the missing modal data based on blood biochemical text guidance factors and the remaining modal data to compensate for the image features of the corresponding missing modal data. Based on the blood biochemical text guidance factors, the structural features of bone tumor images and pathological slice image features are cross-guided and fused to obtain multimodal fusion features, including: The blood biochemical text guidance factors and bone tumor image structural features were linearly transformed and multiplied to calculate the image guidance score, which is expressed as: ; in, represents the image-guided score, represents a linear transformation, Indicates blood biochemical text guide factor, Indicates the imaging structural characteristics of bone tumors. It means that two vectors are multiplied element by element. The data length indicating the blood biochemistry text guide factor; Calculate image guidance weights through the softmax function : ; Multiply the image guidance weight with the bone tumor image structure feature after linear transformation, and obtain the image guidance feature through linear transformation. , expressed as: ; The blood biochemical text guidance factor and the pathological section image feature are linearly transformed and multiplied to calculate the pathological guidance score, which is expressed as: ; in, Represents the characteristics of pathological slice images; Calculate pathological guidance weights through the softmax function : ; Multiply the pathology guidance weight with the pathology slice image feature after linear transformation, and obtain the pathology guidance feature through linear transformation. : ; Splicing image-guided features and pathology-guided features into multimodal fusion features; The multimodal fusion features are used through a linear classification layer to calculate the probability of each bone tumor category and obtain the bone tumor classification result.

2. The multimodal bone tumor classification method based on blood biochemical text guidance according to claim 1 is characterized in that: It also includes data preprocessing and data enhancement steps for bone tumor images; Data preprocessing includes: Bone tumor images are uniformly converted into a standardized RGB format. Image normalization is performed to scale pixel values ​​to a predetermined range or to zero mean and unit variance. The contrast between bone and soft tissue is enhanced based on grayscale transformation. Data augmentation includes: Randomly flip horizontally or vertically, randomly rotate, and use channel enhancement techniques to add edge, texture, or contrast enhancement information to set channels of the image.

3. The multimodal bone tumor classification method based on blood biochemical text guidance according to claim 1 is characterized in that: Feature extraction is performed to obtain the structural features of bone tumor images, including: constructing a bone tumor target detection network, generating bone tumor region positioning information based on the bone tumor target detection network, and performing image cropping on the bone tumor image based on the bone tumor region positioning information to obtain a cropped bone tumor image, wherein the cropped bone tumor image only contains the bone tumor; A channel attention residual network is constructed, and the cropped bone tumor image is input into the channel attention residual network for feature extraction to obtain the structural features of the bone tumor image, which is expressed as: ; in, Indicates the imaging structural characteristics of bone tumors. 、 、 Represents the imaging structural characteristics of bone tumors The first, kth, and nth elements of .

4. The multimodal bone tumor classification method based on blood biochemical text guidance according to claim 3 is characterized in that: The construction of the bone tumor target detection network specifically includes: Building a bone tumor target detection network based on the YOLOv5 network, wherein the bone tumor target detection network includes a backbone network, a neck network, and a detection head; The backbone network adopts the CSPDarknet network architecture to extract features from bone tumor images in layers; The neck network is used to fuse features extracted by layering and detect bone tumors of different sizes and shapes. The neck network includes a feature pyramid network and a path aggregation network. The feature pyramid network is used to transmit high-level semantic information through a top-down path, and the path aggregation network is used to enhance the underlying positioning information through a bottom-up path. The detection head output includes the bounding box of the bone tumor location, the target confidence score and the category probability, and filters out redundant candidate boxes through a non-maximum suppression post-processing algorithm to generate bone tumor area positioning information.

5. The multimodal bone tumor classification method based on blood biochemical text guidance according to claim 1 is characterized in that: Feature extraction is performed to obtain pathological slice image features, including: Construct a channel attention residual network, input the pathological slice image into the channel attention residual network for feature extraction, and obtain the pathological slice image features, which can be expressed as: ; in, Represents the pathological slice image features, 、 、 Represents the pathological slice image features The first, kth, and nth elements of .

6. The multimodal bone tumor classification method based on blood biochemical text guidance according to claim 3 or 5, characterized in that: The channel attention residual network includes multiple attention residual groups, a global pooling layer, and a linear layer. Each attention residual group is equipped with a different number of channel attention blocks. Each channel attention block jumps the output and input. The global pooling layer aggregates the output feature map of the last attention residual group in the spatial dimension, converts the feature map of each feature channel into a numerical value, and obtains a reduced-dimensional feature vector. The reduced-dimensional feature vector is passed through the linear layer to obtain the output feature.

7. The multimodal bone tumor classification method based on blood biochemical text guidance according to claim 1 is characterized in that: Generate compensatory features for missing modal data based on blood biochemical text guidance factors and remaining modal data to compensate for image features of the corresponding missing modal data, specifically including: When bone tumor images are missing, an image feature compensation generator is constructed, which takes blood biochemical text guidance factors and pathological slice image features as input to generate image compensation features, which can be expressed as: ; in, Indicates the image compensation feature, represents the image feature compensation generator, Indicates blood biochemical text guide factor, Represents the characteristics of pathological slice images; Using image compensation features as the structural features of bone tumor images; When the pathological section image is missing, a pathological feature compensation generator is constructed, which takes the blood biochemical text guidance factor and the bone tumor image structure feature as input to generate the pathological compensation feature, which is expressed as: ; in, Indicates pathological compensation characteristics, represents the pathological feature compensation generator, Indicates the imaging structural characteristics of bone tumors; The pathology compensation feature is used as the pathology slice image feature.

8. The multimodal bone tumor classification method based on blood biochemical text guidance according to claim 1 is characterized in that: The multimodal fusion features are passed through the linear classification layer to calculate the probability of each bone tumor category and obtain the bone tumor classification result, which is expressed as: ; ; in, represents the bone tumor classification results, represents multimodal fusion features, represents matrix multiplication, represents the connection parameter matrix of the linear classification layer, 2n represents the feature length of the multimodal fusion feature, b represents the number of classification categories, represents the activation function, Represents the mth element of the multimodal fusion feature, Represents the connection parameter matrix The element in row m and column x, represents the probability of being classified as category x.

9. A multimodal bone tumor classification system based on blood biochemical text guidance, characterized by: Used to implement the multimodal bone tumor classification method based on blood biochemical text guidance as described in any one of claims 1 to 8, the system comprises: a multimodal data acquisition module, a blood biochemical text guidance factor extraction module, a modal data missing judgment module, a feature extraction module, a feature compensation generator, a multimodal fusion feature construction module, and a classification module; The multimodal data acquisition module is used to acquire multimodal data, including blood biochemical markers, and one or more of bone tumor images and pathological section images; The blood biochemistry text guiding factor extraction module is used to extract blood biochemistry text guiding factors, specifically including: Screening blood biochemical markers to obtain abnormal blood biochemical markers, constructing abnormal blood biochemical marker description text based on the large language model, extracting blood biochemical text features, and normalizing them to obtain blood biochemical text guidance factors; The modal data missing judgment module is used to judge whether the bone tumor image or pathological section image in the multimodal data is missing. If it is not missing, the feature extraction module performs feature extraction to obtain the bone tumor image structure feature or the pathological section image feature. If it is missing, the feature compensation generator generates compensation features for the missing modal data based on the blood biochemical text guidance factor and the remaining modal data to compensate for the image features of the corresponding missing modal data; The multimodal fusion feature construction module is used to perform cross-guided fusion of bone tumor image structural features and pathological section image features based on blood biochemical text guidance factors to obtain multimodal fusion features; The classification module is used to calculate the probability of each bone tumor category through the linear classification layer of the multimodal fusion features to obtain the bone tumor classification result.

Citation Information

Patent Citations

  • Real-time semantic segmentation method and device for multidirectional feature refinement network, and medium

    CN117197446A

  • Heterogeneous multi-modal image classification method and device based on missing modality

    CN117911785A