Osteosarcoma-oriented modal adaptive multi-modal auxiliary diagnosis method and system

CN122842906APending Publication Date: 2026-09-29BUXIN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611324877.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-28
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供面向骨肉瘤辅助诊断的模态自适应多模态数据处理方法及系统,以解决现有技术难以适应临床诊疗中数据分期分批获得、无法在模态不全时进行有效推理,以及缺少对影像与病理证据进行联合校准的问题

Benefits of technology

[0052]首先,本发明构建了模态自适应推理架构,根据临床信息、数字放射影像和全视野病理切片的实际可用情况自动选择推理路径,在资料不全时仍可输出对应层级的诊断结果,解决了现有方案因输入模态固定而无法适配临床数据分期分批获得的问题;其次,通过临床特征向量生成门控信号对影像多尺度视觉特征进行残差调制,并配合病灶注意力图进行空间增强,使影像分析过程受到临床背景的有效引导,弥补了现有方案将临床信息与影像特征简单拼接、缺乏特征层面交互的不足;采用先分割定位肿瘤候选区域再进行图像块分类和置信度加权聚合的两阶段策略处理全视野病理切片,在降低计算负担的同时避免了无效区域对诊断结果的干扰;此外,当同一患者多张不同投照体位影像预测结果出现冲突时,引入年龄先验对影像临床协同评分进行校准,利用骨肉瘤的流行病学规律降低非高发年龄段的假阳性风险;在全模态齐备时,以验证集性能确定的权重动态融合影像临床协同评分与病理评分,模拟了临床三结合综合判断的诊断逻辑;同时,系统同步输出病灶注意力热图和病理高贡献图像块作为可解释性证据,便于医生复核与转诊沟通。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842906A_ABST
    Figure CN122842906A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of medical artificial intelligence and medical image processing, in particular to a modal adaptive multi-modal auxiliary diagnosis method and system for osteosarcoma, which comprises: receiving consultation data, identifying currently available data modalities, the data modalities including clinical information, digital radiographic images and whole field pathological sections; performing corresponding operations according to the identification results: if only containing clinical information, generating and outputting prompt information indicating insufficient data; if containing clinical information and digital radiographic images and not containing whole field pathological sections, performing image-clinical collaborative reasoning to obtain image-clinical collaborative scores; the present application can automatically select the adaptive reasoning path according to the currently available data modalities of the patient, can still output the diagnosis support of the corresponding level when the modalities are incomplete, can complete the dynamic fusion of image and pathological evidence when the whole modalities are complete, and can adapt to the clinical diagnosis needs of osteosarcoma at each stage from initial diagnosis to final diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of medical artificial intelligence and medical image processing technology, specifically to a modality-adaptive multimodal assisted diagnosis method and system for osteosarcoma. Background Technology

[0002] The diagnosis of osteosarcoma is highly specialized. Its imaging features overlap with those of osteomyelitis, giant cell tumor of bone, metastatic tumors, chondrogenic tumors, and other benign and malignant bone lesions. Relying on a single examination alone can easily lead to misdiagnosis. Clinically, it is usually necessary to make a comprehensive judgment by combining the patient's age, lesion location, disease course, digital radiography signs, and puncture or surgical pathology results. That is, it follows the diagnostic principle of combining clinical manifestations, imaging signs, and pathological evidence.

[0003] In real-world medical procedures, the aforementioned data are often not obtained all at once. Patients typically receive basic clinical information and digital radiographic images during the initial consultation, while full-field pathological slides can only be generated after a biopsy or surgery. Some medical institutions are able to perform imaging examinations but lack osteosarcoma specialists, and some cases have pathological slides but have not yet formed a unified opinion on imaging and pathology consultation. This reality of gradually acquiring data and having incomplete modalities is significantly different from the design assumptions of existing medical artificial intelligence methods.

[0004] Currently available intelligent diagnostic solutions for osteosarcoma can be broadly categorized into several types: one type is image-centric classification methods that combine the depth features of digital radiography, CT, or MRI with some clinical variables to differentiate between benign and malignant bone tumors. While these methods demonstrate the value of imaging features, the guiding role of clinical information in image analysis is limited, and they struggle to cover the entire diagnostic process after pathological evidence is incorporated. Another type is independent analysis systems based on digital pathological images. These systems perform deep learning modeling on full-view pathological slides to complete tumor assessment or prognosis prediction, but they typically do not involve synchronous reasoning with radiographic images and clinical information from the initial diagnosis stage. A third type is multimodal prediction models that use the fusion of radiomics and clinical variables to assess treatment response or prognosis. However, their application scenarios are mostly in the treatment stage rather than in the initial differential diagnosis, and they usually require complete and comprehensive input modalities.

[0005] Overall, most existing related protocols are designed and operated under fixed and complete data conditions, failing to fully consider the actual characteristics of data acquisition in stages and batches in the clinical diagnosis and treatment of osteosarcoma. When pathology has not yet been obtained and only imaging data is available, or when there is inconsistency between imaging and pathology results, it is difficult to form a stratified judgment that conforms to the clinical pathway, and there is also a lack of a mechanism for joint calibration of multi-source evidence. Summary of the Invention

[0006] The purpose of this invention is to provide a modality-adaptive multimodal data processing method and system for the auxiliary diagnosis of osteosarcoma, addressing the problems of existing technologies being unable to adapt to the staggered and batch-based acquisition of data in clinical diagnosis and treatment, unable to perform effective inference when modalities are incomplete, and lacking joint calibration of imaging and pathological evidence. This solution can automatically select an appropriate inference path and output corresponding levels of auxiliary diagnostic results based on the patient's currently available clinical information, digital radiographic images, and full-field pathological slides. When all modalities are complete, it achieves dynamic fusion of imaging-clinical co-scoring and pathological scoring, while providing interpretable evidence such as lesion attention heatmaps and high-contribution pathological image patches, thereby adapting to the actual diagnostic needs of osteosarcoma at each stage from initial diagnosis to definitive diagnosis.

[0007] To achieve the above objectives, in one respect, the present invention provides the following technical solution: a modality-adaptive multimodal assisted diagnostic method for osteosarcoma, comprising:

[0008] Receive medical data and identify currently available data modalities, including clinical information, digital radiographic images, and full-view pathological slides;

[0009] Based on the recognition results, perform the corresponding operation:

[0010] If only clinical information is included, generate and output a message indicating insufficient data;

[0011] If the image contains clinical information and digital radiographic images but not full-field pathological sections, perform image-clinical co-inference to obtain an image-clinical co-score.

[0012] If the pathological slides are included but digital radiographic images are not included, perform pathological single-modal inference to obtain a pathological score.

[0013] If clinical information, digital radiographic images, and full-view pathological slides are included simultaneously, the image-clinical co-inference and the pathological monomodal inference are executed in parallel, and the image-clinical co-inference score and the pathological score are fused to obtain a comprehensive diagnostic score.

[0014] Output auxiliary diagnostic results based on the aforementioned clinical co-score, pathological score, or comprehensive diagnostic score.

[0015] Preferably, the image-clinical collaborative reasoning includes:

[0016] The clinical information is encoded into a fixed-dimensional clinical feature vector;

[0017] Extract the multi-scale visual features F from the digital radiographic image;

[0018] The clinical feature vector is input into a gated generator consisting of a linear layer and a sigmoid function to obtain a gated signal G, and the modulated visual feature F1 is calculated according to the following formula:

[0019]

[0020] From the modulated visual features Generate a single-channel lesion attention map M, and calculate the enhanced visual features using the following formula. :

[0021]

[0022] For the enhanced visual features Global average pooling is performed to obtain a feature vector, which is then input into a fully connected classification layer to calculate and output the clinical collaborative score of the image.

[0023] Preferably, the image-clinical collaborative reasoning further includes:

[0024] When multiple digital radiographs of the same lesion site are taken from different projection angles for the same patient, the clinical co-positive probability S of each digital radiograph is calculated separately.

[0025] If multiple images of the same patient show both positive and negative results, it is determined that there is a conflict in the image perspective.

[0026] When an image view conflict occurs, the age field in the clinical information is retrieved. If the age is greater than a preset age threshold T, the image-clinical co-positive probability S is lowered according to the following formula. The lowered image-clinical co-positive probability is... As the aforementioned clinical co-scoring of images:

[0027]

[0028] Where λ is the preset calibration coefficient and age is the patient's age.

[0029] Preferably, the pathological single-modal reasoning includes:

[0030] The full-view pathological section was segmented to obtain multiple image blocks;

[0031] A pre-trained segmentation network is used to identify tumor regions in each image patch, and candidate image patches with a tumor probability higher than a threshold are retained.

[0032] The candidate image blocks are input into a pathological image classifier to obtain the positive probability of each candidate image block;

[0033] Calculate the confidence score for each candidate image patch, select the top K candidate image patches from highest to lowest confidence score, and calculate the pathological score using the following formula:

[0034]

[0035] in, For pathological scoring, Let be the positive probability of the i-th candidate image patch. Let be the confidence score of the i-th candidate image patch, and TopK is the set of indices for the top K candidate image patches selected from highest to lowest confidence, where K is a preset positive integer. This is a preset smoothing constant.

[0036] Preferably, the step of fusing the imaging clinical co-scoring score with the pathological score to obtain a comprehensive diagnostic score includes:

[0037] According to the formula Calculate the comprehensive diagnostic score, where, For comprehensive diagnostic scoring, The clinical co-scoring of the images. The pathological score is given by α, which is the fusion weight. The value of α is predetermined based on the performance of the validation set.

[0038] Preferably, the auxiliary diagnostic results further include: the lesion attention heatmap corresponding to the digital radiographic image, and the top K candidate image blocks selected from high to low confidence levels in the full-field pathological slices and their spatial locations.

[0039] Preferably, the step of using a pre-trained segmentation network to identify tumor regions in each image patch and retaining candidate image patches with a tumor probability higher than a threshold includes:

[0040] The full-view pathological section is sliced ​​into multiple image blocks by sliding window according to preset magnification and size rules;

[0041] Each image patch is input into a pre-trained segmentation network, which outputs a tumor probability mask. The training process of the segmentation network includes: using the tumor region boundary as a pixel-level label, performing data augmentation using a staining enhancement method based on hematoxylin-eosin color deconvolution, and optimizing the loss function using a weighted sum of cross-entropy loss with class weights and spatial weights and Dess loss.

[0042] Based on the tumor probability mask, combined with the saturation and color ratio rules based on the hue-saturation-brightness color space and the Sobel gradient rule, blank backgrounds, handwriting marks, fat and normal tissue areas are removed, and the remaining areas are used as the candidate image blocks.

[0043] On the other hand, the present invention provides a modality-adaptive multimodal data processing system for the auxiliary diagnosis of osteosarcoma, comprising:

[0044] The modality recognition module is used to receive medical data and identify the currently available data modalities, including clinical information, digital radiographic images, and full-view pathological slides.

[0045] The image-clinical collaboration module is used to perform image-clinical collaboration reasoning and output an image-clinical collaboration score when the available modal contains clinical information and digital radiographic images but does not contain full-field pathological slices.

[0046] The histopathology analysis module is used to perform pathological single-modal inference and output a pathological score when the available modal includes full-field pathological slides but does not include digital radiographic images.

[0047] The modal adaptive decision aggregation module is used to perform corresponding operations based on the recognition results: when only clinical information is included, a data insufficiency prompt is generated; when clinical information, digital radiographic images and full-view pathological slides are included simultaneously, the image clinical collaboration module and the histopathology analysis module are called in parallel, and the image clinical collaboration score and the pathological score are merged into a comprehensive diagnostic score.

[0048] The results output module is used to output auxiliary diagnostic results based on the image clinical co-scoring score, pathological score, or comprehensive diagnostic score.

[0049] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the modality-adaptive multimodal assisted diagnostic method for osteosarcoma as described above.

[0050] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the modality-adaptive multimodal assisted diagnostic method for osteosarcoma as described above.

[0051] Compared with the prior art, the beneficial effects of the present invention are:

[0052] First, this invention constructs a modality-adaptive inference architecture that automatically selects the inference path based on the actual availability of clinical information, digital radiographic images, and full-field pathological slides. Even with incomplete data, it can still output diagnostic results at the corresponding level, solving the problem of existing solutions being unable to adapt to the staging and batching of clinical data due to fixed input modalities. Second, it generates gating signals using clinical feature vectors to perform residual modulation on multi-scale visual features of images, and uses lesion attention maps for spatial enhancement, effectively guiding the image analysis process with the clinical context. This overcomes the shortcomings of existing solutions that simply stitch clinical information and image features together and lack feature-level interaction. Finally, it employs a method of first segmenting and locating tumor candidate regions before image processing. A two-stage strategy of block classification and confidence-weighted aggregation is used to process full-view pathological slices, reducing computational burden while avoiding interference from invalid regions in diagnostic results. Furthermore, when conflicting prediction results occur from multiple images of the same patient taken from different projection positions, an age prior is introduced to calibrate the image-clinical synergy score, leveraging the epidemiological patterns of osteosarcoma to reduce the risk of false positives in non-high-incidence age groups. With all modalities available, the image-clinical synergy score and pathological score are dynamically fused using weights determined by validation set performance, simulating the diagnostic logic of a comprehensive clinical three-way judgment. Simultaneously, the system outputs lesion attention heatmaps and high-contribution pathological image blocks as interpretable evidence, facilitating physician review and referral communication. Attached Figure Description

[0053] Figure 1 This is a schematic diagram of the method flow provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the system architecture provided for an embodiment of the present invention. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0055] Figure 1 This is a flowchart illustrating a modality-adaptive multimodal data processing method for auxiliary diagnosis of osteosarcoma, provided by an embodiment of the present invention. Figure 1As shown, the overall process includes: Step S1, receiving medical data and identifying the currently available data modalities, which include clinical information, digital radiographic images, and full-field pathological slides; Step S2, performing corresponding operations based on the identification results: if only clinical information is included, generating and outputting a prompt message indicating insufficient data; if clinical information and digital radiographic images are included but full-field pathological slides are not included, performing image-clinical collaborative reasoning to obtain an image-clinical collaborative score; if full-field pathological slides are included but digital radiographic images are not included, performing pathological single-modality reasoning to obtain a pathological score; if clinical information, digital radiographic images, and full-field pathological slides are included simultaneously, performing the image-clinical collaborative reasoning and the pathological single-modality reasoning in parallel, and fusing the image-clinical collaborative score with the pathological score to obtain a comprehensive diagnostic score; Step S3, outputting auxiliary diagnostic results based on the image-clinical collaborative score, pathological score, or comprehensive diagnostic score.

[0056] Figure 2 This is a schematic diagram of the system architecture provided in an embodiment of the present invention, such as... Figure 2 As shown, this embodiment provides a modality-adaptive multimodal data processing system for the auxiliary diagnosis of osteosarcoma. The system includes a modality recognition module, an image-clinical collaboration module, a histopathology analysis module, a modality-adaptive decision aggregation module, and a result output module.

[0057] The modality recognition module is used to receive medical data and identify the currently available data modalities, which include clinical information, digital radiographic images, and full-view pathological slides. In actual clinical scenarios, patients usually receive basic clinical information and digital radiographic images during the initial consultation, while full-view pathological slides can only be generated after a biopsy. Therefore, the medical data received by the system may only contain one or two of these modalities. The modality recognition module automatically determines the currently available data modality combination by parsing the fields and file types of the input data.

[0058] The image-clinical collaboration module is used to perform image-clinical collaborative reasoning and output an image-clinical collaborative score when the available modal contains clinical information and digital radiographic images but does not contain full-field pathological sections. This module uses clinical information to guide the image feature extraction process, enabling image analysis to combine the patient's age, gender, disease site and other clinical background, thereby more accurately focusing on the diagnostically relevant areas.

[0059] The histopathology analysis module is used to perform pathological single-modal inference and output pathological scores when the available modal includes full-view pathological sections but not digital radiographic images. This module adopts a two-stage strategy of localization followed by classification. First, tumor-related candidate regions are screened from full-view pathological sections, and then microscopic classification and patient-level aggregation are performed.

[0060] The modality adaptive decision aggregation module is used to perform corresponding operations based on the recognition results of the modality recognition module. When only clinical information is included, a data insufficiency prompt is generated to avoid misusing simple clinical epidemiological priors as independent diagnostic results. When clinical information, digital radiographic images, and full-field pathological slides are included simultaneously, the imaging clinical collaboration module and the histopathology analysis module are called in parallel, and the imaging clinical collaboration score and the pathology score are integrated into a comprehensive diagnostic score.

[0061] The results output module is used to output auxiliary diagnostic results based on clinical co-scores, pathological scores, or comprehensive diagnostic scores. The auxiliary diagnostic results also include lesion attention heatmaps corresponding to digital radiographic images, as well as the top K candidate image blocks and their spatial locations selected from high to low confidence in the full-field pathological slices, which facilitates physician review and referral communication.

[0062] The following is a detailed explanation of the specific implementation of image-clinical collaborative reasoning:

[0063] The image-clinical collaborative reasoning process first encodes clinical information. The system reads structured fields such as the patient's gender, age, and disease location, and generates structured text according to a preset template. For example, for a 25-year-old male patient with a distal femoral lesion, the generated text is "Male, 25 years old, disease location is distal femoral." If a field is missing, such as the disease location not being recorded, the generated text is "Male, 25 years old, disease location missing," meaning the missing marker for that field is retained without forcibly filling in the clinical facts. The structured text is then input into a pre-trained text encoder to extract its global semantic representation. This text encoder can use the transformer-based bidirectional encoder representation model BERT, or it can be replaced with other pre-trained language models, multilayer perceptron encoders, or tabular feature encoders. The representation vector output by the text encoder is passed through a projection network consisting of fully connected layers, layer normalization, and Gaussian error linear unit activation functions, and mapped to a fixed-dimensional 256-dimensional clinical feature vector C. This clinical feature vector C serves as the gating input for image branches and provides a structured basis for subsequent multi-view conflict detection and age prior calibration.

[0064] Next, multi-scale visual feature extraction is performed. The system preprocesses digital radiographic images to a preset resolution of 384×384 pixels. Considering that digital radiographic images are usually grayscale images, the network input channel is adjusted to a single channel. The preprocessed image is input into a visual backbone network, which can use a sliding window transformer network (SwinTransformer) or be replaced by a convolutional neural network, a visual transformer, or other medical image feature extraction networks. The system extracts shallow and deep features from different levels of the visual backbone network. Shallow features contain detailed information such as edges and textures, while deep features contain semantic-level global information. The deep features are upsampled to the same spatial scale as the shallow features through bilinear interpolation, then concatenated along the channel dimension, and finally fused through a 1×1 convolutional layer to obtain the multi-scale visual feature F. The dimensions of the multi-scale visual feature F are C1×H×W, where C1 is the number of channels, and H and W are the spatial dimensions.

[0065] Then, clinical gating modulation is performed. The system inputs the clinical feature vector C into a gating generator composed of a linear layer and a sigmoid activation function. The linear layer maps the clinical feature vector C into a vector with the same number of channels as the multi-scale visual feature F. The sigmoid function compresses each element of this vector to between 0 and 1, obtaining the gating signal G. The gating signal G has a shape of C1×1×1 and is multiplied channel-by-channel with the multi-scale visual feature F through a broadcast mechanism. The modulated visual features... Calculate using the following formula:

[0066] The calculation method uses residual multiplication modulation, where F×G represents scaling each channel of the multi-scale visual feature F by the gate signal G, and the retained F terms ensure that the original visual information is not completely covered. This residual structure enables clinical information to enhance the visual feature channels that are relevant to diagnosis, while maintaining the original response of irrelevant channels.

[0067] Next, lesion attention enhancement is performed. The system sets up an independent segmentation branch within the classification network. After clinical gating modulation, the visual features F1 are input to this segmentation branch. The segmentation branch consists of several convolutional layers, batch normalization layers, and the ReLU linear rectified activation function, ultimately outputting a single-channel feature-level logistic graph. This logistic graph is then processed by the Sigmoid activation function to generate a lesion attention map M with a feature size ranging from 0 to 1. The classification features are spatially enhanced according to the following formula: In this residual space enhancement method, the factor (1+M) amplifies the features within the attention coverage area, while the features in the uncovered area retain their original values, thus avoiding the complete masking of key background or boundary features due to inaccurate early segmentation.

[0068] After obtaining enhanced visual features Subsequently, a clinical co-score of the imaging was calculated. Specifically, post-visual features were enhanced. The input is a global average pooling layer, which performs average compression along the spatial dimension to obtain a feature vector in the channel dimension. This feature vector is then passed through a fully connected layer and a sigmoid activation function to output a scalar value between 0 and 1, which is the clinical co-score of the digital radiograph. When there are multiple digital radiographs of the same lesion site for the same patient at different projection angles, each digital radiograph independently performs the entire process of clinical information encoding, multi-scale visual feature extraction, clinical gating modulation, and lesion attention enhancement to obtain its own score.

[0069] During the training phase, to align with the ground truth annotations for calculating the segmentation loss, the system uses bilinear interpolation to upsample the logistic graph of the aforementioned feature sizes to the original resolution of the input image (384×384 pixels), and then applies the Sigmoid activation function again to generate the final predicted segmentation mask. The annotation data required for training this segmentation branch is a binary lesion mask, which is obtained by having doctors use annotation tools to draw polygonal outlines of lesion regions in digital radiographs and exporting them in GeoJSON format. The system reads the polygon coordinates from the GeoJSON file, combines them with the image's edge-filling strategy, and draws and generates a ground truth mask aligned with the input image size. To improve the model's robustness, the system supports converting the precise polygonal mask into a bounding box mask with random perturbations (including random scaling and translation) during training. The loss function of the segmentation branch is a weighted sum of the binary cross-entropy loss and the Dice loss, which guides the parameter updates of the segmentation branch.

[0070] When multiple digital radiographs of the same lesion site from different projection angles are available for the same patient, the system also performs multi-view conflict detection and age prior calibration. The system performs the aforementioned steps on each digital radiograph to obtain the clinical co-positive probability S for each digital radiograph. If both positive and negative results are found in the classification results of multiple images of the same patient, the system determines that an image view conflict has occurred. Image view conflicts may arise from the different degrees of lesion display under different projection angles, or from degenerative changes, metastatic lesions, etc., presenting similar signs to osteosarcoma under a single view. When an image view conflict is detected, the system obtains the age field from the clinical information. If the patient's age is greater than the preset age threshold T, age prior calibration is triggered. Osteosarcoma has a clear epidemiological characteristic and is highly prevalent in adolescents. When middle-aged and elderly patients present with similar imaging signs, it is more likely to be degenerative changes, bone metastases, or other non-osteosarcoma lesions. Based on this medical prior, the system lowers the clinical co-positive probability S.

[0071] In a preferred embodiment, the calibration formula adopts an exponential decay form:

[0072] ;

[0073] Where S represents the positive probability of clinical co-occurrence in imaging before calibration. Here, 'age' represents the calibrated probability, 't' is the patient's age, 'T' is a preset age threshold, typically determined through cross-validation on the training set within the range of 30 to 35 years, and 'λ' is a preset calibration coefficient, ranging from 0 to 1, determined based on validation set performance or preset medical rules. When the patient's age does not exceed the threshold T... The probability remains unchanged when the patient's age exceeds the threshold T by more than zero; the down-regulation magnitude is greater when the patient's age exceeds the threshold T by more than zero; the calibration coefficient λ controls the intensity of the down-regulation; this prior calibration mechanism is not limited to the above exponential decay form, and in actual implementation it can also be replaced by piecewise function, continuous decay function, lookup table function or calibration model learned through additional validation set; the prior variable is not limited to age, and can also be extended to clinical information such as the site of onset, duration of symptoms, laboratory indicators or previous tumor history.

[0074] The following is a detailed explanation of the specific implementation of pathological single-modal inference. To address the issues of large full-field pathological slide sizes and numerous invalid regions, a two-stage scheme is adopted for pathological single-modal inference.

[0075] The first stage is tumor region localization. The system performs sliding window segmentation on the full-field pathological slide according to preset magnification and size rules to obtain multiple image blocks. In a preferred embodiment, for the original pathological slide scanned under 40x objective lens and physical resolution of 0.25 μm / pixel, a fixed window size of 256×256 pixels is used for sliding segmentation. The image blocks obtained from the segmentation are input into a pre-trained segmentation network, which outputs a pixel-level tumor probability mask. The segmentation network can adopt the U-Net structure, or it can be replaced by U-Net++, DeepLab, MaskR-CNN, or other network models that can output pixel-level segmentation results.

[0076] The training process of the segmentation network is as follows: Training data is generated by pathologists using professional software to precisely delineate the boundaries of tumor regions in full-view pathological sections. The system then generates grayscale mask images with the same size as the image blocks as pixel-level annotations. For data augmentation, the system employs a staining enhancement method based on hematoxylin-eosin (H&E) color deconvolution. By perturbing and recombining the H&E staining vectors, the system simulates color changes in pathological sections under different staining conditions, enhancing the model's robustness to staining differences. Regarding the loss function, the system uses a weighted sum of cross-entropy loss with class and spatial weights and Dessian loss. Class weights are used to alleviate the class imbalance between tumor and background regions, while spatial weights emphasize pixels in the tumor boundary region, making the network focus more on the interface between tumor and normal tissue. During the inference phase, the segmentation network outputs the probability value of each pixel belonging to the tumor region, forming a tumor probability mask of the same size as the input image patch. Based on this tumor probability mask, the system first filters out regions with tumor probabilities higher than a preset threshold. To further improve the filtering accuracy, the system also combines heuristic rules based on the hue-saturation-brightness HSV color space and the Sobel gradient rule: low saturation regions usually correspond to blank backgrounds, specific color ratio ranges correspond to densely populated regions of cell nuclei with deep hematoxylin staining, and low Sobel gradient regions correspond to flat, textureless backgrounds or adipose tissue. Combining pixel-level tumor probabilities and the above image heuristic rules, the system accurately removes blank backgrounds, blue handwriting markings, adipose tissue, normal tissue, and other low-information regions, retaining the remaining regions as candidate image patches.

[0077] The second stage involves micro-level classification and patient-level aggregation. Candidate image patches obtained in the first stage are input into an independent pathological image classifier to obtain the positive probability of each candidate image patch. The pathological image classifier can use a sliding window transformer network (SwinTransformer), or it can be replaced with a convolutional neural network or other suitable pathological image classification networks.

[0078] For each candidate image patch, the system calculates its confidence score using the following formula: , Let be the probability that the i-th candidate image patch is predicted as positive by the pathological image classifier; the confidence score measures the certainty of the model's judgment on that image patch: when When the confidence level is close to 0 or 1, it is close to 0.5, indicating that the model's judgment is clear; when... When the confidence level is close to 0.5, it is close to 0, indicating that the model has difficulty in making a judgment. The system sorts all candidate image patches from high to low confidence and selects the top K candidate image patches from high to low confidence, where K is a preset positive integer, for example, K=50. The index set of these K candidate image patches is denoted as TopK. The pathological score is calculated according to the following formula:

[0079]

[0080] in, For pathological scoring, Let be the positive probability of the i-th candidate image patch. Let be the confidence score of the i-th candidate image patch. The preset smoothing constant typically takes the value of [value missing]. This is used to prevent the denominator from being zero. The aggregation formula uses confidence as the weight to perform a weighted average of the positive probability. Image patches with high confidence have a larger proportion in the score, while image patches that the model judges to be blurry contribute very little to the score, so that the final score is dominated by the region that the model judges to be the most clear. This aggregation method is not limited to the above-mentioned Top-K confidence weighting. In actual implementation, it can also be replaced by attention multi-instance learning, max pooling, quantile pooling, Bayesian aggregation, or learning-based patient-level aggregator.

[0081] The following is a detailed explanation of the specific implementation of modality adaptive decision aggregation. The system first parses the input data through the modality recognition module to determine whether clinical information, digital radiographic images, and full-view pathological slides are available, and triggers different inference branches based on the recognition results.

[0082] When only clinical information is available, the system does not perform inference calculations or make a definitive diagnosis due to the lack of direct radiographic or pathological visual evidence. Instead, it directly generates and outputs a message indicating insufficient data, such as "Insufficient input data at present, unable to perform intelligent assisted diagnosis of osteosarcoma. Please supplement with digital radiographic images or full-field pathological slides." This avoids misusing simple clinical epidemiological priors as independent diagnostic results.

[0083] When the system includes clinical information and digital radiographic images but not full-field pathological sections, it invokes the imaging-clinical collaboration module, executes the aforementioned imaging-clinical collaboration reasoning steps, and outputs an imaging-clinical collaboration score. .

[0084] When the system includes full-field pathological sections but not digital radiographic images, it invokes the histopathology analysis module, executes the aforementioned pathological single-modality inference steps, and outputs a pathological score. .

[0085] When clinical information, digital radiographic images, and full-field pathological slides are simultaneously included, the system concurrently calls the imaging-clinical collaboration module and the histopathology analysis module to obtain the imaging-clinical collaboration score. and pathology score The results are then processed through full-modal fusion; the comprehensive diagnostic score is calculated using the following formula:

[0086]

[0087] in, For comprehensive diagnostic scoring, The clinical co-scoring of the images. For the pathological score, α is the fusion weight, which ranges from 0 to 1. Its value is pre-determined by performing a grid search within the range of 0.05 to 0.95 with a preset step size of 0.05, aiming to maximize the area under the receiver operating characteristic (AUC) curve of the internal validation set. If age prior calibration is triggered during the image-clinical collaborative inference process, i.e., multi-view image conflict is detected and the patient's age exceeds a preset threshold, then the calibrated probability S' is first used as the new... Then, substituting the above fusion formula, the system outputs a positive or negative auxiliary judgment result based on a preset binary classification threshold of 0.5, and provides continuous diagnostic probability values ​​and interpretable evidence. Interpretable evidence includes the lesion attention heatmap corresponding to the digital radiographic image, and the top K candidate image blocks selected from the full-field pathological slices according to confidence level from high to low, as well as their spatial coordinates in the original slice.

[0088] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A modality-adaptive multimodal assisted diagnostic method for osteosarcoma, characterized in that, include: Receive medical data and identify currently available data modalities, including clinical information, digital radiographic images, and full-view pathological slides; Based on the recognition results, perform the corresponding operation: If only clinical information is included, generate and output a message indicating insufficient data; If the image contains clinical information and digital radiographic images but not full-field pathological sections, perform image-clinical co-inference to obtain an image-clinical co-score. If the pathological slides are included but digital radiographic images are not included, perform pathological single-modal inference to obtain a pathological score. If clinical information, digital radiographic images, and full-view pathological slides are included simultaneously, the image-clinical co-inference and the pathological monomodal inference are executed in parallel, and the image-clinical co-inference score and the pathological score are fused to obtain a comprehensive diagnostic score. Output auxiliary diagnostic results based on the aforementioned clinical co-score, pathological score, or comprehensive diagnostic score.

2. The modality-adaptive multimodal assisted diagnostic method for osteosarcoma according to claim 1, characterized in that, The image-clinical collaborative reasoning includes: The clinical information is encoded into a fixed-dimensional clinical feature vector; Extract the multi-scale visual features F from the digital radiographic image; The clinical feature vector is input into a gated generator consisting of a linear layer and a sigmoid function to obtain a gated signal G, and the modulated visual feature F1 is calculated according to the following formula: ; From the modulated visual features Generate a single-channel lesion attention map M, and calculate the enhanced visual features using the following formula. : ; For the enhanced visual features Global average pooling is performed to obtain a feature vector, which is then input into a fully connected classification layer to calculate and output the clinical collaborative score of the image.

3. The modality-adaptive multimodal assisted diagnostic method for osteosarcoma according to claim 2, characterized in that, The image-clinical co-inference also includes: When multiple digital radiographs of the same lesion site are taken from different projection angles for the same patient, the clinical co-positive probability S of each digital radiograph is calculated separately. If multiple images of the same patient show both positive and negative results, it is determined that there is a conflict in the image perspective. When an image view conflict occurs, the age field in the clinical information is retrieved. If the age is greater than a preset age threshold T, the image-clinical co-positive probability S is lowered according to the following formula. The lowered image-clinical co-positive probability is... As the aforementioned clinical co-scoring of images: ; Where λ is the preset calibration coefficient and age is the patient's age.

4. The modality-adaptive multimodal assisted diagnostic method for osteosarcoma according to claim 1, characterized in that, The pathological single-modal reasoning includes: The full-view pathological section was segmented to obtain multiple image blocks; A pre-trained segmentation network is used to identify tumor regions in each image patch, and candidate image patches with a tumor probability higher than a threshold are retained. The candidate image blocks are input into a pathological image classifier to obtain the positive probability of each candidate image block; Calculate the confidence score for each candidate image patch, select the top K candidate image patches from highest to lowest confidence score, and calculate the pathological score using the following formula: ; in, For pathological scoring, Let be the positive probability of the i-th candidate image patch. Let be the confidence score of the i-th candidate image patch, and TopK is the set of indices for the top K candidate image patches selected from highest to lowest confidence, where K is a preset positive integer. This is a preset smoothing constant.

5. The modality-adaptive multimodal assisted diagnostic method for osteosarcoma according to claim 1, characterized in that, The process of fusing the imaging-clinical synergistic score with the pathological score to obtain a comprehensive diagnostic score includes: According to the formula Calculate the comprehensive diagnostic score, where, For comprehensive diagnostic scoring, The clinical co-scoring of the images. The pathological score is given by α, where α is the fusion weight.

6. The modality-adaptive multimodal assisted diagnostic method for osteosarcoma according to claim 1, characterized in that, The auxiliary diagnostic results also include: the lesion attention heatmap corresponding to the digital radiographic image, and the top K candidate image blocks selected from the full-field pathological slices according to confidence level from high to low, and their spatial locations.

7. The modality-adaptive multimodal assisted diagnostic method for osteosarcoma according to claim 4, characterized in that, The process of using a pre-trained segmentation network to identify tumor regions in each image patch and retaining candidate image patches with a tumor probability higher than a threshold includes: The full-view pathological section is sliced ​​into multiple image blocks by sliding window according to preset magnification and size rules; Each image patch is input into a pre-trained segmentation network, which outputs a tumor probability mask. The training process of the segmentation network includes: using the tumor region boundary as a pixel-level label, performing data augmentation using a staining enhancement method based on hematoxylin-eosin color deconvolution, and optimizing the loss function using a weighted sum of cross-entropy loss with class weights and spatial weights and Dess loss. Based on the tumor probability mask, combined with the saturation and color ratio rules based on the hue-saturation-brightness color space and the Sobel gradient rule, blank backgrounds, handwriting marks, fat and normal tissue areas are removed, and the remaining areas are used as the candidate image blocks.

8. A modality-adaptive multimodal assisted diagnostic system for osteosarcoma, characterized in that, include: The modality recognition module is used to receive medical data and identify the currently available data modalities, including clinical information, digital radiographic images, and full-view pathological slides. The image-clinical collaboration module is used to perform image-clinical collaboration reasoning and output an image-clinical collaboration score when the available modal contains clinical information and digital radiographic images but does not contain full-field pathological slices. The histopathology analysis module is used to perform pathological single-modal inference and output a pathological score when the available modal includes full-field pathological slides but does not include digital radiographic images. The modal adaptive decision aggregation module is used to perform corresponding operations based on the recognition results: when only clinical information is included, a data insufficiency prompt is generated; when clinical information, digital radiographic images and full-view pathological slides are included simultaneously, the image clinical collaboration module and the histopathology analysis module are called in parallel, and the image clinical collaboration score and the pathological score are merged into a comprehensive diagnostic score. The results output module is used to output auxiliary diagnostic results based on the image clinical collaborative score, pathological score, or comprehensive diagnostic score.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the modality-adaptive multimodal assisted diagnostic method for osteosarcoma as described in any one of claims 1 to 7.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the modality-adaptive multimodal assisted diagnostic method for osteosarcoma as described in any one of claims 1 to 7.