X-ray-based disease-specific visual and language large model construction method and system

CN122822374APending Publication Date: 2026-09-25SHANGHAI TENTH PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610698780.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0009]本发明的目的是针对现有技术中的不足,提供一种基于X线的疾病专用视觉及语言大模型构建方法和系统,以解决针对现有技术存在的报告异质性、标签稀缺性、领域不匹配和输出可解释性的技术障碍的技术问题

Benefits of technology

[0020]本发明采用以上技术方案,通过获取多源多模态医学持续预训练数据集和专属骨骼肌肉X线微调数据集;获取多模态基础模型,多模态基础模型由视觉编码器、基于多层感知机的投影层和大语言模型组成;利用多源多模态医学持续预训练数据集,对多模态基础模型的视觉编码器、投影层和大语言模型进行全参数微调的持续预训练,得到医学领域适配模型;利用专属骨骼肌肉X线微调数据集,对医学领域适配模型进行全参数微调的领域特定微调,得到骨肌疾病领域专用视觉及语言大模型,与现有技术相比,具有如下技术效果:生成结构化、细粒度诊断输出的骨肌疾病领域专用视觉及语言模型,以实现超越粗粒度异常筛查的稳健诊断级解读,并具备在多中心、多设备、多解剖部位场景下的良好泛化能力的技术效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822374A_ABST
    Figure CN122822374A_ABST
Patent Text Reader

Abstract

The application relates to an X-ray-based disease special visual and language large model construction method and system, which obtains a multi-source multi-modal medical continuous pre-training data set and a special bone muscle X-ray fine-tuning data set; obtains a multi-modal basic model; uses the multi-source multi-modal medical continuous pre-training data set to perform full-parameter fine-tuning continuous pre-training on a visual encoder, a projection layer and a large language model of the multi-modal basic model, so as to obtain a medical field adaptation model; and the medical field adaptation model is subjected to full-parameter fine-tuning field-specific fine-tuning by using the special bone muscle X-ray fine-tuning data set, so as to obtain a bone muscle disease field special visual and language large model, which has the advantages that the bone muscle disease field special visual and language model can generate a structured and fine-grained diagnosis output, can realize robust diagnosis level interpretation beyond coarse-grained abnormal screening, and has the technical effects of good generalization capability in multi-center, multi-device and multi-anatomical site scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical information technology applications, and in particular to a method and system for constructing a large-scale visual and linguistic model of a disease based on X-rays. Background Technology

[0002] Disability-Adjusted Life Years (DALYs) resulting from musculoskeletal system (MSK) diseases are a leading cause of the global disease burden, accounting for 6.8%. According to the Global Burden of Disease 2021 study, the disability burden caused by MSK diseases continued to increase between 1990 and 2020 and is projected to continue until 2050. With an aging population, the demand for imaging diagnostic services is constantly increasing. While X-ray imaging remains the preferred imaging method for MSK diseases due to its high throughput and accessibility, the quality of interpretation varies significantly among different institutions and physicians.

[0003] Clinical narrative reports often suffer from inconsistencies in structure, terminology, and anatomical granularity, posing challenges to data processing and external validation of AI systems. International radiology guidelines (such as the European Society of Radiology's 2023 updated white paper on structured reporting) advocate for structured reporting to improve diagnostic comparability, but clinical adoption remains limited and inconsistent.

[0004] The existing technology faces the following three core obstacles: (1) Coarse-grained labeling problem: Widely used MSK benchmark datasets (such as the MURA dataset) only provide research-level "normal / abnormal" binary labels, lacking a standardized mapping from image discovery to specific disease entities, which limits fine-grained phenotypic analysis and diagnostic inference capabilities. Most existing artificial intelligence research in the MSK field focuses on single, narrow problems, such as fracture detection, osteoporosis grading, or benign / malignant differentiation; each is trained and evaluated independently, resulting in limited task coverage.

[0005] (2) Limitations of Classification Output: Real MSK diagnosis requires nuanced descriptive reasoning that far exceeds the capabilities of fixed-label classification systems. Simple multi-label outputs cannot capture the complex relationships between anatomical location, injury morphology, and severity (e.g., they cannot distinguish between "comminuted distal radius intra-articular fracture" and a simple "wrist fracture"). Free-text reporting remains the clinical standard, but its high heterogeneity in vocabulary, structure, and reference style hinders high-fidelity data integration and interoperable label extraction.

[0006] (3) Domain mismatch problem: General vision and language models are mainly trained on non-medical corpora or chest-centric datasets, which makes it difficult to handle MSK-specific ontology and detailed anatomical knowledge, and there is a significant risk of distributional bias. Rigorous multi-center external validation on diverse real-world MSK cohorts is urgently needed. In addition, traditional discriminative vision models are limited by predefined classification labels and lack comprehensive language generation capabilities, and cannot output coherent and grammatically correct image descriptions required for standardized structural diagnosis.

[0007] Existing research progress includes: Nguyen et al. proposed a deep learning method for X-ray fracture detection combining YOLACT++ and CLAHE; Alshamrani et al. designed a bone X-ray diagnostic method for osteoporosis based on ROI segmentation and SVM classification; Englebert et al. explored visual-language fusion pre-training of bone X-rays and French radiology reports. However, the above studies are all limited to single-task systems, and no end-to-end diagnostic visual-language model for whole-body musculoskeletal X-ray images has yet been proposed.

[0008] Currently, no effective solutions have been proposed to address the technical obstacles of existing technologies, such as report heterogeneity, label scarcity, domain mismatch, and output interpretability. Summary of the Invention

[0009] The purpose of this invention is to address the shortcomings of existing technologies by providing a method and system for constructing a disease-specific visual and linguistic large model based on X-rays, thereby solving the technical problems of report heterogeneity, label scarcity, domain mismatch, and output interpretability in existing technologies.

[0010] To achieve the above objectives, the technical solution adopted by the present invention is as follows: This invention provides a method for constructing a disease-specific visual and language large-scale model based on X-rays, comprising: acquiring a multi-source, multimodal medical continuous pre-training dataset and a dedicated musculoskeletal X-ray fine-tuning dataset; acquiring a multimodal base model, which consists of a visual encoder, a projection layer based on a multilayer perceptron, and a large language model; using the multi-source, multimodal medical continuous pre-training dataset, continuously pre-training the visual encoder, projection layer, and large language model of the multimodal base model with full parameter fine-tuning to obtain a medical domain-adapted model; and using the dedicated musculoskeletal X-ray fine-tuning dataset, performing domain-specific fine-tuning of the medical domain-adapted model with full parameter fine-tuning to obtain a dedicated visual and language large-scale model for musculoskeletal diseases.

[0011] Optionally, acquiring the multi-source, multimodal medical continuous pre-training dataset and the dedicated musculoskeletal X-ray fine-tuning dataset includes: during the acquisition of the multi-source, multimodal medical continuous pre-training dataset, collecting multimodal medical question-and-answer and image description datasets, endoscopic image datasets, and medical classification and segmentation datasets; for non-multimodal data, generating instruction data responses for medical tasks by combining background knowledge retrieved from medical knowledge bases and professional literature databases using a large language model; during the acquisition of the dedicated musculoskeletal X-ray fine-tuning dataset: acquiring multi-center clinical X-ray image data and corresponding image diagnosis reports from multi-center hospital clinical system PACS; acquiring public musculoskeletal X-ray datasets and performing diagnostic-level re-annotation; and generating instruction tuning data through task-oriented transformation.

[0012] Furthermore, optionally, the method further includes: performing standardized preprocessing on multicenter clinical X-ray image data, wherein the standardized preprocessing on multicenter clinical X-ray image data includes: parsing DICOM format files, retaining core image information and removing privacy-related fields; linearly normalizing pixel values ​​and converting the format; generating a unique content fingerprint for each image based on a hash function, and comparing hash values ​​across dataset partitions to remove duplicate samples.

[0013] Optionally, the method also includes: establishing a structured diagnostic labeling system, which includes: using a large language model to perform structured parsing of the image diagnostic report text, transforming the free text diagnostic conclusions into standardized diagnostic labels; having senior radiologists with clinical experience conduct dual review and verification by combining image visual features with pre-labeled labels; and developing comprehensive labeling guidelines covering fracture morphology, joint alignment abnormalities, degenerative changes, and soft tissue findings for the public MURA dataset, with at least two committee-certified musculoskeletal radiologists independently reviewing and assigning labels under blinded conditions, and arbitrating any disagreements by senior radiologists.

[0014] This invention provides a method for generating diagnostic information for musculoskeletal diseases based on X-rays, comprising: acquiring musculoskeletal X-ray images to be processed; inputting the musculoskeletal X-ray images to be processed into a pre-constructed visual and linguistic large-scale model for musculoskeletal diseases, and obtaining a structured diagnostic output including anatomical location, core disease diagnostic conclusions and auxiliary information; wherein, the pre-constructed visual and linguistic large-scale model for musculoskeletal diseases is constructed by the aforementioned method for constructing disease-specific visual and linguistic large-scale models based on X-rays.

[0015] Optionally, the method further includes: scoring the structured diagnostic output based on a structured diagnostic labeling system, with scoring dimensions including: anatomical location as the first scoring interval, core disease diagnosis as the second scoring interval, and completeness of auxiliary information as the third scoring interval; wherein, a large language model is used for automated batch scoring, supplemented by manual sampling evaluation by physicians, and intragroup correlation coefficients are used to quantitatively analyze the consistency between AI scoring and manual scoring.

[0016] Further, optionally, an intraclass correlation coefficient (ICC) assessment model based on a two-way mixed-effects model is used to generate inter-rater consistency between the scores and physician review scores, calculated independently for the scoring dimensions of core disease diagnosis, anatomical location, and completeness of auxiliary information; based on the central limit theorem, the 95% confidence interval of the population mean is estimated using the sample standard error, wherein the formula for calculating the confidence interval is: , in, ; It is the first The score for each question It is the average score.

[0017] Optionally, the method further includes: Attention maps of each attention layer in a large-scale visual and language model specifically designed for musculoskeletal diseases are obtained. Attribution correlation maps are calculated based on correlation propagation, and these maps are upsampled to the original image size to generate a visual heatmap. The calculation method for the attribution correlation maps is as follows: in This refers to the Hadamard product. Let A be the gradient. It is the model's output for the category to be visualized. This indicates that the mean is taken along the head dimension; the update rule for self-attention is... The aggregation rule for bimodal attention is: The final correlation map is upsampled to the original image size using bilinear interpolation to generate a visual heatmap, which is used to show the correlation between image regions and each generated token.

[0018] This invention provides an X-ray-based intelligent diagnostic system for musculoskeletal diseases, comprising: a data acquisition module for acquiring X-ray images of musculoskeletal organs to be diagnosed; a preprocessing module for performing standardized preprocessing on the X-ray images, including privacy field stripping, pixel normalization, and format conversion; an inference module containing a pre-constructed large-scale visual and linguistic model specific to the musculoskeletal disease domain, for receiving pre-processed X-ray images and structured queries, and outputting a structured diagnostic report containing anatomical location, core disease diagnosis conclusions, and auxiliary information; wherein the pre-constructed large-scale visual and linguistic model specific to the musculoskeletal disease domain is constructed using an X-ray-based disease-specific visual and linguistic model construction method; a visualization module for generating an attention heatmap based on correlation propagation to display the attention distribution of the model in the inference module on image regions; and an output module for outputting a structured diagnostic report.

[0019] Optionally, the inference module supports anatomical location identification query, anomaly detection query, disease type classification query, and open diagnostic report generation query modes. Anatomical location identification query identifies the anatomical location corresponding to the input image from a candidate set containing multiple body parts. Anomaly detection query determines whether there is a specified type of abnormality in the image. Disease type classification query performs fine-grained classification of fracture types and bone and joint diseases. Open diagnostic report generation outputs a complete structured radiology diagnostic conclusion, following the full classification system of orthopedic diseases, covering core disease categories such as traumatic bone injury, degenerative bone and joint diseases, metabolic bone diseases, and neoplastic bone lesions.

[0020] This invention employs the above technical solution by acquiring a multi-source, multimodal medical continuous pre-training dataset and a dedicated musculoskeletal X-ray fine-tuning dataset; acquiring a multimodal base model, which consists of a visual encoder, a projection layer based on a multilayer perceptron, and a large language model; using the multi-source, multimodal medical continuous pre-training dataset, continuously pre-training the visual encoder, projection layer, and large language model of the multimodal base model with full parameter fine-tuning to obtain a medical domain-adapted model; and using the dedicated musculoskeletal X-ray fine-tuning dataset, performing domain-specific fine-tuning of the medical domain-adapted model with full parameter fine-tuning to obtain a large-scale visual and language model specifically for the musculoskeletal disease domain. Compared with existing technologies, this invention has the following technical effects: generating a structured, fine-grained diagnostic output visual and language model specifically for the musculoskeletal disease domain, achieving robust diagnostic-level interpretation beyond coarse-grained anomaly screening, and possessing good generalization ability in multi-center, multi-device, and multi-anatomical site scenarios. Attached Figure Description

[0021] Figure 1 This is a flowchart illustrating a method for constructing a disease-specific visual and language large model based on X-rays according to an embodiment of the present invention. Figure 2This is a flowchart illustrating a method for generating diagnostic information for musculoskeletal diseases based on X-rays, according to an embodiment of the present invention. Figure 3 This is a schematic diagram illustrating a case study and heatmap visualization example in an X-ray-based method for generating diagnostic information for musculoskeletal diseases according to an embodiment of the present invention. Figure 4 This is a flowchart of a method for generating diagnostic information for musculoskeletal diseases based on X-rays according to an embodiment of the present invention. Figure 5 This is a schematic diagram of an X-ray-based musculoskeletal disease diagnostic information generation system according to an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.

[0023] Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without any inventive effort. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, any changes to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.

[0024] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0025] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units (elements) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or apparatus. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms “multiple” / “several” used in this application refer to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can indicate: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following objects are in an "or" relationship. The terms "first," "second," and "third" used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.

[0026] An illustrative embodiment of the present invention, such as Figure 1 As shown, Figure 1 This is a flowchart illustrating a method for constructing a large-scale visual and language model for diseases based on X-rays, according to an embodiment of the present invention. The method for constructing a large-scale visual and language model for diseases based on X-rays provided in this application includes: Step S102: Obtain the multi-source, multimodal medical continuous pre-training dataset and the dedicated musculoskeletal X-ray fine-tuning dataset; Optionally, step S102, obtaining the multi-source multimodal medical continuous pre-training dataset and the dedicated musculoskeletal X-ray fine-tuning dataset, includes: during the process of obtaining the multi-source multimodal medical continuous pre-training dataset, collecting multimodal medical question-and-answer and image description datasets, endoscopic image datasets, and medical classification and segmentation datasets; for non-multimodal data, generating instruction data responses for medical tasks by combining background knowledge retrieved from medical knowledge bases and professional literature databases using a large language model; during the process of obtaining the dedicated musculoskeletal X-ray fine-tuning dataset: obtaining multi-center clinical X-ray image data and corresponding image diagnosis reports from multi-center hospital clinical system PACS; obtaining public musculoskeletal X-ray datasets and performing diagnostic-level re-annotation; and generating instruction tuning data through task-oriented transformation.

[0027] Furthermore, optionally, the method for constructing a disease-specific visual and language large model based on X-rays provided in this application embodiment further includes: standardizing preprocessing of multi-center clinical X-ray image data, wherein the standardizing preprocessing of multi-center clinical X-ray image data includes: parsing DICOM format files, retaining core image information and removing privacy-related fields; linearly normalizing pixel values ​​and converting the format; generating a unique content fingerprint for each image based on a hash function, and comparing hash values ​​across dataset partitions to remove duplicate samples.

[0028] Optionally, the method for constructing a disease-specific visual and linguistic large model based on X-rays provided in this application embodiment further includes: establishing a structured diagnostic labeling system, wherein establishing the structured diagnostic labeling system includes: using a large language model to perform structured parsing of the image diagnostic report text, transforming the free text diagnostic conclusions into standardized diagnostic labels; having senior radiologists with clinical experience conduct dual review and verification by combining image visual features with pre-labeled labels; and developing comprehensive labeling guidelines covering fracture morphology, joint alignment abnormalities, degenerative changes, and soft tissue findings for the public MURA dataset, with at least two committee-certified musculoskeletal radiologists independently reviewing and assigning labels under blinded conditions, and any disagreements being arbitrated by senior radiologists.

[0029] Specifically, in this application embodiment, the total amount of the multi-source multimodal medical continuous pre-training dataset exceeds 1,054,083 samples, covering multiple imaging modalities (endoscopy 31.2%, CT 25.5%, MRI 15.8%, X-ray 8.1%, etc.) and multiple disease domains; In this embodiment, by specifically selecting a multimodal distribution containing 8.1% X-ray data, rather than using only X-ray data, the technical problems of poor generalization ability of models trained with single-modal data, inability to handle differences in X-ray images from different devices and body positions, and insufficient diagnostic reasoning ability due to the lack of support from other modal medical knowledge, are addressed. Through cross-learning of multimodal data, the model can learn more general medical image feature representations, significantly improving the model's adaptability to X-ray images from different centers and devices. At the same time, by combining anatomical knowledge from modalities such as CT and MRI, the model has a deeper understanding of musculoskeletal anatomy and can more accurately identify complex anatomical variations and lesions.

[0030] The embodiment of this application provides a method for constructing a disease-specific visual and linguistic large model based on X-rays. The general multimodal model is only trained on a non-medical corpus and lacks basic medical knowledge and cross-modal medical understanding. Directly using it for musculoskeletal diagnosis results in serious domain bias. Before entering the musculoskeletal specialty fine-tuning, the model has mastered general medical imaging features and professional medical language expressions, which significantly reduces the difficulty of subsequent specialty fine-tuning and avoids model overfitting.

[0031] In a preferred example, the multi-source, multimodal medical continuous pre-training dataset in this application embodiment includes: The multimodal medical question-and-answer and image description data in this application embodiment includes: collected and organized medical image-text correspondence datasets such as PMC-VQA (31,495 entries), SLAKE (8,835 entries), VQA-RAD (523 entries), PubMedVQA (646,741 entries), ROCO (79,789 entries), MedICaT (86,821 entries), and Medpix (2,050 entries), as well as endoscopic image datasets such as Kvasir-Capsule (15,568 entries), Kvasir-Instrument (472 entries), Nerthus (3,867 entries), and WCE-CCDD (5,200 entries). All data underwent deduplication, filtering, and format standardization, and were subsequently reviewed by professional medical personnel.

[0032] In this embodiment, the instruction data for non-multimodal medical data includes: for medical classification and segmentation data such as brain tumor MRI (brain-tumour-MRI-scan), dermatology (Dermnet), endoscopic lesion detection (EDD2020), skin lesions (ISIC2018), lung and colon histopathology (LC25000), chest X-ray (Montgomery-County-CXR-Set), fundus (Retina), and skin cancer (SkinCancer), question-answer pairs oriented towards medical tasks are generated through a large language model. Simultaneously, medical background knowledge highly compatible with the target modality and disease knowledge is retrieved from multi-source medical knowledge bases such as WiNGPT and Doubao, as well as professional literature databases. The GPT-4o model combines this background knowledge to generate high-quality instruction data responses, ensuring the knowledge density and coverage breadth of the pre-trained data. The instruction data responses at least include: image features, disease knowledge, and differential diagnosis. The following is a specific example of prompt word generation for the multi-source, multimodal medical continuous pre-training dataset in this application embodiment: Knowledge generation prompt: "Based on the specified modality {Modality} and target disease {TargetDisease}, please generate content in the following format: "Image characteristics and related knowledge" — "Image features:" — "Disease knowledge:" Diagnostic Reasoning Hints: "Task Description: Given a medical image, please describe from a doctor's perspective how to identify {TargetDisease} from the image, based on the relevant knowledge Context of the medical image. Answer Guidance: (1) The answer should not start by stating the disease, but rather analyze and draw a conclusion; (2) The answer should include a detailed description of the image; (3) Which relevant knowledge about {TargetDisease} is related to the image; (4) Which diseases need to be identified; (5) Give the conclusion that the image is {TargetDisease}." In a preferred example, the acquisition of the dedicated musculoskeletal X-ray fine-tuning dataset in this application embodiment mainly includes the following paths: Path 1: Multicenter Hospital Clinical Dataset We acquired multicenter clinical X-ray imaging data and corresponding imaging diagnostic reports from five independent hospital clinical PACS systems, spanning from January 2009 to June 2025. We performed image-level and research-level deduplication on the multicenter clinical X-ray imaging data using the SHA-256 cryptographic hash function to ensure that the training set, validation set, and test set are strictly mutually exclusive and without redundancy at the patient level. The public MURA dataset was acquired and re-annotated at the diagnostic level. Comprehensive annotation guidelines covering fracture morphology, joint alignment abnormalities, degenerative changes, and soft tissue findings were developed. Labels were independently reviewed and assigned by at least two committee-certified musculoskeletal radiologists under blinded conditions. Disagreements were arbitrated by senior radiologists, and Cohen's kappa quantitative assessment was used to evaluate inter-agency consistency. When the consistency fell below a preset threshold of 0.75, the annotation guidelines were iteratively refined. It should be noted that the combination of Cohen's kappa quantitative assessment and iterative refinement of the guidelines ensured high-quality and consistent labels, and solved the problem of labels being disconnected from image features. This is not a simple secondary annotation.

[0033] The task-oriented transformation generates instruction optimization data encompassing three tasks: anatomical site identification, fracture type classification, and abnormality detection. The dedicated musculoskeletal X-ray fine-tuning dataset is labeled using a three-dimensional structured diagnostic labeling system covering anatomical site localization, core disease diagnosis, and the completeness of auxiliary information. This addresses the problem that existing musculoskeletal datasets only provide binary labels of normal / abnormal or single disease classification labels, failing to capture the complex relationships between anatomical location, injury morphology, and severity. Consequently, trained models can only output coarse-grained classification results and cannot generate clinically usable diagnostic reports. This new dataset enables the model to output structured diagnostic reports that conform to clinical standards. It can not only determine the presence of abnormalities but also accurately locate lesion sites, clarify disease types, and provide auxiliary information such as classification and disease progression, achieving a leap from abnormality screening to diagnostic-level interpretation.

[0034] The multicenter clinical dataset includes X-ray images and corresponding imaging diagnoses from the PACS systems of five independent hospital clinical centers (including three tertiary-level hospitals), spanning nearly 16 years (January 2009 to June 2025), covering skeletal lesions in key anatomical locations such as the skull, elbow, lumbar spine, ankle, and foot. Each imaging report was written by a senior radiologist with more than 5 years of clinical experience to ensure the reliability of data labels. Skeletal X-ray images with clear positive and negative imaging findings were selected based on clearly defined inclusion criteria. The training set sample sizes for each center were 14,044 (Hospital 1, median age 68.0 years [49.0-79.0], male proportion 37.7%), 1,804 (Hospital 2, median age 37.0 years [15.0-54.0]), 2,015 (Hospital 3, median age 48.0 years [31.0-61.0]), 1,420 (Hospital 4, median age 48.0 years [33.0-64.0]), and 290 (Hospital 5, median age 66.0 years [55.0-74.0]), respectively; the test set sample sizes for the five centers were 2,200, 174, 236, 252, and 52, respectively.

[0035] Image preprocessing pipeline: The processing steps for raw DICOM files are as follows: Step (i) parses the DICOM file, retains the core image information, and removes all privacy-related fields (patient name, date of birth, ID number, examination number, etc.). Step (ii) extracts the pixel array, maps the pixel values ​​to the 0-255 range through linear normalization, converts it to uint8 format, and stores it as a JPG file; Step (iii) Calculate the SHA-256 cryptographic hash value for each image to generate a unique content fingerprint. Compare the image-level and research-level hash values ​​across training, validation, and test set partitions. Manually review and delete duplicate samples to ensure that all data partitions are mutually exclusive and without redundancy.

[0036] This application's embodiments utilize image-level and research-level dual deduplication based on SHA-256 hashing and patient-level data partitioning mutual exclusion to solve the problems commonly found in existing medical datasets, such as sample duplication and patient overlap between the training and test sets, leading to inflated model generalization capabilities and inability to be stably applied in real multi-center scenarios. It completely eliminates the risk of data leakage, ensures the objectivity of model performance evaluation, and enables the model to maintain stable diagnostic accuracy on external test sets from five independent hospitals, with cross-center score fluctuations significantly lower than those of general models. For the generation of instructions based on non-multimodal data, a process involving a large language model, background knowledge from a medical knowledge base, and secondary review by professional medical personnel is employed, rather than directly using a large language model for generation. This addresses the problem that medical instructions generated by general-purpose large language models often contain professional errors, outdated knowledge, and do not conform to clinical practice, leading to incorrect medical knowledge being learned by the model when directly used for training. By retrieving the latest medical knowledge base and professional literature, the accuracy of the instruction data is ensured; the secondary review by professional medical personnel further filters out errors and non-standard content, significantly improving the knowledge density and clinical relevance of the pre-training data.

[0037] The embodiments of this application can eliminate image differences between different devices and centers through standardized preprocessing, while avoiding the decline in model generalization ability caused by data leakage.

[0038] Diagnostic label generation process: The DeepSeek 671B model (i.e., the large language model in this embodiment) is used to perform structured parsing of the diagnostic conclusion text and convert it into standardized diagnostic labels; professional clinicians simultaneously receive images and pre-labeled labels, and verify the accuracy and clinical applicability of the labels through image-label dual review, so as to achieve a high degree of consistency between clinical diagnostic report text, image visual features and model training labels, and avoid the problem of labels originating from reports but not matching image features.

[0039] Examples of prompts for generating diagnostic labels in this application embodiment are as follows: "You are a senior radiologist. The patient's imaging diagnosis is (text information (diagnosis conclusion)). Please generate the patient's disease name. Please output the disease name directly. If there are multiple diseases, separate them with commas and include their serial numbers. Remember not to predict the disease." Path 2: Diagnostic-level re-annotation of the MURA dataset: All images in the MURA dataset were re-examined to create a refined set of diagnostic labels, replacing the original binary normal / abnormal labels. Comprehensive labeling guidelines covering fracture morphology, joint alignment abnormalities, degenerative changes, and soft tissue findings were developed to standardize diagnostic criteria for each anatomical region. Two committee-certified musculoskeletal radiologists independently reviewed each study under blinded conditions (without access to original MURA labels, model output, or patient metadata), assigning structured diagnostic labels based on predefined categories and decision rules. Disagreements were arbitrated by senior radiologists, who recorded the reasons for the final labels. Cohen's kappa method was used to quantify inter-investigator consistency; when consistency fell below a preset threshold, the labeling guidelines were iteratively refined. All image readings were performed under a unified, standardized reading protocol to ensure reproducibility.

[0040] Path 3: Task-oriented transformation and instruction construction of public datasets: Anatomical Part Recognition: Based on the UNIFESP X-ray body part classification dataset, anatomical part identification questions were generated using DeepSeek, and standard answer pairs were formed by matching the original task answers to construct training data for 21 body part recognition tasks. Multiple choice for fracture types: Based on the Bone Break Classification Image Dataset (covering 10 fracture types), DeepSeek was used to generate multiple choice questions for fracture type differential diagnosis (4-choose-1 format, the rest are distractors), and training data for fine-grained fracture classification ability was constructed. Binary classification of fractures: Based on Bone Fracture Multi-Region X-ray Data (covering fracture and non-fracture images of all anatomical sites throughout the body), an assessment paradigm in the form of yes / no questions is constructed; Out-of-distribution test datasets: Heel spurs (Heel Dataset, 200 cases), knee osteoarthritis (Digital Knee X-ray Images, 400 cases), bone tumors (Bone Tumor Dataset, 750 cases, positive / negative 1:1), and fractures (FracAtlasDataset, 1,434 cases). Binary single-choice questions for differential diagnosis of the corresponding types were generated using DeepSeek.

[0041] In this embodiment, step S102 constructs a multi-source heterogeneous medical continuous pre-training dataset with over 1 million samples, as well as a dedicated MSK X-ray fine-tuning dataset covering multi-center clinical data, diagnostically re-annotated public datasets, and task-oriented synthetic instruction data. A multi-stage image standardization and hash deduplication pipeline is established to perform privacy desensitization, pixel normalization, and format conversion on DICOM images. Based on the SHA-256 cryptographic hash function, image-level and research-level dual deduplication is implemented in the training set / validation set / test set partitions to ensure that there is no cross-contamination in the dataset partitions.

[0042] Furthermore, in this embodiment of the application, a three-dimensional structured diagnostic labeling system covering anatomical location, core disease diagnosis, and complete auxiliary information is established. This system is automatically pre-labeled using the DeepSeek 671B large language model and combined with a clinician image-label synchronous dual review mechanism to generate high-quality, highly consistent standardized supervisory signals.

[0043] Based on the above, this application utilizes the DeepSeek 671B large language model to perform structured parsing of the image diagnostic report text, and uses preset diagnostic tags to generate prompt words to transform the free text diagnostic conclusions into standardized diagnostic tags. Senior radiologists with more than 5 years of clinical experience simultaneously receive the images and pre-labeled tags, and verify the accuracy and clinical applicability of the tags through image-tag dual review, achieving a high degree of consistency between the clinical diagnostic report text, image visual features, and model training tags, avoiding the problem of tags originating from the report but not matching the image features. For the public MURA dataset, a comprehensive annotation guideline covering fracture morphology, joint alignment abnormalities, degenerative changes, and soft tissue findings is developed. At least two committee-certified musculoskeletal radiologists independently review and assign tags under blinded conditions. Disagreements are arbitrated by senior radiologists, and Cohen's kappa quantifies inter-agency consistency. When the consistency falls below a preset threshold of 0.75, the annotation guideline is iteratively refined until the preset consistency standard is reached.

[0044] This application's embodiments address the problems of single-source data—limited sample size and central bias in single clinical datasets, poor annotation quality and lack of real-world clinical scenario data in single public datasets, and lack of real-world image features in single synthetic datasets—by constructing a three-in-one fine-tuning dataset composed of multi-center clinical data, re-annotated public data, and task-oriented synthetic data. These solutions prevent the training of clinically applicable models with strong generalization capabilities. Multi-center clinical data ensures the clinical authenticity and diversity of the data; re-annotated public data supplements rare and standardized cases; and task-oriented synthetic data specifically enhances the model's capabilities for specific tasks. The combination of these three elements ensures that the fine-tuning dataset possesses both sufficient sample size and good diversity and representativeness.

[0045] Step S104: Obtain the multimodal base model, which consists of a visual encoder, a projection layer based on a multilayer perceptron, and a large language model. Specifically, in step S104, the multimodal base model is based on the Qwen3-VL-8B-Instruct base model architecture. This 8B parameter variant is selected to take into account native dynamic resolution image input support, multi-layer visual token injection fusion, leading performance in authoritative benchmarks, and optimal performance-computing power balance suitable for deployment in the limited computing power environment of hospitals. The process of selecting the basic model architecture in this embodiment is as follows: This application uses Qwen3-VL-8B-Instruct as the basic model architecture. This model consists of a visual encoder, a projection layer based on a multilayer perceptron (MLP), and a large language model. The comprehensive reasons for choosing this 8B parameter variant are as follows: This application's embodiments leverage the technical advantages of Qwen3-VL-8B-Instruct to natively support dynamic resolution image input, eliminating the need for fixed-resolution preprocessing of the original image; it can convert visual information extracted from different layers of the visual encoder into dedicated tokens, which are then injected into the corresponding layers of the large language model to achieve cross-layer visual-language alignment; and it boasts leading performance on multiple authoritative benchmarks. The practical advantage of using Qwen3-VL-8B-Instruct in this application is that the 8B parameter quantity achieves an optimal balance between performance and computational cost, making the model easier to deploy in hospital clinical auxiliary diagnostic systems with limited computing infrastructure, and thus having broad clinical application feasibility.

[0046] It should be noted that the multimodal basic model in this application embodiment is illustrated using Qwen3-VL-8B-Instruct as a preferred example, in order to implement the method for constructing a large-scale visual and language model for diseases based on X-rays provided in this application embodiment, and no specific limitations are made.

[0047] This application selects Qwen3-VL-8B-Instruct as the base model instead of other general multimodal models. This is a targeted choice based on the specific needs of musculoskeletal X-ray diagnosis, rather than a conventional engineering choice. It solves the technical problems that other models often use fixed-resolution input, requiring cropping or scaling of the original X-ray images, resulting in the loss of key lesion information; other models often use single-layer visual token fusion, which cannot fully utilize the features extracted from different layers of the visual encoder, resulting in poor cross-modal alignment; and large-parameter models (such as those with more than 70 bytes) have excessively high computational costs, making them unsuitable for deployment in hospitals with limited computing power. In this embodiment, Qwen3-VL-8B-Instruct natively supports dynamic resolution image input, eliminating the need for preprocessing of the original X-ray images and fully preserving the detailed information of the lesions. Its multi-layer visual token injection fusion mechanism can inject low-level features (such as edges and textures) and high-level features (such as anatomical structures and lesion morphology) extracted from different layers of the visual encoder into the corresponding layers of the large language model, achieving more refined cross-modal alignment. The 8B parameter count achieves an optimal balance between performance and computational cost, enabling real-time inference on a single H200 GPU, meeting the needs of high-throughput clinical diagnosis.

[0048] Step S106: Using the multi-source multimodal medical continuous pre-training dataset, perform full parameter fine-tuning continuous pre-training on the visual encoder, projection layer and large language model of the multimodal base model to obtain a medical domain-adapted model. Step S108: Using a dedicated musculoskeletal X-ray fine-tuning dataset, perform full-parameter fine-tuning of the medical domain adaptation model to obtain a large-scale vision and language model specifically for musculoskeletal diseases.

[0049] Combining steps S106 and S108, this embodiment of the application obtains a large-scale visual and language model specifically for the musculoskeletal disease domain through a two-stage full-parameter fine-tuning training strategy: The first stage utilizes a multi-source, multi-modal medical continuous pre-training dataset for continuous pre-training, adapting the model from a general representation to a medical domain distribution. Specifically, a multi-source, multimodal medical continuous pre-training dataset was used to fine-tune the full parameters of three modules: the visual encoder, the MLP projection layer, and the large language model. After loading the Qwen3-VL-8B-Instruct pre-trained weights, the cross-entropy loss was calculated using task data and backpropagation was performed. Gradient updates were then applied to all parameters, enabling the model to adapt from a general representation depth to the visual feature distribution and specialized language expressions in the medical field.

[0050] Training configuration: 4 H200 GPUs, 3 epochs of training, sequence length 8192.

[0051] The second stage utilizes the MSK X-ray instruction dataset (i.e., the dedicated musculoskeletal X-ray fine-tuning dataset in this application embodiment) to perform domain-specific fine-tuning, enabling the model to deeply adapt to the professional ontology of musculoskeletal image diagnosis. Among them, based on the medical basic capabilities established through continuous pre-training, the model is fine-tuned with all parameters based on the constructed musculoskeletal X-ray instruction dataset, so that the model fully meets the professional needs of clinical musculoskeletal X-ray diagnosis and deeply adapts to MSK specific ontology and detailed anatomical knowledge.

[0052] Training configuration: 8 H200 GPUs, 4 epochs of training, sequence length 16384.

[0053] Fine-tuning data to generate prompts: Fracture type classification: "Based on the provided imaging images, determine the specific type of fracture? A: Greenstick fracture B: Impacted fracture C: Bone fissure D: Avulsion fracture. Please only reply with the corresponding letter; no additional analysis or explanation is required." Fracture binary classification: "Based on the given image, determine whether a fracture exists? A: Yes B: No. Please only reply with the option, no additional analysis required." Anatomical Location Recognition: "Please select the corresponding body part from the following options (abdomen, ankle, cervical spine, chest, clavicle, elbow, foot, finger, forearm, hand, hip, knee, lower leg, lumbar spine, pelvis, shoulder, sinus, skull, thigh, thoracic spine, wrist). If no corresponding part is found, other parts will be output." This application adopts a two-stage full-parameter fine-tuning strategy instead of single-stage fine-tuning or freezing partial parameter fine-tuning. This solves the technical problems of single-stage fine-tuning, which directly uses musculoskeletal specialty data to train a general model, which can easily lead to the model catastrophically forgetting general knowledge and failing to fully learn the professional knowledge of musculoskeletal specialty; and the fine-tuning method of freezing the visual encoder or large language model, which can only adjust a small number of parameters and cannot deeply adapt to the visual features and language expressions of the musculoskeletal domain, resulting in a low upper limit of model performance. The embodiments of this application enable the model to smoothly transition from a general representation to a medical domain representation through continuous pre-training in the first stage, avoiding catastrophic forgetting; the second stage of domain-specific fine-tuning further adapts the model to musculoskeletal specialty knowledge on the basis of basic medical capabilities, enabling the model to master the professional ontology and detailed anatomical knowledge of musculoskeletal imaging diagnosis; full-parameter fine-tuning can adjust all parameters of the model, fully releasing the model's potential, and making the model's performance on musculoskeletal diagnostic tasks significantly better than the frozen fine-tuning method.

[0054] Meanwhile, different training parameter configurations were used in the two stages (stage 1: 4 H200 images, 3 epochs, sequence length 8192; stage 2: 8 H200 images, 4 epochs, sequence length 16384). This was designed for the training objectives of different stages: the first stage aimed to enable the model to quickly adapt to general knowledge in the medical field, with a large amount of data but low complexity per sample, so a shorter sequence length and fewer epochs were used to improve training efficiency; the second stage aimed to enable the model to deeply learn fine-grained diagnostic knowledge in musculoskeletal specialties, with more complex diagnostic descriptions per sample, so a longer sequence length and more epochs were used to ensure that the model could fully learn fine-grained features and knowledge.

[0055] This application, through two-stage different sequence length configurations (8192 to 16384), is a targeted design for the complexity of musculoskeletal diagnostic texts, rather than a conventional engineering adjustment, enabling the model to better handle fine-grained diagnostic descriptions.

[0056] The method for constructing a disease-specific visual and linguistic large model based on X-rays provided in this application addresses the problem that general fine-tuning methods, which use a uniform sequence length, cannot simultaneously meet the needs of learning general medical knowledge and fine-grained diagnostic descriptions for musculoskeletal specialties; and that single-stage fine-tuning can easily lead to the model forgetting general knowledge or failing to deeply adapt to specialized knowledge. In the embodiments of this application, the shorter sequence length of the first stage is suitable for rapid learning of general medical knowledge, while the longer sequence length of the second stage can handle complex fine-grained descriptions in musculoskeletal diagnosis, such as comminuted distal radius intra-articular fracture with displacement, so that the model can deeply adapt to the professional ontology of musculoskeletal imaging diagnosis while retaining general medical capabilities.

[0057] This invention employs the above technical solution by acquiring a multi-source, multimodal medical continuous pre-training dataset and a dedicated musculoskeletal X-ray fine-tuning dataset; acquiring a multimodal base model, which consists of a visual encoder, a projection layer based on a multilayer perceptron, and a large language model; using the multi-source, multimodal medical continuous pre-training dataset, continuously pre-training the visual encoder, projection layer, and large language model of the multimodal base model with full parameter fine-tuning to obtain a medical domain-adapted model; and using the dedicated musculoskeletal X-ray fine-tuning dataset, performing domain-specific fine-tuning of the medical domain-adapted model with full parameter fine-tuning to obtain a large-scale visual and language model specifically for the musculoskeletal disease domain. Compared with existing technologies, this invention has the following technical effects: generating a structured, fine-grained diagnostic output visual and language model specifically for the musculoskeletal disease domain, achieving robust diagnostic-level interpretation beyond coarse-grained anomaly screening, and possessing good generalization ability in multi-center, multi-device, and multi-anatomical site scenarios.

[0058] An illustrative embodiment of the present invention, such as Figure 2 As shown, Figure 2 This is a flowchart illustrating a method for generating diagnostic information for musculoskeletal diseases based on X-rays according to an embodiment of the present invention. The method for generating diagnostic information for musculoskeletal diseases based on X-rays provided in this application includes: Step S202: Acquire the X-ray image of the musculoskeletal system to be processed; Step S204: Input the musculoskeletal X-ray image to be processed into a pre-constructed large-scale visual and linguistic model for musculoskeletal diseases to obtain a structured diagnostic output containing anatomical location, core disease diagnosis conclusions and auxiliary information. Among them, the pre-constructed large-scale visual and language models specifically for the musculoskeletal disease domain are as described above. Figure 1 The embodiment shown is constructed using a disease-specific visual and language large model based on X-rays.

[0059] Optionally, the X-ray-based method for generating diagnostic information for musculoskeletal diseases provided in this application further includes: scoring the structured diagnostic output based on a structured diagnostic labeling system, with scoring dimensions including: anatomical location as the first scoring interval, core disease diagnosis as the second scoring interval, and completeness of auxiliary information as the third scoring interval; wherein, a large language model is used for automated batch scoring, supplemented by manual sampling evaluation by physicians, and intragroup correlation coefficients are used to quantitatively analyze the consistency between AI scoring and manual scoring.

[0060] Specifically, in this application embodiment, the first score range can be 0-30 points; in this application embodiment, the core disease diagnosis score range can be 0-50 points; and in this application embodiment, the auxiliary information completeness score range can be 0-20 points. This application's embodiments utilize the GPT-4o mini model to achieve automated batch scoring, completing the scoring judgment according to three core dimensions: Anatomical location (0-30 points): 30 points for matching the bone or joint name with the standard answer; 20 points for locating to the same bone / finger / toe without specifying the segment; 10 points for locating only to the same limb or adjacent joint; 0 points for incorrect location. Core disease diagnosis (0-50 points): 50 points are awarded if the core diagnosis is consistent with the standard answer or has the same medical meaning; 40 points are awarded if it is very close but has slight differences; 30 points are awarded if the abnormality is identified but the core diagnosis is not accurately given; 0 points are awarded if it does not conform to the standard answer. Completeness of auxiliary information (0-20 points): Based on the secondary information items clearly stated in the standard answer, 5 points will be deducted for each missing core secondary information item; covering surgical procedures and treatment methods (external fixation with plaster cast, internal fixation with plate, intramedullary nail fixation, etc.), disease course and time status (postoperative, old, acute, etc.), fracture or injury classification (closed, open, comminuted, compression, greenstick, etc.) and other important modifying information (multiple, displacement / non-displacement, etc.); for cases where the model's answer is more redundant than the standard answer, points will be deducted according to the following levels: no redundancy (no points deducted), mild redundancy (1-3 points deducted), moderate redundancy (4-8 points deducted), and severe redundancy (9-15 points deducted), with a total deduction of no more than 15 points.

[0061] The scoring results are output in JSON format to ensure the accuracy and comparability of the scores across all dimensions.

[0062] The weight setting logic in this application embodiment is as follows: first determine the anatomical location (localization), then determine the nature of the disease (qualitative analysis), and finally supplement details such as classification and surgical procedure (refinement), which conforms to the actual operation logic of imaging diagnosis.

[0063] In this application embodiment, the model test prompt used for evaluation can be: "Please complete the lesion labeling task from the diagnostic perspective of a professional radiologist, based on X-ray bone images, strictly following the comprehensive classification system of orthopedic diseases (including core categories such as traumatic bone injury, degenerative osteoarthritis, metabolic bone disease, and neoplastic bone lesions). Classification accuracy: Ensure that the lesion type and disease name are standardized and unambiguous; Labeling format: Only output the disease name, multiple lesions are presented independently in the format of "serial number: anatomical location + lesion name", and different lesions are separated by commas; Normal situation: If no obvious abnormalities are found in the skeletal system, uniformly label it as "normal or negative". In this embodiment, the model outputs a structured diagnostic report in the format of sequence number: anatomical location + lesion name, rather than a free text report. This is because: it solves the problems of inconsistent structure, terminology, and anatomical granularity in free text reports, making them unsuitable for direct use in clinical data statistics and subsequent artificial intelligence research; at the same time, the quality of free text reports is difficult to quantify and evaluate; structured diagnostic reports, by adhering to a unified format and terminology standard, facilitate the organization, analysis, and sharing of clinical data; furthermore, the structured output format allows us to use a three-dimensional quantitative scoring system to accurately evaluate model performance, providing the possibility for continuous model optimization.

[0064] Further, optionally, an intraclass correlation coefficient (ICC) assessment model based on a two-way mixed-effects model is used to generate inter-rater consistency between the scores and physician review scores, calculated independently for the scoring dimensions of core disease diagnosis, anatomical location, and completeness of auxiliary information; based on the central limit theorem, the 95% confidence interval of the population mean is estimated using the sample standard error, wherein the formula for calculating the confidence interval is: , in, ; It is the first The score for each question It is the average score.

[0065] Specifically, in this embodiment, 200 unique patients are randomly selected from the reserved test set at the patient level. The test set and training data are strictly separated at the patient level to ensure that patients who provided expert review do not appear in any part of the training pipeline. Each patient is selected for a representative study, without stratified sampling by subgroup.

[0066] A review panel of three committee-certified physicians independently reinterpreted each study under blinded conditions (without access to model output or other reviewers' assessments), accessing only the original X-ray images and corresponding imaging findings. For each case, reviewers assessed whether the AI ​​system correctly identified the anatomical location of the disease, whether there were any imaging abnormalities, and whether the diagnostic conclusions generated by the model were accurate. When reviewers disagreed, a consensus was reached through collective discussion.

[0067] In this embodiment, the anomaly detection and anatomical part recognition tasks can be performed as follows: for each input image, output the single most matching body part from a candidate set of 21 body parts; if no match is found, output "other parts". For anomaly detection and other multiple-choice questions: a unified binary classification option setting rule is adopted, with each question configured with two mutually exclusive options, A and B, and the model outputs the corresponding single-option answer. Accuracy is used to evaluate the model performance for the above tasks.

[0068] The statistical analysis method in this application uses the intraclass correlation coefficient (ICC) based on a two-way mixed-effects model (absolute consistency) to assess the consistency between the model-generated score and the physician's review score, and calculates them independently in three scoring dimensions: core disease diagnosis, anatomical location, and completeness of auxiliary information.

[0069] Estimate the 95% confidence interval of the overall mean score based on the central limit theorem: Treat each problem in the evaluation as an independent sample drawn from a hypothetical infinite population of problems. Let the evaluation consist of n independent problems, and the score of the i-th problem be... The sample mean is: Then the standard error estimate of the overall mean is: For the Bernoulli variable (s_i in {0,1}), the standard error simplifies to: , The 95% confidence interval is: .

[0070] The embodiments of this application employ a three-dimensional quantitative scoring system based on the ICC consistency assessment and the 95% confidence interval estimation of the central limit theorem using a two-way mixed-effects model. This system addresses the technical problems of existing model assessments, which often rely on a single accuracy index and cannot comprehensively evaluate the quality of structured diagnostic outputs; as well as the low efficiency and high subjectivity of manual assessments. It enables automated batch assessment of model diagnostic performance, while ensuring the credibility of the assessment results through physician sampling assessment and consistency analysis. It can accurately quantify the differences in the model's capabilities across the three dimensions of localization, qualitative analysis, and detailed description, providing a clear direction for model optimization.

[0071] This application embodiment constructs an AI-physician joint assessment system, uses GPT-4o mini to achieve automated batch scoring of diagnostic performance, and supplements it with physician random sampling scoring and ICC consistency analysis, taking into account both assessment efficiency and result reliability.

[0072] Optionally, the method for generating diagnostic information for musculoskeletal diseases based on X-rays provided in this application embodiment further includes: Attention maps of each attention layer in a large-scale visual and language model specifically designed for musculoskeletal diseases are obtained. Attribution correlation maps are calculated based on correlation propagation, and these maps are upsampled to the original image size to generate a visual heatmap. The calculation method for the attribution correlation maps is as follows: in This refers to the Hadamard product. Let be the gradient with respect to A. It is the model's output for the category to be visualized. This indicates that the mean is taken along the head dimension; the update rule for self-attention is... The aggregation rule for bimodal attention is The final correlation map is upsampled to the original image size using bilinear interpolation to generate a visual heatmap, which is used to show the correlation between image regions and each generated token.

[0073] Specifically, the interpretability attention heatmap visualization in this application embodiment can be as follows: The basic calculation formula for the attention mechanism is as follows: Where Q is the query matrix, K and V are key / value matrices, A is the attention map, and h is the number of attention heads. For the embedding dimension, k and q represent the number of image tokens and text tokens, respectively. The final attribution attention map is calculated using the attention maps from each attention layer: ; in This refers to the Hadamard product. , Let be the mean across the attention head dimension, (.) + It is a ReLU nonlinear expression.

[0074] Initialization: Relevance matrix of self-attention and R_ Initialized as an identity matrix, the correlation matrix R_qk of the bimodal attention is initialized to a zero matrix. The correlation graph is iteratively updated using the following formula: ; ; ; right and R_ After normalization, the aggregation rule for bimodal attention units is as follows: ; The relevance scores of image patches in the final relevance map are first reconstructed into a grid that matches the layout of the original image (heatmap basis), and then upsampled to the original image size through bilinear interpolation to generate a visual heatmap that shows the region in the image that is most relevant to each generated token.

[0075] This application employs an attention heatmap based on correlation propagation, rather than a simple attention map visualization. This is because it addresses the technical problem that traditional attention map visualization only shows the model's attention to image regions, failing to explain the causal relationship between those regions and specific diagnostic conclusions. Clinicians still cannot understand why the model made that diagnosis. The correlation propagation method can calculate the contribution of each image region to each generated diagnostic token, intuitively showing what regions the model saw and the causal relationship leading to the diagnostic conclusions. This makes the model's decision-making process completely transparent and verifiable, significantly improving clinicians' trust in the model.

[0076] The attention heatmap visualization mechanism based on correlation propagation in this application solves the technical problem that existing deep learning models are black boxes, unable to explain the basis of diagnostic decisions, and clinicians find it difficult to trust the model's output, thus limiting clinical applications. It can intuitively show the correlation between the image region that the model focuses on and each generated diagnostic token, making the model's decision-making process transparent and verifiable, helping clinicians quickly understand the model's diagnostic basis, and facilitating the discovery of model errors and shortcomings.

[0077] Case 1: Imaging after left ankle internal fixation surgery In postoperative imaging of the left ankle, the patient underwent internal fixation treatment, with X-ray images showing multiple screws and plates implanted. The visualization heatmap generated by MSKVLM showed a highly focused area of ​​attention on the internal fixation implants (screws, plates) and the surrounding adjacent bone structures, consistent with typical postoperative imaging findings. This indicates that the model has good sensitivity in recognizing postoperative structural reconstructions in complex postoperative imaging contexts, providing transparent and verifiable diagnostic support in postoperative follow-up and complication identification.

[0078] Case 2: Imaging of a right calcaneal fracture In the imaging analysis of a right calcaneal fracture, the heatmap generated by MSKVLM showed a highly focused area of ​​attention on the posterior part of the calcaneus and near the articular surface, closely matching the actual anatomical location of the fracture. X-ray images revealed discontinuities in the calcaneal structure, accompanied by fracture lines and localized changes in bone density, clearly corresponding to the highlighted areas in the model. The model demonstrated excellent lesion localization capabilities in complex foot and ankle structures, providing structured diagnostic support in preoperative or postoperative imaging and helping clinicians understand the model's decision-making rationale.

[0079] like Figure 3 As shown, Figure 3 This is a schematic diagram of a case study and heatmap visualization example in an X-ray-based method for generating diagnostic information for musculoskeletal diseases according to an embodiment of the present invention, wherein the model attention distribution is shown in post-internal fixation images and calcaneal fracture images.

[0080] Combination Figure 1 and Figure 2 In this embodiment, the diagnostic label synthesis and review mechanism based on a large language model is used. Clinical diagnostic conclusions are mostly unstructured natural language text. Manually extracting labels is time-consuming, laborious, and prone to inconsistencies. By using a large language model to structurally decompose diagnostic conclusions and generate standardized labels, unstructured text can be quickly converted into a label format that the model can recognize. However, pre-annotation has limitations such as misunderstandings of professional terminology, missed diagnoses of rare diseases, and the inability to verify labels in conjunction with visual image features. This embodiment achieves a high degree of consistency between clinical diagnostic report text, image visual features, and model training labels through a mechanism that combines LLM pre-annotation with synchronous dual review of images and labels by clinicians. This fundamentally solves the problem of labels originating from reports but not matching image features, ensuring the integrity and logical consistency of the dataset.

[0081] Furthermore, most existing MSK imaging AIs are trained and evaluated for a single stenosis problem, and feature representations cannot be shared between tasks. In practice, the deployment cost is high and the coverage is limited. However, in the embodiments of this application, MSKVLM unifies anomaly detection, anatomical site recognition and diagnostic label generation into a single vision-language model framework. Without sacrificing the baseline performance of anomaly detection, it significantly expands the task coverage and output granularity, and achieves the clinical benefits of unified deployment of multiple tasks.

[0082] The fundamental reason for the insufficient performance of general VLM in the MSK subspecialty field in the existing technology is that it has not been adapted for training specific to the ontology and anatomical details of musculoskeletal images. However, the embodiments of this application effectively solve the problem of insufficient adaptability and performance limitations of general VLM in the MSK subspecialty field by constructing a large-scale, high-quality domain-specific dataset and performing two-stage full parameter fine-tuning, which fully verifies the scientific nature, rationality and practical effect of the domain-specific training strategy.

[0083] Furthermore, the performance patterns of anatomical sites and their clinical significance reflect the dual impact of data distribution and image complexity. Long bones and small joints (such as elbows, wrists, and fingers) showed high localization and core diagnostic scores; weight-bearing joints (hips, knees, and ankles) also performed strongly, confirming the advantage of deep learning in fracture identification and osteoarthritis-related tasks, which can approach the level of specialists; axial bone regions (cervical, thoracic, lumbar spine, and pelvis) performed well, with moderate cross-center variability, indicating robustness to differences in clinical practice across institutions; the localization of bones with complex shapes, such as the scapula, needs improvement, suggesting the need for more labeled and training data for irregular bone morphologies.

[0084] In summary, the experimental comparison results of the embodiments of this application are as follows: (1) Diagnostic label scoring of multicenter test set On an external test set of five independent tertiary medical centers (with significant differences in age distribution among centers, P=0.04; no significant difference in gender distribution, P=0.32), MSKVLM in this embodiment achieved the highest diagnostic label scores across all five institutions, with an average final score of 77.32 (95% CI: 76.37-78.28). Hospital 1 performed best (77.94 points), significantly outperforming all comparative models: Qwen3-VL-8B (67.11 points), InternVL3.5-8B (66.64 points), MedGemma-4B (50.17 points), Lingshu-7B (60.61 points), and Hulu-Med-7B (62.17 points). The comparative models showed greater score fluctuations across different hospitals, indicating that the label structure modeling capability needs improvement; while the cross-institutional stability of MSKVLM demonstrates good diagnostic label parsing generalization ability.

[0085] (2) Three-dimensional quantitative assessment of each anatomical site MURA test set: MSKVLM showed stable performance across multiple anatomical regions. In terms of anatomical location scores, the tibia (25.78 points) and elbow (25.19 points) scored the highest, demonstrating high recognition accuracy in long bone regions. Regarding core disease diagnosis, the wrist (40.80 points) and elbow (38.29 points) scored the highest, reflecting diagnostic advantages in differentiating small joint lesions. The score for completeness of auxiliary information remained stable at approximately 19 points, indicating room for further optimization in auxiliary information extraction. Overall, its diagnostic performance in joint-related regions was superior to smaller structures such as the fingers and radius / ulna.

[0086] In the in-hospital test set, the cervical and lumbosacral spine scored highest in core disease diagnosis (38.89 points) and auxiliary information completeness (20.00 points), with an average final score of 85.56 points, demonstrating excellent performance in spinal structure recognition and lesion extraction. The clavicle (83.26 points), skull (81.01 points), and chest (80.19 points) also had high overall scores; the scapula (63.85 points) had a relatively low score in anatomical location (13.08 points), suggesting that there is still room for improvement in the recognition of complex skeletal structures.

[0087] Regarding the physician's corrected diagnostic performance, as shown in Table 1: Table 1 Inter-assessor reliability: The consistency of auxiliary diagnostic information was nearly perfect (kappa=0.88), the anatomical location was substantially consistent (kappa=0.75), and the core disease diagnosis was moderate to substantially consistent (kappa=0.67), which verified the scientific validity and reliability of the GPT-4o automatic scoring method.

[0088] As shown in Table 2, in the comparative test on the public dataset: Table 2 Compared with single-task dedicated models: In the fracture multi-region classification task, the existing best single-task dedicated model has an accuracy of 96.7%, while the MSKVLM in this embodiment reaches 98.0%, surpassing the single-task dedicated model. Overall, the embodiment in this application takes into account both multi-task coverage and single-task competitiveness within a single vision-language model framework.

[0089] By clearly defining the multi-center test scores and task accuracy, the effectiveness of the X-ray-based musculoskeletal disease diagnostic information generation method provided in this application embodiment can be demonstrated in practical applications.

[0090] In summary, such as Figure 4 As shown, Figure 4It is a working flow chart of an X-ray-based musculoskeletal disease diagnosis information generation method according to an embodiment of the present invention, which includes a network architecture and visualization framework, a multi-stage data preparation and training process, and downstream evaluation tasks; the details are as follows: As Figure 4 shown, the basic architecture, the entire training iteration process, and the supporting task and evaluation system of the musculoskeletal (MSK) specific medical multimodal large model MSKVLM are presented. The whole is divided into three core modules A, B and C, and the process logic is as follows: Module A: Multimodal Model Basic Architecture: It is the underlying core architecture of MSKVLM, which realizes cross-modal fusion of medical images and text and question-and-answer output. The core process is as follows: Dual-channel input: One channel is the medical image to be analyzed (such as a foot X-ray film) as the visual input Xi; the other channel is the user's inquiry / diagnosis question text (such as "What is the diagnosis result of this image") as the language input Xq.

[0091] Visual feature processing: The input medical image Xi enters the visual encoder to generate visual features Zi; then the visual features are mapped to hidden layer features Hi adapted to the language model through the projection layer W, and input to the deep module of the language model.

[0092] Language feature processing and fusion: Qwen3 (Tongyi Qianwen 3) is used as the base language model fθ. The input question text Xq is processed by the language model to generate hidden layer features Hq, which completes cross-modal fusion with the projected visual features Hi.

[0093] Result output: The language model fused with visual and text information finally outputs the corresponding natural language answer Xa (such as "fracture of the right calcaneus") to complete the multimodal question answering task of medical images.

[0094] Module B: Complete model training process and data set construction. It is the training iteration link of MSKVLM, which is divided into two major stages: continuous pre-training and fine-tuning. It completely covers the whole process of data set construction, model iteration and data splitting. The core steps are as follows: Base model starting point: The general multimodal large model qwen3-vl (Qwen3-Vision Version) is used as the initial base model.

[0095] Continuous pre-training stage: Data set sources: MedAI multimodal data set and MedAI unimodal data set are used as pre-training data sources.

[0096] Data processing flow: First, unified data preprocessing is completed; among which, single-modal data is used to construct question-answer pairs through a large model, and all pre-training data are reviewed by medical experts to finally form a compliant and continuously pre-training dataset.

[0097] Model iteration: The qwen3-vl model is continuously pre-trained in the medical field using the above pre-trained dataset to obtain the intermediate model wingpt-vl adapted to the medical field.

[0098] Fine-tuning datasets are constructed in three main categories, each of which undergoes standardization processing: The hospital's own dataset: First, the data is de-identified and the images are preprocessed. Then, based on the impression part of the image report, a question-answer pair is built through a large model. Finally, it is manually labeled and reviewed by doctors.

[0099] Mura's public dataset: First, manual annotation and physician review are completed, and then question-answer pairs are built using a large model.

[0100] Other publicly available fine-tuning datasets include the UNIFESP X-ray body part classification dataset, fracture classification image dataset, and multi-region fracture X-ray dataset, all of which use a large model to construct question-answer pairs based on data labels. These three types of processed data together constitute the final fine-tuning dataset.

[0101] Fine-tuning training and final model output: Dataset split: The fine-tuning dataset is split in an 8:2 ratio, with 80% of the data used for fine-tuning training and 20% used for model testing.

[0102] Model fine-tuning: Using the split training dataset, the intermediate model wingpt-vl is fine-tuned to finally obtain the dedicated multimodal large model MSKVLM for musculoskeletal imaging scenes.

[0103] Accompanying test set splitting: A separate test dataset is simultaneously split off for subsequent performance testing and effect verification of MSKVLM.

[0104] Module C: Model Task Design and Evaluation System: Defines the three core implementation tasks of MSKVLM and the corresponding standardized evaluation process, as follows: Task 1: Imaging Diagnosis Task Task input: X-ray image and the question "Based on this image, what is the possible diagnosis?"

[0105] Model output: the corresponding disease diagnosis result (e.g., right calcaneal fracture).

[0106] Assessment process: Professional physicians complete quantitative scoring according to three scoring criteria: ① whether the anatomical location description is correct; ② the accuracy of the primary diagnosis; ③ whether the secondary diagnosis is complete, and output the quantitative score results for each item.

[0107] Task 2: Anatomical Site Recognition Task input: X-ray image and question "Accurately identify the body part in the image based on the anatomical region in the X-ray image".

[0108] Model output: The name of the body part being imaged (e.g., calcaneus).

[0109] Assessment process: A professional physician completes three compliance judgments: ① whether the anatomical location is correctly identified; ② whether the imaging abnormalities are correctly identified; ③ whether the primary diagnosis is accurate. Each item outputs a binary judgment result of "yes / no".

[0110] Task 3: Fracture Classification Task input: X-ray images and multiple-choice questions on fracture classification (e.g., "A. fissure fracture B. impacted fracture C. comminuted fracture"), requiring you to determine the specific type of fracture.

[0111] Model output: The corresponding answer option (e.g., C).

[0112] Evaluation process: The automated program determines whether the model's output answer is correct and directly outputs the "yes / no" result.

[0113] The X-ray-based musculoskeletal disease diagnostic information generation method provided in this application addresses four core technical obstacles in existing technologies: report heterogeneity, label scarcity, domain mismatch, and output interpretability. It constructs a closed-loop technical system encompassing high-quality dataset construction, two-stage differentiated full-parameter fine-tuning, multi-dimensional performance evaluation, and interpretability visualization. These technical features are interconnected and synergistic, jointly solving the systemic challenges of AI-based musculoskeletal X-ray image diagnosis. Firstly, it addresses the lack of medical knowledge in general models through multi-source, multi-modal medical continuous pre-training datasets, laying the foundation for cross-modal medical understanding. Secondly, it addresses the issues of coarse-grained labels and label-image disconnect in existing datasets through a dedicated musculoskeletal X-ray fine-tuning dataset, a three-dimensional structured labeling system, and full-process annotation quality control, providing high-quality domain supervision signals for the model. Thirdly, it addresses the mismatch between general models and the musculoskeletal domain distribution through a two-stage differentiated full-parameter fine-tuning strategy, enabling the model to gradually adapt to the transfer from general medical knowledge to musculoskeletal specialty knowledge. Finally, it addresses the issues of incomplete model performance evaluation and opaque decision-making processes through a three-dimensional quantitative evaluation system and a correlation propagation visualization mechanism, ensuring the clinical usability and credibility of the model.

[0114] This invention employs the above technical solution, acquiring X-ray images of musculoskeletal diseases to be processed; inputting these images into a pre-constructed visual and linguistic model specifically for musculoskeletal diseases, resulting in a structured diagnostic output including anatomical location, core disease diagnosis conclusions, and auxiliary information; wherein, the pre-constructed visual and linguistic model specifically for musculoskeletal diseases is constructed using the aforementioned X-ray-based method for constructing disease-specific visual and linguistic models. Compared with existing technologies, this invention has the following technical effects: generating a structured, fine-grained diagnostic output visual and linguistic model specifically for musculoskeletal diseases, achieving robust diagnostic interpretation beyond coarse-grained anomaly screening, and possessing good generalization ability in multi-center, multi-device, and multi-anatomical location scenarios.

[0115] An illustrative embodiment of the present invention, such as Figure 5 As shown, Figure 5 This is a schematic diagram of an X-ray-based musculoskeletal disease diagnostic information generation system according to an embodiment of the present invention. The X-ray-based musculoskeletal disease diagnostic information generation system provided in this application includes: Data acquisition module 51 is used to acquire X-ray images of musculoskeletal diseases to be diagnosed; preprocessing module 52 is used to perform standardized preprocessing on the X-ray images, including privacy field stripping, pixel normalization, and format conversion; inference module 53 contains a pre-built large-scale visual and linguistic model specifically for musculoskeletal diseases, used to receive pre-processed X-ray images and structured queries, and output a structured diagnostic report containing anatomical location, core disease diagnosis conclusions, and auxiliary information; wherein, the pre-built large-scale visual and linguistic model specifically for musculoskeletal diseases consists of... Figure 1 The large-scale visual and language model for diseases based on X-rays is constructed as shown; the visualization module 54 is used to generate an attention heatmap based on correlation propagation to show the attention distribution of the model to the image region in the inference module; the output module 55 is used to output a structured diagnostic report.

[0116] Optionally, the inference module 53 supports anatomical location identification query, anomaly detection query, disease type classification query, and open diagnostic report generation query modes; anatomical location identification query identifies the anatomical location corresponding to the input image from a candidate set containing multiple body parts; anomaly detection query determines whether there is a specified type of abnormality in the image; disease type classification query performs fine-grained classification of fracture types and bone and joint diseases; open diagnostic report generation outputs a complete structured radiology diagnostic conclusion, following the full classification system of orthopedic diseases, covering core disease categories such as traumatic bone injury, degenerative bone and joint diseases, metabolic bone diseases, and neoplastic bone lesions.

[0117] This invention employs the above technical solution, comprising a data acquisition module for acquiring X-ray images of musculoskeletal diseases to be diagnosed; a preprocessing module for standardizing and preprocessing the X-ray images, including privacy field stripping, pixel normalization, and format conversion; an inference module containing a pre-constructed large-scale visual and linguistic model specific to the musculoskeletal disease domain, used to receive pre-processed X-ray images and structured queries, and output a structured diagnostic report containing anatomical location, core disease diagnosis conclusions, and auxiliary information; wherein, the pre-constructed large-scale visual and linguistic model specific to the musculoskeletal disease domain is constructed using a disease-specific visual and linguistic model construction method based on X-rays; a visualization module for generating an attention heatmap based on correlation propagation to display the attention distribution of the model on image regions in the inference module; and an output module for outputting a structured diagnostic report. Compared with existing technologies, this invention has the following technical effects: generating a structured, fine-grained diagnostic output of a large-scale visual and linguistic model specific to the musculoskeletal disease domain, achieving robust diagnostic-level interpretation beyond coarse-grained anomaly screening, and possessing good generalization ability in multi-center, multi-device, and multi-anatomical site scenarios.

[0118] The above description is merely a preferred embodiment of the present invention and does not limit the implementation and protection scope of the present invention. Those skilled in the art should realize that any equivalent substitutions and obvious changes made based on the description and illustrations of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for constructing a large-scale visual and language model for diseases based on X-rays, characterized in that, include: Acquire multi-source, multimodal medical continuous pre-training datasets and dedicated musculoskeletal X-ray fine-tuning datasets; A multimodal base model is obtained, which consists of a visual encoder, a projection layer based on a multilayer perceptron, and a large language model; Using the multi-source multimodal medical continuous pre-training dataset, the visual encoder, projection layer, and large language model of the multimodal base model are continuously pre-trained with full parameter fine-tuning to obtain a medical domain-adapted model. Using the dedicated musculoskeletal X-ray fine-tuning dataset, the medical domain adaptation model is subjected to domain-specific fine-tuning with full parameter fine-tuning, resulting in a large-scale vision and language model specifically for the musculoskeletal disease domain. The acquisition of the multi-source, multimodal medical continuous pre-training dataset and the dedicated musculoskeletal X-ray fine-tuning dataset includes: In the process of acquiring the multi-source multimodal medical continuous pre-training dataset, multimodal medical question answering and image description datasets, endoscopic image datasets, and medical classification and segmentation datasets are collected; for non-multimodal data, a large language model is used in conjunction with background knowledge retrieved from medical knowledge bases and professional literature databases to generate instruction data responses oriented towards medical tasks. In the process of acquiring the dedicated musculoskeletal X-ray fine-tuning dataset: multi-center clinical X-ray image data and corresponding image diagnosis reports from the multi-center hospital clinical system PACS are acquired; public musculoskeletal X-ray datasets are acquired and re-annotated at the diagnostic level; and instruction optimization data is generated through task-oriented transformation. A large-scale visual and language model specifically designed for the musculoskeletal disease domain was obtained through a two-stage, full-parameter fine-tuning training strategy. The first stage utilizes the aforementioned multi-source multimodal medical continuous pre-training dataset for continuous pre-training, adapting the multimodal base model from a general representation to a medical domain distribution. Specifically, using the multi-source multimodal medical continuous pre-training dataset, the three modules—visual encoder, MLP projection layer, and large language model—are fine-tuned with all parameters. After loading the Qwen3-VL-8B-Instruct pre-trained weights, the cross-entropy loss is calculated using task data and backpropagation is performed. Gradient updates are then applied to all parameters, enabling the multimodal base model to deeply adapt from a general representation to the visual feature distribution and professional language expressions of the medical domain. The second stage utilizes the proprietary musculoskeletal X-ray fine-tuning dataset for domain-specific fine-tuning, enabling the medical domain-adapted model to deeply adapt to the professional ontology of musculoskeletal imaging diagnosis. Specifically, based on the medical foundational capabilities established through continuous pre-training, the medical domain-adapted model undergoes full-parameter fine-tuning using the constructed proprietary musculoskeletal X-ray fine-tuning dataset, ensuring that the medical domain-adapted model fully meets the professional needs of clinical musculoskeletal X-ray diagnosis and deeply adapts to the MSK-specific ontology and detailed anatomical knowledge.

2. The method for constructing a large-scale visual and language model for diseases based on X-rays according to claim 1, characterized in that, The method further includes: The multicenter clinical X-ray image data undergoes standardized preprocessing, wherein the standardized preprocessing of the multicenter clinical X-ray image data includes: The DICOM format file is parsed, retaining core image information and removing privacy-related fields; pixel values ​​are linearly normalized and converted; a unique content fingerprint is generated for each image based on a hash function, and hash values ​​are compared across dataset partitions to remove duplicate samples.

3. The method for constructing a large-scale visual and language model for diseases based on X-rays according to claim 1, characterized in that, The method further includes: Establish a structured diagnostic labeling system, wherein establishing the structured diagnostic labeling system includes: The text of the imaging diagnostic report was structured and parsed using a large language model, transforming the free text diagnostic conclusions into standardized diagnostic labels. Senior radiologists with clinical experience conducted a double review and verification process, combining the visual features of the images with the pre-labeled labels. For the public MURA dataset, a comprehensive labeling guideline covering fracture morphology, joint alignment abnormalities, degenerative changes, and soft tissue findings was developed. The guideline was independently reviewed and labeled by at least two committee-certified musculoskeletal radiologists under blinded conditions, with any disagreements being arbitrated by senior radiologists.

4. A method for generating diagnostic information for musculoskeletal diseases based on X-rays, characterized in that, include: Acquire X-ray images of the musculoskeletal system to be processed; The musculoskeletal X-ray images to be processed are input into a pre-constructed large-scale visual and linguistic model specifically for musculoskeletal diseases, resulting in a structured diagnostic output that includes anatomical location, core disease diagnosis conclusions, and auxiliary information. The pre-constructed visual and language large model for musculoskeletal diseases is constructed using the X-ray-based visual and language large model construction method described in any one of claims 1 to 3.

5. The method for generating diagnostic information for musculoskeletal diseases based on X-rays according to claim 4, characterized in that, The method further includes: The structured diagnostic output is scored based on a structured diagnostic labeling system, and the scoring dimensions include: Anatomical location is the first scoring interval, core disease diagnosis is the second scoring interval, and completeness of auxiliary information is the third scoring interval. Among them, large language models are used for automated batch scoring, supplemented by manual sampling assessment by physicians, and intragroup correlation coefficients are used to quantitatively analyze the consistency between AI scores and manual scores.

6. The method for generating diagnostic information for musculoskeletal diseases based on X-rays according to claim 5, characterized in that, The inter-rater consistency between the scores generated by the ICC assessment model based on the two-way mixed effects model and the physician review scores was calculated independently for the scoring dimensions of core disease diagnosis, anatomical location and completeness of auxiliary information. Based on the central limit theorem, the 95% confidence interval of the population mean is estimated using the sample standard error. The formula for calculating the confidence interval is as follows: , in, ; It is the first The score for each question It is the average score.

7. The method for generating diagnostic information for musculoskeletal diseases based on X-rays according to claim 4, characterized in that, The method further includes: Attention maps of each attention layer of the large-scale visual and language model specifically designed for musculoskeletal diseases are obtained. Attribution correlation maps are calculated based on correlation propagation, and these maps are upsampled to the original image size to generate a visual heatmap. The calculation method for the attribution correlation maps is as follows: in This refers to the Hadamard product. Let be the gradient with respect to A. It is the model's output for the category to be visualized. This indicates that the mean is taken along the head dimension; the update rule for self-attention is... The aggregation rule for bimodal attention is: The final correlation map is upsampled to the original image size using bilinear interpolation to generate a visual heatmap, which is used to show the correlation between image regions and each generated token.

8. An intelligent diagnostic system for musculoskeletal diseases based on X-rays, characterized in that, include: The data acquisition module is used to acquire X-ray images of the musculoskeletal system to be diagnosed. The preprocessing module is used to perform standardized preprocessing on the X-ray images, including privacy field stripping, pixel normalization, and format conversion; The inference module includes a pre-built visual and linguistic model specifically for musculoskeletal diseases, used to receive pre-processed X-ray images and structured queries, and output a structured diagnostic report containing anatomical location, core disease diagnosis conclusions, and auxiliary information; wherein, the pre-built visual and linguistic model specifically for musculoskeletal diseases is constructed by the method described in any one of claims 1 to 3; The visualization module is used to generate an attention heatmap based on correlation propagation, which shows the distribution of the model's attention to the image region in the inference module; The output module is used to output the structured diagnostic report.

9. The X-ray-based intelligent diagnostic system for musculoskeletal diseases according to claim 8, characterized in that, The reasoning module supports anatomical location identification query, anomaly detection query, disease type classification query, and open diagnostic report generation query modes. The anatomical location identification query identifies the anatomical location corresponding to the input image from a candidate set containing multiple body parts. The anomaly detection query determines whether there is a specified type of abnormality in the image. The disease type classification query performs fine-grained classification of fracture types and bone and joint diseases. The open diagnostic report generation outputs a complete structured radiological diagnostic conclusion, following the full classification system of orthopedic diseases, covering core disease categories such as traumatic bone injury, degenerative bone and joint diseases, metabolic bone diseases, and neoplastic bone lesions.