Method for establishing medical image data set with uniform structure
Through the private AI model and BERT-NER model, unstructured medical image reports and DICOM data are processed, and a unified structure of medical image data sets are generated, which solves the problems of high construction cost, poor consistency and weak cross-modal analysis capabilities in the existing technology, and realizes efficient data integration and analysis.
Patent Information
- Application Number
- CN202510615877.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing medical imaging data sets have high cost of building, poor data consistency, and weak cross-modal analysis capabilities, resulting in poor training effects of artificial intelligence models and limited clinical application value.
The text of the unstructured image report is processed into structured data through a private AI model, and merged with the DICOM metadata to generate an SR file to establish a medical image data set with a unified structure. Use the BERT-NER model to map user inputs into medical feature encoding, supporting cross-modal queries.
It realizes the standardized integration of medical imaging data, ensures the integrity and consistency of data, improves the operability and parsability of data, enhances the interoperability between different systems, and supports efficient data analysis and clinical decision-making.
Smart Images

Figure CN120148727A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of medical image processing and analysis, and specifically relates to a method for establishing a medical image dataset with a unified structure. Background Art
[0002] The standardized integration and efficient management of medical image data are the core challenges for realizing precision medicine and artificial intelligence-assisted diagnosis. The difficulty in data standardization lies in the randomness of annotation and the lack of unified data structure. In current clinical practice, imaging diagnosis reports are mostly stored independently in the RIS system in the form of unstructured free text, lacking an automated association mechanism with DICOM format image data. Traditional manual annotation methods have low processing efficiency, and the differences in understanding of text descriptions by different physicians result in insufficient consistency of annotation results. In addition, existing imaging annotation work usually needs to be carried out separately outside the diagnostic business process, further exacerbating the problems of low data annotation efficiency and uneven quality. These factors together lead to technical bottlenecks in medical image datasets, such as high construction costs, poor data consistency, and weak cross-modal analysis capabilities, seriously restricting the training effect and clinical application value of artificial intelligence models.
[0003] Due to the low processing efficiency of unstructured reports in the prior art, different doctors may use different terms and description methods when writing imaging reports, making it difficult to ensure the consistency and accuracy of data. There is a lack of an effective automated association mechanism between image data and corresponding diagnostic reports, resulting in a serious information island phenomenon and inconsistent terminology systems between different medical systems, thus making it difficult for traditional technologies to achieve automated association and integration of images and reports. In addition, although existing images and reports can help doctors learn about symptoms, the prior art cannot push images and reports of different levels of difficulty to doctors according to their proficiency, lacking personalized design for doctors' learning. Summary of the Invention
[0004] To solve at least one aspect of the technical problems in the background art, this application provides a method for establishing a medical image dataset with a unified structure.
[0005] The technical solution adopted in this application is as follows: The first aspect embodiment of this application provides a method for establishing a medical image dataset with a unified structure, including: Based on the unstructured imaging report and the corresponding DICOM format imaging data, the unstructured imaging report text, DICOM images, and DICOM metadata are extracted; the unstructured imaging report text is processed by a private AI model to obtain structured data, and the structured data is merged with the DICOM metadata to obtain an SR file, where the SR file includes the DICOM metadata, medical feature codes, and unique identifiers; based on the unique identifier, the SR file is associated with the DICOM images to form a structured dataset; the medical feature codes in the SR file are associated with the unique identification code, and the user input is mapped to the medical feature code through a BERT-NER model to query the structured dataset.
[0006] According to an embodiment of the present application, the unstructured imaging report text is processed by a private AI model to obtain structured data, and the structured data is merged with the DICOM metadata to obtain an SR file, where the SR file includes the DICOM metadata, medical feature codes, and unique identifiers, specifically: Based on medical standard terms, a structured template is defined; the private AI model is supervised and fine-tuned with an initial dataset manually labeled, and the output of the private AI model is iteratively optimized to match the medical standard terms; the unstructured imaging report text is parsed by the private AI model into the structured data that conforms to the structured template; according to the DICOM SR standard, a data template is defined to determine that the structured data matches the unique identifier of the DICOM metadata; based on the DICOM SR standard, the SR file is initialized; the basic information in the DICOM metadata is filled into the basic metadata fields of the SR file, and the basic information includes at least one of patient information and examination parameters; the structured data is traversed to extract medical feature codes and related information; the medical feature codes are formatted according to the requirements of the DICOM SR standard; the formatted medical feature codes and related information are embedded into the content sequence of the SR file; in addition to the basic metadata fields, the DICOM metadata is also added to the relevant fields of the SR file; the fields are checked, and a DICOM verification tool or library is used to check whether the generated SR file conforms to the DICOM SR standard.
[0007] According to an embodiment of the present application, based on the unique identifier, the SR file is associated with the DICOM images to form a structured dataset, specifically: Confirm that the unique identifier of the SR file is consistent with the unique identifier of the DICOM image; construct a structured report form to store relevant information of the SR file; construct a feature coding table to store the medical feature codes extracted from the SR file; upload the SR file to the PACS system, use the unique identifier to find and associate the corresponding DICOM image, or establish an association between the structured report form and the DICOM image metadata by using the unique identifier as a foreign key; associate the DICOM image with the SR file through the unique identifier; associate the SR file with the medical feature codes in the feature coding table through the unique identifier; support cross-modal queries based on symptoms through the medical feature codes, and establish a composite index on the feature codes to accelerate queries based on the medical feature codes.
[0008] According to an embodiment of the present application, the association between the medical feature codes in the SR file and the unique research instance identifier, and mapping the user input to the medical feature codes through the BERT-NER model to query the structured data set are specifically as follows: Use the BERT-NER model to parse the natural language description input by the user to identify key medical entities; convert the key medical entities into the corresponding medical feature codes; use the medical feature codes to construct an SQL query statement to retrieve relevant unique identifiers; based on the unique identifiers, find the corresponding SR file in the structured report form; use the same unique identifier to find the relevant DICOM image in the PACS system or the DICOM image database; integrate the SR file and the DICOM image to form a structured data set.
[0009] According to an embodiment of the present application, it further includes: When obtaining the unstructured image report and the corresponding DICOM format image data, record the learning identity identifier of the doctor; evaluate the basic ability of the doctor through a standardized test, and convert the test result into a proficiency score through a machine learning model; record the operation behavior data of the doctor, and use an online learning algorithm to update the proficiency score in real time; preset structured templates and rules for different proficiency scores.
[0010] According to an embodiment of the present application, the presetting of structured templates and rules for different proficiency scores is specifically as follows: If the proficiency is the first-level proficiency, split the structured data into standardized fields and add annotation descriptions; if the proficiency is the second-level proficiency, reduce the detail level of the structured data, provide a summary structured report, and retain the key conclusions.
[0011] According to an embodiment of the present application, the processing of the unstructured imaging report text by the private AI model to obtain structured data further includes: Dynamically generate prompting words according to proficiency to control the detail level of the output of the private AI model; if the proficiency is the first-level proficiency, fine-tune the private AI model using preset granularity annotation data; if the proficiency is the second-level proficiency, fine-tune the private AI model using summary annotation data.
[0012] According to an embodiment of the present application, the merging of the structured data and the DICOM metadata to obtain the SR file further includes: When generating the SR file, add a proficiency identification field to mark the structured degree of the SR file.
[0013] According to an embodiment of the present application, after associating the medical feature encoding in the SR file with the unique identification code and mapping the user input to the medical feature encoding through the BERT-NER model to query the structured data set, it further includes: recommending differentiated learning content according to the proficiency score of the doctor, specifically: If the proficiency is the first-level proficiency, preferentially push structured reports and basic cases, combined with interactive learning tools; if the proficiency is the second-level proficiency, provide mixed content and introduce complex cases and diagnostic reasoning training; recommend cases of similar difficulty based on the doctor's historical learning behavior.
[0014] According to an embodiment of the present application, after recommending differentiated learning content according to the proficiency score of the doctor, it further includes: Use behavior indicators and ability indicators as learning effect quantification indicators, dynamically adjust the differentiated learning content based on the learning effect quantification indicators; adjust the keyword weights retrieved according to the doctor's proficiency, and provide an explanation of the retrieval results for doctors with second-level proficiency.
[0015] Due to the adoption of the above technical solutions, the beneficial effects obtained by the present application are: This application ensures that all relevant information is extracted from the original medical imaging examinations, including detailed imaging report texts, actual DICOM images, and DICOM metadata containing patient information, examination details, etc. This provides a complete data basis for subsequent data processing and analysis. By extracting unstructured imaging report texts and DICOM images and metadata simultaneously, the integrity and consistency of the data are ensured, enabling each imaging report to be accurately matched with its corresponding medical image. This application converts unstructured imaging reports into structured data, enabling key information (such as lesion location, mass size, etc.) to be accurately identified and utilized by the system, greatly improving the operability and analyzability of the data. By combining the structured data with DICOM metadata to generate SR files, the effective combination of imaging data and report content is achieved, facilitating unified management and retrieval. Using medical feature coding to standardize the representation of specific medical terms not only enhances the interoperability between different systems but also provides strong support for subsequent data analysis and clinical decision-making. This application uses a unique identifier (UID) to ensure that each SR file can be accurately matched with its corresponding DICOM image, avoiding data confusion or incorrect association. Through the association relationship established by the UID, users can conveniently and quickly query the relevant SR files and their corresponding DICOM images according to their needs, greatly facilitating the doctor's workflow. The formed structured dataset not only contains detailed medical images but also includes processed structured report content, further enhancing the value and usability of the data. Through the BERT-NER model in this application, users only need to input a natural language description to quickly locate the required medical feature coding and query the structured dataset accordingly, greatly simplifying the process of complex queries. The BERT-NER model can accurately identify and convert key medical entities in the user input into corresponding medical feature coding, reducing the errors that may occur in manual searches and improving the accuracy of query results. This query method based on natural language understanding and medical feature coding can provide customized search results according to the specific needs of different users, better meeting the actual work needs of clinicians. Brief Description of the Drawings
[0016] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings: Figure 1 It is a schematic flowchart of a method for establishing a medical imaging dataset with a unified structure provided by an embodiment of the present application; Figure 2 It is a schematic specification of key information extraction from an unstructured report of a method for establishing a medical imaging dataset with a unified structure provided by an embodiment of the present applicationFigure 1 ; Figure 3 It is a specification schematic for key information extraction of unstructured reports in a method for establishing a medical image dataset with a unified structure provided by an embodiment of the present application. Figure 2 . Detailed implementation manners
[0017] Embodiment 1 As Figure 1 shown, an embodiment of the first aspect of the present application provides a method for establishing a medical image dataset with a unified structure, including: S100. Based on the unstructured image report and the corresponding DICOM format image data, extract the unstructured image report text, DICOM image, and DICOM metadata.
[0018] As described above, for the acquisition of the unstructured image report text, directly obtain the relevant image reports from the radiology information system (RIS). These reports usually exist in the form of free text and contain information such as doctors' descriptions, analyses, and diagnostic opinions on the image results. This part of the content is unstructured, meaning it has no fixed format or classification method, so subsequent processing is required to be effectively used for data analysis or machine learning model training. Regarding the acquisition of DICOM images, also obtain the medical digital imaging and communications (DICOM) format image data corresponding to the above reports from the RIS system. DICOM is a standard format for storing, exchanging, and transmitting medical images and their related information in the medical field. Through this process, actual medical image data can be ensured, and these images are the basis for medical diagnosis. When it comes to the extraction of DICOM metadata, that is, extract the relevant metadata information from the DICOM format image data. The metadata includes but is not limited to the patient's personal information (such as name, age, gender, etc.), examination parameters (such as scan type, scan time, etc.), and other information that helps to understand and use these image data. These metadata are very important for determining the source of the image, understanding its background information, and associating the image with a specific patient. By extracting this information, the context environment of each image data can be more comprehensively understood, providing necessary support for subsequent data processing and analysis.
[0019] For example, obtain the relevant imaging report of this patient from the Radiology Information System (RIS). This report is the doctor's description of the results of this mammogram, containing information such as breast density, lesion location, mass shape, and margin status. These descriptions are unstructured free text, such as "A round mass with clear margins was found in the upper outer quadrant of the right breast." Obtain the corresponding DICOM-format imaging data from the same system. This means obtaining the actual mammogram image files, which were taken under specific conditions, such as using a certain type of X-ray device and scanned according to set parameters. It is necessary to extract metadata from the DICOM imaging data. These metadata include basic information about the patient, such as name, age, gender, and medical record number; they also include specific parameters about this examination, such as the scanning method used, scanning time, exposure conditions, etc. For example, it can be learned that this patient is a 45-year-old female, and the examination was conducted on March 8, 2025, using Digital Breast Tomosynthesis (DBT). Through this process, not only is a detailed unstructured imaging report text obtained - that is, the doctor's written description of the mammogram results, but also the actual DICOM images, which are the mammogram image files themselves, and various related metadata information. These metadata help us understand the background information of the images, such as the patient's personal information and technical parameters during the examination, thus providing a comprehensive information basis for subsequent possible data analysis, diagnostic support, or research. In this way, the unstructured text description, specific medical images, and their related detailed background information can be effectively combined to prepare for further applications.
[0020] It should be noted that in a specific implementation scenario, based on the above solution, advanced natural language processing techniques can also be adopted to deeply analyze the unstructured imaging report text. For example, through Named Entity Recognition (NER) technology, key medical terms and concepts, such as disease names, anatomical locations, lesion characteristics, etc., are automatically identified and extracted from these free texts. This helps to transform the originally unstructured information into structured data for subsequent data retrieval and analysis. Use a specific standard terminology system in the medical field (such as SNOMED CT, RadLex, etc.) to semantically annotate the extracted key information and convert it into standard codes. This process can not only improve the consistency and accuracy of information but also promote interoperability between different systems, making it more convenient to share and integrate data from different medical institutions.
[0021] In a specific implementation scenario, based on the above solution, a knowledge graph in the field of medical imaging can be constructed based on the extracted structured data and the information after encoding conversion. The knowledge graph can display the complex relationships between medical concepts, such as the relationship between diseases and symptoms, the comparison of the effects of treatment plans, etc., thereby providing a strong knowledge basis for clinical decision support.
[0022] S200. Process the unstructured imaging report text through a private AI model to obtain structured data, and merge the structured data with the DICOM metadata to obtain an SR file, where the SR file includes the DICOM metadata, medical feature encoding, and unique identifiers.
[0023] As described above, the process of processing unstructured imaging report text through a private AI model to obtain structured data involves several key steps. First, the original imaging reports obtained from the Radiology Information System (RIS) exist in free text form, which contain various descriptions, analyses, and conclusions of the imaging results by doctors. Since this information is unstructured, it is difficult to directly use it for data analysis or machine learning. The unstructured imaging report text is processed using a private AI model deployed on a local server. This private AI model is trained and fine-tuned with a large amount of labeled data, and the specific process is as follows: Based on authoritative medical domain standard terminology systems (such as SNOMED CT, RadLex, etc.), combined with the actual needs of clinical diagnosis and treatment, a standardized extraction template is constructed. This template precisely defines the medical entities to be extracted from the imaging report and their logical relationships, forming a normative framework for extracting key information from unstructured reports. This step ensures the accurate identification and extraction of key information in the subsequent processing. According to the above standardized template, prompt engineering is carried out to design professional prompt words suitable for large language models, and the task instructions for the model to process unstructured reports are clarified. At the same time, some unstructured reports are manually labeled based on the template to generate an initial labeled dataset, providing a basis for supervised learning in model training. This step is very important because it directly determines whether the model can correctly understand the task and execute it efficiently. Using the prompt words and the initial labeled dataset, the first round of SFT is performed on the large language model. Through this round of optimization, the model can start processing a large number of unstructured medical imaging reports and generate a large-scale structured report, thus forming a rich labeled dataset and further expanding the data scale of supervised learning. An iterative training strategy is adopted, and multiple rounds of supervised fine-tuning are continuously repeated. Each iteration inputs the professional prompt words designed in the previous round and the large-scale labeled data generated into the model. By continuously calibrating the consistency between the model output and medical professional standards, the model performance is gradually optimized, and finally a target private AI model that meets the structured requirements of medical imaging reports is generated. This method helps to continuously improve the accuracy and reliability of the model in specific tasks. The target model after multiple rounds of fine-tuning is deployed to the local server to build a local processing system. Based on this system, the obtained unstructured imaging reports can be structurally processed to generate structured imaging reports that meet clinical standards, providing standardized data support for the construction of medical imaging datasets. It is specifically optimized for the structured processing task of imaging reports as follows: According to the characteristics of imaging reports, specific prompt words are designed to guide the model to understand specific task requirements, such as identifying lesion locations and describing mass characteristics. During the processing, medical terms are converted into a unified coding form (such as using SNOMED CT or RadLex coding) to improve the interoperability and data sharing efficiency between different medical systems.By continuously adjusting the model parameters and reviewing the results, a quality control mechanism is established to ensure the quality of the model output. At the same time, the model is continuously improved based on the feedback in actual applications. The model can understand medical terms and identify key information in the imaging report, such as lesion location, mass size, and morphology. The processing process includes but is not limited to extracting specific medical entities and their logical relationships, converting unstructured text into a data format with a clear structure, such as organizing information according to a predefined template, ensuring that the data in each part is clear and easy to parse. Once the structured data is obtained, the next step is to merge it with the metadata extracted from DICOM images. DICOM metadata includes basic patient information (such as name, age, gender), examination parameters (such as scanning method, time), and other information that helps to understand and use the imaging data. In this way, the content of the structured imaging report can be associated with the patient's detailed information and the specific circumstances of image acquisition, forming a more comprehensive dataset. Generate an SR file that complies with the DICOM SR (Structured Reporting) standard specification. Such an SR file not only contains the above-mentioned DICOM metadata but also incorporates medical feature codes and unique identifiers converted from structured data. Among them, the medical feature codes are the results of standardizing medical terms in the report using internationally or domestically recognized medical terminology standards (such as SNOMED CT, ICD, RadLex, etc.), which helps to interoperability and data sharing between different medical systems.
[0024] The unique identifiers include: Patient UID (Patient ID), which is used to uniquely identify each patient. Regardless of any activities or records of the patient in the medical institution (such as examinations, diagnoses, treatments, etc.), they can be associated through this ID. Ensure that all medical information related to this patient can be accurately filed and retrieved, avoiding confusion of records of different patients.
[0025] Study Instance UID, each medical imaging examination has a unique Study Instance UID, which runs through the entire examination process, including all imaging series and images. Ensure that all relevant data for each examination (including DICOM images and SR files) can be accurately associated together, supporting data consistency and integrity.
[0026] Series Instance UID, which is used to identify different image series in the same examination. For example, in a chest CT scan, there may be multiple different series, and each series has its own Series Instance UID. It helps to distinguish and manage images of different types or angles within the same examination, facilitating more precise data management and querying.
[0027] SOP Instance UID (SOP Instance UID for images), each individual DICOM image has a unique SOP Instance UID, which identifies the specific image file. It ensures the uniqueness of each image file within its series and examination, facilitating the precise positioning and management of individual image files. SR Document Instance UID (SR Document Instance UID), a unique identifier is assigned to each structured report (SR file). This enables each SR file to be uniquely identified in the system and to establish connections with other associated DICOM images and other SR files.
[0028] For example, an unstructured imaging report of the patient is obtained from a Radiology Information System (RIS). The report content may be as follows: "A round mass with clear margins was found in the upper outer quadrant of the right breast, approximately 1.2×0.6 cm in size, suspected of benign calcification." The above unstructured imaging report is processed using a private AI model deployed on a local server. This model is trained and fine-tuned with a large amount of labeled data and is specifically used to identify and extract key information from medical imaging reports. In this example, the model can identify the following information: lesion location: upper outer quadrant of the right breast, mass shape: round, margin status: clear, size: 1.2×0.6 cm, possible feature: benign calcification. This information is converted into a structured format. For example, "lesion location", "mass shape", etc. are used as independent data fields, and standard medical terminology coding (such as RadLex coding) is used to represent each feature. At the same time, metadata is extracted from the DICOM image file of the same patient. This metadata includes, but is not limited to: patient name: Zhang San, patient ID: 000123456, gender: female, age: 45 years old, examination date: March 8, 2025, scanning method: Digital Breast Tomosynthesis (DBT). Next, the structured data obtained through the private AI model is combined with the DICOM metadata to generate a structured report (SR) file that complies with the DICOM SR standard specification. This SR file not only contains the above DICOM metadata but also includes medical feature coding and a unique identifier. For example, for the key information mentioned above, "lesion location" is encoded with a specific RadLex coding, and "mass shape" is also encoded accordingly. In addition, the SR file will also contain an SR file instance UID to ensure that each SR file has a unique identity in the entire dataset, facilitating subsequent data management and retrieval.
[0029] It should be noted that in a specific implementation scenario, on the basis of the above solution, an automated rule-based review system can also be developed. This system can check whether the generated structured data meets predefined standards (such as the consistency of medical terms, the accuracy of logical relationships, etc.). If any abnormalities or non-compliant situations are found, the system will mark them for manual review. The model is continuously improved using a feedback loop. Whenever an error or adjustment is needed, this information can be fed back into the training dataset to fine-tune the private AI model and improve its accuracy and robustness.
[0030] S300. Based on the unique identifier, associate the SR file with the DICOM image to form a structured dataset.
[0031] As described above, the SR file instance UID is a crucial identifier that ensures each SR file has a unique identity within the entire dataset. The UID follows specific encoding rules, enabling it to accurately locate and distinguish each report during subsequent data management and retrieval. For DICOM images, they also have their unique UIDs, which contain information about key data points such as patient information and examination instances. Through the examination instance UID, one or more SR files can be precisely matched with the corresponding DICOM images. When generating an SR file, a unique UID is assigned to each file. This UID is the SR file instance UID, which can identify the SR file itself, and the examination instance UID contains information related to the associated DICOM image. Similarly, DICOM images also come with an examination instance UID, which is automatically generated during image creation and runs through the entire medical process. The examination instance UID is used to establish data association relationships in the PACS (Picture Archiving and Communication Systems). Specifically, based on the examination instance UID, the SR file is linked to the corresponding DICOM image. This association method ensures that whenever accessing the image data of a certain patient, the detailed structured report of the image can be obtained synchronously. When the SR file and the DICOM image are successfully associated, a structured dataset containing rich information is formed. This dataset not only includes the original medical images but also the processed structured report content, which covers detailed medical findings, measurement data, and other important information. In addition, since all data is standardized and associated through UIDs, the consistency and traceability of the data are greatly improved.
[0032] For example, after performing a mammography examination, the system generates image data in DICOM format. These image data contain detailed medical images and some metadata, such as the patient's name (Li Hua), gender (female), age (48 years old), medical record number (123456789), and examination date (April 20, 2025), etc. Each DICOM image has a unique identifier (examination instance UID). For example, the main UID for this examination may be "1.3.6.1.4.1.9590.100.1.2.1234567890". The doctor writes an unstructured image report based on the examination results, describing some key points found during the examination, such as: "A round mass with clear edges was found in the upper outer quadrant of the right breast, approximately 1.2×0.6 cm in size, suspected of benign calcification." This report is then converted into a structured format by a private AI model and an SR file compliant with the DICOM SR standard is generated. This SR file also has a unique identifier (examination instance UID), such as "2.16.840.1.113669.632.20.123456789012345678901". In the PACS system, based on the unique identification information of the examination instance UID, the generated SR file is associated with the corresponding DICOM image. Specifically, since both contain the same examination instance UID, they can be easily identified and matched. In this example, "1.3.6.1.4.1.9590.100.1.2.1234567890" exists as the examination instance UID in both the DICOM image and the SR file, ensuring their accurate correspondence. Once this association is established, a structured data set containing rich information is formed. This data set not only includes the original mammography images, but also contains detailed structured report content, such as detailed descriptions of the lesion location, mass size, edge status, and possible characteristics, etc. All this information is uniformly managed and stored for easy subsequent query and analysis.
[0033] It should be noted that in a specific implementation scenario, based on the above solution, automated scripts or tools can also be developed to automatically check whether the examination instance UIDs of the SR file and the DICOM image match each time an association is established, and confirm that all necessary metadata fields exist and are correct. Create detailed log records for each data association operation, including information such as the operation time, the person performing the operation, and the association result. This helps to track potential problems and ensure the integrity and traceability of the data.
[0034] S400. Associate the medical feature codes in the SR file with the unique identification codes, and map the user input to the medical feature codes through the BERT-NER model to query the structured data set.
[0035] As described above, first, when generating an SR file that complies with the DICOM SR standard specification, each medical feature (such as lesion location, mass size, etc.) is assigned a specific medical feature code. These codes are usually based on internationally or domestically recognized standard terminology systems (such as SNOMED CT, RadLex, etc.). At the same time, the SR file itself and its associated DICOM images have unique identifiers (study instance UID) to uniquely identify them in the entire dataset. In the SR file, the medical feature codes not only record the specific medical information but also are associated with the study instance information through the study instance UID. For example, a certain SR file may contain a description of "a round mass with clear margins found in the upper outer quadrant of the right breast", where key features such as "right breast", "upper outer quadrant", "clear margins", and "round mass" are all converted into corresponding medical feature codes, and these codes are closely linked to the study instance UID of this SR file. When users want to query specific medical images or reports, they may enter some natural language keywords, such as "right breast mass". At this time, the system will use the BERT-NER model to process the user input. The BERT-NER model is an advanced natural language processing technology that can identify and extract named entities in the text and convert them into a predefined coding format. Specifically, the BERT-NER model analyzes the natural language keywords entered by the user, identifies the key medical terms among them, and maps these terms to the corresponding medical feature codes according to the trained model parameters. For example, "right breast mass" can be parsed into multiple medical features including "right breast" and "mass", and the corresponding codes are found, such as "RID29896" representing "right breast" and "RID39055" representing "mass". Once the medical feature codes corresponding to the user input are obtained, the system will perform a search in the structured database based on these codes. The system first looks for records containing all the target codes in the feature table (sr_features), and then aggregates these features through the lesion table (sr_lesions) to ensure that all features belong to the same lesion (lesion_id). Subsequently, the system associates with the report table (sr_reports) according to the structured report identifier (sr_id) to which the lesion belongs to locate the specific SR file and its corresponding DICOM image (through the study_instance_uid field). For example, if the user input is a query request about "right breast mass", the codes obtained after being processed by the BERT-NER model (such as "RID29896" and "RID39055"). The system then retrieves the records containing this set of codes in the sr_features table and filters out the lesions that contain these features in the same lesion structure through the lesion_id to which the features belong.Next, associate through the sr_id in the sr_lesions table with the sr_reports table to further locate the study_instance_uid corresponding to the report, obtain the complete structured report content and its corresponding DICOM images, and finally return the information required by the user.
[0036] For example, after a mammography examination, the image report written by a doctor is converted into a structured format by a private AI model, and an SR file compliant with the DICOM SR standard is generated. The SR file details that "a well-defined round mass was found in the upper outer quadrant of the right breast, with a size of approximately 1.2×0.6 cm". In this process, key medical features (such as "right breast", "upper outer quadrant", "well-defined", "round mass") are converted into specific medical feature codes (for example, "RID29896" represents "right breast", "RID5707" represents "well-defined", "RID34240" represents "high density", etc.). At the same time, the SR file has a unique identifier (UID), such as "2.16.840.1.113669.632.20.123456789012345678901". In the PACS system, based on the unique identification information of the examination instance UID, the generated SR file is associated with the corresponding DICOM image. In this way, whether viewing the original mammography image or the detailed structured report content, it can be accessed through the same examination instance UID. A doctor wants to search for relevant case materials regarding "right breast mass". He enters the natural language keyword "right breast mass". After the system receives this query request, it uses the BERT-NER model to analyze the input natural language keyword. The BERT-NER model can identify the key medical terms therein - "right breast" and "mass", and map these terms to the corresponding medical feature codes. For example, "right breast" is mapped to "RID29896", and "mass" may correspond to one of multiple codes, such as "RID39055". Query the structured dataset: With these medical feature codes, the system searches for records in the sr_features table that match "RID29896" and "RID39055". Subsequently, these features are aggregated by lesion_id, and further filtered to select lesion records that contain all the target codes in the same lesion structure. The system then associates through the sr_id field in the sr_lesions table to the sr_reports table to locate the corresponding report content and its associated DICOM image data. In this example, the system will return all relevant information regarding the mammography of Ms. Wang, including the specific image file and the detailed structured report. Finally, the doctor can view Ms. Wang's mammography image and also access the detailed structured report to learn detailed information such as the specific location (upper outer quadrant of the right breast), shape (round), edge status (well-defined), and size (1.2×0.6 cm) of the mass. This enables the doctor to quickly and accurately obtain the required information to support more precise diagnosis or treatment decisions.
[0037] It should be noted that in a specific implementation scenario, based on the above solutions, image content analysis (such as analyzing DICOM images through a deep learning model) and text descriptions (i.e., reports in SR files) can be combined to provide a more comprehensive understanding of the context. For example, when processing a query about a "right breast mass", in addition to identifying keywords in the text description, the features in the actual image can also be analyzed to ensure more accurate results. Knowledge graphs or semantic networks are used to enrich the understanding of medical terms, enabling the system to not only identify specific entities but also understand the relationships between these entities. For example, the relationship between the "right breast" and "benign calcification" can help the system better understand and respond to complex query requests.
[0038] In some embodiments of the present application, the unstructured image report text is processed by a private AI model to obtain structured data, and the structured data is merged with the DICOM metadata to obtain an SR file. The SR file includes the DICOM metadata, medical feature codes, and unique identifiers, specifically: Based on medical standard terms, a structured template is defined; the private AI model is fine-tuned in a supervised manner using an initial dataset with manual annotations, and the output of the private AI model is iteratively optimized to match the medical standard terms; the unstructured image report text is parsed by the private AI model into the structured data that conforms to the structured template; according to the DICOM SR standard, a data template is defined to ensure that the structured data matches the unique identifier of the DICOM metadata; based on the DICOM SR standard, an SR file is initialized; the basic information in the DICOM metadata is filled into the basic metadata fields of the SR file; the structured data is traversed to extract medical feature codes and related information; the medical feature codes are formatted according to the requirements of the DICOM SR standard; the formatted medical feature codes and related information are embedded into the content sequence of the SR file; In addition to the basic metadata fields, the DICOM metadata is also added to the relevant fields of the SR file; the fields are checked, and DICOM verification tools or libraries are used to check whether the generated SR file conforms to the DICOM SR standard.
[0039] As described above, based on authoritative medical standard terminology systems (such as SNOMED CT, RadLex, etc.), a structured template is defined. This template details the key information to be extracted from unstructured imaging reports and its format, such as lesion location, mass size and shape, etc. This step ensures that key information can be accurately identified and extracted during subsequent processing and converted into a data format with a clear structure. The private AI model is fine-tuned in a supervised manner using an initial dataset with manual annotations. This process includes: selecting some unstructured imaging reports for manual annotation to generate an initial annotated dataset. Using these annotated data to perform the first round of fine-tuning (Supervised Fine-Tuning, SFT) on the large language model. The optimized model is used to process a large number of unstructured medical imaging reports to generate large-scale structured reports. An iterative training strategy is adopted, repeating multiple rounds of supervised fine-tuning to continuously calibrate the consistency between the model output and medical professional standards, gradually optimizing the model performance, and finally generating a target private AI model that meets the requirements of structuring medical imaging reports. The fine-tuned private AI model is used to parse the text of unstructured imaging reports and convert it into structured data that conforms to the above-defined structured template. For example, a description like "A round mass with clear margins is found in the upper outer quadrant of the right breast" is converted into structured data containing specific fields such as "Lesion location: Upper outer quadrant of the right breast", "Mass shape: Round", "Margin status: Clear". Define a data template according to the DICOM SR standard to ensure that the structured data matches the unique identifier in the DICOM metadata. This means that when generating the SR file, all relevant information (examination instance UID) must have a consistent and unique identifier for subsequent data management and retrieval. Based on the DICOM SR standard, initialize a new SR file. This file follows the data storage format of the DICOM standard, ensuring that each byte and each field conforms to the standard definition and specification. Fill the basic information in the DICOM metadata into the basic metadata fields of the SR file. This information includes but is not limited to the patient's name, age, gender, medical record number, examination date, and scanning method, etc. These metadata are crucial for understanding the source, background of the image and its association with the patient. Traverse the structured data and extract the medical feature codes and related information. For example, descriptions such as "Right breast" and "Round mass" are converted into specific medical feature codes (such as "RID29897" represents "Right breast", "RID39055" represents "Mass"). Format the extracted medical feature codes according to the requirements of the DICOM SR standard. This step ensures that the codes can be correctly embedded in the content sequence of the SR file and can achieve interoperability between different systems. Embed the formatted medical feature codes and related information into the content sequence of the SR file.In this way, the SR file not only contains the basic metadata, but also detailed medical findings, measurement data, and other important information. In addition to the basic metadata fields, other relevant DICOM metadata are added to the relevant fields of the SR file. This further enriches the content of the SR file, making it a comprehensive document reflecting the patient's examination. Finally, the fields are checked, and the generated SR file is checked using a DICOM validation tool or library to see if it conforms to the DICOM SR standard. This step ensures the quality and consistency of the SR file, enabling it to be seamlessly shared and used in different medical information systems.
[0040] In some embodiments of the present application, associating the SR file with the DICOM image based on the unique identifier to form a structured data set specifically includes: Confirming that the unique identifier (examination instance UID) of the SR file is consistent with the unique identifier (examination instance UID) of the DICOM image; constructing a structured report form to store the relevant information of the SR file; constructing a feature coding table to store the medical feature codes extracted from the SR file; uploading the SR file to the PACS system, using the unique identifier (examination instance UID) to find and associate the corresponding DICOM image, or establishing an association between the structured report form and the DICOM image metadata through the unique identifier (examination instance UID) as a foreign key; associating the DICOM image with the SR file through the unique identifier (examination instance UID); associating the SR file with the medical feature codes in the feature coding table through the unique identifier (examination instance UID); supporting cross-modal queries based on symptoms through the medical feature codes, and establishing a composite index on the feature codes to accelerate queries based on the medical feature codes.
[0041] As described above, it is necessary to confirm whether the unique identifier (examination instance UID) in the SR file is consistent with the unique identifier (examination instance UID) in the DICOM image. These UIDs are assigned to each individual examination instance at the time of generation and remain unchanged throughout the medical process. By comparing the examination instance UIDs, it can be ensured that the SR file and the corresponding DICOM image indeed belong to the same examination. Since an SR report may involve multiple lesions, in order to manage the structured information at the level of each lesion, the system needs to construct a lesion table (sr_lesions). This table is used to record the structured location and attribution information of each lesion in the report. The main fields include: lesion unique identifier (lesion_id), the structured report ID to which it belongs (sr_id), lesion sequence number (lesion_order), and lesion text description (description, optional). This table establishes a foreign key association with the structured report table through sr_id to ensure that each lesion can be accurately traced back to its source report. The system also needs to construct a feature coding table (sr_features) to store all the medical feature codes extracted from the SR file. For example, descriptions such as "right breast" and "round mass" will be converted into standardized medical feature codes (e.g., "RID29896" represents "right breast" and "RID39055" represents "mass"). The sr_features table is associated with the lesion table (sr_lesions) through a foreign key to ensure that each feature code can be traced back to its corresponding lesion structure (through lesion_id) and to ensure that each feature code can be traced back to its source SR file. Next, the SR file is uploaded to the PACS system. During the upload process, the unique identifier (examination instance UID) is used to find and associate the corresponding DICOM image. This means that when the SR file is uploaded, the system will automatically find the matching DICOM image according to the examination instance UID and establish an association between the two. The purpose of doing this is to ensure that when accessing the imaging data of a certain patient each time, the detailed structured report of the image can be obtained synchronously. Through the examination instance UID, not only can a direct association be established between the DICOM image and the SR file, but also an indirect association can be established between the SR file and the medical feature codes in the feature coding table. Specifically: Association between DICOM images and SR files: By examining the instance UID, it is possible to determine which DICOM images are related to a specific SR file. Association between SR files and feature encodings: By using the instance UID as a foreign key, a connection can be established between the structured report table and the feature encoding table, enabling the tracking and querying of the medical feature encodings contained in each SR file. To support symptom-based cross-modal queries, creating a composite index on the feature encodings is a crucial step. A composite index typically includes multiple fields, such as the medical feature encoding (radlex_code) and the study instance UID (study_instance_uid). This index structure significantly improves the speed and efficiency of queries based on specific medical features, allowing users to quickly locate all relevant records containing a specific symptom, whether they are images or structured reports. Through the above steps, especially by creating a composite index in the feature encoding table, queries based on medical feature encodings can be significantly accelerated. For example, if a doctor wants to find all cases with "benign calcification", the system can quickly locate all records containing the corresponding feature encoding through the index and return the relevant DICOM images and SR files, providing strong support for clinical decision-making.
[0042] In some embodiments of the present application, the association between the medical feature encoding in the SR file and the study instance unique identifier is achieved by mapping the user input to the medical feature encoding through a BERT-NER model to query the structured dataset, specifically as follows: Use the BERT-NER model to parse the natural language description input by the user and identify the key medical entities; convert the key medical entities into the corresponding medical feature encodings; use the medical feature encodings to construct an SQL query statement to retrieve the relevant unique identifiers; based on the unique identifiers, find the corresponding SR file in the structured report table; use the same unique identifiers to find the relevant DICOM images in the PACS system or DICOM image database; integrate the SR file and the DICOM images to form a structured dataset.
[0043] As described above, when the user inputs a natural language description (e.g., "right breast mass"), the system uses the BERT-NER (Bidirectional Encoder Representations from Transformers - Named Entity Recognition) model to parse this description. The BERT-NER model can identify and extract key medical entities in the text, such as "right breast" and "mass". These key medical entities are the core content of the user's query and represent the specific medical features they hope to search for. The system converts the identified key medical entities into corresponding medical feature codes. For example, "right breast" may be converted to "RID29896", and "mass" may be converted to "RID39055". These codes are usually based on internationally or domestically recognized standard terminology systems (such as SNOMED CT, RadLex, etc.), ensuring interoperability and consistency between different systems. Using the medical feature codes obtained from the above conversion, the system will construct a structured SQL query statement to retrieve lesion records containing all target codes and finally return the corresponding study instance UID (study_instance_uid). Specifically, the system will look for records containing the specified RadLex code in the feature code table (sr_features), and aggregate by lesion_id to determine if there is a lesion that simultaneously has all the specified codes. If the condition is met, the system further obtains the corresponding sr_id through the lesion table (sr_lesions) and associates it with the structured report table (sr_reports) to extract its study_instance_uid. For example, when the query conditions are "RID29896" (right breast) and "RID39055" (mass), the system will retrieve all records that contain both of these codes in the same lesion structure and return their corresponding study instance UIDs. Based on the study instance UIDs obtained from the previous step, the system will look for the corresponding SR files in the structured report table (sr_reports). Since each SR file has a unique study instance UID, this step can accurately locate all SR files related to the user's query. These SR files contain detailed imaging report content, such as lesion location, mass size, margin status, etc. At the same time, using the same study instance UID, the system will look for relevant DICOM images in the PACS system or DICOM image database. Since the study instance UID remains consistent throughout the medical process, it can ensure that DICOM images that exactly match a specific SR file are found. This step enables doctors to not only view the detailed structured report but also view the actual medical images, thus obtaining more comprehensive information to support clinical decision-making.Finally, the system integrates the retrieved SR files and DICOM images to form a complete structured dataset. This dataset not only contains the original medical images but also includes detailed structured report content and all medical feature codes extracted therefrom. This integration method greatly improves the usability and accessibility of the data, allowing doctors to quickly and accurately obtain the required information to support a more efficient diagnosis and treatment process.
[0044] In some embodiments of the present application, it further includes: recording the learning identity identifier of the doctor when obtaining the unstructured imaging report and the corresponding DICOM format imaging data; evaluating the basic capabilities of the doctor through a standardized test and converting the test results into a proficiency score through a machine learning model; recording the operation behavior data of the doctor and using an online learning algorithm to update the proficiency score in real time; presetting structured templates and rules for different proficiency scores.
[0045] As described above, when obtaining the unstructured imaging report and the corresponding DICOM format imaging data, record the doctor's identity information (such as ID or username). This step ensures that all subsequent operations can be traced back to a specific doctor, providing a basis for personalized evaluation and feedback. Integrate a user authentication module in the system to automatically record the executor of each operation. This information can be stored in a log file or database for subsequent analysis. Evaluate the basic capabilities of the doctor through a standardized test, such as imaging interpretation skills, diagnostic accuracy, etc. This helps to understand the doctor's professional level and provides a basis for personalized training and development. Organize online tests or simulated case analyses regularly, and the test content covers multiple aspects such as medical knowledge and imaging interpretation skills. The test results are analyzed through a machine learning model to convert the quantitative scores into more intuitive proficiency scores. For example, a model trained based on historical data can predict the doctor's performance in actual work according to the doctor's performance. Record the actual operation behavior data of the doctor (such as report writing quality, diagnostic speed, patient feedback, etc.) and use an online learning algorithm to update the proficiency score in real time. This method can reflect the doctor's real ability and progress, and timely discover areas that need improvement. The system continuously collects the doctor's operation data, including but not limited to report generation time, accuracy of terms used, patient treatment effect, etc. These data are input into a pre-trained online learning model, and the model will continuously adjust the doctor's proficiency score according to the new data. For example, if a certain doctor's diagnostic accuracy has been significantly improved in recent times, their score will increase accordingly.
[0046] Based on the proficiency scores of different doctors, the system presets specific structured templates and rules for them. For doctors with less experience, more detailed and guiding templates can be provided; while for those with rich experience, more concise and efficient templates can be used. Develop a multi-level structured template system, where each level corresponds to a different proficiency score range. For example, junior doctors may use templates with more prompts and instructions, while senior doctors can use concise templates. In addition to templates, different rule sets can also be set for doctors with different proficiencies. For example, for novice doctors, the system may require them to make a second confirmation before submitting a report; while for senior doctors, they are allowed to submit directly.
[0047] In some embodiments of the present application, for different proficiency scores, structured templates and rules are preset, specifically: If the proficiency is the first-level proficiency, the structured data is split into standardized fields and annotation descriptions are added; if the proficiency is the second-level proficiency, the detail level of the structured data is reduced, and a summary-style structured report is provided, retaining the key conclusions.
[0048] As described above, for doctors at the first-level proficiency (usually novice or less experienced doctors), the system provides a very detailed structured template and adds detailed annotation descriptions next to each field. This can help them better understand the meaning of each item and ensure the accuracy and completeness of the report.
[0049] The structured data is split into multiple standardized fields. For example, information such as "lesion location", "mass size", "margin status" in an imaging report is listed separately instead of being described as a whole. Detailed annotation descriptions are added to each field to help doctors understand the specific meaning of each field and its filling requirements. For example, add an annotation next to the "lesion location" field: "Please select the correct location according to anatomical standard terms, such as the upper outer quadrant of the right breast." Provide specific examples or guiding statements to help doctors fill in each field correctly. For example, "If a mass is found, measure its maximum diameter and record it here." Suppose Dr. Li is a newly recruited novice doctor. When writing an imaging report of mammography, he uses a template designed for the first-level proficiency. This template will clearly list all the fields that need to be filled in and provide detailed annotations and examples next to each field to help Dr. Li complete the report accurately. For doctors at the second-level proficiency (usually those with rich experience and able to complete tasks quickly and accurately), the system provides a simplified structured template that highlights the key conclusions. This kind of template reduces unnecessary details, enabling doctors to complete the report more efficiently while ensuring that key information is not omitted.
[0050] Compared with the first-level template, the second-level template reduces the level of detail of the structured data. For example, instead of filling in the detailed anatomical positions and specific dimensions one by one, a comprehensive description is directly given.
[0051] Although the level of detail is reduced, the most important conclusion part is still retained. For example, in the report, it is only necessary to point out that "there is a round mass in the upper outer quadrant of the right breast, with clear edges and suspected benign calcification", without the need to describe each measurement value in detail.
[0052] The system can automatically fill in some common information based on historical data and common patterns, further reducing the workload of doctors. For example, the system can automatically generate some common diagnostic suggestions based on previous cases.
[0053] Suppose Dr. Zhang is an experienced doctor. When writing an imaging report for mammography, he uses a template designed for the second-level proficiency. This template only requires him to fill in a few key fields, such as the lesion location, the shape of the mass, and the preliminary diagnosis conclusion, while other details are automatically processed or simplified by the system. This enables Dr. Zhang to complete the report faster and focus on complex case analysis.
[0054] Through this hierarchical structured template design, the system can provide personalized support according to the proficiency of different doctors: For novice doctors, the detailed field splitting and annotation instructions help them understand and learn, ensuring the quality and accuracy of the report.
[0055] For senior doctors, the simplified summary-style report improves work efficiency, enabling them to complete routine tasks faster and concentrate on dealing with complex problems.
[0056] This method not only improves the working experience of doctors but also optimizes the efficiency and quality of the entire medical process. It embodies the concept of personalized education and support, helping doctors obtain the greatest support and growth space at their respective career development stages through technical means. In addition, such a design also helps medical institutions better manage and cultivate talents, improving the overall service level.
[0057] In some embodiments of the present application, the processing of the unstructured imaging report text by the private AI model to obtain structured data further includes: dynamically generating prompt words according to the proficiency to control the level of detail of the output of the private AI model; if the proficiency is the first-level proficiency, fine-tuning the private AI model using preset granularity annotation data; if the proficiency is the second-level proficiency, fine-tuning the private AI model using summary-style annotation data.
[0058] As described above, by dynamically generating different prompt words according to the proficiency of doctors, it is ensured that the content output by the private AI model can meet the needs of doctors at different levels. The prompt words can guide the model to process the unstructured imaging report text more precisely and generate appropriate structured data. For novice or less experienced doctors (first-level proficiency), the prompt words will be more detailed and specific, helping the model generate structured data with more details. For example, the prompt words may require the model to provide specific anatomical descriptions when identifying the lesion location and mark the specific dimensions of each feature. Example prompt words: "Please describe in detail the location, size, shape, and edge status of the lesion. For example, a round mass with a diameter of 1.2 cm is found in the upper outer quadrant of the right breast, and the edge is clear." For experienced doctors who can complete tasks quickly and accurately (second-level proficiency), the prompt words will be more concise, emphasizing key conclusions. This helps the model generate concise summary-style structured data. Example prompt words: "Please summarize the main findings and give a preliminary diagnosis conclusion. For example, there is a round mass in the right breast suspected of being a benign calcification." Adjusting the detail level of the output of the private AI model according to the proficiency of doctors enables the generated structured data to not only meet the actual needs of doctors but also improve work efficiency.
[0059] For doctors with first-level proficiency, the model output should be as detailed as possible. For example, when processing a mammography report, the model not only needs to identify "a mass in the right breast" but also provide detailed information such as specific dimensions, shape, and edge status. Output example: "A round mass with a diameter of 1.2 cm is found in the upper outer quadrant of the right breast, with a clear edge and suspected of being a benign calcification." For doctors with second-level proficiency, the model output should be more concise, highlighting key conclusions. For example, only the main findings and preliminary diagnosis conclusions need to be pointed out. Output example: "There is a round mass in the right breast suspected of being a benign calcification." Fine-tune the private AI model with annotation data of different granularities according to the proficiency of different doctors to ensure that the model can better adapt to the needs of specific users.
[0060] For novice doctors, the model is fine-tuned using detailed annotated data. These annotated data usually include very detailed fields such as lesion location, mass size, margin status, etc. This fine-tuning method ensures that the model can generate detailed structured data to help novice doctors understand and learn. Example annotated data: "Lesion location: upper outer quadrant of the right breast; Mass size: 1.2 cm in diameter; Shape: round; Margin status: clear; Preliminary diagnosis: suspected benign calcification." For senior doctors, the model is fine-tuned using more concise summary-style annotated data. These annotated data only contain key conclusions, reducing unnecessary details. This fine-tuning method improves the model's ability to generate summary-style structured data, enabling senior doctors to complete reports faster. Example annotated data: "Main finding: There is a round mass with suspected benign calcification in the right breast." Suppose Dr. Li is a newly recruited novice doctor, while Dr. Zhang is an experienced senior doctor. When they write imaging reports for mammography, the system automatically adjusts the behavior of the private AI model according to their proficiency: Dr. Li (first-level proficiency): The system generates detailed prompt words to guide the model to generate structured data containing detailed information. The model is fine-tuned using preset granularity annotated data to ensure that the output is detailed and easy to understand. Output example: "A round mass with a diameter of 1.2 cm is found in the upper outer quadrant of the right breast, with clear margins and suspected benign calcification." Dr. Zhang (second-level proficiency): The system generates concise prompt words that emphasize key conclusions. The model is fine-tuned using summary-style annotated data to ensure that the output is concise and to the point. Output example: "There is a round mass with suspected benign calcification in the right breast." By dynamically adjusting the prompt words and fine-tuning strategies according to the doctor's proficiency in this way, the system can better support the work needs of doctors at different levels: For novice doctors, detailed prompt words and fine-tuning with detailed annotated data help them understand and learn, ensuring the quality and accuracy of the report. For senior doctors, concise prompt words and fine-tuning with summary-style annotated data improve work efficiency, enabling them to complete routine tasks faster and focus on complex problems.
[0061] In some embodiments of the present application, the step of merging the structured data with the DICOM metadata to obtain the SR file further includes: when generating the SR file, adding a proficiency identification field to mark the degree of structuring of the SR file.
[0062] As described above, by adding a dedicated field to the SR file to mark its degree of structuring, it can clearly reflect which level of doctor proficiency the file is generated based on. This not only helps to understand the quality and detail of the report but also supports personalized data management and applications.
[0063] During the process of generating SR files, the system will automatically add a "proficiency identifier" field based on the proficiency score of the doctor who processes the imaging report. This field can be a simple numerical value (e.g., 1 represents the first level of proficiency, 2 represents the second level of proficiency), or it can be a descriptive label (e.g., "beginner", "advanced"). Example field name: "proficiency_level". Example values: First level of proficiency: "1" or "beginner", Second level of proficiency: "2" or "advanced".
[0064] By marking the level of structuring of SR files, it can help medical institutions manage and utilize these files more effectively. For example, when conducting data analysis, clinical decision support, or research, appropriate files can be selected according to different levels of structuring. If an SR file is generated based on the first level of proficiency, then its level of structuring will be very high, containing information such as detailed anatomical locations, specific dimensions, morphological features, etc. This type of file is suitable for situations where in-depth understanding of case details is required. Example content: "Lesion location: upper outer quadrant of the right breast; Mass size: 1.2 cm in diameter; Morphology: round; Margin status: clear; Preliminary diagnosis: suspected benign calcification." If an SR file is generated based on the second level of proficiency, then its level of structuring is relatively low, only retaining the key conclusions and main findings. This type of file is suitable for quick browsing and preliminary judgment. Example content: "Main finding: There is a round mass with suspected benign calcification in the right breast." Suppose a large number of SR files are stored in the PACS system of a hospital, and these files are generated by doctors with different levels of proficiency. Now, the hospital hopes to classify and further analyze these files: The hospital can classify all files into two categories according to the "proficiency identifier" field in the SR files: one category is detailed reports generated by doctors with less experience (first level of proficiency), and the other category is concise reports generated by senior doctors (second level of proficiency). This classification helps to quickly find the required information in daily work. For example, during routine examinations, concise reports can be viewed first to improve efficiency; while during complex case discussions, detailed reports can be consulted in depth. For scientific research or quality control projects, researchers can screen out specific types of SR files for analysis according to the "proficiency identifier" field. For example, compare the differences in accuracy, completeness, and consistency of reports generated at different proficiency levels. This type of analysis helps to identify potential problem areas and provide targeted training and support for doctors.
[0065] The system can provide personalized feedback and educational suggestions for each doctor according to the "proficiency identifier" field in the SR file. For example, for novice doctors, their skills can be improved through the notes and guidance in detailed reports; while for senior doctors, they can be encouraged to maintain high-quality standards while simplifying reports.
[0066] By adding a proficiency identification field and marking the degree of structuring when generating the SR file, the system can better adapt to the needs of different doctors and provide more flexibility and precision for subsequent data management and applications. This method not only improves the working experience of doctors but also optimizes the efficiency and quality of the entire medical process. It embodies the concept of personalized education and support, helping doctors obtain the greatest support and growth space at their respective career development stages through technical means. In addition, such a design also helps medical institutions better manage and cultivate talents and improve the overall service level.
[0067] In some embodiments of the present application, after associating the medical feature codes in the SR file with the unique identification code and mapping the user input to the medical feature codes through the BERT-NER model to query the structured data set, it further includes: recommending differentiated learning content according to the proficiency score of the doctor. Specifically, if the proficiency is the first-level proficiency, structured reports and basic cases are preferentially pushed, combined with interactive learning tools; if the proficiency is the second-level proficiency, mixed content is provided, and complex cases and diagnostic reasoning training are introduced; cases of similar difficulty are recommended based on the doctor's historical learning behavior.
[0068] As described above, by analyzing the proficiency scores of doctors, the system can recommend learning content that is most suitable for their current level for each doctor, thus promoting their professional growth and development. For doctors at the first-level proficiency, the system will preferentially push structured reports and basic cases containing detailed information. These reports usually include information such as detailed anatomical locations, specific dimensions, morphological features, etc., which helps doctors understand the meaning of each item and ensure the accuracy and completeness of the report. Example learning content: Structured report: "Lesion location: upper outer quadrant of the right breast; mass size: 1.2 cm in diameter; morphology: round; margin status: clear; preliminary diagnosis: suspected benign calcification." Basic case: "A 45-year-old female patient was found to have a round mass about 1.2 cm in diameter in the right breast, with clear margins." To help novice doctors better understand and master knowledge, the system also provides a series of interactive learning tools. These tools can include simulation exercises, real-time feedback mechanisms, and case discussion platforms, etc. Example tools: Doctors can practice image interpretation through a virtual environment, and the system will give immediate feedback and suggestions. When writing a report, the system can provide immediate prompts and corrections based on the doctor's input to help them improve accuracy.
[0069] For doctors at the second level of proficiency, the system provides more diverse learning content, including both basic cases and complex cases. The design of this mixed content aims to maintain doctors' mastery of basic knowledge while challenging them to apply their skills in more complex scenarios. Example learning content: Complex case: "A 60-year-old male patient with a long history of smoking, chest CT shows multiple pulmonary nodules, some with spiculated margins, suspected of malignant tumor." Mixed content example: Combining basic cases with complex cases to help doctors transition between different difficulty levels. In addition to providing complex cases, the system also introduces a diagnostic reasoning training module. These trainings aim to improve doctors' clinical thinking ability and problem-solving skills, enabling them to make more accurate judgments when facing complex medical conditions. Example training: Doctors need to gradually derive possible diagnostic conclusions based on the provided medical records and imaging results and explain the basis. The system will give feedback and guidance according to the doctors' answers. By analyzing doctors' historical learning behaviors, the system can more precisely recommend learning cases that match their current level. This not only improves learning efficiency but also enhances doctors' learning interest and motivation. The system records each doctor's learning activities, including the types of cases they view, the quantity and quality of exercises completed, the discussion topics they participate in, etc. Based on this data, the system can analyze doctors' learning preferences and areas of expertise. Example data points: Doctor A often consults cases related to breast cancer and performs well in relevant exercises. Doctor B pays more attention to the diagnosis of cardiovascular diseases and performs excellently in the diagnostic reasoning training in this field.
[0070] Based on the above analysis results, the system will recommend cases with similar difficulty levels and types to each doctor according to their historical learning behaviors. For example, if a doctor has mainly dealt with basic cases of breast cancer and performed well in the past few months, the system may start to recommend some slightly more complex breast cancer cases to gradually improve their skills. Example recommendation: For Doctor A, the system may recommend some complex cases involving multiple breast nodules or accompanying other complications. For Doctor B, the system may recommend some cardiovascular disease cases involving acute myocardial infarction or other emergencies.
[0071] By means of this method of recommending differentiated learning content based on doctors' proficiency scores, the system can better support the professional growth and development of doctors: for novice doctors, it provides detailed structured reports and basic cases, combined with interactive learning tools, to help them lay a solid foundation and gradually improve their skills. For senior doctors, it offers mixed content and diagnostic reasoning training, which can not only consolidate basic knowledge but also challenge their application abilities in complex scenarios. Recommending cases of similar difficulty based on historical learning behaviors makes the learning content more in line with doctors' actual needs, improving learning efficiency and interest. This method not only enhances doctors' work experience but also optimizes the efficiency and quality of the entire medical process. It embodies the concept of personalized education and support, using technical means to help doctors obtain the greatest support and growth space at their respective career development stages. In addition, such a design also helps medical institutions better manage and cultivate talents, improving the overall service level.
[0072] In some embodiments of the present application, after recommending differentiated learning content according to doctors' proficiency scores, it further includes: Taking behavioral indicators and ability indicators as learning effect quantification indicators, and dynamically adjusting the differentiated learning content based on the learning effect quantification indicators; Adjusting the keyword weights retrieved according to doctors' proficiency, and providing explanations of retrieval results for doctors with secondary proficiency.
[0073] As mentioned above, by analyzing doctors' behavioral indicators (such as the number of exercises completed, the frequency of participating in discussions) and ability indicators (such as diagnostic accuracy, report quality), the system can quantify doctors' learning effects and dynamically adjust the recommended learning content accordingly. The system will record the specific learning activities of each doctor, including the number of cases they view, the number of exercises completed, and the online discussions they participate in. Example behavioral indicators: Doctor A completed 10 basic case exercises and participated in 3 online discussions in the past week. Doctor B completed 5 complex case exercises and submitted 2 detailed case analysis reports in the past month. The system will also evaluate doctors' practical operation abilities, such as diagnostic accuracy, report writing quality, etc. Example ability indicators: Doctor A's diagnostic accuracy rate in basic case exercises is 90%. Doctor B's report quality score in complex case exercises is 85 points (out of 100).
[0074] Based on the above behavior and ability indicators, the system will dynamically adjust the recommended learning content. If a doctor performs well in a certain field, the system may reduce the push of basic content in that field and increase more challenging cases; conversely, if a doctor performs poorly in a certain field, the system will increase basic exercises and guidance in that field. Example adjustment: If doctor A performs very well in basic breast cancer cases, the system may start to recommend some slightly more complex breast cancer cases or introduce diagnostic reasoning training. If doctor B performs averagely in complex cardiovascular disease cases, the system may increase basic knowledge review and basic exercises in that field. Adjust the weight of search keywords according to the proficiency of doctors so that search results are more in line with their actual needs and professional level. This helps to improve the efficiency and accuracy of information acquisition. For doctors with the first level of proficiency, the system will give priority to displaying materials and cases related to basic content during the search process. For example, when a doctor enters "breast lump", the system will give priority to displaying basic case reports and structured templates containing detailed anatomical location, specific size, morphological characteristics and other information. Sample search result: "A 45-year-old female patient was found to have a round mass with a diameter of about 1.2 cm and clear margins in the right breast." Structured template: "Lesion location: right upper outer quadrant of breast; Mass size: 1.2 cm in diameter; Shape: round; Margin status: clear." For doctors with the second level of proficiency: the system will prioritize the display of materials related to complex cases and advanced diagnoses during the search process. In addition, the system will provide explanations of the search results to help doctors better understand and apply this information. Example search results and explanations: Search results: "A 60-year-old male patient with a long history of smoking, chest CT showed multiple lung nodules, some with spiculated margins, suspected of malignant tumors." Explanation: "This case involves multiple lung nodules with spiculated margins, which usually indicate a high possibility of malignancy. It is recommended to make a comprehensive judgment based on the patient's medical history and other imaging examination results." For doctors with the second level of proficiency, the system not only provides search results, but also comes with concise explanations to help doctors quickly understand and apply this information. The system can automatically generate short explanatory text based on the search results, highlighting key conclusions and precautions. This explanation can be based on a predefined knowledge base or automatically generated through natural language processing technology. Example explanation: When a doctor searches for "pulmonary nodules", in addition to returning relevant cases, the system will also provide an explanation: "In this case, the patient has multiple pulmonary nodules, some with spiculated edges, suggesting a high possibility of malignancy. Further PET-CT scans or tissue biopsies are recommended to clarify the diagnosis." By using behavioral indicators and ability indicators as quantitative indicators of learning effectiveness, and dynamically adjusting differentiated learning content based on this, the system can more accurately meet the learning needs of doctors and promote their professional growth and development. At the same time, the retrieval keyword weights are adjusted according to the proficiency of doctors, and retrieval result explanations are provided for doctors at the second level of proficiency, making information acquisition more efficient and targeted.
[0075] Embodiment 2 Now, in combination with a specific embodiment - the construction of a mammography image dataset, the technical details of the method and device will be elaborated in detail to fully demonstrate its innovative value and practical feasibility.
[0076] Step 1: Image and report acquisition. Obtain mammography image reports and corresponding DICOM - format image data from the RIS system, where the image reports are in the form of unstructured free text. The mammography DICOM - format image data in the RIS system. The unstructured image reports of mammography in the RIS system.
[0077] Step 2: Report structuring. Perform structuring of the image reports based on a private AI model deployed on a local server; the local private AI model is obtained through supervised fine - tuning with labeled data, and a specific task instruction for image report structuring is set for this private AI model. Step 2 is divided into the following small steps: S1: Based on the authoritative medical domain standard terminology system (consensus on mammography examination and diagnosis), combined with the actual needs of clinical diagnosis and treatment, construct a template for extracting key information from mammography reports. This template precisely defines the medical entities to be extracted and their logical relationships, forming a normative framework for extracting key information from unstructured reports ( Figure 2 、 Figure 3 ) S2: According to the standardized template in S1, carry out prompt engineering, design professional prompt words adapted to the large - language model, and clarify the task instructions for the model to process unstructured reports. At the same time, based on the template, some unstructured reports are manually annotated to generate an initial annotated dataset, providing a basis for supervised learning in model training. Since different doctors may adopt different description habits and terms when writing image reports, a comparison table of mammography report templates is created according to the situation of our hospital, and the prompt words of the private AI model are standardized according to the comparison table; S3: Use the prompt words and initial annotated dataset in S2 to perform the first - round SFT on the large - language model. The optimized model is used to process a large number of unstructured medical image reports, generate a large - scale structured report, form a rich annotated dataset, and expand the scale of supervised learning data; S4: Adopt an iterative training strategy. Incorporate the specialized prompting words designed in S2 to adapt to large language models, clarify the task instructions for the model to process unstructured reports, and input the large-scale labeled data generated in S3 into the model. Repeat multiple rounds of supervised fine-tuning. Continuously calibrate the consistency between the model output and medical professional standards, gradually optimize the model performance, and finally generate the target private AI model that meets the structured requirements of medical imaging reports; S5: Deploy the target model after multiple rounds of fine-tuning to a local server and build a local processing system. Use the final model deployed on the local server to perform structured processing on the unstructured mammography imaging reports obtained in Step 1 to obtain structured mammography imaging reports.
[0078] Step 3: Data Association and SR Generation: Use a specific data integration script to associate and integrate the structured mammography medical imaging reports output by the private AI model with the metadata in the DICOM-format mammography images (such as patient information, examination parameters, etc., which can be used as data for association) to generate a structured report file that complies with the DICOM SR standard specification. This file uses the DICOM standard data storage format (i.e., the file extension is.dcm). Step 3 is divided into the following small steps: S6: Generate a series of key basic metadata for the SR file. Among them, the unique identifier is the key identifier to ensure that each SR file has a unique identity in the entire dataset. It follows specific coding rules so that each report can be accurately located and distinguished in subsequent data management and retrieval; S7: Carefully read the content and its nested structure of the structured mammography imaging report generated in Step S5. This requires in-depth analysis of every detail in the report, including various medical findings, relevant measurement data, and their logical relationships, etc. Then, accurately add this content to the corresponding positions in the SR file to ensure the integrity and accuracy of the information. During this process, to achieve the standardization and interoperability of medical information, each medical term in the report content is uniformly represented using RadLex term coding. By using these standard term codings, different medical systems and software can more accurately understand and exchange medical information, avoiding misunderstandings and errors caused by inconsistent terms; S8: Extract key metadata information from the DICOM - formatted mammography image data obtained in Step 1. These metadata include, but are not limited to, patient information (such as name, age, gender, medical record number, etc.), examination parameters (such as scanning method, scanning time, scanning parameter settings, etc.), and other data that can be used as a basis for association. These metadata are crucial for understanding the source, background, and patient - relatedness of the images. Then, accurately reference the extracted metadata information into the SR file. By establishing this data association, the SR file not only contains detailed structured report content but also can be closely linked to the corresponding DICOM image data, forming a complete medical information unit. This association relationship provides a more comprehensive and accurate data foundation for subsequent medical image analysis, diagnostic decision - making, and data mining, etc.; S9: Finally, based on the work completed in the previous steps, generate a mammography structured report file that fully complies with the DICOM SR standard specification. This process strictly follows the data storage format requirements for SR files in the DICOM standard, ensuring that every byte and every field of the file conforms to the standard definition and specification. The final generated file has an extension of.dcm, which is the file format specified by the DICOM standard and can be recognized and processed by a wide range of medical devices, image - processing software, and information systems. By generating such a standard - compliant structured report file, the standardized storage and exchange of medical image data and related reports are achieved, providing solid data support for information sharing, quality control, and clinical research in the medical field.
[0079] Step 4: Dataset construction and association. Transmit the mammography structured report file that complies with the DICOM SR standard specification and is stored in the.dcm format back to the PACS system. By establishing data association relationships (such as based on unique identification information such as patient ID, examination time, etc.) in the PACS system, form a medical image dataset in which medical images (DICOM format) and medical image structured reports (compliant with the DICOM SR standard and stored in the.dcm format) are mutually associated. Step 4 is divided into the following sub - steps: S10: Design of the associated database table structure. In the original medical imaging system, a structured report table sr_reports, a lesion table sr_lesions, and a feature coding table sr_features are added. The three are hierarchically associated through primary and foreign keys to construct a "report → lesion → feature" relationship. The sr_reports table stores the metadata of the SR files generated in step three (such as SOPInstance UID, Study Instance UID), the complete text content, and the creation timestamp to ensure the traceability of the reports. Each report is uniquely identified by the primary key sr_id. The sr_lesions table is used to record the structural information of multiple lesions extracted from the reports. Each record is associated with the sr_reports table through the foreign key sr_id to implement a one-to-many mapping relationship between the report and the lesions. This table can include fields such as lesion number, lesion location, and automatic summary. The sr_features table is used to store the standardized medical feature information extracted for each lesion. Each record in the table is associated with the sr_lesions table through the foreign key lesion_id, thus ensuring that all feature information has a clear structural attribution relationship and can be traced back to the original SR file.
[0080] S11: Creation of a composite index. A composite index idx_sr_features is created in the sr_features table, using the B+ tree index structure to accelerate feature retrieval. The index fields are radlex_code and lesion_id; S12: Construction of a three-level association system of "imaging → report → feature" to support cross-modal joint queries. The imaging data is associated through the study_instance_uid of the sr_reports table, and the feature dimension is associated through the radlex_code of the sr_features table.
[0081] Step 5: Image retrieval based on symptom features. Use the NER model based on the BERT architecture to achieve accurate conversion from natural language keywords to feature codes, and use the converted codes to perform an association query to obtain the target image. Step 5 is divided into the following sub-steps: S13: "Natural language - coding" mapping engine: The named entity recognition model based on BERT achieves accurate mapping from natural language keywords to RadLex codes through three stages: First, use large-scale texts such as medical literature and clinical reports to pre-train the BERT model parameters , learn the general medical semantic representation; then fine-tune the model on the labeled dataset, and optimize the parameters by minimizing the cross-entropy loss function to make the conditional probability distribution output by the model approximate the true coding distribution; finally, for the input keyword sequence , the model calculates the RadLex code with the highest output probability to achieve end-to-end conversion from unstructured text to standardized feature encoding, providing structured input for image retrieval; S14: The sr_features table details the RadLex code information of the symptom features corresponding to each examination instance. After converting natural language to RadLex codes through the BERT-based named entity recognition model, using these codes to search in the sr_features table can accurately locate the corresponding examination instance. The images table records the identifiers of the examination instances to which each image file belongs. After determining the examination instance in the first-step association, using the identifier of this examination instance to search in the images table can find the corresponding image file. Through such two-level association, it is possible to efficiently and accurately find the corresponding medical image file starting from the symptom features described in natural language, providing strong data support for clinical diagnosis and research.
Claims
1. A method for establishing a medical image dataset with a unified structure, characterized in that: include: Based on the unstructured image report and the corresponding DICOM format image data, the unstructured image report text, DICOM image and DICOM metadata are extracted; Processing the unstructured imaging report text through a private AI model to obtain structured data, merging the structured data with the DICOM metadata to obtain an SR file, wherein the SR file includes the DICOM metadata, a medical feature code, and a unique identifier; Based on the unique identifier, associating the SR file with the DICOM image to form a structured data set; The medical feature code in the SR file is associated with the unique identification code, and the user input is mapped to the medical feature code through a BERT-NER model to query the structured data set.
2. The method according to claim 1, characterized in that The unstructured image report text is processed by a private AI model to obtain structured data, and the structured data is merged with the DICOM metadata to obtain an SR file, wherein the SR file includes the DICOM metadata, the medical feature code and the unique identifier, specifically: Define structured templates based on standard medical terminology; Performing supervised fine-tuning on the private AI model through a manually annotated initial data set, and iteratively optimizing the output of the private AI model so that the output matches the medical standard terminology; Parsing the unstructured image report text into the structured data conforming to the structured template through the private AI model; Determining, according to the DICOM SR standard, a data template, that the structured data matches the unique identifier of the DICOM metadata; Initialize the SR file based on the DICOM SR standard; Filling basic information in the DICOM metadata into the basic metadata field of the SR file, the basic information including at least one of patient information and examination parameters; Traversing the structured data to extract medical feature codes and related information; Formatting the medical feature code according to the requirements of the DICOM SR standard; Embedding the formatted medical feature code and related information into the content sequence of the SR file; In addition to the basic metadata fields, the DICOM metadata is also added to the relevant fields of the SR file; The fields are checked and a DICOM validation tool or library is used to check whether the generated SR file complies with the DICOM SR standard.
3. The method according to claim 1, characterized in that The step of associating the SR file with the DICOM image based on the unique identifier to form a structured data set is as follows: Confirming that the unique identifier of the SR file is consistent with the unique identifier of the DICOM image; Constructing a structured report table to store relevant information of the SR file; Constructing a feature coding table to store the medical feature codes extracted from the SR file; Uploading the SR file to the PACS system, using the unique identifier to find and associate the corresponding DICOM image, or, using the unique identifier as a foreign key, establishing an association between the structured report form and the DICOM image metadata; Associating the DICOM image with the SR file through the unique identifier; Associating the SR file with the medical feature code in the feature code table through the unique identifier; The medical feature coding is used to support cross-modal queries based on disease symptoms, and a composite index is established on the feature coding to accelerate queries based on the medical feature coding.
4. The method according to claim 1, characterized in that: The step of associating the medical feature code in the SR file with the unique identification code of the research instance and mapping the user input to the medical feature code through the BERT-NER model to query the structured data set is specifically as follows: Use the BERT-NER model to parse the natural language description entered by the user and identify key medical entities; Converting the key medical entities into corresponding medical feature codes; Using the medical feature code, construct an SQL query statement to retrieve the related unique identifier; Based on the unique identifier, searching the structured report table for the corresponding SR file; Using the same unique identifier to search for the related DICOM image in a PACS system or a DICOM image database; The SR file and the DICOM image are integrated to form a structured data set.
5. The method according to claim 1, characterized in that Also includes: When obtaining the unstructured image report and the corresponding DICOM format image data, recording the learning identity of the doctor; Assess doctors’ basic competencies through standardized tests and convert test results into proficiency scores through machine learning models; Recording the doctor's operation behavior data, and using an online learning algorithm to update the proficiency score in real time; Structured templates and rules are preset for different proficiency scores.
6. The method according to claim 5, characterized in that The structured templates and rules are preset for different proficiency scores, specifically: If the proficiency is the first level proficiency, splitting the structured data into standardized fields and adding annotations; If the proficiency level is the second level proficiency, the detail level of the structured data is reduced, a summary structured report is provided, and key conclusions are retained.
7. The method according to claim 1, characterized in that The method of processing the unstructured image report text by a private AI model to obtain structured data further includes: Dynamically generate prompt words based on proficiency to control the level of detail output by the private AI model; If the proficiency is the first level of proficiency, fine-tuning the private AI model using the annotated data of a preset granularity; If the proficiency is the second level proficiency, the private AI model is fine-tuned using summary-type annotated data.
8. The method according to claim 1, characterized in that The step of merging the structured data with the DICOM metadata to obtain the SR file further includes: When generating the SR file, a proficiency identification field is added to mark the structured degree of the SR file.
9. The method according to claim 1, characterized in that: After associating the medical feature code in the SR file with the unique identification code and mapping the user input to the medical feature code through the BERT-NER model to query the structured data set, the method further includes: recommending differentiated learning content according to the doctor's proficiency score, specifically: If the proficiency level is the first level, structured reports and basic cases will be pushed first, combined with interactive learning tools; If the stated proficiency level is Level 2, then mixed content is provided and complex cases and diagnostic reasoning training are introduced; Recommend cases of similar difficulty based on the physician’s historical learning behavior.
10. The method according to claim 9, characterized in that After recommending differentiated learning content based on the doctor's proficiency score, it also includes: Taking the behavior index and the ability index as quantitative indicators of learning effect, and dynamically adjusting the differentiated learning content based on the quantitative indicators of learning effect; Adjust the search keyword weights according to the doctor's proficiency, and provide search result explanations to doctors with the second level of proficiency.
Citation Information
Patent Citations
Multi-mode medical image and report data management method and system
CN109961828A
Medical diagnosis and treatment system
CN111292821A
DICOM-based SR structured report generation method and system, and device
CN111475552A
Systems and methods for improved analysis and generation of medical imaging reports
CN112868020A
Magnetic resonance diagnosis auxiliary teaching system and method based on large language model
CN117037994A