A digital fusion management method and system based on a medical large model

By employing a digital fusion management approach based on a large medical model, and utilizing multi-set prior information and multimodal feature fusion technology, the problem of insufficient information richness in medical image data processing has been solved. This has enabled efficient transmission and intuitive presentation of image information, thereby improving diagnostic and treatment efficiency and patient communication.

CN121171554BActive Publication Date: 2026-05-05INNER MONGOLIA HUAXUN SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INNER MONGOLIA HUAXUN SOFTWARE CO LTD
Filing Date
2025-09-19
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing medical imaging data processing workflows suffer from insufficient content richness due to information dimension compression, information gaps between radiologists and clinicians, and communication gaps between the healthcare system and patients. Current technologies are unable to effectively solve the problems of insufficient image information richness and information transmission.

Method used

We adopt a digital fusion management approach based on a large medical model. By acquiring multiple sets of prior information, including clinically relevant texts, preliminary diagnostic information, and information generated from a historical case knowledge base, we use a bidirectional multi-head cross-attention module to perform multimodal feature fusion. Combined with a visual question-answering model and hierarchical retrieval, we generate structured clinical context information and provide intuitive analytical conclusions and differentiated analysis reports.

Benefits of technology

It significantly improves the richness, accuracy, and transmission efficiency of medical imaging information, ensures the integrity and usability of information, solves the problem of transmission gaps between different roles of imaging information, and improves diagnostic and treatment efficiency and patients' ability to understand examination results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121171554B_ABST
    Figure CN121171554B_ABST
Patent Text Reader

Abstract

This invention provides a digital fusion management method and system based on a large medical model, relating to the intersection of medical information technology and artificial intelligence. The digital fusion management method specifically includes: acquiring and generating multiple sets of prior information, including: prior information generated based on clinically relevant text, second and third prior information generated based on preliminary diagnostic information, and fourth prior information generated based on historical case knowledge base retrieval; processing examination videos based on the first set of prior information to locate one or more candidate video segments, and verifying the candidate video segments based on the second set of prior information to obtain a final set of retrieved segments; and generating a first conclusion for answering clinical questions and a second conclusion for reviewing preliminary diagnostic information based on the final set of retrieved segments and the multiple sets of prior information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the intersection of medical information technology and artificial intelligence, specifically to a system and method for processing, analyzing, and visualizing medical image data using deep learning models, and more particularly to a digital fusion management method and system based on a large-scale medical model. This invention utilizes multimodal information fusion and natural language processing technologies to improve the richness, accuracy, and transmission efficiency of medical image information. Background Technology

[0002] With the development of modern medical imaging technologies such as computed tomography (CT) and magnetic resonance imaging (MRI), massive amounts of high-dimensional medical imaging data have been generated in clinical practice. These data, with their high resolution and three-dimensional volumetric form, provide unprecedented windows for observing the internal structures of the human body. Currently, the standardized processing procedure for this imaging data typically involves radiologists using specialized imaging workstations (such as PACS systems) to observe, analyze, and interpret the acquired two-dimensional slice sequences or three-dimensional data volumes, and then writing an imaging examination report consisting primarily of descriptive text, supplemented by several two-dimensional static keyframe images. This report, as the core deliverable, is transmitted to the clinical physician who ordered the examination and is ultimately included in the patient's medical record.

[0003] However, this long-established standardized process suffers from inherent and deep-seated technical bottlenecks in information processing and transmission. First, the severe compression of information dimensions leads to insufficient content richness. Original CT or MRI data is a three-dimensional data volume containing hundreds or even thousands of two-dimensional images, encompassing the complete spatial morphology, topological relationships, and detailed internal texture features of lesions or anatomical structures. But in the final report, this rich three-dimensional information is reduced to descriptive language summarized by radiologists through subjective observation, and a few isolated, non-interactive two-dimensional screenshots. This conversion from high-dimensional data to low-dimensional text and static images inevitably results in the loss of a large amount of potentially valuable objective information. Clinicians cannot intuitively, dynamically, and comprehensively examine the three-dimensional overall appearance of the lesion and its adjacent relationships with surrounding blood vessels, tissues, and other key structures through the report.

[0004] Meanwhile, significant information gaps exist between different roles. The first gap exists between radiologists and clinicians. When clinicians order examinations, their core objective is to verify or rule out a specific clinical suspicion, which constitutes the clinical intent of the examination. Radiologists, on the other hand, based on their professional knowledge, comprehensively scan the imaging data and report their findings. Due to limitations in professional perspective and information carrier (i.e., the aforementioned reduced-dimensional report), the imaging findings cannot always perfectly and comprehensively respond to the clinical intent; some details crucial to clinical decision-making may be simplified or not fully presented in the report. This forces clinicians to often return to the PACS system to interpret the images themselves if they want to obtain more comprehensive information, which not only places high demands on clinicians' image interpretation skills but also significantly reduces diagnostic efficiency.

[0005] The second information gap exists between the healthcare system and patients. Traditional imaging reports are filled with technical jargon and abbreviations, rendering them incomprehensible to patients without a medical background. Patients, as the primary responsible parties for their own health, are unable to glean any effective or understandable information from these crucial test results, significantly exacerbating their anxiety and increasing the cost and difficulty of doctor-patient communication. Existing AI-assisted technologies mostly focus on single aspects such as automatic lesion detection, segmentation, or benign / malignant classification, aiming to assist doctors in making judgments but failing to systematically address the lack of richness and communication gaps in post-processing and information transmission of imaging information. Therefore, there is an urgent need in this field for a new technological solution that can transcend traditional reporting models, deeply mine and utilize the rich information in raw imaging data, and organize and present it in a more intelligent, intuitive, and adaptable manner to different user needs, thereby comprehensively enhancing the application value of medical imaging information. Summary of the Invention

[0006] This invention provides a digital fusion management method based on a large medical model, which specifically includes the following steps:

[0007] Multiple sets of prior information are acquired and generated, including: prior information generated based on clinically relevant text, second and third prior information generated based on preliminary diagnostic information, and fourth prior information generated based on historical case knowledge base retrieval.

[0008] Guided by the first part of prior information, the inspected video is processed to locate one or more candidate video segments, and the candidate video segments are verified based on the second part of prior information to obtain the final set of retrieved segments.

[0009] Based on the final set of retrieved fragments and the multiple sets of prior information, a first conclusion is generated to answer clinical questions, and a second conclusion is generated to review the preliminary diagnostic information.

[0010] In the step of generating the fourth prior information, the step of constructing the joint query vector includes: generating a text feature sequence based on the first prior information. A visual feature sequence is generated based on the third prior information and the corresponding image data. The text feature sequence and visual feature sequence are deeply fused using a bidirectional multi-head cross-attention module, which includes:

[0011] a) Perform text-to-visual attention computation to generate a text feature sequence modulated by visual information. ;

[0012] b) Perform visual-to-text attention computation to generate a sequence of visual features modulated by textual information. ;

[0013] The and After residual connection and layer normalization with their respective original inputs, the joint query vector is generated through feature pooling and projection.

[0014] The fourth prior information generated based on the historical case knowledge base includes similarity retrieval:

[0015] A two-stage hybrid retrieval and rearrangement method is employed, the method comprising:

[0016] In the first stage, based on the joint query vector, an approximate nearest neighbor search algorithm is used to recall a query vector containing... A set of candidate cases;

[0017] In the second stage, a hierarchical and metadata-aware re-ranking algorithm is used to calculate a comprehensive score for each case in the candidate case set and then re-rank them. The formula for calculating the comprehensive score is as follows:

[0018] in, Candidate Case Overall score These are preset weighting coefficients. , and These represent vector similarity, clinical metadata similarity, and diagnostic information hierarchical similarity, respectively.

[0019] The similarity of the diagnostic information hierarchy The calculation method is as follows:

[0020] The preliminary diagnosis of the query case and the gold standard for the diagnosis of the candidate case are mapped to the corresponding nodes in the medical ontology graph, respectively.

[0021] The similarity between the two nodes is quantified by calculating the structural relationship between them in the graph:

[0022] in, and These are the nodes representing the query diagnosis and candidate diagnoses in the graph. Calculate the depth of the node. Find the lowest common ancestor of two nodes.

[0023] The step of generating the first conclusion is implemented through a reference-guided visual question-answering model:

[0024] Problem feature vectors and reference eigenvectors Perform element-wise multiplication to generate a guiding query vector. ;

[0025] Using the guiding query vector For video feature vectors Attention-weighted features are obtained by performing attention weighting. :

[0026]

[0027]

[0028] in It is a learnable alignment weight matrix. It is attention weight;

[0029] The problem feature vector, attention-weighted video feature vector, and reference feature vector are fused to generate the final fused feature for answer decoding.

[0030] The steps for generating the second conclusion include:

[0031] Based on the second and fourth prior information, a unified feature list containing all imaging features to be verified is constructed. ;

[0032] For each feature in the list, a binary classifier is used to calculate the probability that it exists in the final set of retrieved fragments. ;

[0033] Based on the source of each feature and its corresponding probability If the value exceeds a preset threshold, the feature is categorized into one of the following: consistent items, potential differences, or supplementary findings, in order to generate the differential analysis report.

[0034] The location of one or more candidate video segments is achieved through a localization model guided by prior information, which performs the following calculation steps:

[0035] a) The guiding feature vector representing the search intent Visual temporal feature sequence of the inspection video The fusion process is performed to generate a guided, enhanced sequence of visual features. The calculation formula is as follows:

[0036]

[0037] in, It is the length of the visual temporal feature sequence. The operation copies the guiding feature vector. Second-rate, The operation involves concatenation along the feature dimension;

[0038] b) The guided enhanced visual feature sequence The input is fed into a temporal context model to generate a localization feature sequence modeled with temporal relationships. ;

[0039] c) Based on the location feature sequence For each time point Decode the confidence level of the candidate video segments:

[0040] in, It is the confidence score predicted from the location feature sequence at time point t. It is the Sigmoid activation function.

[0041] This specification also proposes a digital fusion management system based on a large medical model, which includes:

[0042] Prior information generation module: acquires and generates multiple sets of prior information, including: prior information generated based on clinically relevant text, second and third prior information generated based on preliminary diagnosis information, and fourth prior information generated based on historical case knowledge base retrieval;

[0043] Video retrieval module: Based on the first part of prior information, the inspected video is processed to locate one or more candidate video segments, and the candidate video segments are verified based on the second part of prior information to obtain the final set of retrieved segments;

[0044] Conclusion generation and verification module: Based on the final set of retrieved fragments and the multiple sets of prior information, it generates a first conclusion to answer clinical questions and a second conclusion to verify preliminary diagnostic information.

[0045] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned digital fusion management method based on a large medical model.

[0046] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned digital fusion management method based on a large medical model.

[0047] Compared with existing technologies, the digital fusion management method based on a large medical model disclosed in this invention has achieved significant beneficial effects in improving the richness, accuracy, intuitiveness, and transmission efficiency of medical imaging examination information.

[0048] This invention fundamentally expands the information foundation of image analysis, greatly enhancing the richness and completeness of information. Existing technologies typically analyze a single imaging examination as an isolated event, with information sources limited to the current image data and a brief text report. This invention, however, through a systematic process, structurally integrates information from four dimensions and heterogeneous sources: first prior information representing clinical intent, second and third prior information representing initial expert observations, and fourth prior information representing historical experience and collective wisdom. By unifying the digital modeling of the patient's clinical background, the physician's initial judgment, and reference knowledge from a vast amount of historical cases, this invention constructs an extremely rich, multi-dimensional contextual environment for each image analysis. This transformation from information silos to a fused information field provides a solid foundation for subsequent in-depth analysis and precise information presentation.

[0049] This invention significantly improves the accuracy and reliability of dynamic image evidence retrieval through a multi-dimensional prior guidance and cross-validation mechanism. Existing technologies, when associating image data with reports, often rely on manual location or simple content matching, which is prone to bias and lacks effective verification methods. This invention proposes a closed-loop retrieval strategy of "guidance-location-verification." It utilizes a portion of prior information (such as clinical intent and spatial annotation) as a strong guiding signal to accurately locate candidate segments in complex time-series video data; more importantly, it uses another portion of independent prior information (such as detailed textual descriptions from radiologists and baseline features of reference cases) to rigorously verify the semantic content of candidate segments. This non-cyclic cross-validation logic ensures that the ultimately adopted video segments are not only locationally relevant, but their inherent visual content is also highly consistent with expert descriptions and historical experience. This mechanism greatly improves the data accuracy of the extracted dynamic visual evidence, providing highly reliable input for subsequent analysis.

[0050] This invention transforms the complex process of image analysis into intuitive and interpretable core conclusions, significantly improving the usability and insight of information. The system is not merely a data presenter, but also an information extractor and translator. Through a reference-guided visual question-and-answer model, it can generate direct and focused answers (Conclusion 1) to the specific questions of greatest concern to clinicians, combining video evidence and reference cases, transforming obscure visual features into clear and valuable criteria for judgment. Simultaneously, through differential analysis reports (Conclusion 2), the system can objectively compare the machine's independent analysis results with the preliminary conclusions of human experts, highlighting consistent, differing, and supplementary findings in a structured manner. This design transforms raw data, which requires significant time for manual interpretation, into highly condensed, intuitive, and easily understandable intelligent information that directly supports clinical thinking, fully releasing the intrinsic value of the information.

[0051] This invention effectively addresses the information gap prevalent in existing medical processes by deeply processing and role-based presentation of information. For clinicians, the system provides a highly integrated analytical summary that incorporates multidimensional prior knowledge and dynamic evidence, ensuring smooth and accurate information transmission between examination objectives, findings, and in-depth analytical conclusions. For patients, the system automatically simplifies and visualizes professional conclusions, transforming previously obscure professional reports into safe and personalized interpretive materials. Attached Figure Description

[0052] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 This is a flowchart of the present invention for digital fusion management based on a large medical model. Detailed Implementation

[0054] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0055] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. This application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0056] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this application, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number and aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.

[0057] It should be specifically noted that the methods and systems disclosed in this invention are essentially computer information processing technologies. Their ultimate goal is to improve the richness, accuracy, and transmission efficiency of medical image-related information data. They do not involve, nor are they intended to, perform any form of disease diagnosis or provide treatment recommendations. Any information output by this invention, including the generated first and second conclusions, is provided as an objective information verification and decision support reference to qualified medical professionals, who will then combine it with all clinical information to make a final professional judgment. This invention itself does not directly generate diagnostic or treatment actions.

[0058] Furthermore, the processes described in the embodiments of this invention involving the acquisition, processing, or retrieval of any historical medical data are all implemented in strict compliance with relevant laws and regulations (such as the Personal Information Protection Law), after obtaining explicit authorization and consent from the patient or relevant rights holder, and under conditions that ensure data anonymization and security. Therefore, the methods of this invention are designed and implemented with full consideration for the protection of personal privacy and do not involve any infringement on the personal privacy of patients.

[0059] Additionally, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that practice can be carried out without these specific details.

[0060] This invention provides a digital fusion management method based on a large medical model, which specifically includes the following steps:

[0061] Multiple sets of prior information are acquired and generated, including: prior information generated based on clinically relevant text, second and third prior information generated based on preliminary diagnostic information, and fourth prior information generated based on historical case knowledge base retrieval.

[0062] Guided by the first part of prior information, the inspected video is processed to locate one or more candidate video segments, and the candidate video segments are verified based on the second part of prior information to obtain the final set of retrieved segments.

[0063] Based on the final set of retrieved fragments and the multiple sets of prior information, a first conclusion is generated to answer clinical questions, and a second conclusion is generated to review the preliminary diagnostic information.

[0064] This invention discloses a method for collecting and integrating multi-source heterogeneous prior information to generate structured clinical context information. This method is the initial stage of the overall system process described in this invention. Its purpose is to transform unstructured clinically relevant text from different sources into a machine-understandable, structured set of clinical entities, providing accurate and reliable contextual evidence for subsequent similar case retrieval and image analysis.

[0065] In this embodiment, the system first needs to obtain three types of raw text data related to a single imaging examination from a medical information system (such as HIS, EMR). These three data sources are: patient self-report text, denoted as... The patient's self-report text contains a conversational description of symptoms, personal feelings, and medical history information; the physician's clinical record, written by the attending clinician during the diagnostic process, is recorded as... The doctor's clinical record includes more specialized medical terminology, preliminary diagnosis, physical examination findings, etc. A document specifically issued by the clinician outlining the purpose of this imaging examination is recorded as follows: The purpose of the inspection document directly states the objectives that this inspection aims to observe or eliminate.

[0066] because , , These three types of texts differ significantly in writing style, terminology, and information density, and are all unstructured or semi-structured. This embodiment uses a deep learning language model pre-trained in a specific domain to fuse and extract information from these texts.

[0067] The specific processing procedure is as follows:

[0068] S1: Text Preprocessing and Concatenation

[0069] For the three original texts obtained , , Standardized preprocessing is performed, including removing irrelevant formatting marks and standardizing punctuation. Then, the three cleaned texts are concatenated in a pre-defined order (patient's self-report, doctor's record, examination purpose) to form a complete clinical document to be processed. .

[0070] S2: Named Entity Recognition Based on Clinical BERT Model

[0071] In order to process clinical documents To accurately extract core clinical information, this embodiment employs a named entity recognition task based on a Clinical-BERT model. This model was pre-trained on medical literature, clinical guidelines, and anonymized electronic medical record corpora, enabling it to understand medical-specific grammar, terminology, and contextual relationships.

[0072] S2.1 The clinical documents to be processed According to the input requirements of the BERT model, word segmentation is performed to obtain a word segmentation result. A sequence of lexical units .

[0073] S2.2 will sequence The input is fed into the clinical BERT model. This model is denoted as a nonlinear transformation function. For each word in the input sequence In clinical practice, the BERT model combines its contextual information to generate a high-dimensional, semantically rich contextual embedding vector. The output of the entire sequence is an embedding vector matrix. ,in , is the dimension of the embedded vector.

[0074]

[0075] in, This represents the main structure of the clinical BERT model, which includes a multi-layer Transformer encoder. This represents the context embedding vector matrix output by the model.

[0076] After obtaining the context embedding vector for each word in S2.3, entity labeling is performed on each word through a fully connected neural network layer attached to the top of the BERT model. In this embodiment, the predefined entity label categories follow the IOBES annotation system and mainly include Symptom (SYM), Disease (DIS), Body Part (BOD), Examination (EXA), Drug (DRU), and Time Information (TIM).

[0077] For the first in the sequence Each word element has its corresponding embedding vector. After calculation by the fully connected layer, a value about all The score vector of each possible entity label .

[0078]

[0079] in, It is the first Each term corresponds to the original score of all entity tags. It is the weight matrix of the fully connected layer. It is the bias vector of the fully connected layer. It is the total number of categories of predefined entity tags.

[0080] S2.4 To obtain the probability distribution of each label, the system processes the score vector. Normalization is performed using the Softmax function. The predicted labels for each word are probability The calculation is as follows:

[0081]

[0082] in, It is the first Predicted entity labels for each word element. It is a fractional vector The first in Each component corresponds to a label. The score.

[0083] Ultimately, for each lexical The label with the highest probability is selected as its predicted label.

[0084] S3: Structured Information Generation

[0085] After predicting the labels for all lexical units, the system decodes and merges consecutive lexical units belonging to the same entity according to the rules of the IOBES annotation system, forming complete clinical named entities. For example, for the consecutive lexical units "chest tightness" and "shortness of breath" identified as "B-SYM" and "I-SYM", the system will merge them into a single entity of type "symptoms" called "chest tightness and shortness of breath".

[0086] Ultimately, the result of this stage is the generation of a structured set of clinical entities, which constitutes the first prior information, denoted as . . It is a collection of multiple tuples, each tuple representing an identified clinical entity, with the structure (entity text, entity type, start position, end position).

[0087] For example, given an input clinical document, the first prior information generated... It might look like this:

[0088]

[0089]

[0090]

[0091] .

[0092] Next, this embodiment will transform the semi-structured preliminary diagnostic information generated by radiologists in the traditional workflow into a fully machine-readable and analyzable structured multimodal data set.

[0093] S2.1 Formal Collection and Encapsulation of Textual Conclusions

[0094] To ensure the integrity and traceability of text information, and to guarantee its unique identification and retrieval within the system, this invention collects text reports written by radiologists and encapsulates them together with necessary metadata into a structured data object.

[0095] The implementation of this embodiment relies on deep integration with the hospital’s existing Reporting Information System (RIS) or Picture Archiving and Communication System (PACS).

[0096] Specifically, S2.1.1 involves deploying an API event listening module that continuously monitors the status of the reporting system. This module is activated when a radiologist completes an image review on the reporting interface. and diagnostic opinions When a field is written and key operations such as saving, submitting, or approving are performed, this event will be captured by the listening module.

[0097] After capturing the event (S2.1.2), an atomic data encapsulation operation is performed. Contextual metadata related to the report is synchronously retrieved from the RIS / PACS system, and this metadata, along with the text content, is organized into a unified structured data object. The generated second prior information is denoted as... Preferably, JSON format is used. A specific Example as follows:

[0098] {"reportID": "RPT20250903-00123","patientID": "PAT-654321","studyInstanceUID": "1.2.840.113619.2.400...","radiologistID": "DOC-789","timestamp": "2024-09-03T10:45:00Z","imagingModality": "CT","textContent": {"findings": "...","impression": "..."}}

[0099] Among them, the findings and impression in the textContent field correspond to... and .

[0100] S2.2 Analysis and Multiple Representation of Image Annotation Evidence

[0101] Overcoming the randomness and inaccuracy of manual annotation, this invention transforms the varied graphic annotations manually drawn by radiologists on two-dimensional images into structured data containing precise spatial positioning, standard geometric shapes, and various mathematical expressions.

[0102] S2.2.1 Coordinate system normalization based on affine transformation

[0103] In order to eliminate geometric distortions caused by software view scaling, translation and other operations, this invention maps the annotations drawn by radiologists on the monitor screen to the pixel coordinate system of the image data itself.

[0104] When the doctor makes annotations, the sequence of original trajectory points annotated in the screen coordinate system is captured. : ; This represents the first point in the set. It is a two-dimensional coordinate point. This represents the last point in the set. This is the total number of points originally captured. If the system records 200 points while the doctor is drawing a circle, then... The value is 200. This point is the last point the doctor made when he lifted his pen.

[0105] At the same time, obtain the current view affine transformation matrix from the graphics rendering engine. This matrix encapsulates combined translation, rotation, and scaling operations. Using the view affine transformation matrix, each screen coordinate point... Convert to pixel coordinates on image slices This transformation process can be represented using homogeneous coordinates:

[0106] ;

[0107] ;

[0108] in, It is an input screen coordinate point. It is the corresponding pixel coordinate point of the output. It is The matrices together determine scaling and rotation. It is a translation vector.

[0109] After this step, the original screen trajectory point sequence The initial polygon vertex sequence converted to pixel coordinates .

[0110] S2.2.2 Initial Polygon Vertex Sequence Perform geometric shape extraction

[0111] The input for geometry extraction is the initial vertex sequence. and a distance tolerance threshold For a sequence of initial vertices The first point and end point Find the distance connecting line segments The midpoint with the largest vertical distance .

[0112] , This indicates the calculation of vertical distance;

[0113] If the maximum distance is greater than the threshold Then see As a key feature point and retain it, then with Divide the original curve into two segments as the boundary. and And recursively call these two segments respectively. If the maximum distance is less than the threshold If all intermediate points can be approximated by line segments formed by the first and last points, then they are discarded.

[0114] With point From and Taking the perpendicular distance of a defined line segment as an example, the calculation formula is as follows:

[0115]

[0116] After geometric shape extraction, a refined polygon vertex set is obtained, which significantly reduces the number of vertices but retains the core shape of the original annotation. .

[0117] S2.2.3 Binary Mask Generation Based on Ray Casting

[0118] Based on the refined polygon vertex set Generate a binary mask matrix .

[0119] Ultimately, for each annotation made by a radiologist, the system generates a highly information-rich structured object. The collection of all labeled objects constitutes the third prior information. Q represents the number of two-dimensional images manually drawn by the radiologist.

[0120] Compared to the vague text descriptions and static screenshots in traditional reports, the structured data generated in this embodiment accurately records the real spatial location, pixel-level range, and standard geometric shape of the diagnostic annotations, providing information fidelity and richness for subsequent quantitative analysis.

[0121] Next, this embodiment will analyze the current case (based on prior information). (Definition) A method that matches and extracts one or more of the most similar reference cases whose results have been validated from a large historical case knowledge base.

[0122] In clinical practice, physicians' diagnostic decision-making process relies heavily on past experience and memory of typical cases. This embodiment aims to digitize, scale, and refine this case-based reasoning model using artificial intelligence technology. It utilizes the unique clinical textual information of the current case (first prior information). ) and image annotation information (second prior information) By effectively combining these elements, a comprehensive case profile that can be understood by computers is formed. Based on this profile, clinical twin cases that can be used for reference can be efficiently and accurately found from massive amounts of data. Through this embodiment, the system will generate crucial fourth prior information.

[0123] S3: Similar Case Retrieval and Reference System Construction Based on Large Models

[0124] S3.1: Constructing a Joint Query Vector. To achieve accurate cross-modal retrieval, the first step is to transform textual and visual information, which vary in origin and form, into a unified and measurable multimodal feature vector. This step details how to extract deep semantic features from clinical text and visual features from image annotations, and how to intelligently combine them into a joint query vector that comprehensively represents the core features of the current case through an attention fusion mechanism.

[0125] S3.1.1 Clinical text feature extraction, first prior information It contains structured clinical entities such as "cough," "two weeks," and "nodule in the upper lobe of the right lung," which are key to defining the clinical context of a case. In order for the model to understand the logical relationships between these scattered entities, they need to be organized into a natural language sequence with context.

[0126] from Extract the core entity text and, based on a preset template, generate the query text in the information set, denoted as . .

[0127] For example: "The patient complains of a cough for two weeks, and a nodule is clinically suspected in the upper lobe of the right lung." Then, it invokes information related to the first prior. Clinical BERT models with the same extraction process right Encode the text and extract the output vector corresponding to the [CLS] tag in the final hidden layer as the text feature vector representing the core semantics of the entire text. .

[0128] in, It is the output text feature vector, whose dimension is... The hidden layer dimension is consistent with that of the BERT model.

[0129] S3.1.2 Image annotation visual feature extraction, third prior information The image provides the precise location and outline of the lesion. However, finding visually similar cases requires not only outline information but also an understanding of the radiographic features such as texture and density exhibited by the pixels within the outline. Therefore, it is essential to extract the depth visual features of the region from the raw image data.

[0130] Verification information Select one or more core annotation objects. Using its corresponding binary mask matrix The minimum bounding box is calculated, and the labeled region image patch is cropped from the original image slice containing the bounding box. To capture fine features, this embodiment uses a Transformer as the visual encoder, denoted as . .

[0131] in Representing three-dimensional volume data The first in Zhang slice, Indicates based on the mask matrix The image is cropped within a specified range. It is the output visual feature vector. For output dimensions.

[0132] S3.1.3 Multimodal attention feature fusion,

[0133] The multimodal attention feature fusion input is a text feature sequence from S3.1.1. and visual feature sequences from S3.1.2 .

[0134] Simply concatenating textual and visual feature vectors fails to capture the deep, intrinsic connection between the two. For example, the description of "rough edges" in clinical text should be strongly correlated with specific image patches in imaging features that represent "irregular boundaries and spiky appearance." Conversely, specific visual features presented in images should, in turn, enhance or confirm the weight of corresponding descriptions in the text.

[0135] S3.1.3.1 Bidirectional Multi-Head Cross-Attention Calculation

[0136] a) Text-to-visual attention aims to use textual information to guide the understanding of visual features. That is, for each word in the text sequence (such as "rough"), the model will focus on the most relevant image patch in the visual sequence and use this relevant visual information to enrich the representation of the current word.

[0137] In this calculation, the text feature sequence As a query, visual feature sequence It serves as both a key and a value. The computation process follows a standard multi-head attention mechanism. This mechanism first linearly projects the query, key, and value onto... Each head has a dimension of [missing information]. Attention weights are calculated independently within each head and summed using a weighted average. Finally, the results from all heads are concatenated and linearly projected.

[0138] Among them, the The size is calculated as follows:

[0139]

[0140] in: It's about the number of heads. It is the first The learnable projection matrix corresponding to each size, It is the final output projection matrix. It is a new text feature sequence after being adjusted by visual information.

[0141] B) Visual-to-text attention aims to use visual information to guide the understanding of textual features. That is, for each image patch in the visual sequence, the model focuses on the words in the text sequence that can most accurately describe it, and uses these related textual semantics to enrich the representation of the current image patch.

[0142] Visual feature sequence As a query, text feature sequence It can be used as both a key and a value. Finally obtained That is, a new visual feature sequence after being adjusted by text information.

[0143] S3.1.3.2 Residual Connectivity and Layer Normalization

[0144] The attention calculation outputs in both directions are residually connected to their original inputs, and then layer normalization is performed.

[0145]

[0146]

[0147] Among them, the layer normalization function .

[0148] S3.1.3.3 Feedforward Network (FFN), two normalized feature sequences and Each is passed through a feedforward network with the same structure. This feedforward network consists of two linear layers and a ReLU activation function.

[0149] Similarly, the output of FFN also undergoes a residual connection and layer normalization: in, These are learnable parameters of FFN.

[0150] S3.1.3.4 Feature Pooling and Final Projection

[0151] After the above complex fusion process, two deeply fused feature sequences were obtained. and To perform similarity retrieval, these sequence information need to be aggregated into a single, fixed-dimensional joint query vector.

[0152] This embodiment employs a pooling strategy that combines specific markers and global averaging.

[0153] The system extracts the final output vector of the [CLS] token (i.e., the first token) representing global information from the text sequence. And perform average pooling on all vectors of the visual sequence. Then, these two vectors are concatenated and passed through a final linear projection layer to obtain the final query vector.

[0154]

[0155] in and These are the learnable parameters of the final projection layer, and the output...

[0156] S3.2: Semantic retrieval of similar cases. Traditional similarity retrieval, while efficient, has a fundamental flaw: it uses a uniform distance measurement across a single vector space, failing to distinguish the importance of different feature dimensions in specific clinical scenarios. It also completely ignores the valuable, structured metadata (such as patient age and gender) and hierarchical relationships of diagnostic information within the case. For example, a traditional search might match a lung nodule in an 80-year-old high-risk male smoker to a 30-year-old healthy woman, despite their strikingly similar imaging morphology, their clinical significance being entirely different. Such superficially similar but fundamentally misleading matching results have extremely low clinical reference value and may even be misleading.

[0157] To overcome the aforementioned shortcomings, this embodiment proposes a two-stage hybrid retrieval and reordering method. In the first stage, the HNSW algorithm is used to efficiently recall a candidate set that is initially relevant in terms of multimodal features from a massive case database. In the second stage, a hierarchical and metadata-aware reordering algorithm is used to perform a secondary refinement and ranking of the cases in the candidate set.

[0158] S3.2.1 Efficient Candidate Set Recall Based on HNSW

[0159] First, utilize the joint query vector generated in S3.1 In a pre-built case vector database that has been indexed using the HNSW algorithm In this process, a fast ANN search is performed. However, unlike traditional methods, the goal of this stage is not to find a unique best match, but to recall a subset of matches. (For example, The set of candidate cases that are closest in vector space. .

[0160] in, It is the set of identifiers for the candidate cases to be recalled. It is the preset candidate set size.

[0161] S3.2.2 Hierarchical and Metadata-Aware Reordering Algorithm

[0162] The candidate set Each case in A more refined second-stage scoring process was conducted. The final comprehensive score was then determined. It consists of three weighted components:

[0163] in These are preset weighting coefficients that satisfy... These controls the importance of three different similarities in the final decision.

[0164] Vector similarity components This component preserves the similarity of the underlying multimodal features and represents the degree of matching between the image and text features.

[0165]

[0166] Clinical metadata similarity component This component incorporates structured clinical information, making the matching more consistent with clinical logic.

[0167] The system extracts query cases and candidate cases. A set of key metadata fields (e.g., age, gender, smoking history, etc.), totaling Each field has a normalized similarity score, which is then summed using weighted averages.

[0168]

[0169] in, It is the first Importance weights of individual metadata fields . This is the first The similarity function is defined for each field. For numerical data (such as age), a Gaussian kernel function is used: For categorical data (such as gender), use exact matching: if else .

[0170] Diagnostic information hierarchical similarity components Disease diagnosis itself has a hierarchical structure (for example, malignant tumors are a broad category, lung cancer is a subcategory, and lung adenocarcinoma is a further subcategory). The similarity between two diagnoses should not be a simple "yes" or "no," but should be related to their distance in the medical ontological knowledge graph.

[0171] To retrieve the preliminary diagnosis of the case (from) and ) and candidate cases The gold standard for diagnosis is mapped to corresponding nodes in an internationally standardized medical ontology graph, preferably using SNOMED CT or ICD-10. Then, the similarity between these two nodes is quantified by calculating their structural relationship in the graph. This embodiment uses a depth-based calculation based on the lowest common ancestor.

[0172] in, and These are the nodes of the query diagnosis and candidate diagnosis in the ontology graph, respectively. Calculate the depth of a node in the graph (distance to the root node). Find the lowest common ancestor node of two nodes. The higher this value, the more similar the two diagnoses are in terms of classification.

[0173] S3.2.3 Determining the final reference case by traversing the candidate set. For each of the cases, calculate its final comprehensive score. Then, the case with the highest overall score is selected as the final best reference case. .

[0174]

[0175] S3.3: Constructing reference diagnostic information

[0176] Retrieved case identifiers It is merely an index. To make it truly serve as a reference, the complete, authoritative, and final information on the case is retrieved from the hospital's original database and transformed into structured data compatible with other prior information formats in this system.

[0177] Use identifiers With the patient's authorization, the corresponding complete historical medical record is retrieved from the database. From the gold standard documents of this medical record, such as pathology reports and follow-up records, a set of predefined core reference information is extracted. This information is encapsulated into a structured JSON object, serving as the final fourth prior information, denoted as […]. .

[0178] A specific Example as follows:

[0179] {"referenceCaseID": "CASE-HIST-09876","similarityScore": 0.92,"groundTruth": {"finalDiagnosis": "Lung adenocarcinoma (AC)","pathologyReport": "Microscopic examination revealed atypical glandular infiltrative growth...","genomicMarker": "EGFR L858R mutation"},"keyImagingFeatures":{"volume_mm3": 2450.0,"sphericity": 0.78,"spiculation_level": "high","referenceImageUID": "1.2.840.113619.2.300..."}}

[0180] The following embodiment discloses a video target segment retrieval and verification method guided by multidimensional prior information. In the previous embodiments, the system has successfully structured static and discrete information from clinicians, radiologists, and historical similar cases into a series of precise prior information (…). However, a complete image examination is a sequential data stream.

[0181] This embodiment utilizes acquired multidimensional prior information as navigation signals to automatically and accurately locate key video segments corresponding to the lesions or abnormal areas described by this prior information in long-term, complex inspection videos. Furthermore, to ensure the accuracy of the retrieval results, this embodiment introduces a semantic consistency-based verification mechanism to perform secondary filtering on the located segments, ensuring that the final output video segments are truly substantial and dynamic evidence.

[0182] S4: Multidimensional Prior-Guided Video Retrieval and Verification

[0183] S4.1: Constructing multimodal temporal features. In order for computers to perform temporal analysis on videos, the original pixel data of the video and the discrete prior information as guiding signals must first be converted into feature sequences suitable for processing by the temporal model.

[0184] S4.1.1 Visual temporal feature extraction: Whether it is continuous tomographic scanning of CT or real-time ultrasound exploration, the data is essentially a sequence of images arranged in time. In order to capture the dynamic features of lesions changing with the viewing angle or time, the entire three-dimensional volume data or video stream must be transformed into a time series of feature vectors.

[0185] First, the original three-dimensional volume data (Or video stream) is divided on the timeline into There are several time segments of equal length, which may overlap. For each segment, a pre-trained 3D convolutional neural network on a medical image video dataset is used to extract its spatiotemporal features. in, It is the input raw video stream. This represents a feature extractor for a three-dimensional convolutional neural network. It is the final output visual temporal feature sequence. It is the length of the sequence (i.e., the number of segments). It is the dimension of the feature vector of each segment.

[0186] S4.1.2 Prior information guides feature encoding, prior information It contains rich descriptions of the target lesion's location, morphology, nature, and clinical background, which are the core clues guiding video retrieval. These scattered clues must be integrated and encoded into a single guiding vector that represents the retrieval intent.

[0187] From prior information and In the process, all text entities and phrases related to the lesion description are extracted and concatenated into a guiding summary text. Then, the clinical BERT model is invoked. The abstract is encoded, and the output vector corresponding to its [CLS] tag is extracted as a unified guiding feature vector. .

[0188]

[0189] S4.2: Event Localization Guided by Prior Information

[0190] The system deeply integrates non-temporal retrieval intent (guided feature vector) with temporal video content (visual temporal features), and learns the global context through a self-attention mechanism, ultimately predicting possible target segments at each time point.

[0191] S4.2.1 Fusion of Guiding Information and Temporal Characteristics

[0192] To ensure that each time segment of the video can perceive the search intent, a guiding feature vector will be used. The sequence is replicated and expanded along the time dimension, making its length equal to that of the visual temporal feature sequence. length The two are identical, and then they are concatenated along the feature dimension to form a guided enhancement visual feature sequence. .

[0193] in, vector copy Second-rate.

[0194] S4.2.2 Locating Transformer encoding and proposal header prediction will guide the enhanced visual feature sequence. The input is fed into a localization Transformer consisting of multiple stacked standard Transformer encoders. This model is able to capture long-range temporal dependencies between video clips, while the computation at each location is modulated by guiding information.

[0195] in, It is a sequence of localization features after temporal context modeling.

[0196] Location feature sequence The input is fed into a proposal generation head. This head consists of three parallel one-dimensional convolutional networks, each responsible for predicting one parameter. For each time point in the sequence... The head will predict three values: center point offset. Logarithmic scale length and confidence score .

[0197] Based on the three predicted values, the specific event proposal is decoded.

[0198]

[0199]

[0200] in It is the center of the current time step. It is preset and time-point The relevant prior anchor length, It is the Sigmoid activation function, which maps the confidence score to the (0,1) interval.

[0201] Ultimately, a series of... Candidate video segments defined by triples.

[0202] S4.3: Semantic Validation and Filtering of Retrieved Fragments

[0203] To ensure the accuracy of the final retrieval results, this step introduces a "generation-comparison" verification loop. The system generates a text description for each high-confidence candidate fragment, and then compares this machine-generated description with our existing prior information on semantic similarity, retaining only those fragments with highly consistent content.

[0204] For each confidence level Exceeding the preset threshold Candidate segments:

[0205] S4.3.1 Fragment Description Generation

[0206] According to its From the original visual temporal features Extract the corresponding feature subsequences. Input these subsequences into a pre-trained video description generation model, which can be a GRU-based or small Transformer-based decoder. The model generates a natural language text describing the content of the segment, denoted as the generated description. .

[0207] S4.3.2 Semantic Similarity Verification

[0208] The system starts from prior information and In the process, the text describing the characteristics of the lesions is extracted again to form a target description. A sentence-based BERT model is used. , respectively and Encoded as fixed-dimensional sentence embedding vectors and .

[0209]

[0210]

[0211] Calculate the cosine similarity between the two embedded vectors as the semantic consistency score. .

[0212] S4.3.3 Final Segment Filtering

[0213] Only semantic consistency score Exceeding the preset verification threshold Only candidate fragments are considered valid search results. All verified fragments are collected into a final search fragment set. middle.

[0214] Next, this embodiment discloses a method for deep understanding and conclusion generation based on retrieved video segments. The system has successfully located and verified a set of key video segments highly correlated with prior information from the complete inspection video. However, these video clips are still raw visual data streams and have not yet been transformed into meaningful information that doctors and patients can directly use.

[0215] Therefore, this embodiment aims to perform in-depth analysis and reasoning on these dynamic video evidences, and to automatically generate two complementary pieces of information that serve different clinical goals by comprehensively utilizing all prior information.

[0216] S5: Information Generation Based on Video Segment Understanding

[0217] S5.1: Answering Reference Clinical Questions

[0218] To generate information capable of answering clinical questions, this embodiment proposes a reference-guided visual question-answering model. This model does not simply answer a question about the video content; instead, it incorporates similar cases as strong prior knowledge into the question-answering reasoning process. This ensures that the generated answer is based not only on objective evidence from the current video but also on empirical references from historical gold-standard cases.

[0219] S5.1.1 Multimodal input representation: To answer a clinical question, the model needs to understand three core elements: the question itself (from outpatient doctors), visual evidence (from retrieved video clips), and expert experience available for reference (from similar cases).

[0220] Extract the core examination objectives of the outpatient doctor from the first prior information, denoted as [text missing]. Using clinical BERT models We encode it. Specifically, we extract the output vector corresponding to the [CLS] label in its final hidden layer as the problem feature vector. .

[0221] The final set of retrieved fragments The visual temporal features of all video clips are concatenated and aggregated through a temporal attention pooling layer to obtain a single video feature vector representing the entire key event. .

[0222] Extract key baseline truth texts (such as final diagnosis and pathological description) from the fourth prior information, denoted as... Using the same method Encode it and extract the output vector corresponding to its [CLS] label as the reference feature vector. .

[0223]

[0224] in This represents the operation of extracting the first (i.e., [CLS]) vector from the model's output sequence.

[0225] S5.1.2 Reference-guided attention fusion

[0226] When answering the question, "Does this nodule have spiculation?", if the reference case (a confirmed malignant nodule) clearly indicates that "spiculation is a key feature," then the model should pay more attention to visual features related to "spiculation" when analyzing the current video. This guidance mechanism is achieved through attention fusion.

[0227] This embodiment uses a bilinear attention model to fuse these three feature vectors. The question vector... and reference vector Perform element-wise multiplication to generate a guiding query vector that includes the common interests of both parties. Then, this guiding vector is used to explore video features. .

[0228]

[0229] Calculate the attention weights between the guide vector and the video vector. This weight is then used to adjust the video features, resulting in attention-weighted video features. .

[0230]

[0231]

[0232] in It is a weight matrix used to align two vector spaces. It is a function that passes through the Sigmoid function. Normalized scalar attention weights.

[0233] Finally, all the information is fused to generate the final fused feature used for answer decoding. .

[0234]

[0235] S5.1.3 Conclusion Text Generation

[0236] Final fusion features The input is fed into a GRU-based text decoder. The generated natural language sentence is conclusion 1, denoted as... .

[0237] For example: "In response to the query 'Does the nodule have spiculations?', the analysis of the video clips, combined with the characteristics of similar confirmed lung adenocarcinoma cases, showed that the nodule presented clear spiculations and pleural traction signs, which are consistent with the key imaging features of the reference case."

[0238] S5.2.1 Construction of Image Feature List

[0239] To objectively review a report, a clear review standard is needed. This standard should include all the features already mentioned in the report, as well as features deemed appropriate after referring to a more authoritative case.

[0240] Prior information The text conclusion Perform named entity recognition to extract all entities of type "Disease / Symptom (DIS / SYM)" and construct a list of reported features. .

[0241] Prior information The text "keyImagingFeatures" and "groundTruth" are parsed to extract descriptive imaging features, which are then used to construct a reference feature list. .

[0242] The two lists are merged and duplicates are removed to form the final unified feature list. .

[0243] S5.2.2 Verification of Video Features

[0244] The validation task was transformed from open-ended questions into a series of specific true / false questions. For each feature on the list (such as "pleural traction"), the model needed to find evidence in the video and answer "yes" or "no".

[0245] For a unified feature list Each feature text in Encode it into a feature query vector The query vector is then compared with the video feature vector. The concatenated vectors are then fed into a lightweight binary classifier consisting of two fully connected layers, specifically designed to predict the probability of feature presence.

[0246]

[0247] in, This represents the multilayer perceptron classifier. It is the Sigmoid function. It is the first Each feature is predicted to have a probability of being present in the video.

[0248] S5.2.3 Generate a Differentiation Analysis Report

[0249] Traverse the list Each feature in the video is analyzed based on its source and the probability of it being verified in the video. They are categorized into three types:

[0250] Agreement: If the feature and .

[0251] Potential difference term): If feature but .

[0252] Additional discovery items: If features and and .

[0253] These three types of information are summarized to form a structured audit report, which is Conclusion 2, denoted as... .For example:

[0254] {"agreement": [ {"feature": "spiculated sign", "confidence": 0.95}, ... ],"discrepancy": [ {"feature": "predominantly solid component", "confidence": 0.21, "suggestion": "video shows a tendency towards mixed ground glass density"} ],"supplemental": [ {"feature": "pleural traction", "confidence": 0.88, "source": "refer to case E4"} ]}

[0255] This specification also proposes a digital fusion management system based on a large medical model, which includes:

[0256] Prior information generation module: Acquires and generates multiple sets of prior information, including: prior information generated based on clinically relevant text, second and third prior information generated based on preliminary diagnostic information, and fourth prior information generated based on historical case knowledge base retrieval.

[0257] Video retrieval module: Based on the first part of prior information, the inspected video is processed to locate one or more candidate video segments, and the candidate video segments are verified based on the second part of prior information to obtain the final set of retrieved segments;

[0258] Conclusion generation and verification module: Based on the final set of retrieved fragments and the multiple sets of prior information, it generates a first conclusion to answer clinical questions and a second conclusion to verify preliminary diagnostic information.

[0259] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned digital fusion management method based on a large medical model.

[0260] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned digital fusion management method based on a large medical model.

[0261] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0262] In this specification, the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the descriptions of the embodiments described later are relatively simple, and relevant parts can be referred to the descriptions of the foregoing embodiments.

[0263] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A digital fusion management method based on a large medical model, characterized in that, The method includes: Multiple sets of prior information are acquired and generated, including: first prior information generated based on clinically relevant text, second and third prior information generated based on preliminary diagnostic information, and fourth prior information generated based on historical case knowledge base retrieval. Guided by the first part of prior information, the inspected video is processed to locate one or more candidate video segments, and the candidate video segments are verified based on the second part of prior information to obtain the final set of retrieved segments. Based on the final set of retrieved fragments and the multiple sets of prior information, a first conclusion is generated to answer clinical questions, and a second conclusion is generated to review preliminary diagnostic information. The location of one or more candidate video segments is achieved through a localization model guided by prior information, which performs the following calculation steps: a) The guiding feature vector representing the search intent Visual temporal feature sequence of the inspection video The fusion process is performed to generate a guided, enhanced sequence of visual features. The calculation formula is as follows: in, It is the length of the visual temporal feature sequence. The operation copies the guiding feature vector. Second-rate, The operation involves concatenation along the feature dimension; b) The guided enhanced visual feature sequence The input is fed into a temporal context model to generate a localization feature sequence modeled with temporal relationships. ; c) Based on the location feature sequence For each point in time Decode the confidence level of the candidate video segments: in, It is the confidence score predicted from the location feature sequence at time point t. It is the Sigmoid activation function.

2. The digital fusion management method based on a large medical model according to claim 1, characterized in that, The step of generating the fourth prior information, and the step of constructing the joint query vector, includes: generating a text feature sequence based on the first prior information. A visual feature sequence is generated based on the third prior information and the corresponding image data. The text feature sequence and visual feature sequence are deeply fused using a bidirectional multi-head cross-attention module, which includes: a) Perform text-to-visual attention computation to generate a text feature sequence modulated by visual information. ; b) Perform visual-to-text attention computation to generate a sequence of visual features modulated by textual information. ; The and After residual connection and layer normalization with their respective original inputs, the joint query vector is generated through feature pooling and projection.

3. The digital fusion management method based on a large medical model according to claim 2, characterized in that, The fourth prior information generated based on the historical case knowledge base includes similarity retrieval: A two-stage hybrid retrieval and rearrangement method is employed, the method comprising: In the first stage, based on the joint query vector, an approximate nearest neighbor search algorithm is used to recall a query vector containing... A set of candidate cases; In the second stage, a hierarchical and metadata-aware re-ranking algorithm is used to calculate a comprehensive score for each case in the candidate case set and then re-rank them. The formula for calculating the comprehensive score is as follows: in, Candidate Case Overall score These are preset weighting coefficients. , and These represent vector similarity, clinical metadata similarity, and diagnostic information hierarchical similarity, respectively.

4. The digital fusion management method based on a large medical model according to claim 3, characterized in that: The similarity of the diagnostic information hierarchy The calculation method is as follows: The preliminary diagnosis of the query case and the gold standard for the diagnosis of the candidate case are mapped to the corresponding nodes in the medical ontology graph, respectively. The similarity between two nodes is quantified by calculating the structural relationship between them in the graph: in, and These are the nodes representing the query diagnosis and candidate diagnoses in the graph. Calculate the depth of the node. Find the lowest common ancestor of two nodes.

5. The digital fusion management method based on a large medical model according to claim 4, characterized in that: The step of generating the first conclusion is implemented through a reference-guided visual question-answering model: Problem feature vectors and reference eigenvectors Perform element-wise multiplication to generate a guiding query vector. ; Using the guiding query vector For video feature vectors Attention-weighted features are obtained by performing attention weighting. : in It is a learnable alignment weight matrix. It is attention weight; The problem feature vector, attention-weighted video feature vector, and reference feature vector are fused to generate the final fused feature for answer decoding.

6. The digital fusion management method based on a large medical model according to claim 5, characterized in that: The steps for generating the second conclusion include: Based on the second and fourth prior information, a unified feature list containing all imaging features to be verified is constructed. ; For each feature in the list, a binary classifier is used to calculate the probability that it exists in the final set of retrieved fragments. ; Based on the source of each feature and its corresponding probability If the value exceeds a preset threshold, the feature will be categorized into one of the following: consistent items, potential differences, or supplementary findings, in order to generate a differential analysis report.

7. A digital fusion management system based on a medical big data model, the system being used to execute a digital fusion management method based on a medical big data model as described in any one of claims 1-6, characterized in that, The system includes: Prior information generation module: acquires and generates multiple sets of prior information, including: prior information generated based on clinically relevant text, second and third prior information generated based on preliminary diagnosis information, and fourth prior information generated based on historical case knowledge base retrieval; Video retrieval module: Based on the first part of prior information, the inspected video is processed to locate one or more candidate video segments, and the candidate video segments are verified based on the second part of prior information to obtain the final set of retrieved segments; Conclusion generation and verification module: Based on the final set of retrieved fragments and the multiple sets of prior information, it generates a first conclusion to answer clinical questions and a second conclusion to verify preliminary diagnostic information.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements a digital fusion management method based on a medical big data model as described in any one of claims 1-6.

9. A computer-readable storage medium storing a computer program that, when executed by a processor, implements a digital fusion management method based on a medical big data model as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Atlas-driven intelligent medical image retrieval method and system

    CN120407828A

  • Medical image report generation method and system based on large language model

    CN120412875A