Training method, interpretation method and device for dynamic ultrasound image interpretation model
By training a deep learning model and adding a rule constraint layer, the problem of relying on doctors' professional judgment for interpreting critical ultrasound images was solved, achieving highly accurate and stable dynamic ultrasound image interpretation to assist in clinical diagnosis.
Patent Information
- Application Number
- CN202511282948.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Current interpretation of critical care ultrasound images relies on the professional judgment of doctors, which is highly subjective and has unstable interpretation results. Furthermore, the training of doctors varies greatly, and the operation and interpretation are separated, making it difficult to identify complex lesions and subtle pathological changes.
By acquiring dynamic ultrasound sample data, combining expert experience with annotations, training a deep learning model, adding a rule constraint layer, and forming a dynamic ultrasound image interpretation model, the model outputs interpretation data that conforms to the expert-defined pathophysiological phenotype rules and diagnostic logic.
It improves the accuracy and reliability of ultrasound image interpretation, enhances the ability to perceive dynamic images, reduces reliance on doctors' professional judgment, and ensures the stability of interpretation results and the effectiveness of clinical applications.
Smart Images

Figure CN120766067B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of dynamic ultrasound technology, specifically to a training method, interpretation method, and apparatus for a dynamic ultrasound image interpretation model. Background Technology
[0002] Critical care ultrasound, as an important tool for interpreting the pathophysiology of critically ill patients, has been widely used in clinical practice. It can reflect the functional status of vital organs such as the heart and lungs in real time and dynamically, providing crucial information for the diagnosis, treatment decisions, and prognostic assessment of critically ill patients.
[0003] Currently, mainstream dynamic ultrasound image interpretation relies on the professional knowledge and clinical experience of critical care physicians. Critical care physicians need to undergo systematic ultrasound training to master image acquisition and interpretation skills. However, the existing model has many limitations. On the one hand, there is significant variation in physician training, with operation and interpretation separated. Training focuses on operational skills but lacks in-depth training in image interpretation, resulting in limited physicians' ability to identify complex lesions, rare signs, or special sections. On the other hand, image interpretation itself is complex; the same ultrasound manifestation may correspond to multiple etiologies, requiring comprehensive judgment based on clinical context; some early or minor pathological changes have subtle ultrasound manifestations that are difficult to identify; critically ill patients experience rapid changes in their condition, with dynamic changes in ultrasound manifestations, requiring profound pathophysiological knowledge and clinical experience for accurate interpretation.
[0004] Current methods for interpreting critical care ultrasound images rely solely on the doctor's professional judgment, which is highly subjective and results in inconsistent interpretation outcomes. Summary of the Invention
[0005] In view of this, the embodiments of this application aim to provide a training method, interpretation method and device for a dynamic ultrasound image interpretation model, so as to perform dynamic ultrasound interpretation, avoid the problems of relying on the professional judgment of doctors, being greatly affected by subjectivity and having unstable interpretation results.
[0006] This application provides a training method for a dynamic ultrasound image interpretation model, including:
[0007] Acquire dynamic ultrasound sample data;
[0008] Based on expert experience, the dynamic ultrasound sample data is annotated to obtain training data; wherein, the annotation includes: temporal key points, dynamic ROI trajectories, logical chain labels, and interpretation data labels used to characterize the expected model output;
[0009] Based on the training data, a pre-built deep learning model is trained to obtain a dynamic ultrasound image interpretation model.
[0010] The interpretation data is the output of the dynamic ultrasound image interpretation model, including dynamic change analysis data, and the interpretation data conforms to the expert-defined pathophysiological phenotype rules and diagnostic logic.
[0011] In some embodiments, a pre-built deep learning model is trained based on the training data to obtain a dynamic ultrasound image interpretation model, including:
[0012] Based on the dynamic ultrasound sample data and the time-series key points and dynamic ROI trajectories in the corresponding annotations in the training data, the deep learning model is trained so that the deep learning model can be used to identify dynamic ROI trajectories and dynamic change analysis data to obtain a preliminary model.
[0013] Based on expert experience, a rule constraint layer is added to the preliminary model to obtain the target model;
[0014] The rule constraint layer is used to interpret dynamic ROI trajectories and dynamic change analysis data to obtain pathological information that conforms to expert definitions.
[0015] The dynamic ultrasound image interpretation model is obtained by training the target model based on the training data.
[0016] In some embodiments, it also includes:
[0017] Incremental learning is performed on the dynamic ultrasound image interpretation model.
[0018] In some embodiments, it also includes:
[0019] The dynamic ultrasound image interpretation model was clinically validated.
[0020] If the verification fails, the dynamic ultrasound image interpretation model will be retrained.
[0021] In some embodiments, the dynamic ultrasound sample data further includes: dynamic ultrasound data of the heart, lungs, blood vessels, gastrointestinal tract, kidneys, and organs such as the brain.
[0022] This application also provides a training device for a dynamic ultrasound image interpretation model, comprising:
[0023] The acquisition module is used to acquire dynamic ultrasound sample data;
[0024] The annotation module is used to annotate the dynamic ultrasound sample data based on expert experience to obtain training data; wherein, the annotation includes: temporal key points, dynamic ROI trajectories, logical chain labels, and interpretation data labels used to characterize the expected model output;
[0025] The training module is used to train a pre-built deep learning model based on the training data to obtain a dynamic ultrasound image interpretation model.
[0026] The interpretation data is the output of the dynamic ultrasound image interpretation model, including dynamic change analysis data, and the interpretation data conforms to the expert-defined pathophysiological phenotype rules and diagnostic logic.
[0027] This application also provides a device for interpreting dynamic ultrasound images, comprising:
[0028] The acquisition module is used to acquire dynamic ultrasound data;
[0029] The interpretation module is used to input the dynamic ultrasound data into a preset dynamic ultrasound image interpretation model to obtain interpretation data;
[0030] The dynamic ultrasound image interpretation model is obtained through the training method described above, and is used to interpret the dynamic ultrasound data to obtain interpretation data that conforms to expert experience.
[0031] This application also provides an electronic device, including:
[0032] A processor, and a memory for storing a processor-executable program;
[0033] The processor is configured to implement, by running the program in the memory, the training method for the dynamic ultrasound image interpretation model as described above, or the interpretation method for the dynamic ultrasound image as described above.
[0034] This application provides a training method for a dynamic ultrasound image interpretation model, comprising: acquiring dynamic ultrasound sample data; annotating the dynamic ultrasound sample data based on experience to obtain training data; wherein the annotation includes: temporal key points, dynamic ROI trajectories, logical chain labels, and interpretation data labels used to characterize the expected model output; training a pre-constructed deep learning model based on the training data to obtain a dynamic ultrasound image interpretation model; wherein the interpretation data is the output of the dynamic ultrasound image interpretation model, including: dynamic change analysis data, and the interpretation data conforms to expert-defined pathophysiological phenotype rules and diagnostic logic. This setup, combined with experience-based annotation, provides high-quality training data, enabling the model output to conform to clinically relevant diagnostic results. It captures dynamic physiological processes, improving the ability to recognize subtle changes in dynamic images. By acquiring dynamic change analysis data, and ensuring that the interpretation data conforms to expert-defined pathophysiological phenotype rules and diagnostic logic, model-based interpretation avoids reliance on physicians' professional judgment, which is highly subjective and leads to unstable interpretation results. Attached Figure Description
[0035] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0036] Figure 1 This is a flowchart illustrating a training method for a dynamic ultrasound image interpretation model provided in one embodiment of this application.
[0037] Figure 2 This is a flowchart illustrating a training method for a dynamic ultrasound image interpretation model provided in one embodiment of this application.
[0038] Figure 3 This is a schematic diagram of the structure of a dynamic ultrasonic analysis device provided in one embodiment of this application.
[0039] Figure 4 This is a schematic diagram of the structure of a dynamic ultrasonic analysis device provided in one embodiment of this application.
[0040] Figure 5 This is a schematic diagram of an electronic device structure provided in one embodiment of this application. Detailed Implementation
[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0042] Figure 1 This application provides a method for training a dynamic ultrasound image interpretation model, comprising:
[0043] Step S101: Obtain dynamic ultrasound sample data;
[0044] Specifically, dynamic ultrasound sample data is collected from hospital imaging databases, clinical research datasets, or sample libraries provided by ultrasound equipment manufacturers. This data covers a variety of clinical scenarios, such as dynamic ultrasound data of the heart, lungs, blood vessels, gastrointestinal tract, kidneys, and organs such as the brain. Dynamic ultrasound sample data is usually stored in the form of video files, with common file formats including DICOM (Digital Imaging and Communications) standard format, AVI, MP4, etc.
[0045] The collected dynamic ultrasound sample data are preprocessed, including noise removal, image contrast enhancement, and image size normalization, to improve data quality and consistency and ensure the accuracy and efficiency of subsequent processing.
[0046] Step S102: Based on experience, the dynamic ultrasound sample data is annotated to obtain training data; wherein, the annotation includes: temporal key points, dynamic ROI trajectory, logical chain labels, and interpretation data labels used to characterize the expected model output;
[0047] Specifically, ultrasound experts with extensive clinical experience were invited to participate in the annotation process. These experts possess profound professional knowledge and practical experience in the field of critical care ultrasound diagnosis, and are able to accurately identify and interpret various features and pathophysiological information in dynamic ultrasound images.
[0048] Key temporal points: Marking key time points in dynamic ultrasound videos, such as the starting frame of cardiac systole and the peak frame of diastole. These key points help the model capture important changes in dynamic physiological processes, such as cardiac contraction and relaxation.
[0049] Dynamic ROI Trajectory: This involves annotating the motion trajectory of a region of interest (ROI) in dynamic ultrasound video. For example, it can annotate the motion trajectory of the endocardium throughout the cardiac cycle, or the movement path of the interventricular septum. This helps the model focus on the dynamic changes of key anatomical structures, improving the ability to identify subtle pathological features.
[0050] Logical chain labeling: Establishes the logical relationship between dynamic ultrasound image features and clinical diagnosis. For example, it labels the pathophysiological significance of a specific ROI trajectory feature and the correlation between this feature and other image features, forming a complete diagnostic logical chain.
[0051] Interpreting Data Labels: Define the interpretation results that the model expects to output, including dynamic change analysis data (such as systolic left ventricular ejection fraction, diastolic mitral valve blood flow velocity, etc.) and corresponding diagnostic conclusions (such as decreased cardiac function, diastolic dysfunction, etc.), and convert them into a label format that the model can recognize.
[0052] Step S103: Based on the training data, train the pre-constructed deep learning model to obtain a dynamic ultrasound image interpretation model.
[0053] The interpretation data is the output of the dynamic ultrasound image interpretation model, including dynamic change analysis data, and the interpretation data conforms to the expert-defined pathophysiological phenotype rules and diagnostic logic.
[0054] Deep learning models can be designed with architectures suitable for processing dynamic image data. Common choices include 3D Convolutional Neural Networks (3DCNNs), Recurrent Neural Networks (RNNs), and their variants (such as LSTM and GRU). 3DCNNs can effectively extract spatiotemporal features from dynamic ultrasound images, while RNN-type models excel at processing sequential data and can capture the changing patterns of dynamic ultrasound images over time.
[0055] Input training data: Annotated dynamic ultrasound sample data is input into the deep learning model. The model learns from a large amount of training data, automatically extracts features, and establishes mapping relationships to achieve accurate interpretation of dynamic ultrasound images.
[0056] Loss function definition: Define an appropriate loss function to measure the difference between the model output and the expected result. For example, for classification tasks, the classification cross-entropy loss function can be used; for regression tasks, the mean squared error loss function can be used. At the same time, a regularization term can be introduced to prevent overfitting and improve the model's generalization ability.
[0057] Optimization Algorithm: Optimization algorithms (such as gradient descent, Adam optimizer, etc.) are used to minimize the loss function. Through iterative optimization, the model continuously adjusts its internal parameters, gradually improving its ability to fit training data and predict new data.
[0058] Model Validation and Tuning: During training, the model is periodically evaluated using a validation set to monitor performance metrics (such as accuracy, recall, and F1 score). Based on the validation results, the model's structure and hyperparameters (such as learning rate and batch size) are adjusted and optimized to improve its performance and generalization ability.
[0059] Model Output and Validation: After training, the resulting dynamic ultrasound image interpretation model can analyze and interpret input dynamic ultrasound images, outputting interpretation data that conforms to expert-defined pathophysiological phenotype rules and diagnostic logic. The model's accuracy and reliability are validated through evaluation on a test set, ensuring its effectiveness in practical clinical applications.
[0060] Finally, we reiterate the beneficial effects of this method: improving the accuracy of ultrasound image interpretation, enhancing the ability to perceive dynamic images, improving the interpretability of diagnoses, and continuously optimizing through clinical feedback, thereby promoting the intelligent development of medicine.
[0061] Specifically, based on the training data, a pre-built deep learning model is trained to obtain a dynamic ultrasound image interpretation model, including:
[0062] Based on the dynamic ultrasound sample data and the time-series key points and dynamic ROI trajectories in the corresponding annotations in the training data, the deep learning model is trained so that the deep learning model can be used to identify dynamic ROI trajectories and dynamic change analysis data to obtain a preliminary model.
[0063] Based on expert experience, a rule constraint layer is added to the preliminary model to obtain the target model;
[0064] The rule constraint layer is used to interpret dynamic ROI trajectories and dynamic change analysis data to obtain pathological information that conforms to expert definitions.
[0065] Specifically, in the solution provided in this application, the model is first trained into a preliminary model that can "see the trajectory and calculate the indicators", and then transformed into a target model that can "draw conclusions according to clinical standards" through a rule constraint layer.
[0066] In practical applications, the process of training a deep learning model based on training data to obtain a dynamic ultrasound image interpretation model is as follows:
[0067] A preliminary model was trained based on temporal key points and dynamic ROI trajectories.
[0068] 1. Input data: The dynamic ultrasound video (continuous frames), labeled temporal key points (such as the systolic start frame and diastolic peak frame), and labeled dynamic ROI trajectories (pixel-level coordinate sequence of the endocardium, interventricular septum, valve edges, etc. in the whole video) in the training data are used as input.
[0069] 2. Model Architecture Selection: A deep learning architecture suitable for processing dynamic image data, such as a 3D convolutional neural network (3DCNN), is adopted. 3DCNN can simultaneously extract spatial and temporal features of images, making it suitable for the analysis of dynamic ultrasound videos.
[0070] 3. Feature Extraction: 3DCNN extracts features from the input dynamic ultrasound video by capturing the spatiotemporal features in the video through multi-layer convolution operations, including grayscale changes in ultrasound images, edge information, and change patterns in the dynamic process.
[0071] 4. Application of attention mechanism: Introducing an attention mechanism into the network enables the model to automatically focus on the key regions indicated by the dynamic ROI trajectory, enhancing the learning ability of important features while suppressing interference from irrelevant regions.
[0072] 5. Training Objective: By defining appropriate loss functions, such as the classification cross-entropy loss function (used to determine whether the time-series key points are accurately identified) and the ROI trajectory regression loss function (used to accurately fit the dynamic ROI trajectory), the model is trained so that it can accurately identify the dynamic ROI trajectory, thus obtaining a preliminary model.
[0073] The steps above allow the deep learning model to learn to "understand" what is happening in a dynamic echocardiogram video. During training, the video itself, key time points marked by experts in each cardiac cycle (e.g., the onset of systole, diastolic peak), and the continuous motion trajectories (dynamic ROIs) of structures such as the endocardium, interventricular septum, and valves are all fed into the network. Through 3D convolution and attention mechanisms, the network learns both spatial and temporal features, ultimately accurately reproducing these trajectories. It can also calculate the most clinically relevant dynamic indicators, such as systolic left ventricular ejection fraction (LVEF) and diastolic mitral valve velocity (E / A). After this step, a "preliminary model" is obtained, possessing a fairly high image resolution capability, but not yet truly "thinking like an expert."
[0074] To ensure that the model's conclusions are consistent with clinical experience, we added a "rule constraint layer" after the initial model to incorporate clinical experience into the network.
[0075] Specifically, expert experience can be written into differentiable or executable rules, for example:
[0076] If LVEF < 50%, trigger "Reduced Contraction Function".
[0077] If E / A > 2 and e' < 8 cm / s → "diastolic dysfunction" is triggered.
[0078] If the ratio of right ventricular end-diastolic area to left ventricular end-diastolic area is greater than 1.0 and IVC is greater than 2 cm, "right ventricular volume overload" is triggered.
[0079] These rules can be implemented using logic gates, differentiable piecewise functions, or knowledge distillation loss.
[0080] The rule constraint layer essentially transforms the empirical judgments used by experts in daily diagnosis—such as "if LVEF < 50%, it suggests systolic dysfunction" or "if E / A > 2 and e' < 8 cm / s, it suggests diastolic dysfunction"—into differentiable rule nodes. After the network outputs dynamic indicators, these nodes immediately check them: any result that deviates from the rules incurs additional penalties and propagates back, forcing the model to correct itself in the next iteration. Through this fine-tuning, we obtain the final "target model." It not only provides accurate trajectories and indicators but also outputs pathological conclusions that fully conform to expert definitions, achieving seamless translation from images to clinical language.
[0081] Ultimately, the resulting dynamic ultrasound image interpretation model can accurately interpret input dynamic ultrasound images and output interpretation data that conforms to expert-defined pathophysiological phenotype rules and diagnostic logic, providing reliable auxiliary support for clinical diagnosis.
[0082] Furthermore, the solution provided in this application also includes: incremental learning of the dynamic ultrasound image interpretation model.
[0083] Specifically, in clinical applications, new dynamic ultrasound image data is continuously collected. This data comes from different patients, different equipment, and different clinical scenarios, enriching the model's understanding of various situations. The collected new data is preprocessed and labeled to ensure that its format and content are consistent with the training data, so that it can be smoothly used for incremental learning of the model.
[0084] New data is input into an existing dynamic ultrasound image interpretation model, and the model's parameters are updated using the new data. During this process, the model automatically adjusts its internal weights and biases based on the features and annotations in the new data to adapt to the distribution and characteristics of the new data. Depending on the specific circumstances, an online learning approach can be chosen, where only one or a few new samples are used to update the model at a time, allowing for rapid adaptation to new data; alternatively, a mini-batch learning approach can be chosen, where a certain amount of new data is collected periodically and then used together to update the model, thus better balancing learning efficiency and model stability.
[0085] The updated model is evaluated using a validation set, focusing primarily on metrics such as accuracy, recall, and F1 score on the new data. It's also crucial to observe whether the model's performance degrades on the original data to ensure that incremental learning hasn't led to catastrophic forgetting. If the model's performance on the new data is unsatisfactory, or if catastrophic forgetting occurs, further adjustments and optimizations are necessary. This might include adjusting the learning rate, adding regularization terms, and modifying the model structure to improve incremental learning effectiveness and generalization ability.
[0086] Once the model achieves satisfactory performance after incremental learning, the updated model is deployed to a real-world clinical application environment, replacing the original model and providing doctors with more accurate and reliable assistance in interpreting dynamic ultrasound images. After deployment, its performance in real-world applications is continuously monitored, and feedback from doctors and users is collected to promptly identify potential problems and provide a reference for the next incremental learning iteration.
[0087] The solution provided in this application also includes: clinically validating the dynamic ultrasound image interpretation model; if the validation fails, the dynamic ultrasound image interpretation model is retrained.
[0088] Specifically, dynamic ultrasound image data should be collected from actual clinical scenarios, covering a variety of pathological conditions and clinical situations to ensure the comprehensiveness and representativeness of the validation results. Simultaneously, this data should be accurately labeled and confirmed by experts to facilitate comparison with the model's output.
[0089] Identify metrics for evaluating model performance, such as accuracy, recall, F1 score, and area under the ROC curve (AUC). These metrics reflect the effectiveness and reliability of the model in clinical applications from different perspectives. Accuracy measures the model's ability to make correct diagnoses, recall focuses on the model's ability to identify actual cases, F1 score considers both accuracy and recall, and AUC assesses the model's ability to distinguish between different categories.
[0090] The prepared validation data is input into the dynamic ultrasound image interpretation model, and the model outputs diagnostic interpretation results for this data. The model's output results are compared with the expert's diagnostic results, and the aforementioned validation metrics are calculated to quantify the model's performance.
[0091] If a model fails clinical validation, a detailed analysis of the validation results is necessary to identify the reasons for its performance deficiencies. Possible causes include insufficient training data, imbalanced data distribution, an unreasonable model structure, and overfitting or underfitting. For example, if the model has low accuracy in diagnosing certain pathological types, it may be due to a lack of data of that type in the training dataset, resulting in insufficient learning of these features by the model.
[0092] Based on the problem analysis, supplement the training data to enrich the model's learning materials. If performance issues are due to insufficient training data, add more diverse dynamic ultrasound image data, especially data on case types where the model performed poorly during validation. Simultaneously, re-label and reorganize the training data to ensure its quality and consistency.
[0093] The dynamic ultrasound image interpretation model is retrained using supplemented and adjusted training data. During retraining, adjustments can be made to the model's structure or training parameters, such as increasing the model's depth or width, changing the learning rate, or adjusting the regularization term, to improve model performance. For example, if the model suffers from overfitting, the weight of the regularization term can be increased to make the model smoother and reduce overfitting to noise in the training data.
[0094] The retrained model is then clinically validated again, and the above validation process is repeated until the model's performance meets the requirements for clinical application. This process may require multiple iterations, continuously optimizing and validating the model to ensure its accuracy and reliability in real-world clinical applications.
[0095] The calculated validation metrics are analyzed to determine whether the model meets the requirements for clinical application. If the validation metrics reach predetermined thresholds, such as an accuracy rate of over 90% and an AUC of over 0.9, the model is considered to have passed clinical validation and can be used for actual clinical auxiliary diagnosis. If the validation metrics are unsatisfactory, the model is considered to have failed validation.
[0096] Specifically, the dynamic ultrasound sample data may be, but is not limited to, cardiac dynamic ultrasound data.
[0097] The solution provided in this application will be described below with reference to specific embodiments:
[0098] This invention aims to establish a specialized model training method that employs expert thought chain embedding. Modeling experts interpret cognitive paths, train AI to identify subtle signs and perform dynamic temporal analysis, and finally output standardized reports. Furthermore, it combines clinical information and patient conditions for differential support and feedback learning, while expert quality control supports selection, establishing a competitive learning path between large models and vertical models.
[0099] In this application, the Expert Chain-of-Thought (Expert CoT) refers to the explicit simulation and recording of the thought process of experts when solving problems, and integrating it into the training process of a dynamic ultrasound image interpretation model. This chain of thought helps the model more accurately understand and interpret complex features and pathophysiological information in dynamic ultrasound images. The specific content and function of the Expert Chain of Thought are as follows: The Expert Chain of Thought is a technology that simulates the thinking process of experts when solving complex problems. By progressively constructing a logical chain from problem to answer, it enables AI to understand problems more deeply and provide solutions. Unlike traditional pattern recognition or statistical learning methods, the Expert Chain of Thought emphasizes the transparency and interpretability of the reasoning process, making the AI's decision-making process closer to the thinking style of human experts. In dynamic ultrasound image interpretation, the Expert Chain of Thought helps the model identify and understand subtle features and dynamic changes in images, thereby enabling more accurate diagnosis and analysis.
[0100] I. Dynamic Annotation Process: A standardized interpretation template for dynamic videos is adopted, which is used by experts to interpret dynamic images based on the standardized interpretation template, ensuring annotation based on dynamic video information and ensuring the standardization of interpretation content.
[0101] For example:
[0102] Step 1: Experts define the critical care ultrasound pathophysiology standard fields.
[0103] Experts developed standard fields for identifying different pathophysiological abnormalities in ultrasound for severe cases and trained them for recognition, for example:
[0104] 1. Systolic left ventricular ejection fraction (LVEF) <50% → triggers the conclusion of "cardiac dysfunction";
[0105] 2. A diastolic mitral valve flow velocity E / A ratio > 2 → triggers the conclusion of "diastolic dysfunction and increased left atrial pressure";
[0106] 3. Abnormal ROI dynamic trajectory (such as central concavity of the interventricular septum) → triggers the conclusion of "pressure-induced right ventricular dysfunction".
[0107] Step 2: Dynamic ROI and Labeling
[0108] The experts annotated the ultrasound video as follows:
[0109] 1. Key temporal points (such as the starting frame of cardiac systole and the peak frame of diastole).
[0110] 2. Dynamic ROI trajectory (such as the full motion cycle trajectory of the endocardium, continuous coordinate points of the interventricular septum motion path, and dynamic changes in lung ultrasound signs such as Cheyne-Stokes lung recruitment and dynamic bronchial inflation signs): Its advantage is that it can avoid the artifact interference of single-frame images. The position coordinates of each frame image can be deduced from the clear edge trajectory images of the preceding and following time sequences and the motion features can be formed coherently. The second point is that the trajectory based on the whole cardiac cycle has different characteristics under different pathophysiological states, and learning can be carried out based on these characteristics.
[0111] 3. Logical chain tags (the binding relationship between conclusions and ROI / time series, such as the pathophysiological characteristics represented by the continuous coordinate point motion trajectory features of the ventricular septum motion path, and the relationship between ventricular septum motion and systolic / diastolic phases to determine the severity of pathophysiology).
[0112] The annotation results include spatiotemporal joint information, rather than static annotations of independent frames. This avoids local image blurring and artifact interference, and the spatiotemporal joint information provides clues to specific pathophysiological features.
[0113] Step 3: Experts define pathophysiological phenotype rules and diagnostic logic
[0114] Experts develop dynamic diagnostic logic chains for target symptoms (such as shock), for example:
[0115] 1. The diameter of the vena cava is <1.5cm, and the short axis of the vena cava is droplet / linear, and the ratio of the right ventricular end-diastolic area to the left ventricular end-diastolic area is <0.6, and the left ventricular systolic function is >50%, and CO is <4 → triggering the conclusion of "left-right ventricular mismatch, low blood volume, high left ventricular dynamics, and low output";
[0116] 2. The ratio of right ventricular end-diastolic area to left ventricular end-diastolic area >1, and the diameter of the vena cava >2cm, with the short axis of the vena cava being perfectly circular, and the left ventricular systolic function >50%, and CO <4 → triggers the conclusion of "left-right ventricular mismatch, right ventricular volume overload, and left ventricular low volume and low output";
[0117] 3. If the diameter of the vena cava is >2cm, the ratio of the right ventricular end-diastolic area to the left ventricular end-diastolic area is <0.6, the left ventricular ejection fraction (LVEF) is <50%, and CO is <4, the conclusion of "left-right ventricular mismatch, volume overload, decreased left ventricular systolic function, and low output" is triggered.
[0118] 4. The diameter of the vena cava is >1.5cm, and the ratio of the right ventricular end-diastolic area to the left ventricular end-diastolic area is <0.6, and the left ventricular ejection fraction (LVEF) is >50%, and the coronary artery velocity (CO) is >6, and the nasal cavity shows a low-tension spectrum / decreased resistance index → triggering the conclusion of "matching left and right ventricles, not low volume, not poor left ventricular systolic function, not low output, low tension".
[0119] Step 4: Post-clinical feedback learning
[0120] The initially trained model is trained on clinical cases, and the conclusions after the judgment are fed back to the clinic. Clinicians and experts conduct a comprehensive evaluation to consider whether the judgment is consistent or not, and combine the given targeted treatment and indicators such as effect and outcome to provide feedback on the degree of consistency of the initial judgment, thereby conducting post-clinical feedback learning in real-world scenarios.
[0121] For example, for a target symptom (such as shock), the model has already triggered the conclusions of "left-right ventricular matching, low blood volume, high left ventricular dynamics, and low output" based on "vena cava diameter <1.5cm, and vena cava short axis droplet / linear, and right ventricular end-diastolic area ratio <0.6, and left ventricular systolic function >50%, and CO <4". Clinicians then judge and correct the details, such as whether the phenotype conforms to or does not conform to the clinical findings. If it does not conform to the clinical findings, the phenotype features are divided into different short fields, and the fields that do not conform are selected for reverse correction. The correction content is recorded and learned. If it conforms to the clinical findings, the treatment effect is observed to see if it meets expectations. If it does conform, reinforcement learning is performed, and if it does not conform, the specific content that needs to be corrected is retrieved.
[0122] II. The model training process includes:
[0123] Phase 1: ROI Conclusion Association Pre-training Phase
[0124] The input to the model at this stage is: dynamic ultrasound video + expert-annotated ROI trajectory + conclusion label; the model's network is a 3D CNN that extracts spatiotemporal features → an attention mechanism that focuses on the ROI region → a classification layer that outputs the conclusion probability; the model's loss function includes: classification cross-entropy loss + ROI trajectory regression loss (L1 Loss).
[0125] Phase 2: Expert Rule Knowledge Distillation
[0126] The input to this stage of the model is: the input from stage 1 and expert rules (such as "if the expansion rate of ROI_A > X%, then the weight is increased by α"). The network of this stage of the model is based on the stage 1 model, with the addition of a rule constraint layer (such as logic gate modules). The loss function of this stage of the model includes: adding a rule violation penalty term (such as doubling the loss value when there is a rule conflict).
[0127] Phase 3: Incremental Learning and Clinical Validation
[0128] Dynamic updates: New case data, after being reviewed by experts, triggers incremental training of the model;
[0129] Validation mechanism: The Kappa consistency coefficient between the model output conclusions and the expert diagnosis must be ≥0.9.
[0130] Furthermore, refer to Figure 2 This application provides a method for interpreting dynamic ultrasound images, including:
[0131] Step S201: Acquire dynamic ultrasound data;
[0132] Collect dynamic ultrasound data from hospital imaging databases, clinical research datasets, or sample libraries provided by ultrasound equipment manufacturers.
[0133] Dynamic ultrasound data is usually stored in the form of video files, with common file formats including DICOM, AVI, MP4, etc.
[0134] The collected dynamic ultrasound data is preprocessed, including noise removal, image contrast enhancement, and image size normalization, to improve data quality and consistency and ensure the accuracy and efficiency of subsequent processing.
[0135] Step S202: Input the dynamic ultrasound data into a preset dynamic ultrasound image interpretation model to obtain interpretation data;
[0136] The dynamic ultrasound image interpretation model is obtained through the training method described above and is used to interpret the dynamic ultrasound data to obtain interpretation data that conforms to the expert's thought process.
[0137] Specifically, the first step is to ensure that a dynamic ultrasound image interpretation model, trained using the methods described above, already exists. This model, trained on professionally labeled data, is capable of processing dynamic ultrasound data and outputting interpretation results that conform to expert thought processes.
[0138] The preprocessed dynamic ultrasound data is input into the model.
[0139] The model analyzes and interprets the input dynamic ultrasound data, and outputs corresponding interpretation data based on the knowledge and logic learned during its training process.
[0140] The model outputs interpreted data: This data includes dynamic change analysis and conforms to expert-defined pathophysiological phenotype rules and diagnostic logic. Specifically, this data may involve the analysis of dynamic changes in different parts of the image, the identification and judgment of specific pathological features, etc., providing doctors with valuable diagnostic reference information.
[0141] By following the steps above, a well-trained dynamic ultrasound image interpretation model can be used to accurately interpret actual dynamic ultrasound data, assisting doctors in making diagnostic and treatment decisions.
[0142] Specifically, the dynamic ultrasound data includes: dynamic cardiac ultrasound data;
[0143] The interpreted data includes: diagnostic indicator data and diagnostic conclusions corresponding to the diagnostic indicator data;
[0144] The diagnostic indicator data includes at least dynamic change analysis data;
[0145] The dynamic change analysis data includes at least one of the following: systolic left ventricular ejection fraction, diastolic mitral valve blood flow velocity, and ROI dynamic trajectory.
[0146] Specifically, dynamic ultrasound data can be, but is not limited to, cardiac dynamic ultrasound data. Cardiac dynamic ultrasound data consists of a series of continuous ultrasound image frames, typically recorded in video format, reflecting the morphological and kinematic changes of organs such as the heart during dynamic processes. Each image frame contains a pixel value matrix representing the reflection intensity of tissue interfaces, which can be used to distinguish different tissue structures, such as myocardium, heart chambers, and valves. Each image frame has a corresponding timestamp, accurately recording the acquisition time during the dynamic process, which helps analyze the temporal characteristics of dynamic physiological processes such as the heart's systolic and diastolic cycles. It includes cardiac dynamic ultrasound images from different sections, such as the long-axis section, short-axis section, and four-chamber section of the heart. Different sections can display the morphology and motion of various parts of the heart, providing multi-angle perspectives for a comprehensive assessment of cardiac function.
[0147] The diagnostic indicators in the data interpretation include:
[0148] Dynamic change analysis data includes systolic left ventricular ejection fraction, diastolic mitral valve blood flow velocity, and ROI dynamic trajectory.
[0149] Other diagnostic data: In addition to the dynamic change analysis data mentioned above, it may also include other diagnostic features and indicators extracted from dynamic ultrasound data, such as heart rate, heart rhythm, size and volume of each chamber of the heart, etc.
[0150] The diagnostic conclusions are based on the analysis of diagnostic indicator data, including: cardiac function assessment (whether left ventricular systolic function is normal, whether diastolic function is impaired, etc.), valvular function assessment (whether heart valves are stenotic or insufficient, etc.), diagnosis of cardiomyopathy (whether myocardial ischemia, infarction, or cardiomyopathy exists, etc.), diagnosis of heart disease (such as coronary artery disease, dilated cardiomyopathy, rheumatic valvular heart disease, etc.), assessment of disease severity (such as heart disease stage, severity of disease, etc.), and treatment recommendations (drug treatment plan, surgical indications, etc.).
[0151] The apparatus embodiments of this application can be used to execute the method embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the method embodiments of this application.
[0152] Figure 3 The diagram shown is a block diagram of a training device for a dynamic ultrasound image interpretation model provided in one embodiment of this application. Figure 3 As shown, the device includes:
[0153] Acquisition module 31 is used to acquire dynamic ultrasound sample data;
[0154] Annotation module 32 is used to annotate the dynamic ultrasound sample data based on expert experience to obtain training data; wherein, the annotation includes: temporal key points, dynamic ROI trajectory, logical chain labels, and interpretation data labels used to characterize the expected model output;
[0155] Training module 33 is used to train a pre-built deep learning model based on the training data to obtain a dynamic ultrasound image interpretation model.
[0156] The interpretation data is the output of the dynamic ultrasound image interpretation model, including dynamic change analysis data, and the interpretation data conforms to the expert-defined pathophysiological phenotype rules and diagnostic logic.
[0157] Figure 4 The diagram shown is a block diagram of a dynamic ultrasound image interpretation device according to an embodiment of this application. Figure 4 As shown, the device includes:
[0158] Acquisition module 41 is used to acquire dynamic ultrasound data;
[0159] The interpretation module 42 is used to input the dynamic ultrasound data into a preset dynamic ultrasound image interpretation model to obtain interpretation data;
[0160] The dynamic ultrasound image interpretation model is obtained through the training method described above, and is used to interpret the dynamic ultrasound data to obtain interpretation data that conforms to the expert's thought process.
[0161] Below, for reference Figure 5 This describes an electronic device according to embodiments of the present application. Figure 5 A block diagram of an electronic device according to an embodiment of this application is illustrated.
[0162] like Figure 5 As shown, the electronic device 500 includes one or more processors 510 and memory 520.
[0163] The processor 510 may be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and may control other components in the electronic device 500 to perform desired functions.
[0164] The memory 520 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 510 may execute the program instructions to implement the training method of the dynamic ultrasound image interpretation model, the dynamic ultrasound image interpretation method, and / or other desired functions of the various embodiments of this application described above. Various contents, such as category correspondence, may also be stored in the computer-readable storage medium.
[0165] In one example, the electronic device 500 may also include an input device 530 and an output device 540, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0166] In addition, the input device 530 may also include, for example, a keyboard, mouse, interface, etc. The output device 540 can output various information to the outside, including analysis results, etc. The output device 540 may include, for example, a display, speaker, printer, and communication network and its connected remote output devices, etc.
[0167] Of course, for the sake of simplicity, Figure 5 Only some of the components of the electronic device relevant to this application are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device may include any other suitable components depending on the specific application.
[0168] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the training method of the dynamic ultrasound image interpretation model or the dynamic ultrasound image interpretation method according to various embodiments of this application as described in the "Exemplary Methods" section of this specification.
[0169] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0170] Furthermore, embodiments of this application may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the steps in the training method of the dynamic ultrasound image interpretation model or the dynamic ultrasound image interpretation method according to various embodiments of this application as described in the "Exemplary Methods" section of this specification.
[0171] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0172] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A training method for a dynamic ultrasound image interpretation model, characterized in that, include: Acquire dynamic ultrasound sample data; Based on expert experience, the dynamic ultrasound sample data is annotated to obtain training data; wherein, the annotation includes: temporal key points, dynamic ROI trajectories, logical chain labels, and interpretation data labels used to characterize the expected model output; Based on the training data, a pre-built deep learning model is trained to obtain a dynamic ultrasound image interpretation model. The interpretation data is the output of the dynamic ultrasound image interpretation model, including dynamic change analysis data, and the interpretation data conforms to the expert-defined pathophysiological phenotype rules and diagnostic logic. Based on the training data, a pre-built deep learning model is trained to obtain a dynamic ultrasound image interpretation model, including: Based on the dynamic ultrasound sample data and the time-series key points and dynamic ROI trajectories in the corresponding annotations in the training data, the deep learning model is trained so that the deep learning model can be used to identify dynamic ROI trajectories and dynamic change analysis data to obtain a preliminary model. Based on expert experience, a rule constraint layer is added to the preliminary model to obtain the target model; The rule constraint layer is used to interpret dynamic ROI trajectories and dynamic change analysis data to obtain pathological information that conforms to expert definitions. The dynamic ultrasound image interpretation model is obtained by training the target model based on the training data.
2. The training method for the dynamic ultrasound image interpretation model according to claim 1, characterized in that, Also includes: Incremental learning is performed on the dynamic ultrasound image interpretation model.
3. The training method for the dynamic ultrasound image interpretation model according to claim 1, characterized in that, Also includes: The dynamic ultrasound image interpretation model was clinically validated. If the verification fails, the dynamic ultrasound image interpretation model will be retrained.
4. The training method for the dynamic ultrasound image interpretation model according to claim 1, characterized in that, Also includes: The dynamic ultrasound sample data includes dynamic ultrasound data of the heart, lungs, blood vessels, gastrointestinal tract, kidneys, and brain.
5. A method for interpreting dynamic ultrasound images, characterized in that, include: Acquire dynamic ultrasound data; The dynamic ultrasound data is input into a preset dynamic ultrasound image interpretation model to obtain interpretation data; The dynamic ultrasound image interpretation model is obtained by training the dynamic ultrasound image interpretation model as described in any one of claims 1 to 4, and is used to interpret the dynamic ultrasound data to obtain interpretation data that conforms to the expert's thought process.
6. The method for interpreting dynamic ultrasound images according to claim 5, characterized in that, The dynamic ultrasound data includes: cardiac dynamic ultrasound data, lung dynamic ultrasound data, vascular dynamic ultrasound data, gastrointestinal dynamic ultrasound data, renal dynamic ultrasound data, and cranial dynamic ultrasound data. The interpreted data includes: diagnostic indicator data and diagnostic conclusions corresponding to the diagnostic indicator data; The diagnostic indicator data includes at least dynamic change analysis data; The dynamic change analysis data includes at least one of the following: systolic left ventricular ejection fraction, diastolic mitral valve blood flow velocity, and ROI dynamic trajectory.
7. A training device for a dynamic ultrasound image interpretation model, characterized in that, include: The acquisition module is used to acquire dynamic ultrasound sample data; The annotation module is used to annotate the dynamic ultrasound sample data based on expert experience to obtain training data; wherein, the annotation includes: temporal key points, dynamic ROI trajectories, logical chain labels, and interpretation data labels used to characterize the expected model output; The training module is used to train a pre-built deep learning model based on the training data to obtain a dynamic ultrasound image interpretation model. The interpretation data is the output of the dynamic ultrasound image interpretation model, including dynamic change analysis data, and the interpretation data conforms to expert-defined pathophysiological phenotype rules and diagnostic logic; Based on the training data, a pre-built deep learning model is trained to obtain a dynamic ultrasound image interpretation model, including: Based on the dynamic ultrasound sample data and the time-series key points and dynamic ROI trajectories in the corresponding annotations in the training data, the deep learning model is trained so that the deep learning model can be used to identify dynamic ROI trajectories and dynamic change analysis data to obtain a preliminary model. Based on expert experience, a rule constraint layer is added to the preliminary model to obtain the target model; The rule constraint layer is used to interpret dynamic ROI trajectories and dynamic change analysis data to obtain pathological information that conforms to expert definitions. The dynamic ultrasound image interpretation model is obtained by training the target model based on the training data.
8. A device for interpreting dynamic ultrasound images, characterized in that, include: The acquisition module is used to acquire dynamic ultrasound data; The interpretation module is used to input the dynamic ultrasound data into a preset dynamic ultrasound image interpretation model to obtain interpretation data; The dynamic ultrasound image interpretation model is obtained by training the dynamic ultrasound image interpretation model as described in any one of claims 1 to 4, and is used to interpret the dynamic ultrasound data to obtain interpretation data that conforms to the expert's thought process.
9. An electronic device, characterized in that, include: A processor, and a memory for storing a processor-executable program; The processor is configured to implement, by running a program in the memory, a training method for a dynamic ultrasound image interpretation model as described in any one of claims 1 to 4, or a method for interpreting dynamic ultrasound images as described in claim 5 or 6.
Citation Information
Patent Citations
Pelvic floor ultrasonic image recognition and analysis system based on deep learning
CN118628461A
Early diagnosis system and method for gonitis based on multi-modal deep learning
CN119495419A