An Automatic Echocardiography Recording System and Method with Distortion Resistance

By extracting explicit anatomical and implicit spatiotemporal features from cardiac ultrasound image frames and probe motion data, the quality of the cross-section is dynamically determined, solving the problem of cross-section distortion caused by probe jitter and cardiac motion, and improving the stability and diagnostic quality of cardiac ultrasound recordings.

CN122296950APending Publication Date: 2026-06-30RENJI HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RENJI HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
Filing Date
2026-05-29
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing automated cardiac ultrasound recording systems cannot effectively distinguish between probe jitter and sectional distortion caused by the heart's own motion, resulting in recording interruptions or uneven quality. They also lack comprehensive multimodal signal closed-loop feedback and dynamic prediction mechanisms.

Method used

By acquiring continuous cardiac ultrasound image frames and probe motion data, explicit anatomical quality features and implicit spatiotemporal fusion features are extracted to generate section quality deviation indicators. The system dynamically determines stable, transitional, or distorted states and performs video recording control based on these indicators, including compensation processing and distortion alerts.

Benefits of technology

It achieves precise decoupling of probe jitter and cardiac motion, dynamic adaptive judgment, improves video recording continuity and diagnostic quality, reduces differences in operator experience, and ensures uniformity of video recording quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122296950A_ABST
    Figure CN122296950A_ABST
Patent Text Reader

Abstract

This invention discloses an automatic cardiac ultrasound recording system and method with anti-distortion capabilities. The method includes: acquiring synchronized cardiac ultrasound image frames and probe motion state data; extracting explicit anatomical quality features and implicit spatiotemporal fusion features; generating predicted spatiotemporal features based on historical implicit features; obtaining a section quality deviation index by combining acoustic coupling quality indices; dynamically determining stable and distortion boundaries based on the deviation index within a sliding window; classifying the current frame as stable, transitional, or distorted; recording when the frame is stable and the phase meets the conditions; recording when the frame is transitional and the predicted and current features are adaptively fused and compensated based on observational uncertainty; and pausing recording and providing attribution prompts when the frame is distorted. This invention achieves precise decoupling of probe motion and cardiac motion through IMU and image fusion, and employs a dynamic threshold and residual calibration mechanism to effectively improve the continuity and diagnostic quality of automatic recording, making it suitable for intelligent control of cardiac ultrasound examinations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of ultrasound imaging, and more particularly to an automatic cardiac ultrasound anti-distortion recording system and method. Background Technology

[0002] In the field of cardiac ultrasound, existing technologies primarily focus on image acquisition, processing, and analysis to address sectional distortion caused by factors such as dynamic cardiac motion, probe jitter, and signal noise. These issues often affect image stability, diagnostic accuracy, and clinical efficiency, leading to increased manual intervention or decreased video quality. Existing technologies have proposed various methods to mitigate these challenges, such as achieving automated imaging and distortion correction through signal processing, real-time tracking, and multimodal fusion.

[0003] For example, patent application US20210085294A1 discloses a method for processing cardiac ultrasound data to determine cardiac mechanical wave information. This method involves receiving a three-dimensional data frame time series generated from ultrasound signals, where each frame contains voxel values ​​representing the acceleration components of the heart's position; identifying the frame with the maximum value for each voxel; generating a three-dimensional time propagation dataset; and calculating the time derivative to produce a three-dimensional velocity vector field. The method also includes acquiring Doppler datasets, calculating the velocity time derivative, and using electrocardiogram (ECG) to select data frame windows to improve signal quality. It improves the signal-to-noise ratio by processing distortion through clutter filtering (high-pass filtering) and spatiotemporal filtering, and automatically detects wave propagation using segmentation techniques, thereby achieving efficient automated imaging. However, this method mainly relies on post-processing optimization and lacks real-time closed-loop feedback and probe motion decoupling, failing to adequately address instantaneous sectional drift caused by probe micro-vibration.

[0004] For example, patent application US20240000433 discloses a system and method for cardiac imaging using ultrasound. This focuses on reducing distortion caused by patient motion or probe instability through advanced signal processing algorithms, thereby improving the spatiotemporal stability and diagnostic accuracy of cardiac ultrasound images. While this technology emphasizes reducing the need for manual adjustments, its implementation does not elaborate on multimodal signal synchronization or predictive compensation mechanisms, potentially leading to insufficient adaptability in complex clinical scenarios.

[0005] For example, patent application CN119970093A discloses a cardiac ultrasound imaging method and system that addresses the impact of cardiac physiological motion on elastography. First, ultrasound images of the heart are acquired, and the myocardium (e.g., the segment to be observed) is identified and tracked in real time. Then, elastography is performed. Considering the influence of cardiac chamber structure on shear wave propagation, acoustic radiation force is applied to the myocardium (e.g., geometric center or center of gravity) based on real-time position to generate shear waves, ensuring effective generation and a high signal-to-noise ratio. This includes three elastography modes: acoustic radiation force-induced, autonomous physiological shear wave, and hybrid mode; automatic mode selection based on myocardium or cross-sectional type; prompts for probe adjustment during multi-section scanning; and timed wave generation at stable cardiac phases (e.g., end-diastole, end-systole). This method mitigates motion distortion and improves image quality and repeatability through real-time tracking and position optimization, but it does not fully integrate IMU attitude signals or operator intent recognition, limiting the continuity of automatic recording in dynamic scanning mode.

[0006] Furthermore, patent application WO2022183264 discloses an ultrasound-based method and device for continuous monitoring of cardiac parameters. This emphasizes the processing of ultrasound signals to extract cardiac parameters, including multimodal integration, but specific details are limited, and it fails to comprehensively address the issues of section prediction and compensation.

[0007] For example, patent application US10595731B2 discloses a method and system for tracking and scoring arrhythmias. It uses multimodal signals such as ECG, photoplethysmography (PPG), and motion sensors to detect arrhythmias (such as atrial fibrillation) through machine learning, generates a heart health score, and provides improvement suggestions. While it supports continuous monitoring and automatic video recording, it primarily focuses on electrical signals rather than ultrasound imaging and lacks direct processing of sectional distortion.

[0008] Despite the progress these existing technologies have made in the stability and automated processing of cardiac ultrasound images, they still have limitations, such as the lack of comprehensive multimodal signal closed-loop feedback, heterogeneous feature decoupling and dynamic prediction mechanisms, and the inability to effectively distinguish between distortions caused by probe operation and anatomical changes, resulting in video interruption or uneven quality. Summary of the Invention

[0009] This invention aims to solve the technical problems of existing automatic cardiac ultrasound recording systems, such as misjudgment of cross-sectional distortion, recording interruption, and uneven quality caused by the inability to distinguish between probe jitter and the heart's own movement. It provides a system and method that can decouple the distortion source in real time, dynamically determine the cross-sectional state, and perform adaptive recording control.

[0010] The objective of this invention is achieved through the following solution.

[0011] An automatic recording method for anti-distortion echocardiography is applied to an echocardiography imaging system, the echocardiography imaging system including an ultrasound probe, a probe motion sensor, a processing unit, and a recording storage unit, the method comprising:

[0012] Acquire continuous echocardiogram image frames and probe motion state data synchronized with the echocardiogram image frames, and obtain acoustic coupling quality index and cardiac cycle phase information based on the echocardiogram image frames;

[0013] Exclusive anatomical quality features are extracted from the cardiac ultrasound image frames, and latent spatiotemporal fusion features are generated based on the cardiac ultrasound image frames and the probe motion state data.

[0014] Based on the implicit spatiotemporal fusion features of historical moments, predictive spatiotemporal features of the current moment are generated, and based on the predicted spatiotemporal features, the implicit spatiotemporal fusion features of the current moment, the explicit anatomical quality features, and the acoustic coupling quality index, the cross-sectional quality deviation index of the current frame is obtained.

[0015] Based on the cross-sectional quality deviation index of multiple consecutive frames within a preset time window, the stability judgment boundary and the distortion judgment boundary are dynamically determined, and the current frame is judged as a stable state, a transitional state, or a distorted state based on the stability judgment boundary and the distortion judgment boundary.

[0016] Recording control is executed based on the state of the current frame, where:

[0017] If the current frame is determined to be in a stable state, and the cardiac cycle phase information meets the preset recording triggering conditions, the recording storage unit is controlled to write the current frame or the buffer image frame corresponding to the current frame.

[0018] If the current frame is determined to be in a transitional state, the observation uncertainty of the current frame is determined based on the probe motion state data and the image quality information of the cardiac ultrasound image frame. The predicted spatiotemporal features and the implicit spatiotemporal fusion features at the current moment are adaptively fused according to the observation uncertainty to generate compensation information for the current frame. The current frame is then compensated based on the compensation information and written to the video storage unit.

[0019] If the current frame is determined to be distorted, pause or stop recording and output a distortion warning message.

[0020] The core technical logic of the above method lies in: by decoupling the image content from the physical movement of the probe synchronously, explicit anatomical features provide interpretable morphological constraints, implicit spatiotemporal fusion features capture data-driven dynamic patterns, and then combining acoustic quality gating, a closed-loop control link that can perceive cross-sectional deviations in real time and respond in a graded manner is constructed.

[0021] In some embodiments, the probe motion state data includes linear motion information and angular motion information of the ultrasound probe; the linear motion information and angular motion information are acquired by an inertial measurement unit disposed on the ultrasound probe; the cardiac ultrasound image frames and the probe motion state data are synchronized with timestamps, so that each cardiac ultrasound image frame corresponds to the probe motion state data within the acquisition time window of that cardiac ultrasound image frame. This hardware-level synchronization ensures precise temporal alignment between image content and probe posture, which is the data foundation for achieving distortion source separation.

[0022] In some embodiments, the acoustic coupling quality index is determined based on the superficial region echo information, deep region echo information, echo attenuation consistency between the superficial and deep regions, and image information content in the cardiac ultrasound image frame. The acoustic coupling quality index is configured to decrease when there is insufficient effective echo in the superficial region, insufficient effective echo in the deep region, abnormal echo attenuation relationship between the superficial and deep regions, or image information content is lower than a preset requirement. Furthermore, the decrease in the acoustic coupling quality index increases the section quality deviation index or makes the current frame more likely to be classified as a transitional or distorted state. This index gates the image validity at the physical acoustic level, preventing low signal-to-noise ratio or invalid frames from entering the subsequent prediction model, thereby preventing the algorithm from generating false compensation.

[0023] In some embodiments, the explicit anatomical quality features include at least two of the following: left ventricular contour deviation features, used to characterize the degree of deviation of the left ventricular endocardial contour from the standard sectional morphology; atrioventricular junction position deviation features, used to characterize the degree of offset of the atrioventricular junction region from a preset reference position; and valve motion orderliness features, used to characterize the consistency or dispersion of valve region motion in consecutive echocardiogram image frames; wherein, the explicit anatomical quality features are used to constrain whether the current frame conforms to the anatomical morphological requirements of the target cross-section of echocardiography. These features have clear clinical physical significance, can still provide morphological constraints when image texture is blurred, and can cross-validate distortion attribution with probe motion data.

[0024] In some embodiments, implicit spatiotemporal fusion features are generated based on the cardiac ultrasound image frames and the probe motion state data, including:

[0025] Image encoding is performed on the cardiac ultrasound image frames to obtain image representation features;

[0026] Motion encoding is performed on the probe motion state data to obtain motion characterization features;

[0027] The adaptive fusion weights are determined based on the image representation features and the motion representation features;

[0028] The image representation features and the motion representation features are fused based on the adaptive fusion weights to generate the implicit spatiotemporal fusion features;

[0029] The adaptive fusion weights are used to adjust the contribution ratio of image information and probe motion information to the implicit spatiotemporal fusion features. This adaptive gating mechanism dynamically adjusts the fusion weights based on the current image signal-to-noise ratio and the intensity of motion: when the probe shakes violently or the image noise is high, the image feature weights are automatically reduced, and motion data is used instead to maintain feature continuity; conversely, anatomical details are enhanced, thereby improving the anti-interference ability of the features.

[0030] In some embodiments, obtaining the slice quality deviation index of the current frame includes:

[0031] The temporal prediction bias is obtained based on the difference between the implicit spatiotemporal fusion features at the current moment and the predicted spatiotemporal features.

[0032] Based on the aforementioned explicit anatomical quality characteristics, anatomical morphological deviations are obtained;

[0033] Based on the acoustic coupling quality index, the acoustic quality modulation factor is obtained;

[0034] The cross-sectional quality deviation index is obtained based on the time-series prediction deviation, anatomical morphology deviation, and acoustic quality modulation factor.

[0035] The acoustic quality modulation factor is used to enhance the impact of timing prediction bias and / or anatomical deviation on the cross-sectional quality deviation index when acoustic coupling quality deteriorates. This nonlinear coupling function achieves a triple guarantee of "physical gating + data-driven bias + rule-based penalty": the acoustic quality factor acts as a global gating factor, exponentially amplifying the bias when probe slippage or coupling failure occurs, causing the system to immediately enter a distorted state; while under normal coupling, the bias is mainly determined by timing and anatomical morphology, thus balancing sensitivity and robustness.

[0036] In some embodiments, dynamically determining the stability determination boundary and the distortion determination boundary includes:

[0037] Within the preset time window, the central trend information and dispersion information of the cross-sectional quality deviation index of multiple consecutive frames are statistically analyzed.

[0038] The stability determination boundary and the distortion determination boundary are determined based on the central trend information and the dispersion information, wherein the stability determination boundary is lower than the distortion determination boundary.

[0039] If the cross-sectional quality deviation index of the current frame is lower than the stability determination boundary, the current frame is determined to be in a stable state.

[0040] If the cross-sectional quality deviation index of the current frame is between the stability determination boundary and the distortion determination boundary, the current frame is determined to be in a transition state.

[0041] If the cross-sectional quality deviation index of the current frame is higher than the distortion judgment boundary, the current frame is judged as distorted. This dynamic boundary mechanism abandons the fixed threshold and uses a sliding window to statistically analyze the mean and standard deviation of the deviation index in real time, so that the judgment boundary adaptively drifts with the current patient's heart rate, tissue acoustic characteristics and operator habits, thereby solving the problem of misjudgment caused by cross-individual baseline inconsistency.

[0042] In some embodiments, if the current frame is determined to be in a transitional state, the observation uncertainty of the current frame is determined based on the probe motion state data and the image quality information of the cardiac ultrasound image frame, including:

[0043] The degree of motion disturbance is determined based on the motion amplitude of the probe motion state data.

[0044] The image reliability is determined based on at least one of the following: grayscale distribution, texture sharpness, local contrast, or image variance of the cardiac ultrasound image frame.

[0045] The observation uncertainty is determined based on the degree of motion disturbance and the image reliability.

[0046] Specifically, when the degree of motion perturbation increases or the image confidence decreases, the observation uncertainty increases, making the compensation information more dependent on the predicted spatiotemporal features; conversely, when the degree of motion perturbation decreases or the image confidence increases, the observation uncertainty decreases, making the compensation information more dependent on the implicit spatiotemporal fusion features of the current moment. This mechanism employs a Kalman-like gain dynamic tradeoff between the confidence of predicted and observed values: under high noise or strong jitter, the compensation information relies more on historical pattern predictions; when the image is clear, it fully adopts the current observation features, thereby achieving adaptive repair.

[0047] In some embodiments, if the current frame is determined to be in a distorted state, the method further includes distortion attribution processing:

[0048] When the probe motion status data meets the preset motion abnormality conditions, the current distortion is attributed to probe slippage distortion, and a probe position adjustment prompt is output.

[0049] When the probe motion state data does not meet the preset motion anomaly conditions and the acoustic coupling quality index is lower than the preset acoustic quality requirements, the current distortion is attributed to acoustic coupling distortion or sound shadow occlusion distortion, and a delay waiting mechanism is activated.

[0050] When subsequent frames return to a stable or transitional state within the preset waiting time corresponding to the delay waiting mechanism, recording resumes;

[0051] If subsequent frames fail to recover to a stable or transitional state within the preset waiting time, the current video file is truncated or a new video file is generated. This attribution logic automatically distinguishes between three types of distortion—probe slippage, acoustic occlusion, and depth attenuation—by analyzing motion data, acoustic quality indicators, and image grayscale features, and provides targeted recovery strategies for each, thereby improving the system's clinical interpretability and operational guidance value.

[0052] This embodiment also provides an automatic recording system for cardiac ultrasound with distortion resistance. The system includes an ultrasound probe, a probe motion sensor, and a video storage unit, comprising:

[0053] The multimodal acquisition module is used to acquire continuous cardiac ultrasound image frames, probe motion state data synchronized with the cardiac ultrasound image frames, and obtain acoustic coupling quality index and cardiac cycle phase information based on the cardiac ultrasound image frames.

[0054] The feature analysis module is used to extract explicit anatomical quality features based on the cardiac ultrasound image frames and generate implicit spatiotemporal fusion features based on the cardiac ultrasound image frames and the probe motion state data.

[0055] The prediction and evaluation module is used to generate the predicted spatiotemporal features of the current moment based on the implicit spatiotemporal fusion features of historical moments, and to obtain the cross-sectional quality deviation index of the current frame based on the predicted spatiotemporal features, the implicit spatiotemporal fusion features of the current moment, the explicit anatomical quality features, and the acoustic coupling quality index.

[0056] The state determination module is used to dynamically determine the stability determination boundary and the distortion determination boundary based on the cross-sectional quality deviation index of multiple consecutive frames within a preset time window, and to determine the current frame as a stable state, a transitional state or a distorted state based on the stability determination boundary and the distortion determination boundary.

[0057] An adaptive compensation module is used to determine the observation uncertainty of the current frame based on the probe motion state data and the image quality information of the cardiac ultrasound image frame when the current frame is determined to be in a transitional state, and to adaptively fuse the predicted spatiotemporal features and the implicit spatiotemporal fusion features at the current moment according to the observation uncertainty, so as to generate compensation information for the current frame.

[0058] The recording execution module is used to execute recording writing when the current frame is determined to be in a stable state and the cardiac cycle phase information meets the preset recording trigger conditions; to execute recording writing after compensating the current frame based on the compensation information when the current frame is determined to be in a transitional state; and to pause or stop recording writing and output distortion prompt information when the current frame is determined to be in a distorted state.

[0059] The multimodal acquisition module is communicatively connected to the ultrasound probe and probe motion sensor. The input of the feature analysis module is connected to the output of the multimodal acquisition module. The input of the prediction and evaluation module is connected to the output of the feature analysis module. The input of the state determination module is connected to the output of the prediction and evaluation module. The input of the adaptive compensation module is connected to the state determination module, the multimodal acquisition module, and the feature analysis module, respectively. The input of the recording execution module is connected to the state determination module and the adaptive compensation module, respectively. The output of the recording execution module is connected to the recording storage unit. This system, through clear module division and a hard real-time data path, achieves closed-loop control of the entire chain from acquisition, analysis, prediction, determination, compensation, to recording, with an end-to-end latency of less than 20ms, meeting the clinical requirements for real-time cardiac ultrasound imaging.

[0060] Compared with the prior art, the beneficial effects of the present invention include at least the following:

[0061] 1. Precise decoupling of distortion sources: By fusing synchronous data from ultrasound images and inertial measurement units (IMUs), explicit anatomical features and implicit spatiotemporal features are separated into two channels, effectively distinguishing probe jitter from the heart's own motion and avoiding misjudgment.

[0062] 2. Dynamic adaptive judgment: Based on nonlinear coupled energy function and sliding window statistics, the stable / distortion judgment boundary is dynamically generated without the need for a fixed threshold, adapting to different patients and operating habits, and combining sensitivity and robustness.

[0063] 3. Transitional intelligent repair: When there is slight distortion, a residual calibration mechanism based on observation uncertainty (Kalman-like gain) is introduced to compensate the image before writing it, which significantly improves the continuity of recording and the effective diagnostic time, and avoids simply discarding frames.

[0064] 4. Closed-loop continuous optimization: Successfully recorded stable state segments are used as positive samples to fine-tune the prediction model online, achieving personalized adaptation that becomes more accurate with use.

[0065] 5. Standardized clinical operation: Automatically completes distortion attribution (slippage, occlusion, attenuation), phase-locked recording, and interactive prompts, reducing differences in operator experience and ensuring the uniformity of diagnostic recording quality. Attached Figure Description

[0066] Figure 1 This is a flowchart of the present invention;

[0067] Figure 2 This is a system diagram of the present invention. Detailed Implementation

[0068] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided to make the invention more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0069] Example 1

[0070] like Figure 2 As shown, the cardiac ultrasound anti-distortion automatic recording system of the present invention includes a multimodal sensing module, a feature analysis module, a prediction and evaluation module, a state determination module, an adaptive compensation module, and a recording execution module. Each module is coupled and connected via a high-speed data bus, shared memory, and atomic operation mechanisms, forming a hard real-time processing link with an end-to-end latency of less than 20ms. This architecture ensures that every step from image acquisition to video recording meets the latency constraints of real-time cardiac ultrasound imaging, and each module has independent fault tolerance capabilities. The hardware carrier, functional configuration, and data flow of each module are as follows:

[0071] (1) Multimodal sensing module

[0072] Deployed at the front end of the ultrasound probe and in the FPGA synchronization unit, it is configured to perform multi-source data synchronous acquisition, timestamp alignment, and acoustic coupling quality index pre-calculation in step S100. Specifically, it includes:

[0073] A phased array ultrasonic probe with integrated six-degree-of-freedom MEMS IMU is used to acquire triaxial acceleration and angular velocity data in real time (sampling rate ≥100Hz).

[0074] Beam synthesizer for outputting B-mode cardiac ultrasound image sequences (30-60fps).

[0075] Physiological signal acquisition interface, used to connect to an external ECG module to acquire electrocardiogram signals;

[0076] The FPGA hardware timestamp generation unit is used to perform microsecond-level synchronization and alignment of image frames and motion data, and encapsulate them into atomic data packets. Write to the circular timing buffer queue.

[0077] The module communicates with the feature analysis module through the PCIe Gen4 interface. Data packets are transmitted via zero-copy through a shared memory pointer, ensuring that the acquisition-processing link latency is ≤3ms.

[0078] (2) Feature Analysis Module

[0079] Running on an embedded GPU computing cluster (such as NVIDIA Jetson AGX Orin), configured to perform explicit and implicit dual-channel feature extraction and joint embedding in step S200. Specifically, this includes:

[0080] The explicit anatomical feature extraction submodule extracts the left ventricular contour, atrioventricular junction, and valve region mask based on the U-Net semantic segmentation network (ResNet-34 backbone, TensorRTFP16 deployment), and calculates the left ventricular contour deviation features. Characteristics of the deviation of the atrioventricular junction With the orderly characteristics of valve movement Weighted aggregation into a score for overt anatomical abnormalities ;

[0081] Implicit Spatiotemporal Fusion Feature Extraction Submodule: Extracts image representations using a lightweight CNN (a cropped version of MobileNetV3) and a fully connected motion encoder, respectively. With motion representation Learnable gating networks Adaptive fusion as implicit spatiotemporal fusion features .

[0082] The module performs dual-channel extraction in parallel via CUDA streaming, with a single-frame inference latency of ≤10ms. The output feature vector is directly fed into the prediction and evaluation module via shared memory.

[0083] (3) Prediction and evaluation module

[0084] Based on a pre-built LSTM inference engine and a dedicated GPU computing core, it is configured to perform autoregressive prediction and nonlinear coupled energy calculation in step S300. Specifically, this includes:

[0085] Autoregressive prediction submodule: This module converts historical V... The sequence is input into a single-layer LSTM network (128-dimensional hidden layers, TensorRTFP16 engine), and the output is the predicted feature at the current time step. and predicted covariance ;

[0086] Energy Coupling Calculation Submodule: According to the formula:

[0087] Real-time calculation of cross-sectional quality deviation index, among which =1.5、 =0.7、 =0.3、 =1.0, and with embedded `clip` logic constraints. Prevent overflow.

[0088] The module is bound to an independent GPU Stream and executed in parallel with the feature extraction pipeline, with an end-to-end computation latency of ≤8ms.

[0089] (4) Status determination module

[0090] Integrated into the state decision coprocessor, it is configured to execute the sliding window statistics, dynamic threshold generation, and three-state transition determination in step S300. Specifically, it includes:

[0091] The sliding window statistics submodule maintains a circular deviation queue with a capacity of N = 50~100 frames, and updates the mean in real time through atomic write operations. with standard deviation ;

[0092] Dynamic threshold generation submodule: By , ( =1.0、 =3.0) Generate stable / distortion decision boundaries;

[0093] The state transition and instruction issuance submodule will: Compared with the two-boundary three-segment method, the output discrete state label Simultaneously, it sends hard real-time control commands to downstream modules:

[0094] ( ).

[0095] The module sends instructions to the recording execution module via the AXI bus, ensuring that the decision-execution latency is less than 1ms.

[0096] (5) Adaptive compensation module

[0097] Deployed on a standalone GPU Stream, configured to perform the observation uncertainty assessment, dynamic gain calculation, and feature-level residual calibration of step S400. Specifically, this includes:

[0098] Observation covariance construction submodule: by Calculate the observation noise covariance scalar ( = =0.5);

[0099] Dynamic gain calculation submodule: based on Kalman-like gain formula Calculate the fusion weights and embed saturation limiting constraints. [0.05, 0.95];

[0100] Residual calibration and parametric regression submodule: Execution It outputs affine / grayscale compensation parameters through a lightweight fully connected regression head. The original frame is processed by the GPU CUDA Image Warp and LUT hardware unit. Perform subpixel-level correction.

[0101] The module has an end-to-end latency of ≤5ms and an embedded SSIM / PSNR quality gate. If SSIM < 0.92, the original frame is automatically rolled back and marked as `Compensation_Reverted`.

[0102] (6) Video recording execution module

[0103] The NVMe storage controller is coupled with the DICOM packaging engine and configured to perform phase-locked write, transition-state compensated frame serialization, distortion blocking, and attribution interaction in step S500. Specifically, this includes:

[0104] Phase-locked write submodule: only when Falling into the end-diastolic trigger window ( When [0.85, 1.0]) and the state is stable, backtrack the circular queue K=20 frames of historical data, splice them in chronological order and inject them into the DICOM standard header file, and asynchronously write them to high-speed solid-state storage through a double buffering mechanism;

[0105] Transitional state serialization submodule: Directly writes the compensated frame I'_t, and inserts the `Frame_Type=Compensated` flag when there are more than 15 consecutive transitional states;

[0106] Distortion blocking and homing factor module: Upon receiving the `RECORD_HALT` command, immediately close the file handle and execute `fsync` to force a disk flush, based on | |、 The grayscale mean is used to perform a three-level attribution determination, and tactile / visual alarms are output through the interactive terminal and a diagnostic log is generated.

[0107] The module adopts a producer-consumer model and a circular queue lifecycle management, with constant memory usage, eliminating fragmentation during long-term operation.

[0108] Inter-module communication and timing constraints

[0109] The above six modules achieve low-latency, highly reliable data flow and control synchronization through the following mechanisms:

[0110] Data path: Multimodal perception module → Feature analysis module → Prediction and evaluation module → State determination module → Adaptive compensation module / recording execution module. The entire process uses shared memory pointers for zero-copy transmission to avoid data transfer overhead between the host and the device.

[0111] Control path: The status determination module sends hard real-time commands to the recording execution module through the AXI bus to ensure that the decision-execution latency is <1ms;

[0112] Timing constraints: Acquisition synchronization ≤3ms, feature extraction ≤10ms, energy calculation ≤8ms, residual calibration ≤5ms, video recording and writing are asynchronous and do not block the main loop, and the end-to-end processing delay is strictly controlled within 20ms to meet the clinical needs of real-time cardiac ultrasound imaging.

[0113] Fault tolerance: Each module is embedded with timeout detection, numerical overflow protection and degradation rollback strategies (such as IMU disconnection → gating clamping, inference timeout → zero-order hold, compensation over-limit → original frame rollback) to ensure medical device-level operational robustness.

[0114] like Figure 1 As shown, a method for automatic recording of cardiac ultrasound with distortion resistance includes the following steps:

[0115] Step S100: Synchronous encapsulation and acoustic quality pre-calculation of multi-source heterogeneous data

[0116] This step constructs time-aligned atomic input data packets and performs physical-level acoustic contact quality prediction, providing standardized, low-latency data stream input for subsequent feature extraction and state decision-making.

[0117] (1) Hardware-level timestamp synchronization and atomic data packet construction

[0118] Data Source and Preprocessing Link: The six-degree-of-freedom MEMS IMU built into the ultrasonic probe outputs raw triaxial acceleration and triaxial angular velocity signals in real time at a sampling rate of ≥100Hz. After processing by a first-order low-pass digital filter (cutoff frequency 30Hz, used to filter out high-frequency electrical noise and high-frequency mechanical resonance in the probe circuit), the probe motion state data is formed. The beam synthesizer outputs a B-mode grayscale image sequence at 30-60 fps, which is then logarithmically compressed and dynamically adjusted (DR=40-60dB) to form a cardiac ultrasound image frame I_t.

[0119] Asynchronous Sampling Alignment Mechanism and Engineering Implementation: Addressing the inherent asynchronous nature of IMU high-frequency sampling and low-frequency image frame rates, the system generates a unified timestamp at the microsecond level via the FPGA hardware clock. Within each image exposure window, the system performs integral smoothing on the IMU sampling points over that window's time span, extracting the equivalent motion vector at the midpoint of the exposure to achieve precise spatiotemporal alignment between physical motion and image anatomical content. If the FPGA hardware synchronization link malfunctions (e.g., clock drift exceeds limits or communication is interrupted), the processing unit automatically downgrades to a software timestamp alignment mechanism based on linear interpolation, compensating for delays strictly controlled within 5ms to ensure uninterrupted data link operation.

[0120] Cardiac cycle phase analysis strategy: If the system detects an external ECG signal, it filters out baseline drift using a high-pass filter (cutoff frequency 0.5Hz), then employs an improved Pan-Tompkins algorithm to detect the R-wave peak value, calculates the normalized position of the current time within the adjacent RR interval, and outputs the phase. [0, 1); If there is no ECG signal, the system performs autocorrelation analysis on the segmented area sequence of the left ventricular cavity in continuous image frames, extracts the dominant frequency period through Fast Fourier Transform (FFT), and infers the current cardiac phase based on this. Phase estimation error tolerance ≤ ±10%.

[0121] Buffer scheduling and output: aligning synchronization Encapsulated as atomic data packets The data is written to a circular timing buffer queue with a capacity of a preset number of frames (preferably 100 frames). This queue is managed by a circular pointer, and new data packets overwrite the oldest data packets according to the writing sequence. This provides short-term memory support for the sliding window statistics in step S300 and the residual calibration in step S400, enabling continuous stream processing with fixed memory usage and avoiding memory fragmentation during system operation.

[0122] (2) Acoustic coupling quality index Computation and Gating Logic

[0123] Region definition and feature extraction: in image frames In the pixel coordinate system, a shallow near-field region R with a preset fixed depth ratio is defined. (Located in the top 0-20% depth range of the image) and the deep far-field region (Located in the bottom 60-80% depth range of the image). The average grayscale value of the two regions is calculated in real time. and The image grayscale level is standardized to L=256.

[0124] Parameter adaptive calibration: reference threshold , and reference attenuation ratio Automatic calibration is performed during each power-on initialization phase of the device: The system prompts the operator to stably attach the probe to the standard phantom or the patient's chest wall, and acquires 30 consecutive frames of images under stable contact conditions. The average gray values ​​and their ratios of near and far fields are calculated respectively, and a ±10% engineering tolerance is introduced to generate calibration parameters to adaptively compensate for differences in subcutaneous fat thickness, probe aging degree, and device gain settings among different patients.

[0125] The acoustic coupling quality index is calculated in real time according to the following nonlinear coupling formula:

[0126] To prevent the algorithm from generating hallucinations when the probe is not in contact with the skin (air coupling), when there is insufficient coupling agent, or when the ribs completely block the view, the system pre-calculates the acoustic coupling quality index. , is used to characterize the effective acoustic contact level of the current ultrasound image. Its value is strictly normalized and restricted to the interval (0,1] to serve as the physical gating factor for the subsequent energy function.

[0127] In a preferred embodiment, the acoustic coupling quality index is jointly determined by the near-field effective echo term, the far-field effective echo term, the depth attenuation consistency term, and the image information entropy term, and its calculation formula is as follows:

[0128]

[0129] The specific calculation logic for the above items is as follows:

[0130] Near-field effective echo term : Used to characterize whether effective acoustic contact exists in shallow areas, formula:

[0131] Far-field effective echo term : Used to characterize the presence of effective echoes in deep regions (to rule out severe acoustic shadowing), the formula is:

[0132]

[0133] Depth attenuation consistency term A(t): Used to characterize whether the relationship between near-field and far-field echo attenuation conforms to normal physical acoustic propagation laws, the formula is:

[0134]

[0135] Normalized image information entropy : Used to characterize the effective information content of the current frame, the formula is:

[0136] .

[0137] Parameter description and physical meaning:

[0138] This refers to the superficial area of ​​the ultrasound image. This refers to the deep region of an ultrasound image. This represents the average gray value of the corresponding area.

[0139] and These are the effective echo intensity reference thresholds for the shallow and deep regions, respectively; R_0 is the reference value for the near-field to far-field echo intensity ratio under normal acoustic coupling conditions. These three parameters can be obtained by statistically averaging multiple consecutive frames (e.g., 30 frames) of steady-state ultrasound images acquired during the system initialization phase, thus enabling adaptive environmental calibration of the equipment.

[0140] To prevent extremely small positive numbers from being divided by zero (e.g., taking...) This is used to prevent division by zero anomalies when the grayscale mean of a region is zero.

[0141] To prevent the exponent from reaching zero, a lower cutoff value is set (e.g., 0.05) to ensure the stability of subsequent calculations. =0.05 is the lower limit for truncation of the product term to prevent a single extreme value from causing the overall index to drop to zero and triggering subsequent numerical overflow;

[0142] This is the decay consistency penalty coefficient (e.g., 1.0). =1.0 is the decay consistency penalty coefficient. The output is strictly limited to the interval (0, 1], and is a dimensionless probabilistic index.

[0143] L is the number of gray levels in the image (for a standard 8-bit ultrasound image, L=256 is usually used).

[0144] When the probe detaches from the skin, the coupling layer breaks, strong reflections from the ribs block the light, or there is abnormal attenuation of the far-field echo occurs during clinical operation. C Or A(t) will drop rapidly, forcing the acoustic coupling quality index to decrease. This will decrease. This will directly lead to a decrease in the acoustic gating term in the nonlinear coupling energy function. It increases exponentially, thereby forcibly amplifying the energy deviation of the system when irreversible distortion occurs at the physical level, greatly improving the sensitivity and robustness of transition state and distortion state determination.

[0145] This is directly used as the global physical gating factor for calculating the section quality deviation index in subsequent step S300. When probe slippage, coupling agent breakage, strong rib reflection obstruction, or abnormal far-field attenuation occurs, a decrease in any of these factors will lead to… The decrease is product-like. This decreasing trend is directly mapped to the nonlinear amplification of the energy gating term in the algorithm control link, establishing a hard-coupled feedback of "physical contact degradation → dynamic tightening of the decision boundary", blocking the flow of low signal-to-noise ratio or non-anatomical information frames into the subsequent prediction model from the data source, and avoiding the algorithm from generating false structural illusions.

[0146] Step S200: Explicit and Latent Dual-Channel Feature Extraction and Joint Embedding

[0147] This step loads the atomic data packet output from step S100 into the GPU shared memory of the processing unit, and performs the extraction of explicit anatomical quality features and implicit spatiotemporal fusion features through a parallel computing pipeline. The dual-channel design aims to decouple clinically interpretable geometric constraints from data-driven spatiotemporal pattern representations, providing structured, low-latency, and physically meaningful feature inputs for subsequent nonlinear energy calculations.

[0148] (1) Channel A: Extraction of explicit anatomical quality features (rule and geometry driven path)

[0149] Image preprocessing and segmentation inference deployment: Frames of cardiac ultrasound images The input is a pre-configured U-Net semantic segmentation network. To suppress texture aliasing caused by ultrasonic speckle noise and high-frequency jitter, block-matching 3D filtering (BM3D, matching block size 8*8, search window 21*21, noise variance based on local gradient adaptive estimation) is performed on I_t before inference. The U-Net network uses ResNet-34 as the encoder backbone, compiled into an FP16 precision inference engine by TensorRT, and deployed on an embedded GPU computing cluster, with the forward propagation delay of a single frame strictly controlled within 8ms. The network outputs a binary mask of the left ventricular endocardium, semantic labels of the atrioventricular junction region, and a mask of the valve region of interest (ROI). If the Dice confidence of a single frame inference is lower than 0.75 or the computation times out (>12ms), the system automatically enables a temporal smoothing mechanism, reuses the dominant feature vector of the previous frame and marks it as a low-confidence state to prevent feature mutations from causing divergence in subsequent energy calculations.

[0150] Left ventricular contour deviation characteristics Extraction: Based on the binary mask of the left ventricular endocardium, a set of discrete pixels at the edge of the endocardium is extracted. Algebraic least squares is used to fit the point set to a standard elliptical geometric model. The normalized Hausdorff distance from each pixel in the point set to the boundary of the fitted ellipse is calculated and divided by the length of the image diagonal pixels to achieve spatial dimensionlessness. The output is... The physical mapping logic is as follows: based on the anatomical prior of the standard apical section, the left ventricular cavity presents a stable, approximately elliptical shape on the ideal imaging plane; when the probe slips out of plane or the section deviates significantly from the standard anatomical axis, the endocardial contour will exhibit irregular distortion or truncation, leading to a significant increase in the fitting residual and the Hausdorff distance. This feature directly quantifies the standardization of the geometry of the current frame, providing a rigid anatomical constraint for subsequent steps.

[0151] Atrioventricular junction position deviation characteristics Extraction: Locate the spatial coordinates of the attachment points of the anterior and posterior mitral leaflets in the semantic tags, and calculate the geometric midpoint of the line connecting the two attachment points. Using the preset anatomical center of the image (or the boundary reference coordinates when the previous frame was in a stable state) as the reference origin, calculate the Euclidean distance between the midpoint and the reference origin, and output it after image diagonal normalization. The physical mapping logic is as follows: in a standardized cardiac section, the atrioventricular junction should be located near the central axis of symmetry of the image; if the probe undergoes lateral translation or tilting scan, the midpoint of the junction will experience a systematic shift. This feature is related to the probe motion state data acquired in step S100. Forming cross-validation: when Continue to rise and When unidirectional angular velocity accumulation is observed, it can be accurately attributed to probe misalignment rather than deformation of the heart itself.

[0152] orderly characteristics of valve movement Extraction: Within the valve ROI mask, a dense optical flow field is calculated based on three consecutive frames (using the Farneback algorithm with three Gaussian pyramid layers). The Shannon entropy values ​​for all vector directions within this optical flow field are statistically analyzed, and after logarithmic normalization, the output is... The physical mapping logic is as follows: During a normal physiological cycle, the opening and closing motion of the valves has a clear direction and periodicity, the optical flow vector distribution is concentrated, and the entropy value is low; when the probe jitters momentarily or acoustic blockage causes tissue overlap, the local pixel motion exhibits disordered diffusion, the dispersion of the optical flow vector direction increases sharply, and the entropy value increases accordingly. This feature effectively captures the dynamic structural blur caused by operational micro-jitter, compensating for the lag in the response of static geometric features to instantaneous distortion.

[0153] Feature aggregation and output: The above three features are aggregated according to their clinical diagnostic weights to obtain the overt anatomical abnormality score. (Preferred) =0.4, =0.4, =0.2, which satisfies =1). The final output is the dominant anatomical quality feature vector:

[0154] To the characteristic buffer pool, It will be directly fed into step S300 as a calculation primitive for the anatomical morphology deviation term.

[0155] (2) Channel B: Implicit Spatiotemporal Fusion Feature Extraction (Data-Driven and Adaptive Gated Circuit)

[0156] Heterogeneous coding and feature subspace alignment: normalized Compared with probe motion state data after first-order low-pass filtering (cutoff frequency 30Hz) The data is fed into a lightweight coding network in parallel. The image branch uses a cropped version of MobileNetV3 to extract deep texture, echo attenuation patterns, and contextual representations, outputting image representation features. The motion branch maps six-DOF accelerations and angular velocities to a latent space of the same dimension through a three-layer fully connected network (layerNorms are embedded layer by layer), outputting motion representation features. This subspace alignment strategy eliminates the heterogeneity between pixel grayscale values ​​and physical motion dimensions, enabling strict matching of cross-modal features in the mathematical dimension, and providing a computable scalar basis for subsequent weighted fusion.

[0157] The computational logic of the adaptive gating fusion mechanism: To dynamically balance the contribution ratio of image anatomical information and probe physical motion information, a learnable gating network is introduced to calculate adaptive fusion weights. and After splicing along the channel dimension, input to the gating layer:

[0158]

[0159] in It is the Sigmoid activation function. and These are weights pre-trained offline. Based on... Perform element-wise weighted fusion to generate implicit spatiotemporal fusion features:

[0160]

[0161] The engineering significance of this gating mechanism lies in the fact that when the probe experiences severe shaking or the image signal-to-noise ratio drops sharply, the network automatically reduces... The value suppresses the interference of image texture noise, and instead relies on the high-frequency motion trajectory of the IMU to maintain feature continuity; when the probe is held stably and the image is clear, Approaching 1 enhances the representation of anatomical details. This dynamic weight adjustment makes latent features robust against interference and does not depend on explicit rule thresholds.

[0162] Engineering Deployment and Hardware-Level Fault Tolerance Strategy: This fusion network employs a self-supervised contrastive learning strategy during the training phase (based on public datasets such as EchoNet-Dynamic), and only performs forward propagation during the inference phase, without introducing gradient computation overhead. When an IMU hardware communication interruption or data stream verification failure is detected, the system uses a low-level watchdog signal to... Forced clamping to 1 smoothly degenerates to a pure image feature extraction mode, ensuring uninterrupted data transmission. The final output is a normalized implicit spatiotemporal fusion feature vector. The data is then written into the time-series buffer queue for use by the autoregressive model in step S300.

[0163] (3) Mapping of dual-channel output to downstream control link

[0164] After this step is completed, the system synchronously outputs the explicit anatomical quality feature vector. (including) Features of implicit spatiotemporal fusion Both are transmitted to the computational core in step S300 via independent data paths: It directly participates in the calculation of the anatomical penalty term of the nonlinear coupled energy function, providing interpretable morphological constraints; The sequence is then input into the autoregressive prediction model to construct the time series prediction baseline and uncertainty estimation. The dual-channel separation architecture ensures that the algorithm complements the "rule verification" and "data prediction" processes, avoiding misjudgments caused by single-mode failures and providing high-confidence input for subsequent three-state transitions and compensation control.

[0165] The essence of dual-channel design is to decouple interpretable clinical anatomical rules (explicit features) from data-driven spatiotemporal patterns (latent features), so that subsequent distortion judgment has both rule constraints and predictive flexibility, reducing the risk of single-modality failure.

[0166] Step S300: Nonlinear Coupled Energy Calculation and Dynamic Threshold Generation

[0167] This step, based on the dual-channel feature vector output from step S200 and the acoustic quality index from step S100, constructs a cross-sectional quality deviation index in real time on an independent GPU computing core of a heterogeneous computing platform. It also dynamically generates a decision boundary through a sliding statistical window, completing a hard real-time mapping from continuous feature stream to discrete state labels, and providing deterministic decision instructions for the recording control and compensation module.

[0168] Autoregressive inference and confidence estimation for predicting spatiotemporal features

[0169] Data scheduling and queue reading: From the time-series buffer queue written in step S200, the implicit spatiotemporal fusion feature sequences of the most recent L frames (preferably L=10) are continuously read using CUDA zero-copy memory mapping technology with a circular pointer. This read operation is strictly synchronized with the image acquisition frame rate to avoid control cycle mismatch caused by data transfer between the host and device.

[0170] Model Deployment and Inference Execution: The feature sequences are input into a pre-built Long Short-Term Memory (LSTM) network. This network has 128 hidden layers and is compiled into an FP16 precision Engine file using TensorRT, then deployed on an embedded GPU computing cluster. The network employs a single-layer autoregressive structure, with the forward propagation time per frame strictly controlled to within 8ms, ensuring that the control period (30-60Hz) perfectly matches the image frame rate.

[0171] Prediction uncertainty approximation: The spatiotemporal characteristics of the prediction at the current time are output. Simultaneously, during the inference phase, the system enables Monte Carlo Dropout (Dropout rate = 0.1, performing 10 independent forward propagations), statistically analyzes the variance distribution of the output features, and approximates the prediction covariance matrix. The covariance matrix represents the model's confidence in the current time-series baseline and is directly used as the prior input for the dynamic calibration gain calculation in step S400.

[0172] Anomaly detection and fallback: If GPU inference times out (>12ms) or NaN / Inf anomalies appear in the feature sequence, the processing unit immediately triggers the zero-order hold (ZOH) backup strategy, reusing the predicted features from the previous stable frame. and will The prediction is forcibly amplified to a preset upper limit (representing a low-confidence state). This mechanism prevents abnormal prediction values ​​from contaminating subsequent energy calculation links and ensures uninterrupted system real-time performance.

[0173] Nonlinear coupling calculation of cross-sectional quality deviation index

[0174] Component construction and physical mapping: The system calculates three independent components sequentially according to the following control logic, and aggregates them into the current frame's slice quality deviation index through a nonlinear coupling function. :

[0175] 1. Time series forecast bias : Calculate the current implicit spatiotemporal fusion features With predicting spatiotemporal features The square of the L2 norm between them, i.e. This value quantifies the temporal continuity of the characteristics. When the probe experiences slight jitter or rapid scanning, causing instantaneous motion abrupt changes, the prediction residual increases significantly.

[0176] 2. Anatomical morphological deviation : Based on overt anatomical abnormality scoring Calculate the exponential penalty term, i.e. , where tau=1.0. The exponential mapping achieves non-linear amplification of geometric distortion. When the cross-section deviates significantly from the standard anatomical structure, this dominant deviation index ensures "zero tolerance" for non-standard cross-sections.

[0177] 3. Acoustic quality modulation factor Based on acoustic coupling quality index Calculate the global gating term, i.e. ,in [1.0, 2.0] (preferably 1.5). This factor establishes a hard coupling relationship between physical contact quality and algorithm sensitivity: when the probe detaches from the skin, the coupling agent breaks, or there is acoustic shadowing, Sudden drop, It increases sharply inversely, forcibly raising the overall deviation index.

[0178] Coupled function execution and numerical stability: The cross-sectional quality deviation index is calculated in real time according to the following formula:

[0179]

[0180] in =0.7, =0.3, which satisfies + =1.0. All input features have been normalized in step S200. The output is a dimensionless non-negative scalar. To prevent the exponent from causing GPU floating-point overflow under extreme distortion conditions, hardware-level clip logic is embedded within the computation core to limit... The maximum value is 104. This computation task is bound to a separate GPU Stream and executed in parallel with the feature extraction pipeline, with an end-to-end computation latency of ≤10ms.

[0181] Dynamic Boundary Generation and State Transition Control

[0182] Sliding window statistics and baseline maintenance: The system maintains a circular deviation queue with a capacity of N frames (preferably N=50~100, which can be adaptively adjusted according to real-time heart rate) in shared memory. Each time a new baseline is calculated... This involves pushing data into the queue via an atomic write operation, overwriting the oldest data. Based on all valid values ​​in the queue, the central trend information (mean) is calculated in real time. Information on dispersion (standard deviation) ).

[0183] Dynamic generation of decision boundaries: Stable decision boundaries are dynamically generated based on the statistical distribution characteristics of the current operating conditions. Distortion Judgment Boundary :

[0184]

[0185] The threshold coefficient satisfies 0 < < Preferred =1.0, =3.0. This generation logic abandons fixed empirical thresholds, allowing the judgment boundary to adaptively drift according to the current patient's cardiac rhythm, subcutaneous tissue acoustic characteristics, and operator's grip habits, fundamentally solving the problem of misjudgment caused by inconsistencies in cross-individual baselines.

[0186] Status determination and control command issuance: [This will determine the current status and control command issuance.] Perform a three-stage comparison with the double boundary and output discrete state labels. {Stable state, Transitional state, Distortion state}, and simultaneously send hard real-time control commands to downstream modules:

[0187] like Once the state is determined to be stable, a `RECORD_ENABLE` command is sent to the recording execution module;

[0188] like If the condition is determined to be in a transitional state, a `COMPENSATE_TRIGGER` command is sent to the adaptive compensation module, along with the current condition. and ;

[0189] like If the system detects a distortion state, it sends a `RECORD_HALT` command to the recording execution module and triggers the distortion factorization program.

[0190] Baseline drift reset mechanism: If M consecutive frames (preferably M=20) If the value exceeds three times the initial calibration value, it is determined that the data distribution drift is caused by switching the inspection object or large-scale probe movement. The system automatically clears the sliding window queue, discards historical statistics, and uses the current value as the basis for calculation. Reinitialize the starting point and This ensures that the dynamic threshold is always anchored to the current real physical conditions, avoiding the accumulation of errors over long periods of operation.

[0191] Step S400: Residual calibration and image compensation based on dynamic uncertainty

[0192] When step S300 outputs a discrete state label as a transition state and issues a `COMPENSATE_TRIGGER` instruction, this module initiates a dynamic uncertainty assessment and feature-level residual calibration process in the independent GPU computing core to generate geometric / grayscale compensation parameters for the current frame, ensuring recording continuity under mild distortion conditions.

[0193] Construction of observation noise covariance matrix

[0194] Data Input and Feature Mapping: Receive probe motion state data synchronized in step S100. With cardiac ultrasound image frames .extract Six-degree-of-freedom composite amplitude and calculate Gray-scale variance of the whole frame or local ROI As a proxy indicator of image sharpness.

[0195] Covariance calculation model: Based on the mutually exclusive belief logic between physical signal and image quality, a scalar of observation noise covariance is constructed. :

[0196]

[0197] in , =0.5 is the balance coefficient. Zero-prevention. Its engineering and physical logic lies in the fact that an increase in probe motion amplitude or a decrease in image variance (texture blurring / contrast reduction) both indicate a decrease in the reliability of the observed data. The system linearly superimposes uncertainty measures accordingly. To adapt to the real-time requirements of embedded platforms, the high-dimensional matrix is ​​reduced to a diagonal scalar covariance, i.e. This avoids computational bottlenecks caused by online inversion.

[0198] (2) Dynamic calibration gain calculation and numerical stability control

[0199] Gain calculation chain: Read the predicted covariance scalar from the output of the LSTM model in step S300. (Approximate estimation by Monte Carlo Dropout), calculate dynamic calibration gain according to the Kalman-like minimum variance estimation principle. :

[0200]

[0201] Control logic and boundary constraints: This gain naturally satisfies [0, 1]. When motion perturbation is severe or the image is severely degraded, Increase The system automatically reduces the weight of observed features, and the compensation results rely more on historical patterns to predict the baseline; when the probe is stable and the image is clear, Decrease The system fully incorporates current observational characteristics. To prevent numerical oscillations caused by the denominator approaching zero under extreme conditions, a saturation limiting logic is embedded in the calculation kernel to force... [0.05, 0.95].

[0202] (3) Characteristic-level residual fusion and compensation parameter generation

[0203] Residual calibration execution: Predicted and observed features are weighted and fused using dynamic gain to obtain the repaired latent spatiotemporal features.

[0204]

[0205] Compensated parameter regression mapping: To avoid the risk to diagnostic accuracy caused by directly decoding and generating entirely new anatomical structures, the system will... Input a lightweight fully connected regression head (2 layers, <50K parameters), output the affine correction and grayscale compensation parameter set for the current frame. These represent the lateral / vertical translation, in-plane rotation angle, gain compensation coefficient, and contrast correction factor, respectively. The regression head is jointly trained with the encoder / prediction model. The loss function includes parameter smoothing constraints (L2 regularization) and anatomical edge preservation terms (Sobel gradient loss) to ensure that the output parameters only undergo rigid / affine level correction and global grayscale equalization without changing the original tissue texture topology.

[0206] Abnormal rollback strategy: If the rollback head output exceeds the device's physical adjustment limit (e.g., ... Pixels or If the current frame exceeds the repairable range of the transition state (dB), the system immediately discards the compensation parameters, marks the current frame as `RAW_FALLBACK`, and directly enters the original frame writing or truncation logic in step S500.

[0207] (4) Diagnostic authenticity assurance and hardware scheduling

[0208] Raw frame anchoring mechanism: All compensation parameters apply only to the raw cardiac ultrasound image frames buffered in step S100. Subpixel-level geometric resampling and grayscale mapping are performed through the GPU’s built-in CUDA Image Warp and Look-Up Table (LUT) hardware units to generate compensated image frames I'_t.

[0209] Quality gate verification: calculation and Structural similarity (SSIM) and peak signal-to-noise ratio (PSNR) are evaluated. If SSIM < 0.92 or PSNR < 28 dB, it indicates overcompensation or feature mismatch, and the system automatically rolls back to the original frame. Furthermore, the `Quality_Flag=Compensation_Reverted` flag is added to the DICOM metadata to ensure that clinical diagnoses are always based on the original, unaltered echo data. This step maintains a strictly controlled end-to-end latency of less than 5ms, seamlessly integrating with the recording frame rate.

[0210] Step S500: Video recording execution and state-driven control strategy

[0211] This step receives the state decision command from step S300 and the compensated image stream from step S400. Based on the cardiac cycle phase and the circular buffer queue, it performs differentiated video recording, distortion blocking, and attribution interaction to form a complete clinical operation control closed loop.

[0212] (1) Steady-state phase locking and backtracking writing

[0213] Trigger condition determination: When the `RECORD_ENABLE` command is received and the current cardiac cycle phase is... Falling into the preset trigger window (preferably end-diastolic, corresponding to) When the value is [0.85, 1.0], the recording start condition is met.

[0214] Buffer backtracking control: To meet the requirement of continuous cardiac cycle in echocardiography diagnosis, the system does not immediately write the current frame. Instead, it reads the circular time-series buffer queue maintained in step S100 and extracts the original / compensated image sequence of the previous K frames (preferably K=20) at the current time. This backtracking operation is performed using the modulo operation of the circular pointer. Achieve direct addressing with zero memory copy.

[0215] DICOM Encapsulation Writing: The backtracking sequence is concatenated with the current stable frame in chronological order, and a DICOM standard header file (containing metadata such as patient ID, device gain, probe frequency, and acoustic quality index sequence) is injected. This is then asynchronously written to the video storage unit via a high-speed NVMe interface. The writing process is independent of the main processing thread and employs a double-buffering mechanism to avoid I / O blocking that could cause control cycle synchronization issues.

[0216] (2) Transitional state repair frame serialization writing

[0217] Compensated Frame Flow Control: When the state is in a transitional state, the system skips the phase trigger window limitation and directly outputs the compensated image frame from step S400. Or, raw frames marked as `Compensated` are pushed into the video write queue.

[0218] Continuity guarantee mechanism: To prevent diagnostic consistency degradation caused by continuous writes during transition states, the system records the count of consecutive transition state frames. .like > 15. Automatically insert the `Frame_Type=Interpolated / Compensated` flag into the DICOM sequence for subsequent clinical image reading systems to automatically highlight and prompt, balancing video continuity and diagnostic transparency.

[0219] (3) Distortion-state blocking, attribution analysis and interactive prompts

[0220] Hard write cutoff: When the `RECORD_HALT` command is received, the recording execution module immediately closes the current DICOM file handle and executes the file system `fsync` to force a disk flush, preventing data corruption or index breakage.

[0221] Multimodal distortion attribution logic: The system performs a three-level attribution determination based on the real-time features of steps S100-S300.

[0222] 1. Probe slip distortion: If | | > 0.5g and angular velocity If the system detects severe hand tremors or probe slippage, it will output a red "Probe Position Adjustment" message via the interactive terminal and trigger a 200ms tactile vibration alarm from the built-in linear motor.

[0223] 2. Acoustic coupling / sound image distortion: If | |Below the exercise limit but If the value is less than 0.15, it is determined to be due to coupling agent breakage or strong gas reflection from the ribs / lungs. A system startup delay mechanism (preferably 2 seconds) is implemented, during which the feature calculation link remains running but is not written to the file.

[0224] 3. Signal attenuation / gain mismatch: If the global grayscale mean of the image is lower than 5 consecutive frames... and If the beamsmears slowly, it is determined to be due to deep attenuation or improper equipment gain settings. The system automatically sends a +3dB gain compensation command to the beam combiner and marks it as "Pending operator confirmation".

[0225] Recovery and truncation strategy: During the delay or gain compensation period, if the state recovers to a stable or transitional state for 3 consecutive frames, the system automatically creates a new DICOM file to continue recording; if the state does not recover within the timeout period, the current recording segment is truncated, and "Recording has been terminated: Acoustic signal is unavailable" is displayed on the interactive terminal, and a diagnostic log containing distortion timestamps and attribution types is generated.

[0226] (4) Storage resource scheduling and closed-loop signal flow

[0227] Circular queue lifecycle management: The video recording write pointer and feature extraction pointer are decoupled, adopting a producer-consumer model. After writing is complete, the corresponding buffer slot is marked as `FREE`, for cyclic overwriting by step S100. Memory usage remains constant, preventing fragmentation during long-term operation.

[0228] Control command closed loop: The attribution log and video recording status flags output in this step are transmitted back to the interactive terminal for display in real time. At the same time, they are fed into the online model iteration module of subsequent steps as optional inputs, forming a complete medical device control closed loop of "acquisition → decision → compensation / execution → feedback → optimization".

[0229] This application achieves, for the first time, real-time separation of probe physical motion and cardiac physiological motion in automated cardiac ultrasound recording through hardware-level synchronous fusion of ultrasound images and IMU six-degree-of-freedom motion data. This is achieved by combining explicit anatomical features (left ventricular contour, atrioventricular junction, and orderly valve motion) with implicit spatiotemporal features (adaptive gating fusion of images and motion data) in a dual-channel decoupling design. This fundamentally solves the problem of misjudgment caused by sectional distortion due to the inability of existing technologies to distinguish between "probe jitter" and "cardiac motion." Based on this, a nonlinear coupled energy function is constructed, fusing temporal prediction bias (LSTM autoregression), anatomical morphology bias (exponential penalty term), and acoustic quality gating (physical contact factor). A stable / distortion judgment boundary is statistically dynamically generated based on a sliding window, eliminating the need for preset fixed thresholds. This allows for adaptation to different patient heart rates, subcutaneous tissue acoustic characteristics, and operator grip habits, exhibiting both high sensitivity and strong robustness. Meanwhile, this application defines a "transition state" and introduces a residual calibration mechanism (Kalman-like gain) based on observation uncertainty. When there is slight distortion, the image is written after feature-level fusion and compensation, avoiding the loss of diagnostic information caused by simply discarding frames or stopping recording, and significantly improving the continuity of automatic recording and effective diagnostic time.

[0230] Furthermore, this application establishes a closed-loop continuous optimization capability: successfully recorded stable-state segments are used as high-quality positive samples to fine-tune the LSTM prediction model online, allowing the energy function benchmark to gradually adapt to the current patient and operator characteristics, achieving personalized optimization that becomes more accurate with use, overcoming the shortcomings of fixed parameters in existing open-loop systems. At the clinical operation level, the system automatically completes distortion attribution (probe slippage, acoustic obstruction, depth attenuation), end-diastolic phase-locked recording, and multimodal interactive prompts, significantly reducing recording quality fluctuations caused by differences in operator experience, ensuring the standardization and overall efficiency of cardiac ultrasound examinations. In summary, this application forms a complete anti-distortion automatic recording technology solution from signal decoupling, dynamic judgment, intelligent repair, closed-loop optimization to clinical standardization, and has significant clinical application value.

[0231] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the invention without departing from the principles and spirit of the invention, and all such changes should fall within the protection scope of the claims of the present invention.

Claims

1. A method for automatic recording of cardiac ultrasound without distortion, applied to a cardiac ultrasound imaging system, the cardiac ultrasound imaging system comprising an ultrasound probe, a probe motion sensor, a processing unit, and a video storage unit, characterized in that, The method includes: Acquire continuous echocardiogram image frames and probe motion state data synchronized with the echocardiogram image frames, and obtain acoustic coupling quality index and cardiac cycle phase information based on the echocardiogram image frames; Exclusive anatomical quality features are extracted from the cardiac ultrasound image frames, and latent spatiotemporal fusion features are generated based on the cardiac ultrasound image frames and the probe motion state data. Based on the implicit spatiotemporal fusion features of historical moments, predictive spatiotemporal features of the current moment are generated, and based on the predicted spatiotemporal features, the implicit spatiotemporal fusion features of the current moment, the explicit anatomical quality features, and the acoustic coupling quality index, the cross-sectional quality deviation index of the current frame is obtained. Based on the cross-sectional quality deviation index of multiple consecutive frames within a preset time window, the stability judgment boundary and the distortion judgment boundary are dynamically determined, and the current frame is judged as a stable state, a transitional state, or a distorted state based on the stability judgment boundary and the distortion judgment boundary. Recording control is executed based on the state of the current frame, where: If the current frame is determined to be in a stable state, and the cardiac cycle phase information meets the preset recording triggering conditions, the recording storage unit is controlled to write the current frame or the buffer image frame corresponding to the current frame. If the current frame is determined to be in a transitional state, the observation uncertainty of the current frame is determined based on the probe motion state data and the image quality information of the cardiac ultrasound image frame. The predicted spatiotemporal features and the implicit spatiotemporal fusion features at the current moment are adaptively fused according to the observation uncertainty to generate compensation information for the current frame. The current frame is then compensated based on the compensation information and written to the video storage unit. If the current frame is determined to be distorted, pause or stop recording and output a distortion warning message.

2. The method according to claim 1, characterized in that, The probe motion state data includes the linear motion information and angular motion information of the ultrasound probe; the linear motion information and angular motion information are acquired by an inertial measurement unit installed on the ultrasound probe; the cardiac ultrasound image frame and the probe motion state data are synchronized through a timestamp, so that each cardiac ultrasound image frame corresponds to the probe motion state data within the acquisition time window of that cardiac ultrasound image frame.

3. The method according to claim 1, characterized in that, The acoustic coupling quality index is determined based on the superficial region echo information, deep region echo information, echo attenuation consistency between the superficial and deep regions, and image information content in the cardiac ultrasound image frame. The acoustic coupling quality index is configured to decrease when there is insufficient effective echo in the superficial region, insufficient effective echo in the deep region, abnormal echo attenuation relationship between the superficial and deep regions, or image information content is lower than a preset requirement. Furthermore, the decrease in the acoustic coupling quality index increases the section quality deviation index or makes the current frame more likely to be judged as a transitional or distorted state.

4. The method according to claim 1, characterized in that, The explicit anatomical quality features include at least two of the following: left ventricular contour deviation features, used to characterize the degree of deviation of the left ventricular endocardial contour from the standard section shape; atrioventricular junction position deviation features, used to characterize the degree of offset of the atrioventricular junction region from the preset reference position; and valve motion orderliness features, used to characterize the consistency or dispersion of valve region motion in consecutive cardiac ultrasound image frames; wherein, the explicit anatomical quality features are used to constrain whether the current frame conforms to the anatomical morphology requirements of the target section of cardiac ultrasound.

5. The method according to claim 1, characterized in that, Implicit spatiotemporal fusion features are generated based on the cardiac ultrasound image frames and the probe motion state data, including: Image encoding is performed on the cardiac ultrasound image frames to obtain image representation features; Motion encoding is performed on the probe motion state data to obtain motion characterization features; The adaptive fusion weights are determined based on the image representation features and the motion representation features; The image representation features and the motion representation features are fused based on the adaptive fusion weights to generate the implicit spatiotemporal fusion features; The adaptive fusion weights are used to adjust the contribution ratio of image information and probe motion information to the implicit spatiotemporal fusion features.

6. The method according to claim 1, characterized in that, Obtain the slice quality deviation index for the current frame, including: The temporal prediction bias is obtained based on the difference between the implicit spatiotemporal fusion features at the current moment and the predicted spatiotemporal features. Based on the aforementioned explicit anatomical quality characteristics, anatomical morphological deviations are obtained; Based on the acoustic coupling quality index, the acoustic quality modulation factor is obtained; The cross-sectional quality deviation index is obtained based on the time-series prediction deviation, anatomical morphology deviation, and acoustic quality modulation factor. The acoustic quality modulation factor is used to improve the influence of the timing prediction bias and / or anatomical morphology bias on the cross-sectional quality bias index when the acoustic coupling quality is reduced.

7. The method according to claim 1, characterized in that, Dynamically determine the stability and distortion criteria boundaries, including: Within the preset time window, the central trend information and dispersion information of the cross-sectional quality deviation index of multiple consecutive frames are statistically analyzed. The stability determination boundary and the distortion determination boundary are determined based on the central trend information and the dispersion information, wherein the stability determination boundary is lower than the distortion determination boundary. If the cross-sectional quality deviation index of the current frame is lower than the stability determination boundary, the current frame is determined to be in a stable state. If the cross-sectional quality deviation index of the current frame is between the stability determination boundary and the distortion determination boundary, the current frame is determined to be in a transition state. If the cross-sectional quality deviation index of the current frame is higher than the distortion determination boundary, the current frame is determined to be in a distorted state.

8. The method according to claim 1, characterized in that, If the current frame is determined to be in a transitional state, the observation uncertainty of the current frame is determined based on the probe motion state data and the image quality information of the cardiac ultrasound image frame, including: The degree of motion disturbance is determined based on the motion amplitude of the probe motion state data. The image reliability is determined based on at least one of the following: grayscale distribution, texture sharpness, local contrast, or image variance of the cardiac ultrasound image frame. The observation uncertainty is determined based on the degree of motion disturbance and the image reliability. Specifically, when the degree of motion disturbance increases or the image credibility decreases, the observation uncertainty increases, and the compensation information becomes more dependent on the predicted spatiotemporal features; when the degree of motion disturbance decreases or the image credibility increases, the observation uncertainty decreases, and the compensation information becomes more dependent on the implicit spatiotemporal fusion features at the current moment.

9. The method according to claim 1, characterized in that, If the current frame is determined to be in a distorted state, the method further includes distortion attribution processing: When the probe motion status data meets the preset motion abnormality conditions, the current distortion is attributed to probe slippage distortion, and a probe position adjustment prompt is output. When the probe motion state data does not meet the preset motion anomaly conditions and the acoustic coupling quality index is lower than the preset acoustic quality requirements, the current distortion is attributed to acoustic coupling distortion or sound shadow occlusion distortion, and a delay waiting mechanism is activated. When subsequent frames return to a stable or transitional state within the preset waiting time corresponding to the delay waiting mechanism, recording resumes; If subsequent frames do not recover to a stable or transitional state within the preset waiting time, the current recording file is truncated or a new recording file is generated.

10. A cardiac ultrasound anti-distortion automatic recording system, the system comprising an ultrasound probe, a probe motion sensor, and a video storage unit, characterized in that, include: The multimodal acquisition module is used to acquire continuous cardiac ultrasound image frames, probe motion state data synchronized with the cardiac ultrasound image frames, and obtain acoustic coupling quality index and cardiac cycle phase information based on the cardiac ultrasound image frames. The feature analysis module is used to extract explicit anatomical quality features based on the cardiac ultrasound image frames and generate implicit spatiotemporal fusion features based on the cardiac ultrasound image frames and the probe motion state data. The prediction and evaluation module is used to generate the predicted spatiotemporal features of the current moment based on the implicit spatiotemporal fusion features of historical moments, and to obtain the cross-sectional quality deviation index of the current frame based on the predicted spatiotemporal features, the implicit spatiotemporal fusion features of the current moment, the explicit anatomical quality features, and the acoustic coupling quality index. The state determination module is used to dynamically determine the stability determination boundary and the distortion determination boundary based on the cross-sectional quality deviation index of multiple consecutive frames within a preset time window, and to determine the current frame as a stable state, a transitional state or a distorted state based on the stability determination boundary and the distortion determination boundary. An adaptive compensation module is used to determine the observation uncertainty of the current frame based on the probe motion state data and the image quality information of the cardiac ultrasound image frame when the current frame is determined to be in a transitional state, and to adaptively fuse the predicted spatiotemporal features and the implicit spatiotemporal fusion features at the current moment according to the observation uncertainty, so as to generate compensation information for the current frame. The recording execution module is used to execute recording writing when the current frame is determined to be in a stable state and the cardiac cycle phase information meets the preset recording trigger conditions; to execute recording writing after compensating the current frame based on the compensation information when the current frame is determined to be in a transitional state; and to pause or stop recording writing and output distortion prompt information when the current frame is determined to be in a distorted state. The multimodal acquisition module is communicatively connected to the ultrasound probe and the probe motion sensor. The input terminal of the feature analysis module is connected to the output terminal of the multimodal acquisition module. The input terminal of the prediction and evaluation module is connected to the output terminal of the feature analysis module. The input terminal of the state determination module is connected to the output terminal of the prediction and evaluation module. The input terminal of the adaptive compensation module is connected to the state determination module, the multimodal acquisition module, and the feature analysis module, respectively. The input terminal of the video recording execution module is connected to the state determination module and the adaptive compensation module, respectively. The output terminal of the video recording execution module is connected to the video storage unit.

Citation Information

Patent Citations

  • CN119970093A

  • US10595731B2

  • US20210085294A1

  • US20240000433A1

  • WO2022183264A1