Ultrasound image processing device and method suitable for ultrasound-guided interventional surgical robot

By separating and extracting the static and dynamic features of the cardiac ultrasound image in an ultrasound guided interventional surgical robot and combining the processing, the problems of traditional monitoring of pericardial effusion are solved, and real-time dynamic monitoring of pericardial effusion is achieved.

CN120381302AActive Publication Date: 2025-07-29FUWAI HOSPITAL CHINESE ACAD OF MEDICAL SCI & PEKING UNION MEDICAL COLLEGE +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510885339.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Traditionally, the monitoring of pericardial effusion relies on manual measurement of ultrasound images, which has problems such as low efficiency, strong subjectivity and insufficient dynamic tracking capabilities.

Method used

Ultrasound-guided interventional surgical robot is used to continuously acquire cardiac ultrasound images at multiple time points, segment and extract static and dynamic features, and combine lightweight processing to monitor the changes in the volume of pericardial fluid in real time.

Benefits of technology

It improves the targeted nature of feature extraction, avoids information interference, saves computing resources, and realizes real-time dynamic monitoring of pericardial effusion on equipment with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120381302A_ABST
    Figure CN120381302A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the field of image analysis, in particular to an ultrasonic image processing device and method suitable for an ultrasonic guided interventional surgical robot. The ultrasound image processing device comprises: an acquisition module for continuously acquiring cardiac ultrasound images of a patient at a plurality of time points; the preprocessing module is used for segmenting to obtain a cardiac ultrasound image and a segmentation result; the basic feature extraction module is used for extracting static features of a segmentation result; the associated feature extraction module is used for extracting dynamic features of the segmentation result; the feature fusion module is used for combining the static features and the dynamic features and determining the volume change trend of the pericardial effusion; a module is optimized for incorporating lightweight deployment so that it can process ultrasound images in real time. By respectively extracting the static features and the dynamic features of the ultrasonic image, information of two different states is more effectively captured, so that information interference between different modules is avoided, and the pertinence of feature extraction is improved; information of all dimensions is prevented from being processed at the same time, and computing resources are saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of image analysis, and particularly to an ultrasonic image processing device and method applicable to an ultrasonic-guided interventional surgical robot. Background Art

[0002] The statements herein merely provide background information related to the present application and do not necessarily constitute prior art.

[0003] Under normal circumstances, the pericardial cavity contains a small amount of fluid, which mainly plays a role in lubrication and relieving friction. If the amount of fluid in the pericardial cavity increases and exceeds the normal range, it is called pericardial effusion, and the amount of pericardial effusion can be visualized by ultrasonic imaging. Summary of the Invention

[0004] A brief overview of the present application is given below to provide a basic understanding of certain aspects of the present application. It should be understood that this overview is not an exhaustive overview of the present application. It is not intended to identify the key or important parts of the present application, nor is it intended to limit the scope of the present application. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description to be presented later.

[0005] On the one hand, an embodiment of the present application provides an ultrasonic image processing device applicable to an ultrasonic-guided interventional surgical robot. The ultrasonic image is used to reflect the state of pericardial effusion, and it includes: an acquisition module for continuously acquiring cardiac ultrasonic images of a patient at multiple time points; a preprocessing module for segmenting the cardiac ultrasonic images to obtain a segmentation result; a basic feature extraction module for extracting static features of the segmentation result; a correlation feature extraction module for extracting dynamic features of the segmentation result; a feature fusion module for combining the static features and the dynamic features and determining the volume change trend of pericardial effusion according to the combination; and a deployment optimization module for lightweighting the combination so that it can process ultrasonic images in real time.

[0006] On the other hand, an embodiment of the present application provides an ultrasonic image processing method applicable to an ultrasonic-guided interventional surgical robot. The ultrasonic image is used to reflect the state of pericardial effusion, and it includes the following steps: S10: Continuously acquiring cardiac ultrasonic images of a patient at multiple time points; S20: Segmenting the cardiac ultrasonic images to obtain a segmentation result of the cardiac ultrasonic images; S30: Extracting static features of the segmentation result; S40: Extracting dynamic features of the segmentation result; S50: Combining the static features and the dynamic features and determining the volume change trend of pericardial effusion according to the combination; S60: Lightweighting the combination so that it can process cardiac ultrasonic images in real time.

[0007] In the ultrasonic image processing device applicable to an ultrasonic-guided interventional surgical robot provided by an embodiment of the present application, the cardiac ultrasonic images acquired in multiple time periods contain both spatial information and time information. By separately extracting the static features and dynamic features of the ultrasonic images, such separate processing can more effectively capture the information in these two different states, so as to avoid information interference between different modules and improve the pertinence of feature extraction. At the same time, it avoids processing all-dimensional information simultaneously, saving computing resources.

[0008] The ultrasonic image processing method applicable to an ultrasonic-guided interventional surgical robot provided by an embodiment of the present application, by separately extracting static features and dynamic features, makes the feature extraction targeted, avoids interference between multi-dimensional information, improves the computing efficiency, so as to realize real-time processing of the image sequence on edge devices with limited resources, thereby realizing dynamic monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to further elaborate the above and other advantages and features of the present application, the following specifically describes the specific embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings, together with the following detailed description, are included in this specification and form a part of this specification. Elements having the same function and structure are denoted by the same reference numerals. It should be understood that these drawings only depict typical examples of the present application and should not be regarded as limiting the scope of the present application.

[0010] Figure 1 is a structural block diagram of the ultrasonic image processing device provided by an embodiment of the present application; Figure 2 is a schematic flowchart of the ultrasonic image processing method provided by an embodiment of the present application.

[0011] Description of the reference numerals: 10, acquisition module; 20, preprocessing module; 30, basic feature extraction module; 40, associated feature extraction module; 50, feature fusion module; 60, deployment optimization module; 100, ultrasonic image processing device. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] Hereinafter, exemplary embodiments of the present application will be described in conjunction with the accompanying drawings. For the sake of clarity and conciseness, not all features of the actual embodiments are described in the specification. However, it should be understood that many implementation-specific decisions must be made during the development of any such actual embodiment in order to achieve the developer's specific goals, for example, to comply with those limitations related to the system and the business, and these limitations may vary with different embodiments. In addition, it should also be understood that although the development work may be very complex and time-consuming, for those skilled in the art who benefit from the content of the present application, such development work is only a routine task.

[0013] Here, it should also be noted that in order to avoid obscuring the present application due to unnecessary details, only the device structures and / or processing steps closely related to the solution according to the present application are shown in the drawings, while other details less relevant to the present application are omitted.

[0014] The following disclosure provides multiple different embodiments or examples for implementing the present application. To simplify the disclosure of the present application, components and methods of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In the description of the embodiments of the present application, the meaning of "multiple" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0015] Pericardial effusion is an important indication of heart diseases. Traditional monitoring relies on manual measurement of the effusion area in ultrasonic images, which has problems such as low efficiency, strong subjectivity, and insufficient dynamic tracking ability.

[0016] An embodiment of the present application provides an ultrasonic image processing device applicable to an ultrasonic-guided interventional surgical robot. The ultrasonic image is used to reflect the state of pericardial effusion. Figure 1 The structural block diagram of the ultrasonic image processing device provided by the embodiment of the present application is shown, as Figure 1 shown. It includes: an acquisition module 10 for continuously acquiring cardiac ultrasonic images of a patient at multiple time points; a preprocessing module 20 for segmenting the cardiac ultrasonic images to obtain a segmentation result; a basic feature extraction module 30 for extracting static features of the segmentation result; a correlation feature extraction module 40 for extracting dynamic features of the segmentation result; a feature fusion module 50 for combining the static features and the dynamic features and determining the volume change trend of pericardial effusion according to the combination; and a deployment optimization module 60 for lightening the combination so that it can process ultrasonic images in real time.

[0017] In the ultrasonic image processing device 100 applicable to an ultrasonic-guided interventional surgical robot provided by the embodiment of the present application, the cardiac ultrasonic images acquired in multiple time periods contain both spatial information and time information. By separately extracting the static features and dynamic features of the ultrasonic images, such separate processing can more effectively capture the information of these two different states, so as to avoid information interference between different modules and improve the pertinence of feature extraction; at the same time, avoid processing all dimensions of information simultaneously and save computing resources.

[0018] In some embodiments, the preprocessing module 20 is configured to be able to receive cardiac ultrasonic images and is configured to be able to output the segmentation result of the ultrasonic images. The segmentation result includes a binary mask of the area where pericardial effusion is located in the cardiac ultrasonic image and a sequence of predicted values of the volume of pericardial effusion in the cardiac ultrasonic image.

[0019] In some embodiments, the cardiac ultrasound image and the segmentation result satisfy the following relationship: .

[0020] Where T represents the number of time points; t represents the time period serial number; represents the features of the cardiac ultrasound images corresponding to T time periods respectively, and the features include texture, edge, and morphological feature information; Model represents the relationship between; represents the pixel-level segmentation result of the pericardial effusion region in the t-th time period, and the output result is a binary mask; represents the sequence of predicted values of the pericardial effusion volume from the (t + 1)-th to the (t + 10)-th time periods, with the unit of mL, and is used to dynamically monitor the growth trend of the effusion.

[0021] In some embodiments, the output result of the preprocessing module 20 is the segmented cardiac ultrasound image of the pericardial effusion and the predicted sequence values of the pericardial effusion volume, so as to facilitate subsequent analysis of the pericardial effusion region and prediction of the pericardial effusion volume.

[0022] In some embodiments, the basic feature extraction module 30 includes: a first basic feature extraction unit for extracting the color and edge features of the cardiac ultrasound image; a second basic feature extraction unit for extracting the texture features of the cardiac ultrasound image; a third basic feature extraction unit for extracting the morphological features of the cardiac ultrasound image, and the first basic feature extraction unit, the second basic feature extraction unit, and the third basic feature extraction unit are arranged in parallel.

[0023] In some embodiments, by setting multiple parallel extraction units for extracting static features of the cardiac ultrasound image, the features extracted by each unit are complementary, which is convenient for feature splicing, and then a consistent and comprehensive feature map is output.

[0024] In some embodiments, the first basic feature extraction unit includes one pooling layer and two dilated convolutional layers to enhance the gray-scale distribution contrast and edge features in the region where the pericardial effusion is located, and the extraction of the edge features satisfies the following relationship: .

[0025] Where represents the original pixel matrix of the input cardiac ultrasound image; represents the max pooling operation to reduce the resolution of the cardiac ultrasound image, thereby enhancing the edge saliency; represents the dilated convolution with a dilation rate of 2 to expand the receptive field to capture long-distance edge features; and represent the convolutional kernel weight matrices for extracting edge features in different directions; Represents the output edge feature map, which characterizes the gray-scale difference boundary between pericardial effusion and surrounding tissues.

[0026] In some embodiments, setting one pooling layer can reduce the feature map size while retaining the most obvious color and edge features, reducing the computational amount and accelerating the subsequent processing speed; setting two dilated convolutional layers can avoid information loss caused by pooling and capture a larger range of context information.

[0027] In some embodiments, the second basic feature extraction unit includes three convolutional layers and one channel attention layer to quantify the internal intensity heterogeneity of pericardial effusion. The extraction of texture features satisfies the following relationship: .

[0028] .

[0029] Among them, Represents the feature map output by the third convolutional layer, which characterizes the local texture details of the region where the pericardial effusion is located; Represents global average pooling, which compresses the feature map into a channel description vector; Represents the fully connected layer weight, which generates the channel attention weight; Represents the Sigmoid activation function, which maps the weight to the interval [0,1]; Represents the channel attention weight, which quantifies the importance of different texture feature channels; Represents the weighted texture feature map, which characterizes the response of the heterogeneous region.

[0030] In some embodiments, setting three convolutional layers can achieve a gradual understanding of the image from simple to complex to gradually extract multi-level features; setting one channel attention layer can achieve intelligent feature screening and highlight key features. Such a combination can balance the depth of the convolutional layer and the position of the attention, being able to understand complex patterns and avoid insufficient computing power, making the attention judgment more accurate.

[0031] In some embodiments, the third basic feature extraction unit includes four convolutional layers and a dynamic receptive field adjustment module. The convolutional kernel size decreases gradually to capture the global geometric shape of the region where the pericardial effusion is located. The extraction of morphological features satisfies the following relationship: .

[0032] .

[0033] Among them, Represents the feature map output by the fourth convolutional layer, which characterizes the geometric shape information of the encoded effusion region; Represents the dynamic convolutional kernel weight, which is used to calculate the feature point offset; Represents the spatial offset of the feature points, and adjusts the receptive field to adapt to the morphological changes of the effusion; and Represents the initial feature point position and the coordinates of adjacent points; Represents the learnable weight coefficient, which controls the contribution of different positions to the morphological features; Represents the original pixel matrix of the input cardiac ultrasound image; Represents the number of feature points, which is used to describe the total number of feature points considered or calculated during the feature extraction process; n represents the number of the feature point, which is used to distinguish different feature points in the set; Represents the output morphological feature map, which characterizes the global contour and size changes of the pericardial effusion area.

[0034] In some embodiments, setting four convolutional layers can achieve progressive analysis from global to local; setting a dynamic receptive field adjustment module can adapt to targets of different sizes. This combination of hierarchical feature extraction and adaptive perception not only optimizes the computational efficiency, making the feature extraction process suitable for edge device deployment, but also enhances the robustness of the model to complex scenarios and improves the accuracy and stability of real-time inference.

[0035] In some embodiments, the extraction of associated features satisfies the following relationship: .

[0036] Wherein, Represents the associated feature map extracted through 3D convolution operation, which contains dynamic information in the time series; Represents the 3D convolution operation, which can process information in both spatial and time dimensions simultaneously; Represents the sequence of feature maps from t-k to t time, which is the input of the 3D convolution, where t represents the time point and k represents the size of the time window; Represents the weight of the 3D convolution kernel, and the weight performs a dot product operation with the input feature map during the convolution process to extract associated features.

[0037] In some embodiments, the associated feature extraction module 40 may include an edge-texture associated feature extraction unit, a spatial associated feature extraction unit, and a temporal associated feature extraction unit.

[0038] The spatial correlation feature extraction unit may include two convolutional layers and one spatial attention layer for modeling the edge continuity of the fluid accumulation region, where the convolutional kernel size can be set to 1×1; the edge-texture correlation feature extraction unit may include zero pooling layers and three cross-channel convolutional layers. By completely discarding the pooling layers, the spatial information loss caused by downsampling is thoroughly avoided, the original resolution of the feature map is maintained, and cross-channel feature interaction is achieved while fusing the output of the spatial correlation feature extraction unit to model the edges and textures of the fluid accumulation area; the temporal correlation feature extraction unit can be set to perform 3D convolution using three, four, and five temporal convolutional layers respectively, which can capture the high-frequency changes between consecutive frames and expand the receptive field to achieve multi-granularity temporal modeling, thereby extracting the dynamic features of the changes in the fluid accumulation thickness and area between consecutive frames.

[0039] In some embodiments, the feature fusion module 50 includes: a spatial modeling unit for generating a spatial feature map by adaptively weighting and fusing static features; a temporal modeling unit for combining the spatial feature map with dynamic features to obtain the relationship of pericardial effusion between ultrasonic images.

[0040] In some embodiments, by separately processing static features and dynamic features, the spatial modeling can refine local details, single-frame image processing does not require cross-frame calculation, reducing memory occupancy and achieving lightweight spatial modeling; and the temporal modeling can capture global dynamics, only performing serialization processing on the associated features to avoid the high complexity and large computational amount of full spatio-temporal dimension processing.

[0041] In some embodiments, the relationship of pericardial effusion between ultrasonic images satisfies the following expression: 。

[0042] Where, represents the query matrix, characterizing the feature information that needs to be focused on currently; represents the key matrix, used to match with the query matrix to determine the weights of the focus points; represents the value matrix, containing the actual information content, and performing weighted summation according to the attention weights; represents the dimension of the key, used to scale the dot product result to prevent gradient disappearance or explosion; represents a normalization function that converts the input value into a probability distribution to ensure that the sum of all attention weights is 1; represents the attention mechanism, used to calculate the correlation between different inputs, thereby dynamically adjusting the weights of the information.

[0043] In some embodiments, the associated feature maps extracted through 3D convolution operations can be input into a spatio-temporal Transformer module, and the long-term dependencies of the pericardial effusion evolution between adjacent cardiac ultrasound images can be modeled through the multi-head self-attention mechanism, thereby breaking through the processing limitations of local feature relationships, modeling the global interactions among all positions and lines in space, providing a more powerful modeling ability for multi-line convolution feature fusion, and facilitating a clear understanding and reflection of complex spatio-temporal relationships.

[0044] In some embodiments, the spatio-temporal Transformer module includes a dual-branch structure, namely, a spatial encoding branch and a temporal encoding branch. Among them, the spatial encoding branch can use progressive downsampling convolutional layers to extract the effusion edge and echo intensity features in a single-frame ultrasound image; the temporal encoding branch can model the correlation between adjacent-frame ultrasound images through the multi-head self-attention mechanism, thereby capturing the dynamic change law of the effusion volume.

[0045] In some embodiments, the feature fusion module 50 further includes a dynamic prediction unit for predicting the volume change trend of pericardial effusion in a future ultrasound image sequence; among them, the predicted values of pericardial effusion for the next 10 frames satisfy the following relational expression: 。

[0046] 。

[0047] represents a recurrent neural network for capturing long-term dependencies in sequential data; represents the hidden state at time t, the internal state of the LSTM, which contains all the information up to the current moment; represents the cell state at time t, the core memory unit of the LSTM, used to store long-term dependency information; represents the hidden state at time t - 1; represents the cell state at time t - 1; represents the fused feature, combining the information of spatial and temporal features, as the input of the LSTM; represents the output weight matrix for converting the hidden state into the final predicted value; represents the predicted values of the effusion volume for the next 10 frames, calculated based on the hidden state and output weight at the current moment.

[0048] In some embodiments, the dynamic prediction unit can predict the changing trend of pericardial effusion volume in the next 10 frames based on the fused features of the current frame, i.e., the effusion segmentation mask of the current frame, through the LSTM network. This helps to predict the changing trend of future effusion volume and estimate the probability of abnormal fluctuations, providing strong support for the dynamic monitoring and abnormal detection of pericardial effusion.

[0049] In some embodiments, the deployment optimization module 60 may include a lightweight modeling unit, a dynamic quantization unit, and an operator fusion optimization unit.

[0050] The lightweight modeling unit is used to train a lightweight student model according to the relationship between pericardial effusions in ultrasonic images and the output of the relationship formula for predicting pericardial effusion values in the next 10 frames, so that the lightweight student model can achieve real-time processing of the image sequence on the edge computing device through the dynamic quantization unit and the operator fusion optimization unit.

[0051] The dynamic quantization unit can dynamically adjust the quantization parameters according to the actual distribution of the input data, thereby reducing the model size and calculation, and realizing the lightweight of the student model; the operator fusion optimization unit can combine multiple calculation operations into one composite operation, thereby reducing the number of memory accesses and calculation overhead. Through dynamic quantization and operator fusion optimization operations, the student model can achieve efficient and real-time monitoring of pericardial effusion on the wearable ultrasonic device.

[0052] In some embodiments, the acquisition module 10, the preprocessing module 20, the basic feature extraction module 30, the associated feature extraction module 40, and the feature fusion module 50 can be integrated into the ultrasonic monitoring component through the deployment optimization module 60, so that the wearable electrocardiogram monitoring device can achieve real-time processing of the cardiac ultrasonic image sequence, facilitating timely feedback on the cardiac function of the patient.

[0053] Another aspect of the embodiments of the present application provides an ultrasonic image processing method applicable to an ultrasonic-guided interventional surgical robot, where the ultrasonic image is used to reflect the state of pericardial effusion. Figure 2 The flowchart of the ultrasonic image processing method provided by the embodiments of the present application is shown, as Figure 2 shown, which includes the following steps: S10: Continuously acquire cardiac ultrasonic images of a patient at multiple time points; S20: Segment the cardiac ultrasonic images to obtain the segmentation result of the cardiac ultrasonic images; S30: Extract the static features of the segmentation result; S40: Extract the dynamic features of the segmentation result; S50: Combine the static features and the dynamic features, and determine the changing trend of the pericardial effusion volume according to the combination; S60: Lightweight the combination so that it can process cardiac ultrasonic images in real time.

[0054] The ultrasonic image processing method provided by the embodiments of the present application for an ultrasonic-guided interventional surgical robot makes the feature extraction targeted by separately extracting static features and dynamic features, avoiding interference between multi-dimensional information, improving the calculation efficiency, so as to realize real-time processing of the image sequence on edge devices with limited resources, thereby realizing dynamic monitoring.

[0055] In some embodiments, in step S20, the segmentation result includes a binary mask of the region where pericardial effusion is located in the cardiac ultrasound image and a sequence of predicted values of the volume of pericardial effusion in the cardiac ultrasound image.

[0056] In some embodiments, the cardiac ultrasound image and the segmentation result satisfy the following relationship: 。

[0057] Where, T represents the number of time points; t represents the time period serial number; represents the features of the cardiac ultrasound images corresponding to T time periods respectively, and the features include texture, edge, and morphological feature information; Model represents the relationship between; represents the pixel-level segmentation result of the pericardial effusion region in the t-th time period, and the output result is a binary mask; represents the sequence of predicted values of the volume of pericardial effusion from the (t + 1)-th to the (t + 10)-th time periods, in mL, for dynamically monitoring the growth trend of the effusion.

[0058] In some embodiments, multi-line convolutional neural network calculations can be performed on the pixel-level segmentation result. In step S30, the following steps are further included: S31: Extract the color and edge features of the cardiac ultrasound image; S32: Extract the texture features of the cardiac ultrasound image; S33: Extract the morphological features of the cardiac ultrasound image.

[0059] In some embodiments, in step S31, the contrast of the gray distribution and the edge features of the region where pericardial effusion is located are enhanced through one pooling layer and two dilated convolutional layers, and the extraction of the edge features satisfies the following relationship: 。

[0060] Where, represents the original pixel matrix of the input cardiac ultrasound image; represents the max pooling operation to reduce the resolution of the cardiac ultrasound image, thereby enhancing the edge saliency; represents the dilated convolution with a dilation rate of 2 to expand the receptive field to capture long-distance edge features; and represent the convolutional kernel weight matrices for extracting edge features in different directions; It represents the edge feature map of the output, characterizing the grayscale difference boundary between pericardial effusion and surrounding tissues.

[0061] In some embodiments, in step S32, the internal intensity heterogeneity of pericardial effusion is quantified through three convolutional layers and one channel attention layer, and the extraction of texture features satisfies the following relationship: 。

[0062] 。

[0063] Among them, represents the feature map output by the third convolutional layer, characterizing the local texture details of the region where the pericardial effusion is located; represents global average pooling, which compresses the feature map into a channel description vector; represents the weights of the fully connected layer, generating channel attention weights; represents the Sigmoid activation function, mapping the weights to the interval [0, 1]; represents the channel attention weights, quantifying the importance of different texture feature channels; represents the weighted texture feature map, characterizing the response of the heterogeneous region.

[0064] In some embodiments, in step S33, the global geometric shape of the region where the pericardial effusion is located is captured through four convolutional layers and a dynamic receptive field adjustment module. Among them, the convolutional kernel size gradually decreases, and the extraction of morphological features satisfies the following relationship: 。

[0065] 。

[0066] Among them, represents the feature map output by the fourth convolutional layer, characterizing the geometric shape information of the encoded effusion region; represents the dynamic convolutional kernel weights, used to calculate the feature point offset; represents the spatial offset of the feature points, adjusting the receptive field to adapt to the morphological changes of the effusion; and represent the initial feature point position and the coordinates of adjacent points; represents the learnable weight coefficient, controlling the contribution of different positions to the morphological features; represents the output morphological feature map, characterizing the global contour and size changes of the pericardial effusion region.

[0067] In some embodiments, in step S40, the extraction of dynamic features satisfies the following relationship: 。

[0068] Among them, Represents the associated feature map extracted through 3D convolution operation, which contains dynamic information in the time series; Represents the 3D convolution operation, which can process information in both spatial and temporal dimensions simultaneously; Represents the sequence of feature maps from t - k to t time, which is the input of the 3D convolution, where t represents the time point and k represents the size of the time window; Represents the weight of the 3D convolution kernel. The weight performs a dot - product operation with the input feature map during convolution to extract dynamic features.

[0069] In some embodiments, in step S50, the following steps are further included: S51: Generate a spatial feature map by adaptively fusing static features; S52: Combine the spatial feature map with the dynamic features to obtain the relationship of pericardial effusion between ultrasonic images.

[0070] In some embodiments, the relationship of pericardial effusion between ultrasonic images satisfies the following expression: .

[0071] Where, Represents the query matrix, which characterizes the feature information that needs to be focused on currently; Represents the key matrix, which is used to match with the query matrix to determine the weight of the focus point; Represents the value matrix, which contains the actual information content and is weighted and summed according to the attention weight; Represents the dimension of the key, which is used to scale the dot - product result to prevent gradient vanishing or explosion; Represents a normalization function, which converts the input value into a probability distribution to ensure that the sum of all attention weights is 1; Represents the attention mechanism, which is used to calculate the correlation between different inputs, thereby dynamically adjusting the weight of information.

[0072] In some embodiments, in step S50, step S53 is further included: Predict the volume change trend of pericardial effusion in the future ultrasonic image sequence; where, the predicted values of pericardial effusion in the next 10 frames satisfy the following relational expression: .

[0073] .

[0074] Where, Represents the recurrent neural network, which is used to capture the long - term dependencies in the sequence data; Represents the hidden state at time t, which is the internal state of the LSTM and contains all the information up to the current moment; Represents the cell state at time t, which is the core memory unit of the LSTM and is used to store long-term dependency information; Represents the hidden state at time t-1; Represents the cell state at time t-1; Represents the fused feature, which combines spatial and temporal feature information and serves as the input to the LSTM; Represents the output weight matrix, which is used to convert the hidden state into the final predicted value; Represents the predicted value of the pericardial effusion volume for the next 10 frames, calculated based on the hidden state and output weight at the current time.

[0075] For the embodiments of the present application, it should also be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other to obtain new embodiments.

[0076] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. The protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. An ultrasonic image processing device applicable to an ultrasonic-guided interventional surgical robot, wherein the ultrasonic image is used to reflect the state of pericardial effusion, and is characterized in that, It includes: An acquisition module, configured to continuously acquire echocardiogram images of a patient at multiple time points; A preprocessing module, configured to segment the obtained echocardiogram images to obtain a segmentation result; A basic feature extraction module, configured to extract static features of the segmentation result; A correlation feature extraction module, configured to extract dynamic features of the segmentation result; A feature fusion module, configured to combine the static features and the dynamic features, and determine the volume change trend of the pericardial effusion according to the combination; A deployment optimization module, configured to lightweight the combination so that it can process the ultrasonic images in real time.

2. The apparatus according to claim 1, wherein The preprocessing module is configured to be able to receive echocardiogram images and is configured to output a segmentation result of the ultrasonic images, The segmentation result includes a binary mask of the region where the pericardial effusion is located in the echocardiogram image and a sequence of predicted values of the volume of the pericardial effusion in the echocardiogram image.

3. The device according to claim 2, characterized in that, The echocardiogram image and the segmentation result satisfy the following relationship: ; wherein, T represents the number of time points; t represents the period serial number; Represent the features of the cardiac ultrasound images corresponding to T time periods respectively, and the features include texture, edge, and morphological feature information; Model representation The relationship between; Represents the pixel-level segmentation result of the pericardial effusion region in the t-th period, and the output result is the binary mask; Indicates the predicted pericardial effusion volume sequence from the (t + 1)-th to the (t + 10)-th period, with the unit of mL, and is used to dynamically monitor the effusion growth trend.

4. The device according to claim 3, characterized in that The basic feature extraction module includes: A first basic feature extraction unit, configured to extract color and edge features of the echocardiogram image; A second basic feature extraction unit, configured to extract texture features of the echocardiogram image; A third basic feature extraction unit, configured to extract morphological features of the echocardiogram image, The first basic feature extraction unit, the second basic feature extraction unit, and the third basic feature extraction unit are configured to be parallel.

5. The apparatus according to claim 4, wherein The first basic feature extraction unit includes one pooling layer and two dilated convolutional layers to enhance the gray-scale distribution contrast and edge features of the region where the pericardial effusion is located, and the extraction of the edge features satisfies the following relationship: ; Among them, represents the original pixel matrix of the input cardiac ultrasound image; Indicates a max pooling operation to reduce the resolution of the cardiac ultrasound image, thereby enhancing edge saliency; Indicates a dilated convolution with a dilation rate of 2 to expand the receptive field and capture long-range edge features; and represents the convolutional kernel weight matrix and is used to extract edge features in different directions; Indicates the output edge feature map, representing the grayscale difference boundary between the pericardial effusion and the surrounding tissues.

6. The apparatus according to claim 4, wherein The second basic feature extraction unit includes three convolutional layers and one channel attention layer to quantify the internal intensity heterogeneity of the pericardial effusion, and the extraction of the texture features satisfies the following relationship: ; ; Among them, represents the feature map of the output of the third-layer convolution, characterizing the local texture details of the region where the pericardial effusion is located; Represents global average pooling, which compresses the feature map into a channel description vector; Represents the fully connected layer weights and generates channel attention weights; Represents the Sigmoid activation function, mapping the weights to the interval [0, 1]; Indicates the channel attention weight, quantifying the importance of different texture feature channels; Indicates the weighted texture feature map, representing the response of the heterogeneous region.

7. The apparatus according to claim 4, wherein The third basic feature extraction unit includes four convolutional layers and a dynamic receptive field adjustment module, and the convolutional kernel size decreases gradually to capture the global geometric morphology of the region where the pericardial effusion is located, and the extraction of the morphological features satisfies the following relationship: ; ; Among them, represents the feature map of the output of the fourth-layer convolution, which characterizes the geometric shape information of the encoded effusion region; Represents the dynamic convolution kernel weights, which are used to calculate the feature point offsets; representing a spatial offset of the feature point and adjusting a receptive field to adapt to morphological changes of the effusion; and represent the initial feature point position and the coordinates of adjacent points; Represents learnable weight coefficients that control the contributions of different positions to morphological features; represent the original pixel matrix of the input cardiac ultrasound image; Indicates the number of the feature points, which is used to describe the total number of the feature points considered or calculated during the feature extraction process; n represents the number of the feature points, used to distinguish different feature points in the set; It represents the output morphological feature map, which characterizes the global contour and dimensional variation of the pericardial effusion region.

8. The device according to claim 1, characterized in that, The extraction of the correlation features satisfies the following relationship: ; Among them, represents the associated feature map extracted by 3D convolution operation, which contains dynamic information in the time series; Represents a three-dimensional convolution operation that can process information in both spatial and temporal dimensions; Denote the sequence of feature maps from time t - k to t, which is the input of 3D convolution, where t represents the time point and k represents the size of the time window; Represents the weights of the 3D convolution kernel, and the weights perform dot product operations with the input feature map during the convolution process to extract the associated features.

9. The device according to any one of claims 1-8, characterized in that The feature fusion module includes: A spatial modeling unit, configured to generate a spatial feature map by adaptively weighting and fusing the static features; A temporal modeling unit, configured to combine the spatial feature map and the dynamic features to obtain the relationship of the pericardial effusion between the ultrasonic images.

10. The apparatus according to claim 9, wherein The relationship of the pericardial effusion between the ultrasonic images satisfies the following expression: ; Among them, represents a query matrix, characterizing the feature information that needs to be concerned about currently; Indicates a key matrix for matching with the query matrix to determine the weights of the points of interest; Represents a value matrix, containing actual information content, and performs a weighted sum according to the attention weights; Represents the dimension of the key, which is used to scale the dot product result and prevent the vanishing or exploding of gradients; Represents a normalization function that converts the input value into a probability distribution, ensuring that the sum of all attention weights is 1; Represents an attention mechanism that is used to calculate the correlation between different inputs, thereby dynamically adjusting the weights of information.

11. The device according to claim 9, characterized in that, The feature fusion module further includes a dynamic prediction unit for predicting the volume change trend of the pericardial effusion in the future ultrasonic image sequence; wherein, the predicted values of the pericardial effusion in the next 10 frames satisfy the following relational expression: ; ; Represents a recurrent neural network, used to capture long-term dependencies in sequential data; Represents the hidden state at time t, which is the internal state of the LSTM and contains all the information up to the current moment; represents the cell state at time t, and the core memory unit of the LSTM is used to store long-term dependence information; represents the hidden state at time t-1; represents the cell state at time t-1; represent the fusion features, combining the information of spatial and temporal features, as the input to the LSTM; Represents the output weight matrix, which is used to convert the hidden state into the final predicted value; Indicates the predicted value of the fluid accumulation volume for the next 10 frames, calculated based on the hidden state and output weights at the current moment.

12. An ultrasonic image processing method applicable to an ultrasonic-guided interventional surgical robot, where the ultrasonic image is used to reflect the pericardial effusion state, and is characterized in that, It includes the following steps: S10: Continuously acquire the cardiac ultrasonic images of the patient at multiple time points; S20: Segment the cardiac ultrasonic images to obtain the segmentation results of the cardiac ultrasonic images; S30: Extract the static features of the segmentation results; S40: Extract the dynamic features of the segmentation results; S50: Combine the static features with the dynamic features, and determine the volume change trend of the pericardial effusion according to the combination; S60: Lightweight the combination so that it can process the cardiac ultrasonic images in real time.

13. The method according to claim 12, characterized in that, In step S20, the segmentation results include the binary mask of the area where the pericardial effusion is located in the cardiac ultrasonic image and the sequence of predicted values of the volume of the pericardial effusion in the cardiac ultrasonic image.

14. The method according to claim 13, characterized in that, The cardiac ultrasonic image and the segmentation results satisfy the following relationship: ; wherein, T represents the number of time points; t represents the time period serial number; represent the features of the cardiac ultrasound images corresponding to T time periods respectively, where the features include texture, edge, and morphological feature information; Model representation The relationship between; Represents the pixel-level segmentation result of the pericardial effusion area in the t-th period, and the output result is the binary mask; Indicates the predicted pericardial effusion volume sequence from the (t + 1)-th period to the (t + 10)-th period, with the unit of mL, and is used to dynamically monitor the effusion growth trend.

15. The method according to claim 14, wherein Perform multi-line convolutional neural network calculation on the pixel-level segmentation results, In step S30, the following steps are further included: S31: Extract the color and edge features of the cardiac ultrasonic image; S32: Extract the texture features of the cardiac ultrasonic image; S33: Extract the morphological features of the cardiac ultrasonic image.

16. The method according to claim 15, wherein In step S31, enhance the gray-scale distribution contrast and edge features of the area where the pericardial effusion is located through one pooling layer and two dilated convolutional layers, and the extraction of the edge features satisfies the following relationship: ; Among them, represents the original pixel matrix of the input cardiac ultrasound image; Indicates a max pooling operation to reduce the resolution of the cardiac ultrasound image, thereby enhancing edge saliency; Indicates a dilated convolution with a dilation rate of 2 to expand the receptive field and capture long-distance edge features; and represents the convolutional kernel weight matrix and is used to extract edge features in different directions; It represents the output edge feature map, which characterizes the gray-scale difference boundary between the pericardial effusion and the surrounding tissues.

17. The method according to claim 15, wherein In step S32, quantify the internal intensity heterogeneity of the pericardial effusion through three convolutional layers and one channel attention layer, and the extraction of the texture features satisfies the following relationship: ; ; Among them, represents the feature map of the third-layer convolution output, characterizing the local texture details of the region where the pericardial effusion is located; Denotes global average pooling, which compresses the feature map into a channel description vector; Represents the fully connected layer weights and generates channel attention weights; Represents the Sigmoid activation function, mapping the weights to the interval [0, 1]; Indicates the channel attention weight, quantifying the importance of different texture feature channels; Indicates the weighted texture feature map, representing the response of the heterogeneous region.

18. The method according to claim 15, wherein In step S33, capture the global geometric morphology of the area where the pericardial effusion is located through four convolutional layers and a dynamic receptive field adjustment module, wherein the size of the convolutional kernel decreases gradually, and the extraction of the morphological features satisfies the following relationship: ; ; Among them, represents the feature map of the output of the fourth-layer convolution, which characterizes the geometric shape information of the encoded effusion area; Represents the dynamic convolution kernel weights, which are used to calculate the feature point offsets; Indicates the spatial offset of the feature points and adjusts the receptive field to adapt to the morphological changes of the effusion; and represent the initial feature point position and the coordinates of adjacent points; Represents the learnable weight coefficient, controlling the contribution of different positions to the morphological features; represents the original pixel matrix of the input cardiac ultrasound image; Denotes the number of the feature points, which is used to describe the total number of feature points considered or calculated during the feature extraction process; n represents the serial number of the feature point, which is used to distinguish different feature points in the set; It represents the morphological feature map of the output, characterizing the global contour and dimensional variation of the pericardial effusion region.

19. The method according to claim 12, wherein In step S40, the extraction of the dynamic features satisfies the following relationship: ; Among them, represents the associated feature map extracted by 3D convolution operation, which contains the dynamic information in the time series; Represents a three-dimensional convolutional operation that can process information in both spatial and temporal dimensions simultaneously; Denote the sequence of feature maps from time t - k to t, which is the input of 3D convolution, where t represents the time point and k represents the size of the time window; Represents the weights of the 3D convolution kernel, and the weights perform dot product operations with the input feature map during convolution to extract the dynamic features.

20. The method according to any one of claims 12-19, characterized in that, In step S50, the following steps are further included: S51: Generate a spatial feature map by adaptively weighting and fusing the static features; S52: Combine the spatial feature map with the dynamic features to obtain the relationship of the pericardial effusion between the ultrasonic images.

21. The method according to claim 20, wherein The relationship of the pericardial effusion between the ultrasonic images satisfies the following expression: ; Among them, represents a query matrix, characterizing the feature information that needs to be concerned about currently; Indicates a key matrix for matching with the query matrix to determine the weights of the focus points; Represents a value matrix, containing actual information content, and performs a weighted sum according to the attention weights; Represents the dimension of the key, used to scale the dot product result and prevent gradient vanishing or explosion; Represents a normalization function that converts the input value into a probability distribution, ensuring that the sum of all attention weights is 1; Represents an attention mechanism, which is used to calculate the correlation between different inputs, thereby dynamically adjusting the weights of information.

22. The method according to claim 12, wherein In step S50, step S53 is further included: Predict the volume change trend of the pericardial effusion in the future ultrasonic image sequence; wherein, the predicted values of the pericardial effusion in the next 10 frames satisfy the following relational expression: ; ; Among them, represents a recurrent neural network for capturing long-term dependencies in sequential data; Represents the hidden state at time t, which is the internal state of the LSTM and contains all the information up to the current time; represents the cell state at time t, and the core memory unit of the LSTM is used to store long-term dependency information; represents the hidden state at time t-1; represents the cell state at time t-1; represent the fusion features, which combine the information of spatial and temporal features, as the input to the LSTM; Represents the output weight matrix for converting the hidden state into a final predicted value; Indicates the predicted value of the pericardial effusion volume for the next 10 frames, calculated based on the hidden state and output weights at the current moment.

Citation Information

Patent Citations

  • System, method and computer-accessible medium for ultrasound analysis

    CN110461240A

  • Ultrasonic heart regurgitation automatic capture method and system, and ultrasonic imaging equipment

    CN111820947A

  • Ultrasonic cardiogram remote monitoring system for heart failure patient

    CN119837562A

  • Multi-mode breast volume ultrasonic focus grading method, medium and terminal

    CN119963887A

  • Ultrasonic diagnostic apparatus, image processing device, and image processing method

    JP2017121520A