Image processing method, system, device and equipment for motion trajectory planning of cardiac ultrasound robot
By generating fused attention features of global context and local discriminative differences in cardiac ultrasound images through a differential attention module, the problem of focusing on subtle lesions in existing technologies is solved, and precise control and automated scanning of the cardiac ultrasound robot's motion trajectory are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2026-01-08
- Publication Date
- 2026-04-17
AI Technical Summary
Existing medical image analysis methods struggle to effectively focus on subtle and critical discriminative regions in cardiac ultrasound images and are easily interfered with by a large amount of irrelevant background information, resulting in insufficient accuracy and automation in cardiac ultrasound robotic scanning.
A differential attention module is used to perform standard attention and differential attention operations in parallel, generating fused attention features of global context information and local discriminative difference information. The joint torque sequence is directly generated through the task output head to control the motion trajectory of the cardiac ultrasound robot.
It significantly enhances the relevance and effectiveness of feature representation, improves the accuracy, stability and operational safety of automated cardiac ultrasound scanning, and reduces the probability of robot malfunctions caused by image noise or complex backgrounds.
Smart Images

Figure CN121882091A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotic ultrasound imaging technology, and in particular to an image processing method, system, device and equipment for motion trajectory planning of cardiac ultrasound robots. Background Technology
[0002] Ultrasound imaging is a commonly used imaging technique in the diagnosis of cardiac diseases, enabling real-time observation of cardiac structure and function. Currently, deep learning-based medical image analysis methods, especially convolutional neural networks and the visual Transformer architecture, have been widely applied to the automated analysis of cardiac ultrasound images.
[0003] However, cardiac ultrasound images are characterized by high resolution, complex backgrounds, and a typically small proportion of critical lesion areas (such as minute lesions or specific anatomical structures). While existing medical image analysis methods can model long-range dependencies, their standard attention mechanisms often struggle to effectively focus on these subtle yet crucial discriminative regions during global computation. They are easily influenced by a large amount of irrelevant background information, leading to model distraction and inaccurate localization. This limits the feasibility of using their output to guide precise and automated cardiac ultrasound robotic scanning.
[0004] Therefore, how to accurately perceive and locate subtle lesions, improve the robustness of cardiac ultrasound image analysis, and enhance its practical value in robot-assisted diagnosis are technical problems that urgently need to be solved. Summary of the Invention
[0005] The purpose of this application is to provide an image processing method, system, device, and equipment for cardiac ultrasound robot motion trajectory planning, which can accurately capture subtle lesion structures in high-resolution ultrasound images and improve the accuracy and interpretability of cardiac ultrasound robot motion trajectory planning.
[0006] To solve the above-mentioned technical problems, this application is implemented as follows: A first aspect of this application discloses an image processing method for motion trajectory planning of a cardiac ultrasound robot, the method comprising: Acquire cardiac ultrasound images and preprocess the cardiac ultrasound images to extract image features; The image features are input into the differential attention module, and standard attention operations and differential attention operations are performed in parallel to generate a first attention feature representing global context information and a second attention feature representing local discriminative difference information, respectively. The first attention feature and the second attention feature are then concatenated and fused in the attention head dimension to obtain the fused attention feature. Based on the fused attention features, a target output is generated through the task output head, and the target output includes at least a joint torque sequence for controlling the cardiac ultrasound robot; The motion trajectory of the cardiac ultrasound robot is generated based on the joint torque sequence.
[0007] Optionally, the differential attention operation includes: Based on the input image features, generate a query matrix, a key matrix, and a value matrix; According to preset rules, a first query submatrix and a first key submatrix are determined from the query matrix and the key matrix respectively, and a preset second query submatrix and a second key matrix are obtained, wherein the second query submatrix and the second key matrix are identical matrices; Calculate the first relevance matrix of the first query submatrix and the first key submatrix, and the second relevance matrix of the second query submatrix and the second key matrix; Calculate the difference matrix between the first correlation matrix and the second correlation matrix; The second attention feature is calculated based on the difference matrix and the value matrix.
[0008] Optionally, determining the first query submatrix and the first key submatrix from the query matrix and the key matrix respectively according to preset rules includes: Calculate the vector norm of each column of the query matrix as a query norm sequence, and calculate the vector norm of each column of the key matrix as a key norm sequence; The importance weights of each dimension are determined by multiplying the corresponding elements of the query norm sequence and the key norm sequence. Based on the importance weights, important dimensions are selected from the query matrix and the key matrix respectively to form the first query submatrix and the first key submatrix.
[0009] Optionally, before or after concatenating and fusing the first attention feature and the second attention feature in the attention head dimension, the first attention feature and / or the second attention feature may be weighted and adjusted using learnable weight parameters.
[0010] Optionally, the differential attention modules are multiple and connected in series; The i-th differential attention module takes the fused attention feature output by the (i-1)-th differential attention module as input and outputs a new fused attention feature; i is an integer greater than 1, and the input of the first differential attention module is the image feature.
[0011] Optionally, the differential attention module is a neighborhood differential attention module or a deformable differential attention module; The neighborhood differential attention module is configured to compute attention within a local neighborhood window of the image features; The deformable differential attention module is configured to dynamically adjust the receptive field position of attention computation through a learnable offset network.
[0012] A second aspect of this application discloses a cardiac ultrasound robotic system, the system comprising: Ultrasound image acquisition device, used to acquire cardiac ultrasound images; The processing module is configured to execute the image processing method for motion trajectory planning of a cardiac ultrasound robot as described in the first aspect of the embodiments of this application, so as to generate the motion trajectory of the cardiac ultrasound robot. A robotic arm device is used to perform cardiac ultrasound scanning actions according to the said motion trajectory.
[0013] A third aspect of this application discloses an image processing apparatus for motion trajectory planning of a cardiac ultrasound robot, the apparatus comprising: The feature extraction module is used to acquire cardiac ultrasound images and preprocess the cardiac ultrasound images to extract image features. The image processing module inputs image features into the differential attention module, performs standard attention operations and differential attention operations in parallel, generates a first attention feature representing global context information and a second attention feature representing local discriminative difference information, respectively, and then concatenates and fuses the first attention feature and the second attention feature in the attention head dimension to obtain a fused attention feature. A sequence output module is used to generate a target output through a task output head based on the fused attention features, wherein the target output includes at least a joint torque sequence for controlling the cardiac ultrasound robot; The trajectory generation module is used to generate the motion trajectory of the cardiac ultrasound robot based on the joint torque sequence.
[0014] A fourth aspect of this application discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the image processing method for cardiac ultrasound robot motion trajectory planning described in the first aspect of this application.
[0015] The embodiments of this application have the following advantages: In this embodiment, the standard attention and differential attention operations are executed in parallel by the differential attention module, and the two attention features generated, which respectively represent global contextual information and local discriminative difference information, are spliced and fused in the attention head dimension. This enables the model to synergistically utilize global information and local subtle differences, effectively suppressing the interference of large background areas in cardiac ultrasound images, guiding the model to more accurately focus on discriminative small lesions or specific anatomical structures, thereby significantly enhancing the pertinence and effectiveness of feature representation.
[0016] Based on the aforementioned high-level semantic features that integrate global consistency and local discriminativity (fusion attention features), a joint torque sequence for controlling the cardiac ultrasound robot is directly generated through the task output head. This joint torque sequence originates from a more discriminative feature representation, resulting in a more accurate output motion trajectory and stronger anti-interference capability. This reduces the probability of robot malfunctions due to image noise or complex backgrounds, thereby improving the accuracy, stability, and operational safety of the automated cardiac ultrasound scanning process.
[0017] In this way, the method integrates image processing and robot trajectory planning into one framework. By directly decoding the fused high-level semantic features (fused attention features) into joint-level control parameters (joint torque sequence), it eliminates the cumbersome multi-stage processing or manual feature engineering in the middle, improves the system response efficiency, and ensures the consistency and coherence from visual understanding to physical action. This enables the robot's motion trajectory to respond more directly and accurately to the key anatomical and pathological information identified in the image. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart of the steps of an image processing method for motion trajectory planning of a cardiac ultrasound robot provided in an embodiment of this application; Figure 2 This is a schematic diagram of a differential attention module provided in an embodiment of this application; Figure 3 This is a schematic diagram of a cardiac ultrasound robot system provided in an embodiment of this application; Figure 4 This is a schematic diagram of an image processing device for motion trajectory planning of a cardiac ultrasound robot provided in an embodiment of this application; Figure 5This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0020] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] This application addresses the challenges of accurately locating lesions in high-resolution, fine-grained ultrasound images due to their small size and susceptibility to background noise. It proposes an image processing method for motion trajectory planning in cardiac ultrasound robots. The core concept involves combining standard attention and differential attention mechanisms within a Transformer architecture. Standard attention captures global contextual information, while differential attention focuses on extracting discriminative differences between local regions. These two mechanisms are fused at the attention head dimension, enabling the model to utilize both global semantics and local details simultaneously. This significantly suppresses interference from irrelevant backgrounds during training and inference, forcing attention to focus more precisely on minute lesions or key anatomical structures, thereby improving the quality of feature representation and the interpretability of decisions. Finally, the enhanced features are used to directly generate the joint torque sequence of the cardiac ultrasound robot, achieving end-to-end precise control from image understanding to motion planning, and improving the stability and reliability of robot scanning.
[0022] The image processing method, system, apparatus, and device for motion trajectory planning of cardiac ultrasound robots provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.
[0023] Reference Figure 1 As shown, Figure 1 This is a flowchart illustrating the steps of an image processing method for motion trajectory planning of a cardiac ultrasound robot, as provided in an embodiment of this application. Figure 1 As shown, an image processing method for motion trajectory planning of a cardiac ultrasound robot provided in this application embodiment may include steps S110 to S140: Step S110: Acquire cardiac ultrasound images and preprocess the cardiac ultrasound images to extract image features.
[0024] Echocardiographic images can be acquired using an ultrasound image acquisition device (such as an ultrasound probe). Specifically, during a cardiac ultrasound scan performed by a cardiac ultrasound robot, cardiac ultrasound images are acquired using an ultrasound image acquisition device.
[0025] Optionally, after acquiring cardiac ultrasound images, preprocessing operations can be performed on the cardiac ultrasound images to eliminate differences between devices and improve the model's generalization ability; among them, preprocessing operations may include, but are not limited to: image size normalization (e.g., adjusting to a fixed resolution), pixel value standardization (e.g., normalizing to the [0, 1] interval), and data augmentation (e.g., random rotation, flipping).
[0026] The echocardiogram image (or a pre-processed version) is fed into a feature extraction backbone network (e.g., a convolutional neural network or a shallow Visual Transformer (ViT) module). This feature extraction backbone network is responsible for extracting more expressive image features from the echocardiogram image. These image features are typically a feature tensor with spatial and channel dimensions, which serves as the base input for subsequent attention mechanisms.
[0027] Step S120: Input the image features into the differential attention module, perform standard attention operation and differential attention operation in parallel to generate a first attention feature representing global context information and a second attention feature representing local discriminative difference information, respectively, and then concatenate and fuse the first attention feature and the second attention feature in the attention head dimension to obtain the fused attention feature.
[0028] This step focuses on key regions (such as small lesions or key anatomical structures) in cardiac ultrasound images using a differential attention module. This module internally deploys two parallel processing pathways: a standard attention operation pathway and a differential attention operation pathway. The differential attention module processes the input image features based on these two pathways to obtain fused attention features.
[0029] The standard attention operation path employs the classic self-attention mechanism to calculate the correlation weights between all locations (or feature vectors) within an image feature set. This mechanism effectively models global contextual information, understands long-range dependencies between different parts of the image, and thus generates a first attention feature representing the global contextual information.
[0030] Differential Attention Pathway: This pathway employs a differential attention mechanism, highlighting subtle discriminative differences between local regions by calculating the difference between a baseline relevance and a target relevance. Specifically, attention may be calculated from a baseline state (e.g., using a fixed matrix or selectively dimensionality-reduced features) and then compared with the target attention calculated from the original or processed features. This mechanism effectively suppresses stable, undiscriminative information in the background, amplifying and focusing on subtle differences between lesion areas and normal tissues, or between different key anatomical structures, thus representing a second attention feature characterizing local discriminative difference information.
[0031] Subsequently, the first attention feature (global context) and the second attention feature (local difference) output from the two pathways are concatenated and fused at the attention head level. This fusion method integrates at the feature channel level, simultaneously preserving and synergistically utilizing global semantic understanding and local difference perception information. The resulting fused attention feature retains the necessary overall semantics while significantly enhancing the representation of discriminative regions, and its corresponding heatmap will show more concentrated and precise focus on lesion areas.
[0032] Step S130: Based on the fused attention features, generate a target output through the task output head, the target output including at least a joint torque sequence for controlling the cardiac ultrasound robot.
[0033] The task output header consists of one or more neural network layers (such as fully connected layers, multilayer perceptrons, or more complex decoders) adapted to a specific task. In practical applications, the task output header can be designed with different structures according to specific application requirements.
[0034] Based on the fused attention features rich in discriminative information generated in step S120, the task output head is trained to directly map the desired target output. In this embodiment, a key target output is the joint torque sequence used to control the cardiac ultrasound robot (i.e., the fused attention features can be mapped to a set of low-dimensional, specific joint torque sequences through the task output head). Each value in this joint torque sequence corresponds to the torque command (the magnitude of the torque or force to be applied) for controlling one or more joints (such as the axes of a robotic arm) of the cardiac ultrasound robot at the current moment; the joint torque sequence is the underlying control signal driving the motion of the cardiac ultrasound robot.
[0035] Step S140: Generate the motion trajectory of the cardiac ultrasound robot based on the joint torque sequence.
[0036] This step is used to convert control commands into physical actions. Specifically, based on the joint torque sequence generated in step S130, combined with the dynamic model, kinematic constraints, and safety parameters (such as joint limits and velocity limits) of the cardiac ultrasound robot, the joint torque sequence is converted into a smooth, continuous, and executable motion trajectory in three-dimensional space.
[0037] This motion trajectory guides the ultrasound probe through complex movements such as movement, tilting, and pressure on the patient's body surface to complete a preset cardiac ultrasound scanning task (e.g., standard section scanning) or dynamically adjusted based on real-time image analysis results. Ultimately, the cardiac ultrasound robot automatically executes this motion trajectory to acquire high-quality, repeatable cardiac ultrasound images.
[0038] The technical solution adopted in this application executes standard attention and differential attention operations in parallel through a differential attention module. The generated attention features, representing global contextual information and local discriminative differences respectively, are then concatenated and fused at the attention head dimension. This allows the model to synergistically utilize global information and subtle local differences, effectively suppressing interference from large background areas in cardiac ultrasound images. This guides the model to more accurately focus on discriminative micro-lesions or specific anatomical structures, significantly enhancing the specificity and effectiveness of feature representation. Based on the aforementioned high-level semantic features (fused attention features) that integrate global consistency and local discriminativity, a joint torque sequence for controlling the cardiac ultrasound robot is directly generated through the task output head. This joint torque sequence originates from a more discriminative feature representation, resulting in a more accurate output motion trajectory and stronger anti-interference capability. This reduces the probability of robot malfunctions due to image noise or complex backgrounds, thereby improving the accuracy, stability, and operational safety of the automated cardiac ultrasound scanning process. In this way, the method integrates image processing and robot trajectory planning into one framework. By directly decoding the fused high-level semantic features (fused attention features) into joint-level control parameters (joint torque sequence), it eliminates the cumbersome multi-stage processing or manual feature engineering in the middle, improves the system response efficiency, and ensures the consistency and coherence from visual understanding to physical action. This enables the robot's motion trajectory to respond more directly and accurately to the key anatomical and pathological information identified in the image.
[0039] In an optional embodiment, the differential attention operation includes steps A1 to A2: Step A1: Generate a query matrix, a key matrix, and a value matrix based on the input image features.
[0040] Specifically, a learnable linear transformation layer projects the input image features into three distinct matrices: a query matrix, a key matrix, and a value matrix. The query matrix initiates the "query," the key matrix provides the matching "key," and the value matrix contains the substantive information to be aggregated. These three matrices form the foundation for all subsequent relevance calculations and information aggregation.
[0041] Step A2: Determine the first query submatrix and the first key submatrix from the query matrix and the key matrix respectively according to the preset rules, and obtain the preset second query submatrix and the second key submatrix, wherein the second query submatrix and the second key matrix are identical matrices.
[0042] The first query submatrix and the first key submatrix are selected from the query matrix and the key matrix respectively based on preset rules (e.g., selecting feature dimensions with high importance according to vector norm). These represent the "target" information extracted from the input features that needs to be evaluated.
[0043] The second query submatrix and the second key submatrix are specifically defined as identity matrices, which are square matrices with 1s on the diagonal and 0s on the rest. Using them as the query and key essentially defines a neutral, unbiased reference benchmark, assuming that all feature locations have the same fundamental correlation with a "standard unit vector".
[0044] Optionally, determining the first query submatrix and the first key submatrix from the query matrix and the key matrix respectively according to preset rules may include steps A21 to A23: Step A21: Calculate the vector norm of each column of the query matrix as the query norm sequence, and calculate the vector norm of each column of the key matrix as the key norm sequence.
[0045] Vector norms (such as the 2-norm, or Euclidean length) are scalar values that measure the size or strength of a vector. By calculating the vector norm of each column of the query matrix, a query norm sequence is obtained, which initially reflects the activation strength or information content of each feature dimension on the query side. Similarly, calculating the vector norm of each column of the key matrix yields a key norm sequence. The query norm sequence initially reflects the activation strength or information content of each feature dimension on the query side, and the key norm sequence initially reflects the activation strength or information content of each feature dimension on the key side. Dimensions with larger norm values generally indicate that the feature represented by that dimension is more significant or important at the current position.
[0046] For example, taking a vector norm of 2-norm as an example, the calculation of the vector norm of each column of the query matrix and the calculation of the vector norm of each column of the key matrix can be expressed as follows:
[0047] in, Indicates the calculation of L2 norm; This indicates a query for the i-th column of matrix Q; The vector norm (2-norm) of the i-th column of the query matrix Q is represented. Represents the i-th column of the key matrix K; Let represent the vector norm (2-norm) of the i-th column of the key matrix K.
[0048] Step A22: Determine the importance weight of each dimension based on the element-wise product of the corresponding elements of the query norm sequence and the key norm sequence.
[0049] To more comprehensively evaluate the overall importance of a feature dimension in attention interactions, it needs not only to perform strongly as a query or key alone, but also to carry important information in both roles. Therefore, the query norm sequence and key norm sequence obtained in step A21 are multiplied element-wise. Here, "element-wise" refers to the two norm values representing the same feature dimension. The result of the multiplication forms a new sequence called the importance weight sequence. This importance weight comprehensively considers a dimension's potential in both initiating and responding to queries (Q). The higher the importance weight value, the more likely the dimension is to play a key role in establishing effective attention associations.
[0050] For example, the importance weight of the i-th dimension It can be represented as:
[0051] Step A23: Based on the importance weights, select important dimensions from the query matrix and the key matrix respectively to form the first query submatrix and the first key submatrix.
[0052] Specifically, based on the importance weights calculated in step A22, the dimensions are sorted in descending order of weight value. Then, according to a preset retention ratio (e.g., retaining the top 1 / 2 of the dimensions with the highest weights), the indices of the important dimensions to be retained are determined. Finally, based on these indices, the corresponding column vectors are extracted from the original query matrix and concatenated to form the first query submatrix. Similarly, column vectors with the same index are extracted from the original key matrix K and concatenated to form the first key submatrix. This forms the first query submatrix. and the first key matrix That is, the subset of feature dimensions that are most likely to contain discriminative information, which is automatically selected by the model.
[0053] For example, if the first half is selected as the most important half of the parameters to be retained, the construction process of the first query submatrix and the first key submatrix is as follows: (1) Set the retention ratio, that is:
[0054] in, This means that the top 50% of the most important dimensions will be selected from all feature dimensions to construct the submatrix.
[0055] (2) Total number of dimensions obtained:
[0056] Where d is the total number of dimensions obtained, that is, the number of columns of the query matrix Q. Since the query matrix Q and the key matrix K usually have the same dimensions, this value also represents the total number of dimensions of the feature space.
[0057] (3) Calculate the number of dimensions to be retained:
[0058] in, The number of dimensions to be retained is the actual number of dimensions to be retained, and its value is the retention ratio. The product of the product and the total number of dimensions d is then rounded using the round function to ensure that an integer number of dimensions is obtained.
[0059] (4) Determine the indexes for important dimensions:
[0060] in, Here, w represents the importance index for the dimension of importance, calculated in step A22. This formula selects the highest-valued importance weight from all importance weights. Each element is recorded, along with its index position, forming an index for the important dimensions. The index of the important dimension corresponds to the feature dimension that is determined to be the most important.
[0061] (5) Construct the first query submatrix With the first key matrix :
[0062] in, This means retrieving all rows from the query matrix Q, but only keeping the columns whose indexes belong to the query matrix Q. Those columns thus constitute the first query submatrix ; This means retrieving all rows from the key matrix K, but only retaining the columns whose indices belong to the key matrix K. Those columns thus constitute the first key matrix. Thus, the first query submatrix With the first key matrix It consists of the feature dimensions with the highest importance in the original matrix, and will be used as the target feature pairs in differential attention calculation.
[0063] Thus, this implementation provides a learnable, data-driven dimensionality filtering mechanism that, by combining the norm information on both sides of the query and the key to evaluate importance, can accurately identify the feature subspace that contributes most to establishing a strong discriminative attention association. Compared to random selection or fixed rules, this method enables the differential attention mechanism to adaptively focus on the most informative parts of the input features, thereby more effectively constructing the target matrix pair (the first query submatrix) for differential comparison. With the first key matrix This enhances the targeting and efficiency of differential attention operations, helping the model to more accurately extract key differential features related to lesions from echocardiogram images, suppress irrelevant noise, and lay a solid foundation for the subsequent generation of high-quality fused attention features.
[0064] Step A3: Calculate the first correlation matrix of the first query submatrix and the first key submatrix, and the second correlation matrix of the second query submatrix and the second key submatrix.
[0065] The calculation of the first relevance matrix involves operations (such as matrix multiplication and scaling) to calculate the relevance weights between the first query submatrix and the first key submatrix. This first relevance matrix reflects the interaction strength between positions based on the actual input features.
[0066] The second relevance matrix is calculated by performing an operation to calculate the relevance weights between the second query submatrix and the second key submatrix. Since the second query submatrix and the second key matrix are identical matrices, this step produces a uniform or fixed-pattern base attention distribution, representing the background or default relevance level without interference from specific input information.
[0067] Step A4: Calculate the difference matrix between the first correlation matrix and the second correlation matrix.
[0068] Specifically, the second correlation matrix is subtracted from the first correlation matrix to obtain the difference matrix. This operation removes the attention background of the uniform or fixed pattern defined by the identity matrix; therefore, the values in the difference matrix can characterize the significant deviation or specific correlation of the actual input features relative to the neutral benchmark; positive values in the difference matrix indicate that the feature correlation is stronger than the benchmark, which may correspond to key discriminative regions (such as lesions); negative values or values close to zero in the difference matrix may correspond to irrelevant background.
[0069] Step A5: Calculate the second attention feature based on the difference matrix and the value matrix.
[0070] This step is used to complete information aggregation. The difference matrix obtained in step A4 (as the recalibrated attention weights) is applied to the value matrix generated in step A1. The value matrix is then weighted and summed using the difference matrix, and the final output is the second attention feature. This second attention feature does not contain uniform background attention, but strongly focuses on feature interactions that differ significantly from the neutral benchmark. It is rich in local discriminative difference information and plays a crucial role in highlighting small lesions and clearly distinguishing tissue boundaries in cardiac ultrasound images.
[0071] For example, calculating the second attention feature It can be represented as:
[0072] in, These represent the generation of a query matrix, a key matrix, and a value matrix based on the input image features, respectively. Represents the first query submatrix. Denotes the first key matrix. and Used to capture the most critical and discriminative feature interactions in the input; This represents the second query submatrix. Denotes the second key matrix. and It represents a neutral and uniform reference benchmark, and the attention distribution it generates simulates an association pattern in a state without information or background. This represents the first correlation matrix; This represents the second correlation matrix; This represents the Softmax function, which is used to normalize the relevance scores into a probability distribution, ensuring that the sum of the attention weights is 1. It represents a linear transformation of the value matrix V, containing the original feature information that needs to be aggregated.
[0073] The technical solution adopted in this application implements an efficient and specific differential attention implementation path. By introducing an identity matrix as a stable benchmark and performing differential comparison with the target feature matrix (first query sub-matrix and first key sub-matrix) extracted from the input, it can stably and effectively filter out general, non-discriminatory attention background noise, enabling the model to learn and amplify local feature differences that are important for specific tasks (such as lesion recognition). This enhances the model's ability to perceive subtle targets in cardiac ultrasound images, provides a guarantee for generating high-quality feature representations, and improves the accuracy and robustness of subsequent tasks such as trajectory planning.
[0074] like Figure 2 The diagram illustrates the structure of the differential attention module in this embodiment. Specifically, the differential attention module processes image features through the following steps: First, based on the input image features, a query matrix, a key matrix, and a value matrix are generated. According to preset rules, a first query submatrix and a first key submatrix are determined from the query matrix and the key matrix, respectively, and a preset second query submatrix and a second key submatrix are obtained, wherein the second query submatrix and the second key matrix are identical matrices.
[0075] Subsequently, the first correlation matrix of the first query submatrix and the first key submatrix is calculated, and the second correlation matrix of the second query submatrix and the second key matrix is calculated; further, the difference matrix between the first correlation matrix and the second correlation matrix is calculated; based on the difference matrix and the value matrix, differential attention is calculated to obtain the second attention feature representing the local discriminative difference information.
[0076] Simultaneously, based on the query matrix, key matrix, and value matrix, a first attention feature representing global context information is obtained through standard attention computation. Specifically, quasi-self-attention computation can be performed based on the second relevance matrix and value matrix to obtain the first attention feature.
[0077] Finally, the first attention feature and the second attention feature are concatenated and fused in the attention head dimension to obtain the fused attention feature.
[0078] Thus, through the above structure, the differential attention module can synergistically utilize global semantic information and local subtle differences to effectively suppress interference from large background areas in cardiac ultrasound images, guiding the model to more accurately focus on discriminative micro lesions or specific anatomical structures, thereby significantly improving the targeting and effectiveness of feature representation.
[0079] Understandably, during model training, by using appropriate initialization methods (such as steps A21 to A23 above) and learnable parameters, the model can be made as close as possible to standard attention in the early stages of training. Through gradient backpropagation, the parameters are gradually learned and optimized, enabling the model to continuously learn the characteristics of differential attention, and ultimately achieve an organic combination of standard attention and differential attention.
[0080] In one alternative embodiment, before or after splicing and fusing the first attention feature and the second attention feature in the attention head dimension, the first attention feature and / or the second attention feature are weighted and adjusted by a learnable weight parameter.
[0081] Specifically, before the splicing and fusion step, independent learnable weight parameters can be applied to the first attention feature and the second attention feature to scale or transform them respectively, so as to obtain the weighted first attention feature and the weighted second attention feature; then, the weighted first attention feature and the weighted second attention feature are spliced together in the attention head dimension to obtain the fused attention feature.
[0082] After the splicing and fusion step, the first attention feature and the second attention feature can be spliced directly at the attention head dimension to form a spliced feature. Then, the spliced feature is recalibrated through a transformation controlled by learnable weight parameters (such as a lightweight fully connected layer or an element-wise attention gating mechanism) to adjust the contribution of different feature channels (especially channels from different attention sources) to obtain the fused attention feature.
[0083] The technical solution of this application introduces learnable weight parameters to adjust the weights of the first attention features and / or the second attention features, enabling the model to automatically and dynamically adjust the relative importance or contribution ratio of global contextual information (first attention features) and local discriminative difference information (second attention features) in the final fused features. For example, for tasks that require emphasizing the overall structure (such as standard section recognition), the weight of global features can be increased; while for tasks that require locating small lesions, the weight of local difference features can be enhanced. Furthermore, in the early stages of model training, by appropriately initializing these weight parameters (e.g., making the initial fusion close to standard attention), a stable learning starting point can be provided for the model; as training progresses, the model gradually learns the optimal fusion strategy.
[0084] In one optional embodiment, there are multiple differential attention modules connected in series; wherein the i-th differential attention module takes the fused attention feature output by the (i-1)-th differential attention module as input and outputs a new fused attention feature; i is an integer greater than 1, and the input of the first differential attention module is the image feature.
[0085] In this embodiment, multiple differential attention modules with identical or similar structures are sequentially connected to form a processing pipeline or stacked network layer to process the image features of cardiac ultrasound images step by step and level by level. Specifically, the image features extracted in step S110 are input to the first differential attention module, which performs its internal standard attention and differential attention operations and performs splicing and fusion to output the first fused attention feature. Next, the second differential attention module receives the fused attention feature output by the first module as its input, and performs attention calculation and fusion again based on the first fused attention feature to output the second fused attention feature, which contains deeper and more abstract semantic information. This process can be repeated. The i-th differential attention module receives the output fused attention feature of its predecessor (the (i-1)-th differential attention module) as input and generates its own output fused attention feature. In this way, the original image features are transformed and enhanced layer by layer and progressively.
[0086] Thus, by processing image features using multiple differential attention modules, with each layer increasing in depth, subsequent modules operate based on the fused attention features of the previous layer, which have already integrated local and global information. This allows for the capture of more macroscopic semantic relationships and more complex patterns (such as the morphological relationships of the entire heart chambers). Furthermore, through this multi-level attention mechanism, key regions are repeatedly and progressively focused on and confirmed, achieving more precise lesion localization and more robust segmentation and classification, thus improving the accuracy of image analysis. Finally, the fused attention features output by the last (or last few) differential attention modules are used to generate the robot's joint torque sequence. This feature has undergone multiple rounds of collaborative optimization of global and local information, removing a large amount of noise and integrating semantic information from multiple levels, from low to high. Generating control commands based on this fused attention feature ensures that the visual understanding upon which trajectory planning decisions are based is profound and stable, fundamentally improving the accuracy of the robot's motion trajectory.
[0087] In one optional embodiment, the differential attention module is a neighborhood differential attention module or a deformable differential attention module; wherein, the neighborhood differential attention module is configured to compute attention within a local neighborhood window of image features; and the deformable differential attention module is configured to dynamically adjust the receptive field position of attention computation through a learnable offset network.
[0088] Specifically, the neighborhood differential attention module divides the input image features into multiple non-overlapping or overlapping local windows (neighborhoods), and attention computation (including standard attention and differential attention) is restricted to within each window. For example, a feature location is only correlated with other locations within the same window. After processing, windows can exchange information through sliding or shifting operations to establish longer-range dependencies. In this way, neighborhood differential attention reduces computational complexity to a linear level by decomposing global computation into multiple parallel local computations, thereby improving the model's processing efficiency and making it more feasible to apply to real-time or near-real-time analysis of high-resolution images, which is crucial for the response speed of robotic systems. At the same time, local focusing itself is also more conducive to capturing subtle local structural differences.
[0089] The core of the deformable differential attention module lies in the introduction of a learnable offset network. For each query location in the input features, this offset network predicts a set of position offsets based on its surrounding context features. These offsets are used to dynamically and adaptively shift attention sampling points to more relevant feature locations when computing key-value pairs; subsequently, attention computation (including standard attention and differential attention) is performed at these predicted, irregular locations. In this way, the deformable differential attention module endows the model with the ability to dynamically shape its receptive field, enabling it to flexibly adjust and precisely cover irregular lesion areas or key anatomical boundaries, much like a visual focus. This enhances the model's adaptability to target geometric transformations and improves its ability to extract clean lesion features from complex backgrounds, thereby providing the robot with more accurate spatial positioning information for trajectory planning.
[0090] It is understandable that when there are multiple differential attention modules connected in series, these modules can be of the same type (e.g., all neighborhood differential attention modules, or all deformable differential attention modules). Alternatively, the multiple differential attention modules can be different modules; that is, some can be neighborhood differential attention modules, and others can be deformable differential attention modules. In some embodiments, adjacent differential attention modules are set as different modules (e.g., the module following a neighborhood differential attention module is a deformable differential attention module).
[0091] If the differential attention module is a neighborhood differential attention module, then its internal standard attention can be called "neighborhood attention" and differential attention can be called "neighborhood differential attention". Similarly, if the differential attention module is a deformable differential attention module, then its internal standard attention can be called "deformable attention" and differential attention can be called "deformable differential attention".
[0092] In an optional embodiment, the target output in step S130 further includes at least one of the following: Item B-1: Semantic segmentation results of pixel-level division of cardiac anatomical structures or lesion regions; Item B-2: The classification result of determining the overall category of the cardiac ultrasound image, which includes disease diagnosis, standard cardiac section recognition, and image quality correlation.
[0093] In this embodiment, to adapt to various ultrasound imaging tasks, multiple task output heads are designed to accommodate image classification and semantic segmentation outputs. Specifically, the semantic segmentation result of item B-1 is typically achieved through a semantic segmentation output head, which can be a decoder based on a U-Net structure. It receives fused attention features from the differential attention module as input, and through upsampling, cross-layer connections (fusion with shallow features from the encoder), and pixel-by-pixel classification, finally outputs a semantic segmentation result that performs pixel-level division of cardiac anatomical structures or lesion regions.
[0094] For example, the decoder can be represented as:
[0095] in, This represents the upsampling function, which takes features from a deeper layer (the (l+1)th layer) as input. Spatial resolution is obtained by upsampling using an upsampling function. Upsampling functions typically refer to operations such as transposed convolution or bilinear interpolation upsampling, the purpose of which is to increase the height and width of the feature map to match the resolution required by the l-th layer. This represents the features of the corresponding layer l in the autoencoder (or backbone network). These features have high spatial resolution and contain more details and location information (such as edges and textures). Feature fusion typically involves concatenation along the channel dimension or element-wise addition to generate fused features. . The decoding features of layer l are used for the next shallower layer (l 1) The upsampling and fusion process, This represents a convolutional block consisting of several convolutional layers, activation functions, and normalization layers.
[0096] The semantic segmentation result can be represented by a semantic segmentation map with the same resolution as the echocardiogram image. In this semantic segmentation map, each pixel is assigned a category label to accurately delineate specific cardiac anatomical structures (e.g., left ventricular endocardium, left ventricular myocardium, right ventricular cavity, left atrium, etc.) or abnormal lesion areas (e.g., myocardial infarction area, intracardiac thrombus, valvular calcification, etc.). This semantic segmentation result can be used to quantitatively calculate key cardiac function parameters (such as left ventricular ejection fraction, ventricular volume, etc.), providing doctors with objective diagnostic basis. At the same time, the information such as visceral anatomical structures or lesion areas segmented in the semantic segmentation result can also serve as higher-level semantic feedback to assist or verify the generation logic of joint torque sequences and enhance the interpretability of decisions.
[0097] The classification result of item B-2 can be achieved through a classification output head, which is usually a global pooling layer (such as average pooling) followed by one or more fully connected layers. This output head gathers global information from the fused attention features and then maps it to one or more class probability distributions to obtain the classification result that determines the overall category of the cardiac ultrasound image.
[0098] Disease diagnosis is used to determine whether the image presents a specific cardiac disease, such as outputting a category like "normal" or "dilated cardiomyopathy," providing a high-level pathological target for robotic scanning. Standard cardiac section recognition automatically identifies which standard ultrasound section (e.g., "apical four-chamber view") the current image belongs to, a crucial foundation for robotic navigation and standardized scanning. Image quality assessment determines whether the image is clear and usable, or what quality issues exist (e.g., "clear image," "poor probe contact," "severe artifacts"). This provides real-time feedback to the robot to adjust probe posture, pressure, or scanning parameters. Thus, the classification results provide the entire system with a high-level semantic understanding and task context; for example, identifying a standard section ensures the robot is in the correct scanning reference position; diagnosing a specific disease triggers a refined scanning plan for that disease; and assessing poor image quality allows for immediate adaptive adjustments to the robot, significantly improving the intelligence and robustness of the automated scanning process.
[0099] The technical solution adopted in this application can generate various target outputs. These outputs not only directly serve medical diagnosis (providing segmentation and classification results) but also create a synergistic enhancement effect with robot control functions. For example, high-quality semantic segmentation results can provide more refined spatial constraints for trajectory planning, while accurate classification results provide high-level semantic guidance and quality control for control strategies. This constitutes a fully functional and interconnected intelligent auxiliary system for cardiac ultrasound robots, encompassing everything from microscopic pixel-level analysis to macroscopic task decision-making and final physical execution, thus enhancing its practicality, reliability, and added value in clinical practice.
[0100] This application also provides a cardiac ultrasound robotic system, see reference. Figure 3 As shown, Figure 3 This is a schematic diagram of a cardiac ultrasound robot system provided in an embodiment of this application. The system includes: Ultrasound image acquisition device, used to acquire cardiac ultrasound images; The processing module is used to execute the image processing method for motion trajectory planning of cardiac ultrasound robot described in the above embodiments to generate the motion trajectory of cardiac ultrasound robot. A robotic arm device is used to perform cardiac ultrasound scanning actions according to the said motion trajectory.
[0101] In this embodiment, the ultrasound imaging acquisition device is responsible for acquiring real-time echocardiographic images of the patient's heart. The ultrasound imaging acquisition device typically includes an ultrasound probe, an ultrasound transmitting / receiving circuit, a beamformer, and a signal processing unit. The probe physically contacts or approaches the patient's body surface, emitting ultrasound waves and receiving echoes, ultimately forming a dynamic or static sequence of echocardiographic images.
[0102] The processing module is the runtime entity of the software algorithm on hardware (such as a high-performance computer, graphics processing unit, or dedicated AI chip). The processing module is configured to load and run the computer program corresponding to the "Image Processing Method for Motion Trajectory Planning of a Cardiac Ultrasound Robot" described in any of the foregoing embodiments. By receiving cardiac ultrasound images from the ultrasound image acquisition device, it sequentially performs: feature extraction, differential attention fusion processing (which may include multi-layer, neighborhood, or deformable attention), and task inference (generating joint torque sequences, optionally including segmentation or classification results). Its final core output is the motion trajectory of the cardiac ultrasound robot, which typically exists in the form of a series of timestamps corresponding to the robot arm's target pose, velocity, or directly as a sequence of underlying joint torque commands.
[0103] The robotic arm receives the motion trajectory from the processing module and, based on the robot's kinematics and dynamics model, decomposes the trajectory into the real-time position, speed, or torque settings of each joint motor. It then drives the robotic arm to move, thereby precisely guiding the ultrasound probe along the planned trajectory to move, tilt, and maintain appropriate contact pressure on the patient's body surface, thus completing an automated cardiac ultrasound scan.
[0104] Thus, the cardiac ultrasound robot system of this application embodiment achieves end-to-end automation from image acquisition to intelligent analysis and precise execution, reducing manual operation and improving the consistency and efficiency of the scanning process. Thanks to the differential attention model in the processing module's precise focusing on minute lesions and its ability to perceive image quality, the system-generated trajectory guides the probe to more accurately cover the target anatomical structure and can adaptively fine-tune based on real-time image feedback, enhancing the medical value of the scan. The robotic arm's motion precision and stability far exceed those of manual handheld operation, maintaining a stable and repeatable scanning posture for extended periods, reducing image quality fluctuations caused by operator fatigue or tremors, while also ensuring patient safety through built-in safety constraints (such as force feedback and obstacle avoidance).
[0105] This application also provides an image processing device for motion trajectory planning of a cardiac ultrasound robot, referring to... Figure 4 As shown, Figure 4 This is a schematic diagram of an image processing device for motion trajectory planning of a cardiac ultrasound robot, provided in an embodiment of this application. The device includes: The feature extraction module 410 is used to acquire cardiac ultrasound images and preprocess the cardiac ultrasound images to extract image features. The image processing module 420 inputs image features into the differential attention module, performs standard attention operations and differential attention operations in parallel, generates a first attention feature representing global context information and a second attention feature representing local discriminative difference information, respectively, and then concatenates and fuses the first attention feature and the second attention feature in the attention head dimension to obtain a fused attention feature. The sequence output module 430 is used to generate a target output through the task output head based on the fused attention features, the target output including at least a joint torque sequence for controlling the cardiac ultrasound robot; The trajectory generation module 440 is used to generate the motion trajectory of the cardiac ultrasound robot based on the joint torque sequence.
[0106] It is understood that the image processing device for cardiac ultrasound robot motion trajectory planning in the embodiments of this application can realize the image processing method for cardiac ultrasound robot motion trajectory planning in the above embodiments. The image processing device for cardiac ultrasound robot motion trajectory planning has the same advantages as the above-mentioned image processing method for cardiac ultrasound robot motion trajectory planning compared with the prior art, and will not be repeated here.
[0107] This application also provides an electronic device, see embodiments thereof. Figure 5 , Figure 5 This is a schematic diagram of an electronic device provided in an embodiment of this application. For example... Figure 5 As shown, the electronic device 500 includes a memory 510 and a processor 520. The memory 510 and the processor 520 are connected via a bus for communication. The memory 510 stores a computer program that can run on the processor 520 to implement the steps of the image processing method for cardiac ultrasound robot motion trajectory planning described in the embodiments of this application.
[0108] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0109] This application describes embodiments of methods and apparatus according to flowchart illustrations and / or block diagrams. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0110] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0112] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0113] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0114] The foregoing has provided a detailed description of an image processing method, system, apparatus, and device for motion trajectory planning of a cardiac ultrasound robot. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An image processing method for motion trajectory planning of a cardiac ultrasound robot, characterized in that, The method includes: Acquire cardiac ultrasound images and preprocess the cardiac ultrasound images to extract image features; The image features are input into the differential attention module, and standard attention operations and differential attention operations are performed in parallel to generate a first attention feature representing global context information and a second attention feature representing local discriminative difference information, respectively. The first attention feature and the second attention feature are then concatenated and fused in the attention head dimension to obtain the fused attention feature. Based on the fused attention features, a target output is generated through the task output head, and the target output includes at least a joint torque sequence for controlling the cardiac ultrasound robot; The motion trajectory of the cardiac ultrasound robot is generated based on the joint torque sequence.
2. The method according to claim 1, characterized in that, The differential attention operation includes: Based on the input image features, generate a query matrix, a key matrix, and a value matrix; According to preset rules, a first query submatrix and a first key submatrix are determined from the query matrix and the key matrix respectively, and a preset second query submatrix and a second key matrix are obtained, wherein the second query submatrix and the second key matrix are identical matrices; Calculate the first relevance matrix of the first query submatrix and the first key submatrix, and the second relevance matrix of the second query submatrix and the second key matrix; Calculate the difference matrix between the first correlation matrix and the second correlation matrix; The second attention feature is calculated based on the difference matrix and the value matrix.
3. The method according to claim 2, characterized in that, Determining the first query submatrix and the first key submatrix from the query matrix and the key matrix respectively according to preset rules includes: Calculate the vector norm of each column of the query matrix as a query norm sequence, and calculate the vector norm of each column of the key matrix as a key norm sequence; The importance weights of each dimension are determined by multiplying the corresponding elements of the query norm sequence and the key norm sequence. Based on the importance weights, important dimensions are selected from the query matrix and the key matrix respectively to form the first query submatrix and the first key submatrix.
4. The method according to claim 1, characterized in that, Before or after concatenating and fusing the first attention feature and the second attention feature in the attention head dimension, the first attention feature and / or the second attention feature are weighted and adjusted using learnable weight parameters.
5. The method according to any one of claims 1-4, characterized in that, The differential attention modules are multiple and connected in series; The i-th differential attention module takes the fused attention feature output by the (i-1)-th differential attention module as input and outputs a new fused attention feature; i is an integer greater than 1, and the input of the first differential attention module is the image feature.
6. The method according to claim 5, characterized in that, The differential attention module is a neighborhood differential attention module or a deformable differential attention module; The neighborhood differential attention module is configured to compute attention within a local neighborhood window of the image features; The deformable differential attention module is configured to dynamically adjust the receptive field position of attention computation through a learnable offset network.
7. The method according to claim 1, characterized in that, The target output also includes at least one of the following: Semantic segmentation results for pixel-level division of cardiac anatomical structures or lesion regions; The classification result of the overall category of the cardiac ultrasound image, which includes disease diagnosis, standard cardiac section recognition, and image quality correlation.
8. A cardiac ultrasound robotic system, characterized in that, include: Ultrasound image acquisition device, used to acquire cardiac ultrasound images; The processing module is configured to execute the image processing method for motion trajectory planning of a cardiac ultrasound robot as described in any one of claims 1-7, so as to generate the motion trajectory of the cardiac ultrasound robot. A robotic arm device is used to perform cardiac ultrasound scanning actions according to the said motion trajectory.
9. An image processing device for motion trajectory planning of a cardiac ultrasound robot, characterized in that, The device includes: The feature extraction module is used to acquire cardiac ultrasound images and preprocess the cardiac ultrasound images to extract image features. The image processing module inputs image features into the differential attention module, performs standard attention operations and differential attention operations in parallel, generates a first attention feature representing global context information and a second attention feature representing local discriminative difference information, respectively, and then concatenates and fuses the first attention feature and the second attention feature in the attention head dimension to obtain a fused attention feature. A sequence output module is used to generate a target output through a task output head based on the fused attention features, wherein the target output includes at least a joint torque sequence for controlling the cardiac ultrasound robot; The trajectory generation module is used to generate the motion trajectory of the cardiac ultrasound robot based on the joint torque sequence.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the image processing method for motion trajectory planning of a cardiac ultrasound robot as described in any one of claims 1-7.